RFC Implementation Notes
Notes keyed to RFC sections.
Notes keyed to RFC sections. Documents interpretations, deviations, and implementation choices made during development.
Supported standards at a glance
Consolidated map of the RFCs, SAFIs, and features rustbgpd implements. The per-RFC sections below carry the conformance detail, interpretations, and deviations; docs/interop.md has the interop matrix, docs/receipts.md the receipts index, and docs/reference/limitations.md the boundaries and non-goals.
| Area | Standards | Scope |
|---|---|---|
| Core BGP | RFC 4271, RFC 6793 (4-byte ASN), RFC 7606 (revised error handling) | FSM + UPDATE validation with treat-as-withdraw / attribute-discard, dual-stack IPv4/IPv6 unicast (SAFI 1) |
| MP-BGP + extensions | RFC 4760, RFC 7911 (Add-Path), RFC 8654 (Extended Messages), RFC 8950 (Extended Next Hop) | Multiprotocol negotiation and modern capability set |
| Route refresh / filtering | RFC 2918, RFC 7313 (Enhanced RR), RFC 5291/5292 (ORF) | Receive-side Address-Prefix ORF |
| Communities | RFC 1997 (well-known), RFC 4360 (Extended), RFC 8092 (Large) | Match plus policy set/remove; NO_ADVERTISE, NO_EXPORT, and NO_EXPORT_SUBCONFED egress enforcement for unicast, VPNv4/VPNv6, labeled-unicast, RTC, BGP-LS, EVPN, and FlowSpec |
| eBGP default policy | RFC 8212 | Opt-in ebgp_requires_policy: an eBGP direction with no explicit operator policy runs a reserved internal deny |
| Route reflection | RFC 4456, RFC 9107 (ORR, ADR-0095) | Per-client best paths via BGP-LS-sourced SPF |
| Route server (IXP) | RFC 7947 (ADR-0039/0101), RFC 8195 | Transparent redistribution, §2.3.2 per-client best-path, member-set control communities (per-target announce/prepend steering, scrubbed on egress) |
| Graceful restart | RFC 4724 (GR helper), RFC 9494 (LLGR) | Stale retention across all RR families; role-derived forwarding-state bits |
| VPN / MPLS families (RR / controller-feed only, ADR-0077) | RFC 4364/4659 VPNv4/v6 (SAFI 128), RFC 4684 RT-Constrain (SAFI 132), RFC 8277 labeled-unicast (SAFI 4), RFC 9552 BGP-LS (SAFI 71/72) | RD/label/next-hop/RT preserved verbatim; no VRF import, no MPLS FIB, no local BGP-LS production |
| EVPN (Linux/VXLAN alpha) | RFC 7432, RFC 9135/9136 (symmetric IRB), RFC 9012/8365 (VXLAN encap), RFC 9251 (SMET relay) | Route types 1–6; Type 6 SMET relay only; RR + VTEP + multi-homing building blocks; RFC 9721 §5.1/§6.2 local-move cascade (partial) |
| Origin / path security | RFC 6811 + RFC 8210 (RPKI/RTR), ASPA, RFC 9234 (Roles + OTC, ADR-0071) | Origin validation, AS-path verification, leak prevention |
| Transport security | RFC 5925 (TCP-AO), TCP MD5, RFC 5082 (GTSM) | TCP-AO: static-neighbor and direct dynamic-prefix keyrings on Linux; add-only successor installation, observation-gated successor selection/deprecation, then deprecated unselected-MKT deletion on separate SIGHUP generations; RPKI cache (RTR) sockets take the same MD5 or TCP-AO material |
| FlowSpec / blackhole | RFC 8955/8956 and RFC 9117 (FlowSpec, SAFI 133), RFC 7999 (BLACKHOLE) | Opt-in FlowSpec feasibility; opt-in BLACKHOLE Linux FIB discard |
| Liveness | RFC 5880/5881/5882/5883 (BFD), RFC 9384 (BFD Down Cease subcode), RFC 9687 (Send Hold Timer) | Single-hop and multihop async BFD for static neighbors; typed Cease/10 teardown on a genuine BFD Down |
| Maintenance | RFC 8326 (Graceful Shutdown), RFC 9003 (Extended Admin Shutdown Communication) | Receiver gating + initiator toggle |
| Monitoring | RFC 7854/8671/9069 (BMP trio), RFC 9972 (selected BMP statistics), RFC 6396 (MRT TABLE_DUMP_V2), RFC 7951 (gNMI/OpenConfig JSON) | Pre-policy / post-policy / Loc-RIB BMP views |
RFC 1997 — well-known community export coverage
- IPv4/IPv6 unicast, VPNv4/VPNv6, labeled-unicast, RTC, BGP-LS SAFI 71/72,
EVPN, and FlowSpec routes carrying
NO_ADVERTISEin their stored route attributes are ineligible toward every target before export policy, and the post-policy route is checked again before Adj-RIB-Out commit. A permit policy cannot remove the community to bypass the restriction; rustbgpd also conservatively suppresses a modified route when export policy adds it. NO_EXPORT(0xFFFFFF01) andNO_EXPORT_SUBCONFED(0xFFFFFF03) are enforced at eBGP egress for IPv4/IPv6 unicast, VPNv4/VPNv6, labeled-unicast, RTC, BGP-LS, EVPN, and FlowSpec: a route received with either community is suppressed at staging toward every eBGP peer whoseinterpret_rfc1997knob is on (the default for plain eBGP and iBGP peers; route-server clients default to transparent pass-through, matching common IXP route-server practice — setinterpret_rfc1997 = trueon an RS client to opt in). iBGP targets — including RR reflection and ORR — are never suppressed: RFC 1997 permits intra-AS advertisement. rustbgpd has no confederation support, so theNO_EXPORT_SUBCONFEDegress set collapses to theNO_EXPORTone and both are enforced identically.- Unlike
NO_ADVERTISE, theNO_EXPORTcheck is source-route-only. A policy that ADDSNO_EXPORTproduces a route that is still delivered — attaching the community toward a peer that should honor it is the standard route-server/export-policy idiom, and rustbgpd's own RFC 9494 §4.6 LLGR form attachesNO_EXPORT(withLOCAL_PREF0) post-staging in transport. A policy that REMOVESNO_EXPORTcannot bypass the restriction: the source-route check runs before export policy. The explain ladder reports the suppression on its own stableno_exportgate rung (besideno_advertise). - The
NO_EXPORTgate deliberately outranks RFC 9494 LLGR export eligibility: a route RECEIVED already carryingNO_EXPORT(the upstream §4.6 stale form isLLGR_STALE+NO_EXPORT) is suppressed toward eBGP peers in honor mode even where the peer advertised the LLGR capability for the family. RFC 1997 is a MUST; LLGR eligibility only makes a stale route deliverable, it does not override the community's egress restriction. - The same predicate covers single-best plus Add-Path for unicast, VPN, and labeled routes; grouped/private and RFC 7947 per-client-best shapes apply to unicast, grouped/private applies to VPN, and RFC 9107 ORR applies to unicast and labeled-unicast. FlowSpec, RTC, both BGP-LS SAFIs, and EVPN use their single-best export paths. Existing advertisements are withdrawn and logical Adj-RIB-Out state is cleared.
- Add-Path and per-client-best remove scoped candidates before ranking, so
surviving siblings compact normally. They also skip policy-modified
candidates whose result carries
NO_ADVERTISE. ORR single-best first selects its per-vantage best and suppresses that winner without falling back to a different route, whether the community arrived on the source or from policy. Ranked ORR is the asymmetric case: an Add-Path peer bound to a vantage ranks with the ORR comparator over candidates already filtered forNO_ADVERTISE, so the runner-up fills the vacated rank — the no-fallback shape applies only to ORR single-best. - For RTC (SAFI 132),
NO_ADVERTISEsuppression has a wider blast radius than for other families. An RT-membership NLRI suppressed by community policy is withdrawn like any other route, and a receiver that filters VPN or EVPN advertisements by RT-Constrain membership then prunes every route carrying that Route Target; rustbgpd's own RFC 4684 outbound gate reacts the same way toward a peer whose membership no longer covers an RT. The mechanics are correct and fail-closed — the amplification is inherent to SAFI 132, where one NLRI stands for the whole class of VPN routes carrying that RT.
RFC 7947 §2.3.2 / RFC 8195 — route-server control communities
- A member steers per-target redistribution with communities keyed on the
TARGET peer's ASN:
0:PEER/RS:0:PEER(do not announce toPEER),0:RS/RS:0:0(announce to none) overridable per target byRS:PEER/RS:1:PEER, andRS:101|102|103:PEER(prepend the announcing member's leftmost ASN 1–3× towardPEER;RS:10x:0= every target). Standard and RFC 8195 large forms compose; extended-community control forms are deliberately not implemented. The related draft-ietf-grow-ixp-ext-comms-03 expired on 2026-06-12 without RFC publication; it is background for this implementation choice, not an RFC requirement. Full matrix and evaluation ladder: the route-server cookbook. - Enforcement is gated per session by
rs_control_communities(defaultroute_server_client), evaluated pre-policy on the source route like the RFC 1997 gates (its ownrs_controlexplain rung), and covers the unicast export shapes: single-best, Add-Path, and per-client-best (suppressed candidates are removed before ranking). Non-unicast families are out of scope — RFC 7947 is an IPv4/IPv6 unicast IXP profile. - Acted-on control communities are scrubbed from the wire-bound attribute set
toward enabled sessions (standard admin
0/RSoutside the0xFFFF____well-known space; largeRS:{0,1,101,102,103}:*); disabled sessions keep RFC 7947 §2.2 byte-level transparency. Enabled sessions stay in shared update-groups: the filter is route-granular at emit (ADR-0101 Decision 3), so only routes carrying a control-form community pay per-target divergence while untagged routes share staging and encoding fleet-wide.
RFC 8212 — Default eBGP route handling
- RFC 8212 makes an eBGP route without an explicit import policy ineligible
for the decision process and keeps a route without an explicit export
policy out of that peer's Adj-RIB-Out. rustbgpd implements the boundary
behind the opt-in
[global] ebgp_requires_policyknob (restart-required): when on, an eBGP session that resolves no explicit operator policy in a direction runs a reserved internal deny-all chain (rfc8212_missing_import_policy/rfc8212_missing_export_policy— both reserved names that operator policy cannot shadow) in that direction. The two directions are independent, and the session stays Established, so the gap is repairable without transport churn. - Default: epoch-scoped (ADR-0119). Under
config_epoch = 2an omittedebgp_requires_policyresolves to the RFC-mandated secure default — effectivetrue, sourceepoch_2_default— so epoch-2 configs comply with the RFC's default-deny requirement without writing the knob. Epoch-less and epoch-1 omission remain permit-all permanently: the historical default is permit-all when a session resolves no policy chain, and flipping it under existing configs would make an upgrade silently drop routes in a way indistinguishable from a policy change — the warn-first posture RFC 8212 Appendix A.1 describes for implementations with an installed base. Explicittrue/falsekeeps its stated value in every epoch.rustbgpd --checkwarns on every eBGP neighbor missing explicit policy whether enforcement is on or off, every shipped starter config with eBGP neighbors sets the knob, and all starters passrustbgpd --check --strict(fenced bytests/starter_configs_check_strict.rs). - Migration surface (ADR-0119). The knob's posture is carried by a root
config_epoch = 1 | 2key alongside a lossless record of whetherebgp_requires_policywas written at all. Epoch 1 pins the legacy permit-all default; epoch 2 activates the RFC-aligned default for the omitted boolean. A config that states neither resolves toLegacyOmissionand raises therfc8212_secure_default_readyadvisory:rustbgpd --checkwarns and exits 0,--check --strictexits 1.rustbgpd --migrate-config pin-legacy|prepare-secure --offlineperforms the in-place move to an explicit posture. - Full knob semantics, what counts as explicit policy, and the observability
surfaces (neighbor detail, Prometheus, doctor, export explain) are in
docs/reference/configuration.md under
"
ebgp_requires_policy— RFC 8212 explicit policy on eBGP".
Inbound filtered-replacement semantics
- Import-policy denial of a replacement retires the exact previously accepted
FlowSpec
(AFI, rule), EVPN route key, VPN(RD + prefix, path_id), labeled-unicast(prefix, path_id), RTC(NLRI, path_id), or BGP-LS(family, NLRI, path_id)identity. First-seen denials remain filter-only, explicit overlapping withdrawals are deduplicated, and other AFI or Add-Path siblings remain intact. AS_PATH-loop and route-reflector-loop rejection likewise retires an exact previously accepted VPN, labeled-unicast, RTC, or BGP-LS replacement. First-seen and repeated rejected announcements remain silent, distinct validMP_UNREACH_NLRIwithdrawals still propagate, and exact Add-Path siblings remain intact.
RFC 9552 and RFC 9107 — ORR topology scope
- ORR builds only the default BGP-LS topology. Link and Prefix NLRIs use a
single descriptor MT-ID; Node membership comes only from Multi-Topology TLV
263 inside BGP-LS Attribute 29. IS-IS interprets the lower 12 bits, while
OSPF requires the reserved high bits clear and a value in
0..=127. - Absent MT-ID is default. A Node list containing zero is default even when it also advertises non-zero memberships. Valid non-zero-only objects are excluded; duplicate, wrongly placed, structurally invalid, or protocol-uninterpretable topology data is excluded fail-closed before graph insertion.
- Flex-Algorithm definition/prefix inputs and Flex Prefix-SID/SR-Algorithm values do not select another SPF. They are reported as ignored aggregate input while the valid base default object and classic IGP/Prefix Metric stay active.
- An IPv6 prefix carrying an SRv6 Locator TLV (1162) contributes reachability only with a valid Prefix Metric TLV (1155), as required by RFC 9514 §5.1 and verified erratum 7737. Locator-only advertisements remain available for raw BGP-LS reflection. That TLV is inapplicable to IPv4 prefixes and does not change their ORR reachability. Ordinary prefixes retain metric zero when the Prefix Metric is absent or unreadable.
RFC 9234 — Roles and Only-to-Customer
- OPEN negotiation reads every received capability code 9. The wire decoder
keeps an unassigned value (5-255) or a length other than 1 as an unknown
code-9 capability with its raw bytes. Negotiation treats that capability as
a received Role that matches no Table 2 pair, so a session with a configured
local Role is rejected with Role Mismatch (2/11), with or without
strict_role. Section 4.2's multiple-capability rule compares raw bytes and applies with or without a local Role:[Customer, 7]or[7, 8]is rejected with 2/11. Without a local Role, a single unassigned or wrong-length Role capability is accepted and no remote role is recorded. - The wrong-length case has no RFC 9234 MUST, so its handling is a local
decision. FRR (
bgp_capability_role) and BIRD (bgp_read_capabilities) reject a Role length other than 1 as a malformed OPEN (2/0) even without a local Role. rustbgpd keeps one rule for both invalid forms: such a capability carries no Table 2 value, so it fails the Role check (2/11) only when a local Role is configured, and is otherwise ignored like any capability the speaker does not act on (RFC 5492 section 3). Both implementations also compare the raw value against Table 2, so an unassigned value is a Role Mismatch there too when a local Role is configured. BIRD additionally rejects value 255, which it uses internally for "no Role", even without one. bgp_role_mismatch_totalreports the first assigned Role value in the rejected OPEN asremote_role, so[Customer, 7]counts asremote_role="customer". When the OPEN carries Role capabilities but no assigned value, only unassigned or wrong-length ones, the label isremote_role="unrecognized"and the warning log carries the first raw value asremote_role_raw, for example[7].remote_role="none"means the OPEN carried no Role capability, as with an absent Role understrict_role.- The configured local Role is session-stamped into the RIB before
PeerUp, so the first Adj-RIB-Out build and every subsequent export use the same RFC 9234 relationship semantics. Update-group identity includes that role; peers with different OTC egress behavior cannot share advertised state. - E2 suppression for IPv4/IPv6 unicast happens after export-policy modifications but before grouped or private Adj-RIB-Out commit. This covers single-best, ORR, Add-Path, and per-client-best selection. A route that was previously advertised and becomes OTC-blocked is withdrawn and removed from logical advertised state; a newly blocked route is never committed.
- Transport retains the E2 check as a defense-in-depth encoder guard and owns
the established
bgp_otc_routes_blocked_total/OTC_ROUTE_BLOCKEDdiagnostic publication. The RIB passes rejected route context explicitly, so metrics/events and export-explain reflect the same pre-commit decision. Backpressured grouped peers retain at most one pending diagnostic per route; resync rebuilds that residue from current denials so withdrawn sources do not leak stale events into an unrelated later update. - RFC 9234 section 5 applies only to IPv4/IPv6 unicast SAFI 1 here. FlowSpec, EVPN, VPN, labeled-unicast, RTC, and BGP-LS are not subject to the OTC gate.
Milestone 0 — RFC 4271 Sections
§4.2 — OPEN Message
- Hold Time negotiation: Use the minimum of local and remote proposed
hold times. If the negotiated value is non-zero and less than 3 seconds,
send NOTIFICATION (2, 6) — Unacceptable Hold Time. Zero means no
keepalives (supported but discouraged in config docs).
With
min_hold_time, zero and peer proposals below the configured floor are rejected with the same NOTIFICATION before negotiation. - BGP Identifier: RFC 6286 defines this as any non-zero unsigned 32-bit integer rendered in dotted-quad form; multicast, loopback, and other non-unicast-shaped values are valid. Zero and an iBGP peer using the local identifier get Bad BGP Identifier. Equal identifiers are allowed for eBGP; during a connection collision, the connection initiated by the speaker with the larger AS number is preserved, including for four-octet ASNs (§2.3).
- Version: Only BGP-4 (version 4). Any other version gets NOTIFICATION (2, 1) — Unsupported Version Number, with data field containing the supported version (4).
- My AS: 4-byte ASN support via RFC 6793 capability. If the peer does not advertise 4-byte ASN capability, we use 2-byte AS in OPEN and set AS_TRANS (23456) if our ASN > 65535.
§4.3 — UPDATE Message
- Wire-level decode implemented in M0. Full processing (NLRI, path attributes, validation, RIB population) implemented in M1.
- NLRI uses prefix-length encoding: 1 byte prefix length + ceil(len/8) bytes of address. Host bits are masked off on decode.
- Path attribute TLV: flags(1) + type(1) + length(1 or 2) + value. Extended Length flag (0x10) controls 2-byte length field.
- 2-byte vs 4-byte AS_PATH encoding controlled by
four_octet_ascapability negotiated in OPEN. - Structural decode (can I read these bytes?) separated from semantic validation (is the attribute set RFC-compliant?). See ADR-0012.
- Outbound IPv4/IPv6 unicast and IPv4/IPv6 FlowSpec announcements and withdrawals are chunked by the peer's negotiated 4096/65535-byte message limit. MP chunking starts with at most 1,024 entries, grows through exactly built and size-checked candidates up to a 4,096-entry probe ceiling, and retains successful/failed bounds; it does not promise to fill every Extended Message. Structured FlowSpec construction is fallible; an individually unencodable NLRI fails the session rather than partially committing a batch.
- EVPN MP_REACH/MP_UNREACH is likewise chunked to the negotiated ceiling. A single announcement or withdrawal that still cannot fit invokes the Cease/8 outbound-saturation teardown, preventing a live session from retaining logical Adj-RIB-Out state that never reached the wire.
- Before any announcement enters Adj-RIB-Out, the RIB probes its exact one-route wire form through an immutable snapshot of that session's live encoder and negotiated 4096/65535-byte ceiling. This post-policy check covers unicast, FlowSpec, EVPN, BGP-LS, VPN, labeled-unicast, and RT-Constrain. A failure is rejected before commit; if the identity was previously advertised, the same transition emits its withdrawal. Grouped peers keep a sparse per-member rejection overlay, so peers with different negotiated ceilings retain exact individual advertised views while sharing the staged group table. Recompute/resync retries rejected routes, and a source withdrawal retires the rejection without a redundant wire withdrawal. Transport's Cease/8 path remains the final defense if a route-bearing envelope lacks the matching snapshot or a live encoder still finds an impossible single-route UPDATE.
- FlowSpec identity is
(AFI, rule)throughout Adj-RIB-In, Loc-RIB, Adj-RIB-Out, recompute, distribution, and withdrawal. AFI is never inferred from an optional destination-prefix component: legal destination-less IPv4 and IPv6 rules remain distinct.
§4.4 — KEEPALIVE Message
- Sent at negotiated hold_time / 3 interval.
- If hold_time is 0, no KEEPALIVEs are sent or expected.
§4.5 — NOTIFICATION Message
- Error codes are the typed
NotificationCodeenum: codes 1-8 have named variants (code 7 isRouteRefreshMessagefor RFC 7313 errors, from wire 0.22.0), andUnknown(u8)preserves any other byte, including 9. Subcodes are namedu8constants per code, such ascease_subcode::OUT_OF_RESOURCESandroute_refresh_subcode::INVALID_MESSAGE_LENGTH. - On send: log structured event, then close the TCP connection.
- On receive: log structured event, transition FSM to Idle.
§6.8 — BGP Identifier Collision
- If an OPEN is received from a peer with the same BGP Identifier as an existing session, the collision resolution procedure applies: compare local and remote BGP Identifiers as unsigned 32-bit integers. The connection initiated by the higher ID is kept. If the identifiers are equal for an eBGP collision, RFC 6286 §2.3 keeps the connection initiated by the speaker with the larger AS number, including a four-octet AS number.
- The comparison applies only while the configured session is in OpenConfirm or OpenSent. Its state is read when the collision is resolved. A valid OPEN on either connection can trigger resolution while an inbound candidate is pending; an identity already known at accept can also be used.
- The OpenSent case is optional in the RFC: "A BGP speaker MAY also examine
connections in an OpenSent state if it knows the BGP Identifier of the
peer by means outside of the protocol." rustbgpd deliberately counts the
identifier in the inbound connection's OPEN as that knowledge, as
bgp_collision_detectin FRR 10.7.1 does. - The inbound candidate withholds its KEEPALIVE and cannot reach Established before the manager's verdict. This hold applies only to the candidate; the configured primary sends its KEEPALIVE after a valid OPEN normally. It does not guarantee that simultaneous active opens keep one connection without a retry.
- The speakers can observe different states on the same connection. If
rustbgpd still sees OpenConfirm when the remote identifier wins, it closes
its primary. The remote speaker may already have received that primary's
KEEPALIVE and reached Established, causing it to retain that connection
and close the other one. Both connections can therefore close, followed by
a reconnect attempt. This follows from the state-dependent rules in
RFC 4271 §6.8
and the OPEN/KEEPALIVE transitions in
§8.2.2.
Configuring the other speaker as passive, where supported, avoids
simultaneous active opens; FRR provides
neighbor PEER passive. - A configured session in Idle, Connect or Active has no connection to collide with (RFC 4271 §6.8), so the inbound connection replaces it without a Cease; its outbound connect attempt and reconnect timer stop. Two exceptions keep the configured session and close the inbound connection with Cease 6/7 instead, when they already hold as the OPEN is resolved: the neighbor is administratively disabled, or BFD is holding BGP down for it. While the neighbor waits out an escalated NOTIFICATION reconnect backoff, the inbound connection is closed without a NOTIFICATION.
- A configured session that is already Established keeps its connection; the inbound one is closed with Cease 6/7.
- If the configured session's state cannot be read within the peer-query deadline when the inbound connection's OPEN arrives, it is treated as possibly Established: it keeps its connection and the inbound one is closed with Cease 6/7. A timeout at TCP accept is handled differently; see the last item.
- Every Cease 6/7 above is sent when collision resolution runs using a known peer identity. Disabling the neighbor, or a BFD down that holds BGP, also tears down a pending inbound candidate at that moment, independently of any OPEN. That teardown is an ordinary stop: a candidate that has reached Established sends Cease 6/2 (Administrative Shutdown), and one in any earlier state, OpenSent and OpenConfirm included, closes its TCP connection without a NOTIFICATION.
- The Cease 6/7 outcomes above apply to an inbound candidate session, which
exists only after the TCP connection is accepted. At accept, before any
candidate exists, the raw TCP connection is closed without a NOTIFICATION,
and before any BGP message is exchanged, when any of these holds:
- the neighbor is disabled or BFD-held;
- the configured session is Established or in NOTIFICATION backoff;
- the query for its state times out.
§8 — Finite State Machine
- All six states implemented: Idle, Connect, Active, OpenSent, OpenConfirm, Established.
- All timers modeled as inputs (not spawned internally): ConnectRetry, Hold, Keepalive.
- DelayOpen timer: not implemented in v1 (RFC 4271 §8 optional).
- Exponential backoff on connect retry:
base * 2^counter, capped at 300s, reset on ManualStart or reaching Established. - DampPeerOscillations (RFC 4271 §8.1.1) is not implemented as the
optional attribute with its IdleHoldTimer. The transport layer applies a
fixed equivalent: each consecutive fall to Idle caused by a NOTIFICATION
(sent or received, including an OPEN exchange that ends in one) doubles
the Idle reconnect wait from
connect_retry_secs, capped at 300s and never below the configured interval. The streak clears after the session stays Established for five minutes, on an explicit operator enable, or on an administrative reset. TCP connection misses in Connect or Active keep the fast initial retries followed by the existing exponential curve; a TCP loss after Established uses the configured fixed Idle interval. Neither advances the NOTIFICATION streak, and there is no configuration knob for its curve. - Initial hold timer before OPEN negotiation: 240s (RFC 4271 "large value"), replaced by negotiated value once OPEN exchange completes.
handle_eventnever returnsResult— every (State, Event) pair produces a well-defined output. Invalid events in any state produce a NOTIFICATION (FSM Error) and transition to Idle.SessionDownis normally emitted when leaving Established state. A BFD Down event in OpenSent or OpenConfirm also emits it after the typed Cease notification so the shared transport teardown path clears cached notification state; other failed handshakes are not surfaced as session-down events.StateChangedaction emitted on every state transition for telemetry.
§8.2 — Timers (transport implementation)
- Timers are
Option<Pin<Box<Sleep>>>in the transport layer.Nonemeans the timer is stopped;Somemeans it is running. - A freestanding
poll_timerfuture is used intokio::select!to avoid&mut selfborrow conflicts with other select branches. - When a timer fires, the transport clears the slot (
= None) before feeding the event to the FSM. The FSM may restart the timer via aStartTimeraction in the same event cycle.
§8.2 — TCP Connection Management (transport implementation)
- Transport uses
Option<TcpStream>for connection state.Nonewhen disconnected; the TCP read branch ofselect!is disabled via a guard (if stream.is_some()). InitiateTcpConnectionaction triggersTcpStream::connectwith a configurable timeout. The result is returned as a follow-up FSM event (TcpConnectionConfirmedorTcpConnectionFails).- Send failures (OPEN, KEEPALIVE) are treated as TCP failures: the
stream is dropped and
TcpConnectionFailsis queued. CloseTcpConnectiondrops the stream and clears the read buffer.
§10 — Error Handling
- Every error condition maps to a specific NOTIFICATION code/subcode.
- No generic error paths. Each failure has a unique structured event.
RFC 6793 — 4-Byte ASN Support
Capability Advertisement
- Capability code 65, length 4, containing our 4-byte ASN.
- If peer advertises this capability, 4-byte AS_PATH segments are used.
- If peer does not advertise it, 2-byte AS encoding is used with AS_TRANS (23456) substitution where necessary.
AS_TRANS Handling
- On a session without capability 65, inbound
AS_PATH/AS4_PATHandAGGREGATOR/AS4_AGGREGATORpairs are normalized to one four-octet logical path and one typed aggregator before loop detection, policy, and RIB consumers run. The compatibility attributes do not enter long-lived route state. - When encoding for a 2-byte-only peer, every non-mappable ASN becomes
AS_TRANSinAS_PATHand a compensatingAS4_PATHcarries the complete final logical path. A non-mappable aggregator is projected as six-octetAGGREGATORwithAS_TRANSplus an eight-octetAS4_AGGREGATOR; mappable paths and aggregators gain no redundant compatibility attribute. - A four-octet-capable peer receives ordinary four-octet
AS_PATHand eight-octetAGGREGATORonly. Received types 17/18 are discarded after the raw RFC 9774 inspection and cannot be propagated onward. - Live sending and exact-export sizing share this target-specific encoder, so generated AS4 bytes count against the negotiated message ceiling before Adj-RIB-Out advances. See ADR-0114.
RFC 7607 — AS 0 Rejection
- Strict decoding rejects AS 0 in
AS_PATH,AS4_PATH,AGGREGATOR, andAS4_AGGREGATOR. Revised decoding maps ordinaryAS_PATHto treat-as-withdraw and discards the other three attributes. - AS 0 never participates in RFC 6793 reconstruction or canonical route state. Canonical encoding rejects it before deriving type 17/18 compatibility attributes.
- Policy cannot introduce AS 0: a TOML
set_as_path_prependwith ASN 0 fails config validation, a literal.rpolprepend as 0is a compile error, a parameter that resolves to AS 0 in a prepend is rejected when the daemon attaches the chain, and a computed prepend operand that evaluates to 0 denies the route. See policy actions.
AS_PATH Element Ceiling (max_as_path_length)
[global] max_as_path_length(default 750,0disables) bounds the number of AS numbers accepted in a receivedAS_PATH, counted across every segment with prepends andAS_SETmembers included. A longer path carrying reachable NLRI is treat-as-withdraw through the RFC 7606 validation path (subcode 11): its routes are withdrawn, the session stays Established, andbgp_update_malformed_total{disposition="treat_as_withdraw"}increments. Without reachable NLRI, RFC 7606 section 5.2 requires a session reset. The log line carries the ceiling, never the path.- The ceiling is an operational guard (draft-ietf-grow-bgpopsecupd), not an
RFC requirement; OpenBGPD applies the same 750-element limit. The wire
crate exposes
validate_as_path_ceiling;0leaves it off.
RFC 9774 — AS_SET / AS_CONFED_SET Deprecation
- The revised inbound attribute decoder inspects raw AS_PATH and AS4_PATH segment framing for AS_SET and AS_CONFED_SET before duplicate discard, and assigns RFC 7606 treat-as-withdraw on every session and family. When the UPDATE carries no reachable NLRI, the existing RFC 7606 §5.2 composition escalates that disposition to a session reset.
- Surviving
AS4_PATHvalues are then parsed for RFC 6793 reconstruction. AnAS_CONFED_SEQUENCEis removed from type 17 and processing continues; this narrow rule does not add confederation segments to the route model.
Interpretation Decisions
These are deliberate choices where the RFC is ambiguous or permits multiple behaviors. Each is documented here for auditability.
Partial Bit Policy
The Partial bit (flag 0x20) is meaningful only on an optional transitive attribute, and rustbgpd handles three cases distinctly. None is configurable in v1.
- Attributes rustbgpd does not parse — Partial is OR'd on re-advertisement. This covers an unrecognized optional transitive attribute and, since v0.67.0, an IANA-assigned optional transitive type whose payload semantics are deliberately unsupported: Tunnel Encapsulation (23), IPv6 Address Specific Extended Community (25), PE Distinguisher Labels (27), Community Container (34), D-PATH (36), SFP (37), BFD Discriminator (38), NHC (39), BGP Prefix-SID (40), BIER (41), and ATTR_SET (128). Registry recognition fences those types' class and outer framing; it is not a claim to understand the value, so the conservative signal still applies. All other flags and the value bytes are preserved unchanged, including an explicitly received Extended Length encoding — re-emitting a one-octet length under a retained Extended Length flag would corrupt the attribute boundary.
- Recognized optional transitive attributes — a received Partial is
preserved, never invented. AGGREGATOR (7), Communities (8), Extended
Communities (16), PMSI Tunnel (22), Large Communities (32), and
Only-to-Customer (35) decode into typed values that carry the received
Partial bit as state (
PathAttribute::CommunitiesPartialand its siblings,Aggregator::partial), so it survives policy and the RIB and is re-emitted on egress under both compact and Extended Length framing. rustbgpd validated these values, so it never adds Partial to one that arrived without it — including one that export policy modified, since the mutating accessors keep the typed variant. Locally originated values are canonical and carry Partial clear. - Everything else — Partial is never set. Well-known attributes (Optional=0) are never marked Partial. Optional non-transitive attributes MUST carry Partial clear (RFC 4271 §4.3); rustbgpd enforces that on receipt for MP_REACH_NLRI, MP_UNREACH_NLRI, the BGP-LS Attribute, and the four assigned optional non-transitive types (Traffic Engineering 24, AIGP 26, BGPsec_PATH 33, Edge Metadata 42). Unrecognized optional non-transitive attributes are ignored on receipt and omitted by the defensive encoder (RFC 4271 §5), so they have no egress form to mark.
Rationale: Partial states that a speaker on the path did not fully process the attribute. Setting it on anything rustbgpd only reflected byte-for-byte is the correct conservative signal to downstream peers, and matches FRR, BIRD, and most production implementations. Setting it on a value rustbgpd did parse and validate would be a false claim, which is why the recognized types carry the received bit through rather than OR one in.
NHC opaque transit
NHC (39) is framing-validated opaque data, with unsupported payload semantics. Import and export next-hop rewrites preserve its payload; rustbgpd neither constructs NHC characteristics nor rebuilds the next-hop header. draft-ietf-idr-nhc-07 §1.2 limits supporting-speaker requirements to implementations supporting and enabling the specification, and §2.2 explicitly permits opaque propagation through unsupported speakers without updating NHC. A supporting receiver must compare the encoded next hop before using the characteristics (§2.3). The rewrite requirement therefore does not require NHC removal under rustbgpd's current support boundary. This remains an Internet-Draft, not a published RFC. The registry boundary records framing/error handling and the existing transport retention evidence.
Cease Subcode 8 (Out of Resources)
rustbgpd sends NOTIFICATION Cease with subcode 8 (Out of Resources, RFC 4486 §3) when one session's outbound path cannot continue: the bounded outbound writer queue cannot admit pending output before its resource deadline, or a committed outbound update cannot be sent exactly (a missing, incompatible, or foreign export snapshot, or structurally unsendable output). The session is torn down; the maximum-prefix limit uses Cease subcode 1 instead. There is no fallback to another subcode.
Message Size Limits (RFC 4271 + RFC 8654)
RFC 4271 §4.1 defines a 4096-byte maximum unless Extended Messages (RFC 8654) are negotiated. rustbgpd enforces negotiated limits:
- Inbound: Message length > negotiated max is rejected with NOTIFICATION (1, 2) — Bad Message Length. The raw length value is included in the NOTIFICATION data field.
- Outbound: Encode attempts beyond the negotiated max return an internal encode error and the message is not sent.
- Negotiation behavior: Sessions start at 4096-byte framing. If both peers advertise capability code 6, max message length is raised to 65535 for that session; on session-down it resets to 4096.
Hold Time Floor
If the negotiated hold time is non-zero and less than 3 seconds,
rustbgpd sends NOTIFICATION (2, 6) — Unacceptable Hold Time. This
prevents pathologically short hold times that would cause false flaps.
RFC 4271 recommends a minimum of 3 seconds; we enforce it.
An optional per-neighbor or peer-group min_hold_time raises that acceptance
floor and also rejects a zero proposal with NOTIFICATION (2, 6).
RFC 7606 — Revised BGP UPDATE Error Handling
Implemented on the inbound UPDATE path: a malformed path attribute no longer
tears the session down by default. wire::validate::ErrorDisposition
classifies every attribute error as attribute-discard, treat-as-withdraw, or
session-reset; the transport applies the strongest disposition found in the
message (§3 (h)).
- Treat-as-withdraw (§2): the UPDATE's routes — body NLRI and every MP family — are withdrawn from the Adj-RIB-In (previously accepted routes for the same NLRI are removed), announcements are dropped, explicit withdrawals still apply, and the session stays Established. Covers ORIGIN, AS_PATH (§7.2 — no longer a session reset), NEXT_HOP, MULTI_EXIT_DISC, communities of all sizes, attribute-flag conflicts (§3 (c)), missing mandatory attributes (§3 (d)), and attribute-section framing overruns (§4), except visible MP_REACH_NLRI / MP_UNREACH_NLRI framing failures whose embedded NLRI cannot be parsed (§5.3).
- Attribute-discard (§2): ATOMIC_AGGREGATE with a non-zero length and
AGGREGATOR with a length other than 6/8 (§7.6, §7.7) are dropped and the
UPDATE proceeds. Malformed LOCAL_PREF / ORIGINATOR_ID / CLUSTER_LIST from
an external neighbor are discarded (§7.5, §7.9, §7.10); from an internal
neighbor they are treat-as-withdraw. Duplicate attributes keep the first
occurrence and discard the rest (§3 (g)).
§7.9 and §7.10 are not limited to malformed attributes: a well-formed
ORIGINATOR_ID or CLUSTER_LIST from an external neighbor is also discarded,
for every address family, before the route is stored and selected, and the
RFC 4456 reflection-loop check does not act on it. These removals are counted
in
bgp_path_attribute_discarded_total{type_code="9"|"10"}; pre-policy BMP still mirrors the UPDATE as received. The discard does not cover a wrong Optional/Transitive flag class on these attributes: that is a §3 (c) treat-as-withdraw from any neighbor, and the UPDATE's routes are withdrawn before the discard applies. - Session-reset is retained only where the NLRI cannot be trusted: UPDATE section-length inconsistencies (§3 (b), unchanged), a structurally unparseable or duplicated MP_REACH_NLRI / MP_UNREACH_NLRI (§7.11, §3 (g)), syntactically incorrect NLRI or Withdrawn Routes fields (§5.3), and the §5.2 escalation (a treat-as-withdraw-class error in an UPDATE that encodes no reachable NLRI). A visible MP attribute with incomplete header/value framing uses UPDATE Optional Attribute Error (3/9) with exactly the received attribute bytes, as recommended by RFC 4760 §7; ordinary attribute-section overruns remain treat-as-withdraw. A complete MP attribute shorter than its minimum value uses Attribute Length Error (3/5); other intact structural MP decode failures (including AFI/SAFI decoding, next-hop length/framing, and embedded-NLRI syntax) use 3/9 with the exact complete attribute bytes. A byte-complete next hop that parses but fails semantic address validation (for example, an unspecified IPv6 address) is treat-as-withdraw with Invalid NEXT_HOP (3/8), not session-reset. Duplicate MP attributes remain 3/1 with empty data, and flag conflicts remain 3/4 with exact attribute data.
Assigned attributes whose payload is not parsed (v0.67.0):
Fifteen IANA-assigned path attribute types are recognized by registry entry only. The class fence and the framing walk below are the whole of that recognition — no typed model, no semantic, and no support claim for any of them. Before v0.67.0 an unparsed assigned type was either silently ignored or retained opaquely regardless of how its flags and framing were encoded.
- Class fencing. Each of the fifteen carries its registered Optional/Transitive class through decode. A class conflict is treat-as-withdraw (§3 (c)) instead of being silently accepted, with one exception: an AIGP attribute carrying Transitive is attribute-discard, because RFC 7311 §3.2 is more specific than §3 (c). Correctly classed values keep their ordinary treatment — the four optional non-transitive types (Traffic Engineering 24, AIGP 26, BGPsec_PATH 33, Edge Metadata 42) are ignored on receipt and never retained, stored, or emitted, including on egress; the eleven optional transitive types (23, 25, 27, 34, 36-41, 128) are retained opaquely and re-advertised with the Partial bit.
- Structural framing. Ten of the transitive types additionally get a bounded, syntax-only walk of their outer framing before opaque propagation: Tunnel Encapsulation (23) must carry at least one Tunnel TLV, and each Tunnel TLV and its sub-TLV stream must frame exactly; IPv6 Address Specific Extended Community (25) must be a non-empty multiple of 20 octets; Community Container (34) walks container headers, its Type 1 subtypes — which may not repeat — the atom TLV stream, and atom prefix lengths; D-PATH (36) walks non-empty 7-octet domain segments; SFP (37) requires a Hop TLV and walks its sub-TLVs; BFD Discriminator (38) requires a Source IP TLV of length 4 or 16; NHC (39) walks the next hop and requires at least one characteristic; BGP Prefix-SID (40) length-checks the Label-Index and Originator SRGB TLVs and validates RFC 9252 Service nesting; BIER (41) must carry at least one TLV and walks its sub-TLV nesting; and ATTR_SET (128) requires its 4-octet Origin AS followed by an exactly framed embedded attribute stream that contains no MP_REACH_NLRI or MP_UNREACH_NLRI. Values that pass stay opaque bytes.
- Framing dispositions split by type. A framing failure in BFD Discriminator (38), NHC (39), generic BGP Prefix-SID (40), or BIER (41) is attribute-discard: the malformed value cannot affect route selection, so the UPDATE's routes survive without it. Recognized SRv6 L3/L2 Service TLV malformation instead follows RFC 9252 §7 treat-as-withdraw. A framing failure in Tunnel Encapsulation (23), IPv6 Address Specific Extended Community (25), Community Container (34), D-PATH (36), SFP (37), or ATTR_SET (128) is treat-as-withdraw. A class conflict on any of the ten is treat-as-withdraw regardless — §3 (h) takes the stronger of the two actions. Zero-length BIER is the case that changed disposition in v0.67.0: it was previously retained as though it held a valid TLV sequence, and is now attribute-discard.
- Duplicate Community Container (34) is treat-as-withdraw rather than the §3 (g) keep-first discard, because the container model has no defined merge for a second instance. Both the strict and the revised decoder enforce it.
- Unrecognized optional non-transitive attributes (no registry entry) are ignored on receipt by both decoders and omitted by the defensive encoder per RFC 4271 §5, instead of being retained and re-advertised. Unknown optional transitive attributes are still preserved and propagated with Partial.
- The strict (RFC 4271) decoder reports each of these as a typed UPDATE attribute error — Attribute Flags Error for a class conflict, Attribute Length Error for a framing failure — rather than reaching an unreachable branch. Its residual arm for a correctly classed optional non-transitive assigned type returns Optional Attribute Error; that arm is defensive, since the decode loop drops those attributes before value decoding.
Interpretation decisions:
- An unrecognized attribute claiming to be well-known (Optional=0, unknown type) is handled treat-as-withdraw. RFC 7606 §3 (c) only literally covers flag conflicts on recognized attributes, but resetting the session for a tunneled unknown attribute is exactly the amplification §1 warns about; FRR makes the same call.
- Flag conflicts (§3 (c)) are treat-as-withdraw for every attribute, including ATOMIC_AGGREGATE and AGGREGATOR — their §7.6/§7.7 attribute-discard covers length malformations only. On MP_REACH_NLRI / MP_UNREACH_NLRI a flag conflict stays session-reset (§5.3 lists inconsistent flags among what makes the MP attribute itself incorrect).
- Zero-length Communities / Extended Communities are malformed (§7.8 / §7.14 require a non-zero multiple of 4 / 8) — treat-as-withdraw.
- A zero-length CLUSTER_LIST is malformed (Attribute Length Error) — attribute-discard from an external neighbor and treat-as-withdraw from an internal one (§7.10).
- An UPDATE whose only attributes were discarded as malformed is not mistaken for an End-of-RIB marker (RFC 4724 §2 detection requires a clean decode).
- Malformed attributes are logged at
warnwith peer, attribute type, and disposition. The §6 debugging facility is a DEBUG-level dump emitted on every malformed event: the entire UPDATE message as hex (untruncated — message size is protocol-bounded at 4096 bytes, or 65535 with Extended Messages) plus an explicit enumeration of the NLRI involved (prefixes and Add-Path path IDs, per family, announcements and withdrawals). DEBUG-only so a hostile peer cannot spam operators at info/warn; nothing is rendered when DEBUG is disabled.bgp_update_malformed_total{peer,disposition}counts each malformed UPDATE once under the strongest applied disposition:attribute_discard,treat_as_withdraw, orsession_reset.bgp_update_malformed_causes_total{peer,type_code,reason,disposition}counts the individual reported causes, several of which can occur in one UPDATE; see ingress rejection metrics. - AS4_PATH / AS4_AGGREGATOR are parsed as short-lived compatibility inputs and normalized into the typed path / aggregator model. Malformed type 17/18 values are attribute-discard; conflicting Optional/Transitive flags remain the stronger treat-as-withdraw action. Raw types 17/18 never survive in the canonical route.
Milestone 1 — RFC 4271 Sections
§5.1.1 — ORIGIN Attribute
- Decoded from 1-byte value: 0=IGP, 1=EGP, 2=INCOMPLETE.
- Well-known mandatory. Flags must be Optional=0, Transitive=1.
§5.1.2 — AS_PATH Attribute
- Segments decoded as type(1) + count(1) + ASNs(2 or 4 bytes each).
- Segment types: AS_SEQUENCE (2), AS_SET (1).
- Empty segments (count=0) are rejected as malformed (NOTIFICATION 3,11).
- 4-byte ASN encoding used when
four_octet_ascapability is negotiated.
§5.1.3 — NEXT_HOP Attribute
- 4 bytes decoded as IPv4 address.
- Validated: 0.0.0.0, 127.0.0.0/8, 224.0.0.0/4, 255.255.255.255 are all rejected with NOTIFICATION (3, 8) — Invalid NEXT_HOP Attribute.
- Mandatory for eBGP with NLRI. Not required for iBGP (may be omitted or set by the transport layer).
§5.1.4 — MULTI_EXIT_DISC (MED) Attribute
- 4 bytes decoded as u32.
- Optional non-transitive. Used in best-path step 4 (deterministic always-compare mode).
§5.1.5 — LOCAL_PREF Attribute
- 4 bytes decoded as u32.
- Well-known mandatory (iBGP scope). Used in best-path step 1 (highest wins, default 100).
§6.3 — UPDATE Message Error Handling
- All validation checks produce specific NOTIFICATION subcodes:
- (3,1) Malformed Attribute List — duplicate type codes
- (3,2) Unrecognized Well-known Attribute — Optional=0 + unknown type
- (3,3) Missing Well-known Attribute — ORIGIN, AS_PATH, NEXT_HOP (eBGP)
- (3,4) Attribute Flags Error — well-known with wrong Optional/Transitive
- (3,8) Invalid NEXT_HOP Attribute — reserved/multicast/loopback address
- (3,11) Malformed AS_PATH — empty segment
- Validation is separate from decode (ADR-0012). Withdrawal-only UPDATEs (zero attributes) pass decode fine and skip validation.
§9.1 — Adj-RIB-In
- Per-peer
AdjRibInstores routes keyed by(Prefix, u32)(prefix + path_id for Add-Path support). - Insert replaces existing route for the same prefix.
- Withdraw removes by prefix, returns whether the route existed.
- PeerDown clears all routes for that peer.
- Single
RibManagertokio task owns all Adj-RIB-In state (ADR-0013).
§9.1.2.2 — Decision Process tie-breaking
- Step (f), "prefer the route received from the speaker with the lowest BGP Identifier", is implemented for the unicast, VPN, labeled-unicast, FlowSpec, BGP-LS, and RT-Constrain decision chains. The identifier a route is compared by is its ORIGINATOR_ID when present, otherwise the BGP Identifier from the advertising peer's OPEN (RFC 4456 §9). The step runs after eBGP-over-iBGP (and the RFC 9107 ORR interior-cost step when one applies) and before the shorter-CLUSTER_LIST comparison, which in turn precedes step (g), lowest peer address.
- Locally originated routes have no advertising speaker and no BGP Identifier. The step ranks every one of them ahead of every session-learned route, and two of them tie, so every route has a position and the ordering stays transitive. The daemon's own BGP Identifier is not substituted. FRR, BIRD, and GoBGP all prefer local routes over learned ones before this step; rustbgpd keeps them on the same side without moving that preference earlier.
- Explain output and BMP path marking report the step as
lower_originator_idwhen both routes carried ORIGINATOR_ID and aslower_bgp_identifierotherwise; both map to the same path-marking "router ID" reason code. - RFC 5004 (retain the existing external best path across identifier changes) is not implemented.
RFC 4760 — Multiprotocol Extensions for BGP-4
§3 — MP_REACH_NLRI (Type 14)
Wire layout:
AFI (2 bytes) | SAFI (1) | NH-Len (1) | Next Hop (variable) | Reserved (1) | NLRI (variable)- Flags: Optional + Non-Transitive (
0x80). The Extended Length bit may be added for framing; Partial is invalid. - For IPv6 unicast, next-hop length is 16 bytes (global IPv6 address) or
32 bytes (global + link-local). When 32 bytes, rustbgpd takes the first 16
as the primary
next-hop and preserves the trailing 16 in
link_local_next_hop(round-tripped through wire / RIB / MRT since v0.11.0); ADR-0069 resolves a link-local next-hop as a scoped next-hop for unnumbered IPv4-over-IPv6 and Linux FIBdev. - NLRI: same prefix-length encoding as IPv4, but up to 128 bits (16 bytes of address data).
- When
MP_REACH_NLRIis present in an UPDATE, the body NEXT_HOP attribute (type 3) is not required — the next-hop is carried inside the MP attribute.validate_update_attributes()requires NEXT_HOP only for an eBGP UPDATE that carries body NLRI (is_ebgp && has_body_nlri).
§4 — MP_UNREACH_NLRI (Type 15)
Wire layout:
AFI (2 bytes) | SAFI (1) | Withdrawn Routes (variable)- Flags: Optional + Non-Transitive (
0x80). As with MP_REACH_NLRI, the Extended Length bit may be added for framing and Partial is invalid. - Withdrawn routes use the same prefix-length encoding as announced NLRI.
- An empty NLRI field after AFI/SAFI is End-of-RIB only when that family was negotiated on the session and is not IPv4 unicast (RFC 4724 §2). IPv4 unicast always uses the minimum-length empty UPDATE marker, including sessions that negotiated Extended Next Hop.
The current typed MP decoder supports IPv4/IPv6 unicast, IPv4/IPv6 FlowSpec, L2VPN EVPN, BGP-LS/BGP-LS VPN, VPNv4/VPNv6, IPv4/IPv6 labeled-unicast, and IPv4 RT-Constrain. This is an implementation boundary, not a claim to support every IANA AFI/SAFI assignment.
AFI/SAFI Negotiation
- MP-BGP capabilities are advertised in OPEN via
Capability::MultiProtocol. intersect_families()computes the intersection of locally configured families (fromPeerConfig.families) and the peer's advertisedMultiProtocolcapabilities. Only negotiated families are processed.- Result stored in
NegotiatedSession.negotiated_families. - If neither side advertises IPv4 unicast MP-BGP capability, IPv4 unicast is still implicitly supported (RFC 4760 §8 backward compat: body NLRI is always IPv4).
IPv6 NLRI Encoding
- Same wire format as IPv4: 1 byte prefix length + ceil(len/8) bytes of address. Maximum prefix length is 128 (vs 32 for IPv4).
- Host bits are masked off on decode (same as
Ipv4Prefix::new()). Ipv6Prefixtype mirrorsIpv4Prefix: public fieldsaddr: Ipv6Addrandlen: u8.
Outbound UPDATE Splitting
- IPv4 routes use body NLRI (WITHDRAWN + NLRI fields in the UPDATE body).
- IPv6 routes use
MP_REACH_NLRI/MP_UNREACH_NLRIin the path attributes with empty body NLRI. - A single UPDATE carries only one address family.
MpReachNlriandMpUnreachNlriare not stored onRoute.attributes— they are per-UPDATE framing, rebuilt on each outbound send.
eBGP NEXT_HOP for IPv6
- eBGP next-hop rewrite:
MpReachNlri.next_hopis set to the local socket's IPv6 address (same pattern as IPv4 eBGP next-hop rewrite). - iBGP: next-hop passed through unchanged.
Interpretation Decisions
Attribute Ordering
RFC 4271 §4.3 states well-known attributes should appear before optional
attributes. rustbgpd accepts out-of-order attributes but emits a
structured warning event. A future strict_attribute_order config option
may reject them, but this is not v1 scope.
RFC 4724 — Graceful Restart Mechanism for BGP
rustbgpd implements the receiving speaker role and a planned-restart restarting speaker role. As a receiver, when a peer that previously advertised the Graceful Restart capability goes down, rustbgpd preserves that peer's routes as stale rather than immediately withdrawing them. The bounded restarting-speaker behavior is described in §4.1 and ADR-0040.
§3 — Graceful Restart Capability
- Capability code 64. Wire format: 2-byte flags/time + N × 4-byte per-family entries.
restart_state(R-bit): indicates the sender has restarted and may have preserved forwarding state. 12-bitrestart_timefield.- Per-family: AFI (2) + SAFI (1) + flags (1). Bit 0x80 =
forwarding_preserved. - The daemon sets the GR Forwarding State bit and LLGR F bit independently of
Restart State, using one committed local-role snapshot per OPEN. Both bits
are set for families with no configured kernel installer, following the
control-plane-only guidance in RFC 9494 §5.
They remain clear for each unicast family selected by
fib_tables, for both unicast families whenhonor_blackholeandinstall_blackhole_discardare enabled, and for EVPN when an L2 instance or IP-VRF is configured. This does not claim kernel-state preservation or restore. - Committed role changes and rollback affect the next OPEN, including an already-running session's reconnect. Staged candidates do not change these bits or restart unchanged peers. Restart-required blackhole settings stay pinned to the running process. The standalone FSM retains its conservative default of advertising both bits clear.
- If runtime FIB ownership becomes uncertain, the next OPEN also keeps F=0 for the union of committed and potentially installed unicast families. This uses the settlement watchdog's terminal fence, including lost acknowledgements and failed compensation; it does not publish a candidate as committed. A late acknowledgement after that fence cannot erase the conservative union. Failed or interrupted EVPN convergence similarly keeps the EVPN bit clear until acknowledged convergence proves its role; validation and no-op requests cannot clear that uncertainty. Other families keep their own role-derived bits.
- If a peer sends multiple GR capabilities (malformed OPEN), only the first is used. A warning is logged.
- Capability decode is bounded to the enclosing optional-parameter slice — a malformed capability length cannot consume beyond the parameter.
§4.1 — Procedures for the Restarting Speaker
Restarting-speaker mode is implemented (ADR-0040). After a coordinated
shutdown, a marker file is written to runtime_state_dir. On startup, if the
marker is present and not expired, static peers from config are offered R=1 in
OPEN. The per-family forwarding bits follow the role rules above, independently
of R. Dynamic gRPC-added peers always get R=0. Before sessions start,
the RIB freezes the resolved static GR peer/family roster. Per-family Loc-RIB
selection and outbound initial table/EoR are held until current-session EoRs
arrive from every eligible waiter or the marker-bounded selection timer
expires. Peer Restart State and absent GR families exclude that peer/family;
superseded-session EoRs are rejected.
When warm_cache_checkpoint_on_shutdown = true, a successful bounded
checkpoint publication binds its generation into the restart marker
(ADR-0104). Linux marker v3 also binds the expiry to a complete boot and time
namespace CLOCK_BOOTTIME domain; wall-only v1/v2 remain compatibility
fallbacks when that domain is unavailable. Checkpoint failure retains a
generationless marker. This does not extend RFC 4724 semantics: startup never
restores or advertises cached routes. Configured kernel installers still
advertise forwarding-state bits clear.
§4.2 — Procedures for the Receiving Speaker
GR trigger: On SessionDown, GR is entered when the peer previously
advertised GR capability (peer_gr_capable) AND local config has
graceful_restart = true. The R-bit is NOT checked — it indicates
restart state in the NEW OPEN after reconnection, not in the dying session.
A peer whose GR capability lists no usable family also enters retention when
it advertised LLGR and llgr_stale_time is non-zero locally (RFC 9494 §4.2,
below).
Family handling: ALL families from the peer's GR capability are retained
as stale (not just those with forwarding_preserved=true). The
forwarding_preserved flag affects forwarding decisions, not route
retention. Routes for negotiated families NOT in the peer's GR capability
are withdrawn immediately.
Stale route demotion: Route.is_stale flag. Best-path step 0 (before
LOCAL_PREF) prefers non-stale over stale. This is more aggressive than the
RFC suggestion (step 7 or later) but matches GoBGP and FRR behavior.
Two-phase timer:
- Initial timer =
restart_time(peer's advertised value). This is the window for the peer to re-establish the TCP session. - On
PeerUpduring GR, timer resets togr_stale_routes_time(local config, default 360s). This is the window for the peer to send End-of-RIB markers.
PeerUp during GR: Routes are NOT cleared of stale flags. The timer is
reset. Outbound state is re-registered. Stale flags are cleared only by
per-family End-of-RIB, not by session re-establishment. The exception is a
retained family that the new OPEN does not list: RFC 4724 §4.2 requires its
stale routes to be removed immediately "if a specific address family is not
included in the newly received Graceful Restart Capability, or if the Graceful
Restart Capability is not received in the re-established session at all", and
rustbgpd removes them at PeerUp. The Forwarding State bit in the new
capability is not checked (see All GR Families Retained below).
End-of-RIB: Clears stale flag for the indicated address family. Recomputes best paths (previously-demoted routes may now win). If all families have received EoR, GR completes and state is cleaned up.
Timer expiry: Remaining stale routes are swept as withdrawals. GR
state is cleaned up. bgp_gr_timer_expired_total metric incremented.
End-of-RIB Detection
- IPv4: empty UPDATE (no NLRI, no withdrawn, no attributes)
- IPv6: UPDATE with only empty
MP_UNREACH_NLRI
End-of-RIB Sending
After sending the initial table to a new peer, EoR markers are sent for
each negotiated family via OutboundRouteUpdate.end_of_rib.
Metrics
bgp_gr_active_peers— gauge, set on GR entry, cleared on completion or timer expirybgp_gr_stale_routes— gauge per peer, updated on GR entry, per-family EoR, and completion/expiry. It counts every route held for retention, GR-stale or LLGR-stale.bgp_gr_timer_expired_total— counter, incremented on timer expiry
Interpretation Decisions (RFC 4724)
Stale Demotion Placement
RFC 4724 suggests demotion "in its decision process" without specifying where. rustbgpd places it at step 0 (before LOCAL_PREF), meaning a stale route always loses to any non-stale alternative regardless of other attributes. This matches GoBGP and FRR and is the safest behavior for a receiving speaker.
RFC 9494 §4.3/§4.4 — LLGR_STALE Routes Are Least Preferred
RFC 9494 §4.3: "A BGP speaker that has advertised the Long-Lived Graceful Restart Capability to a neighbor MUST perform the following upon receiving a route from that neighbor with the LLGR_STALE community or upon attaching the LLGR_STALE community itself per Section 4.2: Treat the route as the least preferred in route selection". §4.4: "A least preferred route MUST be treated as less preferred than any other route that is not also least preferred. When performing route selection between two routes when both are least preferred, normal tiebreaking applies."
Every family's ranker (unicast and each route-reflection family) places a
route in the least-preferred tier when this speaker holds it LLGR-stale or
when it carries the LLGR_STALE community, at the same step 0 as GR stale
demotion (see Stale Demotion Placement above). Two least-preferred routes,
including one tagged on receipt and one tagged locally, compare by the
remaining steps. Best-path explain reports llgr_stale_community when the
losing route is least preferred only because of a received community, and
stale_preference for local GR or LLGR stale state.
The rule applies whether or not this speaker advertised LLGR to the neighbor
the route came from. §4.3 mandates the behavior when LLGR was advertised and
states no requirement otherwise, so applying it unconditionally satisfies the
MUST. A route tagged LLGR_STALE then ranks the same on every LLGR-capable
router in the AS, whatever each router's per-session LLGR configuration,
which is the consistency §5.2 identifies as the defense against forwarding
loops. FRR also ranks on the community without a session check
(bgp_path_info_cmp in bgpd/bgp_route.c). Unlike that implementation, two
least-preferred routes here fall back to normal tie-breaking, as §4.4
requires.
The check reads the route's community list, like the other
attribute-derived steps (LOCAL_PREF, AS_PATH, MED); no per-route field is
added, and the list is read only when the route is not already LLGR-stale.
Unicast Loc-RIB selection ranks each candidate once per recompute, so it
reads each list once rather than on every comparison. The other families'
rankers and single pairwise comparisons, such as best-path explain, read it
on each comparison.
All GR Families Retained
RFC 4724 §4.2: "the receiving speaker MUST retain the routes received from
the restarting speaker for all the address families that were previously
received in the Graceful Restart Capability." The forwarding_preserved
flag does NOT gate route retention — it indicates whether the data plane
was preserved for forwarding decisions.
The same applies on re-establishment. RFC 4724 §4.2 and RFC 9494 §4.2 also require removing a family's stale routes when the Forwarding State bit (GR) or F bit (LLGR) is clear in the newly received capability. rustbgpd does not check either bit and keeps the stale routes until End-of-RIB or the timer. This receiving-side deviation is independent of the role-derived bits in rustbgpd's own outgoing OPEN. Removal on reconnect is limited to families the new OPEN does not list at all.
RFC 9494 §4.2 — LLGR Families Outside the GR Capability
RFC 9494 §4.2: "If the Graceful Restart Capability that was received does not list all AFIs/SAFIs supported by the session, then the GR Restart Time shall be deemed zero for those AFIs/SAFIs that are not listed." §4.1 names "omitting all AFIs/SAFIs from the GR Capability" as the way to skip the GR phase, and only an absent GR capability makes LLGR "disregarded" (§4.1, §4.5).
At session down, a family in the peer's LLGR capability but not in its GR
capability is retained and enters the LLGR phase immediately: its routes
become LLGR-stale, receive the LLGR_STALE community, and are swept at that
family's Long-Lived Stale Time. Families in both capabilities run the GR phase
first. A peer that sends a GR capability with an empty family list plus LLGR
is LLGR-capable when at least one family has a non-zero Long-Lived Stale Time,
and RFC 8538 Notification GR applies to its retention too. A family with both a
zero Restart Time and a zero Long-Lived Stale Time gets no retention ("none of
these procedures would apply"). FRR's default helper-mode OPEN (a GR
capability with no families and an LLGR capability with a zero stale time)
therefore gets no route retention from rustbgpd. Notification GR is negotiated
independently through the two advertised N bits (RFC 8538 §4), so that helper
still receives Hard Reset for protective max-prefix and BFD teardown. The
Hard Reset encapsulates the original Cease reason and its data, preventing
the helper from retaining rustbgpd's routes after protective teardown.
An LLGR capability without any GR capability is ignored.
On re-establishment, a family already in the LLGR phase whose tuple the new
LLGR capability omits, or whose new OPEN has no LLGR capability (or no GR
capability, which makes LLGR disregarded), has its stale routes removed at
PeerUp (RFC 9494 §4.2: "a specific address family is not included in the
newly received LLGR Capability, or the LLGR and accompanying GR Capability are
not received in the re-established session at all"). The F-bit clause is not
implemented; see All GR Families Retained above. If End-of-RIB is still missing
when the re-armed GR timer expires, promotion to LLGR follows the new OPEN's
LLGR capability and Long-Lived Stale Times, not the previous session's. A
family whose Long-Lived Stale Time is zero is purged at the end of its GR
phase instead of being promoted.
RFC 9494 §5 — Per-AFI/SAFI Configuration
RFC 9494 §5: "Implementations MUST NOT enable these procedures by default.
They MUST require affirmative configuration per AFI/SAFI in order to enable
them." LLGR is disabled by default (llgr_stale_time = 0), but enabling it is
per neighbor: a non-zero llgr_stale_time advertises LLGR for every family
the neighbor negotiates GR for. This is a deviation from the per-AFI/SAFI
requirement.
gr_stale_routes_time Cap
gr_stale_routes_time is capped at 3600 seconds (1 hour). This is an
implementation safety limit, not an RFC constraint. A misconfigured value
should not keep stale routes for days.
Kernel Forwarding-State Preservation
Restarting-speaker support does not preserve or adopt kernel forwarding state. Configured kernel installers continue to advertise F=0; a verified restore/adoption design would be needed to change that claim. The R-bit marker behavior from ADR-0040 and ADR-0104's publication-only checkpoint remain separate from control-plane-only families advertising F=1.
RFC 2918 — Route Refresh Capability
- Capability code 2, unconditionally advertised.
- Inbound: on receiving ROUTE-REFRESH, re-advertise the requested family from Adj-RIB-Out.
- Operator inbound refresh:
SoftResetInsends ROUTE-REFRESH to the peer. - Operator outbound refresh:
RefreshOutboundre-emits rustbgpd's current exportable outbound inventory to one peer; it does not send ROUTE-REFRESH. - See ADR-0027.
RFC 5291 / RFC 5292 — Outbound Route Filtering, Address-Prefix ORF
- Capability code 3, Address-Prefix ORF-Type 64.
- rustbgpd implements the receive side: per-neighbor / peer-group
prefix_orf_receive = trueadvertises willingness to receive Address-Prefix ORF entries and applies the peer-pushed filter before export policy. - ORF filters use prefix-list semantics: sequence order, first match wins, implicit deny on a non-empty list, permit-all when empty.
- The initial advertisement for an ORF-negotiated family is gated until the
peer's first ROUTE-REFRESH;
DEFERinstalls state and waits for a later immediate or plain refresh to sweep advertisements and withdrawals. - Unknown
When-to-refreshvalues are invalid control input: rustbgpd resets the negotiated Address-Prefix ORF list for that family/type and forces a safe resync instead of treating the value as defer-like state. - A well-formed entry whose maximum length is below its own prefix length
(for example
10.0.0.0/8 le 4) can never match; it is installed as sent and reported with awarnlog line naming the peer, family, and window, because rejecting it would flush the list and fail open to permit-all. - Address-Prefix ORF entries are decoded only for IPv4/IPv6 unicast. L2VPN and unknown future SAFIs are preserved as raw ORF groups until their family-specific prefix encodings and export semantics are implemented.
- See ADR-0075.
RFC 7313 — Enhanced Route Refresh
- Capability code 70, unconditionally advertised.
- BoRR/EoRR markers demarcate the refresh window.
- Inbound BoRR marks existing routes as refresh-stale; EoRR sweeps unreplaced routes. 5-minute timeout on the refresh window.
- Outbound: Enhanced peers get BoRR → routes → EoRR; legacy peers get routes → EoR.
- A GR restarter's initial flood remains unmarked until its first EoR, including ORF-deferred and backpressured initial dumps. Later refreshes use the normal BoRR/EoRR brackets.
- For peers advertising Enhanced Route Refresh, subtype 1/2 bodies whose length is not four bytes close the session with ROUTE-REFRESH Message Error / Invalid Message Length (7/1). Data contains the complete received PDU, including the BGP header, when it fits the peer's receive limit. This permits offending PDUs up to 4075 bytes, or 65514 bytes when the peer advertised Extended Messages. Larger PDUs produce empty Data and a log of the received length: RFC 7313 section 5's full-PDU requirement cannot be met within RFC 8654 section 4's message-size limit in this exceptional case. Data is never truncated and the peer's receive limit is never exceeded; our own Extended Messages advertisement does not increase that limit. Unknown identifiable subtypes are ignored before ORF decoding. A body too short to contain a subtype retains the generic framing error; subtype 0 and peers without ERR retain their prior rules.
- Joint behavior with GR/LLGR retention: routes flagged
GR-stale or LLGR-stale are NOT snapshotted at BoRR, so EoRR (or the
window timeout) never purges them. A restarting peer's refresh replay
is not authoritative while it is still converging — RFC 4724 §4.1
retains stale paths until End-of-RIB or the restart timer, RFC 9494
§4.2 until the LLGR timer; those remain the only removal points. A
re-advertisement inside the window still clears both staleness kinds
via implicit replace. The reverse ordering (GR entry during an open
window) is a session-down, which drops all refresh windows for the
peer. Combination matrix in
handle_begin_route_refresh. - See ADR-0038.
RFC 4360 — Extended Communities
- Type code 16. Two-octet AS (subtypes 0x02 RT, 0x03 RO) and four-octet AS (subtypes 0x02 RT, 0x03 RO) encodings.
- Policy matching uses logical RT/RO equivalence across encodings.
- See ADR-0025, ADR-0026.
RFC 8092 — Large Communities
- Type code 32. 12 bytes: Global Administrator (4) + Local Data Part 1 (4) + Local Data Part 2 (4).
- Zero-length Large Communities attribute rejected at wire decode.
- Policy:
LC:G:L1:L2format inmatch_community,set_community_add,set_community_remove. - See ADR-0031.
RFC 8654 — Extended Message Support
- Capability code 6, unconditionally advertised.
- When both peers advertise, max message length is raised from 4096 to 65535 bytes for that session. Resets to 4096 on session-down.
ReadBuffer.set_max_message_len()dynamically resizes on negotiation.- See ADR-0032.
RFC 9072 — Extended OPEN Optional Parameters Length
- Outbound OPEN messages retain the RFC 4271 format while the complete Optional Parameters field is at most 255 octets. Larger fields use Optional Parameter type 255 as the extended-length marker, a 16-bit aggregate length, and 16-bit lengths for each enclosed Optional Parameter.
- Extended-format OPEN messages are accepted even when their aggregate length is 255 or less, including zero. The non-extended length octet is ignored after the type-255 marker is recognized, as required by RFC 9072 §2.
- The Capabilities Optional Parameter remains type 2, and individual capability TLVs retain their 8-bit value lengths. Unknown capabilities are preserved; an unknown Optional Parameter type is rejected with OPEN Message Error / Unsupported Optional Parameter (2/4) and empty Notification Data.
- RFC 8654 does not extend OPEN: both classic and RFC 9072 encodings remain subject to the 4096-byte message maximum.
RFC 7911 — Add-Path
Experimental Paths-Limit
Capability code 76 implements the tuple format from the expired
draft-abraitis-idr-addpath-paths-limit-04. Each AFI/SAFI receiver preference
is applied only to the matching negotiated Add-Path send direction. This is an
experimental interoperability feature, not an adopted IETF standard.
rustbgpd caps what it sends at the peer's advertised limit. It also applies
its own nonzero advertised limit locally to retained path identities per
prefix for negotiated IPv4/IPv6-unicast Add-Path receive; other families have
no local receive enforcement. Rejected identities count only when the
family's received-prefix bound already enables their tracking. This local
cap is defensive; the draft places the sender-side obligation on the peer.
If this local cap shuts down a session, Cease/1 carries no optional RFC 4486
data: that field specifies a prefix upper bound, not paths per prefix. The
configured path cap remains visible in the local warning, counter, and peer
latch reason.
Neighbor output orders rows by numeric AFI/SAFI and carries an optional
normalized limit whose presence distinguishes active unlimited from inactive.
The raw effective_send_max sentinel is gone: PathsLimitState field number
and name 5 are reserved, and optional effective_send_limit (field 6) alone
carries inactive (absent), unlimited (zero), or finite.
Base Add-Path
- Capability code 69. Per-family Send/Receive/Both modes.
- Adj-RIB-In/Out keyed by
(Prefix, u32)for multi-path storage. - Multi-path send: rank-based path IDs (best=1, second=2, ...).
send_maxcaps paths per prefix per peer.- Both IPv4 body NLRI and IPv6 MP_REACH/MP_UNREACH supported.
- A received path identifier carries no preference (§2). Unicast, VPN, and labeled-unicast selection compares it only as the last step, after the peer address, so otherwise-equal routes from one peer rank the same whatever order they arrived in.
- See ADR-0033.
RFC 8950 — Extended Next Hop
- Capability code 5. IPv4-unicast IPv6-next-hop receive support is advertised
automatically when both
ipv4_unicastandipv6_unicastare configured. VPNv4 IPv6-next-hop receive support (tuple 1/128/2) is advertised wheneverl3vpn_ipv4_unicastis configured, independently of those unicast families. - Negotiation: exact 6-byte tuple matching (NLRI AFI, NLRI SAFI, NH AFI).
- When negotiated, an IPv4-unicast route whose exported next hop is IPv6 is
sent in
MP_REACH_NLRI(16 or 32 octets, §3). A route whose exported next hop is IPv4, and every IPv4-unicast withdrawal, stays in the classic UPDATE body (NEXT_HOPplus NLRI, and Withdrawn Routes). §3 keeps that encoding as the existing mode of operation, and it is what OpenBGPD sends. OpenBGPD 9.2 resets the session with UPDATE Message Error / Optional Attribute Error (3/9) on an IPv4-unicastMP_REACH_NLRIwith a 4-octet next hop and on any IPv4-unicastMP_UNREACH_NLRI, even with Extended Next Hop negotiated (rde_get_mp_nexthop,rde_update_dispatch). Earlier rustbgpd releases sent both forms. FRR 10.3.1 source sends every IPv4 withdrawal inMP_UNREACH_NLRIonce Extended Next Hop is negotiated, which the same OpenBGPD check rejects. A scoped link-local (unnumbered) session keeps the MP form for all IPv4 routes and withdrawals, because receivers there ignore IPv4 body NLRI. - IPv4-unicast reflection with an unchanged IPv6 next hop requires the recipient's Extended Next Hop receive capability (§5). Without it, export is suppressed rather than sending classic IPv4 NLRI without NEXT_HOP. Ordinary eBGP, next-hop-self, and explicit IPv4 export-policy rewrites remain eligible when they supply a classic IPv4 NEXT_HOP.
- VPNv4 reflection preserves the 24- or 48-octet next-hop encoding (§3/§5) and requires the recipient's VPNv4 IPv6-next-hop receive capability. A rejected replacement withdraws any previously advertised route; ordinary IPv4-next-hop VPNv4 routes and VPN withdrawals remain eligible without it. The peer's receive capability is not an inbound admission requirement.
- Inbound, IPv4 unicast may use
MP_REACH_NLRI/MP_UNREACH_NLRIon any session that negotiated the family. Extended Next Hop governs only the next-hop encoding, not whether the AFI/SAFI is allowed (§4). A 4-octet IPv4 next hop is accepted without the capability, and IPv4MP_UNREACH_NLRIwithdrawals are always applied. Earlier releases ignored both forms when Extended Next Hop was not negotiated, so a withdrawal sent that way left the route in place. - A 16- or 32-octet (IPv6) next hop on IPv4-unicast NLRI without negotiated
Extended Next Hop is a malformed
MP_REACH_NLRI. RFC 7606 §7.11 judges the next-hop length against the one expected for the AFI/SAFI as modified by the extensions in use, and gives RFC 5549 (now RFC 8950) as the example: only when it is in use may IPv4 unicast carry a 16-octet next hop. A mismatch requires session reset or AFI/SAFI disable. rustbgpd does not implement AFI/SAFI disable, so the session resets with UPDATE Message Error / Optional Attribute Error (3/9) carrying theMP_REACH_NLRIattribute as received, the same NOTIFICATION as other malformedMP_REACH_NLRInext hops. The reset is counted inbgp_update_malformed_total{disposition="session_reset"}. That the decoder could still locate the NLRI does not change the disposition. Earlier releases dropped such an UPDATE with only a log line. - Receiving behavior of other implementations for that case, from their source: ExaBGP 5.0.13 and OpenBGPD 9.2 also reset the session, BIRD 3.3.2 discards the routes, and FRR 10.7.1 and GoBGP v4.9.0 accept them. GoBGP and ExaBGP send IPv4 routes with an IPv6 next hop without checking the capability, so a peer misconfigured that way against a rustbgpd neighbor that does not advertise Extended Next Hop has its session reset.
- See ADR-0037 for the IPv4-unicast behavior.
RFC 6811 — RPKI Origin Validation + RFC 8210 — RTR
- VRP table with sorted-Vec binary search for prefix containment.
Arc<VrpTable>snapshot pattern for lock-free reads.- RTR codec: RFC 8210 v1 plus the 8210bis v2 ASPA PDU. Serial/Reset queries, Serial Notify, expire enforcement. Router Key PDUs (BGPsec) are not implemented.
- RTR cache sockets take the same transport authentication as BGP neighbors:
TCP MD5 (RFC 2385,
md5_password) or a TCP-AO keyring (RFC 5925,tcp_ao) per[[rpki.cache_servers]], mutually exclusive, installed before connect and preflighted against the kernel at startup so a refused key is a startup error rather than a reconnect loop. RTR over TLS or SSH is not implemented. - Cache state is one per-cache epoch
(version, session ID, serial), advanced only at a validated End of Data. Identity mismatches (session ID, RFC 1982 serial regression) force a Reset Query resync — never a splice. Validated data is retained through reconnect/Cache Reset until replaced or expired. Transactions are bounded by deadline and record/byte budgets. - ASPA over RTR v2 uses 8210bis replacement semantics: announce replaces the customer's provider set; withdraw removes the customer ASN.
- Strict acceptance limits (8210bis-27): per-PDU length is capped at 65,535 octets (§5 — an over-limit length field is corrupt framing, Error Report code 0); negotiated-v2 IPv4 and IPv6 Prefix PDUs must carry canonical network addresses (nonzero host bits draw code 0 with the offending frame, close the session, flush that cache's held data, and publish none of the incomplete transaction; v1 still accepts nonzero host bits, while invalid prefix-length and max-length PDUs on either version share the fatal code-0 flush disposition); End of Data timers are bounded to the §6 legal ranges (zeros mean "not provided"; above-maximum values clamp down with a warning; an expire below the 600 s minimum is honored as-is, since expiring early is safe; a refresh/retry not below the expire is lowered under it per the §6 relationship rule); ASPA PDUs must be well-shaped per §5.12 (announce: at least one provider, strictly increasing, no AS 0 among multiple; withdraw: no provider list, PDU length exactly 12) — violations get Error Report code 9.
- Deviation: Duplicate Announcement Received (Error Code 7) and Withdrawal of Unknown Record (Error Code 6) are not detected. The RTR client does not hold the per-cache active record set during a transaction — the VRP manager applies updates by normalizing announce/withdraw merges — so detecting either would require a parallel active-record index in the client. Duplicate announcements and withdrawals of unknown records are normalized silently instead of failing the session.
- Best-path step 0.5: Valid > NotFound > Invalid (between stale demotion and LOCAL_PREF).
match_rpki_validationin policy.- draft-ietf-sidrops-avoid-rpki-state-in-bgp-12 (BCP, RFC Editor queue):
the daemon does not automatically encode RPKI- or ASPA-derived validation
state in any Path Attribute. Validation state lives in the RIB, policy predicates
(
match_rpki_validation,match_aspa_validation,route.rpki/route.aspa), and the API/CLI. Explicit operator policy can attach the RFC 8097OV_*extended-community aliases; type0x43is non-transitive, so ordinary eBGP export strips it under the non-transitive rule in "Extended Communities — non-transitive eBGP export" below unlesssend_non_transitive_extended_communities = true, while route-server-client export preserves it (RFC 7947 §2.2.4 transparency). The hand-written route-server example'shygiene.rpoladds noOV_*community, and neither doesrs-config-renderin either mode: IXP Manager mode embeds that file, and arouteserver mode filters on origin validation without tagging the state. ARouteServer's BIRD and OpenBGPD output tags RFC 8097 state internally, but both daemons strip non-transitive extended communities on eBGP export, so its clients never receive it either. An operator policy that tags routes reaching another AS contradicts the draft's §6: "Operators MUST NOT signal RPKI-derived validation states using BGP Path Attributes carried over EBGP sessions across administrative boundaries." rpol can still addOV_*. - See ADR-0034.
RFC 8955/8956 — FlowSpec
- SAFI 133. IPv4 and IPv6 unicast FlowSpec.
- 13 match component types (destination/source prefix, protocol, ports, ICMP, TCP flags, packet length, DSCP, fragment, flow label).
- Actions via extended communities: traffic-rate, traffic-action, traffic-marking, redirect.
- Typed byte- and packet-rate actions interpret and construct negative rates as zero (RFC 8955 §7.1 and §7.2). As a local choice, NaN and negative zero also become positive zero; positive rates, including positive infinity, are preserved. This affects typed API views and injection only; raw extended-community storage and reflection are unchanged.
- NH length = 0 in MP_REACH_NLRI for FlowSpec.
- An IPv6 destination component with a non-zero offset (RFC 8956 §3.1) carries no destination prefix for policy, RPKI, or validation purposes: the pattern is the address shifted right by the offset, so it names no routable prefix, and RFC 8956 §5 counts only an offset-0 destination toward validation item (a). The rule remains decoded and retained after admission. A received rule with this offset is ineligible for selection when cross-RIB validation is enabled.
- Startup-only
[flowspec] validation = "rfc9117"enables feasibility checks against the selected unicast cover and all admitted, retained more-specific unicast candidates. The default is"off". Relevant unicast changes trigger resumable revalidation, including losing Add-Path changes; infeasible rules remain available in the received-peer diagnostic view. Local injection stays trusted origination. This adds no FlowSpec forwarding dataplane. - See ADR-0035 and the superseding validation decision in ADR-0135.
RFC 7854 — BMP
- BMP exporter (router-initiated). All 6 message types encoded.
- Per-collector TCP client with reconnect/backoff.
- Peer Up replay on collector reconnect.
- Periodic Stats Report (type 7: Adj-RIB-In route count, 60s interval).
- Coordinated Termination on daemon shutdown.
- RFC 9736 (Peer Up message namespace) is satisfied by construction: Peer
Up stays message type 3, the ordinary Peer Up carries no Information
TLVs, and the RFC 9069 Loc-RIB Peer Up carries only the VRF/Table Name
TLV (type 3), which keeps that value in the new Peer Up registry; no
Initiation-only TLV type is emitted in a Peer Up (
encode_peer_up/encode_loc_rib_peer_upincrates/bmp/src/codec.rs). - Raw UPDATE PDU capture via
Bytesrefcount clone (zero overhead when unconfigured). - Per-collector view selection:
monitor = ["rib_in_pre", "rib_out_post", "loc_rib"](default["rib_in_pre"]). The full monitoring trio (7854 + 8671 + 9069) ships on one exporter; M81 receipt. - See ADR-0041 and ADR-0097.
RFC 8671 — BMP Adj-RIB-Out Monitoring
- Post-policy Adj-RIB-Out Route Monitoring, tapped at the transport's
outbound byte funnel (
enqueue_bulk): the BMP message wraps the exact PDU sent on the wire, after transport stamping (ORIGINATOR_ID/CLUSTER_LIST, GShut, LLGR §4.6) — byte-exact by construction, test-pinned for VPN and EVPN. - O-flag 0x10 on Route Monitoring; L=post-policy under O=1; peer identity stays the remote peer's; PeerUp/Down/Stats remain O=0.
- Stats types 15 and 17 (post-policy Adj-RIB-Out counts) from one batched AdjRibOut query per stats tick; unavailable counts are omitted, never a false zero.
- Live-only (no rib-out table dump): AdjRibOut stores routes pre-transport-stamping, so a synthesized dump would not be byte-faithful. Pre-policy rib-out deliberately skipped.
- Saturation semantics: "mirror what was actually sent" is enforced in
both directions. An UPDATE that fails to enqueue on the saturated
writer is never mirrored (it never reached the wire; the
Cease/8teardown ends in a reliably delivered, correctly ordered Peer Down and the re-established session re-floods the view). Conversely, a mirror event dropped on a full BMP channel after the wire send succeeded forces a synthetic Peer Down/Peer Up peer-state reset on the stream (reason 2, FSM code 0) so the divergence is collector-detectable — live-only views can never silently under-report. Per-peer Peer Up/Peer Down delivery to the BMP manager is reliable (never try_send-dropped). - See ADR-0097 (Decisions 1, 3 incl. the saturation amendment).
RFC 9069 — BMP Local RIB Monitoring
- Loc-RIB Route Monitoring synthesized at the RIB recompute commit
seams (
crates/rib/src/bmp_sync.rs); announcement PDUs rebuild MP_REACH the same way the MRT exporter does. - Emulated instance peer per §5.2.1: peer type 3, peer flags forced to
zero (its own registry — never V, even for v6), fabricated sent-OPEN
advertising exactly the streamed capability set, VRF/Table Name
"global", Peer Down reason 6, stats types 8/10. - V1 family scope: IPv4/IPv6 unicast + VPNv4/VPNv6; the fabricated OPEN only promises families that actually stream.
- Collector-connect table sync: chunked non-blocking dump (256-message chunks drained by a per-collector forwarder task with a send timeout) → one End-of-RIB per family → live. Dump/live overlap is the standard BMP initial-sync race, accepted.
- See ADR-0097 (Decisions 1, 2).
RFC 9972 — Advanced BMP Statistics Types
- The periodic peer Stats Report emits the selected RFC 9972 statistics types: post-policy Adj-RIB-In types 20/21/23, inbound policy-rejection type 22, and RPKI Invalid/Valid/NotFound types 35/36/37.
- The family-qualified rows cover negotiated IPv4/IPv6 unicast. Types 20/21/23 are omitted under effective Add-Path receive; type 22 requires authoritative reject retention; types 35/36/37 require an authoritative VRP table. Unavailable observations are omitted rather than reported as zero.
- See the BMP configuration contract for the emission conditions and the M81 interop entry for the collector checks. Other RFC 9972 statistics types are outside this emitted set.
draft-ietf-grow-bmp-tlv-21 — BMPv4 TLV framing (pre-IANA)
- Per-collector
version = 3 | 4, default 3; v3 output byte-identical to prior releases (golden-bytes pinned). - v4: common-header version 4 on every message; Route Monitoring wraps the UPDATE in the mandatory BGP Message TLV (type 4, index 0, §5.2); Stats Reports wrap in the Stats TLV (code 1, §5.4); the TLV-provisioned message types change only their version byte.
- Indexed-TLV (§4.3: 2-byte index after length, excluded from the length value, G-bit) and Group TLV (type 1) encoding infrastructure.
- All draft code points live in
crates/bmp/src/tlv.rswith a renumber note — an IANA renumber at RFC publication is a single-file change. The draft's Appendix A contradicts its normative §5.2.1 on the Group TLV type; we follow the normative text. - See ADR-0097 (Decision 4).
draft-ietf-grow-bmp-path-marking-tlv-05 — Path Marking (pre-IANA)
- Temporarily unavailable: path-marking-05 self-assigns RM TLV type 5, while tlv-21 assigns type 5 to Sequence Number. No ambiguous type-5 Path Marking is emitted; emission remains pending a non-colliding assignment.
- Status bits limited to what an RR can attest: Best + Stale (from the GR/LLGR machinery). FIB/damping/filter bits are never fabricated.
- Reason Code (§3.2) on live unicast announces: re-derived
best-vs-runner-up through the
best_path_cmp_with_reasonexplain ladder over the same candidate pool the recompute used; sole candidate → no reason; decisive steps without a registered draft code are omitted rather than mislabeled. VPN and dump entries carry bits only. - The internal status payload remains available so emission can resume without changing the RIB seam once the drafts converge.
- See ADR-0097 (Decision 5).
RFC 9003 — Extended Admin Shutdown Communication (obsoletes RFC 8203)
- RFC 9003 §2
permits an optional UTF-8 reason with Cease subcode 2 (Administrative
Shutdown) or 4 (Administrative Reset). The gRPC
DisableNeighborandResetNeighborreasons reach the NOTIFICATION data field through transport. - The sender deliberately retains a 128-octet interoperability cap, truncating at a UTF-8 character boundary. This follows the recommendation for peers whose extended support is unknown in §3; the receiver accepts the full 255-octet RFC 9003 limit. The send cap is not the RFC's maximum receive length.
- Receive-side extraction handles both direct administrative subcodes and one RFC 8538 Hard Reset envelope encapsulating either subcode. Other inner errors, nested envelopes, and malformed shutdown strings are not interpreted as a shutdown reason.
RFC 9384 — BFD Down Cease subcode
- A BFD-driven teardown sends NOTIFICATION Cease (6) with subcode 10 — BFD Down — in place of the Administrative Shutdown (subcode 2) rustbgpd previously reused for this cause. The data field is empty: RFC 9384 defines no reason string for the subcode, so the RFC 9003 shutdown communication is neither attached on send nor extracted on receipt.
- Only a genuine transition triggers it. BFD reporting Down or AdminDown for a session that was already held down — a level re-report or resync ack — is release-only and never tears BGP down. A freshly (re)started BFD session legitimately starts Down, so treating a level report as an event would flap every non-strict peer at startup. Only a real Up→Down transition reaches the FSM as a BFD Down event.
- Handshake states, not just Established. BFD Down in OpenSent or
OpenConfirm also sends the typed Cease, and uniquely among failed
handshakes it emits
SessionDownso the shared transport teardown path clears cached notification state (§8 above). In Connect and Active there is no peer to notify, so the session goes silently to Idle. The NOTIFICATION is cached beforeSessionDownis dispatched, so BMP Peer Down and the operator event history both carry the BFD reason. - RFC 8538 interaction. When the peer negotiated the Notification GR N-bit, transport rewrites the outgoing NOTIFICATION as Cease/9 (Hard Reset, RFC 8538 §4) whose data field encapsulates the Cease code and the BFD Down subcode. The reason survives the envelope, and the Hard Reset correctly bypasses the receiver's Notification GR handling: a link BFD has declared dead must not leave the peer holding our routes as stale. The encapsulated tuple is not an administrative one, so no shutdown communication is read out of it either.
- Operator impact: automation that classified BFD-driven teardowns by Cease subcode 2 must re-classify. From v0.67.0 the observed subcode is 10, or 9 wrapping it when Notification GR was negotiated.
RFC 9687 — Send Hold Timer
- A peer that stops draining its TCP socket can no longer wedge a
session forever: each
write_all + flushin the per-peer writer task is bounded by the configuredSendHoldTime(crates/transport/src/session/writer.rs). - Detection shape (documented deviation from §4.3's letter): the RFC models a free-running timer restarted on every sent message; we run a per-write deadline that only ticks while a write is pending. The trigger condition is equivalent — a wedged peer stalls the pending write, which times out — and the variant cannot fire on an idle session, so the §4.3 "stop when negotiated HoldTime is zero" rule is unnecessary and protection stays active with keepalives disabled. FRR's SendQ-progress check is the same shape.
- Expiry actions (§4.3, Event 29 §4.2): teardown reuses the
TCP-failure path — session down, TCP close,
ConnectRetryCounterincrement, transition to Idle — without sending a NOTIFICATION. §4.3 makes the NOTIFICATION optional ("if … doing so will not delay" the teardown); the socket is by definition not draining, and a cancelledwrite_allmay have left a partial PDU on the wire, so injecting one could corrupt framing. FRR/OpenBGPd attempt a best-effort code-8 NOTIFICATION here; we deliberately do not. - Local reporting (§4.3 required log, §5/§9 error code):
warnlog with peer + duration,bgp_send_hold_expirations_total{peer}counter, an operator event-history record carrying error code 8 / subcode 0 ("Send Hold Timer Expired"), and a BMP Peer Down reason 2 (local close, no NOTIFICATION) with FSM event code 29 (SendHoldTimer_Expires).NotificationCode::SendHoldTimerExpired(8) decodes/logs if a peer ever sends it to us. - Default (§6): enabled by default at
max(480, 2 × configured hold_time)seconds. §6 recommends max(8 min, 2 × negotiated hold time); the negotiated value is unknowable at config time but never exceeds the configured one, so the derived default is always ≥ the RFC's recommendation and always satisfies the §4.4SendHoldTime > HoldTimeMUST. Precedent: FRR uses 2 × hold time (not configurable); OpenBGPd uses max(negotiated hold time, 90 s) (not configurable). Per-neighbor + peer-groupsend_hold_timeknob (§6 MAY); 0 disables; non-zero values ≤ the effective hold time are rejected at config load (§4.4). Also settable over gRPC (AddNeighbor/PeerGroupDefinition, same validation) and viarbgp neighbor <addr> add --send-hold-time.
RFC 7432 — EVPN (Phase 1: Route Reflector + Phase 2: Bidirectional VTEP + Phase 3: Multi-homing + Phase 4: IRB foundation)
- AFI 25 (L2VPN) / SAFI 70 (EVPN). Enum variants added to
AfiandSafi; capability negotiation works automatically. - Wire codec for all 5 RFC 7432 route types:
- Type 1 EAD (per-ES when
ethernet_tag == MAX_ET (0xFFFFFFFF), per-EVI otherwise). DistinctEvpnRouteKeyvariants prevent semantic collapse. - Type 2 MAC/IP Advertisement. IP Addr Length is in bits (0 / 32 / 128). Label2 is optional — either 0 or 3 trailing bytes after the primary label.
- Type 3 IMET. IP length is in bits (32 / 128).
- Type 4 ES. IP length in bits.
- Type 5 IP Prefix (RFC 9136). Fixed total length disambiguates IPv4 (34 bytes) from IPv6 (58 bytes) — prefix-length byte alone cannot distinguish since 32 is valid for both.
- Type 1 EAD (per-ES when
- Route Distinguisher (RFC 4364) displays as
<asn16>:<u32>(Type 0),<ipv4>:<u16>(Type 1),<asn32>:<u16>(Type 2). Unknown RD types fall back to hex. EvpnRoutecarries full wire payload (for reflection);EvpnRouteKeyis the hashable identity used as RIB key.- Best-path §15.1: Type 2 routes run a MAC Mobility head (sticky
preserved against displacement by non-sticky; higher sequence wins)
before the standard BGP preference chain. Absence of the MAC Mobility
community →
(sticky=false, seq=0)per §7.7. - Route reflection: RFC 4456 rules (
ORIGINATOR_ID,CLUSTER_LIST, split-horizon) reuse the existing unicastshould_suppress_ibgp_innervia a syntheticRouteprobe — no EVPN-specific reflection logic. Split horizon is keyed on the source peer, not the route's next-hop, so a reflector with one client behind a NAT or a different loopback still suppresses correctly. AS_PATH and RR cluster-loop branches emit a proper EVPN withdrawal toward the looping peer (rather than silently dropping the route in the Adj-RIB-Out), so a client that previously received the route observes a clean retract. - Best-path tie-break: the EVPN best-path chain runs the same ordering as the unicast decision process after the MAC Mobility head — stale flag → LOCAL_PREF → AS_PATH length → ORIGIN → MED → eBGP over iBGP → lowest effective BGP Identifier → shortest CLUSTER_LIST → lowest peer address — so a reflector with multiple equal-AS paths converges deterministically. The identifier step compares ORIGINATOR_ID when present and otherwise the advertising peer's BGP Identifier (RFC 4456 §9), the same substitution the unicast, VPN, labeled-unicast, FlowSpec, BGP-LS, and RT-Constrain chains apply (see the RFC 4271 §9.1.2.2 notes under Milestone 1). Every family uses this order, and every family ranks a locally originated route ahead of every received route at the identifier step: a locally originated VTEP route carries an injection placeholder, not a BGP Identifier.
- Initial dump on session up: when an iBGP EVPN session reaches Established, the existing Adj-RIB-In is replayed to the new peer through the same Adj-RIB-Out path that handles steady-state reflection — no separate "fast-path" code that could skip RFC 4456 attribute attachment. EoR is emitted per family after the dump.
- Enhanced Route Refresh tracking (RFC 7313):
refresh_stale_evpnrecords EVPN keys present in Adj-RIB-In at BoRR time; any key not re-advertised before EoRR is withdrawn at sweep, mirroring the unicastrefresh_stalepath. - Max-prefix accounting counts EVPN keys alongside unicast
prefixes in the per-peer prefix counter, so
max_prefixestriggers Cease/1 (Maximum Number of Prefixes Reached) when a misbehaving VTEP floods Type 2 routes. - Policy context: EVPN routes expose attributes, communities, RTs,
and route type to import and export policy. Type 5 supplies its actual
IP prefix in
RouteContext; Types 1–4 useprefix: None, so prefix predicates do not match them, including a0.0.0.0/0prefix-list clause. Use RT, community, or EVPN route-type predicates for those routes. - Route-target retention: the reflector does not require a received
route's RTs to match a local VNI or VRF. A reflector with no local EVPN
instances retains eligible routes and reflects selected paths, subject
to policy and normal export checks. RT matching for local dataplane
import is a separate operation; no
retain route-target allknob is needed. - Next-hop preservation: EVPN export preserves the VTEP next hop by default, including over eBGP. An export policy specifying a next-hop address can change it; no next-hop-unchanged knob is required.
- GR / LLGR stale handling (RFC 4724 + RFC 9494, Gate 2): EVPN
routes participate in the stale-route pipeline alongside unicast
and FlowSpec. On
PeerGracefulRestart,mark_stale_evpn((L2Vpn, Evpn))flags routes; on GR timer expiry with LLGR-negotiated,promote_to_llgr_stale_evpninjectsCOMMUNITY_LLGR_STALEviaArc::make_mutand records the route key inevpn_llgr_stale_local_tagssoclear_stale_evpn/clear_llgr_stale_evpnon EoR later strip only the locally-injected communities (peer-originated ones are preserved). Routes carryingCOMMUNITY_NO_LLGRare dropped on GR expiry rather than promoted, per RFC 9494 §4.7. Enhanced Route Refresh (RFC 7313) tracks unreplaced EVPN keys inrefresh_stale_evpnand withdraws them on BoRR/EoRR completion. - Type 2 MAC/IP Advertisement interop: validated end-to-end
against FRR 10.7.1 via the M30 containerlab suite
(
tests/interop/m30-evpn-type2-frr.clab.yml). Real kernel VXLAN + bridge per VTEP; MAC injection on one VTEP viabridge fdb addpropagates through the rustbgpd RR to the second VTEP and appears in its EVPN MAC table. Assertions cover RFC 4456ORIGINATOR_IDCLUSTER_LIST, next-hop preservation (VTEP loopback, not RR), VXLAN encap community surfaced through gRPC, and withdrawal propagation on FDB delete.
- MAC Mobility + sticky-MAC preservation interop (RFC 7432 §15.1,
§7.7): validated via the M31 4-node harness
(
tests/interop/m31-evpn-mac-mobility-frr.clab.yml). MAC moved between two originating VTEPs through the RR increments the Mobility sequence on the reflected Type 2 and flips the observing VTEP's best path. Sticky MAC on the first VTEP is not displaced by a non-sticky advertisement from the second VTEP. - Scale validation (Gate 5, M33, 2026-04-24): the RR sustains
50,000 Type 2 MAC/IP routes reflected from two originating peers
to a third observer, followed by 60 s of 1,000 rps withdraw +
re-advertise churn, with no route loss and no session flap. The
load generator is the in-tree
bench/evpn-loadcrate, built directly onrustbgpd-wire— no third-party daemon sits in the measurement path. Seetests/interop/m33-evpn-scale.clab.ymlanddocs/benchmarks.md§ "EVPN RR Scale (M33)". - Controller-driven injection (Gate 6, 2026-04-24): Type 2
MAC/IP and Type 3 IMET routes can be injected via gRPC
(
InjectionService::AddEvpnRoute) and withdrawn viaDeleteEvpnRoute. The service accepts display-form RDs (65000:100,10.0.0.1:100,4200000000:100), parses MAC addresses and host IPs, and assembles anEvpnRibRoutewithRouteOrigin::Localthat flows through the same reflection pipeline as iBGP-learned routes.rbgp evpn add-mac-ip / add-imet / delete-mac-ip / delete-imetCLI subcommands cover the operator-facing surface. Type 5 IP-Prefix injection was deferred at this gate pending use-case signal; it shipped later (v0.25.0, M45). Native Type 1/4 multi-homing origination ships through[[ethernet_segments]]; controller injection for those route types is not exposed. - Multi-homing Type 1 EAD + Type 4 ES reflection interop (RFC 7432 §8):
validated via the M32 4-node harness
(
tests/interop/m32-evpn-multihome-frr.clab.yml). Two FRR VTEPs share an Ethernet Segment on a bond ES interface (samees-id+es-sys-mac→ identical 10-byte ESI); both originate Type 4 ES- Type 1 EAD-per-EVI routes that the rustbgpd RR reflects to a
third observing VTEP. Gated assertions cover that both ESI-sharing
peers' Type 1 EAD + Type 4 ES routes reach the observer with
correct
ORIGINATOR_ID+CLUSTER_LISTand that gRPCListEvpnRoutessurfaces both Route Type 1 and Route Type 4 entries. DF election itself runs on the VTEPs; the RR is path-transparent.
- Type 1 EAD-per-EVI routes that the rustbgpd RR reflects to a
third observing VTEP. Gated assertions cover that both ESI-sharing
peers' Type 1 EAD + Type 4 ES routes reach the observer with
correct
- See ADR-0050.
Phase 2: Bidirectional VTEP (Gates 7a, 7b, 7b+1)
- Gate 7a (v0.13.0, ADR-0052): declarative local-VTEP domain in
crates/evpn—EvpnInstanceTable+[[evpn_instances]]TOML schema + read-onlyEvpnService.ListEvpnInstances. Empty by default; RR-only deployments unchanged. - Gate 7b (v0.14.0, ADR-0054): Linux kernel reconciliation in
the new
crates/evpn-linuxcrate. TheReconcileActor<D: Dataplane>consumes atokio::sync::watch<Arc<DataplaneIntent>>from a daemon-side projection of the RIB's best-path Type 2 routes, and programs/withdraws remote-MAC FDB entries via rtnetlink (single combined-flagRTM_NEWNEIGHwithNTF_SELF | NTF_MASTER | NTF_EXT_LEARNEDandNUD_NOARP | NUD_PERMANENT). Foreign-entry preservation is structural — the delete pass iteratesOwnedSet(rustbgpd-programmed keys), never the kernel snapshot, so kernel-learned local MACs and operator-static FDB entries cannot be deleted by the algorithm. - Gate 7b+1 (v0.15.0, ADR-0055): local-MAC origination
closes the upward flow. New
crates/evpn/src/origination.rsships the pure deterministicLocalMacOriginatorstate machine encoding RFC 7432 §15.1 sequence rules: first-Learned-no-contender ⇒ seq=0 with no extcomm; first-Learned-vs-contender at R ⇒ R+1 with extcomm; remote announces M ≥ N ⇒ bump tomax(M, N) + 1; aged-then-relearn preserves the seq ratchet so a stale peer never wins contention. Local-port move on a previously-advertised MAC bumps the seq AND wakes up the extcomm even without a contender, so peers see the bumped seq on the wire (otherwise a later stale remote at seq=0 would tie our hidden seq=1). Thecrates/evpn-linux/src/linux/notify.rsclassifier subscribes toRTNLGRP_NEIGH(enum group id3, not the legacy bitmaskRTMGRP_NEIGH = 4—Socket::add_membershiptakes the enum value), dropsNTF_EXT_LEARNEDechoes (we programmed those) and VXLAN-port ifindexes (those are remote-MAC echoes), and resolves bridge-port → VNI via a newLinkCache::bridge_port_to_vnimap. The daemon-sidesrc/evpn_originator/actor mirrors the dataplane supervisor on the upward flow and emitsRibUpdate::InjectEvpn/WithdrawEvpn. Self-NH routes are filtered before reaching the state machine via the existingproject_evpn_routes, so the originator never sees its own re-Inject as a contender. - Type 3 IMET origination per L2VNI (Gate 7b+1):
src/evpn_imet.rsemits one Type 3 perEvpnInstanceat startup, withdrawn at coordinated-shutdown. Lifecycle is decoupled from kernel Ready/NotReady — IMET expresses BGP-level VNI membership, not data-plane programmability. Carries the PMSI Tunnel attribute (RFC 6514 §5, path attribute type 22) for ingress replication. - PMSI Tunnel codec (RFC 6514 §5): new
crates/wire/src/pmsi.rsdefinesPathAttribute::PmsiTunnel(PmsiTunnel)with typedPmsiTunnelType(preserves unknown values for forward-compat per RFC 7385) andPmsiTunnelIdentifier(Empty / Ipv4 / Ipv6 / Raw). Wire layout: flags(1) | type(1) | label(3) | tunnel id(variable). Thefor_evpn_ingress_replication(vni, ip)constructor encodes the label field as the raw 24-bit VNI per RFC 8365 §5.1.3 — RFC 8365 redefines the field semantics for EVPN-VXLAN so the full 24 bits are the VNI, not the MPLS-style high-20-bits shift. Matches FRR/Cumulus on the wire and stays consistent withEvpnMacIp.label1. Tunnel identifier carries the originator IP. - Coordinated shutdown ordering (Gate 7b+1): the daemon drains
the EVPN originator (emits Type 2 Withdraws) and withdraws the
IMET keys before sending
PeerManagerCommand::Shutdownso Type 2 / Type 3 Withdraws ride still-open BGP sessions. The dataplane reconciler drains afterward (FDB teardown doesn't need active BGP). - Bidirectional VTEP interop (M37): first validated end-to-end
against Linux 6.17 + FRR 10.3.1 via
tests/interop/m37-evpn-local-origination.clab.yml; the topology now pins FRR 10.7.1, and the hosted gate below runs it against that version. rustbgpd as VTEP originator, FRR as consumer. 4/4 PASS: Type 3 IMET originated at startup, Type 2 originated within ~3 s ofbridge fdb add, Type 2 withdrawn within ~3 s ofbridge fdb del, Type 3 IMET drained on shutdown. Gated in hostedkernel-dataplaneCI alongside the Gate 7b downward path (M36). - Closed after v0.16.0 (v0.17.0 follow-ups):
advertise_svi_macconsumption (origination of the bridge's own MAC on instance-Ready viaInstanceDataplaneStatus.bridge_mac),sticky_macsconfig schema (ADR-0056 — listed MACs originated with the RFC 7432 §15.4 sticky bit), sub-second mobility convergence (Gate 7c — EVPN-keyedEvpnRouteEventbroadcast incrates/rib; the 5 s poll stays asLagged/ cold-start backstop), and MAC-with-IP Type 2 origination via ARP/ND suppression (Gate 7b+2 —AF_INET/AF_INET6classifier incrates/evpn-linux,LocalMacIpOriginatorincrates/evpn, daemon correlation under the FRR-style replace model per RFC 9135 §7.2.3 insrc/evpn_originator/. Operator prerequisite: bridgeneigh_suppress on). - Partially shipped (tracked in ADR-0055 §9 +
docs/how-to/evpn-alpha-soak.md): RFC 7432 §15.1 duplicate-MAC M/N detection defaults to M=180 s / N=5 and can opt into local-originsuppress_localrecovery. Remote-route processing and dataplane loop-protection remain deferred.
Phase 3: Multi-homing foundation (Gate 8, ADR-0057)
- Gate 8 (v0.17.0, ADR-0057): observable DF election +
Type 1/4 origination. New
crates/evpn/src/segment.rscarries theEthernetSegmentdomain type (ESI, member VNIs, DF preference, algorithm, originator IP). Newcrates/evpn/src/df_election.rsships the pure(state, event) → rolesstate machine — RFC 7432 §8.5 service carving (sort candidates by originator IP ascending;vni mod npicks the slot) plus RFC 8584 §2.2 algorithm negotiation (a non-default algorithm runs only when every Type 4 route in the segment advertises it; if even a single Type 4 advertisement is received without the locally configured DF Alg, the RFC 7432 defaultDefaultModuloMUST be used — the algorithm id is not a tiebreak). Newcrates/evpn/src/origination_es.rsships three deterministic Type 1/4 originators:LocalEsOriginator(Type 4 ES),LocalEadPerEsOriginator(Type 1 EAD-per-ES with MAX_ET marker),LocalEadPerEviOriginator(Type 1 EAD-per-EVI, role-aware viaon_vni_role_changed). Daemon orchestrator atsrc/evpn_segment.rssubscribes to the EVPN best-path broadcast (Gate 7c), re-runs election on every Type 4 event, and updates the Prometheus surface (evpn_df_role{esi,vni,role}gauge,evpn_df_role_changes_total{esi,vni}counter). When one originator advertises a segment under several RDs, the route with the lowest RD supplies its DF candidate, so election does not depend on RIB iteration order. - Gate 8 scope was observation only; Gate 8b is now alpha and default-on
with explicit opt-out flags.
The follow-up Gate 8b slices add ESI Label / ES-Import RT
origination, DF-role-aware Type 2 ESI attachment, the
BUM-suppression kernel primitive behind
apply_bum_enforcement, aliasing projection, and a receive-side EAD-per-ES mass-withdraw filter. The 24 h MAC-churn soak passed 2026-05-16 (docs/soaks/soak-gate8b-mac-churn-24h.md), which unblocks the production-default flip, and the flip landed:apply_bum_enforcement/apply_aliasing_ecmpdefault totruesince v0.23.0, with explicit= falseas the documented opt-out. - M38 smoke (
tests/interop/m38-evpn-df-election.clab.yml): 2-PE rustbgpd segment, asserts (1) PE1 elected DF, (2) PE2 elected NonDF, (3) PE2 promotes to DF after PE1 shutdown, (4)evpn_df_role_changes_totaladvances on the promotion.
Phase 4: Symmetric Interface-less IRB end-to-end (Gate 9, ADR-0058)
- Gate 9 shipped end-to-end in v0.18.0:
[[evpn_ip_vrfs]]TOML schema (parsed insrc/config/schema.rsasEvpnIpVrfConfig),[[evpn_instances]].ip_vrfbinding,IpVrf/IpVrfTabledomain types undercrates/evpn/src/ip_vrf/,IpVrfStatusreadiness probe (seven ADR-0058 §3 predicates),LinuxDataplane::probe_ip_vrfs+ IP-VRF / L3 VXLAN device dumps. Slice 6 PR A (#77) added per-IP-VRF kernel-route observation + Type 5 origination viaRibUpdate::InjectEvpn. Slice 6 PR B (#78) added remote Type 5 import + L3 FIB programming through the transactionalL3OwnedStatemodel (per-prefix install state + shared kernel-neighbor / L3VXLAN-FDB refcount, value-aware drift detection, four-phase apply ordering: route-remove → resolution-add → route-add → resolution-remove, Router MAC conflict detection, foreign-state preservation). PR #79 addsRTNLGRP_IPV4_ROUTE/RTNLGRP_IPV6_ROUTEmulticast for sub-second tenantip addr delwithdraw.DataplaneReport.ip_vrf_statusDataplaneReport.ip_vrf_installed_routespropagate the verdict to subscribers;rbgp evpn vrfs [NAME]+EvpnService.ListIpVrfs/EvpnService.GetIpVrfgRPC surfaces let operators read it without scraping logs. M39 hosted kernel-dataplane CI validates the bidirectional Type 5 path against FRR 10.7.1.
- Shipped since Gate 9: auto-derived Route Targets (RFC 8365 §5.1.2.1 for L2VNI / AS:VNI for L3VNI, v0.25.0, M39b cross-vendor smoke) and Type 5 gRPC injection (v0.25.0, M45). Hosted kernel-dataplane CI gates M36–M43 incl. M39b.
- Shipped after Gate 9: receive-side overlay-index recursion for imported Type 5 routes, native GW-IP + ESI overlay-index Type 5 origination (ADR-0087), single-active ESI overlay-index Type 5 receive, and all-active ESI overlay-index Type 5 receive with route-level ECMP plus L3VXLAN FDB-NHG programming. Remaining EVPN work is outside the core overlay-index shape: Linux softswitch local-bias limits, true shared-VNI / non-zero Ethernet Tag service, managed netdev ergonomics, and service-provider route families.
RFC 9012 / RFC 8365 — BGP Encapsulation Ext Community + VXLAN-EVPN
- The BGP Encapsulation extended community (Type 0x03, Subtype 0x0C) carries 4 reserved bytes followed by a 2-byte Tunnel Type. RFC 9012 §4.1 (Figure 14) keeps the RFC 5512 layout unchanged — Reserved (2 octets), Reserved (2 octets), Tunnel Type (2 octets) — so the emitted form is the current standard. See RFC 9012 §4.1. The contrary layout description in the dated ADR-0050 is a historical documentation error; the reserved-plus-tunnel-type wire encoding did not change.
- Tunnel Type values (IANA BGP Tunnel Encapsulation Attribute Tunnel
Types): 8 = VXLAN, 9 = NVGRE, 11 = MPLS in GRE.
as_bgp_encapsulation()returns the u16 tunnel type. - rustbgpd does not yet negotiate a preferred encap. VXLAN is assumed; non-VXLAN values are passed through untouched.
- Tunnel Encapsulation (23): exact Tunnel TLV and sub-TLV framing is validated before opaque transit. rustbgpd does not interpret tunnel endpoints or select tunnels from that attribute; carriage is not RFC 9012 semantic support. The registry boundary distinguishes it from the Encapsulation Extended Community.
- Administrative-domain boundary: rustbgpd originates VXLAN Encapsulation ECs and interprets received tunnel types for local EVPN eligibility, while retaining the community across eBGP. RFC 9012 §11 requires default ingress and egress filtering by speakers that understand it. Current behavior is an explicit compatibility exception for the alpha EVPN profile in operator-controlled administrative domains: default eBGP filtering and a dedicated same-domain permission control are absent. This is a conformance limitation, not an implementation of that RFC requirement.
- Compatibility: RFC 8365 §5.1.3 requires encapsulation signaling for EVPN overlays but provides no exemption from RFC 9012 §11. Removing received communities would also change local eligibility: advertising only non-VXLAN types excludes a route from the local VXLAN profile, whereas absence uses the statically configured VXLAN fallback. A future default-filter change needs explicit same-domain authorization and migration guidance; this documented decision preserves current wire behavior and route-server transparency.
- Automatic RT derivation: L2VNI auto-RT follows RFC 8365 §5.1.2.1; L3VNI auto-RT uses AS:VNI. Both reject ASNs above 65535 with a typed error. A four-octet-AS RT has only a two-octet local administrator and cannot encode an arbitrary three-octet VNI; configure explicit RTs instead.
RFC 9135 / RFC 9136 — Symmetric Interface-less IRB (Gate 9 end-to-end, v0.18.0)
- Type 2 MAC/IP routes with a second MPLS label (Label2) and the
Router MAC ext community (Type 0x06, Subtype 0x03) are decoded
and reflected unchanged. The Router MAC ext community accessor
returns the 6-byte MAC. Slice 6 PR B interprets both:
label2carries the L3VNI on Type 5 origination, and the Router MAC drives the kernel L3 neighbor + L3VXLAN FDB rows on remote Type 5 import. - Gate 9 (ADR-0058) adopts the RFC 9136 §4.4.2 symmetric Interface-less IP-VRF-to-IP-VRF model as the only IRB mode rustbgpd supports. Asymmetric IRB (RFC 9135 §4.1) and the Interface-ful IP-VRF-to-IP-VRF model (RFC 9136 §4.4.1) are explicit non-goals.
- The
[[evpn_ip_vrfs]]config block declares per-tenant IP-VRF / L3VNI state (RD, RTs, VTEP source IP, Router MAC, observed Linux VRF + L3 VXLAN device names, VRF table id).[[evpn_instances]]gains an optionalip_vrf = "..."field binding an L2VNI to one IP-VRF tenant. - A pure-logic IP-VRF readiness probe (
rustbgpd-evpncrate) maps a portable kernel snapshot against the configuredIpVrfand returns anIpVrfStatusverdict; every failing predicate is enumerated. - Shipped in v0.18.0: per-IP-VRF kernel-route observation,
Type 5 origination via
RibUpdate::InjectEvpngated on readiness, remote Type 5 import + L3 FIB programming through the transactionalL3OwnedStatemodel with four-phase apply ordering, Router MAC conflict detection, sub-second withdraw viaRTNLGRP_IPV4_ROUTE/RTNLGRP_IPV6_ROUTEmulticast,rbgp evpn vrfsCLI +ListIpVrfs/GetIpVrfgRPC, M39 hosted smoke against FRR 10.7.1. - Shipped after Gate 9: receive-side overlay-index recursion, auto-derived Route Targets per RFC 8365 §5.1.2.1, native GW-IP + ESI overlay-index Type 5 origination (ADR-0087), single-active ESI overlay-index Type 5 receive, and all-active ESI overlay-index Type 5 receive with route-level ECMP plus L3VXLAN FDB-NHG programming. Remaining EVPN work is outside the core overlay-index shape: runtime mixed-edit tails, Linux softswitch local-bias limits, true shared-VNI / non-zero Ethernet Tag service, managed netdev ergonomics, and service-provider route families.
Later EVPN standards against the VXLAN/Linux lane
These adjacent EVPN standards affect the reflector, interconnect, or VXLAN/Linux VTEP roles. Each row separates attribute preservation from implemented service procedures.
| RFC | Status | Relevance |
|---|---|---|
| RFC 9014 | Not implemented | EVPN overlay interconnect gateway procedures, including overlay-to-MPLS interworking and Interconnect Ethernet Segments, are not implemented. Reflecting supported EVPN routes does not provide a DCI gateway. |
| RFC 9252 | Partial: service-aware reflection | Recognized malformed L3/L2 Service framing follows §7 treat-as-withdraw; see the framing contract. Structurally valid routes with no semantically valid applicable SID remain retained but are excluded from selection, Add-Path, ORR, and ECMP; see the service eligibility contract. Unchanged-next-hop reflection preserves eligible raw attributes; transport regressions cover raw receive/export. PE import, service origination, next-hop rewriting, and SRv6 forwarding are not implemented; VPN and EVPN views may show an optional display-only reconstructed_sid from a single route's transposition, unused by selection. This is not full RFC 9252 service support. |
| RFC 9251 | Partial: Type 6 SMET relay, alpha | Typed Type 6 receive/reflect/withdraw and inspection are implemented. The M113 controlled raw-peer proof checks reflected bytes and recovery with an independent TShark decoder; vendor interoperability is unproven. Types 7/8 remain unsupported typed NLRIs, counted and discarded under RFC 7606 §5.4. No SMET origination, IGMP/MLD proxy, or multicast forwarding; see the SMET boundary. |
| RFC 9746 (Mar 2025; updates RFC 7432, RFC 8365) | Not implemented | Split Horizon Type (SHT) bits in the ESI Label extended community. §2.2: an egress NVE MUST NOT use an SHT other than 00 with VXLAN (tunnel type 8), so local bias is the only multi-homing split-horizon mechanism for VXLAN. This is the normative backing for the Linux softswitch local-bias limitation in docs/reference/limitations.md; the ESI Label decoder reads only the single-active flag. |
| RFC 9785 (Jun 2025; updates RFC 8584) | Partial | Highest-/Lowest-Preference DF election is implemented (df_algorithm), under the same unanimous-or-default negotiation restated in §4.1. Don't-Preempt recovery (§4.3) derives operational preference/DP from remote Type 4 routes after a three-second wait; the §4.3(5) boot timer additionally holds any recovery started within 30 seconds of daemon start until an L2VPN/EVPN session is established and every established one has sent End-of-RIB, bounded by those 30 seconds. DP wins equal-preference ties. Mixed DP does not trigger algorithm fallback. Explicit administrative preference changes override inheritance. No per-Ethernet-Tag algorithm override (§4.2), cross-vendor DP validation, or guarantee when reference routes arrive after the recovery wait (for example, from a session established after the others have synced, or a peer that sends no End-of-RIB within the 30-second bound). The existing configured default preference remains 32768 rather than the RFC's 32767. |
| RFC 9722 (May 2025; updates RFC 8584) | Not implemented | Fast DF recovery: a Service Carving Time extended community on the Type 4 route synchronizes the DF election timer across the segment's PEs so they carve at the same instant. rustbgpd runs each election on its own timer. |
| RFC 9721 (Apr 2025; extends the RFC 7432 and RFC 9135 IRB procedures) | Partial | Extended IRB mobility. Implemented: a local bridge-port move of a MAC advertised as MAC+IP re-advertises every (MAC, IP) Type 2 for that MAC with the MAC Mobility sequence incremented (§5.1 parent/child, §6.2 inheritance; LocalMacIpOriginator::on_local_mac_moved in crates/evpn/src/origination_macip.rs). The MAC-only and per-(MAC, IP) mobility ratchets otherwise remain independent; the local-move cascade is a bump-all operation. §6.4/§6.5 peer-sync, partial: a Type 2 received from a PE on the VNI's own non-zero Ethernet Segment is a peer-sync route, excluded from mobility contention and both duplicate detectors. For an already locally learned MAC, the daemon adopts a higher peer sequence exactly, without adding one, and synchronizes locally learned MAC/IP children to the highest peer or retained local sequence. This applies in either arrival order, requires a matching configured import RT, VNI, tag zero, and a nonlocal next hop, and preserves local sticky state and lifecycle suppression (is_same_segment_peer and the originators' adopt_peer_sequence methods in crates/evpn/src/, coordinated in src/evpn_originator/rib_polling.rs). Originating a MAC or MAC/IP learned only from the ES peer (Peer-Sync-Local) is not implemented. Not applicable: §7/§8.3 RT-5 mobility — Type 5 routes carry no MAC Mobility extended community (RFC 9136 defines none for the route type; crates/evpn/src/ip_vrf/origination.rs builds ORIGIN, AS_PATH, and extended communities only, and the receive-side mobility comparison in crates/rib/src/loc_rib.rs is gated to route type 2). Mobility reaches Type 5 only through GW-IP overlay-index resolution, which already selects the highest-sequence Type 2 and fails closed when distinct MACs tie at the top sequence. §6.1 third rule: a newly activated local IPv4 or IPv6 MAC/IP binding adopts at least one above the effective maximum sequence of different imported remote MACs holding that IP, saturating at u32::MAX. This floor also preserves higher local and exact ES-peer sequences and synchronizes the MAC and its local IP children. Both local arrival orders are supported; duplicate observations, remote-only changes, and suppression recovery do not create another ownership event. Scope recreation refreshes the remote view before replay; initial or failed snapshots defer local activations in a bounded queue, which backpressures further observations when full. This is bounded local activation, not full simultaneous-move convergence. Absent: §6.3/§6.7 stale-entry procedures, full §8.2 duplicate-address procedures, and §6.8 probing. Optional §8.2 detect-only IP accounting is implemented: conflicting local/local or local/remote MAC ownership uses a per-(VNI, IP) M/N window, excludes sticky and duplicate-MAC-quarantined contenders, and reports counters/warnings without suppression or sequence changes. |
| RFC 9786 (Jun 2025) | Not implemented | Port-active multi-homing redundancy mode (per-port active/standby). Demand-shaped alongside the other redundancy modes outside all-active and single-active. |
| RFC 10039 (Sep 2026) | Partial: D-PATH framing only | The D-PATH attribute (36) is retained as an opaque optional-transitive attribute on every family. §4 framing is validated: malformed segments or a total length under 8 octets are treat-as-withdraw; see the D-PATH boundary. No gateway interconnect procedures are implemented: there is no D-PATH origination or prepending, domain-loop detection, D-PATH route selection, attribute propagation modes, or EVPN/IPVPN re-origination. |
The DP interpretation in historical ADR-0057 is superseded: RFC 9785 §4.1 requires DP as an equal-preference tie-break. Section 4.3 preserves a remote reference by changing the recovering PE's advertised preference and DP, so restart recovery uses received routes rather than persistent incumbent memory.
Type 6 SMET reflection
Type 6 Selective Multicast Ethernet Tag (SMET) routes use the normal alpha EVPN RR path: typed receive, policy and selection, reflection, and withdrawal. Route identity includes RD, Ethernet Tag, source, group, and originator IP. Flags are mutable payload outside that identity. The structural codec retains the full flags byte, including reserved high bits. Source and group can be wildcards; a wildcard group requires a wildcard source. Concrete source/group addresses share a family, while the originator address family is independent. See RFC 9251 §9.1 and its default wildcard route. RFC 9625 §3.3 also permits zero flags for SBD-SMET wildcard state when IGMP/MLD reports are not required.
The announcement validator implements the following low-nibble acceptance
matrix. High flag bits do not affect acceptance and remain preserved. A
familyless (*,*) route accepts either family's wildcard profile; its
originator address does not select a profile.
| Source/group shape | Accepted low flag nibble (hex) |
|---|---|
IPv4 (*,G) | 0, 2, 3, 8, A, B, C, D, E, F |
IPv4 (S,G) | 4, 5, C, D |
IPv6 (*,G) | 0, 1, 8, 9, A, B |
IPv6 (S,G) | 2, A |
Familyless (*,*) | 0, 1, 2, 3, 8, 9, A, B, C, D, E, F |
Structural decoding and announcement admission are separate. For a canonical Type 6 NLRI with an invalid announcement flag profile, revised MP_REACH handling records treat-as-withdraw and retains every extracted route key, including valid siblings. MP_UNREACH uses the flags-free identity and accepts those flags so a withdrawal can remove retained state. Missing or extra flag octets and malformed address framing remain session-reset errors; they do not provide a safely extracted canonical route. Warm-state loading uses the same announcement validator before admission.
Type 6 is represented in route and event views, exact explain, generic MRT records, warm state, and BMP output. Filters accept route type 6; see the API contract and operator examples. No multicast service is implied: SMET origination, IGMP/MLD proxy behavior, and multicast forwarding are not implemented. Types 7–11 remain unsupported, counted, and discarded. The M113 controlled raw-peer receipt records 41 protocol phases, 15 TCP connections, and 46 reflected SMET NLRIs, checked against an independent TShark decoder. It covers IPv4 wildcard-source, IPv6 source-specific, and familyless wildcard routes; flags-only replacement, withdrawal fallback, UPDATE-wide treat-as-withdraw and same-session recovery; and six structural-error resets with same-peer reconnect recovery. This is neither vendor interoperability nor scale evidence. Historical Types 1–5 receipts remain scoped to their original route types.
Extended Communities — non-transitive eBGP export
- draft-ietf-idr-rfc4360-bis-09 §6 (RFC Editor queue): received
non-transitive Extended Communities remain available to local policy and
iBGP, but ordinary eBGP export removes them after export policy and before
UPDATE encoding.
send_non_transitive_extended_communities = trueis the explicit per-neighbor / peer-group opt-in for crossing that AS boundary. - RFC 7947 §2.2.4:
route_server_clientexport is exempt and preserves transitive and non-transitive Communities. iBGP also preserves them because it crosses no AS boundary. - The shared classic-unicast and MP attribute-preparation paths apply the same rule. Mixed attributes retain their normal or Partial form with only the transitive values; an attribute with no remaining values is omitted.
- This is send-side behavior only. Inbound decoding, retained attributes, and export-policy semantics are unchanged.
RFC 10005 — Link Bandwidth receiver subset
- §2 / §3.2:
ExtendedCommunity::as_link_bandwidth()accepts the exact transitive (0x00) and non-transitive (0x40) two-octet-AS-specific types with subtype0x04. It returns the raw AS and IEEE-754 bytes/second value;ExtendedCommunity::link_bandwidth()continues to construct non-transitive type0x40. - §4:
Route::link_bandwidth()chooses the lowest finite nonnegative value, independent of transitivity or AS. Positive and negative zero are valid; negative values are ignored. NaN and infinities are also ignored as a local finite-value policy. With no usable value, existing equal-cost fallback wins. - §3.3.2: receive-side inspection does not mutate, reorder, or deduplicate the raw Extended Communities vector, so unchanged-next-hop reflection remains byte-preserving.
The typed accessor remains a receiver subset. Re-advertisement follows the non-transitive eBGP export rule above.
EVPN Extended Communities — typed accessors (RFC 7432 §7.5-§7.8)
Subtypes with typed accessors on ExtendedCommunity (others pass
through as opaque u64):
| Type/Subtype | Name | Payload |
|---|---|---|
| 0x03 / 0x0C | BGP Encapsulation | u16 tunnel type |
| 0x03 / 0x0D | Default Gateway | flag-only (value = 0) |
| 0x06 / 0x00 | MAC Mobility | (sticky: bool, sequence: u32) |
| 0x06 / 0x01 | ESI Label | (single_active: bool, label: u32) |
| 0x06 / 0x02 | ES-Import RT | 6-byte MAC target |
| 0x06 / 0x03 | Router MAC | 6-byte MAC |
- RFC 8214 Layer 2 Attributes (Type 0x06 / Subtype 0x04) deferred — encoding is complex and not needed for Phase 1 RR flow.