Paired route servers — staggered updates and maintenance windows
Operate a resilient pair of IXP route servers.
Operate a resilient pair of IXP route servers.
When this is you: Your exchange runs or should run two route servers. Every member peers with both, and you want the operational discipline that makes the pair actually redundant: independent instances, staggered config rollout, a mechanical inter-RS consistency check, and a drain flow for maintenance.
Why two, and why members peer with both
A route server is a single point of failure for multilateral peering: when it goes down, members lose every route learned through it. The standard remedy is two route servers on separate hosts (ideally separate power/switch domains), with every member holding a session to each. Both carry the same policy and advertise the same routes, so a member losing RS1 converges onto identical paths already learned from RS2 — the fabric forwards through the same next hops throughout (RFC 7947 transparency: the route server is not in the data path). Redundancy only holds if members actually configure both sessions; make dual sessions part of member onboarding, not an option.
Independence rules
- Two instances, two hosts, zero shared runtime state. The instances never talk to each other; consistency comes from feeding both the same policy inputs, not from synchronization.
- If a small deployment must share one host, bind distinct endpoints. Set
[global].listen_addressesto one local address per enabled family on each instance. The same addresses source active opens, so the instances do not collide on TCP/179. Every served family needs distinct addresses on both instances; do not mix wildcard and exact mode on the same port. Addresses must already exist locally because there is no freebind or fallback. This is a restart-required endpoint choice, not per-neighbor source selection; separate hosts remain the resilient design. - One policy source, two renders. With the
IXP filter pipeline, both route servers
render from the same arouteserver
general.yml/clients.yml— run the render per host (or distribute one render's output), and keep each host'srustbgpd --check --strictgate local so a corrupt copy cannot pass. - Distinct
router_ids, same ASN, same member ACLs. Members see two ordinary sessions to one route-server AS. - Independent RPKI feeds. Point each instance at the RTR
cache set directly (
[rpki]takes multiplecache_servers); do not proxy one instance's view through the other's host.
Staggered config updates
Never reload both route servers in the same step — the second instance is your rollback while the first proves the change. Per update:
- Validate everywhere first:
rustbgpd --check --stricton the candidate config on both hosts (exit 0 or stop), and check the reload matrix for whether any touched field is restart-required. - Reload RS1 only (SIGHUP). Parse, validation, or dataset load failures
and rejected family combinations leave runtime untouched. Member, policy,
and dataset changes, including a member joining or leaving, take the
generation route: a later failure restores RS1's prior generation and
rejects the reload, including for a member with
md5_passwordorttl_security. Only the sequential route (for example a TCP-AO rotation, or a changed password or GTSM setting on a member that stays) can halt with known partial changes.rustbgpd --diffprints the route asSIGHUP reload route; inspect the reload result and effective configuration before proceeding (SIGHUP reload routes). - Verify RS1: sessions established (
rbgp summary), spot-check a member's view (rbgp rib sent <member>), no alert movement. - Soak for an operator-chosen window (long enough for a full member announcement cycle and your monitoring interval).
- Reload RS2, re-verify, then run the consistency check below.
The same stagger applies to daemon upgrades, with the maintenance drain around each restart.
Inter-RS consistency: rbgp diff advertised
Two independently rendered route servers can drift — a stale render on
one host, a reload that never happened, a member session down on one
side. The check is mechanical and fail-closed
(docs/how-to/ribdiff.md): snapshot what RS2 advertises,
compare RS1's live Adj-RIB-Out against it, gate on the exit code.
Capture RS2's advertised view from its own BMP feed — wire-true,
including attributes no CLI view renders. Enable the RFC 8671
post-policy stream on RS2 (restart-required, so make it part of the
standing config rather than a per-check toggle; see
[bmp]):
[[bmp.collectors]]
address = "192.0.2.100:11019" # your capture/collector host
monitor = ["rib_out_post"]The capture has to contain a complete Adj-RIB-Out boundary for each member.
RS2 sends an automatic initial dump when the member's session establishes.
A collector that connects later gets the Peer Ups and live updates only
(rib_out_post has no automatic reconnect dump), and from-bmp refuses
such a capture (exit 2, End-of-RIB not seen) rather than emit an incomplete
peer. rbgp neighbor <member> refresh-out reports scheduling and re-sends
routes without an End-of-RIB. A member's own ROUTE-REFRESH is answered
with an End-of-RIB-Refresh marker instead when it negotiated enhanced
route refresh (FRR 10.7.1 does) — a ROUTE-REFRESH message, which BMP route
monitoring never carries.
For automatic capture, start the listener before RS2's member sessions
establish — before RS2 starts (in practice its maintenance-window restart)
or before its member sessions are cleared — and leave it running. Live
updates fold into the same capture. After a dropped BMP connection, a new
capture needs another complete boundary from establishment or the explicit
outbound replay below. Reconnect does not automatically reconstruct the prior
rib_out_post stream.
# Terminal 1: start before RS2's member sessions come up; leave running.
nc -l 11019 > rs2-adjout.bmpA late collector can instead use the experimental
replay-out operation
on RS2, one member at a time, after this outbound-only collector connects:
rbgp neighbor <member> replay-out. Each selected session must be Established,
negotiate only IPv4/IPv6 unicast with at least one family, and have an IP
address unique among managed peers. Replay resets that collector's cached
peer inventory and reannounces the selected member's routes on its live BGP
session. The CLI confirms scheduling only; require terminal BMP EoRs for
every negotiated family before conversion, and reject partial captures.
Once the members have converged, take a fixed copy in another terminal. A copy can end inside a BMP message while the listener appends; if conversion refuses it, take a new copy and retry. Compare only after successful conversion:
# Terminal 2: keep the previous snapshot intact if this capture is refused.
if cp rs2-adjout.bmp rs2-adjout-copy.bmp &&
rbgp diff snapshot from-bmp rs2-adjout-copy.bmp > rs2.ndjson.tmp; then
mv rs2.ndjson.tmp rs2.ndjson &&
rbgp diff advertised --against rs2.ndjson --ignore-attribute unknown
echo $? # 0 in sync, 1 divergent, 2 comparison refused (RS1's gRPC socket)
else
rm -f rs2.ndjson.tmp
echo "Capture copy refused; no comparison performed" >&2
fi--ignore-attribute unknown is needed with RFC 9234 roles configured:
a route server attaches OTC on the wire, so the capture carries it and
gRPC does not expose it; without the flag every route reports as
attribute-changed. This flag excludes all unknown/opaque attributes, not
just OTC: the verdict covers the remaining attributes and is not full wire
attribute equivalence. A refused conversion names the first peer and family
whose End-of-RIB is missing and writes no snapshot; --neighbor (alias
--peer) narrows the snapshot to the members that did complete.
Run it after every staggered rollout completes, and on a schedule
between rollouts from the still-running capture; archive the --json
report. Expected divergence,
not a finding: a member session down on exactly one instance makes
that member's routes one-sided everywhere. Anything else is drift —
diff the two hosts' render receipts and reload timestamps first.
Maintenance-window flow
Taking RS1 down (upgrade, host maintenance) without a member-visible routing gap — RFC 8326 graceful shutdown, initiator side (full semantics and verification in OPERATIONS.md):
- Confirm RS2 is healthy and consistent (the diff above, all member sessions established). Two-instance redundancy means never starting maintenance while the survivor is degraded.
- Drain RS1:
rbgp gshut --all(on a terminal it asks for confirmation;--yesskips the prompt). Outbound paths get theGRACEFUL_SHUTDOWNcommunity; members honoring it demote those paths, so the RS2-learned copies win before anything closes. - Wait for the shift (operator-defined; verify on a member: the RS1-learned paths show demoted preference, best paths point at RS2-learned copies).
- Do the maintenance. Members stay converged via RS2.
- Restore RS1, wait for sessions and full table
(
rbgp summary,rbgp rib sent <member> --countplausible). - Clear the drain:
rbgp gshut --all --clear(the toggle does not persist across restart by design — after a restart-type maintenance, step 6 is a no-op, but run it when the daemon kept running). Re-run the consistency check.
The honest caveat: the drain only moves members that honor
GRACEFUL_SHUTDOWN (rustbgpd receivers opt in via
honor_graceful_shutdown; support varies across member stacks).
Members that ignore it still reconverge onto RS2 when the sessions
close — the drain turns that reconvergence from break-before-make into
make-before-break for everyone who participates.