Runbook: RR pair day-2 operations
Operate a production route-reflector pair safely.
Operate a production route-reflector pair safely.
When this is you: Your redundant route-reflector pair is in production. It follows the route-reflector recipe, and you need to make routine changes without dropping the fleet. Rule of thumb for every task below: change one RR, verify, then the other.
GR settings sanity
The RR is a receiving-speaker GR/LLGR helper: on a client restart it
retains routes for the advertised restart window (RFC 4724), then
demotes to LLGR_STALE (RFC 9494) if llgr_stale_time > 0. Check the
peer-group template carries what you think it does:
[peer_groups.rr-clients]
graceful_restart = true
gr_stale_routes_time = 120
llgr_stale_time = 300During a client restart, watch the helper actually helping:
rbgp metrics | grep -E 'bgp_gr_active_peers|bgp_gr_stale_routes|bgp_gr_timer_expired_total'gr_stale_routes_time is hot-applied per the
reload matrix; graceful_restart itself binds at
OPEN negotiation, so changing it resets the session (see below).
Adding a client
If the fleet auto-accepts via a [[dynamic_neighbors]] range, a new
client inside the range needs nothing. Otherwise:
rbgp neighbor 10.0.0.42 add --remote-asn 65000 --peer-group rr-clients --description new-client
# or widen the accept range:
rbgp dynamic-neighbor add 10.0.9.0/24 --peer-group rr-clientsBoth persist to the daemon's config file — the CONFIG_PATH argument,
or /etc/rustbgpd/config.toml when it is omitted — before the RPC
returns. The static neighbor inherits RR-client, family, and GR settings from
rr-clients without materializing CLI defaults. Verify with rbgp neighbor --wide
— the RRC column marks reflector clients. Uniform clients join the
existing update group automatically; there is nothing to tune.
Config edits: hot vs session reset
Before touching a live RR, classify the edit — the reload matrix is the index, the daemon is the oracle:
rbgp config diff /etc/rustbgpd/candidate.tomlNot sure what a value on the live RR currently is (an inherited
peer-group timer, a defaulted hold_time)? Dump the effective running
config first — defaults resolved, selected default-empty policy lists omitted,
secrets redacted:
rbgp config effectiveThe diff annotates each changed field. Two classes matter here:
- Hot-applied (
description,max_prefixes,gr_stale_routes_time,log_level,remove_private_as, import/export policy and chain refs, ...): applied in place on SIGHUP — session task, TCP connection, and FSM untouched; policy edits trigger Route Refresh / re-emit. - Session reset (
families,add_path,graceful_restart,hold_time,md5_password, ...): applied through an immediate delete + re-add of the session on SIGHUP, so the fleet sees a flap now, not at the next natural reset. Do these one RR at a time and confirm the other RR is carrying full state first (rbgp neighbor --wideprefix counts on the peer RR).
Mixing a hot edit and a reset edit on one neighbor applies both through one session rebuild.
Risky edits: commit-confirm
For anything you would want auto-rolled-back if you cut yourself off, use a config transaction with a confirm timer instead of SIGHUP:
rbgp config plan /etc/rustbgpd/candidate.toml
# plan prints the runtime snapshot token; it is opaque — capture it and pass
# it back verbatim on apply:
RUNTIME_SNAPSHOT_TOKEN="$(rbgp --json config plan /etc/rustbgpd/candidate.toml \
| jq -r .runtime_snapshot_token)"
rbgp config apply /etc/rustbgpd/candidate.toml \
--expected-runtime-snapshot-token "$RUNTIME_SNAPSHOT_TOKEN" \
--confirm-id rr1-edit-$(date -u +%Y%m%d-%H%M) \
--confirm-timeout 120
rbgp config status # pending / confirmed / auto_reverted
rbgp config confirm rr1-edit-... # keep it
rbgp config abort rr1-edit-... # or roll back nowIf the timer expires unconfirmed — or the daemon restarts inside the
window — the pre-commit config is re-applied. While a confirm is
pending, SIGHUP and other config mutators are fenced off. Full
lifecycle and failure states: OPERATIONS.md.
Draining an RR for maintenance
rbgp gshut --all # tag all outbound with GRACEFUL_SHUTDOWN (RFC 8326)
# clients honoring 8326 de-pref this RR; then stop the daemon:
rbgp shutdown # writes the GR marker, notifies peersOn a terminal, both commands ask for confirmation first; --yes skips the
prompt, and non-interactive runs never prompt.
Bring it back, confirm rbgp neighbor --wide converges to the same
prefix counts as its twin, then repeat on the other RR.
Runbook: a peer is flapping
Diagnose a BGP session that keeps flapping.
Runbook: activation exit 5 (manual recovery)
When this is you: rs-config-render activate or rs-config-render ixp-manager-lifecycle run returned exit 5 — the activation command started, but the helper could not prove the daemon settled on the candidate — and every later run (and resume) also returns exit 5.