rustbgpd
Cookbook

Runbook: RR pair day-2 operations

Operate a production route-reflector pair safely.

Operate a production route-reflector pair safely.

When this is you: Your redundant route-reflector pair is in production. It follows the route-reflector recipe, and you need to make routine changes without dropping the fleet. Rule of thumb for every task below: change one RR, verify, then the other.

GR settings sanity

The RR is a receiving-speaker GR/LLGR helper: on a client restart it retains routes for the advertised restart window (RFC 4724), then demotes to LLGR_STALE (RFC 9494) if llgr_stale_time > 0. Check the peer-group template carries what you think it does:

[peer_groups.rr-clients]
graceful_restart = true
gr_stale_routes_time = 120
llgr_stale_time = 300

During a client restart, watch the helper actually helping:

rbgp metrics | grep -E 'bgp_gr_active_peers|bgp_gr_stale_routes|bgp_gr_timer_expired_total'

gr_stale_routes_time is hot-applied per the reload matrix; graceful_restart itself binds at OPEN negotiation, so changing it resets the session (see below).

Adding a client

If the fleet auto-accepts via a [[dynamic_neighbors]] range, a new client inside the range needs nothing. Otherwise:

rbgp neighbor 10.0.0.42 add --remote-asn 65000 --peer-group rr-clients --description new-client
# or widen the accept range:
rbgp dynamic-neighbor add 10.0.9.0/24 --peer-group rr-clients

Both persist to the daemon's config file — the CONFIG_PATH argument, or /etc/rustbgpd/config.toml when it is omitted — before the RPC returns. The static neighbor inherits RR-client, family, and GR settings from rr-clients without materializing CLI defaults. Verify with rbgp neighbor --wide — the RRC column marks reflector clients. Uniform clients join the existing update group automatically; there is nothing to tune.

Config edits: hot vs session reset

Before touching a live RR, classify the edit — the reload matrix is the index, the daemon is the oracle:

rbgp config diff /etc/rustbgpd/candidate.toml

Not sure what a value on the live RR currently is (an inherited peer-group timer, a defaulted hold_time)? Dump the effective running config first — defaults resolved, selected default-empty policy lists omitted, secrets redacted:

rbgp config effective

The diff annotates each changed field. Two classes matter here:

  • Hot-applied (description, max_prefixes, gr_stale_routes_time, log_level, remove_private_as, import/export policy and chain refs, ...): applied in place on SIGHUP — session task, TCP connection, and FSM untouched; policy edits trigger Route Refresh / re-emit.
  • Session reset (families, add_path, graceful_restart, hold_time, md5_password, ...): applied through an immediate delete + re-add of the session on SIGHUP, so the fleet sees a flap now, not at the next natural reset. Do these one RR at a time and confirm the other RR is carrying full state first (rbgp neighbor --wide prefix counts on the peer RR).

Mixing a hot edit and a reset edit on one neighbor applies both through one session rebuild.

Risky edits: commit-confirm

For anything you would want auto-rolled-back if you cut yourself off, use a config transaction with a confirm timer instead of SIGHUP:

rbgp config plan /etc/rustbgpd/candidate.toml
# plan prints the runtime snapshot token; it is opaque — capture it and pass
# it back verbatim on apply:
RUNTIME_SNAPSHOT_TOKEN="$(rbgp --json config plan /etc/rustbgpd/candidate.toml \
  | jq -r .runtime_snapshot_token)"
rbgp config apply /etc/rustbgpd/candidate.toml \
  --expected-runtime-snapshot-token "$RUNTIME_SNAPSHOT_TOKEN" \
  --confirm-id rr1-edit-$(date -u +%Y%m%d-%H%M) \
  --confirm-timeout 120

rbgp config status        # pending / confirmed / auto_reverted
rbgp config confirm rr1-edit-...   # keep it
rbgp config abort rr1-edit-...     # or roll back now

If the timer expires unconfirmed — or the daemon restarts inside the window — the pre-commit config is re-applied. While a confirm is pending, SIGHUP and other config mutators are fenced off. Full lifecycle and failure states: OPERATIONS.md.

Draining an RR for maintenance

rbgp gshut --all           # tag all outbound with GRACEFUL_SHUTDOWN (RFC 8326)
# clients honoring 8326 de-pref this RR; then stop the daemon:
rbgp shutdown              # writes the GR marker, notifies peers

On a terminal, both commands ask for confirmation first; --yes skips the prompt, and non-interactive runs never prompt.

Bring it back, confirm rbgp neighbor --wide converges to the same prefix counts as its twin, then repeat on the other RR.

Source on GitHub

On this page