Operations Guide
Operate, monitor, and troubleshoot a running daemon.
Operate, monitor, and troubleshoot a running daemon.
Practical reference for running rustbgpd in production. For config syntax, see CONFIGURATION.md. For security posture, see SECURITY.md. For the end-to-end install + lifecycle walkthrough (systemd setup, Docker, containerlab quick-start, sample profiles), see deployment.md.
Man pages for both binaries ship in the release tarball under
share/man/ and can be regenerated any time with rbgp man and
rustbgpd --man (roff on stdout; view with rbgp man | man -l -).
Starting the daemon
rustbgpd /etc/rustbgpd/config.tomlOr via systemd (see examples/systemd/rustbgpd.service):
sudo systemctl start rustbgpdThe daemon validates the config file at startup. Validation errors display rustc-style diagnostics showing the offending TOML line with column markers:
error: invalid hold_time 2: must be 0 or >= 3
--> /etc/rustbgpd/config.toml:12:13
|
12 | hold_time = 2
| ^ must be 0 or >= 3The daemon exits with code 1 — it never starts with an invalid config.
On success, logs go to stdout: structured JSON under log_format = "json",
human-readable lines under log_format = "text". If prometheus_addr is
configured, use GET /readyz on that listener as the orchestrator readiness
signal; it returns ready once the PeerManager answers an O(1) actor ping and
the RIB actor answers its O(1) Loc-RIB count query within one shared 200 ms
deadline. The readiness ping does not collect peer inventory or query session
state; GetHealth continues to return that detailed snapshot. The starting rustbgpd log line means process startup reached runtime wiring, not that every
actor has answered a readiness probe yet. Under the shipped Type=notify
systemd unit the daemon also sends READY=1 only after every configured gRPC
listener is bound and that boundary is reached. A gRPC bind failure prevents
configured-peer registration, BGP ingress, telemetry, dial-out, and readiness
publication, then enters shortened coordinated teardown. The daemon sends
STOPPING=1 when coordinated shutdown begins and WATCHDOG=1 from the same
core-actor probe. PID 1 independently enforces the unit's five-minute watchdog
deadline; see deployment.md.
Runtime-config settlement fail-stop
The dedicated operator page — ownership phases, the two bounds, the fence reasons, the exit-70 recovery runbook, and the supervisor contract — is settlement-watchdog.md. This section is the condensed production reference.
Persisted runtime mutations use a cancellation-shielded owner. If creation of
the initial persistence stage fails because the config directory is not
writable, the request is rejected cleanly before runtime state changes and the
daemon remains available. That is different from failure after durable staging
and runtime mutation have begun. A failed rename is NotPublished, with the
old file still authoritative. It returns cleanly only after complete
acknowledged compensation; an unrestorable session identity instead fences as
KnownDivergence. Directory fsync failing after rename is
PublicationAmbiguous: the complete candidate is visible and adopted, no
rollback is attempted, and restart decides durability. A read-only bind-mounted
config file in an otherwise writable directory can stage successfully but
fails replacement as NotPublished.
A recovery-fenced owner immediately makes /readyz fail, closes admission for
new persisted mutations, and retains ownership until process death. Detected
divergence, publication ambiguity, acknowledgement loss, or executor loss exits
with status 70 after a five-second fencing grace. A silent owner is fenced after
the fixed 30-minute owned-settlement budget and exits five seconds later. The
fatal clock is a prestarted OS thread independent of the mutation executor,
readiness propagation, metric collection, logging, and retained-resource locks;
none of those paths can postpone the 30-minute-plus-five-second deadline. Do not
classify 70 as success or suppress its restart: the new process re-establishes
authority from durable state.
While one owner exists, /metrics exposes exactly one series for each family
below; they emit no idle tuple.
| Metric | Type | Meaning |
|---|---|---|
bgp_runtime_config_settlement_active | gauge | 1 while one owner is live or recovery-fenced. |
bgp_runtime_config_settlement_elapsed_seconds | gauge | Seconds since ownership, derived at scrape time from the stored start instant — it keeps growing while a wedged owner's phase stays frozen. |
bgp_runtime_config_settlement_budget_seconds | gauge | The fixed 1800-second fail-stop budget. |
bgp_runtime_config_settlement_fail_stops_total | counter | Becomes 1 when recovery fencing wins; exit 70 follows after the grace. |
All four use the bounded labels kind, phase
(owned_preflight, mutating, or settling_rollback), response_attached
(attached or detached), and fence_reason (none, budget_expired,
executor_lost, known_divergence, publication_ambiguous,
acknowledgement_lost, or operator_forced). Detached means the daemon owns
the work without a live RPC response; it does not weaken the deadline. The
daemon-generated operation ID appears in logs at ownership registration, each
phase transition, settlement, and fencing — never as a metric label.
The daemon warns once at 15, 25, and 29 minutes without extending or resetting
the deadline. BgpRuntimeConfigSettlementSlow fires after ten minutes of owned
settlement and BgpRuntimeConfigSettlementHalfBudget at 50% of the budget; at
either point identify the labeled owner and inspect the owning actor rather
than retrying the mutation. Fencing produces one redacted diagnostic containing
only operation identity/classification, phase, elapsed/budget seconds,
attachment, terminal, optional fence reason, SIGHUP reload_step and
accepted_effect, and exit status. It never includes
config contents, tokens, confirm IDs, comments, credentials, paths, digests,
candidates, or raw error text.
The shipped systemd unit uses Restart=on-failure, but caps recovery at five
starts per ten minutes so a deterministic persistence fault cannot flap every
five seconds forever. TimeoutStopSec=32min lets an explicit stop wait through
the watchdog. A further SIGTERM or SIGINT forces the settling owner to
fail-stop instead (fence_reason="operator_forced", exit 70). What the
supervisor does next depends on how the stop was requested: raw signals to a
running unit end in exit 70, an unclean exit code that Restart=on-failure
restarts, while systemd never automatically restarts a unit stopped
explicitly with systemctl stop, whatever its exit status. Recovery from the
persisted transaction runs on the next actual start either way. A wedge that
consumes the full 30-minute budget intentionally does not reach five starts
in ten minutes; the limit bounds fast deterministic
failures, while the independent fatal clock still bounds each slow wedge. After
inspecting and fixing the config directory, bind mount,
and on-disk authority, recover a rate-limited unit with:
sudo systemctl reset-failed rustbgpd
sudo systemctl start rustbgpdPer-peer log filtering
Set log_level on any neighbor or peer group to override the global log level:
[[neighbors]]
address = "10.0.0.1"
remote_asn = 65001
log_level = "debug"Or filter via RUST_LOG using the per-peer tracing span:
RUST_LOG='info,[peer{peer_addr=10.0.0.1}]=debug' rustbgpd /etc/rustbgpd/config.tomlThe span directive needs the square brackets. RUST_LOG is parsed one
comma-separated directive at a time: a directive that does not parse, such as
the unbracketed peer{peer_addr=10.0.0.1}=debug, is dropped and the valid
directives are kept. The dropped directives are reported with their parse
errors in one warning: ignoring unparseable RUST_LOG directive(s), ... line
on stderr at startup and whenever a successful SIGHUP reload rebuilds the log
filter; the reload also logs that line as a warn event. Only a RUST_LOG
with no valid directive falls back to info. An empty, comma-only or
whitespace-only RUST_LOG is treated as unset.
The span directive and log_level select events emitted inside that peer's
session span: the BGP session task, its connect path, and its socket writer.
They do not select events about the peer that other components emit outside
the span: the TCP listener, the peer manager (lifecycle, inbound
connections, policy apply, snapshots), reload, the RIB, BFD, and the
blackhole route installer.
tracing filter directives cannot match an event field's value, so select
those events in the JSON log instead (this needs log_format = "json").
Daemon events about a BGP peer name its address in the peer event field, and
session-span events also carry the span's peer_addr:
jq -c 'select(.fields.peer == "10.0.0.1"
or any(.spans[]?; .peer_addr == "10.0.0.1"))' rustbgpd.logEvents about a link-local peer configured with an interface show either the
bare address or the scoped label, such as fe80::1%eth0, so match both.
Raise the global level, or the level of the emitting target (for
example rustbgpd::peer_manager=debug), when the out-of-span events you
need are below info.
Config validation
Validate a config file without starting the daemon:
rustbgpd --check /etc/rustbgpd/config.tomlPrints rustc-style diagnostics on error, or config OK on success.
A valid config can still be worth flagging. When it is, the summary reads
config VALID, <n> WARNINGS — NOT a clean check instead of config OK,
the warnings are framed on stderr above it, and the exit code stays 0 —
these are things to look at, not failures. Three conditions are counted today:
- a configured eBGP neighbor or dynamic range with no explicit policy in a
direction — unfiltered with
[global] ebgp_requires_policyoff, carrying no routes in that direction with it on; rfc8212_secure_default_ready— the RFC 8212 posture is still inherited from legacy omission. It clears by writing an explicit pair: rootconfig_epoch = 1with[global] ebgp_requires_policy = false, orconfig_epoch = 2withebgp_requires_policy = true;orr_vantageconfigured while no neighbor negotiates thelinkstatefamily, so the optimal-route-reflection topology feed stays empty.
All are legitimate configurations and none is rejected; none should reach
production unnoticed. Note that (2) fires on the omission alone, however well
policed the config is — a --strict gate on an epoch-less config exits 1 for
that reason and no other.
Add --strict to make any warning exit 1 instead of 0
(rustbgpd --check --strict /etc/rustbgpd/config.toml) — for CI and
deployment gates that must not accept a valid-but-risky config.
--strict without --check is rejected (exit 2). See the validation
workflow table in deployment.md for the full exit-code
contract.
Config diff (dry-run reload)
Preview what a SIGHUP reload would change before sending it:
# Compare proposed config against current config
rustbgpd --diff /tmp/new-config.toml /etc/rustbgpd/config.toml
# JSON output for scripting
rustbgpd --diff /tmp/new-config.toml /etc/rustbgpd/config.toml --jsonWhen the daemon is already running, compare a candidate file against the live runtime snapshot instead of the on-disk file, then plan/apply a supported transaction with an optimistic runtime snapshot token:
rbgp is the supported CLI spelling.
Before the first runtime change. Any mutation that persists rewrites the config file in canonical form: comments and formatting are not preserved, defaults are canonicalized (default-valued optional sections and selected default-empty collections are omitted), and the file ends up owned by the daemon user at mode
0600. See Runtime changes rewrite the config file.
# What is the daemon actually running? Dump the effective config with
# defaults resolved (hold_time, send_hold_time, GR timers, families; selected
# default-empty policy lists omitted) and secrets redacted:
rbgp config effective
rbgp -j config effective
rbgp config diff /tmp/new-config.toml
rbgp --json config diff /tmp/new-config.toml
rbgp config plan /tmp/new-config.toml
rbgp --json config plan /tmp/new-config.toml
# Apply can plan the same candidate automatically, or consume the exact token
# printed by `config plan` when a separate review/approval step is required.
# The runtime snapshot token is opaque — capture it and pass it back verbatim:
RUNTIME_SNAPSHOT_TOKEN="$(rbgp --json config plan /tmp/new-config.toml \
| jq -r .runtime_snapshot_token)"
rbgp config apply /tmp/new-config.toml \
--expected-runtime-snapshot-token "$RUNTIME_SNAPSHOT_TOKEN"
rbgp config apply /tmp/new-config.toml \
--plan-token 550e8400-e29b-41d4-a716-446655440000 \
--expected-runtime-snapshot-token "$RUNTIME_SNAPSHOT_TOKEN"rbgp config effective downloads this byte-exact full document with a finite
384 MiB TOML ceiling (plus the protobuf envelope). The same method-specific
allowance covers the effective-config copy collected by rbgp doctor; other
config RPCs retain the ordinary 4 MiB client decode limit.
rbgp config plan and config apply stream the candidate in bounded frames;
they are not limited by tonic's 4,194,304-byte unary-message ceiling. The plan
receipt includes a single-use plan_token bound to the candidate digest,
length, and runtime snapshot. Supplying it with --plan-token skips the
automatic plan and applies only that exact reviewed candidate. Without it,
config apply first streams a plan for the same bytes and applies only a
COMMITTABLE result. NOOP and REJECTED return a non-mutating status without
calling Apply; the CLI prints that receipt and exits 0 for NOOP and 3 for
REJECTED. Older daemons are supported only when the streaming method is
explicitly UNIMPLEMENTED; Plan and automatically planned Apply then retry the
legacy unary RPC if the encoded request fits its four-MiB ceiling. An explicit
--plan-token fails closed when streamed Apply is unavailable because unary
Apply cannot honor the reviewed binding. Other streaming errors are terminal.
For a safe deploy that should roll back unless explicitly confirmed, provide a confirm handle and timeout on the same apply. The daemon applies the candidate immediately, starts the timer, and rolls back when the timer expires unless the same handle is confirmed:
rbgp config apply /tmp/new-config.toml \
--plan-token 550e8400-e29b-41d4-a716-446655440000 \
--expected-runtime-snapshot-token "$RUNTIME_SNAPSHOT_TOKEN" \
--confirm-id deploy-20260605-1 \
--confirm-timeout 120
rbgp config status
rbgp config confirm deploy-20260605-1
# or, to roll back immediately:
rbgp config abort deploy-20260605-1Confirm handles must be non-empty, at most 128 characters, and free of control
characters. --confirm-timeout requires --confirm-id; the daemon default is
600 seconds and the maximum accepted timeout is 86400 seconds. The full window
starts after Apply commits, so a large candidate does not consume its own
confirmation time while it is still applying. Durable rollback authority is
published first; its stored deadline is derived at that pre-apply publication
and may already be past when a slow Apply returns, while config status
reports the live post-commit deadline. Boot recovery ignores the authority
deadline and reverts unconditionally.
The v3 commit-confirm journal caps the current accepted normalized config it
must retain as rollback authority at 384 MiB. If that prior exceeds the cap,
confirmed apply returns FAILED_PRECONDITION with the actual and limit byte
counts before publishing authority or mutating peer, persisted, or runtime
state. Apply without --confirm-id, or reduce the canonical config size.
While a confirmed transaction is applying or awaiting confirmation, SIGHUP
reload is rejected at preflight with no effect (SIGHUP outcome
rejected_no_effect) and every persisted runtime config mutator (FIB-table and
dynamic-neighbor CRUD, neighbor lifecycle, a further config apply) is rejected
with FAILED_PRECONDITION, so a pending timeout rollback cannot be overwritten
by a later ad hoc change. The fence clears only after the transaction is
terminal, including durable locator removal and parent-directory sync. If that
authority removal fails after a confirm or otherwise successful rollback, the
operation remains nonterminal and config mutations stay blocked; failure only
in the later exact metadata/raw cleanup or pending-directory fsync is
warning-only.
rbgp config status reports the lifecycle outcome: pending (timer running),
confirmed, aborted, auto_reverted (timer expired and the pre-commit
snapshot was re-applied), or one of the two rollback-failure states,
auto_revert_failed / abort_failed — these mean the daemon could not
re-apply the pre-commit snapshot. The transaction then stays pending and the
mutation fence stays closed (the revert journal is retained, so a mutation
accepted on top of the inconsistency would be clobbered by a boot revert);
resolve it by retrying the abort, confirming the candidate, or restarting the
daemon to boot-revert. Session churn is not one of the causes: abort and timer
rollback restore the recorded pre-commit snapshot without checking a runtime
snapshot token, so peers flapping, dynamic peers arriving or expiring, or an
operator disabling a neighbor inside the window cannot make the rollback fail.
(A caller's own Plan→Apply token does go stale on those events; re-plan.) A
rollback failure therefore points at the apply path itself — persistence
unavailable, or the prior snapshot no longer committable — and the abort error
and the daemon log carry the reason. A confirmed apply that itself fails without proof of a
terminal outcome (lost persistence acknowledgement, post-persist finalization
failure, or compound rollback failure) likewise retains the journal and blocks
all config mutations until a restart boot-reverts.
The timer rollback runs under the same runtime-config coordinator as every
other config mutation, so it waits for the current owner (a long apply, a
reload, or a runtime-config read) to finish. It never gives up that wait: the
status stays pending, and once the rollback is ten minutes past its deadline
the daemon logs a warning (repeated every ten minutes) and rbgp config status
says the automatic rollback is waiting for the coordinator, with the deadline it
missed. The rollback runs as soon as the owner releases and needs no operator
action; a confirm or abort issued meanwhile queues behind it for the same owner
and may time out with "coordinator busy".
Commit-confirmed also survives a daemon restart or crash inside the confirm
window. Before the candidate commits, the v3 writer publishes three objects in
order: the exact accepted normalized TOML at
<runtime_state_dir>/commit-confirm-v3-prior.toml, provenance and fixed-file
identity metadata at
<runtime_state_dir>/commit-confirm-v3-metadata.json, then the sole pending
boot authority at
<absolute lexical config path>.commit-confirm-locator.json. Each object uses
write-temp + fsync + rename + parent-directory fsync; raw prior durability
precedes metadata durability, and both precede the locator. A publication
failure refuses the confirmed apply before candidate persistence unless the
locator rename may already be durable, in which case mutations stay fenced
until restart.
Startup checks the config-adjacent locator before opening or parsing candidate
contents. It may inspect launch-target metadata while pinning and binding the
authority. If the locator is present, the daemon verifies it, the complete
metadata and raw prior, config-target binding, retained TOML, external sources,
file identity, and digests,
then directly adopts that accepted snapshot for boot restore. The revert is
unconditional, regardless of how much confirm time remained, because the
operator's confirming session died with the old process (the NETCONF RFC 6241
§8.4 rule: session loss cancels a confirmed commit). The unconfirmed candidate
is saved once as <recorded-target>.unconfirmed; the loud boot banner names
the transaction but redacts locator-carried paths and digests. A damaged,
unsafe, oversized, or
mismatched locator, metadata, or raw prior refuses boot before candidate
contents are opened and before candidate or backup mutation.
Abort, timeout, and boot restore durably restore the verified prior config
while all three objects remain. They then unlink the locator and fsync its
parent; only that durable locator absence is terminal. Confirm performs no
restore and starts by unlinking and syncing the locator. Exact metadata
removal, raw-prior removal, and pending-directory fsync follow all terminal
paths; failure there is a warning and locator-free residue cannot re-arm the
transaction. A later confirmed apply removes only verified residue at those
two fixed v3 names.
Production reads and writes v3 only. Retired authority makes v0.65.0 and every later release refuse boot untouched. Finish before upgrading past v0.64.0, recover with rustbgpd v0.64.0, or delete only after proving it terminal/intended. Never delete a live v3 locator.
The boot-revert save-aside moves the unconfirmed candidate to
<config>.unconfirmed with an atomic hard-link + unlink, so it never clobbers
an existing .unconfirmed. That needs the config file and its directory to be
on a filesystem that supports hard links — every local filesystem does. On a
filesystem without hard-link support (some FUSE mounts, certain NFS setups),
the save-aside fails and the daemon refuses to boot with the journal left
intact, rather than risk an unsafe revert. Keep the config and
runtime_state_dir on a local filesystem (the default); this mainly matters
for containerized deployments that bind-mount the config from an exotic
backend.
The lexical config, metadata, and raw-prior paths are resolved to absolute
identities. A writer or present pending object requires daemon-owned real
parents that are not group- or world-writable. Locator absence carries no
authority and therefore does not impose that storage policy on an ordinary
launch path; confirmed writes remain unavailable until the config directory is
private, writable, and daemon-owned. Locator, metadata, raw-prior, and staging
files must be regular files owned by the daemon UID with mode 0600; symlinks
and special files fail closed. Keep writable state private and on local
storage.
The transactional quartet: check, compare, commit confirmed, rollback
For operators coming from Junos, the four verbs map directly:
| Junos | rustbgpd |
|---|---|
commit check | rbgp config plan (or rustbgpd --check offline) |
show | compare | rbgp config diff (with reload-impact annotations) |
commit confirmed | rbgp config apply --confirm-id ... --confirm-timeout ... |
rollback N | rbgp config rollback N |
The daemon retains a bounded, best-effort history of up to 20 recognized rows
under <runtime_state_dir>/config-history/. Payload-bearing v2 rows and
metadata-only v3 rows share that count and sequence. A new v2 row deduplicates
only when the newest mixed-history row is verified v2 with matching TOML and
manifest. Unreadable/duplicate rows count; retired TOML is ignored and retained. Transactions and gRPC CRUD record after writes. Boot and
successful SIGHUP reload also record the canonical validated snapshot on a
best-effort basis; SIGHUP does not rewrite the operator's file. History
survives restarts:
# What is retained? Newest first; index 0 is the newest config-history row.
rbgp config history
rbgp -j config history
# Preview an eligible row selected from the listing, then restore it
HISTORY_INDEX=3
rbgp config diff --history "$HISTORY_INDEX"
rbgp --json config diff --history "$HISTORY_INDEX"
rbgp config rollback "$HISTORY_INDEX"
# A cautious rollback: auto-reverts the rollback itself unless confirmed
rbgp config rollback "$HISTORY_INDEX" --confirm-id undo-1 --confirm-timeout 120
rbgp config confirm undo-1config diff --history N previews a retained rollback without changing runtime,
persisted config, history, or confirmed-commit state. Choose either a candidate
file or --history, with an index of 1 or higher. The preview uses the rollback
source-provenance checks and prints the same redacted plan, section support,
reload impact, and update-group projection as config plan; retained secrets
are never exported. Missing, unreadable, metadata-only, or source-mismatched
rows fail with a nonzero exit status.
The preview describes the state observed while planning. It does not reserve
the numeric history index or authorize a later rollback; new history rows can
shift indexes, and rollback checks the selected row again. If a config mutation
is in progress when the planner receives the preview, it returns UNAVAILABLE;
retry after the mutation settles. The dedicated
PreviewConfigRollback RPC is outside the v1 contract and available at
sensitive_read tier. Older daemons return UNIMPLEMENTED; the CLI reports an
error without calling a mutation RPC.
config history lists index, timestamp, normalized-TOML content hash,
provenance status, and—when recorded—a config-source hash over that TOML digest
plus the canonical accepted rpol/dataset source roster. It also shows a one-line
summary per entry, never config document contents. The listing also exposes
normalized byte count and a metadata-only reason. Valid v2 rows are recorded;
corrupt or duplicate-sequence rows are unreadable with both digests withheld.
Accepted normalized TOML strictly above 10 MiB produces a metadata_only v3
row, explicitly marked metadata-only and rollback-ineligible in human output
and with rollback_eligible: false in JSON. Exactly 10 MiB remains v2. A v3
row contains only hashes, byte count, recording identity, and a redacted summary;
it never retains normalized TOML or the external-source roster. It is capped
at 64 KiB, with a 4 KiB summary limit. Only an identical verified newest v3
row deduplicates a new metadata record, comparing both hashes and byte count.
An intervening v2 or unreadable row prevents that deduplication. New large rows
can evict old rollback-capable rows; successful eviction counts appear in the
daemon log. Recording failure warns and leaves the accepted config authoritative.
A directory containing over twenty recognized finals fails listing before any
payload decode; a subsequent successful writer can repair the old roster.
Metadata-only rows always refuse rollback before payload or external-source access, planning, confirmed-commit authority, or mutation. Their hashes cannot restore omitted bytes from live files or commit-confirm cleanup residue. Keep an independent deployment copy of large configurations and their source data.
Index 0 means the newest config-history row, not
necessarily the running or currently persisted config. config rollback N
accepts v2 only when one load reproduces TOML/manifest/source digests;
unreadable/mismatched rows fail closed before mutation. An eligible
row is routed through the same transaction path as
config apply: the same plan classification, the same reload-impact and
update-group annotations, and the same receipts and exit codes. Confirmed rollback
remains subject to that planner: a full-snapshot rollback with external policy
inputs is admitted only when the verified history row's .rpol/dataset identity is
byte-identical to the currently accepted one, and an unchanged-external-input
pure-[[fib_tables]] rollback can carry the provenance-verified v2 history
row's exact accepted prior snapshot through v3 authority. Rolling back across
an external-content change stays rejected — reload the sources first.
There is no second apply engine —
a rollback whose entry contains sections the transaction executor cannot
commit live (for example restart-required [global] fields) is rejected
without mutation, exactly like an apply of that file would be. Rolling back
past the retained history fails cleanly, naming how many entries exist.
Each retained row includes the normalized-TOML digest. V2 rows also hash the
canonical accepted .rpol and dataset source roster, lengths, and content
digests, but do not archive those external bytes. V2 rollback performs one
detached load and requires the live sources to reproduce that identity exactly.
Keep the operator-authored TOML, .rpol graph, and datasets under version control or
another coordinated deployment system.
With .rpol or [policy.datasets] files declared, native config
plan/apply/rollback (commit-confirm included) work as long as the external
sources on disk are byte-identical to the accepted ones: the planner captures
every declared external file at plan and apply time and compares its identity
(canonical path, length, content digest — ADR-0130) against the accepted
snapshot before admitting any full-candidate change. Any drift — an edited
dataset, a touched .rpol module, even a comment-only rewrite — rejects the
transaction without mutation, and a missing or unreadable file fails the
candidate load. Changed external content still deploys the old way: update the
TOML and external inputs together and send SIGHUP. gNMI Set is stricter — it
never verifies external inputs, so a gNMI full-candidate change is rejected
whenever they are present, even unchanged. External-input no-ops still return
NOOP; a pure [[fib_tables]] transaction with unchanged external inputs
remains safe and committable because it stages only the FIB table set.
For a static-neighbor edit, change the neighbor in the candidate file (for
example hold_time, max_prefixes, policy-chain refs, or ORF receive), run
config plan, then apply with the returned token. The transaction stages the
candidate on disk first, so a config directory the daemon cannot write refuses
the change before the peer is touched; it then reconfigures that peer using the
same delete/re-add semantics as SIGHUP and rolls back if apply or publication
fails.
The live API returns redacted text / JSON diff buckets and never exports the
daemon's full config snapshot. Transaction apply is intentionally narrower than
SIGHUP: v1 commits one pure runtime family at a time ([[fib_tables]],
[[dynamic_neighbors]], static [[neighbors]] add/delete/modify, or
catalog-only policy/neighbor-set/peer-group/global-chain edits). It can also
commit policy/neighbor-set/peer-group/global-chain edits that only move existing
static neighbors' or accepted dynamic peers' resolved import/export policy
chains; the executor re-applies those chains to live sessions and rolls them
back if apply or persistence fails.
Because import-chain changes are re-evaluated with Route Refresh, every impacted
Established peer must have negotiated Route Refresh or the transaction is
rejected without committing the candidate.
Peer-group/session reshape transactions can also rebuild affected sessions
(for example, a peer-group hold_time edit inherited by existing members):
static neighbors are reconfigured in place with captured prior peer configs and
rollback, and live dynamic sessions accepted by an affected
[[dynamic_neighbors]] range are gracefully reset after persist so they
re-accept under the committed config on reconnect (ADR-0086). Mixed-family
candidates, dynamic-range peer-group reassignments, mixed policy/session
effective impact, and unsupported sections are rejected without mutation.
Output is grouped into two actionable sections plus a per-neighbor
effective-impact view:
- Reload-applied changes —
[[neighbors]]deltas, neighbor sets, named policies, peer groups, global / per-neighbor policy chains, and the hot-applied[global]flagshonor_graceful_shutdownand control-plane-onlyhonor_blackhole. SIGHUP reconciles all of these. - Restart-required changes —
[global]ASN/router-id/cluster-id,[global.telemetry.grpc_*]listener config (including TLS / mTLS),[rpki],[bmp],[mrt], unsupported EVPN shapes, andapply_bum_enforcement. Supported EVPN edits are shape-aware and appear under Reload-applied. - Effectively impacted neighbors (via inheritance) — every
neighbor whose resolved import / export chain would move at reload,
with the upstream change(s) responsible (peer-group / policy /
neighbor-set / global chain). The JSON form carries
kindas"policy_chain"for pure live-policy impact or"session_reshape"when inherited peer-group/session state changes. Catches transitive references: a policy definition edit picked up via the globalimport_chain(chain list itself unchanged) or via a peer-group's chain (peer-group record unchanged) still flags every affected member. Direct.rpolterm edits also append reasons such asimport rpol policy "customer-in(200)" term "customer-routes" changedfor static neighbors and dynamic ranges, in both text and JSON output. These name changed compiled terms when policy calls, ordered term names, defaults, and shared match tables still align. Compiler-generated.1/.2term suffixes are preserved. Changed sets, call arguments, term order, or other ambiguous structural edits retain the broader reasons; this is structural attribution, not a proof of changed route outcomes.
Exit codes: rustbgpd --diff returns 0 = no actionable changes,
1 = actionable changes found, 2 = error (bad config, missing file).
rbgp config diff <candidate> uses 0 = no changes, 2 = changes present, 1 = error.
rbgp config plan and rbgp config diff --history N use 0 = noop,
2 = committable, 3 = rejected, 1 = error.
rbgp config apply and rbgp config rollback use 0 = committed or noop,
3 = rejected, 1 = error. A rejected transaction still prints its full receipt,
including --json output, and changes nothing; the status field tells a
noop from a commit.
Configuration reload (SIGHUP)
sudo systemctl reload rustbgpd
# or: kill -HUP $(pidof rustbgpd)What happens:
- The daemon re-reads the TOML config file and its
.rpoland dataset files from disk, pins restart-required fields to the live values, and classifies the candidate before any credential, listener, session, or catalog effect.rustbgpd --diffandrbgp config diffprint the route for changes visible to the config diff asSIGHUP reload route. The diff reports dataset bindings; SIGHUP also checks staged dataset contents and loader errors when choosing the actual route. - Generation route — a candidate whose reload-applied changes are
static
[[neighbors]],[peer_groups], BFD member attachments, inline policy definitions, neighbor sets, global chains, changed.rpolcontent (imports included), dataset contents, dataset bindings (added, removed, or re-mapped[policy.datasets]entries), or outbound prefix maxima settles as one owned runtime generation. The daemon resolves the complete candidate once and derives one action per static neighbor — unchanged, hot update in place, replace (one delete/re-add with the final policies), add, or remove — while the peer manager resolves every live peer's final chains against the same candidate and applies the changed ones through the rollback-capable policy snapshot. A group reshape plus an explicit edit of one member therefore rebuilds that session exactly once; policy-only and bystander sessions keep their identity. Removed definitions are simply absent from the adopted candidate. The prior config, compiled.rpolregistry, resolved chains, dataset snapshots and loader errors, and captured session configs are retained through the operation: a later failure restores them from memory, the reload reports a clean rejection, the candidate file stays on disk for correction, and an identical retry re-derives the same plan. Dataset contents move through the stable live handles; an added or removed binding rides in the candidate config, so a route-server member join or leave with its own datasets applies on SIGHUP and a late failure restores the prior binding set with the prior config. A newly declared dataset must load cleanly or the candidate is rejected at load.[policy.explain]is carried with such a candidate; an explain-only change stays sequential. - Sequential route — a candidate with no generation-class change
(
[[dynamic_neighbors]], EVPN runtime tables,[[fib_tables]], the honor knobs, TCP-AO rotation, listener MD5/GTSM inventory, explain-only,[gnmi_dialout]) runs the existing per-subsystem steps. A generation-class change combined with a TCP-AO rotation or an in-place listener MD5/GTSM edit also stays on this path when dataset contents, dataset bindings, and BFD attachments are unchanged. An in-place edit changes the password or GTSM setting of a neighbor that stays configured, or of a dynamic range. The daemon logs that the change runs without generation compensation, because the rotation is its own ordered protocol and the session reshape primitive refuses authentication changes. A static neighbor joining or leaving with its ownmd5_passwordorttl_securitytakes the generation route. Its listener entry is installed before the session is added, withdrawn after the session is removed, and restored with the prior generation if the reload fails. - Rejected changes — dataset content/binding or BFD attachment changes
combined with TCP-AO rotation or an in-place listener MD5/GTSM edit reject before
any effect. A generation-class or dataset change combined with
[[dynamic_neighbors]], EVPN runtime tables,[[fib_tables]], orhonor_graceful_shutdown/honor_blackholealso rejects: those families do not retain and restore priors. Apply independently reloadable families in separate reloads. - Automatic Route Refresh on import-policy hot-apply — when a
peer's effective import chain changes (whether triggered by a
SIGHUP reload or a gRPC mutation), the peer manager issues
soft_reset_in(gated on Established) so routes already inAdjRibInget re-evaluated against the new policy. Operators no longer need to runsoftresetmanually after a chain swap.
Outside configuration transactions, a pending neighbor inventory or detail snapshot allows a bounded batch of policy-statistics reads to proceed while session-state replies are outstanding. Additional neighbor snapshots wait their turn; peer mutations wait for the active snapshot to complete or be canceled. The existing request deadlines still apply under sustained read load.
During a forward generation's export-destination prestaging, neighbor inventory
and policy-statistics reads can inspect the installed generation while the RIB
prepares the candidate destination. These reads retain their existing deadlines.
Once session policy application starts, operator reads wait for that
per-session step to finish; while the daemon then awaits the batched RIB
transition for the cohort, they are served again: the peer manager answers
session snapshots and import-statistics collections (each cohort session
already runs its new chains, so a per-session row may show either generation,
exactly as during prestaging), and the RIB answers general queries between its
pre-commit transition polls from the pre-commit state. Reads that arrive during
the RIB's short commit batches wait for the commit, which remains the single
switch point; a route listing started before the commit cannot be continued
across it. Forward API policy transactions and their policy-only
publication-failure compensation use these same admission points.
TestPolicy also uses the operator lane for its live peer context; its route
pages retain their version fence, without a shared generation pin across
peer context and routes.
When a reload's export-policy transition is rejected and rolled back, the peer manager serves session snapshots, import-statistics collections, and dataset status while it awaits enqueue or completion of the batched RIB restore. These are live reads: a failed session restore can leave different sessions on different policies, and peer-manager and RIB reads do not share a generation pin.
The operator-read design also serves
bounded batches between completed serial steps. Individual session
acknowledgements remain fenced until corresponding manager bookkeeping is
complete. SIGHUP compensation remains fenced after an earlier ambiguous
restoration, such as
session knobs restored before a failed RIB refresh leaves manager metadata
unchanged. Successful earlier steps permit live reads during later policy and
dataset restoration; compensation alone does not imply a full fence. During
API publication compensation, chains restore before staged configuration;
reads report the current values of each source. Honor-only SIGHUP changes to
honor_graceful_shutdown and honor_blackhole use the same admission after
each acknowledged import-chain update. Candidate classification keeps those
setters outside policy/dataset generation unwind; hot-knob restoration keeps
its existing fences. Desired configuration publication and ordinary mutations
retain their command ordering.
Forward RFC 8212 preflight, cohort selection, post-application state probes and their clean-state retries, retained-route proofs and per-peer RIB export replacement admit bounded reads. Dataset settlement inherits its owner's admission through capacity and reply waits. Legacy dataset refresh retains its absolute five-second reply deadline; generation settlement retains its existing reply budget that excludes read servicing.
For synchronous RIB policy replacement and export-only dataset reevaluation, the daemon captures numeric export statistics and neighbor RIB summaries once before mutation. Bounded checkpoints serve those frozen values while ordinary RIB queries and mutations stay fenced. The RIB uses one outer executor handoff on the daemon's multi-thread runtime so independent RPC tasks can consume their replies. This is a temporary observation, not an atomic fleet snapshot or persistent cache. Capture cost and session collection can still exhaust a read deadline.
The shared RIB backlog drain processes one existing route chunk or primary update, serves bounded reads where the transaction permits, then yields. Timer, destination-prestage and deferred-registration callers preserve route-before-EoR ordering while allowing reads between units. This removes a full-backlog drain fence; individual unit cost and sustained input can still delay a response or destination preparation.
gNMI neighbor snapshots use the same live operator lane on listeners and dial-out, retaining their two-second budget and snapshot failure behavior. Periodic BMP sampling overlaps its independent session and RIB input waits; each retains its 100 ms bound, including RIB admission and reply. Missing inputs are omitted and output remains nonblocking, with existing drop counts.
Readiness retains its live actor checks and existing transaction seams. Operator deadlines, complete-result requirements, and the clean transition's atomic commit remain unchanged. The paired native rollback receipt records successful baseline and candidate reads, including candidate calls inside restore. It does not establish a baseline deadline failure. For release-level operating evidence, see the flagship soak receipts.
The serial work still scales with fleet size: 1,000 individual 100 ms state probes have a theoretical 100-second ceiling, and 1,000 individual 500 ms hot-apply acknowledgements have a 500-second ceiling per changed direction, before outer ownership limits. These are component bounds, not measured healthy-fleet timings; read service between completed steps avoids one uninterrupted fleet-sized fence.
For a dataset content generation, every file must load successfully before
publication. The daemon retains prior snapshots and loader errors, reserves
generation numbers for both publication and compensation, settles policy
changes, then publishes the prepared batch without yielding. A changed dataset
advances from g to g + 1; compensation restores its prior contents as
g + 2, so generations never move backward. Content-equal reloads do not
advance the dataset generation. Evaluation still pins each dataset separately;
this does not provide an atomic snapshot across multiple datasets.
The generation refreshes the union of peers whose old or candidate chains reference changed datasets, including dynamic peers. Export recomputation preserves installed chain instances and counters. Established import dependents must support Route Refresh. A positively down peer needs no import replay when the RIB confirms it retains no GR/LLGR routes; any remaining export registration is still recomputed. Unknown session state, lost acknowledgement, or unresolved required refresh work cannot produce a successful settlement. Import refresh acknowledgements prove local dispatch, not completion of the remote peer's replay.
gRPC credentials rotate only after the runtime generation is acknowledged, so a rejected or restored candidate has no credential effect. The generation route claims neither atomic wire visibility nor zero session resets: a replaced session is reset once, and restoration after a late failure re-applies the prior session configs through the same primitives.
Where settlement-owned policy work reaches a clean-session-state fence, the state is a fail-closed proof rather than a best-effort hint. The ordinary query is 100 ms. Only a typed timeout is retried, at most once, and all such retries share one absolute two-second window for the operation. A positive non-Established reply, a departed session task, or retry exhaustion rolls the policy change back. RIB compensation is registered as one reverse-order exact batch before rollback Route Refresh work is issued and may use the normal two-minute batch-reply bound. A local timeout does not cancel the queued repair: the late receiver remains daemon-owned and retry flags stay armed until a later operation proves convergence.
On the sequential route, reload halts at the first step failure. A typed
authoritative receipt records the exact successful peer effects, and the
peer-manager snapshot, config bridge/persister, tracing projection, and gNMI
dial-out targets all adopt that same complete or known-partial runtime state
before the owner settles. The operator can then fix the failing TOML and
reload again to converge. The generation route never produces a known-partial
receipt: it settles the candidate or restores the prior generation. A command
that was definitely not accepted before any prior effect is a clean rejection;
an accepted reply loss or non-authoritative reconcile instead recovery-fences
the daemon, makes /readyz fail, rejects another persisted mutation, and exits
70 after the five-second grace. Settlement-facing policy failures use closed
policy_* codes, never raw session, RIB, config, path, or credential-bearing
error text. Per-step operational detail remains in structured
bucket / target / error logs.
bgp_sighup_reload_outcomes_total{outcome} is process-global and
preinitializes five bounded outcomes: complete includes a clean no-op;
known_partial means the retained task settled with an authoritative partial
receipt; rejected_no_effect means it rejected before any runtime effect or
the generation route restored the prior generation after a later failure (a
replaced session may already have reset once);
ignored_in_flight counts each concurrent signal dropped while another reload
owns the lane; and task_failed means the retained Tokio task itself failed.
Recovery-fenced ownership deliberately has no outcome row because it remains
parked. Diagnose that state with bgp_runtime_config_settlement_active,
bgp_runtime_config_settlement_elapsed_seconds,
bgp_runtime_config_settlement_budget_seconds, and
bgp_runtime_config_settlement_fail_stops_total.
Coordinated shutdown closes new SIGHUP and runtime-mutation admission first.
It does not abort an already-owned reload: the daemon waits for its typed
settlement before taking the warm checkpoint or tearing down required actors.
That wait belongs to the settlement watchdog alone. A coordinator permit that
no settlement owner holds (a holder that has taken the permit but has not
registered its owner yet) and the join of a SIGHUP task with no owner each get
five seconds; on expiry the daemon logs runtime config coordinator permit is still held outside settlement ownership or SIGHUP reload task is still running with no settlement owner at ERROR, skips the optional warm checkpoint
because its coordinator fence is missing, and continues teardown. The exit
status is unchanged.
A further SIGINT or SIGTERM after coordinated shutdown has begun (including
the first signal after a Shutdown RPC) means "stop waiting". With no
settlement owner the daemon logs termination signal received during coordinated shutdown at WARN and skips every remaining wait that has no
deadline of its own: the unowned coordinator permit and SIGHUP join above, the
EVPN IMET withdrawal sweep, the peer-manager drain (peers may then see the
session drop without a Cease), the BMP shutdown enqueue, and the RIB event
conversion stage. Cleanup that carries its own deadline still runs, so kernel
routes, FDB entries, and BFD sessions are still withdrawn, and the exit status
is unchanged.
With a settlement owner still registered, the same signal accelerates the
settlement watchdog's fail-stop instead
of skipping it: the owner is fenced as operator_forced, the daemon logs the
signal line at ERROR followed by the usual runtime config settlement fail-stop armed diagnostic, and the process exits 70 after the five-second
grace with the owner's journal and pending transaction left on disk for the
next start. That is not a completed rollback and not a graceful teardown;
peers get no Cease and only the owner's own journaling is recorded. An owner
that settles before the signal is left settled and shutdown continues
normally; an operation that registers after the signal is refused by the
shutdown gate, never fenced. SIGKILL is no longer needed to end a stop that
is waiting on a mutation, and it still leaves the same durable state behind.
Restart-required surfaces (logged at reload, surfaced under
"Restart-required" in --diff): [global] ASN/router-id/cluster-id,
the RFC 8212 config_epoch / ebgp_requires_policy tuple,
[global.telemetry.grpc_tcp] and [global.telemetry.grpc_uds]
listener config (including any TLS / mTLS field), [rpki], [bmp],
[mrt], [flowspec], [event_history], [inbound_admission],
[security.grpc], [managed_netdevs], [[bfd_profiles]] definitions, and
apply_bum_enforcement. The reload matrix is the full
per-field list. EVPN table edits are
coordinator-gated rather than blanket restart-required: SIGHUP uses the
same daemon actor converger as EvpnService.ApplyEvpnRuntime for supported
L2VNI/IP-VRF/ES shapes, additive build-up, atomic tenant teardown,
ip_vrf relink, and L2VNI-only mixed compositions. Static
rustbgpd --diff is shape-aware: supported EVPN edits appear under
Reload-applied, while unsupported mixed edits or L3VNI/device/table IP-VRF
identity changes remain restart-required or rejected. Missing EVPN actors
or actor convergence failure are runtime outcomes; those pin back to the
committed runtime model and are logged. apply_bum_enforcement remains
restart-required because it is a Gate 8b dataplane actor startup flag.
[[fib_tables]] is the exception to those restart-required tables: when the
ADR-0061 FIB reconciler is running (at least one table present at startup),
table add / remove / edit hot-applies on SIGHUP, with the in-memory
snapshot advancing only after the reconciler acks the new desired set. Only
starting the FIB subsystem from an empty config (0→N) still requires a
restart — surfaced under "Restart-required" in --diff as [[fib_tables]] (start FIB from an empty config).
Use rustbgpd --diff to preview changes before reloading; the diff
buckets the changes by Reload-applied / Restart-required and surfaces
a per-neighbor "effective impact" view for transitive references
(policy edit picked up via global import_chain, peer-group's
chain, etc.).
What state persists
| State | Where | When |
|---|---|---|
| Neighbor add/delete/modify via gRPC | Config file (atomic write) | Serialized with SIGHUP reload; the RPC waits for persistence acknowledgement and rolls runtime back if the write is rejected |
| Dynamic-neighbor add/delete via gRPC | Config file (atomic write) | Serialized with SIGHUP reload; the RPC waits for persistence acknowledgement and rolls the matcher back if the write is rejected |
| GR restart marker | <runtime_state_dir>/gr-restart.toml | On coordinated shutdown |
| Optional shutdown warm checkpoint | <runtime_state_dir>/warm-bundle-v1/ | On coordinated shutdown when warm_cache_checkpoint_on_shutdown = true; owner-private post-import-policy Adj-RIB-In snapshot and manifest, never restored on boot |
| General FIB owned-state | <runtime_state_dir>/fib-owned.json | Before a reconcile pass sends route installs or replacements to the kernel, and after each ADR-0061 FIB apply/drain; written with file and directory fsync |
| MRT dump files | [mrt] output_dir | On periodic timer or TriggerMrtDump |
| gRPC UDS socket | <runtime_state_dir>/grpc.sock | Daemon lifetime |
There is no non-persisting mode: the daemon always takes a config path (the
positional argument, /etc/rustbgpd/config.toml by default), and the rows
above that name the config file always write to it.
The GR marker format is versioned. V1 is generationless and wall-clock-only;
v2 adds a required checkpoint generation; v3 adds a complete Linux boot and
time-namespace identity plus an absolute CLOCK_BOOTTIME deadline, with an
optional checkpoint generation. Full v3 protection requires Linux 5.6+ built
with CONFIG_TIME_NS, a readable valid
/proc/sys/kernel/random/boot_id, inspectable /proc/self/ns/time
device/inode, readable valid /proc/self/timens_offsets, and a sampleable
CLOCK_BOOTTIME value, with the overall marker identity and deadline fitting
their serialized representations. A missing or access-restricted procfs input
cannot establish clock-domain continuity; it selects the fallback rather than
proving that no time namespace exists.
Startup trusts the v3 boottime deadline only when the live boot ID, current
time-namespace device/inode, and both boottime-offset components match exactly.
Only that exact-domain path protects the shutdown-to-startup interval from
discontinuous CLOCK_REALTIME steps. Otherwise the daemon uses the marker's
wall deadline, bounded by the current maximum configured restart time. A
forward wall-clock step can therefore shorten or expire that wall fallback,
including the v1/v2 publication fallback. Publication logs a warning when
clock-domain sampling, checked arithmetic, or TOML's signed integer range
forces a complete v1/v2 fallback; partial v3 markers are never published.
Each concurrently running daemon requires its own runtime_state_dir.
Sharing the directory across live daemon processes is unsupported: its restart
marker, optional warm bundle, FIB ownership receipt, and Unix socket all assume
one writer. At startup, an enabled warm checkpoint runs a bounded cleanup of
canonical orphan snapshots and interrupted-write temporary files. A valid
byte-stable manifest protects its selected snapshot; an absent manifest allows
orphan cleanup; and an invalid or changed manifest deletes nothing. Cleanup
errors are warnings and do not disable the next coordinated-shutdown
publication attempt.
Warm checkpoint manifests now use format version 2, which requires corrected
RFC 8050 Add-Path encoding. Version-1 manifests are rejected before snapshot
decoding, including bundles without Add-Path. Regenerate them through a new
coordinated shutdown; the warm-bundle-v1 directory name remains unchanged.
An old manifest can make startup cleanup skip work until the next successful
checkpoint replaces it.
Not restored: routing state, policy evaluation state, RPKI VRP tables, and BMP client state. The optional warm checkpoint persists only eligible post-import-policy Adj-RIB-In views as a future-use artifact; the daemon does not load, select, install, or advertise any route from it. Loc-RIB and Adj-RIB-Out are never checkpointed. The ADR-0061 FIB file is only an ownership receipt for rows rustbgpd already installed; all route selection state is rebuilt from peers after restart.
Runtime changes rewrite the config file
Every persisted runtime mutation — rbgp neighbor add/delete,
rbgp dynamic-neighbor add/delete, rbgp fib-table set/delete, policy and
peer-group RPCs, rbgp config apply, rbgp config rollback — serializes the
daemon's whole config snapshot and replaces the file with it. On the first
such change, expect all of the following:
- comments are gone, and formatting and key order are re-derived;
- fields inside the sections the file keeps are written out with their
resolved values, but optional feature sections still at their defaults
(
[security], an all-default[policy],[flowspec],[managed_netdevs],[event_history],[inbound_admission]) are omitted, and selected default-empty inline-policy community lists are omitted because omission and[]decode identically; - the temp-file + rename leaves the file owned by the daemon user at mode
0600.
The rewritten file carries a header saying so. This is what makes the write atomic and crash-safe, and it is the intended behavior — a persisted mutation never patches your text in place.
Keep the annotated config under version control and treat the daemon's file as generated output. A deployment that only ever edits the file and sends SIGHUP keeps its comments; SIGHUP alone never writes the file.
Upgrading
For native .deb and .rpm installs, follow the canonical
package upgrade and rollback procedure.
It preflights the candidate binary against the live configuration before
the stop, stops rustbgpd cleanly so the restart marker is written, checks the
package manager's preserved-config files, and verifies the restarted daemon.
The package hooks do not stop or restart a running service.
Native TLS expiry warnings remain off on upgrade unless
tls_expiry_warning_seconds is positive. Enabling the setting requires a
restart and makes matching expiry warnings fail --check --strict; ordinary
--check still returns 0. Before rolling back to a binary without this field,
remove it from the configuration, including a default 0 written by a newer
configuration edit.
Finish any pending confirmed transaction before any upgrade. Before upgrading
past v0.64.0, also clear retired v1/v2 authority — v0.65.0 and every later
release refuse to boot while it is present. The candidate binary's --check
validates the config only; it cannot inspect runtime_state_dir or
config-adjacent commit-confirm authority. If retired authority remains or is
inaccessible, recover it with exactly rustbgpd v0.64.0. Delete it only after
proving the transaction is terminal and the current config is intended.
rbgp doctor --pre-upgrade /etc/rustbgpd/config.toml gathers that live
evidence in one read-only run: it is red, with the next action, while a
confirmed transaction is pending, applying, rollback-failed, or ambiguous;
while a runtime-config settlement owner (a transaction, neighbor or FIB
change, or SIGHUP reload) is still settling; when the named file resolves to
a different RFC 8212 epoch/posture than the live daemon runs; and whenever
the evidence is unavailable, denied, or unimplemented. It never confirms,
aborts, rewrites, or stops anything, and a green result is dated
(Pre-upgrade observation as of unix <t> in text output,
observed_at_unix_seconds in JSON): it is an observation at one instant, not a fence —
a transaction can start after it. Continue with the coordinated stop,
verify the service is inactive, then repeat the candidate --check --strict
and any offline authority checks before installing. See
the check reference.
Moving a config between RFC 8212 epochs is a separate, offline step:
rustbgpd --migrate-config pin-legacy|prepare-secure|downgrade-v0.64 --offline [--dry-run] CONFIG_PATH rewrites the posture representation in place. Stop or
quiesce the daemon first — --offline is an operator assertion, not a daemon
probe — and see CONFIGURATION.md → config_epoch for
what each action writes.
When Graceful Restart is enabled (the default), a coordinated stop writes a GR
restart marker. On the next start, the daemon advertises R=1 to static peers,
asking them to retain eligible routes while sessions rebuild. Configured kernel
installers still advertise forwarding_preserved = false; control-plane-only
families advertise F=1 under the role rules.
This is not a forwarding-continuity guarantee.
Use a drained route-server pair or another traffic-shift procedure when
forwarding continuity matters.
In a matching domain, marker v3 makes the shutdown-to-startup deadline
resistant to discontinuous CLOCK_REALTIME steps and includes suspend time
before the new process resolves it. The daemon then uses its normal
process-local monotonic timer for the remaining live window; it does not claim
suspend-inclusive timing after startup.
Rolling back across the marker-v3 or commit-confirm-v3 boundaries needs the state checks in the canonical procedure. A pre-marker-v3 binary rejects a v3 marker and cold-starts without restarting-speaker mode.
For zero-downtime upgrades in a route-server pair, drain traffic to the standby, upgrade, then swap.
Failure modes
gRPC server dies unexpectedly
The daemon treats an unexpected gRPC server exit as fatal and initiates a coordinated shutdown (NOTIFICATION to all peers, GR marker write). This is deliberate: losing the control plane means losing the ability to shut down cleanly later. See ADR-0022.
A gRPC listener whose socket becomes unusable (listener socket unusable; stopping its accept loop) gives its open connections the same one-second
grace as coordinated shutdown, then exits. An open WatchEvents stream or an
idle client connection cannot hold the listener, or the fail-stop, open.
The other listeners then get one further second to drain before the gRPC
server exits. So the fail-stop starts within about 2 s of the failure.
A TLS listener detects the failure only on its next accept, and it stops
accepting while all 64 concurrent handshake slots are busy. If stalled
clients hold every slot, detection waits for the first handshake to finish
or hit its 10 s timeout, which adds up to 10 s.
RIB manager or peer manager exits unexpectedly
The daemon likewise treats any RIB manager or peer manager return or panic as
fatal. It logs RIB manager exited unexpectedly or peer manager task exited unexpectedly, performs the ordinary coordinated shutdown (including peer
NOTIFICATIONs where the actor that sends them is still alive), and exits 1 for
Restart=on-failure. An intentional shutdown still stops the peer manager as
part of the ordered teardown and exits 0.
Metrics/readiness server exits unexpectedly
When prometheus_addr is configured, the HTTP server behind /metrics,
/readyz and /livez is supervised like the gRPC server. Its accept loop ends
only when the listening socket becomes unusable (listener socket unusable; stopping its accept loop). The daemon then logs metrics/readiness server exited unexpectedly, performs the ordinary coordinated shutdown and exits 1
for Restart=on-failure, rather than running on without its health surface.
RPKI subsystem task exits unexpectedly
The daemon treats an unexpected return or panic from the VRP manager, either
validation-table forwarder, or any configured RTR client task as fatal. It logs
RPKI subsystem task exited unexpectedly, performs the ordinary coordinated
shutdown, and exits 1 for Restart=on-failure. Ordinary cache connection
failures and data expiry continue through the retry behavior below.
The daemon emits exactly one structured fatal receipt. task and outcome
name the primary observed result, preferring a recovered panic
(VrpManager, VrpForwarder, AspaForwarder, or RtrClient);
panic_count is the number of recovered panic results and
resolution_complete is boolean. first_observed records the initial
completion boundary as <Task>:<outcome>; it is context, not a claim of causal
order. After the first exit, the supervisor aborts and accounts for the
remaining tasks under a single one-second deadline.
Multiple recovered panics use task=Multiple and outcome=multiple_panics;
panic_classes is the bounded, sorted set of affected task classes and does
not claim causal order. An incomplete receipt instead reports the
still-unaccounted task classes and count.
RPKI cache unreachable
Each RTR client reconnects independently after a fixed retry_interval
(default 600s). If no fresh EndOfData arrives before the effective expire
(the cache-advertised expire, expire_interval (default 7200s) until one
arrives, both capped by max_expire_interval when set), cached VRPs for that
server are discarded. Routes are re-validated against the remaining VRP table.
When all caches are down, the VRP table is empty and all routes have
validation state NotFound. If your policy denies NotFound routes, this
will cause route drops. The recommended policy is to deny Invalid and
prefer Valid, leaving NotFound as a neutral fallback.
An ordinary loss of the RTR session (a TCP failure, a closed connection, or a
non-fatal Error Report) does not change the validation table until the
effective expire passes, so readiness and the VRP count stay green through the
outage. A flush is different: a fatal Error Report (for example Cache Shutdown)
or a corrupt Prefix PDU drops the cache's contribution at once, and readiness
falls to 0 with the session. A Cache Reset is not a flush; the held data stays
until the full table that follows replaces it. Watch the session itself:
bgp_rpki_cache_connected{cache}is1while the daemon has an RTR session to that configured cache and0otherwise, including at startup and while a retained contribution is still in use.- The shipped
RpkiCacheDisconnectedalert fires after 15 minutes at0. rbgp rpki cachesshows each cache asconnected,retained(session down, contribution still in use),syncing, ordisconnected, with the age in seconds of the last accepted End of Data. A retained contribution drops out once that age reachesbgp_rpki_cache_effective_expire_seconds{cache}, unless a reconnect replaces or flushes it first.
rbgp doctor reads the same inventory from the daemon (ListCaches). The
daemon-side rpki.cache.<addr>.session check is green while the session is
established and warns when it is down, stating whether a contribution is
still retained and its age. That check needs a token that can read RPKI
cache state; without one it warns that the state is unavailable. The
daemon-side rpki.vrp_table check is green only when the merged VRP count is
nonzero and every configured cache reports retained accepted complete
End-of-Data readiness; it names configured caches whose readiness is 0
separately from caches missing from the metric snapshot. Doctor also probes
every configured cache from the rbgp process
(rpki.cache.<addr>.reachable_from_cli); that raw connect reflects only the
CLI's network vantage and warns on failure.
bgp_rpki_cache_end_of_data_ready{cache} is 0 at startup, becomes 1 only
after an accepted, complete, validated End of Data (including an empty table),
and remains 1 while that cache's retained contribution stays usable through
disconnect or resynchronization. It returns to 0 after flush or expiry. This
is cache readiness, not connectivity (bgp_rpki_cache_connected), merged-table
availability, policy matchability, or a startup gate.
BMP collector unreachable
Each BMP client reconnects independently with backoff (default
reconnect_interval = 30s). During disconnection, BMP events for that
collector are dropped. No routing state is affected — BMP is purely
observational. On reconnect, the client sends a fresh Initiation message and
the manager replays cached Peer Up state. A configured Loc-RIB view then gets
a fresh table dump closed by End-of-RIB. Adj-RIB-In and Adj-RIB-Out resume
live updates without an automatic reconnect dump. For established IPv4/IPv6
unicast sessions, the experimental
replay-out operation
can rebuild an eligible collector's outbound inventory by reannouncing routes
on the live BGP session. The CLI confirms scheduling; terminal BMP EoRs are
required for a complete capture.
During coordinated shutdown, the BMP manager first queues final Peer Down messages. Connected clients drain that queue, send BMP Termination, and flush. Disconnected clients stop their active connect or backoff wait, and every BMP client shares one aggregate two-second drain budget; a stalled collector cannot multiply daemon shutdown latency by the number of configured collectors.
gNMI dial-out collector unreachable
Each [gnmi_dialout] target dials its collector independently and
reconnects with capped exponential backoff (backoff_initial, doubling to
backoff_max). A collector that is down — at startup or later — never
affects BGP operation; telemetry for that target is simply not delivered
while disconnected. The daemon logs one warn when a target becomes
unreachable and one info when it (re)connects; repeated retries log at
debug only. Watch gnmi_dialout_connected{target} (0/1) for the live
connection state and gnmi_dialout_resync_total{target} for established
sessions that ended and entered the fresh-snapshot reconnect path; failed dials
do not increment it. A connected target whose gnmi_dialout_queue_depth
stays elevated or whose time() - gnmi_dialout_last_publish_timestamp_seconds
keeps growing is not making healthy publish progress even if its connection
gauge remains 1. The timestamp records a response handed to the local gRPC
transport, not a collector acknowledgement; PublishResponse remains reserved
and no remote-apply acknowledgement exists. Every (re)connection restarts the subscription, so the
collector resyncs from a fresh initial snapshot + sync_response — the
disconnect window is not replayed (same contract as a dial-in Subscribe
reconnect; use the durable event cursor for gap-free history).
MRT dump failure
The output directory is created lazily when a dump is due. If it cannot be
prepared as a directory, the manager fails that dump before requesting a
full-table RIB snapshot. If the directory is not writable, or encode/write
fails later, a periodic dump logs the error and skips that cycle. An on-demand
TriggerMrtDump failure is returned to its caller while the RPC remains
connected; only write failures also emit a manager error log. Periodic dumps
continue on the next interval without replaying missed intervals in a catch-up
burst. The daemon does not crash on MRT failures. Dump health is exported as
the mrt_* metrics (see the MRT table under Metrics below):
mrt_dump_failures_total{stage} names the failing stage and
mrt_last_dump_success_timestamp_seconds stops advancing, which the shipped
MrtDumpStale alert turns into a page once the newest dump is older than twice
dump_interval.
If an on-demand request is already canceled when the RIB actor handles it, the actor skips route materialization. Cancellation observed while the manager awaits the RIB reply prevents encode and file publication. Snapshot cloning is synchronous actor work: a clone already running cannot be interrupted.
Peer max-prefix exceeded
When a peer exceeds max_prefixes, max_prefixes_ipv4, max_prefixes_ipv6,
or the pre-policy max_prefixes_received_ipv4 / max_prefixes_received_ipv6
bounds, the daemon sends a NOTIFICATION and tears down the
session. Without negotiated Notification GR this is Cease/1 (Maximum Number of
Prefixes Reached). With the RFC 8538 N-bit, the daemon sends outer Cease/9
(Hard Reset) whose data encapsulates the same Cease/1 reason and RFC 4486 data,
preventing the over-limit routes from being retained as stale. By default the
peer is not automatically re-enabled — use rbgp neighbor <addr> enable or the
gRPC EnableNeighbor RPC to restart it.
max_prefix_action selects a non-teardown response instead. Under "block",
a net-new prefix that would take a slot beyond a full per-family bound is
withheld — never installed, never delivered to the RIB — while the session
stays Established; attribute changes and new Add-Path identities for prefixes
already accepted still pass. The first withheld prefix opens a blocking episode
(one warn log line, bgp_max_prefix_blocking{peer,scope} = 1, the
inbound_prefix_limits[] row in rbgp neighbor <addr> reports
blocking=inbound_prefix_limit_reached) and increments
bgp_max_prefix_blocked_total once for that episode; prefixes withheld later in
the same episode are not counted again. When usage falls back under the bound (a
withdrawal, an enhanced-refresh sweep, or a raised limit), the bound is
removed, or the action leaves block, the episode ends and the daemon sends
one plain ROUTE-REFRESH per affected family so the peer replays what was
withheld. Without negotiated route refresh, peer reannouncement or a session
reset is required;
a replay that overflows again opens a new episode, so a peer churning at its
bound can trigger repeated full replays — block is a containment mode, not a
steady-state one. Under "warning" the daemon only reports and keeps
accepting. Neither mode latches the peer or sends a NOTIFICATION, and
max_prefix_action = "block" requires the aggregate max_prefixes to be
unset.
For negotiated IPv4/IPv6-unicast Add-Path receive, a nonzero
add_path.receive_max also caps retained path IDs per prefix. It counts
accepted IDs plus rejected IDs only when that family's
max_prefixes_received_* bound enables rejected-identity tracking. Without
that bound, a denied path does not occupy a receive-cap slot; rejected-route
diagnostics remain separately capacity-bounded. An existing ID can be
replaced at the cap, and a withdrawal releases its slot. Under "block", a
net-new over-limit ID is withheld; recovery requires peer reannouncement or
a session reset. Under "shutdown", the peer receives Cease/1 and is latched
off as above. Under "warning", the excess is observed but accepted, so
the cap does not bound memory. bgp_add_path_receive_limit_attempts_total
counts net-new ID admission attempts beyond the cap, including repeated
attempts to send a withheld ID. Replacing an already admitted ID does not
count. The counter is labeled by peer, family and action. The
bgp_max_prefix_* gauges and counters continue to describe
unique-prefix bounds, not this path count. A receive_max config edit
rebuilds the session; a lower cap applies to the replacement session, not as
an in-place trim of previously retained paths.
max_prefix_warning_percent adds a threshold to any action: when a scope's
usage reaches that percentage of its bound the daemon emits one warn log line
(max prefix warning threshold crossed), one max_prefix_warning session event
(visible in rbgp watch and the durable event stream), and one
bgp_max_prefix_warning_total{peer,scope} increment, then stays quiet until
usage falls back under the threshold and crosses it again.
Setting the inheritable, non-zero max_prefix_restart_seconds opts that peer
into exactly one restart attempt after the hold-down. A second breach creates a
new hold-down. Failure to deliver the timed session Start command consumes
that attempt and remains latched off until explicit enable; successful delivery
removes the latch and returns the session to ordinary TCP/OPEN retry.
Peers whose hold-downs expire together share one 500 ms command-delivery window,
so a stalled session cannot multiply the restart delay across the due set.
Start delivery failure replaces last_error with the cause and the exact
rbgp neighbor <addr> enable recovery action.
A live edit of max_prefix_restart_seconds while a countdown is armed
reschedules the single pending attempt to now + the new duration; the
superseded deadline never fires. Removing the duration cancels the countdown,
and adding one to an already-latched peer does not retroactively restart it —
explicit enable remains the recovery path in both cases. Peer-group edits
inherited by dynamic peers follow the same rules. Session-generation
replacement, explicit disable, neighbor removal, and dynamic-range replacement
cancel an existing countdown rather than carrying it into new policy.
Before an explicit enable, inspect the latch:
rbgp neighbor <addr>
rbgp --json neighbor <addr>Human output reports Max-Prefix Action (shutdown or restart), the
configured Max-Prefix Restart, an active Max-Prefix Hold-Down countdown,
and Last Error. The JSON equivalents are max_prefix_action,
max_prefix_restart_seconds, max_prefix_restart_remaining_millis, and
last_error. A configured timer has an armed attempt only while the remaining
field is present: after Start delivery failure consumes its one chance, the
effective action is shutdown, the countdown is absent, and explicit enable is
required. Successful delivery clears the latch and hands recovery to ordinary
TCP/OPEN retry; it does not mean that the session is already Established.
Explicit enable clears the latch, including any armed countdown, and requests
an immediate start (subject to strict BFD withholding).
From Prometheus, bgp_max_prefix_latched{peer,interface} (see the
Health metrics) is 1 for exactly this latch. It returns to 0 on
explicit enable, or when the timed hold-down expires into a successful restart
or, under strict BFD, into the BFD withhold (BGP then starts on BFD Up).
The latched peer also reads bgp_peer_admin_enabled = 0, so the shipped
BgpSessionNotEstablished alert stays silent for it. The shipped
BgpMaxPrefixLimitExceeded alert pages on the breach and resolves 10 minutes
later; BgpMaxPrefixLatched keeps firing until the latch clears. For a
max_prefix_action = "block" episode, the shipped BgpMaxPrefixBlocking alert
warns after bgp_max_prefix_blocking has stayed 1 for 5 minutes, and
rbgp doctor warns on each blocking inbound_prefix_limits[] row.
Neighbor detail also reports the session actor's O(1) aggregate
max-prefix-counted NLRI identity count plus unique IPv4- and IPv6-unicast
prefix counts. Each count is paired with its effective finite limit and
remaining headroom;
an absent limit is rendered as unlimited and remains absent (never zero) in
JSON and gRPC. The aggregate includes all max-prefix-counted NLRI, while the
family counts cover unicast only. If the session query times out, the snapshot
is marked stale: configured limits remain visible, but headroom is withheld
rather than derived from placeholder zero counts.
Prometheus exposes the same live actor authority as
bgp_max_prefix_usage, bgp_max_prefix_limit, and
bgp_max_prefix_headroom, keyed by peer and bounded scope
(aggregate, ipv4_unicast, ipv6_unicast, ipv4_unicast_received, or
ipv6_unicast_received). Usage is the session
actor's enforcement count, not an alias for bgp_rib_prefixes: in particular,
unicast Add-Path IDs collapse to one unique prefix while the aggregate also
includes max-prefix-counted non-unicast identities. Limit and headroom series
exist only for finite configured bounds. The two *_received scopes report
the pre-policy count bounded by max_prefixes_received_* (announced prefixes,
accepted or rejected) and exist — usage included — only while that bound is
configured, because rejected identities are not tracked otherwise. All
capacity families are
removed when the session goes down and republished from fresh actor state on
reconnect; GR-retained RIB rows therefore never appear as live session usage.
SRv6 service route is visible but cannot be selected
An accepted route with a semantically unusable applicable SRv6 Service TLV
remains in Adj-RIB-In, but is excluded from best-path selection and export.
For unicast, rbgp rib --prefix <cidr> --explain reports
srv6_sid_invalid; an exact EVPN selector uses
rbgp evpn explain ip-prefix --rd <rd> --prefix <cidr>. This is distinct from
an import-policy rejection. VPN has no received-route query.
This is not malformed-attribute handling. A malformed recognized Service TLV
is treated as withdrawn; a malformed generic Prefix-SID attribute is discarded
while the route remains. See SRv6 Service framing
and service eligibility
for the canonical rules. These reflection checks do not add SRv6 PE import,
service origination, next-hop rewriting, or forwarding. The optional
reconstructed_sid in VPN and EVPN inspection views is display-only and does
not affect selection.
Key metrics to watch
Metrics and HTTP probes are exposed on the Prometheus endpoint if
prometheus_addr is configured. If omitted, metrics are still collected
internally and available via gRPC GetMetrics and GetHealth RPCs, but the
HTTP /livez and /readyz probes are disabled.
Metric names are split by prefix on purpose: bgp_* covers the BGP core
(sessions, RIB, policy, RPKI/ASPA, GR, event stream/outbox, FIB) and
evpn_* covers the EVPN/VTEP dataplane surface — the two subsystems have
different cardinality profiles and are typically dashboarded and alerted
separately. bfd_* names the BFD liveness metrics; bmp_* and mrt_*
cover the BMP exporter and MRT dump export. A ready-to-load
Prometheus alert-rule pack covering the high-signal bgp_* metrics ships
at examples/prometheus/rustbgpd-alerts.yml.
HTTP probes
The telemetry HTTP listener exposes these read-only paths:
| Path | Success | Failure |
|---|---|---|
/metrics | 200 Prometheus text exposition | 500 if metrics encoding fails; 503 when the scrape's wait for the render exceeds 5 s (the render itself keeps running), or at once when 56 scrapes are already in flight (waiting, rendering or writing). Scrapes alone cannot fill the 64-connection budget, so /livez and /readyz stay reachable while a render is stuck |
/livez | 200 ok once the listener accepts connections | No actor checks. While all 64 connections are in use, the listener closes connections that have not sent their request line within 250 ms of being accepted, oldest first, so idle or slow-sending clients cannot keep /livez and /readyz waiting for a connection |
/readyz | 200 ready when PeerManager and RIB respond within 200 ms total | 503 not ready: <reason> |
/dp-readyz (opt-in alpha) | 200 dataplane workers ready after each configured FIB/EVPN worker has completed an initial attempt and continues making progress | 404 when disabled; 503 when not configured, starting, unavailable, closed, or stale |
Readiness is actor responsiveness, not routing policy. A daemon with zero configured peers, zero Established peers, or zero routes can still be ready. Event-history, EVPN, FIB, and peer-count health are surfaced through their own metrics and status commands rather than as v1 readiness gates.
Enable the separate alpha dataplane probe with
dataplane_readiness = true under [global.telemetry], alongside
prometheus_addr. Both the option and the monitored worker inventory are
startup-only. With no FIB tables, EVPN instances, IP-VRFs or managed netdevs
configured at startup, it returns 503 dataplane not configured. A configured
worker whose netlink setup failed remains in the inventory as unavailable.
Removing all its tables at runtime still monitors that worker's empty passes;
adding a previously absent worker requires a restart.
The probe reads compact worker-owned observations: general FIB, EVPN intent producer, and EVPN kernel reconciler. Each records its own work checkpoints, including within multi-operation passes. Forwarding a cached route report cannot refresh any worker's observation. Startup fails until an initial attempt returns, and task exit closes its observation channel immediately. Individual install failures, unresolved next hops, filtered routes and normal retries do not themselves mean that a worker stopped progressing.
The default observation windows follow the workers' existing cadences:
| Worker | Idle cadence | Stale after no checkpoint for |
|---|---|---|
| General FIB | 30 seconds | 60 seconds (idle cadence plus the 30-second planning budget) |
| EVPN intent producer | 5 seconds | 10 seconds (two poll intervals) |
| EVPN kernel reconciler | 60 seconds | 120 seconds (two resync intervals) |
Stale means no observed worker progress, not proof that a task is dead. A single unbounded dependency wait, kernel dump or monolithic computation longer than its observation window can produce this result. The probe does not cancel that work or change its recovery behavior. Empty idle tables remain healthy across their normal cadence, and successful checkpoints restore a stale worker's verdict. The HTTP request performs no route scan, metrics scrape or netlink operation.
This is worker readiness only: it does not prove packet forwarding, route
convergence, or successful installation of every route. It remains outside
the v1 contract and does not change /readyz, GetHealth, the systemd watchdog,
or the daemon's treatment of optional dataplane failure.
Each /readyz response reports the current probe result; the endpoint has no
internal failure-count hysteresis. The shared 200 ms core-actor deadline is
not a bound on total HTTP wall-clock time: runtime scheduling and response
delivery can take longer, and a successful probe observed after the deadline
is rejected.
Consumers choose their own failure policy. Kubernetes defaults to three failed probes, a 10-second interval, and a one-second timeout; an HTTP 503 still fails that probe even when it arrives within one second. The flagship RS soak instead requires HTTP 200 within 250 ms and fails on three consecutive breaches at its default 30-second sampling interval. Missing observations fail its gate. Isolated breaches remain visible findings but do not automatically block a release when all agreed acceptance gates pass. See the readiness acceptance and Kubernetes comparison for the precise observation rules and the separate RR gate.
The peer label
Every peer-labeled series — session, RIB, policy, max-prefix, BFD, BMP —
identifies the peer the same way: by its bare neighbor address, 192.0.2.1
or 2001:db8::1, never the transport endpoint's addr:port. sum by (peer)
therefore returns one series per peer, and a by (peer) join across any two
families matches. The current administrative and session-state gauges add the
configured interface (empty for unscoped peers), so exact
on (instance, peer, interface) joins distinguish the same IPv6 link-local
address configured on multiple interfaces.
Per-peer series lifecycle
All peer-labeled series are removed when the peer is deleted — a static
neighbor delete (CLI/gRPC/config reload) or a dynamic peer's auto-removal
when its session ends. Prometheus marks the removed series stale at the next
scrape, so deleted peers stop appearing in instant queries instead of
freezing at their last value. A session flap or admin disable does not remove
durable counters or history. The live bgp_max_prefix_usage,
bgp_max_prefix_limit, and bgp_max_prefix_headroom gauges are the deliberate
exception: they are removed while the session is down and republished on
Established. If a deleted peer is later re-added, its per-peer counters restart
from zero — PromQL
rate() / increase() treat that as an ordinary counter reset, so
dashboards see no negative-rate artifacts. Process-global counters and
families keyed by other identities (AFI/SAFI, VRF, VNI, BMP collector) are
never removed.
The exact bgp_peer_admin_enabled{peer,interface},
bgp_max_prefix_latched{peer,interface},
bgp_peer_session_established{peer,interface},
bgp_peer_session_state{peer,interface,state}, and
bgp_peer_info{peer,interface,remote_asn,description,peer_group} identity is
removed after its session actor terminates, including abort-on-timeout
teardown. Its matching bgp_session_down_total{peer,interface,reason} history
is reaped at the same identity boundary. If a scoped sibling shares the bare
address, its exact rows and the shared bare-address history remain.
bgp_peer_info holds exactly one row per exact (peer, interface) identity.
Editing a neighbor's description or peer group, or learning the ASN from OPEN
on an accept-any dynamic range, replaces that row rather than adding a second
one beside it, so group_left joins never match two identities for one peer.
Health
| Metric | What it tells you |
|---|---|
bgp_peer_admin_enabled{peer,interface} | Effective administrative state: 1 enabled, 0 disabled. 0 covers both an operator disable and a max-prefix shutdown latch, even though the configuration still says enabled; bgp_max_prefix_latched separates the two |
bgp_max_prefix_latched{peer,interface} | 1 while a max-prefix shutdown latch holds the peer off, from the breach until an explicit enable, or until the timed hold-down expires into a successful restart or, under strict BFD, into the BFD withhold; 0 otherwise. Seeded and reaped with the other exact peer-identity gauges. The shipped BgpMaxPrefixLatched alert holds on it after the 10-minute BgpMaxPrefixLimitExceeded event alert resolves |
bgp_peer_session_established{peer,interface} | Current active-primary session truth: 1 Established, 0 otherwise |
bgp_peer_session_state{peer,interface,state} | Exact active-primary one-hot FSM state; state is idle, connect, active, open_sent, open_confirm, or established |
bgp_peer_info{peer,interface,remote_asn,description,peer_group} | Configured identity of each exact peer, always 1. Join it onto any per-peer family with * on (instance, peer, interface) group_left(remote_asn, description, peer_group) bgp_peer_info so dashboards and alerts name the member instead of the bare address. remote_asn is the configured ASN, or the ASN learned from OPEN for an accept-any dynamic range (0 until then); description falls back to the neighbor address when none is configured, and dynamic peers without a range description carry dynamic:<peer_group>; peer_group is empty for ungrouped neighbors. description and peer_group are scrubbed — control characters dropped, surrounding whitespace trimmed, and bounded to 128 characters |
bgp_session_established_total | Cumulative sessions that reached Established (per-process counter; resets on restart) |
bgp_session_flaps_total | Cumulative session flaps |
bgp_session_down_total{peer,interface,reason} | Established active-primary sessions that ended. local_notification means a locally initiated NOTIFICATION teardown even if best-effort delivery failed; remote_notification means a received NOTIFICATION; local_no_notification includes forced local close such as send-hold expiry; remote_no_notification means remote TCP close without a NOTIFICATION; the remaining bounded values are transport_error and defensive unknown. |
bgp_session_state_transitions_total | FSM state transitions |
The shipped BgpSessionNotEstablished alert requires effective administrative
state 1 and Established state 0 for the same (instance, peer, interface) for
two minutes. It therefore covers never-established and previously-down peers
without paging on disabled peers. A max-prefix latched peer reads effective
administrative state 0, so it is silent here and covered by
BgpMaxPrefixLatched instead. Flap-rate alerting remains based on
bgp_session_flaps_total. Aggregate Established counts and daemon uptime are
also available via ControlService.GetHealth / rbgp health; that RPC uses
the same 200 ms core-actor deadline as /readyz. rbgp health --liveness
uses the Read-tier ControlService.CheckLiveness instead: it proves only that
an authenticated gRPC handler answered and prints alive (JSON: a
pretty-printed object with "alive": true).
HTTP /livez is the corresponding non-disclosing HTTP alternative when the
metrics listener is configured. Neither liveness path checks actor readiness,
BGP convergence, or forwarding. Successful Read authorization audits are DEBUG;
authorization counters still include every request, and sensitive reads,
mutations, denials, and errors retain their existing audit levels.
For a current state view that also catches a broken one-hot invariant:
sum by (instance, peer, interface) (bgp_peer_session_state) == 1To break recent Established-session losses down by their bounded cause:
sum by (instance, peer, interface, reason) (
increase(bgp_session_down_total[5m])
)Routing
Dynamic-neighbor admission capacity is process-global and deliberately
label-free. The three slot gauges are materialized when PeerManager starts:
used counts accepted dynamic peers that still own a slot, limit mirrors
global.dynamic_neighbor_limit, and headroom is the saturating difference.
Ordinary Idle removal and config-rollback reaping return a slot; a terminal
max-prefix peer retained disabled for explicit recovery still owns one.
bgp_dynamic_neighbor_limit_rejections_total increments only when a matching
inbound dynamic connection is dropped because all slots are occupied.
bgp_inbound_connections_dropped_total breaks accept-path drops down by a
bounded reason vocabulary (ADR-0120): unconfigured (source matched no
static neighbor and no dynamic range), rate_limited (the opt-in
[inbound_admission] per-source token bucket was empty),
dynamic_limit (slot saturation, counted alongside the legacy counter), and
notification_backoff (a configured neighbor dialled in while its session
waited out an escalated NOTIFICATION reconnect wait; see
Debugging a session that won't establish).
There is deliberately no per-source label — source cardinality is unbounded
exactly under the floods these drops account for.
| Metric | What it tells you |
|---|---|
bgp_rib_loc_prefixes{afi_safi} | Loc-RIB size (best paths) per table; afi_safi="all" is IPv4 + IPv6 unicast, not a sum of the other values |
bgp_rib_prefixes{peer,afi_safi} | Adj-RIB-In size per peer (received), by table: afi_safi="all" is IPv4 + IPv6 unicast paths combined, each Add-Path path counted individually; evpn and flowspec are those tables. The synthetic peer 0.0.0.0 reports locally injected unicast and FlowSpec rules, including rules not selected into Loc-RIB. VPN, labeled-unicast, BGP-LS and RTC routes are not counted here |
bgp_rib_adj_out_prefixes{peer,afi_safi} | Adj-RIB-Out size per peer (advertised) per table; afi_safi="all" is IPv4 + IPv6 unicast, not a sum of the other values |
bgp_dynamic_neighbor_slots_used | Dynamic peers currently consuming the process-global admission limit, including retained disabled max-prefix recovery targets |
bgp_dynamic_neighbor_slots_limit | Effective process-global dynamic_neighbor_limit |
bgp_dynamic_neighbor_slots_headroom | Saturating limit - used; zero means the next matching dynamic inbound is rejected |
bgp_dynamic_neighbor_limit_rejections_total | Matching inbound dynamic connections rejected because the slot limit was already full |
bgp_inbound_connections_dropped_total{reason} | Accept-path inbound connection drops by bounded reason: unconfigured, rate_limited (ADR-0120 [inbound_admission]), dynamic_limit, or notification_backoff (a configured neighbor held by its escalated NOTIFICATION reconnect wait) |
bgp_max_prefix_usage{peer,scope} | Live session-actor max-prefix enforcement count for aggregate, ipv4_unicast, ipv6_unicast, ipv4_unicast_received, or ipv6_unicast_received; series are absent while the session is down, and the *_received scopes are absent while their max_prefixes_received_* bound is unset |
bgp_max_prefix_limit{peer,scope} | Effective finite bound for the same scope; absent means unlimited, never zero |
bgp_max_prefix_headroom{peer,scope} | Saturating limit - usage for a finite scope; absent when unlimited or disconnected |
bgp_max_prefix_blocking{peer,scope} | 1 while a max_prefix_action = "block" episode is open for a finite scope — net-new prefixes are being withheld until usage falls back under the bound; present for every finite scope, reaped with the other capacity series |
bgp_add_path_receive_limit_attempts_total{peer,family,action} | Net-new received Add-Path ID admission attempts beyond the per-prefix cap for negotiated ipv4_unicast or ipv6_unicast; a repeated withheld ID counts again, but an admitted-ID replacement does not. action is block, shutdown, or warning. No prefix label or per-prefix gauge. Reaped when the configured peer is deleted |
bgp_outbound_prefix_usage{peer,family} | Distinct prefixes admitted into a peer's ADVERTISED unicast state (ADR-0113), family = ipv4_unicast or ipv6_unicast. Post-policy, post-OTC, post-exact-export — the same truth the neighbor API reports, never the shared update-group table's count. Series reaped on session teardown |
bgp_outbound_prefix_limit{peer,family} | Effective finite max_prefixes_out_* for the same family; absent means unlimited, never zero. A family that becomes unlimited drops this series rather than keeping a stale value |
bgp_outbound_prefix_headroom{peer,family} | Saturating limit - usage; absent while the family is unlimited |
bgp_outbound_prefix_blocking{peer,family} | 1 while a blocking episode is open — the peer's advertised view is intentionally diverging from group intent until capacity recovers |
bgp_outbound_prefix_blocked_total{peer,family} | Net-new prefixes dropped from this peer's outbound vector by its configured maximum. Deliberately carries no prefix label: the blocked set is exactly the unbounded quantity the limit exists to contain — use rbgp rib advertised <addr> (or rbgp rib --prefix <P> advertised <addr> --explain for a single prefix's gates) against the export policy to find what is missing |
bgp_rib_attr_intern_global_size | Unique attribute sets in the daemon-wide cross-peer intern table (attribute-memory dedup across ALL peers). Tracks reclaim sweeps and growth under churn; a monotonic slope under steady-state churn indicates an intern leak. Replaces the per-peer bgp_rib_attr_intern_size{peer} gauge |
bgp_messages_received_total | Inbound BGP messages by type |
bgp_messages_sent_total | Outbound BGP messages by type |
bgp_outbound_route_drops_total{peer} | Outbound BGP work dropped because the peer writer channel was full or closed. Unlike inbound RIB backpressure, this is a loss signal, not safe producer parking. Ordinary route updates, peer-refresh responses, and initial dumps enter dirty-peer resync; prefix-limit recovery remains queued for automatic family-scoped retry. Collision-failback inbound ROUTE-REFRESH requests retain one peer/session-generation-scoped intent and retry on the same bounded cadence, so temporary saturation of that request does not increment this counter; a closed matching session is terminal and does increment it |
bgp_peer_outbound_queue_depth{peer} | Coalesced update frames buffered for a peer's outbound writer — the "which clients are behind" signal during convergence. Sampled at batch granularity (once per enqueue batch and once per writer drain pass, never per message). A value pinned near the writer's bulk-buffer capacity marks a slow or stuck client that is not draining our output; a healthy peer's depth returns to 0 after each burst. At large route-reflector fanout, sort peers by this gauge to find the laggard holding up convergence. Series reaped on session teardown |
bgp_peer_update_group{peer} | Which update group a peer currently belongs to — the "which group is this client in" lookup that bgp_update_group_members{group} (member counts) cannot answer. The value is the numeric group id (matches the group label of bgp_update_group_members), stable while the policy content remains live; reinstalling a fully retired policy gets a new ID; the sentinel -1 marks a peer on the per-peer/ungrouped fallback path (peer-context policy, Add-Path send, per-client-best on a VPN/RTC session, ORR vantage, negotiated ORF, or slow-peer isolation). Refreshed on every membership change; the peer's series is removed when its outbound registration ends, including session-down and graceful-restart teardown |
bgp_peer_slow{peer} | 1 while the peer is flagged slow: Established and alive (keepalives flowing, so the RFC 9687 send-hold teardown never fires) but persistently not draining its outbound queue — backlog at or above slow_peer_threshold_pct of the writer buffer for slow_peer_duration seconds. See "Slow peers" below for interpretation and actions. Refreshed on both transitions and on session teardown; series reaped on peer delete |
bgp_route_refresh_in_progress{peer,afi_safi} | Active inbound Enhanced Route Refresh window for a peer/family (1 = active, 0 = inactive) |
bgp_route_refresh_stale_entries{peer,afi_safi} | Routes still awaiting replacement before EoRR or timeout during an inbound Enhanced Route Refresh window |
bgp_rib_route_refresh_actor_duration_seconds{operation} | Wall-clock RIB-actor time for accepted inbound Enhanced Route Refresh work. The closed operation label is begin (BoRR's full stale snapshot/count/gauge update), eorr (an active EoRR's complete stale sweep, recompute, distribution, and cleanup), or timeout (the same active finisher after expiry). Duplicate accepted BoRRs are counted; stale-session markers and EoRR/timeout markers without an active family are not. Uses the standard RIB actor duration buckets. |
bgp_rib_outbound_registered_peers | Global count of peers currently registered for outbound route distribution. It has no peer label and cannot identify an unregistered Established peer. Confirm a suspected gap with the peer's session state and bgp_peer_update_group{peer}: its series is absent when no outbound membership exists; absence can also persist while initial registration is deferred behind queued imports; there is no fixed maximum delay |
bgp_rib_outbound_registration_replaced_total{peer} | PeerUp re-registrations that replaced a still-registered outbound sender for the same address — two sessions overlapped (collision window); the replacement resets the prior session's RIB state and keeps the superseded session live for failover; its PeerDown is matched by session identity |
bgp_rib_stale_peer_down_ignored_total{peer} | PeerDown/PeerGracefulRestart events discarded because their session id didn't match the registered session — a stale teardown from a superseded collision-loser session; the surviving session's state is untouched |
bgp_rib_stale_session_message_ignored_total{peer,kind} | Session-scoped RIB messages discarded by the same session-identity rule — a superseded session's queued message processed after the replacement's PeerUp. kind labels: routes, bgpls, vpn, labeled, rtc, eor, refresh, orf, policy_context, slow_peer. These identify the combined unicast / FlowSpec / EVPN RoutesReceived envelope, the separate BGP-LS, VPN, labeled-unicast, and RT-Constrain route-message variants, End-of-RIB, route-refresh request / RFC 7313 BoRR/EoRR, RFC 5291 ORF push, peer-group policy identity, and slow-peer state respectively. The registered session's state is untouched |
bgp_rib_outbound_registration_failover_total{peer} | Outbound registrations handed to another live session for the same address after the active session's PeerDown/GR-down (the symmetric collision interleaving: the loser's PeerUp replaced the winner's registration before the loser went down). The exact nonzero, unambiguous survivor enters awaiting_refresh, is re-registered, receives the staged initial table without EoR, and is asked for an inbound ROUTE-REFRESH; matching post-failback BoRR/EoRR completes convergence |
bgp_rib_dirty_resync_total{outcome} | Dirty-peer resync timer fires, by cleared / still_dirty |
bgp_rib_ingest_channel_depth | Queued RibUpdate messages (one per inbound UPDATE or control event, not routes) waiting for the RIB manager, sampled once per manager loop iteration: the ingest channel plus at most one update held while a distribution window decides whether to extend. Producers block only when the channel itself is full, which reads capacity + 1 while an update is held and capacity otherwise; a sustained reading at or above the channel capacity means the channel is full or within one message of full |
bgp_rib_policy_transition_in_progress | Whether the RIB actor owns an atomic export-policy transition (1 = in progress). Mutations remain fenced until the transition is terminal; general RIB queries are served from the pre-commit state between the pre-commit polls and wait only through the commit batches, while the dedicated core-readiness lane remains responsive throughout |
bgp_rib_policy_transition_last_duration_milliseconds | Monotonic elapsed duration of the most recently completed atomic export-policy transition; retained across idle periods for post-event diagnosis |
bgp_rib_policy_transition_actor_poll_duration_seconds{poll_kind} | Duration of each real RIB actor transition poll. Bounded poll_kind values are bounded (chunked phase work), prefix_snapshot (the two complete O(table) snapshot polls), finalize (atomic membership/emission commit plus the global dirty/forced retry opportunity), and commit (bounded CommitMembers batches — at most eight members flushed per poll) |
bgp_rib_policy_transition_total{outcome} | Terminal actor-owned policy transitions. committed means the atomic cohort transition committed; fallback_handoff means uncommitted cleanup succeeded and authoritative per-peer apply is required; fallback_cleanup_error means that cleanup failed. Empty, rejected, missing, in-progress, abandoned, continued, pre-ownership, and synthetic actor-exit paths do not increment it |
bgp_rib_outbound_prefix_limit_actor_duration_seconds{operation} | Duration of complete synchronous outbound prefix-limit work on the RIB actor. The closed operation set is apply (one active transaction after its identity/epoch gates, including the live-peer precondition recheck and any successful installation) and recovery (one non-empty scheduled batch that replays at least one live peer/family). An apply whose live precondition recheck rejects still contributes its real scan time; discarded, missing, superseded, idempotent, and empty paths do not contribute samples |
bgp_rib_actor_work_duration_seconds{work_unit} | Wall-clock duration of RIB actor work components. The closed work_unit set is route_chunk (construction and processing of one bounded route chunk, excluding its drained-batch tail), distribute_flush (one outbound pass across the peer set per distribution window: a route batch plus any unicast-only route messages that were already queued when it drained, up to 256 messages, 4,096 routes, 1,024 changed prefixes or 5 ms, so one observation can cover many UPDATE messages and the series count is windows, not messages; includes readiness servicing at peer boundaries), exact_export_retire (retiring exact-export rejections when that window settles), attribute_gc (deadline-triggered attribute-intern collection outside ingest chunks), and flowspec_validation (a receive-side feasibility slice, including completed selection/distribution). Validation slices cap candidate visits; dependency discovery, peer inventory, and distribution add work outside that visit count. Count-triggered collection remains included in its route chunk. Compare the le="0.2" bucket against the series count to count components above 200 ms. These are component timings for correlation with readiness waits, not uninterrupted stalls, a bound on probe latency, or coverage of all actor work |
bgp_rib_readiness_query_wait_seconds{seam} | Wall-clock delay from admission to the dedicated RIB readiness lane until actor service. The closed seam set is actor_loop (ordinary drains, including in-pass ingest servicing), policy_transition_fence (synchronous replacement checkpoints), and selection_release (synchronous selection-deferral release checkpoints). Records service even after a caller times out, but queries canceled before admission or never served contribute no sample. Excludes channel admission wait, the prior peer-manager probe, and reply delivery; it is not full /readyz latency or a probe-timeout rate |
bgp_peer_manager_operator_query_wait_seconds{seam} | Wall-clock delay from send on the peer manager's operator-read lane until actor service begins, including bounded-channel admission wait. The closed seam set is unfenced, prestage (policy preflight, cohort selection, destination prestage and session setup), forward_transition, commit_batches, and rollback. The label is the current command's policy marker, or the latest completed marked command overlapping the wait; intervening ordinary commands preserve it. Markers include trailing command work. A wait spanning several phases is recorded once under one label, not split by cause. Fenced phases produce samples only when reads are eventually drained, so their counts depend on arrival timing and are not phase load. Service after caller timeout still counts; sends canceled before admission and reads never drained do not. Excludes work before the send, service execution, and reply delivery; this is not RPC latency or a timeout rate. The 0.1 s, 0.5 s, and 2 s edges match per-peer, explain, and aggregate budgets; count - bucket{le="2"} counts waits over 2 s |
bgp_orr_input_objects{classification} | Inputs considered by the ORR default-topology builder before NLRI deduplication. Exactly five classifications exist: included_default, excluded_nondefault, malformed_topology, malformed_attribute_29, and default_with_ignored_flex_algo. The Flex series is a subset of included default objects: its base object and classic metric remain usable. All series reset to zero when no vantage is configured |
The shipped alert pack raises BgpPolicyTransitionStalled when
bgp_rib_policy_transition_in_progress == 1 for one minute. Independently,
readiness remains healthy during bounded progress below 30 seconds, then fails
closed with RIB export-policy transition stalled until commit or a cleaned-up
fallback restores it. A selection-deferral release shares that 30-second bound
but fails closed under its own reason, RIB selection-deferral release stalled, and records its readiness waits under the selection_release seam. Use the poll-duration histogram to distinguish many
bounded polls from one long O(table) snapshot or finalization poll; the
retained terminal duration still describes the previous completed transition
while one is active.
A readiness request that arrives after general queries have queued can overtake
them on its dedicated lane. The PeerManager half is a constant-time actor ping
and does not fan out session-state queries. It cannot preempt an O(table)
general query that was already executing when the request arrived; that actor
call must return before either the transition or readiness probe can advance.
Policy artifact freshness
For deployments whose policy is rendered by an external pipeline (IRR toolchain, config generator, cron + SIGHUP), these metrics answer "when did the daemon last accept new policy artifacts" — which is the staleness signal that matters. File mtimes only say when something was written to disk; a render that produces a config the daemon rejects leaves the old generation live, and these timestamps deliberately do not advance on a rejected load.
| Metric | What it tells you |
|---|---|
bgp_policy_generation_loaded_timestamp_seconds | Unix time of the last successful full policy apply — initial load, SIGHUP reload (stamped even when the reloaded content is unchanged: the daemon re-accepted it), or config transaction. Frozen across rejected reloads |
bgp_policy_dataset_loaded_timestamp_seconds{dataset} | Unix time the named [policy.datasets.<name>] external dataset last swapped in a loaded generation (initial load, or a refresh whose content changed). A failed refresh keeps the prior snapshot serving and does not advance this. Series reaped when the dataset is removed from config |
bgp_policy_dataset_refresh_errors_total{dataset} | SIGHUP candidates in which the named, already-running dataset's file failed to load or parse (a newly declared dataset that fails to load rejects the candidate at parse time instead, without counting). The daemon counts the failure before routing the reload, then rejects the whole reload with no runtime effect (SIGHUP reload rejected without runtime effect), so the prior snapshot keeps serving and neither timestamp advances. Series reaped when the dataset is removed from config |
Alerting on pipeline staleness. Export ages, not the raw timestamps, and let Prometheus do the arithmetic. If your pipeline refreshes every 6 hours, page when nothing has been accepted for a few missed cycles:
# Full policy apply older than 3 refresh intervals (here: 3 × 6h).
time() - bgp_policy_generation_loaded_timestamp_seconds > 3 * 21600
# A dataset whose last accepted swap is stale, or whose refreshes are
# actively failing.
time() - bgp_policy_dataset_loaded_timestamp_seconds > 3 * 21600
increase(bgp_policy_dataset_refresh_errors_total[6h]) > 0The dataset timestamp advances only when a refresh swaps in changed content, so pair the dataset-age expression with the refresh-error counter: age alone can also mean "the data legitimately hasn't changed", while age plus rising errors means the pipeline output is being rejected. A malformed dataset file also freezes the full-apply timestamp, because it rejects the whole reload. That timestamp has no data-unchanged caveat — it is stamped on every successful SIGHUP — so it is the primary "pipeline stuck or daemon rejecting everything" pager.
RFC 8212 explicit-policy enforcement
Both 0/1 gauges exist for every configured peer, including zero-valued series
while [global] ebgp_requires_policy = false (ADR-0112). They are read from the
chain each peer currently has installed, so they cannot disagree with the
RFC 8212 Policy block in rbgp neighbor <addr> or with the rbgp doctor
verdict. Series are reaped when the peer is removed.
| Metric | What it tells you |
|---|---|
bgp_rfc8212_missing_import_policy{peer} | 1 when this external session has no explicit operator import policy, so the reserved internal deny is installed and no route it announces becomes eligible, in any negotiated family. 0 otherwise, including for iBGP and while enforcement is off |
bgp_rfc8212_missing_export_policy{peer} | 1 when this external session has no explicit operator export policy, so nothing enters its Adj-RIB-Out in any negotiated family. 0 otherwise, including for iBGP and while enforcement is off |
# Any peer fail-closed on a missing operator policy in either direction.
(bgp_rfc8212_missing_import_policy
+ bgp_rfc8212_missing_export_policy) > 0PromQL's bare or is a left-biased set union, not a boolean OR of sample
values; because both directional series have the same labels, it would hide an
export value of 1 behind an import value of 0. The shipped directional warnings
use == 1 with a five-minute hold. Alert on those rather than on readiness:
/readyz stays green for a healthy daemon whose reserved deny is doing exactly
what it was configured to do.
Ingress rejection / route-leak detection
Mechanisms that block routes for protocol-correctness reasons: most
reject at the session boundary before the route reaches the RIB. Egress OTC
and exact-export checks instead reject an outbound advertisement after policy
but before Adj-RIB-Out commit. If a previously advertised route becomes
blocked, the RIB withdraws it and removes the logical advertised entry. Where a metric
carries a reason label, its values are the canonical contract
pinned in crates/telemetry/src/reason_labels.rs — stable across
releases and shared verbatim by the metric label, the log-line
reason token, and (for OTC) the structured OTC_ROUTE_BLOCKED
event payload, so alert expressions can key on them safely. Metrics
without a reason label encode the mechanism in the metric name.
| Metric | What it tells you |
|---|---|
bgp_otc_routes_blocked_total{peer,reason} | RFC 9234 Only-to-Customer route-leak blocks (ADR-0071). reason is ingress_from_customer_rsclient (OTC-tagged route arrived while we act as Provider / Route Server), ingress_peer_mismatch (lateral Peer session, OTC value is not the peer's ASN), malformed_length (OTC attribute undecodable; announcements dropped treat-as-withdraw style per RFC 7606), or egress_to_upstream_via_otc (post-policy route rejected before grouped/private Adj-RIB-Out commit toward a Provider / Peer / Route Server). Ingress remains per rejected UPDATE decision; egress increments once per transition of a peer/route identity into the blocked disposition, and permit, withdrawal, or session reset re-arms it. Export explain remains current decision truth; transport retains a defense-only check. |
bgp_role_mismatch_total{peer,local_role,remote_role} | OPENs refused with RFC 9234 Role Mismatch (NOTIFICATION 2/11). local_role is the configured Role, or none. remote_role is the first assigned Role in the OPEN (provider, route_server, route_server_client, customer, peer); unrecognized when the OPEN carries Role capabilities with only unassigned values (5-255) or a length other than 1, with the raw value in the warning log's remote_role_raw field; none when the OPEN carries no Role capability (a strict_role rejection). See RFC notes. |
bgp_as_path_loop_detected_total{peer} | Announced NLRI rejected because our own ASN appears in the received AS_PATH (RFC 4271 §9.1.2), counted across every address family, including EVPN and FlowSpec. No reason label — the mechanism is the metric name; withdrawals in the same UPDATE are still processed |
bgp_rr_loop_detected_total{peer} | UPDATEs rejected by route-reflection loop detection (RFC 4456 §8). No reason label; the debug log line emitted with each increment carries reason=originator_id (received ORIGINATOR_ID equals our router-id) or reason=cluster_list (our cluster-id already in CLUSTER_LIST) |
bgp_bgpls_nlri_discarded_total{peer} | Known BGP-LS NLRIs dropped for out-of-order descriptor TLVs (RFC 9552 fault management). The affected NLRI is isolated and the session is preserved; each increment carries a family=bgp_ls debug log line. Fatal BGP-LS framing/length errors are not counted here — they still reset the session |
bgp_evpn_nlri_discarded_total{peer} | Unrecognized or unsupported EVPN typed NLRIs discarded under RFC 7606 §5.4 while supported routes in the same MP attribute and the session are preserved. This existing counter remains the peer aggregate across announcements and withdrawals. Malformed EVPN framing and malformed payloads for supported types are not counted here and retain their existing decode-error handling. |
bgp_evpn_nlri_discarded_by_type_total{peer,route_type} | The same discards separated by decimal wire route type (bounded to one octet). One WARN per type per TCP connection identifies the peer, family=evpn, route_type, and initial discarded count; DEBUG records retain every per-type count per UPDATE. Repeated discards continue incrementing both counters. Ordinary reconnects preserve counters and re-arm warnings; configured-peer deletion reaps all type series. Use this counter to identify unsupported routes sent by a peer, such as multicast types 7–8; these routes do not enter the RIB or get reflected. |
bgp_path_attribute_discarded_total{peer,type_code} | Surviving decoded attributes removed by effective discard_path_attributes, plus well-formed ORIGINATOR_ID (9) and CLUSTER_LIST (10) received from an external neighbor, which RFC 7606 §7.9/§7.10 discard without any configuration (an attribute matching both counts once); type_code is the decimal wire type. Increments once per removed attribute occurrence in an UPDATE, not once per NLRI, and each peer/type/update produces one bounded DEBUG record. RFC 7606-removed malformed attributes do not increment it; their reported causes appear in bgp_update_malformed_causes_total. Ordinary session resets preserve the counter; configured-peer deletion reaps all of its type-code series. |
bgp_update_malformed_total{peer,disposition} | Malformed UPDATE messages by the RFC 7606 disposition applied: attribute_discard (offending attribute dropped, UPDATE proceeds), treat_as_withdraw (every route in the UPDATE handled as withdrawn, session stays Established), or session_reset (NOTIFICATION + teardown, retained where the NLRI cannot be trusted — including the §5.2 escalation when a treat-as-withdraw-class error arrives with no reachable NLRI). One increment per malformed UPDATE, labeled with the strongest-action disposition that governed it (§3 (h)). Each increment is accompanied by a warn log line per malformed attribute and, at DEBUG, the §6 full-message hex capture |
bgp_update_malformed_causes_total{peer,type_code,reason,disposition} | Causes reported while processing malformed UPDATEs, with the final applied disposition on every cause. type_code is decimal 0–255 or none when no attribute type is attributable. Multiple independent causes can occur in one UPDATE or attribute; this is neither an UPDATE nor an NLRI count. Each prohibited AS-set attribute occurrence contributes one as_set_prohibited cause. A missing mandatory attribute caused by the decoder removing that same malformed attribute adds no second cause. Counters survive session resets; configured-peer deletion reaps them. |
bgp_exact_export_rejections_total{peer,family,reason} | Post-policy announcements rejected before Adj-RIB-Out commit because the session's exact one-route encoder could not produce a legal wire message. family is a bounded OpenConfig AFI/SAFI label; reason is encoding, missing_ipv6_next_hop, ipv4_requires_extended_next_hop, or message_too_long. Alert on a sustained increase, then correlate the peer/family with the warning log's bounded route identity and detail. Series are reaped only when the configured peer is deleted. |
bgp_max_prefix_exceeded_total{peer} | max_prefixes ceiling breaches; each increment is followed by max-prefix teardown: bare Cease/1 without Notification GR, or RFC 8538 Hard Reset encapsulating Cease/1 when the N-bit was negotiated (see "Peer max-prefix exceeded" above) |
bgp_max_prefix_blocked_total{peer,scope} | Blocking episodes opened for a scope under max_prefix_action = "block", counted once when the first net-new prefix is withheld by a full bound — not once per withheld prefix. Each increment marks a bgp_max_prefix_blocking{peer,scope} transition to 1; the scope must recover to 0 before another withheld prefix increments the counter again. Read it as how often a peer has driven a bound to full, never as how much it sent: a rate() over it measures episode frequency, the gauge answers whether prefixes are being withheld right now, and the peer's replay after recovery, or rbgp rib received <addr>, shows what is installed. Carries no prefix label: the withheld set is exactly the unbounded quantity the limit exists to contain |
bgp_max_prefix_warning_total{peer,scope} | max_prefix_warning_percent (or the bound itself under max_prefix_action = "warning") crossings, one per crossing per scope; each is paired with one warn log line and one max_prefix_warning session event |
Malformed UPDATE cause reasons are bounded to attribute_list,
unrecognized_well_known, missing_well_known, attribute_flags,
attribute_length, invalid_origin, invalid_next_hop, optional_attribute,
invalid_network, malformed_as_path, as_set_prohibited, as_path_limit,
aspa_first_as_mismatch, srv6_service_tlv, and other. Only typed RFC 9774
AS_SET / AS_CONFED_SET findings use as_set_prohibited; separate malformed
path, flag, length, or other-attribute findings remain visible. Reasons and type
codes never contain raw diagnostic text, ASNs, or prefixes. The theoretical
bound is 257 types × 15 reasons × 3 dispositions = 11,565 series per peer;
only observed combinations are created.
For a cumulative inventory excluding the known prohibited-set cause:
bgp_update_malformed_causes_total{reason!="as_set_prohibited"} > 0This query covers the counter lifetime, not a recent time window. A new cause
series first appears at one: increase() alone can miss that first event until
another increment is observed. Keep the existing BgpMalformedUpdate aggregate
alert for complete recent malformed-message detection; its three disposition
series are initialized at zero. Use the detail counter and warning logs to
identify causes and tune local alert policy.
The shipped BgpMaxPrefixNearLimit example alert warns after a finite scope
has remained at or above 80% usage for ten minutes. This threshold lives in the
editable Prometheus rule, not daemon configuration. Unlimited and disconnected
scopes cannot fire because their limit series is absent. After a teardown,
the BgpMaxPrefixLimitExceeded counter alert pages on the breach and resolves
10 minutes later; BgpMaxPrefixLatched holds while the peer stays latched.
Event Streams
| Metric | What it tells you |
|---|---|
bgp_session_event_source_dropped_total{kind,reason} | Session state_change or notification events dropped before peer-manager publication because the bounded source channel was channel_full or channel_closed. All four bounded series exist at zero from startup. |
bgp_event_stream_lagged_total{service,source} | Events skipped because a live stream subscriber fell behind the bounded broadcast channel. service is watch_events; source is route, session, policy, evpn, dataplane, dataplane_route, or bfd where applicable |
bgp_event_stream_subscribers{service,source} | Current live stream subscriber count by service/source |
bgp_route_event_history_depth | Current number of unicast route events retained for ListRouteEvents / rbgp events history queries |
bgp_route_event_history_capacity | Fixed capacity of the bounded unicast route-event history ring |
These counters identify distinct loss boundaries. A non-zero
bgp_session_event_source_dropped_total means the peer manager never received
the named event kind. For state_change, take a fresh neighbor snapshot before
consuming further incremental state events. For notification, consult the
structured daemon logs for the corresponding BGP NOTIFICATION or local
teardown cause.
WatchEvents is a live tail, not a durable queue. Non-zero
bgp_event_stream_lagged_total means at least one live subscriber fell behind
after the event reached the peer manager; combine a fresh snapshot or ListRouteEvents
query with a new live watch. bgp_event_outbox_dropped_total instead reports
loss at the later durable-outbox or durable-cursor boundary.
gNMI dial-out
| Metric | What it tells you |
|---|---|
gnmi_dialout_connected{target} | 1 while the dial-out Publish stream to this [gnmi_dialout] target is established, 0 while disconnected/retrying. Refreshed on both transitions; the series exists (at 0) from startup even when the collector is down, and is reaped when the target is removed from config (SIGHUP) |
gnmi_dialout_resync_total{target} | Established Publish sessions that ended and entered the reconnect path, where the next connection starts a fresh initial snapshot. Failed dials do not increment it; the series exists at 0 from startup and is reaped with the target |
gnmi_dialout_queue_depth{target} | SubscribeResponse items waiting in the bounded local subscription queue. Sustained non-zero depth identifies a target whose transport is not keeping up; it returns to 0 on disconnect and the series is reaped with the target |
gnmi_dialout_last_publish_timestamp_seconds{target} | Unix time the last SubscribeResponse was handed to the local gRPC transport, or 0 before the first publish. Use time() - <metric> for publish age. This is not a collector acknowledgement; the series is reaped with the target |
BMP
| Metric | What it tells you |
|---|---|
bmp_stream_diverged{peer} | 1 while a dropped per-peer RouteMonitoring event has left collectors' live view incomplete and the session is awaiting its bounded synthetic PeerDown/PeerUp repair. It clears after a successful reset or genuine session teardown and is reaped when the peer is deleted |
bmp_loc_rib_source_drops_total{event,reason} | RFC 9069 Loc-RIB events dropped before they reached the BMP manager, at the internal RIB/PeerManager→BmpManager channel. event is route_monitoring (a Loc-RIB route install/withdraw that will never reach any collector's Loc-RIB view) or stats (one periodic Loc-RIB statistics report skipped; the next tick reports current totals). reason is channel_full (the BMP manager is not draining as fast as RIB churn produces events) or channel_closed (BMP teardown). Both label sets are bounded; the series is process-global, not per-collector. A dropped route_monitoring event silently diverges every connected Loc-RIB collector until its next reconnect-triggered table dump, so alert on any sustained channel_full increase and correlate with the per-collector bmp_collector_drops_total and the bmp_loc_rib_dump_live_buffer_* gauges to find the slow consumer; each increment also emits a warn log line |
bmp_collector_drops_total{collector,phase,reason} | A bounded collector queue or Loc-RIB dump failure. phase="fan_out" with channel_full or channel_closed automatically fences only that collector generation: TCP closes without BMP Termination, the client retries after one second, cached Peer Up state is replayed, and configured Loc-RIB state is rebuilt through a fresh dump and EoR before live deltas. Healthy collectors are unaffected. Repeated increases identify a collector that cannot drain at the live event rate |
bmp_control_event_drops_total{collector,kind,reason} | Collector lifecycle transitions that failed to reach the BMP manager. kind identifies connected, bootstrap-complete, or disconnected; reason is the bounded channel failure. Any increase means manager state may not match the collector connection lifecycle |
The sibling bmp_source_drops_total{peer,reason} counts per-peer loss before
the BMP manager. channel_full / channel_closed cover the
PeerSession/PeerManager→BmpManager path. state_query_timeout means the
periodic statistics tick could not obtain a current session snapshot within
its bounded deadline; the session may still be Established, so this is not a
Peer Down signal and the report is retried on the next tick.
The periodic sampler gathers session, peer-RIB, and Loc-RIB values concurrently under their existing 100 ms input budgets, including RIB channel admission. These are independent observations, not an atomic cross-actor snapshot. Unavailable RIB values are omitted for that tick; BMP output uses nonblocking sends, with the source-drop counters above recording output backpressure.
The shipped alert pack
(examples/prometheus/rustbgpd-alerts.yml)
fires BmpSourceDrops, BmpLocRibSourceDrops, BmpCollectorDrops, and
BmpControlEventDrops on any 10-minute increase of the corresponding loss
counters. BmpStreamDiverged fires when the new gauge remains 1 for five
minutes, identifying a repair that has not completed rather than a historical
drop that already recovered.
MRT
| Metric | What it tells you |
|---|---|
mrt_dump_interval_seconds | The configured [mrt] dump_interval. Exported on every instance: 0 when [mrt] is not configured, otherwise always > 0. The alert pack's MrtDumpStale guards on a value > 0 and fires once the newest dump is older than twice it |
mrt_last_dump_success_timestamp_seconds | Unix time the last periodic or on-demand dump was published successfully; 0 until the first success after start. A failed dump does not advance it, so time() - <this> is the age of the newest dump file |
mrt_last_dump_duration_milliseconds | Wall-clock duration of the last successful dump from trigger to published file (RIB snapshot, encode, write, rename) |
mrt_dump_bytes_written_total | Bytes of dump files published on disk (after compression) |
mrt_dump_failures_total{stage} | Failed dumps by bounded stage: preflight (output directory), snapshot (RIB actor query), encode, write. A caller-canceled on-demand dump is not counted; a periodic failure also logs at error level, an on-demand failure is returned to its caller |
Durable Event Cursor (ADR-0072)
The durable outbox (SubscribeFromEvent RPC, CLI
rbgp events watch --from-event-id <N>, and the
examples/event-bridge/ reference binary) survives daemon restart and
exposes a monotonic event_id cursor. The legacy live surfaces above
(WatchEvents / List*Events) keep their existing
ring-backed behavior and are unaffected by this section.
While rbgp events watch --from-event-id <N> remains running, it reconnects
only after clean EOF or gRPC UNAVAILABLE. Backoff starts at 1 second, doubles
to a 30-second cap, and resets after a complete human or JSON record plus
newline is written and stdout is flushed. Each request preserves the original
category, type, neighbor, family, and prefix filters; only the cursor changes.
The cursor uses the highest flushed top-level BgpEvent.event_id, so lag frames
without that field do not advance it. Every other RPC status and any output
failure is terminal. This cursor lives only for the running CLI process; use
the event bridge when a downstream-confirmed cursor must survive CLI restart.
Cursorless OTC subscriptions and ordinary WatchEvents streams remain
one-shot.
Opt-in — default off as of v0.32.0. The outbox is disabled by default (v0.32.0 benchmarking measured ~62 MB RSS plus roughly double the peak CPU at 2p/100k — too much to impose on operators who never consume the cursor). Enable it explicitly and restart:
[event_history]
enabled = true
max_bytes = 256_000_000 # size retention to your collector's reconnect SLATwo deployment profiles:
- Lean / high-scale (the default):
[event_history].enabled = false. Routing fast and lean;SubscribeFromEventand gNMISubscribe ON_CHANGEreturnFAILED_PRECONDITION, but the liveWatchEvents/List*Eventssurfaces still provide real-time observability. Security-signal caveat: the structuredOTC_ROUTE_BLOCKEDevent (RFC 9234 route-leak prevention — per-decision prefixes, AS_PATH, roles) is emitted only through the durable outbox, so with the lean default it is not available viaSubscribeFromEvent. Blocks are still observable via the always-onbgp_otc_routes_blocked_total{peer,reason}counter, theotc_routes_blockedper-neighbor scalar, and the daemon log — but if you need the rich per-decision event for incident reconstruction or a SIEM feed, enable event-history (the observability/replay profile below). - Observability / replay:
[event_history].enabled = truewithmax_events/max_bytessized for your collector's worst-case reconnect window. Gives restart-safeevent_idcursor replay; budget the ~62 MB RSS + CPU shown indocs/benchmarks.md.
Producer set: route, evpn, session (lifecycle + notification),
policy (config-mutation POLICY_CHANGED events plus
transport-layer OTC_ROUTE_BLOCKED route-leak decisions — both
ride on EVENT_CATEGORY_POLICY, discriminated by BgpEventType),
bfd, and dataplane (FIB + blackhole summaries and per-route
install / withdraw / failure events). All six categories are
produced through the durable outbox when
[event_history].enabled = true. The dataplane poller is
startup-spawned so durable summaries flow regardless of whether
any WatchEvents subscriber is alive.
| Metric | What it tells you |
|---|---|
bgp_event_outbox_committed_total{category} | Events durably committed to the local outbox, per category. Increments inside EHM after the SQLite transaction commits — not on producer enqueue. |
bgp_event_outbox_dropped_total{category, reason} | Events lost before durable storage, or committed events skipped during durable cursor delivery. reason is queue_full, closed, db_error, shutdown_timeout, decode_failure, opaque_codec, or source_lagged. shutdown_timeout means an accepted event remained actor-queued or in an accepted producer handoff when the coordinated shutdown deadline expired; an append already owned by the actor has unknown outcome and is not counted as a definite per-category drop. source_lagged fires when an upstream broadcast receiver (FIB or BFD bridge) reports Lagged(missed) — those missed events never reached the bridge body and therefore never reached EHM; the counter increments by missed. queue_full, db_error, shutdown_timeout, decode/codec failures, and source_lagged flip bgp_event_outbox_degraded to 1; shutdown-time closed drops do not. After a runtime storage failure, events that were already accepted count as db_error and producers count refused events as closed; bgp_event_outbox_storage_failed distinguishes these from shutdown. |
bgp_event_outbox_queue_depth{category} | Accepted pending events in the EHM producer queue by category, including a producer handoff between its admission CAS and permit.send. Climbs before drops start — early-warning signal. |
bgp_event_outbox_db_size_bytes | Combined size of events.db + WAL on disk, refreshed after commits and retention passes. [event_history].max_bytes is the scheduled size-retention target; a bounded pass can finish while the store remains above it. |
bgp_event_outbox_retention_evicted_total{reason} | Events evicted by the retention pass. reason is count_cap or byte_cap. |
bgp_event_outbox_latest_event_id | The latest committed event_id. Forward progress indicator. |
bgp_event_outbox_open_failures_total | DB-open failures across the process lifetime. Typically 0 or 1; non-zero means EHM went into recovery or pass-through at startup. |
bgp_event_outbox_degraded | 1 once the outbox has seen durability-impacting loss, a committed-event delivery skip, or DB open/recovery/quarantine failure since start. Expected shutdown reason=closed drops are excluded. The signal does not auto-clear in v1; restart clears the latch. |
bgp_event_outbox_storage_failed | 1 once the event-history storage thread stopped (exited or panicked) while the daemon was running. The outbox then refuses producer events and durable cursor subscriptions, and bgp_event_outbox_degraded is also 1. Restart the daemon to recover; the signal does not auto-clear. |
bgp_event_outbox_cursor_gap_total | Cursor-gap StreamLagEvent frames sent on SubscribeFromEvent streams: the leading frame when the requested cursor was older than the retention floor, plus one for each time retention evicted events ahead of a replay still in progress. Operator signal that [event_history].max_events / max_bytes is undersized for the collector reconnect SLA. |
FAILED_PRECONDITION on SubscribeFromEvent means one of:
[event_history].enabled = falsein the daemon config — by design; the legacy live surfaces still work, but the durable cursor is intentionally off. Flip totrueand restart.[event_history].required = falseand EHM failed to openevents.dbat startup (permission denied, disk full, corruption).bgp_event_outbox_open_failures_totalwill be ≥ 1 andbgp_event_outbox_degradedis1. Check the daemon log for the reason; fix permissions / free disk / restore from backup, then restart. Pre-1.0,required = trueis the strictest posture — the daemon refuses to start when the outbox cannot be opened.- Event-history startup failed because the allocator anchor was unrecoverable
(for example, a moved
.stale-<ts>quarantine file with no sidecar fallback). Withrequired = false, the daemon continues live-only without EHM; withrequired = true, startup fails. Check the startup log, fix the underlying I/O issue, and restart.
UNAVAILABLE during SubscribeFromEvent replay means the storage actor
stopped accepting work or failed to reply. Inspect daemon health and logs,
restart if necessary, and resume from the last received top-level event_id.
Events already delivered remain valid; the terminal status is emitted once.
UNAVAILABLE at SubscribeFromEvent admission means the storage thread
has already stopped: bgp_event_outbox_storage_failed is 1 and the daemon
log has the error event-history storage stopped (after a panic, also a
crash report). From that point the outbox accepts no producer events, so the
events produced before the restart are lost. Streams that were open when the
storage stopped end with DATA_LOSS, and gNMI Subscribe ON_CHANGE streams,
including new ones, end the same way. Restart the daemon, then reconcile
collectors against authoritative state before resuming from their last
event_id.
Sizing retention — max_events and max_bytes are retention targets
evaluated during each scheduled pass. The count target is evaluated first and
removes at most 5,000 oldest events above the target; the byte target then
removes up to ten 5,000-event batches while the database remains oversized.
Default 100_000 events /
256_000_000 bytes covers a few minutes of even a busy daemon's
event stream. Operators with longer collector-reconnect tolerance
should raise both proportionally; collectors should alert on
bgp_event_outbox_cursor_gap_total > 0 to know when retention is
too small for their SLA. A busy or heavily oversized store can remain
above either target between passes or after one bounded pass. If sustained
event production exceeds the maximum per-pass eviction throughput, the store
can continue growing without a hard ceiling; alert on
bgp_event_outbox_db_size_bytes. SQLite may also hold onto freed pages rather
than immediately shrink the main DB.
External-bus integration — see examples/event-bridge/ for the
reference skeleton. The pattern is:
- Connect with the last
event_idyour downstream sink confirmed durable. - Forward
BgpEventrecords to your sink. - Advance the persisted
last_seen_event_idonly after the sink confirms durable receipt. - Treat a
StreamLagEventas a gap signal, not a stream end — your collector lost events older than the retention floor. It is the leading frame when the cursor was already too old, and it can also arrive mid-replay if retention evicts events ahead of a slow replay;missed_countcovers exactly the ids skipped at that point. - Use
BgpEvent.timestamp, notevent_id, for causal joins across event categories. The durableevent_idis order-of-arrival at the EHM actor, not order-of-occurrence at each producer.
Delayed outbound registration
BgpPeerOutboundUnregistered warns after five minutes of continuously observed
Established sessions without bgp_peer_update_group membership for that
instance and peer address. Any present value, including group 0 or the
private-path sentinel -1, means registered. Queued imports take priority over
initial outbound registration and can defer it without a fixed maximum delay.
Five minutes is an investigation policy, not a guaranteed startup budget or
proof that registration was lost. Inspect peer state and RIB backlog before
attributing the absence to a fault; slow-peer and writer-queue signals can stay
quiet when no outbound registration exists.
rbgp doctor adds a yellow peer.<addr>.outbound check when a successful,
non-stale neighbor/RIB snapshot shows an Established session older than five
minutes with an empty update_group. The message reports session age and
current absence, not the duration of missing registration. Unavailable
snapshots and stale session observations do not establish absence. An empty
field also requires a recognized daemon health version of at least 0.50.0,
a release that exposes membership. Missing, older, malformed, or prerelease
version strings leave registration unknown; any nonempty group remains evidence
of registration. The warning does not change readiness or make doctor exit red.
RIB membership is address-level. Same-address IPv6 link-local sessions on
different interfaces share the membership evidence: one registered sibling can
mask another's missing registration. A scoped doctor check name identifies the
session observed, but does not make the underlying membership interface-scoped.
The aggregate bgp_rib_outbound_registered_peers remains an unlabeled total.
Slow peers (detection and isolation)
A slow peer is a client that is Established and alive — keepalives flow in both directions, so neither the hold timer nor the RFC 9687 send-hold timer fires — but persistently fails to keep up with the update stream we send it. On a route reflector or route server with shared update-groups, one such client can drag the shared staging pass for its whole group. Typical causes: an undersized control plane on the client, a congested or lossy path (small effective TCP window), or a client busy with its own convergence.
Detection is on by default and purely observational: once a peer's
outbound backlog stays at or above slow_peer_threshold_pct (default
50%) of the writer buffer for slow_peer_duration seconds (default
30), the daemon
- raises the
slow_peerflag inrbgp neighbor <ip>,rbgp neighbor --json, and theGetNeighborState/ListNeighborsgRPC surface; therbgp neighbor --widefleet view marks it with!in theSlowcolumn, - sets
bgp_peer_slow{peer}to 1, and - logs a warn (
peer flagged slow: session alive but outbound queue persistently backlogged).
All three clear as soon as the backlog drains below the threshold —
the flag can never latch — and on session teardown. Set
slow_peer_duration = 0 on a neighbor or peer group to disable
detection.
What to do when it fires:
- Confirm with
bgp_peer_outbound_queue_depth{peer}— a persistent plateau (rather than sawtooth bursts) matches the flag. - Check whether the whole group is slow (many peers flagged → suspect the local box or an upstream burst) or one client lags its group-mates (→ suspect that client / its path).
- For a chronically slow client on a shared update-group, enable
slow_peer_isolation = truefor it (or its peer group): the daemon moves the flagged peer onto its own per-peer update path (bgp_peer_update_groupshows the ungrouped sentinel-1, reasonslow_peer) so the rest of the group converges at full speed, and regroups it automatically when it recovers. Isolation trades the shared-encode saving for that one peer against group convergence — it is off by default. - A peer that stops draining entirely is handled by the existing
guards, not this feature: the RFC 9687 send-hold timer tears down a
wedged socket with local error
8/0, without a NOTIFICATION. Large outbound envelopes pause at the bounded writer queue while the session continues serving timers, input, and operator snapshots. If no writer capacity becomes available before the resource-admission deadline, the session sendsCease/Out of Resources(6/8). Admission progress resets this deadline. Its interval is the configured nonzerosend_hold_time, ormax(480, 2 × hold_time)seconds whensend_hold_time = 0; disabling the RFC timer does not disable the resource bound.
Native gRPC certificate expiry
bgp_grpc_tls_certificate_not_after_seconds{kind} reports certificate
notAfter as Unix seconds from the active credential generation. Its only
label is kind:
| Kind | Meaning |
|---|---|
server_leaf | First certificate in the configured server PEM bundle |
server_bundle_min | Earliest date across every certificate supplied in that bundle, including an optional root |
client_ca_bundle_min | Earliest date across every certificate supplied in the client trust bundle |
These are raw certificate dates. A bundle minimum does not determine the effective peer certification path or its handshake cutoff. In particular, the client CA bundle is not an inventory of client leaf certificates, and a trust anchor's date is not an automatic rustls handshake cutoff.
# Active server leaf is within seven days of notAfter, including past dates.
bgp_grpc_tls_certificate_not_after_seconds{kind="server_leaf"} - time() <= 604800Without native TLS, the family has no series. Unavailable metadata is omitted; if any supplied bundle member cannot be inspected, that bundle minimum is omitted rather than calculated over an incomplete subset. Metadata parsing does not add a credential rejection rule. File changes alone do not alter these gauges: successful SIGHUP credential publication replaces the snapshot, while failed reloads retain the prior snapshot. Existing connections cannot overwrite the active values.
Set tls_expiry_warning_seconds = 604800 in
[global.telemetry.grpc_tcp] to request warnings within seven days, including
past dates. The default 0 disables expiry warnings while preserving metrics
and successful-client metadata logs. The setting requires a restart. Warnings
appear at startup, after successful credential reload, and during --check;
ordinary --check returns 0 for warnings, while --check --strict returns 1.
The setting does not reject startup solely because a certificate date is past.
After each successful native TLS handshake, the grpc_tls_client_certificate
log event reports the observed client's leaf certificate_not_after_seconds
when available. A positive warning window also enables the
grpc_tls_client_certificate_expiry warning event. These records describe
clients that completed a handshake, not unseen clients or a guarantee about
future connections. RPC role authorization remains a separate check. There
is no per-client expiry metric or client certificate inventory.
gRPC audit and resource guardrails
ADR-0064 v1 uses the daemon's structured log path plus Prometheus metrics as the operational gRPC authorization audit and resource-guardrail surface. rustbgpd does not run a separate in-daemon audit file writer or remote audit sink in this release. That keeps audit emission on the existing non-blocking logging path instead of adding a second I/O path that could wedge routing-control tasks if a disk, syslog daemon, or collector stalls.
For production, collect stdout/stderr with journald, syslog, or your log agent of choice and apply retention outside the daemon:
- Retain
grpc_authzrecords for at least the same window as management-plane change approvals and incident timelines. Thirty to ninety days is a practical minimum for most environments; regulated deployments should use their own audit policy. - Store gRPC authorization logs where only network operators, security responders, and the log-collection service account can read them. The records mask known credential-bearing fields, but they still expose management-plane method names, principals, listener posture, peer names, and topology context.
- Rotate locally before the filesystem can pressure the daemon host. With
journald, bound
SystemMaxUse/RuntimeMaxUse; with syslog or a file-based collector, use normal logrotate or collector retention controls. - Export logs to remote storage when
operator_onlyactions, role denials, or authentication failures need tamper-resistant evidence.
Useful local queries follow. operator_only decisions log at WARN and most
others at INFO, but a successful read-tier call (CheckLiveness) logs at
DEBUG, so these queries omit such calls unless debug logging is enabled;
bgp_grpc_authz_decisions_total still counts every request. The queries
below need log_format = "json" (the edge and route-server profile
default; the lab profile writes text). The JSON log nests event fields
under .fields, so tier, result and principal are read as
.fields.tier and so on. The unit's journal also holds the plain-text
startup banner from stderr, so the queries read raw lines (jq -R) and skip
any line that is not JSON (fromjson?).
# gRPC authorization records at the configured log level from a systemd unit.
journalctl -u rustbgpd -o cat --since -24h \
| jq -R 'fromjson? | select(.target == "grpc_authz")'
# Operator-only calls, including forwarded and denied attempts.
journalctl -u rustbgpd -o cat --since -24h \
| jq -R 'fromjson? | select(.target == "grpc_authz"
and .fields.tier == "operator_only")'
# Listener-cap, role, and authentication denials to investigate.
journalctl -u rustbgpd -o cat --since -24h \
| jq -R 'fromjson? | select(.target == "grpc_authz"
and (.fields.result == "listener_tier_denied"
or .fields.result == "principal_unmapped"
or .fields.result == "role_tier_denied"
or .fields.result == "authn_failed"))'
# Mutating/operator activity grouped by principal.
journalctl -u rustbgpd -o cat --since -24h \
| jq -nR '[inputs | fromjson? | select(.target == "grpc_authz"
and (.fields.tier == "mutating" or .fields.tier == "operator_only"))]
| group_by(.fields.principal)
| map({principal: .[0].fields.principal, count: length})'TLS handshake failures occur before request authorization and increment
bgp_grpc_tls_handshake_failures_total{reason} once per rejected handshake,
including the ten-second handshake timeout. The fixed reasons are
missing_certificate, certificate_expired, certificate_not_yet_valid,
unknown_issuer, invalid_certificate (other certificate validation failures),
tls_error (other typed TLS errors), io_error (transport failures), and
timeout. All eight series start at zero. These failures do not increment
bgp_grpc_authz_decisions_total; successful handshakes and TCP accept errors
also do not increment the handshake-failure counter. Labels contain no
certificate contents, principal, listener address, or peer address. TLS alert
delivery to the client remains best effort; use the server metric to observe
its rejection reason.
Prometheus should watch handshake failures, authorization volume, and stream pressure:
# Rejected TLS handshakes, including clients that never reach an RPC.
sum by (reason) (increase(bgp_grpc_tls_handshake_failures_total[5m]))
# Any denied management-plane call in the last five minutes.
sum by (tier, result, authn, access_mode) (
increase(bgp_grpc_authz_decisions_total{
result=~"listener_tier_denied|principal_unmapped|role_tier_denied|authn_failed"
}[5m])
)
# Operator-only calls, successful or denied.
sum by (result, authn, access_mode) (
increase(bgp_grpc_authz_decisions_total{tier="operator_only"}[5m])
)
# Slow live-stream consumers missing events.
sum by (service, source) (
increase(bgp_event_stream_lagged_total[5m])
)
# Current live stream fan-out.
sum by (service, source) (bgp_event_stream_subscribers)Resource-abuse posture for v1 is intentionally operational rather than a new
daemon-side rate limiter. The existing controls are listener max_tier, opt-in
per-principal tier enforcement, pagination on large route-list RPCs, bounded
event-history rings, bounded stream broadcasts, subscriber gauges, and lag
counters. Keep accepted clients on a management network and use separate
listeners when monitoring, automation, and operators need different ceilings.
For BMP Loc-RIB collectors,
bmp_loc_rib_dump_live_buffer_depth{collector} reports live rows currently
held behind bootstrap or a table dump, while
bmp_loc_rib_dump_live_buffer_high_watermark{collector} retains that
connection generation's peak. The label is the configured SocketAddr
including port. Current depth returns to zero when storage is released; the
high-water mark resets only after the next validated connection.
The RPCs that deserve the most attention are:
| Class | Examples | Guardrail |
|---|---|---|
| Large sensitive reads | ListReceivedRoutes, ListBestRoutes, ListAdvertisedRoutes, ListEvpnRoutes, ListFlowSpecRoutes, GetMetrics | Prefer pagination or narrow filters, set client deadlines, and alert on sustained sensitive_read volume |
| Live streams | EventService.WatchEvents, EventService.SubscribeFromEvent | Keep clients draining, reconnect after stream_lagged, and alert on subscriber count or lag spikes |
| History queries | ListRouteEvents, ListSessionEvents, ListPolicyEvents | Histories are bounded and process-local; use explicit limits for dashboards |
| Mutating calls | Neighbor, policy, peer-group, injection RPCs | Use listener max_tier, role enforcement, and audit alerts on mutating volume |
| Operator-only calls | Shutdown, TriggerMrtDump, SetGracefulShutdown, selected policy/global changes | Restrict to operator principals/listeners and page on unexpected operator_only activity |
Clients should set realistic deadlines on unary inventory queries and avoid opening idle streams that do not continuously read responses. For long-lived streams, use keepalive settings conservatively; aggressive keepalives can create avoidable load and disconnected streams fail like any other RPC. After a lag warning, treat the stream as a live tail after a gap and refresh state with a snapshot or bounded history query.
RFC 7999 BLACKHOLE discards
| Metric | What it tells you |
|---|---|
bgp_blackhole_discard_installed_total | Successful kernel discard-route installation events |
bgp_blackhole_discard_active | Current receipt-authorized discard rows, including adopted rows pending reaping |
bgp_blackhole_discard_withdrawn_total | Successful daemon-owned kernel discard-route removal events |
bgp_blackhole_discard_adopted_total | Startup adoption events for marker-matching kernel discard routes |
bgp_blackhole_discard_reaped_total | Post-startup cleanup events for adopted-but-unclaimed kernel discard routes |
bgp_blackhole_discard_rejected_total{reason} | Pre-install rejection transitions; reason is bounded to broad_prefix, not_ebgp, active_limit_exceeded, or install_rate_limited |
bgp_blackhole_discard_kernel_failures_total{action} | Kernel-operation failure events; action is bounded to setup, install, remove, or dump |
bgp_dataplane_reconcile_planning_failures_total{actor="blackhole_discard",reason} | Pre-kernel planning aborts; reason is bounded to send_failed, reply_dropped, query_failed, or timeout |
Use rbgp rib blackholes for the current per-prefix decision details.
Reconcile planning is bounded to 30 seconds with two-second RIB query slices.
If planning cannot complete, the previous status and all kernel and ownership
state remain unchanged. Under configured install guardrails, observed route
churn may surface route_churn_deferred; it consumes no install token and a
later event or periodic pass retries it.
General Unicast FIB
These metrics are present when the daemon is built with the ADR-0061 general
FIB runtime. The actor is still default-off; configure at least one
[[fib_tables]] block to start it.
| Metric | What it tells you |
|---|---|
bgp_fib_routes_installed_total | Configured-table routes successfully installed or replaced in the Linux kernel |
bgp_fib_routes_withdrawn_total | Daemon-owned configured-table routes successfully removed from the kernel |
bgp_fib_routes_unresolved | Current desired Add/Replace rows held after Linux returned the family-specific route-level unreachable errno for a target made entirely of unscoped, same-family, non-link-local next hops; one uncovered ECMP member can hold the whole route, and relevant route events plus the periodic reconcile trigger retries |
bgp_fib_routes_rejected_total{reason="foreign_route_exists"} | Desired route suppressed because a kernel row already exists at the same table / metric / prefix and is not daemon-owned |
bgp_fib_routes_rejected_total{reason="owned_route_drifted"} | A row rustbgpd previously owned was externally changed; rustbgpd released ownership and preserved the live kernel row |
bgp_fib_routes_rejected_total{reason="next_hop_family_unsupported"} | Desired route suppressed because the table family and BGP next-hop family do not match |
bgp_fib_routes_rejected_total{reason="link_local_next_hop_scope_missing"} | Desired route suppressed because an IPv6 link-local next-hop was selected without the egress interface needed to resolve it in the Linux FIB |
bgp_fib_routes_rejected_total{reason="peer_not_allowed"} | Desired route suppressed by a [[fib_tables]] peer / peer-group allow-list |
bgp_fib_routes_rejected_total{reason="route_limit_exceeded"} | Desired route suppressed because the table exceeded its max_routes hard cap; existing owned rows are frozen in place |
bgp_fib_kernel_failures_total{action="setup"} | Runtime could not open the Linux FIB programming surface at startup |
bgp_fib_kernel_failures_total{action="dump"} | Runtime could not dump configured route tables during a reconcile pass |
bgp_fib_kernel_failures_total{action="install"} | Kernel rejected an add operation for a reason other than a classified unresolved next hop |
bgp_fib_kernel_failures_total{action="replace"} | Kernel rejected a replace operation for a reason other than a classified unresolved next hop |
bgp_fib_kernel_failures_total{action="remove"} | Kernel rejected a remove operation |
bgp_fib_owned_state_persist_failures_total | A write of <runtime_state_dir>/fib-owned.json failed (for example a full or read-only filesystem). While the write fails, route installs and replacements are held with status failed / owned_state_persist_failed:*; removals continue, and the next reconcile retries the write |
bgp_dataplane_reconcile_planning_failures_total{actor="general_fib",reason} | Pre-kernel planning aborts; the last successful status snapshot and all kernel/ownership state remain unchanged |
bgp_kernel_route_notify_dropped_total{actor,reason="channel_full"} | Kernel route-event wake feed dropped an event before the FIB or BLACKHOLE reconciler could consume it; periodic reconcile remains the repair backstop |
bgp_kernel_route_notify_subscription_failures_total{actor,group} | The FIB or BLACKHOLE reconciler failed to subscribe to an IPv4/IPv6 route multicast group and is running with periodic-only kernel-drift repair |
bgp_netlink_subscription_overruns_total{actor} | The kernel reported a receive-buffer overrun to the link_carrier, general_fib, or blackhole_discard NETLINK_ROUTE actor. Each increment is one overrun notification—not a count of lost events—and proves only that one or more multicast events may have been lost |
bgp_session_notification_outstanding | Current lossless session notifications from sender entry through successful PeerManager dequeue. This includes synchronous in-flight reservations plus queued notifications; handling occurs after the value is decremented. |
bgp_session_notification_outstanding_high_watermark | Monotonic daemon-lifetime high-water mark of that outstanding population. It resets only when the daemon restarts and is not an exact per-flap or per-round peak. |
Alert on recent planning aborts without assuming the status RPC was refreshed:
increase(bgp_dataplane_reconcile_planning_failures_total[10m]) > 0Correlate actor and bounded reason with the structured warning; its stage
field is diagnostic context, not a metric label. General-FIB shutdown
cancellation is not counted.
The three netlink-overrun series are materialized at zero. A non-zero value
does not identify which multicast group lost events and does not prove that a
later 10 s link-carrier poll or 30 s route reconcile repaired the resulting
drift. Preserve the counter as incident evidence; rbgp doctor includes all
three series in system/metrics.prom.
Use rbgp rib fib --json as the per-route companion to these counters.
The actor streams bounded Loc-RIB pages and exactly revalidates owned prefixes
that disappear from provisional intent before any removal. A planning timeout
or dropped query causes no kernel dump or state publication. Route churn
freezes Add/Replace for capped tables, while uncapped tables and exact safe
cleanup continue; peer-group churn similarly freezes programming for tables
whose eligibility depends on allowed_peer_groups.
The most important states to investigate are foreign_route_exists and
owned_route_drifted. foreign_route_exists means rustbgpd never proved
ownership of the live row; owned_route_drifted means rustbgpd previously
owned the key but another writer changed the live kernel row. In both cases,
rustbgpd preserves the row instead of overwriting or deleting it. After an
ungraceful restart, rustbgpd only recovers rows that also appear in
<runtime_state_dir>/fib-owned.json, belong to a table whose
[[fib_tables]] signature is unchanged, and still have the exact kernel
next-hop value the previous instance owned. An owned-state file with an
unsupported (newer) version, or a payload whose shape contradicts its version,
is renamed to fib-owned.json.stale and the daemon starts owning nothing. A
changed [[fib_tables]] signature is handled per table: the receipt is copied
(not renamed) to fib-owned.json.stale, only the changed table's rows drop
out of ownership, and unchanged tables keep their owned routes. Rows of a
table recorded during an interrupted runtime table change but absent from the
booted config are withdrawn by the first reconcile.
Route-safety alerts
The shipped Prometheus rule pack alerts only on actionable route-safety event counters, using a 15-minute increase window:
BgpExactExportRejectedmeans a post-policy announcement was withheld before Adj-RIB-Out commit; correlatefamilyandreasonwith the daemon warning andrbgp rib --prefix <prefix> advertised <peer> --explain.BgpMalformedUpdatereports one or more malformed UPDATEs by their strongest RFC 7606disposition.BgpSelectionDeferralTimedOutmeans a family gate used its configured timer fallback before every planned-restart convergence signal arrived.BgpSelectionDeferralLedgerOverflowmeans the bounded identity ledger fell back to a complete release sweep. Safety is preserved, but table scale and retained-key pressure merit investigation.
Graceful Restart
| Metric | What it tells you |
|---|---|
bgp_gr_active_peers | Peers currently in GR stale-route state |
bgp_gr_stale_routes | Routes currently held stale (GR-stale or LLGR-stale) |
bgp_gr_timer_expired_total | GR timers that expired (routes swept) |
bgp_selection_deferral_active{afi_safi} | Planned-restart family convergence/release gate (1 = active); it remains active while collision failback waits for EoRR even after route selection is staged |
bgp_selection_deferral_waiters{afi_safi} | Frozen-roster peers still blocking family convergence/release, including an awaiting_refresh survivor after route selection is staged |
bgp_selection_deferral_releases_total{afi_safi,reason} | Family gates released after all_eor, collision_refresh, all_excluded, or timer |
bgp_selection_deferral_timeouts_total{afi_safi} | Family gates released by the selection-deferral timer |
bgp_selection_deferral_ledger_overflows_total{afi_safi} | Gated families whose next identity would exceed the process-wide one-million-identity or 64 MiB logical retained-key-data ledger and therefore use a complete release sweep |
Retention during the hold window is an Adj-RIB-In property, and the
routing metrics say so. bgp_rib_prefixes keeps counting the peer's retained
(stale) received routes, alongside bgp_gr_active_peers and
bgp_gr_stale_routes. bgp_rib_adj_out_prefixes drops to zero for that peer
instead: entering GR tears down the session's outbound registration and its
Adj-RIB-Out with it (bgp_rib_outbound_registered_peers decrements for the
same reason), so there is nothing left to advertise until the peer returns and
the table is rebuilt. A zero advertised-count for a peer whose received-count
is still high is the expected shape of a restart in progress, not a lost
table.
For active gates, rbgp neighbor <address> and
NeighborService.GetNeighborState also show the peer's waiter state, stamped
session, blocking-waiter count, and remaining time. A released row retains its
reason for the daemon lifetime. A ledger-overflow warning is emitted once per
family; release then enumerates the complete Adj-RIB-In and Loc-RIB family so
withdrawals that already left Adj-RIB-In are still removed before EoR. The
64 MiB bound is deterministic accounting for retained key data (including
nested FlowSpec terms and BGP-LS payload bytes), not a promise about process
RSS, allocator capacity, or hash-table overhead. An identity that lands exactly
on either cap is retained; the family of the next identity that would exceed a
cap enters overflow fallback.
An ordinary same-address replacement remains an EoR waiter, and stale EoR from
its predecessor is rejected. If collision resolution fails registration back
to the exact nonzero, unambiguous survivor, only that survivor enters
awaiting_refresh; other waiters remain blocking. The current Loc-RIB is staged
as soon as ordinary waiters finish, but downstream EoR and route-refresh
responses stay held until a post-failback BoRR is followed by the matching peer
EoRR. Ordinary EoR, stray EoRR, and the local refresh timeout do not satisfy the
waiter. A full survivor outbound channel retains one session-generation-scoped
request and retries it through the bounded RIB resync cadence. A closed channel
retires the request without spinning; the original selection-deferral deadline
remains the bounded fallback.
The active and waiters gauges, plus ordinary all_eor,
collision_refresh, and all_excluded releases, are dashboard context rather
than alert conditions. This avoids paging on normal planned-restart progress;
use rbgp neighbor <address> to correlate the live per-family waiter state.
BFD
| Metric | What it tells you |
|---|---|
bfd_session_up{peer} | Per-peer BFD session state (1 = Up, 0 = not Up) |
bfd_session_flaps_total{peer} | BFD session flaps (transitions out of Up) per peer |
Removing a BFD attachment on SIGHUP removes its BFD metric series while keeping the neighbor's BGP history. An administrative neighbor disable retains the BFD series at Down; a session flap retains its counters.
Config transactions
| Metric | What it tells you |
|---|---|
bgp_config_transaction_lifecycle_total{operation, outcome} | Confirmed config transaction lifecycle transitions. operation is confirm, abort, or auto_revert; outcome is success or failure. |
Alert on outcome="failure": it means an operator abort or timer-driven
auto-revert attempted rollback but could not complete, and
GetConfigTransactionStatus / rbgp config status will hold the redacted last
failed lifecycle record for triage. The counter intentionally omits
confirm_id, candidate content, peer labels, and free-form error text; those
details stay in the structured daemon log and RPC status.
RPKI / ASPA validation
| Metric | What it tells you |
|---|---|
bgp_rpki_vrp_count{af="ipv4"} | IPv4 VRP entries loaded |
bgp_rpki_vrp_count{af="ipv6"} | IPv6 VRP entries loaded |
bgp_rpki_cache_effective_expire_seconds{cache} | Effective RTR expire per cache (IP:port): the cache-advertised expire after the RFC 8210 two-day maximum and the configured max_expire_interval ceiling. Set at client start and after every End of Data |
bgp_rpki_cache_end_of_data_ready{cache} | Per-cache retained End-of-Data readiness: 0 at startup and after flush/expiry; 1 after validated End of Data, including an empty table, and through reconnect/resync |
bgp_rpki_cache_connected{cache} | Per-cache RTR session state: 1 while a session to that configured cache is established, 0 at startup and whenever it is down, whether or not a contribution is still retained (compare bgp_rpki_cache_end_of_data_ready) |
bgp_aspa_records | ASPA customer records loaded in the merged table. Renamed from bgp_aspa_records_total (a gauge must not carry the counter _total suffix) |
bgp_validation_import_refreshes_total{dependency, outcome} | Inbound Route Refresh work triggered by VRP / ASPA cache updates for peers whose import policy matches validation state. dependency is rpki or aspa; outcome is eligible, refreshed, skipped_not_established, skipped_state_unknown, or failed. A state-query timeout increments both skipped_state_unknown and failed and leaves the peer's refresh intent pending for replay. |
A sudden drop in VRP count likely means a cache connection was lost or the
cache itself has stale data. A non-zero failed outcome on
bgp_validation_import_refreshes_total means the cache update arrived, but one
or more validation-dependent peers could not be refreshed immediately; check the
daemon log for the peer-specific reason.
NotFound includes startup/no validated data and after all applicable retained
cache contributions flush or expire. The readiness gauge cannot be matched in
policy.
For a point check against the daemon's current authoritative table, run:
rbgp rpki validate 203.0.113.0/24 64496
rbgp rpki aspa 64497
rbgp rpki verify-path --role peer --neighbor-asn 64496 "64496 64497"The verdict always uses the complete table. The accompanying list is capped at
256 effective covering VRPs and reports complete plus exact omitted; AS0
rows are shown but never authorize. precondition failed means no first
authoritative snapshot has arrived, while not_found with a complete empty
list means an authoritative snapshot exists but contains no covering VRP. This
command does not diagnose individual cache readiness or provenance; use the
per-cache readiness metrics and RTR logs for those questions.
EVPN VTEP alpha
| Metric | What it tells you |
|---|---|
evpn_local_originations_total{action="inject"} | Locally learned MACs that the originator successfully handed to the RIB as Type 2 advertisements |
evpn_local_originations_total{action="withdraw"} | Locally aged / deleted MACs that the originator successfully handed to the RIB as Type 2 withdraws |
evpn_local_origination_errors_total{action="inject"} | Failed local Type 2 inject attempts: RIB channel closed, RIB rejected the inject, or the reply was dropped |
evpn_local_origination_errors_total{action="withdraw"} | Failed local Type 2 withdraw attempts: RIB channel closed, RIB rejected the withdraw, or the reply was dropped |
evpn_local_observations_dropped_total{reason="channel_full"} | Kernel local-MAC observations classified by the netlink notify loop but dropped because the originator channel was full |
evpn_local_observations_dropped_total{reason="channel_closed"} | Kernel local-MAC observations classified by the netlink notify loop after the originator receiver was gone |
evpn_duplicate_ip_moves_total{vni} | Conflicting IPv4/IPv6 ownership observations counted by the optional detect-only IP detector |
evpn_duplicate_ip_threshold_exceeded_total{vni} | Per-IP M/N threshold crossings; each crossing logs the VNI, IP, observed MAC, threshold, and window without changing routes |
evpn_duplicate_mac_moves_total{vni,mac} | Cross-VTEP MAC mobility contention events detected by the local originator |
evpn_duplicate_mac_first_move_timestamp_seconds{vni,mac} | Unix timestamp of the first observed duplicate-MAC / mobility contention event for that key |
evpn_duplicate_mac_threshold_exceeded_total{vni,mac,action} | RFC 7432 §15.1 M/N threshold crossings. action is detect or suppress_local from the per-instance config |
evpn_duplicate_mac_quarantine_active{vni,mac} | 1 while action = "suppress_local" is actively suppressing local Type 2 originations for that key; returns to 0 after timed recovery |
evpn_ip_vrf_observed_routes{vrf} | Kernel routes currently eligible for Type 5 origination from each configured IP-VRF |
evpn_ip_vrf_observed_routes_filtered_total{vrf,reason} | Kernel routes rejected by the bounded IP-VRF observation classifier |
evpn_ip_vrf_origination_suppressed_total{vrf,reason} | Type 5 candidates withheld because the IP-VRF is not ready or the address family does not match |
evpn_ip_vrf_originated_routes{vrf} | Locally originated Type 5 routes currently advertised from each IP-VRF |
evpn_ip_vrf_installed_routes{vrf} | Remote Type 5 routes currently owned in each IP-VRF kernel table |
evpn_ip_vrf_remote_prefix_drops{vrf,reason} | Current remote Type 5 projection drops by bounded IP-VRF/reason labels. Overlay-index reasons include overlay_index_no_linked_l2vni, unresolved_overlay_index_gateway, and ambiguous_overlay_index_gateway; vrf="_unscoped" means the drop happened before a configured IP-VRF could be selected. |
evpn_foreign_replaces_blocked_total | Installs or replaces withheld because an exact-key foreign kernel row already exists |
evpn_foreign_deletes_skipped_total | Deletes skipped after the live kernel row ceased to be rustbgpd-owned |
evpn_foreign_owned_relinquished_total | Owned keys released after another writer replaced the live kernel row |
evpn_fdb_nhg_drift_members_repaired_total | Per-VTEP nexthop members repaired after kernel drift. Repeated growth means recovery is working but another writer or unstable kernel state needs investigation; correlate with dataplane logs and ip nexthop show. |
evpn_fdb_nhg_drift_groups_replaced_total | FDB nexthop groups re-created or replaced after kernel drift. Investigate sustained increases alongside member repairs for competing kernel writers. |
evpn_fdb_nhg_orphans_cleaned_total | Unreferenced rustbgpd-tagged FDB nexthops removed by ownership-aware garbage collection. Sustained growth outside restart recovery indicates unstable group membership or cleanup churn. |
evpn_fdb_nhg_drift_disabled_total | Permanent dump failures that disabled FDB-NHG drift recovery for this daemon lifetime. Any increase means automatic repair is unavailable; use the accompanying error to correct kernel support, privilege, or netlink-message failures, then restart the daemon. |
evpn_single_active_backup_active | Distinct (VNI, ESI, Ethernet Tag) single-active groups currently retargeted to a backup PE during the post-failover window |
For an instance with duplicate_ip_detection.enabled = true, investigate an
increase in evpn_duplicate_ip_threshold_exceeded_total alongside the
EVPN duplicate-IP threshold exceeded; action is detect-only warning. The
warning identifies the IP and observed MAC; compare the local neighbor table
with received EVPN Type 2 routes before changing host addressing. Counters
are cumulative per VNI rather than per-IP labels. A normal one-off host
rebind can increment evpn_duplicate_ip_moves_total; it need not indicate a
persistent duplicate. The detector neither suppresses routes nor provides an
IP clear/quarantine command. See the configuration reference
for defaults, exclusions, and window semantics.
During M37 or a synthetic MAC-churn soak, the inject and withdraw counters
should follow the bridge fdb add / bridge fdb del cadence. Any non-zero
observation-drop counter means the kernel event reached the notify loop but
not the originator; any non-zero origination-error counter means the
observation reached the originator but did not complete at the RIB boundary.
evpn_duplicate_mac_moves_total and
evpn_duplicate_mac_first_move_timestamp_seconds are intentionally per
(VNI, MAC); alert on threshold crossings rather than on one-off
mobility during planned host moves. Default duplicate_mac_detection
behavior is detect-only. When an instance opts into
action = "suppress_local", active quarantine withdraws/suppresses only
locally-originated Type 2 routes for that MAC and automatically retries
after recovery_seconds; remote EVPN route visibility stays intact, while
dataplane receive-side intent for the quarantined key is filtered out of
the local FDB reconciler. After confirming the loop condition is gone, an
operator can clear one active quarantine immediately:
rbgp evpn clear-duplicate-mac --vni 100 --mac aa:bb:cc:dd:ee:ffThe clear path returns success with cleared=false if no active quarantine
exists. When it clears an active key, the originator resets the active gauge
to 0, republishes the quarantine set, and replays still-live local MAC or
MAC+IP state through the normal recovery path.
rbgp evpn instances also reports each L2VNI's
readiness=ready|not-ready|unbound|unknown and
originated-local-macs=N; rbgp evpn instances --json exposes the
same values as readiness, not_ready_reason, and
originated_local_macs_count.
Key log messages
rustbgpd uses structured logging: JSON under log_format = "json" (the
edge and route-server profile default) or human-readable text under
log_format = "text" (the lab profile default). Key messages to watch for:
| Message | Level | Meaning |
|---|---|---|
starting rustbgpd | INFO | Daemon started successfully |
session established | INFO | BGP session reached Established |
session down | INFO | BGP session left Established |
TCP connect failed | INFO / WARN | First socket failure in a failed-connect episode; WARN after an Established epoch, INFO for a peer that has not established yet |
TCP connect task failed | WARN | First internal connect-task failure in a failed-connect episode |
received SIGTERM / received SIGINT | INFO | Process signal received |
shutdown initiated via gRPC | INFO | Shutdown RPC called |
termination signal received during coordinated shutdown; skipping waits that have no deadline | WARN | A further SIGINT/SIGTERM arrived after shutdown began with no settlement owner; waits without a deadline are skipped |
termination signal received during coordinated shutdown; the owned runtime-config settlement is already fenced and the daemon will fail-stop | WARN | A further SIGINT/SIGTERM arrived after the owner had already been fenced; fence_reason names why, and exit 70 was already in flight |
termination signal received during coordinated shutdown; the owned runtime-config settlement is fenced and the daemon will fail-stop | ERROR | A further SIGINT/SIGTERM arrived while a runtime-config owner was still settling; it is fenced as operator_forced and exit 70 follows after the grace |
runtime config coordinator permit is still held outside settlement ownership | ERROR | Shutdown stopped waiting for an unowned coordinator permit (reason is deadline_expired or second_signal) and continues without the warm checkpoint |
SIGHUP reload task is still running with no settlement owner | ERROR | Shutdown stopped waiting to join the reload task and continues |
gRPC server exited unexpectedly | ERROR | Fatal — coordinated shutdown follows |
RIB manager exited unexpectedly | ERROR | Fatal — coordinated shutdown follows |
peer manager task exited unexpectedly | ERROR | Fatal — coordinated shutdown follows |
RPKI subsystem task exited unexpectedly | ERROR | Fatal — coordinated shutdown follows |
BGP listener task exited unexpectedly | ERROR | Fatal — coordinated shutdown follows |
BGP accept-forwarding task exited unexpectedly | ERROR | Fatal — coordinated shutdown follows |
metrics/readiness server exited unexpectedly | ERROR | Fatal — coordinated shutdown follows |
listener accept failing; backing off | ERROR | A gRPC TCP, gRPC UDS or metrics listener hit resource exhaustion (EMFILE, ENFILE, ENOMEM or ENOBUFS). Accepts back off from 100 ms, doubling to a 1 s cap; the line is logged once per episode and then every 60th failure, with listener, failures and backoff_ms |
listener accept recovered | INFO | The next connection was accepted after a backoff episode; failures counts the episode |
listener socket unusable; stopping its accept loop | ERROR | A management listener socket itself failed (for example EBADF or EINVAL). The daemon then fail-stops through gRPC server or metrics/readiness server supervision |
config reload complete | INFO | SIGHUP reload completed; the generation route logs config reload complete (one runtime generation) |
SIGHUP reload rejected without runtime effect | ERROR | The candidate was rejected before any effect, or a generation-route failure restored the prior generation; the candidate file is unchanged |
reload generation failed; the peer manager restored the prior generation and the candidate file is left for correction | ERROR | A generation-route step failed after effects began and compensation restored the prior generation |
reload generation left runtime state uncertain; fencing | ERROR | Generation-route compensation could not be proved; the daemon recovery-fences |
config reload stopped at this step; settling the acknowledged partial runtime authority before another reload may begin | ERROR | A reload step failed after the coordinator established the resulting authority; inspect the structured bucket, target, and error fields |
GR restart marker | INFO | Restart marker written or read |
published GR restart marker with wall-clock fallback because boottime protection was unavailable | WARN | Clock-domain sampling or representation failed; a complete bounded v1/v2 marker was selected. Check publication_durability on the final publication log for directory-sync status. |
max-prefix limit exceeded | WARN | Peer exceeded prefix limit |
gRPC TCP listener bound to a non-loopback address | WARN | Security posture warning |
Later socket or task retries in the same failed-connect episode stay at DEBUG
to avoid one default-visible record per ConnectRetry interval. Every retry still
refreshes the neighbor's Last Error; a successful TCP connection re-arms both
records for a later outage.
Debugging a session that won't establish
-
Check peer state:
rbgp neighborLook at the FSM state.
Activemeans we're trying to connect but TCP isn't establishing.OpenSent/OpenConfirmmeans OPEN exchange is failing.Idlewith aReconnect Inrow inrbgp neighbor <addr>means the daemon is waiting out the deferred reconnect that follows an unplanned teardown. Each consecutive NOTIFICATION teardown (sent or received, including an OPEN exchange that ends in one, such as an ASN mismatch or a repeated hold-timer expiry) doubles that wait, starting fromconnect_retry_secs(default 5 s) and capped at 300 s. The wait is alsoreconnect_in_secondsin the JSON output and inNeighborState. The API field is absent on older daemons, while JSON omits it when absent or zero. The wait covers both directions. From the second consecutive NOTIFICATION teardown, an inbound connection from the neighbor during the wait is closed without an OPEN and counted asbgp_inbound_connections_dropped_total{reason="notification_backoff"}, with a throttled log line naming the streak and the remaining wait; a pending collision candidate is dropped instead of promoted. After the first teardown an inbound connection is still accepted at once, and the session that replaces the waiting one keeps its streak, so a neighbor that always reconnects to us escalates the same way. The streak clears after the session stays Established for five minutes, onrbgp neighbor <addr> enable, or on an administrative reset. TCP misses inConnectorActivekeep the fast initial retries followed by the existing exponential curve; a TCP loss afterEstablisheduses the configured fixed Idle interval. Neither advances the NOTIFICATION streak. A max-prefix shutdown stays latched until an explicit enable. There is no configuration knob for the NOTIFICATION curve. -
Check logs for the peer:
journalctl -u rustbgpd | grep "10.0.0.2"Look for NOTIFICATION codes, capability mismatches, or hold timer expiry.
-
Common causes:
- TCP not reaching: Firewall, wrong address, peer not listening on 179
- ASN mismatch: Remote peer has a different
remote-asconfigured for us - Router ID collision: Two speakers with the same router ID
- Hold timer zero vs non-zero: One side sends hold_time=0, the other expects keepalives
- Capability mismatch: Check address family negotiation in OPEN logs
- MD5 mismatch: TCP RST with no BGP-level error; check both sides' passwords
- TTL security: Both GTSM speakers must transmit TTL / Hop Limit 255.
With
ttl_security_hops = N, rustbgpd accepts packets at or above255 - (N - 1)(256 - N); the knob changes only this receive floor and cannot make a peer that transmits an ordinary lower TTL compatible. Bounded multihop is supported, with distance-aware GTSM configured on both sides. Packets below rustbgpd's receive floor are discarded by the kernel before BGP processing and therefore produce no BGP NOTIFICATION.
There is no separate
ebgp_multihopmode that disables GTSM. Without GTSM, rustbgpd uses the kernel default outbound TTL and does not require an eBGP peer to be directly connected. That permits routed peers, but an off-subnet address typo is attempted instead of failing an intent check.GTSM setup failures are socket-scoped: an active open fails before connect, a passive accepted socket is discarded, and an affected listener-family socket is rejected. The process fails startup only if no listener family remains usable or strict explicit endpoint binding fails; do not interpret one family/socket rejection as an unconditional daemon startup failure.
-
Verify from the remote side: Check FRR/BIRD/peer logs for their view of the session attempt.
Support bundles with rbgp doctor
rbgp doctor runs red/green triage checks live (daemon reachable and
healthy, peers stuck outside Established with time-in-state, flap loops,
TCP-AO configuration against the daemon's kernel capability probe, RFC 8212
directional policy presence, daemon nofile rlimits, recent panic reports)
plus first-deploy environment probes, and writes one redacted
rustbgpd-doctor-<ts>.tar.gz:
Each lightweight gRPC collection call has a 30-second response deadline,
including response-body transfer, for both TCP and Unix sockets. A timeout
names the RPC and budget, preserves successful evidence, and marks the
affected section incomplete. The effective-config collection has a separate
30-minute-and-30-second allowance; doctor can therefore take longer than
30 seconds overall. No timeout triggers an automatic retry; the one retry
doctor makes is described under daemon.healthy below. Native one-shot
CLI reads, including config status and history, use the same 30-second limit
per call or RIB page. rbgp config effective uses the effective-config
allowance. Long-running config diff and plan, config mutations and streams,
and live watches keep their existing budgets and lifetimes.
daemon.healthy reads GetHealth. When the daemon answers UNAVAILABLE
(a core actor missed the 200 ms probe deadline), doctor waits one second and
asks once more. The miss may be transient, as briefly while a reload settles,
or persistent, from a wedged actor. A healthy second answer is yellow and
names the first miss; a second miss is red. Any other
error, including doctor's own 30-second response deadline, is red at once
without a retry, as is a snapshot reporting healthy=false.
peer.<addr>.rfc8212_policy is the ADR-0112 check. It is green for
not_required — the compatibility default, so it never turns an existing
deployment yellow — and for present. It is red when a direction reports
missing, naming import, export, or both: the reserved internal deny is
installed there and no route crosses it. It is yellow against a daemon that
does not expose the status, because "no evidence" is not "no problem". A red
verdict here never affects /readyz: missing operator policy is a
configuration state to repair, not a daemon that cannot serve traffic.
rpki.invalid_route_policy and aspa.invalid_route_policy are advisory compiled-policy
disposition proofs. Green requires a complete, nonempty response whose aggregate and every
scope are enforced; it does not claim readiness/currentness, traffic, intent, FIB state, or
runtime enforcement. Missing or uncertain evidence is yellow, including UNIMPLEMENTED
from an older daemon, and these warnings retain exit status 0.
Doctor treats a present max-prefix restart countdown, including 0ms, as an
intentional yellow hold-down rather than a stuck-session failure; stale
session evidence still takes precedence. An active outbound prefix-limit
blocking episode is red as
peer.<scoped-address>.outbound_prefix_limit.<family> because that peer is
intentionally withholding routes until capacity recovers. Merely configured,
unlimited, and nonblocking family rows do not add checks. An open inbound
max_prefix_action = "block" episode is yellow as
peer.<scoped-address>.inbound_prefix_limit.<scope>: the session stays
Established while that peer's net-new prefixes are withheld. Nonblocking
inbound scopes add no checks. Link-local identities
retain their %interface scope in both check names and details.
Doctor parses the effective-config document once and reads probe targets from
its documented key paths. With global.listen_addresses, a local check uses
every configured address, including when management uses a Unix socket; it
does not substitute loopback. With the default wildcard listeners, either
IPv4 or IPv6 loopback reaching the listener is sufficient, matching the
daemon's tolerance of an unavailable address family. Local checks identify
management through a Unix socket, localhost, or a loopback IP. These probes
assume that connection reaches a daemon in the CLI's network namespace;
forwarded management sockets do not prove that shared vantage.
From a remote CLI, listener connections are reachability observations; a
failure cannot distinguish a failed bind from routing or filtering. Explicit
daemon loopback addresses cannot be checked remotely. Run doctor on the
daemon host to verify these binds, including for --pre-upgrade diagnostics.
On Linux, local process limits and config freshness belong only to the peer
identified by the connected Unix-domain socket. Doctor checks the process start
time before using /proc evidence and refreshes identity on reconnect. Config
paths and freshness markers are read through that process's filesystem view. Other
rustbgpd instances on the host do not contribute these checks or limits files.
TCP connections and unavailable local identity leave process limits unavailable.
When effective config cannot be read, doctor uses the verified local peer's
config path; an unreachable Unix socket can still use the packaged default file
as local first-deploy input. TCP does not fall back to a CLI-host config file.
First-deploy checks (network probes are bounded to a 2s timeout; all are read-only):
| Check | What it probes | Red/yellow advice |
|---|---|---|
bgp.listener | Daemon up: TCP connect to the configured BGP addresses and listen port. Daemon down: test-bind the port and release it | A failed local explicit-address probe is red; remote CLI reachability failures are yellow. CAP_NET_BIND_SERVICE is needed for ports below 1024; port-in-use on a test-bind is yellow |
rpki.vrp_table | With configured caches and a reachable daemon, requires a nonzero complete IPv4 + IPv6 bgp_rpki_vrp_count snapshot and retained accepted complete End-of-Data readiness for every configured cache | yellow when the merged table is zero/missing/malformed/unavailable or a configured cache is not ready/missing from the readiness snapshot; this is retained readiness, not current RTR connectivity |
rpki.cache.<addr>.session | With configured caches and a reachable daemon, the daemon's ListCaches row for each [rpki] cache_servers entry | yellow when the RTR session is down (the detail says whether a contribution is still retained and its age), when the cache has no inventory row, or when ListCaches fails, for example because the token cannot read RPKI cache state |
rpki.cache.<addr>.reachable_from_cli | TCP connect from the rbgp process to each [rpki] cache_servers entry | yellow on failure because this is CLI-network-vantage evidence, not daemon-side connectivity; use rpki.cache.<addr>.session and rpki.vrp_table for the daemon's state |
bmp.collector.<addr>.reachable_from_cli | TCP connect from the rbgp process to each [bmp] collectors entry | yellow on failure because the daemon may have a different network vantage; inspect rustbgpd and collector logs for actual export state |
gnmi_dialout.<name>.reachable_from_cli | TCP connect from the rbgp process to each [gnmi_dialout] targets entry | yellow on failure because the daemon may have a different network vantage; inspect gnmi_dialout_connected and daemon logs for actual dial-out state |
state_dir.writable / state_dir.disk | runtime_state_dir writability and free space (yellow < 1 GiB, red < 100 MiB) | journal, MRT dumps, crash reports, and the event-history DB write there |
host.run_context | systemd / container / unknown from pid-1 facts | tailors remediation lines (e.g. LimitNOFILE= vs container ulimits) |
daemon.config_freshness.<pid> | whether config file mtime is newer than the daemon's last config-file marker (process start fallback) | yellow means on-disk changes may be pending; green does not prove effective runtime agreement |
Probe targets come from the daemon's effective config when it is up; when
it is down, from the local config file (the path a local daemon process
was started with, else /etc/rustbgpd/config.toml) — parsed for
addresses only, never copied into the bundle.
Pre-upgrade checks (--pre-upgrade CONFIG)
rbgp doctor --pre-upgrade CONFIG adds three read-only checks against
CONFIG, the file the upgraded daemon will boot, under the same 0/1/2 exit
contract. Green is an observation at one instant, never a maintenance fence;
missing evidence is red, never green. The mode resolves nothing on the
operator's behalf.
| Check | Evidence | Red when |
|---|---|---|
upgrade.transaction | GetConfigTransactionStatus, the same RPC as rbgp config status | a confirmed transaction is pending or applying (confirm or abort it: rbgp config confirm <id> / rbgp config abort <id>), rollback-failed (retry abort, confirm, or restart to boot-revert), or ambiguous (restart to boot-revert); the RPC is denied, unimplemented, or fails. Each pending, applying, rollback-failed, or ambiguous detail quotes the daemon's status text as rbgp config status prints it, so an overdue automatic rollback shows that it is waiting for the runtime-config coordinator. Green only for an empty "none" record or a terminal confirmed / aborted / auto_reverted outcome |
upgrade.settlement | bgp_runtime_config_settlement_active from the metrics doctor already collects; the daemon emits it only while an owner is live or recovery-fenced | any series is 1 (wait for the transaction, neighbor or FIB change, or SIGHUP reload to settle; a fence_reason other than none means an exit-70 restart is coming); the metrics RPC failed |
upgrade.posture | CONFIG and the daemon's effective config, both resolved with the config_epoch omitted-versus-explicit rules | the effective epoch/posture pair differs (restarting on the file changes RFC 8212 behavior: keep it only with explicit policy on every eBGP direction, or restore the live tuple after the stop: pin-legacy writes (1, false), prepare-secure writes (2, true); other tuples require setting config_epoch and [global] ebgp_requires_policy explicitly. Repeat candidate --check --strict after any rewrite); CONFIG is unreadable or invalid; the effective config is unavailable |
--json output gains a pre_upgrade object (candidate_config,
observed_at_unix_seconds, ok) and the manifest a pre_upgrade section;
human output ends with the dated observation and the next step. Nothing else
in the doctor output changes outside the mode.
rustbgpd-doctor-<ts>/
├── manifest.json # versions, redaction note, per-section collected/partial/unavailable status, check results
├── config/effective.toml # GetEffectiveConfig dump (resolved defaults, selected empty lists omitted, secrets <redacted>; 384 MiB max)
├── peers/bfd.json # BFD state/diagnostic/strict plus remote-AdminDown bool or null when unknown
├── peers/dynamic-neighbors.json # configured acceptance ranges; descriptions scrubbed client-side
├── peers/neighbors.json # per-peer state, counters, flap/slow-peer status
├── peers/events.json # recent session + policy events (free text scrubbed)
├── logs/tail-1000.jsonl # only with --log-file; see below
├── crashes/panic-*.toml # panic reports swept from <runtime_state_dir>/crash/
└── system/ # environment.json, health.json, global.json, metrics.prom, daemon rlimitssession_events and policy_events are reported independently in the
manifest. A failed history RPC marks only its source partial; the bundle
keeps the successful peer, BFD, and other event-history evidence.
Dynamic-neighbor range inventory is likewise independent evidence. With no
active neighbor sessions, doctor distinguishes zero configured ranges from
configured ranges waiting for a future inbound connection. Zero-and-zero keeps
the existing yellow first-deploy warning; dormant configured ranges are
inventory evidence, not a health failure. If ListDynamicNeighbors fails, the
manifest records the inventory as unavailable and the check stays yellow
instead of fabricating a zero-range result.
Exit codes: 0 no checks red (green and yellow warnings may be present), 1
error, 2 bundle written but one or more checks are red. A down-daemon run produces a
bundle (system facts + crash reports) and the manifest records which
sections are missing.
Without --output the bundle is written to the first writable of: the
working directory, runtime_state_dir, the temp directory — so the
container image, whose working directory is not writable by its nonroot
user, still produces one. The final path is printed on the last line
(bundle in --json). --output <FILE> overrides the choice.
For live inspection, rbgp neighbor <address> reports the effective protected
transport and an explicit TCP-AO health state. Direct dynamic-prefix sessions
derive TCP-AO identity from their validated accepted socket rather than a
per-neighbor key configuration. unavailable means TCP-AO protection is
expected but no socket inspection snapshot is available (the peer may be
disconnected, connecting, or socket inspection may have failed).
Persistent disagreement between TCP_AO_INFO and TCP_AO_GET_KEYS is also
unavailable: rustbgpd clears the whole snapshot instead of publishing an
inconsistent degraded value. healthy means the published live snapshot has
valid current/RNext keys mapped to a nonempty, internally consistent live MKT
inventory, neither active key is deprecated, and there are no authentication
error counters; degraded means a key-validity flag is missing, an active key
is deprecated, or at least one cumulative socket-lifetime error counter is
non-zero. When socket inspection succeeds, connected sessions
also show current/RNext KeyIDs,
packet verification counters, and redacted per-key peer/prefix, directional
IDs, algorithm, selection flags, rollover metadata, and counters. Key bytes,
lengths, hashes, and fingerprints are never returned. TCP_AO_GET_KEYS does
copy raw key bytes into a private temporary buffer; rustbgpd compares accepted
sockets against the complete configured keyring without logging the bytes and
zeroizes every temporary on success, error, retry, and unwind.
Longer-lived TCP-AO keys and TCP-MD5 passwords owned by the internal API,
peer-manager, and transport runtime are also redacted and each clone zeroizes
its allocation when replaced or dropped. The temporary Linux tcp_md5sig
record scrubs its key buffer and length after setsockopt. Parsed TOML,
configuration snapshots, protobuf messages, compiler-created copies, and
kernel-owned keys retain their own lifetimes and are not covered by that
guarantee.
The same neighbor view reports TCP-AO rotation desired, applied, and
phase values. idle means the peer has acknowledged the desired immutable
inventory; add_only is an in-progress successor install; and
add_only_failed includes a secret-free actionable error. selecting means an
installed successor is being assigned as local RNext; awaiting_peer means the
one-shot observation did not yet see the peer use that successor; and
selection_failed is a hard inventory/counter/commit failure. deleting
means an exact deprecated/unselected-MKT deletion generation is moving from the
listener through queued accepted children and every protected primary/pending
session; delete_failed retains that immutable generation and its secret-free
error. An identical SIGHUP retry is safe before mutation, after successful
exact-prior restoration, or when the listener already reached desired. Failed
restoration leaves a non-resumable intermediate inventory: a retry must first
re-prove the exact prior inventory or it is rejected before mutation, and a
daemon restart is required if the inventory remains partial or unprovable.
While awaiting,
desired=N and applied=N-1; a later SIGHUP must present the identical full
desired config and retries that same N. Selection never sets Linux Current,
or commits predecessor deprecation before every affected session shows a
generation-relative increase in the successor's verified-packet counter.
Deletion removes only deprecated MKTs that are neither Current nor RNext and
preserves owner identity, survivor order, key definitions, and the selected
MKT. If listener mutation may have started, affected protected passive accepts
can reject until the same generation is retried or the daemon restarts. If any
changed session may have mutated, the complete changed cohort is discarded; an
incomplete reset aborts every affected session task. A child queued before a
successful listener deletion is repaired only when it exactly matches the
complete immediately previous owner-union inventory. Partial inventories and
older queued generations remain rejected.
The optional per-key vrf_ifindex is Linux's VRF L3-master key selector, not
an IPv6 link-local interface scope. rustbgpd currently installs VRF-unbound
TCP-AO MKTs (TCP_AO_KEYF_IFINDEX clear), so Linux may match them in any L3
master. Scoped link-local routing remains attached to the TCP socket,
and configuration validation forbids reusing the same link-local neighbor
address across interfaces.
Counters and the key inventory are refreshed from the socket for each query;
inspect last_error for setup or connect failures.
What is never collected: the raw daemon config file (the config section is
the daemon's own secret-redacted effective dump — the same document as
rbgp config effective) and bearer-token material. Metrics, event free
text, peer descriptions, crash reports, and log lines are additionally
scrubbed for password/secret/token/bearer lines client-side.
Logs: the daemon logs to stdout (journald under systemd), as JSON or text
per log_format, so no log file is collected by default — the manifest
records that instead. If stdout is redirected to a file, pass it explicitly:
rbgp doctor --log-file /var/log/rustbgpd.jsonl # tails the last 1000 linesCrash reports: a daemon panic writes a small TOML report (panic message,
source location, thread, version — never environment variables or argv)
to <runtime_state_dir>/crash/panic-<ts>-<pid>-<n>.toml, keeping the 10
most recent. Each report is written to a temporary file and renamed into
place, so a report is either complete or absent. rbgp doctor sweeps them into crashes/; a report in a bug
ticket usually pinpoints the crash without a core dump.
Attach the tarball to bug reports — the GitHub bug-report template asks for it.
Common operational tasks
Check session status
rbgp neighbor # summary table (alias: rbgp summary)
rbgp neighbor --wide # adds Source, MsgRcvd, MsgSent, Flaps, RRC, Slow, State/PfxRcd--wide adds a Source column before the classic vendor summary columns.
Static peers say static; dynamic peers name the canonical accepted prefix and
peer group captured at connection time. An older daemon that exposes only
is_dynamic says dynamic (range unavailable) rather than guessing from the
current matcher. The remaining columns are total messages received/sent (all
types, daemon-lifetime — an Established session whose counters stop moving is
wedged), flap count, an RR-client marker, and the overloaded State/PfxRcd
column (a number means Established and shows the prefixes received). Its
Slow column marks a slow-but-established peer with !. Display-only: JSON is
unaffected by --wide and may omit optional false healthy-state fields.
The same captured value appears as Peer Source in
rbgp neighbor <address> and as Source in the TUI peer detail.
Neighbor detail includes an Effective Posture block for the resolved running
NEXT_HOP ownership, RFC 1997 interpretation, route-server control-community
handling, and ORR vantage. This is the quickest inheritance audit for both
static and accepted dynamic peers. During a rolling CLI upgrade,
unknown (not exposed by daemon) means the server predates this nested state;
it must not be interpreted as an explicitly disabled posture. JSON preserves
the distinction by omitting effective_posture.
rbgp neighbor <address> reports actor-authoritative negotiation state only
for the current Established session: negotiated hold time, remote router ID,
four-octet-AS result, mutual families, and usable peer Graceful Restart family
coverage. When GR coverage is usable it also shows the peer-advertised Restart
Time and the effective initial disconnected retention after
gr_peer_restart_time_max; the effective value is absent when the local GR
helper is disabled. It does not substitute configured families or timers.
Human output distinguishes unknown (not exposed by daemon) during a rolling
CLI upgrade, unknown (stale state) after a timed-out actor query,
unavailable (session not Established), unsupported for negotiated families, peer capable; disabled locally, and active helper state. JSON and
gRPC expose the same distinction with optional negotiation_available, an
optional negotiated_session object, an optional nested graceful_restart
object, and optional effective_retention_time_seconds. Scalar presence is
meaningful: a negotiated hold or Restart Time of zero and
four_octet_as = false are explicit values, not missing data.
At startup, the topology banner reports configured static neighbors separately from dynamic-neighbor acceptance ranges. A dynamic-only configuration reports the range count; zero static neighbors and zero ranges is identified as unconfigured rather than mislabeled dynamic-only.
Add a peer at runtime
rbgp neighbor 10.0.0.5 add --remote-asn 65005 --description "new-peer"
rbgp neighbor 203.0.113.2 add --remote-asn 65002 --role provider --strict-roleThe peer is persisted to the config file automatically. --role enables RFC
9234 BGP Roles / OTC route-leak protection for static eBGP peers; the optional
--strict-role flag rejects peers that do not advertise a compatible Role.
Remove a peer
rbgp neighbor 10.0.0.5 deleteSends NOTIFICATION, tears down the session, removes from config.
Manage dynamic-neighbor ranges at runtime
rbgp dynamic-neighbor list
rbgp dynamic-neighbor add 10.0.0.0/24 --peer-group ix-members
rbgp dynamic-neighbor delete 10.0.0.0/24Adds or removes [[dynamic_neighbors]] accept-prefix ranges without a
restart (AddDynamicNeighbor / DeleteDynamicNeighbor, tier mutating).
add validates exactly like config load: the peer group must exist and must
not enable BFD, the prefix must be valid, and the effective prefix must not
duplicate an existing range. delete stops future accepts only —
already-established dynamic peers keep running and drain when they next
return to Idle. Omitting --remote-asn uses the accept-any sentinel (remote_asn = 0);
once a peer is accepted, operational state surfaces the ASN learned from the
peer's OPEN. Changes persist to the TOML file (atomic write) before the RPC
returns and survive a restart.
The live mutation path is serialized with SIGHUP reload, so a reload cannot
drop an accepted-but-not-yet-persisted range.
When rbgp neighbor has no live neighbor rows, human output queries the
range inventory and distinguishes an unconfigured daemon from configured
ranges that have not accepted a peer yet. JSON compatibility is unchanged:
the same empty live-neighbor result is exactly [], with no range lookup or
extra inventory fields; use rbgp dynamic-neighbor list -j for range JSON.
Soft reset (re-evaluate import policy)
rbgp neighbor 10.0.0.2 softresetRe-applies import policy to all routes from this peer without tearing down the session.
Note: as of v0.12.0,
update_runtime_policiesautomatically issues a Route Refresh whenever a peer's effective import chain materially changes (via SIGHUP reload, gRPCSetPolicy,SetPeerGroup, or chain mutations). Operators only need this command after manual ad-hoc edits or to recover from a session-mid-restart at the time of the original reload. Thepending_refreshretry semantics onManagedPeercover most of those edge cases automatically.
Refresh outbound (re-send a peer's Adj-RIB-Out)
rbgp neighbor 10.0.0.2 refresh-outThe outbound sibling of softreset: re-emits this one peer's current
exportable routes through the live export path
(NeighborService.RefreshOutbound) without tearing down or renegotiating
the session. Useful when a peer is suspected of having missed or dropped
advertisements and you want to reconverge it without a flap.
Replay outbound unicast with a completion boundary
rbgp neighbor 10.0.0.2 replay-outThis experimental operation schedules one Established peer's negotiated
IPv4/IPv6 unicast inventory through the live BGP export path and appends
terminal UPDATE EoRs, including for empty families. It requires an eligible
connected rib_out_post BMP collector. There is no all-peer or family-filter
form; serialize full-table use because replay sends an O(table) UPDATE burst
on the live session.
The session must negotiate only IPv4/IPv6 unicast families, with at least one family, and its IP address must be unique among managed peers. A single scoped IPv6 peer is supported; repeated link-local addresses on different interfaces are refused. These prerequisites are checked before monitoring reset or replay traffic, because the reset clears the entire cached peer inventory.
Eligible collectors have a monitor list that includes rib_out_post and
omits rib_in_pre (for example monitor = ["rib_out_post"]; the default
["rib_in_pre"] is not eligible).
Before replay, each receives a monitoring-only Peer Down (reason 5) followed
by the current Peer Up, clearing its previous peer inventory. The BGP session
stays established. Mixed inbound/outbound collectors are excluded because
this operation cannot rebuild their inbound inventory. After enrollment,
ordinary initial-table and peer-refresh unicast EoRs are not mirrored for the
rest of this BGP writer generation, including for collectors excluded from
enrollment. The actual BGP EoRs are still sent. Further complete BMP boundaries require
another explicit replay; a new BGP session restores ordinary EoR mirroring.
The CLI reports scheduling only. Terminal BMP EoRs mark completion after the
session writer completes the replay and terminal wire EoRs; they do not prove
collector receipt or remote BGP processing. A peer or collector generation
change, writer failure, or timeout can leave a scheduled replay incomplete.
Treat missing terminal EoRs as incomplete evidence. Older daemons return
UNIMPLEMENTED; the CLI does not substitute refresh-out.
See the RPC contract
for refusal conditions and the five-second operation bound. This RPC is
outside the v1 contract; refresh-out retains its existing behavior.
Explain an import decision (ADR-0073)
The task-oriented catalog of every explain surface — which question maps to which command, plus the member-support workflow — is explain.md; the sections here carry the full semantics.
Answer "why didn't this prefix come in?" — or "what did the chain do to it when it did?" — from the per-session import-decision cache:
rbgp policy explain --neighbor 10.0.0.2 --prefix 198.51.100.0/24 --direction import
rbgp policy explain --neighbor 10.0.0.2 --prefix 2001:db8::/32 --direction import --json
# Add-Path peer: omit --path-id to see every path, or pin one:
rbgp policy explain --neighbor 10.0.0.2 --prefix 192.0.2.0/24 --direction import --path-id 3An import explain request without --path-id returns every matching path
while there are at most 4096. Above that, the RPC fails with
RESOURCE_EXHAUSTED and the CLI suggests --path-id; it does not return
an incomplete list. A query with --path-id still resolves that one path.
--direction export runs the
export dry run
under the same verb (identical to rbgp rib --prefix <cidr> advertised <peer> --explain without RD, labeled, or source flags); it needs no configuration.
The address family is inferred from the prefix (IPv4 / IPv6 unicast). Each result reports an outcome:
| Outcome | Meaning |
|---|---|
permit / deny | The chain admitted / rejected the prefix; a deny is explainable even though it never reached the RIB. |
withdrawn | Was permitted, then withdrawn by the peer (tombstone; policy context dropped). |
evicted | Was cached but pushed out by the per-peer cap — raise cache_size. Every evicted key is remembered until session reset, so an evicted prefix never answers not_seen. |
stale | A decision exists but the peer's import policy has changed since; the historical decision is shown with its original generation. |
not_seen | The peer hasn't advertised this prefix on the current session (cache resets on flap / restart). This is an evaluated answer: the session is live and the cache is enabled. |
cache_disabled | The session records no decisions ([policy.explain] enabled = false). The CLI renders this as an error with a config hint and exits nonzero — it is never folded into not_seen. |
no_session | No live session with the requested neighbor, so there is no session-local cache to consult. The CLI renders this as an error and exits nonzero. |
A permit / deny result additionally carries a statement trace —
which statement inside the matched chain decided, per policy evaluated:
permit
decision: no policy rejected; chain default permit
...
statements:
[0] policy edge-import statement 1 permit match: prefix 192.0.2.0/24 set: local_pref 100 -> 200One row per policy the chain consulted (a deny ends the trace at the
denying policy — later policies were never evaluated). default-action
rows mean no statement in that policy matched and its default decided.
For a Permit, the decision attribution is chain_default_permit: the
statement rows retain the member walk, but no member rejected the route.
Matched conditions lead with stable labels (prefix, community,
as_path, neighbor_set, rpki, local_pref, …; any for an
unconditional statement) and attribute edits render as
before -> after against the route's pre-policy values. The trace is
re-derived on demand from the cached compact pre-policy context, so it
attaches only to current-generation permit / deny results — a
stale decision's chain no longer exists and a withdrawn tombstone
has dropped the context the re-derivation needs. --json carries
the same trace as a statements array per match.
This is a side-effect-free read: it does not touch the RIB or move any policy counter. The cache is diagnostic session state, not durable history — it resets on peer flap and daemon restart.
Text output ends with the session's eviction count and cache_size
(cache: N decision(s) evicted since session reset (cache_size C)); JSON and
gRPC carry cache_size and evictions_since_reset, absent when the daemon
predates them. bgp_import_explain_cache_evictions_total{peer} counts the same
evictions: a steadily rising value means that peer announces more distinct
prefixes than cache_size holds, so most of its prefixes answer evicted.
Tuning ([policy.explain] in the config, diagnostic retention only —
never affects which routes are accepted). Both settings are global:
there is no per-peer or per-group override.
enabled(defaultfalse) — import explain is opt-in. On a stock daemon the surface answerscache_disabledand the CLI errors with the config lines to add. Turn it on, reload, and let the session re-establish; the value is read when a session is built.cache_size(default4096) — capacity per session, a fabric / partial-table size. For reliable full-table explain, raise it toward the peer's expected retained-prefix count and budget the memory: the number applies to every session, so budget roughlysum over nonempty peer caches (~1 KiB + min(max(1, cache_size), recorded decisions) × ~600 B), plus eviction memory of about 19 B per evicted key (27–34 B with a nonzero Add-Path identifier), capped at 2,097,152 keys per session (about 38 MB, or up to about 72 MB for nonzero Add-Path identifiers). The index grows with entries; the minimal first-insert probe requested 1,428 heap bytes. Actual memory depends on attributes and allocator.cache_sizeis capped at 2,097,152 and zero acts as one. SeeCONFIGURATION.md.
Answer a member's "why is my route filtered?"
The enumeration complement to policy explain: when a route-server
member calls asking why their prefix isn't in the RS — and neither of
you knows exactly which announcement is at fault — list everything of
theirs the import path rejected, tagged with the reason:
rbgp rib received 10.0.0.2 --rejected
rbgp rib received 10.0.0.2 --rejected --json # Alice-LG-style tooling feedEach row carries the rejected prefix (with Add-Path path_id), the
canonical reason token, a detail field, the wire next-hop, the AS path,
and RPKI/ASPA validation states at rejection time:
| Reason | Meaning / detail field |
|---|---|
policy_reject | The import chain denied it (detail: policy:term when a named .rpol term decided, otherwise the matched policy name; anonymous policies remain empty). The field is bounded to 64 bytes. The row's RPKI/ASPA columns show whether validation drove the deny. Follow up with rbgp policy explain on the prefix for the full statement trace. |
otc_route_leak | RFC 9234 Only-to-Customer ingress drop (detail: the canonical OTC sub-reason, e.g. ingress_from_customer_rsclient). |
next_hop_ownership | Strict-peer next-hop ownership gate, RFC 7948 §4.8 / ADR-0107 (detail: e.g. foreign_next_hop; the next-hop column shows the violating value). |
as_path_loop | Our ASN in the received AS_PATH (RFC 4271 §9.1.2). |
rr_loop | Reflection loop, RFC 4456 §8 (detail: originator_id or cluster_list). |
treat_as_withdraw | RFC 7606: a malformed attribute forced the whole UPDATE's routes to be handled as withdrawn. |
Retention is per-session and self-maintaining: an identity that is
later accepted or explicitly withdrawn drops out (the listing
never claims a live route is filtered), and the store resets on session
flap. It is bounded per peer ([policy.reject_retention] capacity, default
1024, LRU on rejection recency). evictions_since_reset = 0 proves the
retained listing is complete for this session; a positive value makes the CLI
warn that older rejects were lost. Absence means an older daemon cannot answer,
not zero. Max-prefix violations don't appear here: exceeding the limit tears
the session down (Cease/1), which is its own, louder signal.
bgp_rejected_routes_retained{peer} gauges the store per peer — a
sustained high value on a member session is the "their filters are
rejecting a lot" signal worth proactive outreach before the support
call. bgp_rejected_route_retention_evictions_total{peer} counts genuine LRU
displacements; warnings fire only on the first and power-of-two evictions per
session. [policy.reject_retention] enabled = false disables retention
entirely (the CLI then reports the disabled state, never an empty
answer); both knobs are restart-required per peer.
Enable / disable a peer
rbgp neighbor 10.0.0.2 enable
rbgp neighbor 10.0.0.2 disable --reason "maintenance"enable also clears the NOTIFICATION reconnect backoff, so the next
attempt starts immediately and any later failure streak restarts from
connect_retry_secs.
Reset a peer session
rbgp neighbor 10.0.0.2 reset
rbgp neighbor 10.0.0.2 reset --reason "maintenance window"One-shot session bounce (NeighborService.ResetNeighbor): sends Cease /
Administrative Reset with the optional RFC 9003 shutdown communication,
and closes the TCP connection. Unlike disable, the peer stays enabled, so
there is no enable to remember afterwards. A static active-open peer retries
on its normal schedule. If an enabled static peer is already Idle, reset clears
any NOTIFICATION backoff and starts the connection immediately; it emits no
teardown event or session-down sample because no session exists. An accepted
dynamic peer is removed on Idle and must dial in again; its next connection
uses the current dynamic-range configuration. Unknown peers fail with
NOT_FOUND and disabled peers with FAILED_PRECONDITION. Routes are not
retained across an Established bounce: peers that negotiated RFC 8538
Notification Graceful Restart receive a Cease / Hard Reset wrapping the
Administrative Reset. The sent event therefore reports the on-wire subcode 9
in that case rather than the inner Administrative Reset subcode 4. Only an
Established active-primary teardown increments
bgp_session_down_total{reason="local_notification"}.
Trigger an MRT dump
rbgp mrt-dumpLive dashboard
rbgp top # default 2s poll
rbgp top -i 5 # 5s poll intervalShows sessions, prefix counts, message rates, RPKI VRP counts, and
streaming route events in a terminal UI. q or Ctrl-C quits; SIGTERM,
SIGINT, and SIGHUP quit the same way and restore the terminal before the
process exits with status 0, and so does a terminal hangup such as a dropped
SSH session. rbgp top needs an interactive terminal on both stdin and stdout;
otherwise it exits 1 with rbgp top needs an interactive terminal on stdin and stdout before connecting. The route-event subscription is
opened only while the events panel is visible (e).
The panel uses the lag-aware WatchEvents route stream, including exact missed
counts and policy-filter source/target/reason/Add-Path context. Admission and
stream errors are shown as a visible status and retried; there is no fallback
stream.
Global identity is retried
on each normal poll until the first successful fetch, then cached; the
Prometheus metrics scrape runs on a separate 60-second cadence. Before either
optional source succeeds, the status line reports global unavailable or
RPKI unavailable without marking the core connection disconnected. A failed
metrics scrape retains a known VRP count and labels it stale; a successful
scrape with no RPKI family clears the count. Stale health or neighbor data is
likewise labeled while the last-good snapshot remains on screen.
When the peer roster is empty, the dashboard fetches dynamic-neighbor range
inventory immediately and then every 60 seconds while it remains empty. It
distinguishes configured dormant ranges, a proven unconfigured daemon, an
initially unavailable inventory, and a retained stale count; failure of this
optional inventory does not mark the core connection disconnected.
Press h for keybindings.
From peer detail, r opens the on-demand route explorer. It shows one
point-in-time server page (100 routes) of one unicast table at a time:
vcycles the table: Best (the global Loc-RIB, with the selected peer retained only as the export-explain target), Received, Advertised, and Rejected (the last three are scoped to the selected peer).ftoggles IPv4 unicast / IPv6 unicast. Toggling drops a prefix filter from the other family./opens the prefix editor for one exact prefix such as203.0.113.0/24.Tabtoggles the longer-prefixes match inside the editor,Enterapplies,Esccancels without changing the active filter, and an empty submission clears it. A prefix that does not parse, or whose family does not match the selected family, is reported inline and sends no request.Space/PgDnandPgUpmove the highlight one screen within the current server page;nandpfollow the server's opaque page tokens to the next or previous server page.Enterevaluates the highlighted Best prefix withExplainAdvertisedRoute;eopens the same editor, prefilled from any selected row, to explain a prefix for the selected peer, including one absent from Adj-RIB-Out. The result shows advertise or deny, ordered gates and reasons, policy attribution, and route modifications.Enterdoes nothing on Received, Advertised, or Rejected rows.rre-runs page 1 of the current table.
The title and status line expose the table, peer scope (or global for Best),
family, active filter, server page index, exact server total_count, and
snapshot fence (page_version).
Best, Received, and Advertised pass the family and prefix filter to the
daemon's ListBestRoutes, ListReceivedRoutes, and ListAdvertisedRoutes
RPCs. Rejected uses ListRejectedRoutes, which has no server-side paging or
filters: the explorer fetches the peer's retained rejections once, applies the
family and prefix filter client-side, and shows page 1 of 1 with the
matched, retained, capacity, and eviction counts.
The explorer does not poll the RIB, walk pages automatically, or retry general
failures. Changing the table, family, filter, or peer, leaving the view, or
removal of the selected peer on a fresh snapshot cancels the in-flight request
and discards any late result for the previous scope; every such change
restarts at page 1 with an empty token stack. Continuation tokens and their
page_version fence stay authoritative, so pages are never stitched across
generations. If a token becomes stale and the daemon returns ABORTED, the
explorer restarts at page 1 once and names it in the status line (table changed — refreshed at page 1); a second ABORTED in the same scope is shown
as an error. All four listing RPCs and the export-explain request use the
existing bearer credentials and require SensitiveRead authorization.
The explorer covers unicast tables only. It does not offer EVPN, VPN, labeled-unicast, or FlowSpec tables, community or origin-ASN query builders, best-path comparison, source-candidate explanations, or streaming updates. Use the CLI explain catalog for those broader questions.
Watch live events
rbgp events watch
rbgp events watch --backfill 50
rbgp events watch --prefix 203.0.113.0/24 --type added,best_changed
rbgp events watch --category session --type established,lost
rbgp events watch --category session --type notification_sent,notification_received
rbgp events watch --category policy --type policy_changed
# OTC route-leak decisions are published only through the durable
# outbox; the CLI automatically routes this filter through
# SubscribeFromEvent in live-only mode. Add `--from-event-id 0` to
# also replay any retained history.
rbgp events watch --category policy --type otc_route_blocked
rbgp events watch --category dataplane --type dataplane_status_changed
rbgp events watch --category dataplane --type dataplane_route_failed --prefix 203.0.113.0/24
rbgp events watch --prefix 203.0.113.0/24 --type policy_filtered
rbgp events watch --category evpn --type evpn_added,evpn_withdrawn,evpn_best_changed
rbgp events watch --category bfd --type bfd_up,bfd_down,bfd_state_changed
rbgp events sessions --neighbor 10.0.0.2 --type established,lost --limit 20
rbgp events policy --neighbor 10.0.0.2 --type policy_changed --limit 20
rbgp events evpn --route-type 2 --rd 65000:100 --limit 20events watch tails the unified EventService.WatchEvents stream. The
live stream carries route add / withdraw / best-change /
export-policy-filtered events plus structured session lifecycle events
(state_changed, established, lost,
peer_enabled, peer_disabled), metadata-only BGP NOTIFICATION
sent/received events (notification_sent, notification_received), opt-in
policy mutation summaries (policy_changed), EVPN route best-path events
(evpn_added, evpn_withdrawn, evpn_best_changed), and dataplane status-row
summary changes for the FIB / BLACKHOLE discard reconcilers
(dataplane_status_changed) plus live per-route ADR-0061 FIB outcomes
(dataplane_route_installed, dataplane_route_withdrawn,
dataplane_route_failed), and BFD session events (bfd_up, bfd_down,
bfd_state_changed). Prefix and family filters match route events and
per-route FIB dataplane events; use --category session with peer and type
filters when watching session events, --category policy to watch policy /
neighbor-set / peer-group / chain mutations accepted by the runtime, or
--category evpn when watching EVPN route changes. Use --category bfd for
BFD up/down/state-change events. Dataplane summary events are peerless and do
not match --neighbor, --family, or --prefix. FIB rejected counts reflect
surfaced status rows; sampled route_limit_exceeded rows are not a global
suppressed-route total. Policy-filtered route events are target-peer
scoped: peer_address remains the source route peer, target_peer_address is
the outbound peer whose export policy denied the route, and --neighbor matches
either side for route history and live route filtering. Policy events describe
runtime apply success;
config-file persistence is separate. Session state-change events use a bounded
observability channel separate from the lossless TCP collision-coordination
path, so a saturated watch stream can miss lifecycle events without blocking
BGP collision handling. If the client falls behind a bounded route, session,
EVPN, per-route dataplane, or BFD source stream, events watch prints a
stream_lagged warning with the missed count; treat subsequent output as a
live tail after a gap. Policy and peerless dataplane lag is metric-only, with
no in-band warning. Use --backfill N
to print recent matching route history before the live tail starts. Backfill
is route-history only; session, policy, EVPN, dataplane, and BFD events are not
backfilled from process-local history through the live stream command. When
EHM is enabled, those categories — including per-route FIB dataplane events —
can instead be replayed durably with SubscribeFromEvent or
rbgp events watch --from-event-id <N>. Use rbgp rib fib for the current
route ownership snapshot after a reconnect.
For the cursorful CLI form, clean EOF and gRPC UNAVAILABLE reconnect from the
highest complete record flushed to stdout, using 1-second exponential backoff
capped at 30 seconds. Other statuses and output failures stop the command; lag
frames without a top-level BgpEvent.event_id do not advance its process-local
cursor. All filters are preserved on reconnect. Cursorless OTC subscriptions
and ordinary WatchEvents streams remain one-shot.
Category/type values are ORed within each dimension and ANDed across them. The
CLI rejects impossible combinations locally before connecting. On the server,
a missing EHM or an EHM in pass-through mode returns FAILED_PRECONDITION
before category/type parsing. Durable lag is global for every category; OTC is
durable-policy only.
Backfilled route events use the same output shape as live route events, but
the command still prints a history block followed by the live tail rather than
merging the two by wall-clock timestamp.
For recent route history without a live tail, use
rbgp events --prefix <PREFIX>. For recent session lifecycle history,
use rbgp events sessions; it reads the peer manager's bounded
process-local history and resets on daemon restart. The CLI returns 100
history entries by default; rbgp events sessions --all requests the full
bounded in-memory window (the API spells it limit = 0). --limit 0 is a
usage error.
For recent runtime policy / neighbor-set / peer-group / chain mutation history,
use rbgp events policy; it reads a separate bounded 4096-event
process-local history from the peer manager. --neighbor matches only
peer-scoped policy events, so global policy and peer-group changes disappear
from an address-filtered query. rbgp events policy --all requests
the full bounded in-memory window.
For recent EVPN route history, use rbgp events evpn; it reads the RIB's
bounded 4096-event process-local EVPN route-event history. --neighbor matches
both the current and previous best-path peer, --route-type accepts route types
1 through 6, and --rd uses the same Route Distinguisher display format as
rbgp evpn.
Pick the right observability surface
Use the narrowest surface for the question you are asking:
| Question | Command / RPC | Notes |
|---|---|---|
| "What is changing right now?" | rbgp events watch / EventService.WatchEvents | Default live route + session stream. Policy, EVPN, dataplane, and BFD streams are opt-in with --category or matching --type. No replay after reconnect. |
| "What just changed for this prefix?" | rbgp events --prefix 203.0.113.0/24 / ListRouteEvents | Exact-prefix route history from the bounded in-memory RIB ring. |
| "Why did this prefix not reach a peer?" | rbgp events watch --neighbor 10.0.0.2 --type policy_filtered --prefix 203.0.113.0/24 / ListRouteEvents | Export-policy denials where the peer is the denied outbound target. |
| "Did FIB apply fail for this prefix?" | rbgp events watch --category dataplane --type dataplane_route_failed --prefix 203.0.113.0/24 / EventService.WatchEvents | Live ADR-0061 route apply outcome; replayable through SubscribeFromEvent when [event_history].enabled = true. |
| "What policy changed recently?" | rbgp events policy / ListPolicyEvents | Recent policy / neighbor-set / peer-group / chain mutation summaries from the bounded peer-manager ring. |
| "What EVPN route changed recently?" | rbgp events evpn --route-type 2 --rd 65000:100 / ListEvpnEvents | Recent EVPN route add / withdraw / best-change history from the bounded RIB ring. |
| "Are BFD sessions up?" | rbgp bfd, rbgp bfd show 10.0.0.2 / BfdService.GetBfdSessions | Snapshot of configured single-hop and multihop BFD sessions, mode, strict flag, state, diagnostic, and presence-aware remote-AdminDown cause. Down plus remote AdminDown means the peer disabled BFD and RFC 5882 permits BGP; an absent cause from an older daemon is shown as unknown. rbgp doctor records the same bool/null in peers/bfd.json. |
| "Did BFD flap right now?" | rbgp events watch --category bfd --type bfd_up,bfd_down,bfd_state_changed / EventService.WatchEvents | Live BFD session events. No bounded BFD history API. |
| "What routes does the general FIB runtime own or reject?" | rbgp rib fib / ListFibRoutes | Snapshot of ADR-0061 configured-table route ownership. |
| "What BLACKHOLE discards are installed or rejected?" | rbgp rib blackholes / ListBlackholeDiscards | Snapshot of RFC 7999 discard programming. |
| "Are EVPN L2/L3 dataplane pieces ready?" | rbgp evpn runtime, rbgp evpn instances, rbgp evpn nexthops, rbgp evpn vrfs | Snapshot of the committed EVPN runtime generation, resolved EVPN config, and latest dataplane reports. |
| "Do I need alerting over time?" | Prometheus /metrics | Use counters/gauges for alerting; pair with CLI/RPC snapshots for row-level detail. |
Streams answer "what happened while I was connected." Snapshot RPCs answer "what does the daemon currently believe or own." Bounded route, session, policy, and EVPN history rings answer recent after-the-fact timeline questions. Per-route/per-MAC dataplane histories remain roadmap items.
Check health
rbgp healthCheck TCP-AO readiness
rbgp globalThe TCP-AO row reports the local kernel capability probe for RFC 5925
TCP-AO support. supported means the daemon's internal socket primitive can
install keys on this host. unsupported / probe_failed means any configured
static-neighbor or direct dynamic-prefix tcp_ao keyring will fail closed
instead of falling back to unauthenticated sessions: listener failures abort
startup, while active-open failures reject that connect attempt and retry
later. SIGHUP can append non-preferred successor keys to unchanged static and
dynamic owner keyrings without changing Current/RNext. A later SIGHUP can
select the successor and observation-gate predecessor deprecation; a
still-later SIGHUP can delete deprecated MKTs that are neither Current nor
RNext. Editing or reordering existing keys, deleting a selected/non-deprecated
key, and protected-owner CRUD remain restart-required.
rbgp neighbor <address> reads TCP-AO KeyIDs and verification counters from
the live connected socket on every query. Inspection failure is reported as
unavailable without falling back to an older snapshot or disturbing the BGP
session. Linux counters are cumulative for the socket lifetime, so degraded
means either the current/RNext key-validity flag is missing or at least one
verification, missing-key, unsigned-required, or dropped-ICMP error has occurred
since that TCP connection was created.
The same neighbor view reports the RIB's effective live distribution mode,
rather than inferring it from configuration or update-group fallback labels.
Down peers show unknown. Paths-Limit rows are numeric AFI/SAFI ordered, and
their effective send value is explicitly inactive, unlimited, or finite.
View received routes from a peer
rbgp rib received 10.0.0.2
rbgp rib received 10.0.0.2 --prefix 203.0.113.0/24
rbgp rib received 10.0.0.2 --origin-asn 64496 --limit 100
# next page: pass the token the previous page printed, with the same filters
rbgp rib received 10.0.0.2 --origin-asn 64496 --limit 100 --page-token '<token>'
rbgp rib received 10.0.0.2 --age
rbgp rib received 10.0.0.2 --countView best routes (Loc-RIB)
rbgp rib
rbgp rib --age
rbgp rib --count--count (also on rbgp rib advertised PEER --count) applies the same
filters as the full view and renders only the total:
Total matching routes: N in human output, {"total_count": N} with
--json.
Find the route covering an address or prefix
rbgp rib lookup 203.0.113.99
rbgp rib lookup 2001:db8::7
rbgp --json rib lookup 203.0.113.64/26rib lookup asks the daemon for the global Loc-RIB longest-prefix match in one
atomic RPC. Bare IPv4 and IPv6 addresses become /32 and /128; an explicit
CIDR keeps its mask. The response is the normal best-path explanation for the
matched prefix, including the winner, every alternative for that prefix, and
their comparison reasons, so the displayed prefix can be less specific than
the input. No covering route is a not found error. not supported by this daemon means that daemon predates the outside-v1 LookupBestPath method; the
CLI does not hide that boundary by scanning a route listing.
Received and advertised route filters belong after the view and peer:
rbgp rib received PEER --prefix 203.0.113.0/24 --longer
rbgp rib received PEER --origin-asn 64496
rbgp rib received PEER --as-path-contains 64496
rbgp rib received PEER --rpki-state invalid --aspa-state unknown
rbgp rib advertised PEER --community 64496:100
rbgp rib advertised PEER --large-community 64496:1:100The parent spelling (rbgp rib --prefix CIDR received PEER) remains accepted
for compatibility. On a continuously changing full table, use a selective
filter or --limit 1..1000. A limited query is exactly one mutation-fenced
RPC and reports whether the exact matching total was truncated. An unbounded
query continues to require a version-consistent walk of every page and returns
an error if the table changes; the CLI does not emit a torn snapshot.
All filter dimensions are AND-composed. Repeated values within the standard-
or large-community dimension are OR-matched. RPKI states are valid,
invalid, and not_found; the last is the route's recorded uncovered-origin
verdict, not evidence about cache readiness. ASPA states are valid,
invalid, and unknown. --as-path-contains requires one canonical nonzero
decimal ASN and matches exact numeric membership in represented AS_SEQUENCE
or AS_SET segments. It does not evaluate a regex or policy. RFC 9774
rejection of newly received AS_SET and AS_CONFED_SET forms remains
unchanged; the inspection filter does not reopen them.
--age appends the time since the route was originally received into the RIB.
It also works on rbgp rib advertised PEER --age, where it remains the
original RIB receive age rather than the advertisement age. Unknown timestamps
render as -; future timestamps (for example, from CLI/daemon clock skew)
clamp to 00:00:00. A remote CLI compares the daemon-supplied epoch timestamp
with the CLI host's clock. The default human table is unchanged. JSON always
returns the raw received_at_epoch_seconds value and is identical with or
without --age.
View general FIB route status
rbgp rib fib
rbgp -j rib fib
rbgp rib fib --table edge --state rejected --reason route_limit_exceeded
rbgp rib fib --prefix 203.0.113.0/24 --neighbor 198.51.100.2
rbgp rib fib --limit 100This reports only the ADR-0061 configured-table runtime, not the ordinary
Loc-RIB. Rows are installed, rejected, failed, or unresolved. The filters compose
with AND semantics. The --prefix filter is exact prefix+length matching, not
longest-prefix or containment matching. Use --limit and the returned
next-page token (--page-token) to page through large surfaced status snapshots. Pagination is
over rows visible to ListFibRoutes; it does not add suppressed-route counts
for sampled route_limit_exceeded rows.
installed/owned: rustbgpd owns the row and the kernel table matches the current best route.unresolved/next_hop_unresolved: Linux returned the family-specific route-level unreachable errno while applying a target made entirely of unscoped, same-family, non-link-local next hops. One uncovered ECMP member can therefore hold the whole route. The desired row is held without counting a rejection/failure; relevant kernel route events and the 30-second periodic reconcile retry it. Withdrawal, target change, foreign-row appearance, or owned drift clears the stale hold.rejected/foreign_route_exists: a kernel row already exists at the same table / metric / prefix but is not owned by this daemon instance. This includes pre-existingRTPROT_BGProws that are absent from<runtime_state_dir>/fib-owned.json, have a mismatched[[fib_tables]]declaration, or otherwise cannot be tied to persisted owned-state; rustbgpd preserves them rather than taking ownership by protocol alone.rejected/owned_route_drifted: rustbgpd had owned state for the row, but a live reconcile found that the kernel row no longer matched the recorded next-hop orRTPROT_BGPprotocol. rustbgpd releases ownership and leaves the row in place; a later BGP withdraw will not delete the replacement.rejected/next_hop_family_unsupported: the configured table family and BGP next-hop family do not match.rejected/peer_not_allowed: the route's source peer did not match the table'sallowed_neighborsorallowed_peer_groupsguardrail.rejected/route_limit_exceeded: the table's eligible route count exceededmax_routes. The table freezes for that pass: existing owned rows stay installed, and growth or replacement is suppressed until the eligible count falls back under the cap. For very large over-cap tables, rejected rows are sampled so status output stays bounded.failed/dump_failed:*,install_failed:*,replace_failed:*, orremove_failed:*: the runtime hit a RIB or kernel boundary error. Checkbgp_fib_kernel_failures_totaland daemon logs for the matching action.failed/owned_state_persist_failed:*: rustbgpd could not record the route in<runtime_state_dir>/fib-owned.jsonbefore installing it, so it held the install. Checkbgp_fib_owned_state_persist_failures_totaland free space or the mount state of the runtime state directory; the next reconcile retries once the write succeeds.
For direct kernel inspection, use the configured table and metric:
ip route show table 1000
ip -6 route show table 1000On coordinated shutdown, the daemon drains only rows still matching its owned next-hop. If a row drifted underneath the daemon, it is preserved and ownership is dropped.
Quick smoke check — one-shot verification that the runtime is live and
programming the kernel (substitute the configured table_id):
rbgp rib fib # per-route owned / rejected / failed state
ip route show table 1000 # the configured table, straight from the kernel
curl -s localhost:9179/metrics | grep '^bgp_fib_' # install / withdraw / reject / kernel-failure countersManage FIB export tables at runtime (ADR-0061)
rbgp fib-table list
rbgp fib-table set edge --table-id 1000 --metric 200 \
--families ipv4_unicast,ipv6_unicast --max-routes 50000
rbgp fib-table delete edgeset is create-or-replace by name and carries the full table definition (not
a patch); optional ECMP caps are --maximum-paths, --maximum-paths-ebgp,
and --maximum-paths-ibgp. Edits hot-apply through the running ADR-0061
reconciler and persist to the TOML config (atomic write). Changing --table-id
/ --metric for an existing name is a table-key move: the old kernel rows
withdraw and the new table back-fills. The mutating verbs require the reconciler to be running (at least
one [[fib_tables]] entry at startup, on Linux); otherwise they fail
FAILED_PRECONDITION — add the first table to the config and restart. See
CONFIGURATION.md for the full [[fib_tables]] lifecycle.
Explain a best-path decision
# Global Loc-RIB view: best route + every losing candidate annotated with
# the decisive comparison reason and the compared values behind it
# (e.g. "local_pref 100 < 200"). The winner carries the step that beat
# the runner-up ("only_path" for a single-path prefix), and each loser
# is classified against the equal-cost multipath cut
# (eligible / relax_only / none). A prefix with no paths is NOT_FOUND.
rbgp rib --prefix 203.0.113.0/24 --explain
# Peer-scoped view: same shape, but every candidate the named peer would
# actually receive gets a non-zero `advertised_path_id` (rank within the
# peer's effective Add-Path send_max). Filtered candidates (export policy
# reject, family mismatch, split-horizon, iBGP / RFC 4456 RR suppression,
# beyond send_max) stay at 0 so the operator can see *why* each isn't
# advertised.
rbgp rib --prefix 203.0.113.0/24 --explain --explain-peer 10.0.0.2Explain an export decision ("why did/didn't route X go to peer Y?")
# The full export gate ladder for one prefix toward one peer, in the
# exact order the live export path evaluates it. Each rung reports
# pass / STOP / n/a with detail; a STOP names the gate that held the
# route back.
rbgp rib --prefix 203.0.113.0/24 advertised 10.0.0.2 --explain
# Negotiated unicast Add-Path: select one exact Adj-RIB-In candidate.
# The inbound ID may be zero; the response reports its independent outbound rank.
rbgp rib --prefix 203.0.113.0/24 advertised 10.0.0.2 --explain \
--source-peer 198.51.100.7 --source-path-id 0
# VPNv4/VPNv6 (SAFI 128): explain the (RD, prefix) identity instead --
# the ladder additionally includes the RFC 4684 RT-Constrain
# membership gate.
rbgp rib --prefix 10.1.0.0/24 advertised 10.0.0.2 --explain --rd 65000:1
# Labeled unicast (SAFI 4, RFC 8277): explain the labeled export ladder
# for the prefix instead of the plain unicast one.
rbgp rib --prefix 10.1.0.0/24 advertised 10.0.0.2 --explain --labeled
# JSON for scripting
rbgp --json rib --prefix 203.0.113.0/24 advertised 10.0.0.2 --explainThe gate ladder, in live evaluation order (unicast single-best):
best_route (Loc-RIB best exists) -> split_horizon (not sent back to
its source) -> rr_reflection (iBGP split horizon / RFC 4456
client/cluster rules) -> family (peer negotiated the AFI/SAFI) ->
llgr (RFC 9494 stale-export restriction) -> orf (peer-pushed
RFC 5291 filter) -> export_policy (per-chain verdict, labeled
policy:term for .rpol members) -> adj_rib_out (diff against the
advertised state: staged_announce = would send, already_advertised
= identical route already advertised, Adj-RIB-Out in sync). Every
adj_rib_out verdict describes local send-side state only: BGP has no
acceptance signal, so an already_advertised pass never means the peer
holds the route — a peer that treats the updates as withdrawn (RFC 7606)
stays Established with none of them. The VPN ladder
follows the live VPN staging order and adds rt_membership
(RFC 4684), as does the EVPN export ladder of rbgp evpn explain; the
labeled-unicast ladder follows the live labeled staging
order (family first, no rt_membership/orf). A family still held by
the initial-ORF gate (RFC 5291 section 6) stops at orf_gate before any
per-prefix work.
Truthfulness: the explanation is produced by a read-only dry run of
the same staging body live distribution executes -- including for
update-grouped peers, which are explained against their group table
with split horizon applied per member exactly like the emit-time
source-flip matrix. The response carries update_group_id when the
peer's export is group-staged. Explain queries never count toward
bgp_policy_routes_total or per-term policy hit counters.
The live commit path then applies one final exact-wire gate through an
immutable snapshot of the session encoder and its negotiated 4096/65535-byte
ceiling. A failed announcement never enters Adj-RIB-Out. If it had been
advertised before an attribute, next-hop, capability, or ceiling change made
it unexportable, that transition sends a withdrawal. For update groups the
shared staged table is unchanged; a sparse per-member overlay removes the
rejected identity from that peer's advertised query and BMP count, allowing a
classic-message peer and an Extended Message peer to remain grouped safely.
Later recompute or dirty resync retries the route, while a source withdrawal
clears the rejection without sending a duplicate withdrawal for a route that
was never advertised. Watch
bgp_exact_export_rejections_total{peer,family,reason} and the corresponding
warning log. Cease/8 still tears down the session if transport encounters an
impossible single-route message or a missing/mismatched encoder snapshot; that
is defense in depth, not the normal rejection path.
Neighbor last_error and session notification history retain the bounded local
cause of that defense-in-depth teardown, outbound queue saturation, and TCP
reader/writer failures. These diagnostics deliberately exclude raw operating-
system errors, task panic payloads, routes, attributes, policy data, and
shutdown communication text. A locally sent Cease/8 retains the category while
its canonical BGP description and wire payload remain unchanged.
Every decoded inbound or attempted outbound BGP NOTIFICATION also emits one
structured INFO record named BGP NOTIFICATION. The record includes peer,
direction, the outer numeric code and subcode, and their human
description. A valid Hard Reset keeps outer Cease/9 while its description
identifies the decoded inner error. Valid RFC 9003 shutdown communication is an
optional reason; it is escaped to printable ASCII and capped at 512 bytes.
Malformed or absent communication is omitted, with malformed input retaining
its separate warning.
An RFC 9107 ORR peer's explain ranks the per-vantage candidate set the
ORR export uses (with per-candidate cost output). When filtered or ignored
topology inputs are present, the stable
orr_topology_input_diagnostics reason is explicitly aggregate and
non-decisive; it does not replace the winner's decisive interior-cost reason.
rbgp orr always prints a Topology inputs: aggregate line; JSON and
ListOrrStatus expose the same five counters. For negotiated IPv4/IPv6
unicast Add-Path send, --source-peer plus --source-path-id makes export
explain follow one exact Adj-RIB-In path through eligibility, policy, compact
outbound ranking, OTC/exact-wire checks, and the rank-specific Adj-RIB-Out
diff. The source flags are paired, require --explain, and conflict with
--rd/--labeled; an unknown source fails rather than falling back to the
winner. The returned outbound Path ID is independently assigned per RFC 7911:
rank 0 means the candidate was filtered, denied, or beyond send_max, while a
post-rank OTC/exact-wire denial retains its attempted rank. Use
rbgp rib --prefix X --explain --explain-peer when the question is instead
which candidates make the peer's whole ranked send view.
Manage policies, peer groups, and neighbor sets
# Read
rbgp policy list
rbgp policy get import-from-transit
rbgp neighbor-set list
rbgp peer-group list
# Write — JSON file matches the proto message shape
rbgp policy set import-from-transit --from-file policy.json
rbgp neighbor-set set transit-peers --from-file ns.json
rbgp peer-group set transit --from-file pg.json
# Apply chains globally or per-neighbor
rbgp policy chain set-import --global import-from-transit
rbgp policy chain set-import import-from-transit --neighbor 10.0.0.2
rbgp policy chain show --neighbor 10.0.0.2
# Bind / unbind neighbors to a peer-group
rbgp peer-group attach 10.0.0.5 --group transit
rbgp peer-group detach 10.0.0.5--from-file accepts JSON whose shape mirrors the proto message
(PolicyDefinition / NeighborSetDefinition / PeerGroupDefinition);
unknown fields are rejected at parse time. Empty
chain set-{import,export} is rejected — use the matching clear-*
subcommand to drop a chain.
A global chain change applies to every neighbor without its own chain and is
persisted. Select the scope with --global or --neighbor; omitting both is
a usage error (exit 2) and changes nothing. When stdin and stdout are both
terminals, a global change names the endpoint and asks [y/N] first;
-y/--yes skips the prompt, and non-interactive runs never prompt. A
declined prompt exits 1 without changing anything.
Graceful shutdown (daemon exit)
rbgp shutdownSends NOTIFICATION to all peers, writes GR marker, exits cleanly. On a
terminal, rbgp shutdown names the endpoint and asks [y/N] first; pass
-y/--yes to skip the prompt. Non-interactive runs never prompt.
RFC 8326 graceful-shutdown community (planned maintenance)
Distinct from the daemon-shutdown RPC above. RFC 8326 lets you drain
traffic ahead of a planned EBGP session shutdown by tagging outbound
paths with the well-known GRACEFUL_SHUTDOWN community
(65535:0 / 0xFFFF_0000); receivers that honor the community
demote LOCAL_PREF to 0 so any non-shutting alternate becomes
preferred. By the time you actually close the session, traffic has
already moved.
Initiator (the side going down for maintenance):
# Start the drain on one peer
rbgp gshut --neighbor 10.0.0.2
# Or drain every currently-managed peer at once
rbgp gshut --all
# Wait for traffic to shift (operator-defined, typically 30s-5min
# depending on convergence in the upstream AS), then proceed with
# the actual maintenance — restart, config edit, etc.
# Clear the community when maintenance ends
rbgp gshut --neighbor 10.0.0.2 --clear
rbgp gshut --all --clearAn all-peers change asks [y/N] first when stdin and stdout are both
terminals; -y/--yes skips it. rbgp gshut without --all or
--neighbor is a usage error (exit 2) and changes nothing.
The toggle is operator-runtime state, not config — it lives on
the ManagedPeer desired-state record, mirrors to the live session,
and survives session flaps mid-maintenance. The toggle does NOT
persist across daemon restart by design (RFC 8326 is a maintenance-
window action, not a steady state).
When the toggle flips, rustbgpd issues a RibUpdate::RefreshPeerOutbound
which queues re-emission of all routes already in AdjRibOut to the
target peer without waiting for an unrelated RIB event. Receiver-side
verification is still required to prove delivery.
rbgp neighbor <peer> always reports the local desired state as
GShut Advertise Intent: true|false|unknown; unknown means the CLI is
connected to an older daemon that does not expose this field. This is only the
local send intent — it does not prove that a route was re-advertised, received,
or converged downstream.
Receiver (the side honoring others' GShut):
Set in [global]:
[global]
honor_graceful_shutdown = trueWhen enabled, an implicit chain-tail rule fires on every EBGP peer's
import chain — see docs/reference/configuration.md for the exact semantics.
iBGP peers are exempt because LOCAL_PREF is preserved within an AS.
Verifying the drain is working:
The community is attached on the wire by the per-peer transport layer
after the RIB-side advertised view is computed, so
rbgp rib advertised does NOT show the GShut community on the
initiator side — the RIB doesn't know about the toggle. The
authoritative checks are:
# Receiver-side: routes from a draining peer that honor the community
# show explicit local_pref_attr = 0 in the RIB (proves the implicit
# chain-tail rule fired). EBGP-received routes have no LOCAL_PREF on
# the wire, so look at local_pref_attr (explicit) rather than
# local_pref (proto3 default).
rbgp --json rib received <draining-peer> | jq '.[] | {prefix, local_pref_attr, communities}'
# Initiator-side: confirm the local desired advertisement state.
rbgp --json neighbor <receiving-peer> | jq -e '.graceful_shutdown_advertise_intent == true'
# Or verify on the *receiving* peer's BGP table — the canonical
# observation. On FRR:
vtysh -c 'show ip bgp <prefix> json' \
| jq '.paths[].community'
# (In a maintenance scenario you usually have control of both ends, so
# the receiver-side check is what matters for correctness.)Interop is validated in M35 (tests/interop/m35-graceful-shutdown-frr.clab.yml)
against FRR 10.7.1 — both legs (FRR → rustbgpd inbound honor +
rustbgpd → FRR outbound advertise + clear) end-to-end.
Explain best-path selection
rbgp rib --prefix 10.0.0.0/24 --explainShows all candidates for a prefix with the decisive comparison reason
for each non-winner (e.g., higher_local_pref, shorter_as_path)
plus the compared values behind it (e.g., local_pref 100 < 200),
the step that selected the winner over the runner-up, and whether each
loser would survive the equal-cost multipath cut.
Looking glass (Birdwatcher-shaped REST subset)
For status, peer, accepted-route, filtered-route, and noexport views in
external looking glass frontends, run the external
examples/birdwatcher-adapter binary. It serves the Birdwatcher-shaped
endpoints (/status, /protocols/bgp with real per-neighbor filtered
counts, /routes/protocol/{id}, /routes/peer/{peer},
/routes/filtered/{id}, /routes/noexport/{id}) from the daemon's gRPC
API. The filtered view surfaces the PolicyService.ListRejectedRoutes
reject-retention store with a synthesized reject-reason large community
per route; the noexport view diffs Loc-RIB best against the peer's
Adj-RIB-Out and names each suppression's export gate via
RibService.ExplainAdvertisedRoute, with its own synthesized
noexport-reason community (mapping tables and Alice-LG config in the
adapter's README). The in-daemon [global.telemetry.looking_glass]
server has been removed.
EVPN Route Reflector + Bidirectional VTEP
rustbgpd has two operational EVPN modes that share the same l2vpn_evpn
session machinery:
- RR mode (Phase 1): empty
[[evpn_instances]]. The daemon reflects EVPN Types 1–6 between iBGP-speaking VTEPs, including RFC 9251 Type 6 SMET relay. It owns no kernel state and runs no DF election. External VTEPs (FRR on SONiC, commercial NOS) handle local origination + forwarding. - Bidirectional VTEP mode (Phase 2 — Gates 7a / 7b / 7b+1):
populated
[[evpn_instances]]. The daemon programs the kernel bridge FDB from received Type 2 routes (downward, ADR-0054) AND originates Type 2 from kernel-learned local MACs plus one Type 3 IMET per configured L2VNI (upward, ADR-0055). Linux-only. Gate 7b+1 ships in v0.15.0.
Phase-2 status: Gates 7a/7b/7b+1/7b+2/7c have shipped the bidirectional L2VNI VTEP loop: declarative instances, downward FDB reconciliation, local MAC and MAC+IP origination, Type 3 IMET, SVI MAC origination, sticky MAC config, and sub-second mobility wakeups. Gate 8/8b adds alpha multi-homing execution: DF election, Type 1/4 origination, production-default BUM suppression with opt-out config, ESI-aware Type 2 origination, aliasing projection, and receive-side mass-withdraw filtering. Gate 9 ships symmetric Interface-less IRB end-to-end in v0.18.0 (RFC 9136 §4.4.2 / ADR-0058):
[[evpn_ip_vrfs]]config schema +[[evpn_instances]].ip_vrfbinding,IpVrfStatusreadiness probe, Linux VRF / L3VXLAN netlink dumps, per-IP-VRF kernel-route observation with conservative classifier, Type 5 origination viaRibUpdate::InjectEvpngated on readiness, remote import + L3 FIB programming through a transactionalL3OwnedStatemodel,RTNLGRP_IPV4/IPV6_ROUTEmulticast for sub-second withdraw,ListIpVrfs/GetIpVrfgRPC +rbgp evpn vrfsCLI, M39 hosted kernel-dataplane CI. ADR-0059 (v0.19.0) adds receive-path aliasing-ECMP via FDB nexthop groups (slices 1-4, M40 FRR-validated); slice 3.5 hardening (PRs #91 / #92 / #93) added theapply_aliasing_ecmpper-instance off-switch, periodicRTM_GETNEXTHOPdrift recovery, and homogeneous IPv6 alias members. Production-default multi-homing enforcement, auto-derived RTs, partial ADR-0063 live EVPN runtime mutation, receive-side RFC 9135 overlay-index recursion, native GW-IP + ESI overlay-index Type 5 origination, single-active and all-active ESI overlay-index Type 5 receive, and controller Gateway Address Type 5 injection have since shipped. Still ahead: remaining ADR-0063 shapes, Linux softswitch local-bias split-horizon, true shared-VNI / non-zero Ethernet Tag service, managed netdev ergonomics, service-provider EVPN breadth, and deeper cross-vendor/scale validation. Seeevpn-enablement.mdfor the gate ladder,evpn-alpha-soak.mdfor the residual alpha-confidence checklist, andevpn-vtep-troubleshooting.mdfor the operator runbook.
Per-neighbor knob
[[neighbors]]
address = "10.0.1.1"
remote_asn = 65000
families = ["l2vpn_evpn"]
route_reflector_client = trueSet route_reflector_client = true on every VTEP peer; the daemon's
own cluster_id (under [global]) drives the RFC 4456 ORIGINATOR_ID
- CLUSTER_LIST stamping.
Inspect the EVPN RIB
rbgp evpn # all EVPN routes
rbgp evpn --route-type 2 # MAC/IP only
rbgp evpn --route-type 6 # SMET relay routes
rbgp evpn --rd 65000:100 # filter by RD
rbgp evpn --neighbor 10.0.1.1 # filter by source peer
rbgp evpn diagnose # alpha VTEP summary
rbgp evpn runtime # committed EVPN generation / mutation state
rbgp evpn clear-duplicate-mac --vni 100 --mac aa:bb:cc:dd:ee:fftunnel_type=8 in the output indicates the RFC 8365 VXLAN
encapsulation extended community is present.
For an exact Type 6 route, use its literal source/group wildcard or IP and originator address. Quote the wildcard for the shell:
rbgp evpn explain smet --rd 65000:100 --source '*' --group 239.1.2.3 \
--originator-ip 2001:db8::1 --advertised-to 10.0.1.2Type 6 supports relay and inspection only. It does not originate SMET routes, run an IGMP/MLD proxy, or program multicast forwarding. The M113 controlled raw-peer proof checks reflection, withdrawal, and error recovery with an independent TShark decoder. Vendor interoperability remains unproven; see the SMET boundary.
Inspect the dataplane (ADR-0059 FDB nexthop groups)
rbgp evpn nexthops # owned FDB-NHG groups / members / MAC refs
rbgp evpn nexthops --json # JSON for scriptingThis is the rustbgpd-owned view of ADR-0059 aliasing-ECMP state —
distinct from the RIB above. Compare its group-id, member
nh_ids, and mac-refs against ip nexthop show / bridge fdb show when debugging multi-homed Type 2 forwarding. The top-line
header reports orphan-nexthops, pending-deletes,
drift-recovery-disabled, l3-orphan-nexthops, and l3-pending-deletes
so the periodic drift-recovery latch and the L2 and L3 (all-active
Type 5) allocator GC backlog are visible without log scraping.
Inject a route from a controller
rbgp evpn add-mac-ip --rd 65000:100 \
--mac 02:00:00:aa:bb:cc --ip 10.0.0.5 \
--label 100 --next-hop 10.0.0.2 \
--rt 65000:100
rbgp evpn delete-mac-ip --rd 65000:100 \
--mac 02:00:00:aa:bb:cc --ip 10.0.0.5
rbgp evpn add-ip-prefix --rd 65000:5000 \
--prefix 10.50.0.0/24 --label 5000 \
--next-hop 192.0.2.10 --router-mac 02:00:00:00:50:00 \
--rt 65000:5000
rbgp evpn delete-ip-prefix --rd 65000:5000 \
--prefix 10.50.0.0/24Two complementary origination paths exist:
- gRPC injection (Phase 1, Gate 6):
InjectionService.AddEvpnRoute/DeleteEvpnRoute(therbgp evpn add-mac-ip / add-imet / add-ip-prefix / delete-*commands above). The controller decides what to originate; rustbgpd reflects + distributes. Type 2 (MAC/IP), Type 3 (IMET), and Type 5 (IP Prefix) are exposed. Type 5 injection uses ESI=0 and Ethernet Tag ID=0. Omitting--gatewaykeeps the Interface-less Gateway IP=0 form; supplying--gatewayinjects a non-zero overlay-index Gateway Address.--router-macis required for the default VXLAN encapsulation path and should be omitted when--no-vxlan-encapis set. Non-zero ESI overlay-index injection and Type 1/4 multi-homing route injection are not exposed. Native Type 1/4 origination is driven by[[ethernet_segments]]. - Kernel-driven origination (Phase 2, Gate 7b+1): with
[[evpn_instances]]populated, the daemon subscribes toRTNLGRP_NEIGH(enum group id 3) and emits Type 2 routes for MACs the kernel learns on non-VXLAN bridge ports, plus one Type 3 IMET per L2VNI at startup. RFC 7432 §15.1 mobility sequencing is automatic. Withdraws fire on FDB age-out /bridge fdb deland on coordinated shutdown.
Common operational signals
- EVPN routes counted toward
max_prefixes. A peer flooding EVPN Type 2 routes will trip the same Cease/MAX_PREFIXES that a peer flooding unicast prefixes would. The cap is the union of unicast unique prefixes + FlowSpec rules + EVPN keys. - GR / LLGR stale handling is implemented for EVPN. When negotiated for the family, restart handling marks retained EVPN routes stale and ranks them below fresh alternatives (RFC 4724 §4.2 / RFC 9494 §4.7). Unit and integration tests cover the transitions; the live FRR VTEP restart evidence remains pending.
- Late-joining peer. A VTEP that connects to a converged RR receives the existing EVPN routes in its initial dump before the EoR marker. (This was not always the case — see commit history for the regression test.)
- MAC mobility selection. Local MAC moves increment the advertised
MAC Mobility sequence, saturating at
u32::MAX. Same-segment peer synchronization can adopt a higher sequence without incrementing it. EVPN selection also considers freshness, sticky status, and the remaining tie-breaks; a higher sequence alone does not guarantee a path change.
For the full enablement story, gate ladder, and known limitations, see docs/project/evpn-enablement.md. For a step-by-step operator checklist, see docs/how-to/evpn-vtep-troubleshooting.md.
Troubleshooting kernel-driven origination (Gate 7b+1)
- Local MAC learned in kernel, but Type 2 not on the wire. Check
in order: (a)
[[evpn_instances]]is populated and the bridge named there exists with a single VXLAN port (probe reportsReadyonly when ADR-0054 §4's five-point check passes); (b) the MAC was learned on a non-VXLAN bridge port — the classifier intentionally drops VXLAN-port ifindexes (those are remote-MAC echoes); (c)RUST_LOG=rustbgpd_evpn_linux=debugshows the classifier hit (cache miss →bridge_port_to_vnidoesn't yet contain the slave ifindex; the supervisor's periodic dump should populate it within 5 s); (d) the BGP session reached Established before the originator emitted the Inject — pre-Established Injects do reach the AdjRibOut and ride the initial dump after the session reaches Established. - Type 3 IMET not visible on a peer. IMET is emitted at startup
for every configured
EvpnInstanceregardless of dataplane Ready/NotReady. If FRR'sshow bgp l2vpn evpn route type multicastdoesn't show it, check that the peer reached Established and that the L2VPN/EVPN family was negotiated (families = ["l2vpn_evpn"]). - Type 2 / Type 3 not withdrawn cleanly on shutdown. The
shutdown order is: (1) drain originator's outstanding Withdraws;
(2) withdraw IMET keys; (3)
PeerManagerCommand::Shutdown. If peers see stale routes after a clean exit, check the structured log for thedraining EVPN originator/withdrawing EVPN Type 3 IMET routeslines firing before any peer-session-shutdown log lines. rtnetlink multicast subscription failed; corresponding upward feed will be silentwithgroup_name="RTNLGRP_NEIGH"in the startup log: local-MAC observations will be silent. The same message withRTNLGRP_IPV4_ROUTE,RTNLGRP_IPV6_ROUTE, orRTNLGRP_LINKsilences that route or link-event feed instead. The daemon lacksCAP_NET_ADMIN. Downward FDB programming also needs the capability; if the dataplane reconciler is working but the originator is silent, the cap is partially granted (rare). Checkgetcapon the binary.