rustbgpd
Cookbook

Controller / monitoring feed (BMP trio, events, MRT)

Export routing state to your monitoring stack.

Export routing state to your monitoring stack.

When this is you: Your controller, analytics pipeline, or NOC needs routing state. It needs to see what this daemon sees — pre-policy, post-policy, and Loc-RIB route streams into BMP collectors, a durable event feed your bridge can replay after a restart, periodic MRT table archives, and a Grafana dashboard on top. rustbgpd exports the full BMP monitoring trio on one exporter (RFC 7854 Adj-RIB-In, RFC 8671 Adj-RIB-Out as byte-exact wire PDUs, RFC 9069 Loc-RIB), selectable per collector.

Proven by: M24 (BMP Initiation / PeerUp / RouteMonitoring ordering vs a BMP receiver) and M81 (the trio plus BMPv4, validated against three independent decoders at once: pmacct, gobmp, and tshark). Config shape derived from tests/interop/configs/rustbgpd-m81-bmp-rr.toml and examples/route-collector/.

Config

The example is an RR that also feeds the monitoring stack; the same [bmp] / [event_history] / [mrt] blocks bolt onto any of the other recipes unchanged.

[global]
asn = 65000
router_id = "10.255.0.1"
listen_port = 179
cluster_id = "10.255.0.1"
# RFC 8212 posture, stated explicitly. It governs eBGP sessions only, so
# these iBGP clients are unaffected; an eBGP neighbor added later needs
# explicit import and export policy.
ebgp_requires_policy = true

[global.telemetry]
prometheus_addr = "127.0.0.1:9179"
log_format = "json"

# Owner-only local socket (default mode 0600): clients are authorized as the
# implicit "local-operator" principal — no [security.grpc] block needed.
[global.telemetry.grpc_uds]
path = "/var/lib/rustbgpd/grpc.sock"

# Durable event outbox (ADR-0072): restart-safe replay cursor for
# SubscribeFromEvent. Opt-in — it costs memory/CPU at scale, which is
# why it is off by default. Enable it here because replay is the
# point of this deployment.
[event_history]
enabled = true

# Periodic MRT TABLE_DUMP_V2 snapshots — readable by bgpdump, BGPKIT,
# and the RouteViews/RIPE RIS toolchains.
[mrt]
output_dir = "/var/lib/rustbgpd/mrt"
dump_interval = 3600
compress = true
file_prefix = "feed"

[bmp]
sys_name = "rustbgpd-feed"
sys_descr = "fabric RR monitoring feed"

# Production collector: BMP v3 (the default), all three RIB views.
# On (re)connect the collector receives Initiation and the cached Peer Up
# state; the loc_rib view also gets a chunked Loc-RIB dump closed by
# End-of-RIB. rib_in_pre and rib_out_post have no reconnect dump and start
# at the next UPDATE. No daemon restart is needed to attach a collector.
[[bmp.collectors]]
address = "10.20.0.10:1790"
reconnect_interval = 5
monitor = ["rib_in_pre", "rib_out_post", "loc_rib"]

# Optional second collector with BMPv4 TLV framing. Code points follow
# draft-ietf-grow-bmp-tlv-21 and remain pre-IANA. Path Marking is temporarily
# unavailable because its draft type 5 collides with Sequence Number. Keep
# production collectors on v3 (the default); point v4 only at tooling that
# tracks the draft (M81 validates the v4 bytes with a raw oracle and tshark).
#[[bmp.collectors]]
#address = "10.20.0.11:1790"
#reconnect_interval = 5
#version = 4
#monitor = ["rib_in_pre", "rib_out_post", "loc_rib"]

[[neighbors]]
address = "10.0.0.11"
remote_asn = 65000
description = "pe-1"
hold_time = 90
route_reflector_client = true
families = ["ipv4_unicast", "ipv6_unicast", "l3vpn_ipv4_unicast"]

[[neighbors]]
address = "10.0.1.2"
remote_asn = 65000
description = "pe-2"
hold_time = 90
route_reflector_client = true
families = ["ipv4_unicast", "ipv6_unicast", "l3vpn_ipv4_unicast"]

Passive iBGP relay

Use a passive relay when two routers feed a collector through rustbgpd and the relay must advertise nothing. The route-collector starter provides the same accept-all import and deny-all export posture for eBGP sources. This minimal IPv4 configuration uses two sources in the relay's own AS and sends only the pre-policy Adj-RIB-In BMP view:

[global]
asn = 65001
router_id = "10.255.0.1"
listen_port = 179
ebgp_requires_policy = true

[global.telemetry]
prometheus_addr = "127.0.0.1:9179"
log_format = "json"

[global.telemetry.grpc_uds]
path = "/var/lib/rustbgpd/grpc.sock"

[bmp]
sys_name = "rustbgpd-passive-relay"

[[bmp.collectors]]
address = "10.20.0.10:1790"
monitor = ["rib_in_pre"]

[policy]
import_chain = ["observe-all"]
export_chain = ["deny-all"]

[policy.definitions.observe-all]
default_action = "permit"

[policy.definitions.deny-all]
default_action = "deny"

[[neighbors]]
address = "10.0.0.2"
remote_asn = 65001
families = ["ipv4_unicast"]

[[neighbors]]
address = "10.0.1.2"
remote_asn = 65001
families = ["ipv4_unicast"]

Keep the deny-all export chain attached: an empty iBGP export chain permits routes. This configuration enables no forwarding backend; received routes remain available to the daemon's RIB queries and BMP feed.

At the collector, clear that peer's route state on PeerDown, as specified by RFC 7854 §4.9. Count that state teardown separately from explicit RouteMonitoring withdrawals. Connect the collector before feeding routes. After a BMP reconnect, cached PeerUp messages do not rebuild the received table: Adj-RIB-In has no automatic reconnect dump. See the BMP reconnect caveat.

The dated withdrawal parity receipt and its runnable lab prove one bounded shape: eight IPv4 /32s per peer, three announce/withdraw rounds and two flaps per peer, with 24 matching wire/BMP withdrawals per peer and zero exported announcements. This is not full-table memory, sustained-load, backpressure or reconnect-completeness evidence.

Verify

$ export RUSTBGPD_ADDR=unix:///var/lib/rustbgpd/grpc.sock
$ rbgp health
$ rbgp neighbor

BMP: on the collector you should see Initiation, the Peer Ups, the chunked Loc-RIB dump and its End-of-RIB (loc_rib only), then live RouteMonitoring deltas for every selected view. From the daemon side the collector connection state is visible in the logs (log_format = "json") and the BMP metrics below. Periodic peer Stats Reports also carry RFC 9972 post-policy accepted-route gauges, exact type-22 IPv4/IPv6-unicast policy reject counts, and exact per-path RPKI Invalid/Valid/NotFound types 35/36/37. The RPKI rows cover negotiated IPv4/IPv6-unicast families, count every Add-Path identity, and include authoritative zero rows; they are omitted when no VRP table is installed. Type 22 is present only while the default-on rejected-route store is authoritative: disabling retention or reaching its capacity omits the rows after the first eviction instead of publishing a partial count.

During a Loc-RIB bootstrap or dump, watch bmp_loc_rib_dump_live_buffer_depth{collector} for the current queued live rows and bmp_loc_rib_dump_live_buffer_high_watermark{collector} for the connection generation's peak. The collector value is the configured socket address including port. A high-water mark near the fixed 8192-row bound means the collector or dump path needs investigation before the next burst.

Event replay (the bridge contract): live tail plus durable replay from a cursor. --from-event-id 0 replays everything retained, then tails; your bridge persists the last event_id it processed and resumes from there after either side restarts:

$ rbgp events watch --from-event-id 0
$ rbgp events watch --category route,session --from-event-id 41236

While this CLI process remains running, a clean stream end or gRPC UNAVAILABLE reconnects from the highest successfully flushed top-level BgpEvent.event_id. Reconnect delay starts at 1 second, doubles to a 30-second cap, and resets after a complete human or JSON record plus newline has been written and stdout has been flushed. The CLI preserves the full filter set on every request. Lag frames without a top-level event ID do not move the cursor; all other RPC statuses and output failures are terminal. Cursorless OTC subscriptions and ordinary WatchEvents streams remain one-shot. This process-local retry does not replace the bridge's downstream-confirmed, persisted cursor across CLI restarts.

The same contract over raw gRPC is EventService.SubscribeFromEvent — see examples/event-bridge/ for a complete bridge (gRPC → JSON lines) you can pipe into Kafka, NATS, or Vector. rustbgpd deliberately is not an event bus; the outbox is a bounded SQLite WAL store with a monotonic cursor (ADR-0072).

MRT: force a dump and inspect it with your usual tooling:

$ rbgp mrt-dump
$ ls /var/lib/rustbgpd/mrt/
feed.20260703.120001.123456789.mrt.gz

Looking glass (optional): for status, peer, accepted-route, filtered-route, and noexport views in an Alice-LG-style frontend, run the examples/birdwatcher-adapter/ against the local Unix socket, or use its dedicated authenticated observer listener pattern for least privilege. (The in-daemon [global.telemetry.looking_glass] server has been removed.) The filtered view is served from PolicyService.ListRejectedRoutes with structured reject reasons; the noexport view diffs the Loc-RIB best set against the peer's Adj-RIB-Out and names each suppression's export gate via RibService.ExplainAdvertisedRoute — see the adapter README for exactly what each view contains.

Watch

Import the overview dashboard per GRAFANA.md; the BMP / event-outbox row is populated once the features above are configured. Key series:

MetricMeaning
bgp_event_outbox_degraded1 = latched durability-impacting loss, committed-event delivery skip, or DB open/recovery/quarantine failure; expected shutdown reason=closed drops are excluded. Inspect the drop reason and daemon log: replay can remain available
bgp_event_outbox_storage_failed1 = the event-history storage stopped at runtime; the outbox refuses events and durable subscriptions until a restart
bmp_collector_drops_totalper-collector queue/dump failures; live fan-out Full/Closed automatically resets only that collector generation, replays cached Peer Up state, and rebuilds configured Loc-RIB state after the one-second reconnect
bmp_source_drops_totalper-peer BMP events dropped at the source tap, including a periodic stats report whose session-state query timed out
bmp_replay_attempts_totalPeerUp-cache replays on collector reconnect
bgp_rib_outbound_registered_peersfeed coverage: peers whose routes the views carry

Failure modes

SubscribeFromEvent / rbgp events watch --from-event-id returns FAILED_PRECONDITION. The daemon is running with [event_history].enabled = false (the default). Enable it and restart — the outbox fields are restart-required.

The bridge missed events during a burst. Live streams emit stream_lagged warnings when a bounded source dropped events for a slow consumer. That is the signal to resume via the durable cursor (--from-event-id <last-processed>) rather than the live ring.

A BMP collector shows nothing after a network blip. The daemon redials with backoff capped by reconnect_interval and replays cached Peer Up state. It performs a fresh EoR-closed table dump only for collectors with monitor = ["loc_rib"]. Adj-RIB-In and Adj-RIB-Out have no automatic reconnect dump. For a late outbound collector, follow the outbound capture procedure. Its experimental replay-out operation reannounces the selected peer's routes on the live BGP session; successful scheduling alone does not prove a complete capture. See Known issues. Check the collector-side listener first, then the daemon log for connection and bootstrap failures.

pmacct rejects the v4 stream (BMPv4 BGP PDU TLV != 1). Known and expected — see the caveat in the config above. Move that collector to version = 3 (or drop the version key; 3 is the default).

The events DB was corrupted by a crash. The bad file is renamed events.db.stale, the daemon continues (pass-through) unless [event_history].required = true, and bgp_event_outbox_degraded latches to 1 until an operator restart. Details: CONFIGURATION.md §event_history.

Source on GitHub

On this page