Controller / monitoring feed (BMP trio, events, MRT)
Export routing state to your monitoring stack.
Export routing state to your monitoring stack.
When this is you: Your controller, analytics pipeline, or NOC needs routing state. It needs to see what this daemon sees — pre-policy, post-policy, and Loc-RIB route streams into BMP collectors, a durable event feed your bridge can replay after a restart, periodic MRT table archives, and a Grafana dashboard on top. rustbgpd exports the full BMP monitoring trio on one exporter (RFC 7854 Adj-RIB-In, RFC 8671 Adj-RIB-Out as byte-exact wire PDUs, RFC 9069 Loc-RIB), selectable per collector.
Proven by: M24
(BMP Initiation / PeerUp / RouteMonitoring ordering vs a BMP receiver)
and M81 (the trio plus BMPv4, validated against three independent
decoders at once: pmacct, gobmp, and tshark). Config shape derived
from
tests/interop/configs/rustbgpd-m81-bmp-rr.toml
and examples/route-collector/.
Config
The example is an RR that also feeds the monitoring stack; the same
[bmp] / [event_history] / [mrt] blocks bolt onto any of the
other recipes unchanged.
[global]
asn = 65000
router_id = "10.255.0.1"
listen_port = 179
cluster_id = "10.255.0.1"
# RFC 8212 posture, stated explicitly. It governs eBGP sessions only, so
# these iBGP clients are unaffected; an eBGP neighbor added later needs
# explicit import and export policy.
ebgp_requires_policy = true
[global.telemetry]
prometheus_addr = "127.0.0.1:9179"
log_format = "json"
# Owner-only local socket (default mode 0600): clients are authorized as the
# implicit "local-operator" principal — no [security.grpc] block needed.
[global.telemetry.grpc_uds]
path = "/var/lib/rustbgpd/grpc.sock"
# Durable event outbox (ADR-0072): restart-safe replay cursor for
# SubscribeFromEvent. Opt-in — it costs memory/CPU at scale, which is
# why it is off by default. Enable it here because replay is the
# point of this deployment.
[event_history]
enabled = true
# Periodic MRT TABLE_DUMP_V2 snapshots — readable by bgpdump, BGPKIT,
# and the RouteViews/RIPE RIS toolchains.
[mrt]
output_dir = "/var/lib/rustbgpd/mrt"
dump_interval = 3600
compress = true
file_prefix = "feed"
[bmp]
sys_name = "rustbgpd-feed"
sys_descr = "fabric RR monitoring feed"
# Production collector: BMP v3 (the default), all three RIB views.
# On (re)connect the collector receives Initiation and the cached Peer Up
# state; the loc_rib view also gets a chunked Loc-RIB dump closed by
# End-of-RIB. rib_in_pre and rib_out_post have no reconnect dump and start
# at the next UPDATE. No daemon restart is needed to attach a collector.
[[bmp.collectors]]
address = "10.20.0.10:1790"
reconnect_interval = 5
monitor = ["rib_in_pre", "rib_out_post", "loc_rib"]
# Optional second collector with BMPv4 TLV framing. Code points follow
# draft-ietf-grow-bmp-tlv-21 and remain pre-IANA. Path Marking is temporarily
# unavailable because its draft type 5 collides with Sequence Number. Keep
# production collectors on v3 (the default); point v4 only at tooling that
# tracks the draft (M81 validates the v4 bytes with a raw oracle and tshark).
#[[bmp.collectors]]
#address = "10.20.0.11:1790"
#reconnect_interval = 5
#version = 4
#monitor = ["rib_in_pre", "rib_out_post", "loc_rib"]
[[neighbors]]
address = "10.0.0.11"
remote_asn = 65000
description = "pe-1"
hold_time = 90
route_reflector_client = true
families = ["ipv4_unicast", "ipv6_unicast", "l3vpn_ipv4_unicast"]
[[neighbors]]
address = "10.0.1.2"
remote_asn = 65000
description = "pe-2"
hold_time = 90
route_reflector_client = true
families = ["ipv4_unicast", "ipv6_unicast", "l3vpn_ipv4_unicast"]Passive iBGP relay
Use a passive relay when two routers feed a collector through rustbgpd and the relay must advertise nothing. The route-collector starter provides the same accept-all import and deny-all export posture for eBGP sources. This minimal IPv4 configuration uses two sources in the relay's own AS and sends only the pre-policy Adj-RIB-In BMP view:
[global]
asn = 65001
router_id = "10.255.0.1"
listen_port = 179
ebgp_requires_policy = true
[global.telemetry]
prometheus_addr = "127.0.0.1:9179"
log_format = "json"
[global.telemetry.grpc_uds]
path = "/var/lib/rustbgpd/grpc.sock"
[bmp]
sys_name = "rustbgpd-passive-relay"
[[bmp.collectors]]
address = "10.20.0.10:1790"
monitor = ["rib_in_pre"]
[policy]
import_chain = ["observe-all"]
export_chain = ["deny-all"]
[policy.definitions.observe-all]
default_action = "permit"
[policy.definitions.deny-all]
default_action = "deny"
[[neighbors]]
address = "10.0.0.2"
remote_asn = 65001
families = ["ipv4_unicast"]
[[neighbors]]
address = "10.0.1.2"
remote_asn = 65001
families = ["ipv4_unicast"]Keep the deny-all export chain attached: an empty iBGP export chain permits routes. This configuration enables no forwarding backend; received routes remain available to the daemon's RIB queries and BMP feed.
At the collector, clear that peer's route state on PeerDown, as specified by RFC 7854 §4.9. Count that state teardown separately from explicit RouteMonitoring withdrawals. Connect the collector before feeding routes. After a BMP reconnect, cached PeerUp messages do not rebuild the received table: Adj-RIB-In has no automatic reconnect dump. See the BMP reconnect caveat.
The dated withdrawal parity receipt and its runnable lab prove one bounded shape: eight IPv4 /32s per peer, three announce/withdraw rounds and two flaps per peer, with 24 matching wire/BMP withdrawals per peer and zero exported announcements. This is not full-table memory, sustained-load, backpressure or reconnect-completeness evidence.
Verify
$ export RUSTBGPD_ADDR=unix:///var/lib/rustbgpd/grpc.sock
$ rbgp health
$ rbgp neighborBMP: on the collector you should see Initiation, the Peer Ups, the
chunked Loc-RIB dump and its End-of-RIB (loc_rib only), then live
RouteMonitoring deltas for every selected view. From the daemon side the
collector connection state is visible in the logs (log_format = "json") and
the BMP metrics below. Periodic peer Stats Reports also carry RFC 9972
post-policy accepted-route gauges, exact type-22 IPv4/IPv6-unicast policy
reject counts, and exact per-path RPKI Invalid/Valid/NotFound types 35/36/37.
The RPKI rows cover negotiated IPv4/IPv6-unicast families, count every
Add-Path identity, and include authoritative zero rows; they are omitted when
no VRP table is installed. Type 22 is present only while the default-on
rejected-route store is authoritative: disabling retention or reaching its
capacity omits the rows after the first eviction instead of publishing a
partial count.
During a Loc-RIB bootstrap or dump, watch
bmp_loc_rib_dump_live_buffer_depth{collector} for the current queued live
rows and bmp_loc_rib_dump_live_buffer_high_watermark{collector} for the
connection generation's peak. The collector value is the configured socket
address including port. A high-water mark near the fixed 8192-row bound means
the collector or dump path needs investigation before the next burst.
Event replay (the bridge contract): live tail plus durable replay
from a cursor. --from-event-id 0 replays everything retained, then
tails; your bridge persists the last event_id it processed and
resumes from there after either side restarts:
$ rbgp events watch --from-event-id 0
$ rbgp events watch --category route,session --from-event-id 41236While this CLI process remains running, a clean stream end or gRPC
UNAVAILABLE reconnects from the highest successfully flushed top-level
BgpEvent.event_id. Reconnect delay starts at 1 second, doubles to a 30-second
cap, and resets after a complete human or JSON record plus newline has been
written and stdout has been flushed. The CLI preserves the full filter set on
every request. Lag frames without a top-level event ID do not move the cursor;
all other RPC statuses and output failures are terminal. Cursorless OTC
subscriptions and ordinary WatchEvents streams remain one-shot. This
process-local retry does not replace the bridge's downstream-confirmed,
persisted cursor across CLI restarts.
The same contract over raw gRPC is
EventService.SubscribeFromEvent — see
examples/event-bridge/ for a
complete bridge (gRPC → JSON lines) you can pipe into Kafka, NATS,
or Vector. rustbgpd deliberately is not an event bus; the outbox is
a bounded SQLite WAL store with a monotonic cursor
(ADR-0072).
MRT: force a dump and inspect it with your usual tooling:
$ rbgp mrt-dump
$ ls /var/lib/rustbgpd/mrt/
feed.20260703.120001.123456789.mrt.gzLooking glass (optional): for status, peer, accepted-route,
filtered-route, and noexport views in an Alice-LG-style frontend, run the
examples/birdwatcher-adapter/
against the local Unix socket, or use its dedicated authenticated observer
listener pattern for least privilege. (The in-daemon
[global.telemetry.looking_glass] server has been removed.) The filtered
view is served from PolicyService.ListRejectedRoutes with structured
reject reasons; the noexport view diffs the Loc-RIB best set against the
peer's Adj-RIB-Out and names each suppression's export gate via
RibService.ExplainAdvertisedRoute — see the adapter README for exactly
what each view contains.
Watch
Import the overview dashboard per GRAFANA.md; the
BMP / event-outbox row is populated once the features above are
configured. Key series:
| Metric | Meaning |
|---|---|
bgp_event_outbox_degraded | 1 = latched durability-impacting loss, committed-event delivery skip, or DB open/recovery/quarantine failure; expected shutdown reason=closed drops are excluded. Inspect the drop reason and daemon log: replay can remain available |
bgp_event_outbox_storage_failed | 1 = the event-history storage stopped at runtime; the outbox refuses events and durable subscriptions until a restart |
bmp_collector_drops_total | per-collector queue/dump failures; live fan-out Full/Closed automatically resets only that collector generation, replays cached Peer Up state, and rebuilds configured Loc-RIB state after the one-second reconnect |
bmp_source_drops_total | per-peer BMP events dropped at the source tap, including a periodic stats report whose session-state query timed out |
bmp_replay_attempts_total | PeerUp-cache replays on collector reconnect |
bgp_rib_outbound_registered_peers | feed coverage: peers whose routes the views carry |
Failure modes
SubscribeFromEvent / rbgp events watch --from-event-id returns
FAILED_PRECONDITION. The daemon is running with
[event_history].enabled = false (the default). Enable it and
restart — the outbox fields are restart-required.
The bridge missed events during a burst. Live streams emit
stream_lagged warnings when a bounded source dropped events for a
slow consumer. That is the signal to resume via the durable cursor
(--from-event-id <last-processed>) rather than the live ring.
A BMP collector shows nothing after a network blip. The daemon
redials with backoff capped by reconnect_interval and replays cached Peer Up
state. It performs a fresh EoR-closed table dump only for collectors with
monitor = ["loc_rib"]. Adj-RIB-In and Adj-RIB-Out have no automatic
reconnect dump. For a late outbound collector, follow the
outbound capture procedure.
Its experimental replay-out operation reannounces the selected peer's routes
on the live BGP session; successful scheduling alone does not prove a complete
capture. See Known issues. Check the
collector-side listener first, then the daemon log for connection and
bootstrap failures.
pmacct rejects the v4 stream (BMPv4 BGP PDU TLV != 1). Known and
expected — see the caveat in the config above. Move that collector to
version = 3 (or drop the version key; 3 is the default).
The events DB was corrupted by a crash. The bad file is renamed
events.db.stale, the daemon continues (pass-through) unless
[event_history].required = true, and bgp_event_outbox_degraded
latches to 1 until an operator restart. Details:
CONFIGURATION.md §event_history.