Prometheus, OTLP and Grafana (pulse-export)¶
pulse-export puts a probed ROS 2 stack on a Grafana dashboard with no rebuild and no ROS
dependency. Like pulse-top it is a pure log consumer: it tails the log files the probe
already writes (the default text format or jsonl), serves them as Prometheus metrics on /metrics, and can push the same series
to an OpenTelemetry collector over OTLP/HTTP. Standard library only; it ships in the
ros2-pulse-top package.
One command: Prometheus + Grafana¶
git clone https://github.com/TanayK07/ros2_pulse && cd ros2_pulse
docker compose -f examples/grafana/docker-compose.yml up --build
# http://localhost:3000 opens on the ros2_pulse dashboard, no login needed to view
The compose file (examples/grafana/) runs three containers:
pulse-export reading the host's /tmp/topic_freq.*.log (the probe's default output, one file
per probed process), Prometheus scraping it every second, and Grafana with the datasource and
dashboard provisioned. Start the probe on the robot as usual (either output format works):
LD_PRELOAD=libros2_pulse.so ros2 launch my_robot bringup.launch.py
Set PULSE_LOG_DIR if the probe writes somewhere other than /tmp (a non-default TMPDIR).
No robot at hand: PULSE_EXPORT_ARGS=--demo docker compose -f examples/grafana/docker-compose.yml up --build
exports pulse-top's scripted demo graph (a /scan stall, a /cmd_vel rate sag, a recv_lag
on /camera/image_raw, /localization going missing).
The example is for a local look: anonymous viewing is on and the Grafana admin password is the default. For an existing Prometheus, skip compose and point a scrape job at the exporter.
Running the exporter directly¶
pip3 install ros2-pulse-top
pulse-export # every $TMPDIR/topic_freq.<pid>.log, port 9464
pulse-export '/var/log/topic_freq.*.log' # a quoted glob; new files are picked up as nodes start
pulse-export --otlp http://localhost:4318 # also push OTLP/HTTP JSON every 15 s
pulse-export --once /tmp/topic_freq.4242.log # print the /metrics text once and exit
| Flag | Default | Meaning |
|---|---|---|
file |
every $TMPDIR/topic_freq.*.log |
Log path or quoted glob, re-expanded on every poll. |
--port |
9464 |
Prometheus port; 0 disables the HTTP server (OTLP only). |
--bind |
0.0.0.0 |
Listen address. Use 127.0.0.1 to keep topic names off the network. |
--poll |
1.0 |
Seconds between log reads. |
--otlp URL |
off | Push OTLP/HTTP JSON here. A bare http://host:4318 gets /v1/metrics appended. |
--otlp-interval |
15 |
Seconds between pushes. |
--otlp-header K=V |
none | Extra request header, repeatable (for example Authorization=Basic ...). |
--lag-tol, --lag-windows |
0.10, 3 |
The recv_lag detector, same rules as pulse-top. |
--demo |
off | Export pulse-top's self-generated demo log. |
--once |
off | Read the log once, print the exposition, exit. |
Metrics¶
Every series is labelled pid, taken from the probe's file name topic_freq.<pid>.log (jsonl
records carry no pid). A log with any other name, such as a shared
ROS_TOPIC_STATS_OUTPUT_FILE, is labelled with its file name instead.
| Metric | Type | Labels | Meaning |
|---|---|---|---|
ros2_pulse_topic_publish_rate_hertz |
gauge | pid, topic, path |
Publish rate, path is inter or intra process. |
ros2_pulse_topic_receive_rate_hertz |
gauge | pid, topic, path |
Subscription callback rate. |
ros2_pulse_topic_max_gap_seconds |
gauge | pid, topic, side |
Largest inter-arrival gap, side is pub or recv. Only with ROS_TOPIC_STATS_JITTER. |
ros2_pulse_topic_recv_lag_deficit_ratio |
gauge | pid, topic |
While a recv_lag warn is active: (pub - recv) / pub. pid is the subscriber's. |
ros2_pulse_warn_active |
gauge | pid, kind, topic, node |
1 per warn in the latest window: topic_rate, topic_gap, node_missing, recv_lag. |
ros2_pulse_node_up |
gauge | pid, node |
1 if listed in the latest window, 0 if the probe reports it missing. |
ros2_pulse_node_last_seen_age_seconds |
gauge | node |
Seconds since a window last listed the node. |
ros2_pulse_process_window_seconds |
gauge | pid |
The process's window period. |
ros2_pulse_process_last_window_age_seconds |
gauge | pid |
Seconds since the process's last window was read. Grows when a process stops flushing. |
ros2_pulse_process_last_window_timestamp_seconds |
gauge | pid |
Probe timestamp of that window, Unix seconds. |
ros2_pulse_windows_total |
counter | pid |
Windows read. |
ros2_pulse_warns_total |
counter | pid, kind |
Probe warns read. Catches a warn that came and went between two scrapes. |
ros2_pulse_lines_skipped_total |
counter | none | Lines that were not probe output in either format (truncated, foreign). |
Useful queries: max by (topic) (ros2_pulse_topic_publish_rate_hertz) is the probe's own
headline rate (one publish can fire both paths, so the busier one, not the sum);
sum by (topic, pid) (ros2_pulse_topic_receive_rate_hertz) is callbacks per subscribing
process; increase(ros2_pulse_warns_total[5m]) > 0 alerts on any probe warn.
What a scrape means¶
A scrape shows each process's latest window, and keeps the probe's rule that absent is not zero:
- A key the probe did not write (no receive side in this process, no gap measured) is an
absent series, never a
0sample. A topic missing from a process's latest window drops out of that process's series. - A process whose last window is older than two of its periods plus one poll (a crash, a hang,
a clean exit) loses its topic, node and warn series, so a dead publisher's last rate is not
scraped as if current. Its
process_last_window_age_secondskeeps growing, which is the series to alert on. After five minutes of silence it is forgotten entirely. - Ages use the exporter's clock at read time, not the probe's
ts_ns, so a log copied from another machine or replayed after the fact does not read as hours stale. - Scrape at or under the window period (
ROS_TOPIC_STATISTICS_PUBLISH_PERIOD, 5 s by default). A slower scrape skips windows; the counters still count them.
OTLP¶
--otlp posts an ExportMetricsServiceRequest in the OTLP/HTTP JSON encoding: gauges as
gauges, counters as cumulative monotonic sums with the _total suffix dropped (an
OTLP-to-Prometheus translation adds it back, so the dashboard queries work on either path),
UCUM units (Hz, s, 1), resource service.name=ros2_pulse and host.name. It uses
urllib, so there is no opentelemetry-sdk to install; a failed push is reported once per
outage on stderr and retried on the next interval. Checked against OpenTelemetry Collector
0.128.0 (partialSuccess empty, no rejected points).