diagnostics-prometheus plugin. It listens to trusted diagnostics plus
internally tagged, dispatcher-owned diagnostic events (queue, memory, and
session-recovery signals), and renders a Prometheus text endpoint at:
text/plain; version=0.0.4; charset=utf-8, the standard
Prometheus exposition format.
For traces, logs, OTLP push, and OpenTelemetry GenAI semantic attributes, see OpenTelemetry export.
Quick start
1
Install the plugin
2
Enable the plugin
- Config
- CLI
3
Restart the Gateway
The HTTP route is registered at plugin startup, so reload after enabling.
4
Scrape the protected route
Send the same gateway auth your operator clients use:
5
Wire Prometheus
diagnostics.enabled defaults to true; set it to false only in tightly constrained environments. When it is false at exporter startup, the plugin still registers the HTTP route, but no diagnostic events or runtime identity are recorded, so the response is empty.Metrics exported
For model-call metrics,
observation_unit="request" measures one observable
provider request. observation_unit="turn" measures a synthetic Claude Code
or Codex CLI agent turn that can contain multiple hidden provider requests.
Keep those series separate when comparing latency.
Gateway RPC metrics cover valid authenticated WebSocket requests, including
subsequent rejections. first_response measures receipt through the first frame
accepted by the sender; unavailable or suppressed sends have no duration sample.
handler measures actual handler invocation through return or throw, and
admission measures receipt through that invocation. queue_wait measures only
operator request start-queue wait, separately from command/session lane metrics.
They measure elapsed time, not CPU time. Early acknowledgments and responses
after handler return are distinct from completed agent work. See
Gateway RPC timing semantics.
Receipt begins after the connected client’s request frame passes validation.
These timings exclude CLI startup, local diagnostics, connection/authentication
setup, and event-loop delay before request dispatch. Histograms record completed
observations: an unfinished handler has no handler-duration sample yet. Compare
request counts, completed timings, and event-loop observations when investigating
a timeout; low handler latency alone does not establish a responsive client path.
RPC method labels contain canonical core method names, other for plugin
methods, or unknown. Outcome totals aggregate by phase and outcome without a
method dimension. Each method with all four timings occupies five aggregate
samples in the shared 2,048-sample cap. A duration histogram occupies one sample
but expands into 19 scrape series (buckets, sum, and count). Existing samples keep
updating when the cap fills; unseen RPC or other operational samples are refused
and increment openclaw_prometheus_series_dropped_total. Monitor that counter:
coverage of every core method can fill the cap, so a zero value matters when
interpreting totals or latency percentiles. Async diagnostic queue saturation can
also drop observations, reported by openclaw_diagnostic_async_queue_dropped_total.
Runtime identity
openclaw_gateway_build_info has value 1 and identifies the process serving
the scrape. Its process_instance_id is the same process-owned UUID returned by
system.info; it changes when the process restarts, including when a PID is
reused. build_id matches the loaded build reported by hello.server.buildId
and is omitted when that provenance is unavailable. Updating files on disk does
not change the running process’s identity.
When diagnostics are enabled, the exporter captures these facts at service startup,
before recording events.
The metric uses one aggregate sample under the existing cap. Older hosts without
the optional runtime-identity capability omit it. The UUID is confined to this
info metric; it is not added to RPC or other metric labels.
Use the info sample from the same scrape to attribute new measurements and split
counter intervals at process changes. It is not a health signal, a request ID,
or an exporter epoch: restarting the exporter in the same process resets its
counters while retaining the process identity. It cannot relabel older samples
or establish complete diagnostic-loss coverage.
Event-loop observation windows
openclaw_liveness_cpu_core_ratio measures whole-process CPU usage in core
equivalents, including worker and native threads, and can exceed 1. Interpret
it alongside main-thread delay and utilization; see
CPU pressure and event-loop delay.
The event-loop histogram records the maximum delay from each completed Gateway
health-monitor window. The counter sums the seconds represented by those
windows. Both are cumulative: a later healthy window does not erase an earlier
high-delay observation. Readiness, status, and scrape requests consume completed
observations without advancing or resetting the sampling window.
The monitor samples elapsed event-loop intervals every 20 milliseconds and
completes a window after at least one second, or sooner for a delay warning.
It preserves the pending interval across ordinary window resets, so reading
health before an overdue sample cannot erase that delay. Histogram counts are window counts, not stall
counts. Histogram quantiles describe window maxima, not the sampled event-loop
delay distribution or its overall p99. These metrics have no request labels or
trace attribution and do not identify the JavaScript function that blocked.
Collection uses the plugin enablement above. It starts when an interested
metrics exporter is running; it does not backfill earlier windows. Intentional
monitor resets discard the unfinished window. Diagnostic queue drops, the
exporter’s series cap, and process restarts can also lose observations. Watch
the existing drop counters and the represented-duration counter when assessing
coverage. Readiness decisions and persistent liveness-warning thresholds are unchanged.
Memory and process churn
openclaw_memory_bytes exposes rss, heap_total, heap_used, external,
array_buffers, worker_heap_total, and worker_heap_used. RSS covers the
whole process. The unprefixed heap and native-buffer values cover the main
isolate; array_buffers is included in external, so do not add them together.
These values do not account for every native allocation or allocator arena.
Worker totals sum completed native heap samples from live Workers created after
the resource registry starts. On Node, this includes direct plugin Workers;
nested Workers and V8’s internal threads are outside the parent registry.
The 30-second diagnostics heartbeat starts a nonblocking refresh, retaining at
most one outstanding request per Worker. Samples expire after 60 seconds and
are removed when the Worker exits. Compare openclaw_worker_heap_sampled_count
with openclaw_worker_count: startup, unavailable APIs, and stalled Workers can
produce partial totals. No heap snapshot or extra sampling timer is created.
openclaw_worker_heap_used_bytes{script="..."} sums fresh heap samples for each
Worker script. Labels use a fixed allowlist of runtime-entrypoint and pooled Worker basenames, normalized to
.js in source and packaged runs; unknown, eval, and directly created Workers
use other. Full paths and eval source are never recorded. A script’s series
disappears when it has no live, fresh samples. Memory-pressure logs include the
same byte counts and Worker coverage counts, plus workerHeaps: the five largest
individual fresh Worker heaps as {script, heapUsed, heapTotal} (bytes), using the same
bounded script names.
openclaw_child_process_spawn_total{family="..."} counts successful launches
through OpenClaw’s shared spawn and exec owners, including brokered launches.
Diagnostics must be enabled. The existing heartbeat publishes accumulated
counts after at least one minute, with debug logs reporting counts and rates
using the actual elapsed interval. Failed launches, direct calls bypassing
these owners, and descendants started by children are excluded. Families are
a fixed executable-name allowlist; unrecognized commands become other.
Arguments and paths are never recorded. For launches per minute, use
60 * rate(openclaw_child_process_spawn_total[5m]); this window accommodates
the minute-batched publication. Neither accounting path changes pressure
thresholds or user-tool execution.
Garbage collection duration
openclaw_gc_duration_seconds records elapsed garbage collection (GC) duration
reported by Node.js for the hosting JavaScript isolate. Each observation is one
GC entry, not CPU time, allocated bytes, or a guaranteed stop-the-world pause.
Compare its bucket counts with event-loop window maxima to investigate GC as a
possible contributor to stalls; matching scrape intervals do not prove causality.
Collection uses the existing diagnostics enablement and starts when the
diagnostics heartbeat observes an interested consumer, such as a metrics exporter. A consumer added
after heartbeat startup may wait until the next 30-second tick, or longer if the
event loop is stalled. Entries preceding observer activation are not backfilled.
Demand is checked when entries are delivered, so a brief consumer gap before the
next heartbeat can still yield delayed observations. Losing the last
consumer suppresses new exports; the observer disconnects at the next heartbeat.
Disabling diagnostics or stopping the heartbeat disconnects it immediately.
The histogram is absent until the first observation, so absence does not prove
zero GC. Queue drops, the series cap, observation gaps and process restarts limit
coverage. Diagnostics disable/re-enable preserves the exporter’s existing
counters; restarting the exporter resets them as usual. No extra timer, GC
trigger, trace attribution or application payload is collected.
Label policy
Bounded, low-cardinality labels
Bounded, low-cardinality labels
Prometheus labels stay bounded and low-cardinality. The exporter does not emit raw diagnostic identifiers such as
runId, sessionKey, sessionId, callId, toolCallId, message IDs, chat IDs, or provider request IDs.Label values are redacted and must match OpenClaw’s low-cardinality character policy. Values that fail the policy are replaced with unknown, other, or none, depending on the metric. Labels that look like scoped agent session keys are also replaced with unknown.Series cap and overflow accounting
Series cap and overflow accounting
The exporter caps retained time series in memory at 2048 series across counters, gauges, and histograms combined. New series beyond that cap are dropped, and
openclaw_prometheus_series_dropped_total increments by one each time.Watch this counter as a hard signal that an attribute upstream is leaking high-cardinality values. The exporter never lifts the cap automatically; if it climbs, fix the source rather than disabling the cap.What never appears in Prometheus output
What never appears in Prometheus output
- prompt text, response text, tool inputs, tool outputs, system prompts
- Talk transcripts, audio payloads, call ids, room ids, handoff tokens, turn ids, and raw session ids
- raw provider request IDs (only bounded hashes, where applicable, on spans — never on metrics)
- session keys and session IDs
- hostnames, file paths, secret values
PromQL recipes
Choosing between Prometheus and OpenTelemetry export
OpenClaw supports both surfaces independently. You can run either, both, or neither.- diagnostics-prometheus
- diagnostics-otel
- Pull model: Prometheus scrapes
/api/diagnostics/prometheus. - No external collector required.
- Authenticated through normal Gateway auth.
- Surface is metrics only (no traces or logs).
- Best for stacks already standardized on Prometheus + Grafana.
Troubleshooting
Empty response body
Empty response body
- Check that
diagnostics.enabledis not set tofalsein config (it defaults totrue). - Confirm the plugin is enabled and loaded with
openclaw plugins list --enabled. - Generate some traffic; counters and histograms only emit lines after at least one event.
403 missing scope: operator.read
403 missing scope: operator.read
The caller authenticated, but its effective operator scopes do not include
operator.read. This happens when an identity-bearing auth mode such as trusted-proxy maps the scraper to a named role whose scope ceiling excludes reads. Grant the scraper role operator.read (or operator.write / operator.admin, which imply it).openclaw_prometheus_series_dropped_total is climbing
openclaw_prometheus_series_dropped_total is climbing
A new attribute is exceeding the 2048-series cap. Inspect recent metrics for an unexpectedly high-cardinality label and fix it at the source. The exporter intentionally drops new series instead of silently rewriting labels.
Prometheus shows stale series after a restart
Prometheus shows stale series after a restart
The plugin keeps state in memory only. After a Gateway restart, counters reset to zero and gauges restart at their next reported value. Use PromQL
rate() and increase() to handle resets cleanly.Related
- Diagnostics export — local diagnostics zip for support bundles
- Health and readiness —
/healthzand/readyzprobes - Logging — file-based logging
- OpenTelemetry export — OTLP push for traces, metrics, and logs