Skip to content

fix(observability): record every documented metric, including notifications - #33

Merged
arg1998 merged 6 commits into
mainfrom
fix/otel-unrecorded-metrics
Sep 29, 2026
Merged

arg1998 merged 6 commits into
mainfrom
fix/otel-unrecorded-metrics

Conversation

@arg1998

@arg1998 arg1998 commented Sep 29, 2026

Copy link
Copy Markdown
Owner

The bug

A pre-release audit found metrics that are defined in the instrument catalogue and promised by spec 10 §7 and the telemetry guide, but never reach an OTLP collector. I re-checked it against main (7873a06):

  • Never recorded: browserhive.attention.wait, browserhive.session.launch.duration, browserhive.ws.connections, browserhive.ws.buffered_bytes, browserhive.ws.frames_dropped, browserhive.db.dropped_writes, browserhive.browser.rss_bytes and browserhive.process.event_loop_lag. The last one was also missing from the audit.
  • Documented attributes never set: closed_reason on session.lifetime, kind on attention.open, table on retention.pruned_rows.
  • The three notification counters were in fact recorded. build-domain passes them into the outbox, the act-button service and the report scheduler, and they showed up in the live check below. They were missing from the guide, though. notifications.reports also never emitted the documented manual outcome, and it emitted an undocumented revised for silent in-app anomaly revisions. Both are fixed. The composition now wires these counters through one helper, notificationCounters(), and the guard test uses the same helper.
  • Found during the live run: the SDK keeps sending the last value of an observable-gauge series after it stops being observed. So a closed session's browser memory, or a closed WebSocket's buffered bytes, would have been exported forever.
  • Before this change the consumers were wired even when --otelSignals left metrics out.

Spec first

The first commit makes spec 10 §7 the single source of truth: every instrument with its type, unit, attributes (closed value sets in parentheses) and when it is recorded. docs/guide/telemetry.md mirrors that table and now includes the notification and process metrics. Spec 09 names the guard test. Changes to the tables:

  • Units were added: ms, By, and UCUM annotations such as {call}.
  • process.* is split into three rows.
  • db.dropped_writes now lists the recorder tables, and every table is reported from the start at 0.
  • The manual outcome is defined for reports.
  • Browser memory is defined as the sum of the tree's RSS, on Linux and macOS.
  • The spec states that observable gauges report only the series observed at each export.

What was wired where

Metric Source
session.launch.duration {channel, stealth} first session.updated that carries launchMs (bus)
session.lifetime {closed_reason} session.closed (bus)
attention.open {kind} attention.* and vault.confirm.* (bus). Only requests seen opening are settled, so the startup reconcile can't push it below zero.
attention.wait {status} attention.resolved, from waited_ms
retention.pruned_rows {table} retention.completed, through a new internal prunedByTable key that the WS schema strips
ws.connections, ws.buffered_bytes (top 50), ws.frames_dropped {screencast, logs, feed} realtime hub, via realtimeMetrics(hub). The hub keeps plain dropped-frame totals.
db.dropped_writes {table} the write queue's new droppedWritesByTable
browser.rss_bytes {session_id} a 10 s sampler. It reads each session's browser pid once, via SessionHandle.browserPid() (a browser-level CDP SystemInfo.getProcessInfo; no page target is touched), then sums the process tree from /proc (Linux) or ps (macOS). Windows gets no data points.
process.event_loop_lag perf_hooks.monitorEventLoopDelay, reporting the p99 since the previous export

Observable gauges are now collected with delta temporality. OTLP gauges don't carry temporality, so the only effect is that series nobody observes anymore disappear. Counters stay cumulative. The app layer imports no infra, and dependency-cruiser is clean.

With OTel off, or with the metrics signal off, nothing is subscribed, sampled or timed. The only additions that always run are plain integer and map increments on the drop paths of the hub and the write queue.

Guard test

packages/browserhive/test/composition/metrics-guard.test.ts works in two steps.

  1. It parses both Markdown tables and requires spec == docs == METRIC_DEFINITIONS, row by row.

  2. A SOURCES table has to cover every catalogue metric exactly once. Each source is driven for real:

    • the bus consumers;
    • the SQLite write queue;
    • a real realtime hub with a socket;
    • the sampler reading a real process tree;
    • the event-loop monitor;
    • buildOps with a real webhook channel, an act-button press and an on-demand digest.

    Everything is exported through the real createTelemetry to an OTLP/HTTP receiver inside the test. Every metric must arrive with its documented kind, unit and attribute keys, and a closed session's gauge series must disappear.

I mutation-checked it. Each of these makes it fail: dropping a callback, dropping an attribute, adding an undocumented attribute, and reverting the delta-gauge change. There are also unit tests for every new piece, and an integration test that reads each real browser's pid and process-tree memory.

Real OTLP check

I ran the built daemon with --otel --otelProtocol=http/json against a small OTLP receiver, with an ntfy channel on FakePlatforms, a blocklist and the fake bw. The traffic was:

  • two MCP sessions;
  • a failing click;
  • a blocked navigate;
  • a vault_fill;
  • request_attention, answered by pressing the ntfy "Mark resolved" act button, plus a press of a token that was never issued;
  • a dashboard login and WebSocket;
  • an on-demand digest;
  • closing a session;
  • then SIGTERM.
22 metrics received in 4 exports:

browserhive.attention.open  [up-down counter, {request}]  1 pts  {"kind":"attention"}=0
browserhive.attention.wait  [histogram, ms]  1 pts  {"status":"resolved"}={"count":1,"sum":202}
browserhive.blocklist.hits  [counter, {hit}]  1 pts  {"source":"tool"}=1
browserhive.browser.rss_bytes  [gauge, By]  1 pts  {"session_id":"otela-2jutnqs5"}=1020932096
browserhive.db.dropped_writes  [counter, {write}]  8 pts  {"table":"sessions"}=0  {"table":"tool_calls"}=0  {"table":"pages"}=0
browserhive.db.size_bytes  [gauge, By]  1 pts  {}=446464
browserhive.db.write_queue.depth  [gauge, {write}]  1 pts  {}=0
browserhive.notifications.actions  [counter, {press}]  2 pts  {"channel_kind":"ntfy","outcome":"unknown"}=1  {"channel_kind":"ntfy","outcome":"done"}=1
browserhive.notifications.deliveries  [counter, {delivery}]  1 pts  {"channel_kind":"ntfy","status":"sent"}=4
browserhive.notifications.reports  [counter, {report}]  2 pts  {"kind":"digest.daily","outcome":"manual"}=1  {"kind":"digest.daily","outcome":"in_app"}=1
browserhive.process.event_loop_lag  [gauge, ms]  1 pts  {}=3.577855
browserhive.process.heap_bytes  [gauge, By]  1 pts  {}=71143067
browserhive.process.rss_bytes  [gauge, By]  1 pts  {}=229859328
browserhive.session.launch.duration  [histogram, ms]  1 pts  {"channel":"chromium","stealth":true}={"count":2,"sum":1162}
browserhive.session.lifetime  [histogram, ms]  1 pts  {"closed_reason":"user"}={"count":1,"sum":30786}
browserhive.sessions.active  [up-down counter, {session}]  1 pts  {"harness":"other"}=1
browserhive.tool_call.duration  [histogram, ms]  6 pts  {"tool":"launch_session"}={"count":2,"sum":1180}  {"tool":"navigate"}={"count":2,"sum":56}  {"tool":"click"}={"count":1,"sum":30007}
browserhive.tool_calls  [counter, {call}]  7 pts  {"tool":"launch_session","ok":true,"harness":"other"}=2  {"tool":"navigate","ok":true,"harness":"other"}=1  {"tool":"click","ok":false,"harness":"other","error_code":"ELEMENT_NOT_ACTIONABLE"}=1
browserhive.vault.fills  [counter, {fill}]  1 pts  {"result":"blocked"}=1
browserhive.ws.buffered_bytes  [gauge, By]  1 pts  {"connection_id":"c-DTjILvV3nD"}=0
browserhive.ws.connections  [up-down counter, {connection}]  1 pts  {}=1
browserhive.ws.frames_dropped  [counter, {frame}]  3 pts  {"channel":"screencast"}=0  {"channel":"logs"}=0  {"channel":"feed"}=0

22 of the 23 metrics arrived live. retention.pruned_rows can't show up in a short run, because retention first runs 6 h after start. The guard test covers it. A ps check of each browser tree (~1.06–1.09 GB) agrees with the metric, and the closed session no longer appeared in the next export.

Local gate

  • bun run check, test:goldens, build, package:check and license:check all pass.
  • test:integration: 68 passed, 4 skipped.
  • e2e, replicated like CI: 15 passed, 1 skipped.
  • The website builds.

Spec 10 §7 and the telemetry guide now list every instrument with its
type, unit, attributes and when it is recorded, including the three
notification counters and the three process gauges, and name the guard
test that keeps the tables, the catalogue and the recorded metrics in
step (spec 09 §3.2).
The adapter only fed eleven of the catalogue's instruments. It now
records the rest from their sources:

- session launch duration (channel, stealth) from the first
  session.updated carrying launch_ms; lifetime by closed_reason
- attention.open by kind (attention and vault confirmations) and the
  attention wait by status, from the broker's events
- retention pruned rows by table, carried internally on
  retention.completed
- WebSocket connections, buffered bytes (top 50) and frames dropped by
  channel, read from the realtime hub's running totals at export
- dropped writes by table, read from the write queue's totals
- each live session's browser process-tree RSS, sampled every 10 s from
  the browser pid (one browser-level DevTools read, cached) and /proc or
  ps
- the p99 event-loop delay since the previous export

Every instrument now carries the unit of its catalogue row. The report
scheduler counts on-demand digests as manual and no longer counts a
silent revision of an in-app anomaly alert. With telemetry off nothing
is subscribed, sampled or timed.
A table-driven guard parses the metric tables of spec 10 §7 and the
telemetry guide and requires them to match the instrument catalogue row
for row, then drives every source the composition root wires (bus
consumers, the SQLite write queue, the realtime hub, the browser-memory
sampler over a real process tree, the event-loop monitor, and the
notification outbox, act buttons and report scheduler built by
buildOps) and requires each instrument to reach a real OTLP/HTTP
receiver with its documented type, unit and attribute keys.

Unit tests cover the process-tree reader, the event-loop monitor, the
sampler, the hub's and the write queue's totals, the catalogue units and
the report outcomes; an integration test reads each real browser's pid
and process-tree memory.
…corder table

The SDK repeats the last value of an observable gauge's series that is
no longer observed, so a closed session's browser memory or a closed
WebSocket's buffered bytes was exported forever. Observable gauges are
now collected with delta temporality, which OTLP gauges do not carry, so
each export holds only what exists; counters stay cumulative.

browserhive.db.dropped_writes now reports every recorder table from the
start at 0, so a rate over the series works before the first drop.

Spec 10 §7 and the guide say both, and that browser memory is the sum
of the processes' RSS. The guard test checks that a closed session's
series disappears and that the spec's recorder tables match.
…orted

With --otelSignals traces,logs the meter is a no-op, so the bus
consumers, the browser-memory sampler and the event-loop monitor are no
longer started for nothing.
@arg1998
arg1998 marked this pull request as ready for review September 29, 2026 17:02
@arg1998
arg1998 merged commit f4ebe1a into main Sep 29, 2026
23 of 25 checks passed
@arg1998
arg1998 deleted the fix/otel-unrecorded-metrics branch September 29, 2026 17:58
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant