fix(observability): record every documented metric, including notifications - #33
Merged
Merged
Conversation
Spec 10 §7 and the telemetry guide now list every instrument with its type, unit, attributes and when it is recorded, including the three notification counters and the three process gauges, and name the guard test that keeps the tables, the catalogue and the recorded metrics in step (spec 09 §3.2).
The adapter only fed eleven of the catalogue's instruments. It now records the rest from their sources: - session launch duration (channel, stealth) from the first session.updated carrying launch_ms; lifetime by closed_reason - attention.open by kind (attention and vault confirmations) and the attention wait by status, from the broker's events - retention pruned rows by table, carried internally on retention.completed - WebSocket connections, buffered bytes (top 50) and frames dropped by channel, read from the realtime hub's running totals at export - dropped writes by table, read from the write queue's totals - each live session's browser process-tree RSS, sampled every 10 s from the browser pid (one browser-level DevTools read, cached) and /proc or ps - the p99 event-loop delay since the previous export Every instrument now carries the unit of its catalogue row. The report scheduler counts on-demand digests as manual and no longer counts a silent revision of an in-app anomaly alert. With telemetry off nothing is subscribed, sampled or timed.
A table-driven guard parses the metric tables of spec 10 §7 and the telemetry guide and requires them to match the instrument catalogue row for row, then drives every source the composition root wires (bus consumers, the SQLite write queue, the realtime hub, the browser-memory sampler over a real process tree, the event-loop monitor, and the notification outbox, act buttons and report scheduler built by buildOps) and requires each instrument to reach a real OTLP/HTTP receiver with its documented type, unit and attribute keys. Unit tests cover the process-tree reader, the event-loop monitor, the sampler, the hub's and the write queue's totals, the catalogue units and the report outcomes; an integration test reads each real browser's pid and process-tree memory.
…corder table The SDK repeats the last value of an observable gauge's series that is no longer observed, so a closed session's browser memory or a closed WebSocket's buffered bytes was exported forever. Observable gauges are now collected with delta temporality, which OTLP gauges do not carry, so each export holds only what exists; counters stay cumulative. browserhive.db.dropped_writes now reports every recorder table from the start at 0, so a rate over the series works before the first drop. Spec 10 §7 and the guide say both, and that browser memory is the sum of the processes' RSS. The guard test checks that a closed session's series disappears and that the spec's recorder tables match.
…orted With --otelSignals traces,logs the meter is a no-op, so the bus consumers, the browser-memory sampler and the event-loop monitor are no longer started for nothing.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The bug
A pre-release audit found metrics that are defined in the instrument catalogue and promised by spec 10 §7 and the telemetry guide, but never reach an OTLP collector. I re-checked it against
main(7873a06):browserhive.attention.wait,browserhive.session.launch.duration,browserhive.ws.connections,browserhive.ws.buffered_bytes,browserhive.ws.frames_dropped,browserhive.db.dropped_writes,browserhive.browser.rss_bytesandbrowserhive.process.event_loop_lag. The last one was also missing from the audit.closed_reasononsession.lifetime,kindonattention.open,tableonretention.pruned_rows.build-domainpasses them into the outbox, the act-button service and the report scheduler, and they showed up in the live check below. They were missing from the guide, though.notifications.reportsalso never emitted the documentedmanualoutcome, and it emitted an undocumentedrevisedfor silent in-app anomaly revisions. Both are fixed. The composition now wires these counters through one helper,notificationCounters(), and the guard test uses the same helper.--otelSignalsleft metrics out.Spec first
The first commit makes spec 10 §7 the single source of truth: every instrument with its type, unit, attributes (closed value sets in parentheses) and when it is recorded.
docs/guide/telemetry.mdmirrors that table and now includes the notification and process metrics. Spec 09 names the guard test. Changes to the tables:ms,By, and UCUM annotations such as{call}.process.*is split into three rows.db.dropped_writesnow lists the recorder tables, and every table is reported from the start at 0.manualoutcome is defined for reports.What was wired where
session.launch.duration{channel, stealth}session.updatedthat carrieslaunchMs(bus)session.lifetime{closed_reason}session.closed(bus)attention.open{kind}attention.*andvault.confirm.*(bus). Only requests seen opening are settled, so the startup reconcile can't push it below zero.attention.wait{status}attention.resolved, fromwaited_msretention.pruned_rows{table}retention.completed, through a new internalprunedByTablekey that the WS schema stripsws.connections,ws.buffered_bytes(top 50),ws.frames_dropped{screencast, logs, feed}realtimeMetrics(hub). The hub keeps plain dropped-frame totals.db.dropped_writes{table}droppedWritesByTablebrowser.rss_bytes{session_id}SessionHandle.browserPid()(a browser-level CDPSystemInfo.getProcessInfo; no page target is touched), then sums the process tree from/proc(Linux) orps(macOS). Windows gets no data points.process.event_loop_lagperf_hooks.monitorEventLoopDelay, reporting the p99 since the previous exportObservable gauges are now collected with delta temporality. OTLP gauges don't carry temporality, so the only effect is that series nobody observes anymore disappear. Counters stay cumulative. The app layer imports no infra, and dependency-cruiser is clean.
With OTel off, or with the metrics signal off, nothing is subscribed, sampled or timed. The only additions that always run are plain integer and map increments on the drop paths of the hub and the write queue.
Guard test
packages/browserhive/test/composition/metrics-guard.test.tsworks in two steps.It parses both Markdown tables and requires spec == docs ==
METRIC_DEFINITIONS, row by row.A
SOURCEStable has to cover every catalogue metric exactly once. Each source is driven for real:buildOpswith a real webhook channel, an act-button press and an on-demand digest.Everything is exported through the real
createTelemetryto an OTLP/HTTP receiver inside the test. Every metric must arrive with its documented kind, unit and attribute keys, and a closed session's gauge series must disappear.I mutation-checked it. Each of these makes it fail: dropping a callback, dropping an attribute, adding an undocumented attribute, and reverting the delta-gauge change. There are also unit tests for every new piece, and an integration test that reads each real browser's pid and process-tree memory.
Real OTLP check
I ran the built daemon with
--otel --otelProtocol=http/jsonagainst a small OTLP receiver, with an ntfy channel onFakePlatforms, a blocklist and the fakebw. The traffic was:vault_fill;request_attention, answered by pressing the ntfy "Mark resolved" act button, plus a press of a token that was never issued;22 of the 23 metrics arrived live.
retention.pruned_rowscan't show up in a short run, because retention first runs 6 h after start. The guard test covers it. Apscheck of each browser tree (~1.06–1.09 GB) agrees with the metric, and the closed session no longer appeared in the next export.Local gate
bun run check,test:goldens,build,package:checkandlicense:checkall pass.test:integration: 68 passed, 4 skipped.