From 956d3557bbbe6327e8bbcb853b1e3926292ce66a Mon Sep 17 00:00:00 2001 From: Amir Ghorbani Date: Sun, 27 Sep 2026 11:10:03 -0400 Subject: [PATCH 1/7] docs(specs): notification contract, delivery outbox and channel decisions (D-32 to D-39) --- specs/00-decisions.md | 123 +++++++++++++++++++- specs/01-overall-architecture.md | 2 +- specs/02-mcp-and-tools.md | 1 + specs/03-admin-backend.md | 137 +++++++++++++++++++---- specs/04-admin-frontend.md | 6 +- specs/05-coding-standards.md | 1 + specs/08-cli-arguments-and-config.md | 19 +++- specs/09-testing.md | 5 +- specs/10-error-handling-and-telemetry.md | 7 +- 9 files changed, 268 insertions(+), 33 deletions(-) diff --git a/specs/00-decisions.md b/specs/00-decisions.md index 36212e5..0f0d656 100644 --- a/specs/00-decisions.md +++ b/specs/00-decisions.md @@ -401,14 +401,15 @@ Details in `04-admin-frontend.md`. **Decision.** - Notifications are produced server-side from the domain event bus, persisted (`notifications` table with read and dismissed state), and delivered in-app over the WS `notifications` topic. -- Producer rules: attention requested; session crashed; lease-expired reap; vault confirm pending; system degraded (error severity); tool errors. +- Producer rules: attention requested; session crashed; lease-expired reap; vault confirm pending; system degraded (error severity); tool errors. Each produced notification is also a versioned `NotificationMessage` (D-32) with a kind, category, severity, state and revision; the resolution of an attention request or a vault confirmation, the recovery of a degradation and a growing tool-error group are new revisions of the same notification, never new notifications. - Tool errors are **grouped per session**: one row per group (`" · N tool errors"`) grows while it is unread, has been idle for less than 5 minutes and is younger than 60 minutes; `notification.updated` carries the full row and clients upsert by id; lists sort by `updated_at`. Session-less caller mistakes (codes whose retry guidance is "different arguments") produce no notification. A failing agent would otherwise flood the inbox and toasts with one row per call, none naming the session. - `/me/preferences` stores the notification toast preferences (`notifications.toasts`, `notifications.types`), which follow the operator across devices; sidebar state and page size are per-device or per-URL. -- External channel adapters (webhook, ntfy, Telegram, Slack, email) are a documented seam (`NotificationChannel.send(payload)`); not built. +- External channels (Telegram, Discord, ntfy, a generic webhook, later more) implement the `NotificationChannel` port and receive the contract through the delivery outbox (D-34). The in-app inbox is itself a channel on that port, delivered inline. **Consequences.** - Read state survives reloads and is shared across tabs. - By default the dashboard does not toast tool errors; the grouped inbox row is the signal. +- A lifecycle revision changes the row's state fields but not its title, body or `updated_at`, so the inbox order does not move when a request is resolved. **Alternatives considered.** - *Client-only notifications derived from the feed.* Lost on reload, different in every tab, and unavailable to external channels. @@ -526,9 +527,9 @@ OS defaults: `~/Library/Application Support/BrowserHive` (macOS), `%LOCALAPPDATA **Context.** Several features are likely but not required for the first release. Building their seams now avoids reshaping public contracts later; building the features now would delay the product. -**Decision.** Not built: managed proxy pool and rotation; foreign fingerprint identities; profile blueprints (named, versioned, encrypted); local vault, TOTP, 1Password; extensions registry; external notification channels; security-intercept rule engine; resource governor and eviction; CAPTCHA detection and solving; Web Bot Auth; Tor egress; benchmark harness; multi-user, organisations and OIDC; Firefox and WebKit engines; standalone binaries. Each has a named seam in `01-overall-architecture.md` §9. +**Decision.** Not built: managed proxy pool and rotation; foreign fingerprint identities; profile blueprints (named, versioned, encrypted); local vault, TOTP, 1Password; extensions registry; security-intercept rule engine; resource governor and eviction; CAPTCHA detection and solving; Web Bot Auth; Tor egress; benchmark harness; multi-user, organisations and OIDC; Firefox and WebKit engines; standalone binaries. Each has a named seam in `01-overall-architecture.md` §9. -**Consequences.** Reserved enum values and config values for these features fail fast with a clear message rather than silently doing nothing. +**Consequences.** Reserved enum values and config values for these features fail fast with a clear message rather than silently doing nothing. External notification channels left this list in 0.2: they are built in stages from the contract and outbox up (D-32 to D-39). ## D-26 Browser choice: bundled by default, installed browsers by choice, never a silent switch @@ -629,3 +630,117 @@ OS defaults: `~/Library/Application Support/BrowserHive` (macOS), `%LOCALAPPDATA **Consequences.** Closed sessions keep showing whether they ran sandboxed and with which browser version, in the API, over WS and in the dashboard. Older readers still open the database (the migration is additive). History from before the upgrade honestly reads "not recorded". **Alternatives considered.** *Inferring the verdict for old rows from the channel and the host*: rejected; the policy's verdict depends on the executable, the host and the mode at the time, none of which is stored. *First launch wins*: rejected in favour of the latest launch, which is what the session actually ran with last. *A new `sandbox: 'sandboxed' | 'unsandboxed' | 'unknown'` wire field*: rejected; the existing optional `browser` already expresses "not recorded" by its absence, and a second field would duplicate it. + +## D-32 The notification contract is producer-owned, versioned and full-state + +**Status:** Accepted + +**Implementation:** the notification foundations (plan N0): contract, producers, JSON Schema, `degrade()`. + +**Context.** Notifications are about to leave the machine: Telegram, Discord, ntfy and a generic webhook first, more later (D-16). Each platform renders differently, some can edit a sent message and some cannot, and a user may build their own consumer from the webhook. If every adapter read domain events or the database, each would re-derive what happened, re-decide what may leave the machine, and break whenever the core changes. The in-app row (`Notification`) was designed for the inbox and carries neither severity, lifecycle, structure nor links. + +**Decision.** +- BrowserHive owns one message contract, `NotificationMessage` (`@browserhive/contracts/notifications`): zod-first, JSON-serialisable, `schema: 1`, with a JSON Schema published at `docs/reference/notification-message.schema.json` and regenerated in CI like the other references. It carries `id`, `revision`, `thread`, `kind`, `category`, `severity`, `state`, `alert`, `at`, `title`, `summary`, semantic `blocks` (a tiny inline AST, never a markdown string), up to 5 `actions` (`act` commands with an `open` fallback, or `open` dashboard paths), routing `entities` and the `privacy` already applied. +- **Producers own it; consumers only render it.** Producers are pure, table-driven functions from observed bus events (spec 03 §9). A platform adapter receives the contract and nothing else: it never reads domain events or the database. **Agents never author notifications**: every message derives from facts BrowserHive observed (D-09, D-12); there is no `notify` tool. +- **Full-state revisions.** A notification keeps its `id` for life; every state change is `revision + 1` and the message is complete at every revision, so a re-send or re-edit is always correct and adapters are idempotent. +- **Redaction happens before the contract** (spec 10 §9): every string a producer copies from an event goes through the `Redactor` (registered secrets and credential patterns) and URLs through `sanitizeUrl`. Content levels (`counts` < `titles` < `full`) are applied by the core per channel, never by an adapter. +- **Versioning.** Additive changes (a new optional field, a new kind, block, inline or command) keep `schema: 1`; consumers MUST ignore what they do not know (an unknown block renders as nothing, an unknown action is skipped). Removing or re-typing a field bumps `schema`, and the generic webhook announces the version it sends. +- A shared, pure `degrade(message, capabilities)` adapts a message to what a renderer supports (tables → lists, images dropped or linked, `act` → `open`, truncation with "… Open in BrowserHive"); renderers never implement fallbacks themselves. + +**Consequences.** Adding a platform is a renderer plus a transport against a fixed input, testable with golden files. The contract is a public compatibility surface: its JSON Schema is diffed in review. The in-app `Notification` DTO keeps its shape and gains the contract's classification fields (`kind`, `category`, `severity`, `state`, `revision`, `thread`) additively. Rows from before schema v5 have no stored message (`message_json` NULL): nothing is fabricated for them. + +**Alternatives considered.** *Adapters over the wire DTO* (the research's envelope): no structure, no lifecycle, and every adapter would derive links and severity itself. *A markdown string body*: every platform escapes markdown differently (Telegram MarkdownV2 reserves 18 characters) and a page title would break messages. *Letting agents send notifications*: a prompt-injected page would become a message from BrowserHive on the operator's phone; `request_attention` already is the agent's way to reach a human. + +## D-33 No hosted infrastructure; bring your own credentials; secrets only as environment variable names + +**Status:** Accepted + +**Implementation:** binding on every channel; the storage (`secret_refs_json`) exists since the notification foundations (N0), the checks arrive with the first channels (N1). + +**Context.** BrowserHive is local-first. A relay, a shared bot or a hosted callback would make BrowserHive a service with an operator, an uptime and a data-protection story, and would see every user's messages. Channel credentials (bot tokens, webhook URLs, access tokens) are live secrets; the database is backed up (`db backup`) and copied around. + +**Decision.** +- BrowserHive provides no server, relay or shared bot, now or later. Everything it does with a platform is an **outbound** connection from the user's machine (HTTP requests, Telegram long polling, the Discord gateway WebSocket, ntfy subscriptions); nothing needs a public URL or an open port. +- Every user brings their own Telegram bot, Discord webhook or bot and ntfy topic, and has full authority over it. +- **Secrets live only in the environment.** A channel stores the *names* of environment variables (`secret_refs_json`, `token=env:BH_TG_TOKEN` in a startup channel), never a value, so a database backup never contains a token. An inline secret is refused: over the API with a validation error, on the command line with exit 64. The dashboard shows whether a variable is set, never its value. Names starting with `BROWSERHIVE_` are refused, because the config loader reserves that prefix. + +**Consequences.** Changing a token is an environment change plus a restart. Features that seem to need an inbound connection are designed without one (act buttons through Telegram callbacks, the Discord gateway or a second ntfy topic, D-38). + +**Alternatives considered.** *A shared BrowserHive bot*: users would share a rate limit and trust a third party with every message. *Storing tokens encrypted in the database*: the key would have to live next to the database, so a backup would still carry both. + +## D-34 Notification delivery is a transactional outbox: at least once, full-state, silent edits, no degradation loop + +**Status:** Accepted + +**Implementation:** the notification foundations (N0): tables, worker, breaker, retention, metrics; platform adapters follow (N1). + +**Context.** An external platform can be down, rate-limited or misconfigured, and BrowserHive can stop at any moment. A delivery made from an in-memory callback is lost on restart, and the producer's in-memory de-duplication set does not survive one either. A channel failure reported as a system degradation would itself produce a notification, delivered through the failing channel: a feedback loop. + +**Decision.** +- A notification change and its delivery jobs (`notification_deliveries`, one row per channel × revision × op) are written in **one transaction**; a worker drains the jobs. `notification_deliveries` is both the outbox and the delivery log; `notification_channel_messages` maps a channel and notification to the platform message it produced, its last revision and its TTL deadline. +- **At least once.** Jobs left `sending` by a crash are retried at the next start. Edits and deletes are idempotent (full-state, D-32); a crash in the middle of a *new* send can duplicate that message, which is documented. +- **Coalescing.** A job renders the notification's current state; a job whose revision a later delivery already covered is `superseded`. At most one edit per message every 3 s. +- **Edits are silent.** Anything that must wake the user is a new message (`alert: true`). When a platform cannot edit, or the message was deleted in the chat, an alerting revision is sent as a new message and anything else is `superseded`. +- **Retries.** Exponential backoff with jitter, honouring `retry_after` / `Retry-After`; a job is `dead` after 8 attempts or 24 h. +- **Circuit breaker.** 5 consecutive failures mark the channel `broken`. That is shown in-app (a `channel.broken` notification and the channel's status) and **never** raised as a system degradation. The loop is cut structurally by kind: `channel.broken` notifications are delivered in-app only and are never enqueued for an external channel. +- **Backlog.** After an outage only the latest revision per notification is sent, and more than 20 pending `info` sends on one channel collapse into the newest one with a "you missed N" note. +- **Suppressed deliveries are logged** with a reason (`filtered`, `quiet_hours`, `throttled`, `channel_paused`, `content_blocked`, `image_blocked`, `edit_unsupported`, `delete_unsupported`, `collapsed`), so "why didn't I get it?" always has an answer. +- With no external channel configured nothing is enqueued, no worker timer runs and the only cost is one indexed read of `notification_channels` at startup. + +**Consequences.** Delivery rows are telemetry-class (30 days, spec 03 §7.1); channels are configuration and never pruned. `browserhive.notifications.deliveries{channel_kind,status}` counts outcomes and every platform call is a span (spec 10). A per-principal routing model is not built: channels are instance-wide and deliveries are enqueued once per produced notification, matching the single shared inbox (spec 03 §9). + +**Alternatives considered.** *Fire-and-forget calls from the producer* (the pre-0.2 seam): no retry, no log, lost on restart. *A retry counter to stop the degradation loop*: it bounds the loop instead of removing it, and a slow loop still spams the other channels. *Delivering the latest revision per thread* (the plan's first wording): a thread can hold several notifications (a session's crash and its wrap-up), so coalescing is per notification. + +## D-35 Message TTL is performed by BrowserHive + +**Status:** Accepted + +**Implementation:** `expires_at` and the sweeper exist since the notification foundations (N0); rules, platform deletes and the wizard arrive with the channels (N1). + +**Context.** Users want notifications that clean themselves up. No platform offers a per-message timer to bots; Telegram's auto-delete timer is a whole-chat setting chosen by the user, and a bot may delete its own messages only within 48 hours of sending them. + +**Decision.** BrowserHive deletes the message itself. A TTL can be set per channel and category; the default is **never** for every category. The deadline is stored as `expires_at` on the channel message when it is sent, so a deletion that fell due while BrowserHive was stopped happens at the next start and is logged as late. "Delete when resolved" is offered per channel and category and is **off** by default. The setup wizard caps Telegram TTLs at 47 h and suggests Telegram's own chat timer as a backstop; a message that became too old while BrowserHive was stopped is logged as `could_not_delete: too_old`. The UI says plainly that a lock-screen preview someone already saw cannot be taken back, and that ntfy.sh drops attachments after 3 h regardless. + +**Consequences.** TTL deletes are ordinary outbox jobs (`op = delete`) with the same retries and log. + +## D-36 Screenshots in notifications are opt-in; vault screenshots are taken before the fill + +**Status:** Accepted + +**Implementation:** with the first external channels (N1); the contract's `image` block and `privacy.hasImage` exist since N0. + +**Context.** A screenshot is the most useful and the most dangerous thing a notification can carry: a logged-in page, an inbox, a balance, or a credential being typed. + +**Decision.** Screenshots are off by default and switched on per channel and category, for three triggers only: an attention request (CAPTCHA requests are attention requests with a CAPTCHA reason; no new detector), a vault fill and a session crash (its last screenshot, when one exists). The vault screenshot is taken **before the fill sequence starts** (the login page and the origin being filled), never during or after a fill and never while a secret window is open. Images are never sent when `recordToolResults=none` or when the channel's content level is below `full`. An optional setting masks form fields (Playwright's screenshot `mask`). The wizard recommends a self-hosted ntfy for images or warns about ntfy.sh's public 3-hour attachment store. + +## D-37 `publicUrl` builds notification links and is trusted automatically + +**Status:** Accepted + +**Implementation:** with the first external channels (N1); the `LinkBuilder` port exists since N0. + +**Context.** "Open session" in a notification is a link, and a `127.0.0.1` link does nothing on a phone. Users who reach their dashboard remotely do it through their own reverse proxy, Cloudflare, Caddy, nginx or a Tailscale name, which today also needs `allowedHosts` and can fail the origin check when a proxy rewrites `Host`. + +**Decision.** A config key `publicUrl` (absolute `http(s)://`; a warning for `http` on a non-loopback host) holds the address where the user made the dashboard reachable. Every open link in a notification is `publicUrl + path`, built by one `LinkBuilder`; without `publicUrl` links use the local address and are labelled "Open on this computer". The `publicUrl` host is added to the Host allow-list and accepted by the origin (CSRF) guard. `doctor` and the System page check it by fetching `/health` and comparing this start's random instance id. Links never carry tokens; opening one still needs a login. BrowserHive provides no proxy or tunnel. + +## D-38 A Discord channel uses webhook mode or bot mode, one per channel + +**Status:** Accepted + +**Implementation:** webhook mode with the first external channels (N1), bot mode with act buttons (N2); `notification_channels.mode` exists since N0. + +**Context.** A Discord webhook takes 30 seconds to create and needs no connection, but its messages cannot carry interactive buttons. A bot needs a Developer Portal application and a gateway connection, and is the only way to receive button presses without a public endpoint (an Interactions Endpoint URL and the gateway are mutually exclusive). + +**Decision.** Each Discord channel uses exactly one mode: **webhook** (default: open links only; act buttons degrade to "Open in BrowserHive") or **bot** (opt-in: act buttons over the gateway WebSocket while BrowserHive runs). Switching mode is an edit of the channel and keeps its rules. The setup explains the difference with pictures of BrowserHive's own messages, never images copied from Discord or the web. + +## D-39 Startup channels come from command-line flags only and are read-only + +**Status:** Accepted + +**Implementation:** the registry merge and clash rule since the notification foundations (N0); the flag and its parser with the first external channels (N1). + +**Context.** Some users want a channel that exists from the first start (a server, a container) without clicking through the dashboard. Nested per-channel rules do not fit the flat config grammar (spec 08 §2), and a second copy of a channel in the config file and the database would be two truths. Process arguments are visible to other users of the machine (`ps`). + +**Decision.** Startup channels are declared only with a repeatable `--notificationChannel` flag (spec 08 §5.7): not in the config file and not in the environment. A startup channel references secrets by environment variable **name** (`token=env:BH_TG_TOKEN`); an inline secret is a usage error (exit 64). At each start the startup channels are projected into `notification_channels` with `source = 'startup'` (their configuration columns rewritten from the flags, their status and failure counters kept), so deliveries keep their foreign keys and the breaker state survives restarts; a startup channel no longer passed is removed with its delivery log. The dashboard and API show them read-only with a "from startup" badge. A startup channel whose name matches a dashboard channel stops startup with `CONFIG_INVALID` (exit 64); neither silently shadows the other. + +**Alternatives considered.** *A `notificationChannels` config key with a URI grammar* (the research's P1): secrets would sit in a file that gets committed, and per-channel rules would need a grammar the config ladder does not have. *Seeding the database from config on first run*: friendlier once, but the file would then lie about what is configured. diff --git a/specs/01-overall-architecture.md b/specs/01-overall-architecture.md index 53d5c38..3b96da4 100644 --- a/specs/01-overall-architecture.md +++ b/specs/01-overall-architecture.md @@ -163,7 +163,7 @@ PR → Changeset → merge to `main` → Changesets bot opens "Version Packages" | `VaultBackend` + `capabilities` | ports | local vault, 1Password, HTTP vault, TOTP | | `OperatorRequestBroker.kind` | domain | security intercept, approval gates, CAPTCHA takeover | | `AuthenticationProvider` chain, `Authorizer`, `tenant_id` | app/auth | better-auth, RBAC, orgs, OIDC/SAML | -| `NotificationChannel.send` | ports | webhook, ntfy, Telegram, Slack, email | +| `NotificationChannel` (`send`/`edit`/`delete`, capabilities) + `LinkBuilder`, fed by the notification outbox (D-34) | ports / app/notifications | Telegram, Discord, ntfy, webhook (next), Slack, Pushover, Teams, Apprise, email (later) | | `AdmissionPolicy` in the create pipeline | domain | resource governor, queueing, eviction | | `InterceptionChain` (blocklist route is the first handler) | domain | security rules, egress firewall, Web Bot Auth signing | | `ToolPack` registry with `requires` | interface/mcp | profile tools, proxy tools, captcha tools | diff --git a/specs/02-mcp-and-tools.md b/specs/02-mcp-and-tools.md index b76c89d..671310e 100644 --- a/specs/02-mcp-and-tools.md +++ b/specs/02-mcp-and-tools.md @@ -367,3 +367,4 @@ await server.stop({ deadlineMs?: 20_000 }); // idempotent; unwinds even after a - `launch_options.chromiumSandbox`: `false` is refused (`UNSAFE_LAUNCH_ARG`, spec 11 §4); `true` is accepted in every `sandbox` mode because it only strengthens the posture, and makes the sandbox a requirement for that session (a host that cannot give it answers `SANDBOX_UNAVAILABLE`, never `INTERNAL_ERROR`). - `list_saved_auths` scopes by the `owner` field in `.meta.json`; a manifest without it (for example one written by hand) is treated as owned by `local`. - `instructions` on the server and `title` on tools are additive; the golden generator normalizes key order so an SDK reordering does not produce a false diff. +- There is no tool that sends a notification, and none will be added to the core catalog: notifications derive only from facts BrowserHive observed (D-32). `request_attention` is the agent's way to reach a human, and it already produces the `attention.requested` notification (03 §9) that external channels deliver. diff --git a/specs/03-admin-backend.md b/specs/03-admin-backend.md index 0eb4920..bfebd8d 100644 --- a/specs/03-admin-backend.md +++ b/specs/03-admin-backend.md @@ -281,7 +281,7 @@ One broker (D-15) backs two resource views; paths stay recognizable. | Method | Path | Auth | Request | Response | |---|---|---|---|---| -| GET | `/notifications` | S | `read` (`all|unread|read`), `type[]`, `since`, `until`, `sort` (`updated_at` default \| `created_at`), cursor | `Page` + `unread_count`. `Notification` = `notification_id, principal_id, type, title, body, session_id, session_slug, target, source_event_id, created_at, updated_at, count, read_at, dismissed_at`; `session_slug` is `null` without a session; `count ≥ 1` is the number of folded occurrences; `updated_at` is the latest occurrence (= `created_at` when `count` is 1). `since`/`until` filter on the sort column and cursors are bound to it, so a growing group moves to the top. `unread_count` counts rows (a group of 12 errors counts 1) | +| GET | `/notifications` | S | `read` (`all|unread|read`), `type[]`, `since`, `until`, `sort` (`updated_at` default \| `created_at`), cursor | `Page` + `unread_count`. `Notification` = `notification_id, principal_id, type, title, body, session_id, session_slug, target, source_event_id, created_at, updated_at, count, read_at, dismissed_at, kind, category, severity, state, revision, thread` (the last six classify the row by the notification contract, §9; rows from before schema v5 read values derived from `type`); `session_slug` is `null` without a session; `count ≥ 1` is the number of folded occurrences; `updated_at` is the latest occurrence (= `created_at` when `count` is 1). `since`/`until` filter on the sort column and cursors are bound to it, so a growing group moves to the top. `unread_count` counts rows (a group of 12 errors counts 1) | | POST | `/notifications/{notification_id}/read` | S | — | `{ok:true}` | | POST | `/notifications/read-all` | S | — | `{ok:true, updated:n}` | | DELETE | `/notifications/{notification_id}` | S | — | `{ok:true}` (dismiss) | @@ -383,7 +383,7 @@ Ordering: `screencast.start`, `screencast.stop` and `screencast.set_size` from o | `pages` | `page.visited` `{row}` fleet-wide (scope `sessions:read`), so overview and websites views update live; also published on `session:` | | `blocklist` | `blocklist.hit` `{row}`, `blocklist.reloaded` `{patterns, skipped}` | | `system` | `system.degraded` `{event: SystemEvent}`, `system.recovered`, `tick`, `capacity` `{live, max}`, `retention.completed` | -| `notifications` | `notification.created` `{notification}` (first occurrence), `notification.updated` `{notification}` (the full row: a group grew — `count`, `title`, `body`, `updated_at`, `source_event_id` changed — or it was read/dismissed; clients upsert by `notification_id` and re-position by `updated_at`) | +| `notifications` | `notification.created` `{notification}` (first occurrence), `notification.updated` `{notification}` (the full row: a group grew — `count`, `title`, `body`, `updated_at`, `source_event_id`, `revision` changed — a lifecycle revision changed `state` and `revision` (`updated_at` unchanged), or it was read/dismissed; clients upsert by `notification_id` and re-position by `updated_at`) | | `logs` | `log.record` `{record}` (droppable, live only) | | `screencast:` | `meta`, `started`, `stopped`, `failed {code}` | @@ -402,7 +402,7 @@ Every event is produced by the app-level event bus (01 §5); the hub only maps b --- -## 7. Data model (schema v4) +## 7. Data model (schema v5) All tables in `browserhive.db` (D-24). Conventions: `TEXT` ids, epoch-ms `INTEGER` columns suffixed `_at`/`_ts`, durations `_ms`, sizes `_bytes`, booleans `INTEGER CHECK IN (0,1)`, enums `TEXT CHECK (col IN (...))` generated from `contracts/enums`, every `session_id` FK `ON DELETE CASCADE`. `WITHOUT ROWID` on tables with a TEXT primary key that are never scanned in insertion order. @@ -513,6 +513,16 @@ ALTER TABLE notifications ADD COLUMN count INTEGER NOT NULL DEFAULT 1 CHECK (cou ALTER TABLE notifications ADD COLUMN group_key TEXT; -- e.g. tool-errors: CREATE INDEX idx_notifications_updated ON notifications(principal_id, updated_at) WHERE dismissed_at IS NULL; CREATE INDEX idx_notifications_group ON notifications(group_key, updated_at) WHERE group_key IS NOT NULL AND read_at IS NULL AND dismissed_at IS NULL; +-- v5 (0005-notification-outbox): the notification contract (§9, D-32). Classification columns are nullable and have no CHECK on the open sets +-- (kind, category), so an older reader in the compatibility window can still insert rows; readers derive NULLs from `type` (§9). +ALTER TABLE notifications ADD COLUMN kind TEXT; -- NotificationKind (attention.requested, tool.errors, …); backfilled from type +ALTER TABLE notifications ADD COLUMN category TEXT; -- needs-you|problems|wrap-ups|reports|system; backfilled from kind +ALTER TABLE notifications ADD COLUMN severity TEXT CHECK (severity IN ('info','warn','error','critical')); +ALTER TABLE notifications ADD COLUMN state TEXT CHECK (state IN ('open','acted','resolved','expired','final')); -- backfilled from the request / degradation it points at +ALTER TABLE notifications ADD COLUMN revision INTEGER NOT NULL DEFAULT 1 CHECK (revision >= 1); +ALTER TABLE notifications ADD COLUMN thread TEXT; -- attention:, vault:, tool-errors:, session:, system: +ALTER TABLE notifications ADD COLUMN message_json TEXT; -- the current NotificationMessage; NULL for rows from before v5 (never fabricated) +CREATE INDEX idx_notifications_thread ON notifications(thread, created_at) WHERE thread IS NOT NULL; CREATE TABLE preferences (principal_id TEXT NOT NULL REFERENCES principals(principal_id) ON DELETE CASCADE, key TEXT NOT NULL, value_json TEXT NOT NULL, updated_at INTEGER NOT NULL, PRIMARY KEY (principal_id, key)) WITHOUT ROWID; -- operations @@ -520,14 +530,61 @@ CREATE TABLE system_events (seq INTEGER PRIMARY KEY, event_id TEXT NOT NULL UNIQ CREATE INDEX idx_system_events_open ON system_events(code) WHERE resolved_at IS NULL; CREATE TABLE artifact_outbox (outbox_id INTEGER PRIMARY KEY, kind TEXT NOT NULL CHECK (kind IN ('trace','screenshot','session_dir','backup')), path TEXT NOT NULL, session_id TEXT, enqueued_at INTEGER NOT NULL, attempts INTEGER NOT NULL DEFAULT 0, last_error TEXT); CREATE TABLE idempotency_keys (key TEXT PRIMARY KEY, principal_id TEXT NOT NULL, route TEXT NOT NULL, response_json TEXT NOT NULL, created_at INTEGER NOT NULL) WITHOUT ROWID; + +-- v5: notification channels and the delivery outbox (§9, D-33, D-34, D-35, D-39) +CREATE TABLE notification_channels ( + channel_id TEXT PRIMARY KEY, name TEXT NOT NULL UNIQUE, + kind TEXT NOT NULL, -- telegram|discord|ntfy|webhook|… (open set: NotificationChannelKind, no CHECK so a platform needs no rebuild) + mode TEXT, -- discord: webhook|bot (D-38); NULL elsewhere + source TEXT NOT NULL DEFAULT 'db' CHECK (source IN ('db','startup')), -- startup = projected from --notificationChannel at each start (D-39), read-only + status TEXT NOT NULL DEFAULT 'active' CHECK (status IN ('active','paused','broken')), + target_json TEXT NOT NULL DEFAULT '{}', -- non-secret coordinates (chat id, topic, server) + secret_refs_json TEXT NOT NULL DEFAULT '{}', -- environment variable NAMES only, never values (D-33) + rules_json TEXT NOT NULL DEFAULT '{}', -- NotificationChannelRules (categories, severity, sessions, quiet hours, content, TTL, …) + failure_count INTEGER NOT NULL DEFAULT 0 CHECK (failure_count >= 0), -- consecutive failures (breaker) + last_error TEXT, last_ok_at INTEGER, last_failure_at INTEGER, + created_at INTEGER NOT NULL, updated_at INTEGER NOT NULL +) WITHOUT ROWID; +CREATE TABLE notification_deliveries ( -- the outbox and the delivery log in one table + seq INTEGER PRIMARY KEY, + channel_id TEXT NOT NULL REFERENCES notification_channels(channel_id) ON DELETE CASCADE, + notification_id TEXT NOT NULL REFERENCES notifications(notification_id) ON DELETE CASCADE, + revision INTEGER NOT NULL CHECK (revision >= 1), -- the revision that caused the job + op TEXT NOT NULL CHECK (op IN ('send','edit','delete')), + status TEXT NOT NULL CHECK (status IN ('pending','sending','sent','retrying','dead','suppressed','superseded')), + reason TEXT, -- suppressed/superseded/dead reason (filtered, quiet_hours, collapsed, too_old, …) + attempts INTEGER NOT NULL DEFAULT 0 CHECK (attempts >= 0), + next_attempt_at INTEGER, last_error TEXT, duration_ms INTEGER, + message_ref_json TEXT, -- the platform message the job produced or addressed + created_at INTEGER NOT NULL, updated_at INTEGER NOT NULL +); +CREATE UNIQUE INDEX idx_notification_deliveries_idem ON notification_deliveries(channel_id, notification_id, revision, op); +CREATE INDEX idx_notification_deliveries_due ON notification_deliveries(next_attempt_at, seq) WHERE status IN ('pending','retrying'); +CREATE INDEX idx_notification_deliveries_sending ON notification_deliveries(updated_at) WHERE status = 'sending'; +CREATE INDEX idx_notification_deliveries_channel ON notification_deliveries(channel_id, seq); +CREATE INDEX idx_notification_deliveries_notification ON notification_deliveries(notification_id, seq); +CREATE INDEX idx_notification_deliveries_updated ON notification_deliveries(updated_at); +CREATE TABLE notification_channel_messages ( -- the platform message each notification became on each channel: edits and TTL + channel_id TEXT NOT NULL REFERENCES notification_channels(channel_id) ON DELETE CASCADE, + notification_id TEXT NOT NULL REFERENCES notifications(notification_id) ON DELETE CASCADE, + thread TEXT NOT NULL, -- the notification's thread (reply-to on platforms with replies) + message_ref_json TEXT NOT NULL, -- opaque platform refs (message id, chat id, sequence id) + last_revision INTEGER NOT NULL CHECK (last_revision >= 1), + sent_at INTEGER NOT NULL, updated_at INTEGER NOT NULL, + expires_at INTEGER, -- TTL deadline (D-35); NULL = never + deleted_at INTEGER, + PRIMARY KEY (channel_id, notification_id) +) WITHOUT ROWID; +CREATE INDEX idx_notification_channel_messages_expiry ON notification_channel_messages(expires_at) WHERE expires_at IS NOT NULL AND deleted_at IS NULL; +CREATE INDEX idx_notification_channel_messages_thread ON notification_channel_messages(channel_id, thread, sent_at); -- created in v1 ahead of use: logs (written only by the optional `--logPersist` durable sink), resource_samples (no writer yet; pruned by retention) --- reserved names, NOT created: proxies, profiles, security_rules, extensions, notification_channels — the migration of the feature that needs one creates it (and may pick another name) +-- reserved names, NOT created: proxies, profiles, security_rules, extensions — the migration of the feature that needs one creates it (and may pick another name) CREATE TABLE logs (seq INTEGER PRIMARY KEY, ts INTEGER NOT NULL, level TEXT NOT NULL, module TEXT NOT NULL, msg TEXT NOT NULL, trace_id TEXT, span_id TEXT, request_id TEXT, session_id TEXT, principal TEXT, fields_json TEXT); CREATE INDEX idx_logs_ts ON logs(ts); CREATE INDEX idx_logs_trace ON logs(trace_id) WHERE trace_id IS NOT NULL; CREATE INDEX idx_logs_session ON logs(session_id, ts) WHERE session_id IS NOT NULL; CREATE TABLE resource_samples (ts INTEGER NOT NULL, session_id TEXT REFERENCES sessions(session_id) ON DELETE CASCADE, cpu_pct REAL, rss_bytes INTEGER, host_free_bytes INTEGER, PRIMARY KEY (ts, session_id)) WITHOUT ROWID; ``` -`meta.min_reader_version = 1`. Migration v1 (`0001-initial`) creates everything above except the v2, v3 and v4 lines; migration v2 (`0002-notification-groups`, `compatible: true`, so the min reader stays 1) adds the notification grouping columns and indexes; migration v3 (`0003-harness-identity`, `compatible: true`) adds the identity columns and indexes, backfills `workspace` from `agent_name` and the two source columns where a value was set, and leaves `sessions.harness` NULL for existing sessions; migration v4 (`0004-session-browser`, `compatible: true`) adds `sessions.sandboxed` and `sessions.browser_version`, both NULL for existing sessions (no guess is backfilled: under `sandbox=auto` the answer depends on the browser and the host). The runner writes a backup before migrating. A dedicated CI test asserts fresh == migrated (D-04); fixtures `v1.db`, `v2.db`, `v3.db` and `v4.db` upgrade to head, and the golden is `schema-v4.json`. +`meta.min_reader_version = 1`. Migration v1 (`0001-initial`) creates everything above except the v2 to v5 lines; migration v2 (`0002-notification-groups`, `compatible: true`, so the min reader stays 1) adds the notification grouping columns and indexes; migration v3 (`0003-harness-identity`, `compatible: true`) adds the identity columns and indexes, backfills `workspace` from `agent_name` and the two source columns where a value was set, and leaves `sessions.harness` NULL for existing sessions; migration v4 (`0004-session-browser`, `compatible: true`) adds `sessions.sandboxed` and `sessions.browser_version`, both NULL for existing sessions (no guess is backfilled: under `sandbox=auto` the answer depends on the browser and the host); migration v5 (`0005-notification-outbox`, `compatible: true`) adds the notification contract columns and the three outbox tables. v5 backfills every existing notification from facts only: `kind` from `type` (`attention` → `attention.requested`, `vault` → `vault.confirm`, `lifecycle` → `session.reaped`, `system` → `system.degraded`, `error` titled "Session crashed" → `session.crashed`, any other `error` → `tool.errors`), `category` and `severity` from `kind` (§9), `state` from the operator request or degradation the row points at (`source_event_id`; `open` while pending/unresolved, `resolved`, `expired` on timeout, `final` when cancelled or gone), `open` for tool-error groups and `final` for the rest, `thread` from the ids the row carries, `revision` 1; `message_json` stays NULL. The runner writes a backup before migrating. A dedicated CI test asserts fresh == migrated (D-04); fixtures `v1.db` to `v5.db` upgrade to head, and the golden is `schema-v5.json`. ### 7.1 Retention classes @@ -538,7 +595,9 @@ CREATE TABLE resource_samples (ts INTEGER NOT NULL, session_id TEXT REFERENCES s | sessions | `sessions` rows | deleted only when all children are gone and `closed_at < now - retentionDays`; **archived sessions exempt** | | artifacts | `trace.zip`, `sessions//`, downloads | follow their session; deletion via `artifact_outbox` (row delete and outbox insert in one transaction; sweeper unlinks with retries; orphan scan weekly) | | connections | `mcp_connections` (with its IP, `User-Agent` and meta bag) | closed rows with `closed_at < now - retentionDays` that no remaining `sessions` row references (a session keeps its client metadata as long as it lives, archived ones included); open rows never. A pruned row's tool calls are older than it, so they are pruned first; a session's own harness survives in `sessions.harness` | -| notifications | `notifications` | 30 d after `dismissed_at`/`read_at`, 90 d otherwise | +| notifications | `notifications` | 30 d after `dismissed_at`/`read_at`, 90 d otherwise (a pruned row takes its deliveries and channel messages with it) | +| notification deliveries | `notification_deliveries` (terminal rows: `sent`, `dead`, `suppressed`, `superseded`), `notification_channel_messages` (deleted, or without a TTL) | telemetry-like, own window: 30 d after `updated_at`; never byte-pruned; `pending`, `sending` and `retrying` jobs and messages with a pending TTL are never pruned | +| configuration | `notification_channels`, `vault_bindings`, `vault_group_policies`, `preferences` | never pruned; `purge` lists them with the other tables (deleting the database loses configured channels and bindings) | | backups | `backups/*.db` | keep last 5 | `retentionDays` must be ≥ 1 (`0` is rejected at config time, see 08 — "keep forever" is expressed by the per-class exemptions below and a large value). The sweep runs every 6 h (`retentionIntervalMs`), catches per-item failures, records a `system.degraded` on repeated failure, never throws, and never runs `VACUUM`: the DB is opened with `auto_vacuum=INCREMENTAL` and the sweep issues `PRAGMA incremental_vacuum(N)` in bounded chunks. `/system.retention` exposes the last run. @@ -557,7 +616,10 @@ interface EventLogRepository { append(event); replay(afterSeq, limit); } interface PrincipalRepository / CredentialRepository / AuthSessionRepository / GrantRepository / AuthEventRepository interface VaultBindingRepository { list(query); get(handle); upsert(binding, ifVersion?); remove(handle); exportAll(); importAll(doc, mode); } interface VaultGroupPolicyRepository { list(); get(groupKey); upsert(policy, ifVersion?); } -interface NotificationRepository { …; list(query /* sort: updated_at|created_at */); findOpenGroup(principalId, groupKey); updateGroup(id, patch /* applies only while unread and undismissed */); } +interface NotificationRepository { …; list(query /* sort: updated_at|created_at */); findOpenGroup(principalId, groupKey); updateGroup(id, patch /* applies only while unread and undismissed; carries revision + message */); findLatestByThread(principalId, thread); revise(id, patch /* state, severity, revision, message; never title/body/updated_at */); } +interface NotificationChannelRepository { list(); get(id); getByName(name); upsert(row); remove(id); setStatus(id, status, at); recordSuccess(id, at); recordFailure(id, at, error): {failureCount}; } +interface NotificationDeliveryRepository { enqueue(rows) /* ignores duplicates of (channel, notification, revision, op) */; due(now, limit); claim(seq, at): boolean /* pending|retrying → sending */; finish(seq, patch); supersedeOlder(channelId, notificationId, revision, at); recoverSending(at): number; pendingByChannel(channelId, severity); list(query); get(seq); } +interface NotificationChannelMessageRepository { get(channelId, notificationId); upsert(row); firstInThread(channelId, thread); dueForDelete(now, limit) /* expired, not deleted, no delete job yet */; markDeleted(channelId, notificationId, at); } interface PreferenceRepository / SystemEventRepository / IdempotencyRepository / ArtifactOutboxRepository interface McpConnectionRepository { insert(row); update(id, patch); get(id); listOpen(); listRecent(limit) /* live first, with session counts */; closeAll(at); } interface UnitOfWork { transaction(fn: (repos: Repositories) => Promise): Promise; } @@ -581,25 +643,60 @@ Writes are enqueued (FIFO, one transaction per drain, statements prepared once); --- -## 9. Notifications (D-16) +## 9. Notifications (D-16, D-32, D-34) -Producer (`app/notifications`) subscribes to the bus and writes every row to **one shared operator inbox** (`principal_id` NULL): v1 has a single operator, the list, unread count and read/dismiss routes are not filtered by principal, and per-operator inboxes wait for multi-user (D-25). `NotificationService`'s `recipients` hook (default `[null]`) is the seam they plug into; composition does not set it. Rows: +Producer (`app/notifications`) subscribes to the bus and writes every row to **one shared operator inbox** (`principal_id` NULL): v1 has a single operator, the list, unread count and read/dismiss routes are not filtered by principal, and per-operator inboxes wait for multi-user (D-25). `NotificationService`'s `recipients` hook (default `[null]`) is the seam they plug into; composition does not set it. External channels are instance-wide: deliveries are enqueued once per produced notification, for the first recipient's row. -| Bus event | Notification | -|---|---| -| `attention.created` | type `attention`, title "Attention requested", body "{reason} · {mode} — agent blocked, lease frozen", target `/sessions/{id}?tab=live` (the dashboard redirects it to `?live=1`) | -| `session.closed` with reason `crash` | type `error`, "Session crashed", target `/sessions/{id}` | -| `tool.called` with `ok=false` | type `error`, **grouped per session** (below): title "{slug} · 1 tool error" / "{slug} · {n} tool errors", body "{tool} · {error_code} ({duration_ms} ms)" (or "{tool} · failed (…)"), `source_event_id` = latest failing call, target `/sessions/{id}?kinds=tool&errors_only=1` | -| `vault.confirm.created` | type `vault`, "Vault fill awaiting confirm", target `/vault?tab=confirm` | -| `session.closed` with `lease_expired` | type `lifecycle`, "Session reaped (lease expired)" | -| `system.degraded` (severity error) | type `system`, message | +### 9.1 Producer rules -**Tool-error grouping** (`app/notifications/producers.ts`): group key `tool-errors:`. A new failure grows the existing row of its group when that row is unread, not dismissed, its `updated_at` is < 5 min ago (`NOTIFICATION_GROUP_IDLE_MS`) and its `created_at` is < 60 min ago (`NOTIFICATION_GROUP_MAX_AGE_MS`): `count` +1, `title`, `body`, `updated_at` and `source_event_id` follow the latest occurrence, and `notification.updated` carries the full row. Otherwise a new row (`count` 1) is created with `notification.created`. Marking read or dismissing therefore starts a fresh group, and a failure run longer than an hour resurfaces hourly. Session-less failures: a caller mistake (an error code with `retryable: 'different_args'`, e.g. `INVALID_ARGUMENTS`, `SESSION_NOT_FOUND`) produces no notification; any other (e.g. `launch_session` → `BROWSER_NOT_INSTALLED`) is grouped under "No session · {n} tool errors" with `session_id`, `session_slug` and `target` null. External `NotificationChannel`s receive created rows only. Every row carries `session_slug` when it has a session. +| Bus event | Kind | Category · severity · state | In-app row | +|---|---|---|---| +| `attention.created` | `attention.requested` | needs-you · warn · open | type `attention`, title "Attention requested", body "{reason} · {mode} — agent blocked, lease frozen", target `/sessions/{id}?live=1` (`&takeover=1` for takeover) | +| `attention.resolved` | revision of the request's notification | state `resolved` (resolved, rejected), `expired` (timeout), `final` (cancelled) | row unchanged except the classification fields | +| `vault.confirm.created` | `vault.confirm` | needs-you · warn · open | type `vault`, "Vault fill awaiting confirm", target `/vault?tab=confirm` | +| `vault.confirm.resolved` | revision | state `resolved` (approved, denied), `expired`, `final` | as above | +| `session.closed` with a crash reason | `session.crashed` | problems · error · final | type `error`, "Session crashed", target `/sessions/{id}` | +| `session.closed` with `lease_expired` | `session.reaped` | problems · warn · final | type `lifecycle`, "Session reaped (lease expired)" | +| `tool.called` with `ok=false` | `tool.errors` | problems · warn · open | type `error`, **grouped per session** (below) | +| `system.degraded` (severity error) | `system.degraded` | system · error · open | type `system`, message, target `/system` | +| `system.recovered` | revision | state `resolved` | as above | +| `notification.channel.changed` to `broken` (internal) | `channel.broken` | system · error · final | type `system`, "Notification channel {name} is failing", target `/system`; **in-app only** (§9.4) | -Deliberately silent: `session.opened`, `page.visited`, `session.removed`, `attention.resolved`, `vault.confirm.resolved`. Rows are broadcast on the `notifications` topic; read/dismiss state is server-side and survives reloads. External channels are a `NotificationChannel` port (`send(payload)`) with no implementations. +Reserved kinds without a producer yet: `session.finished`, `vault.filled` (wrap-ups · info), `digest.daily` (reports · info), `report.anomaly` (reports · warn), `test` (system · info). The kind → category map is fixed in `contracts/notifications`; severity is set per producer. ---- +**Tool-error grouping** (`app/notifications/producers.ts`): group key `tool-errors:`. A new failure grows the existing row of its group when that row is unread, not dismissed, its `updated_at` is < 5 min ago (`NOTIFICATION_GROUP_IDLE_MS`) and its `created_at` is < 60 min ago (`NOTIFICATION_GROUP_MAX_AGE_MS`): `count` +1, `title`, `body`, `updated_at` and `source_event_id` follow the latest occurrence, the revision grows by one, and `notification.updated` carries the full row. Otherwise a new row (`count` 1) is created with `notification.created`. Marking read or dismissing therefore starts a fresh group, and a failure run longer than an hour resurfaces hourly. Session-less failures: a caller mistake (an error code with `retryable: 'different_args'`, e.g. `INVALID_ARGUMENTS`, `SESSION_NOT_FOUND`) produces no notification; any other (e.g. `launch_session` → `BROWSER_NOT_INSTALLED`) is grouped under "No session · {n} tool errors" with `session_id`, `session_slug` and `target` null. Every row carries `session_slug` when it has a session. + +Deliberately silent (no new notification): `session.opened`, `page.visited`, `session.removed`, a clean `session.closed`. `attention.resolved`, `vault.confirm.resolved` and `system.recovered` only revise an existing notification (found by its `thread`); rows from before schema v5 are revised in their classification fields only. Rows are broadcast on the `notifications` topic; read/dismiss state is server-side and survives reloads. + +### 9.2 The message contract + +Every produced or revised notification also stores its current `NotificationMessage` (`message_json`, `@browserhive/contracts/notifications`, JSON Schema in `docs/reference/notification-message.schema.json`, D-32): `schema: 1`, `id` (= `notification_id`), `revision`, `thread`, `kind`, `category`, `severity`, `state`, `alert` (whether this revision should make noise: true for the first revision, false for lifecycle revisions and group growth), `at {created, updated}`, `title` (≤ 120), `summary` (≤ 240), `blocks` (text, heading, fields, quote, list, table, image, code, divider, footer; inline text, bold, italic, code, dashboard-path link, time), `actions` (≤ 5: `act` with a `command {op, args}` and an `open` fallback, or `open` with a dashboard `path`; act ops `attention.resolve`, `vault.confirm.resolve`, `session.extend_lease`, `session.close`), `entities` (`sessionId`, `sessionSlug`, `harness`, `owner`, `tool`, `errorCode`, `domain`, `requestId`) and `privacy {level, hasImage}`. + +- Producers are pure (`buildMessage(draft, …)`); every copied string passes the `Redactor` and URLs pass `sanitizeUrl` before it becomes part of the message. The in-app title and body are the message's `title` and `summary` at creation; lifecycle revisions change the message only. +- Actions exist only while `state = open`; a revision to any other state carries none: the buttons disappear with a silent edit. +- Links are paths (`/sessions/{id}?live=1`); a `LinkBuilder` port turns them into absolute URLs for external channels (`publicUrl`, D-37). The in-app channel needs none. +- `restrictContent(message, level)` derives the lower content levels per channel: `titles` keeps title, summary, `fields` and `footer` blocks and the actions; `counts` keeps only a fixed per-kind title (with the group count), the session slug and the actions. +- `degrade(message, capabilities)` adapts a message to a renderer (D-32); both are pure and tested table-driven. + +Rows whose classification columns are NULL (written by an older reader in the compatibility window) are read with values derived from `type` exactly as migration v5 backfills them. + +### 9.3 Channels and the registry + +`NotificationChannel` (`ports/notification-channel.ts`) is the platform seam: `id`, `name`, `kind`, `capabilities` (rich blocks, tables, images, act buttons, open links, edit, delete, replies, delete window, max title/text length, max buttons), `send(delivery) → {ref}`, `edit(ref, delivery) → {ref}`, `delete(ref)`. A delivery is the restricted, degraded message plus the `LinkBuilder` and, where the platform supports replies, the ref of the first message of the thread. A platform failure is a `ChannelSendError` (`retryable`, `retryAfterMs`, `code`: `rate_limited`, `unavailable`, `timeout`, `auth`, `rejected`, `message_gone`, `too_old`). The in-app channel (`kind: in-app`) implements the same port and is delivered **inline after the commit**: the row is the delivery, so it has no outbox rows. + +`ChannelRegistry` (`app/notifications/channel-registry.ts`) holds the configured channels (`notification_channels`) and builds an adapter for each through factories registered per kind by composition (none are registered yet: without a factory a channel's jobs are suppressed with reason `no_adapter`). At start it projects the startup channels (`--notificationChannel`, 08 §5.7, D-39) into rows with `source = 'startup'`: configuration columns rewritten, status and failure counters kept, rows no longer declared removed; a name that a `source = 'db'` channel already uses stops startup with `CONFIG_INVALID` (exit 64). Channel rows store environment variable names only (D-33). + +### 9.4 The outbox (D-34) +- **Enqueue.** In the same transaction as the notification insert, growth or revision, `planDeliveries` writes one `notification_deliveries` row per external channel: `pending` with `op = send` for the first revision and `edit` for later ones, or `suppressed` with its reason when the channel's rules filter it (`channel_paused` for a paused or broken channel, `filtered` for category, minimum severity, session glob or harness rules, `quiet_hours` outside the channel's hours unless `critical`, `edit_unsupported` for a silent revision on a platform that cannot edit; an alerting revision there becomes a new `send`). `channel.broken` is never enqueued for an external channel: the degradation loop is cut by kind. With no external channel nothing is written. +- **Worker** (`app/notifications/outbox.ts`, injected clock, interval scheduler and jitter; like `RetentionScheduler`). It runs only while at least one external channel exists: a tick every second plus a kick after each enqueue. Each tick claims due jobs (`pending`/`retrying` with `next_attempt_at <= now`, oldest first) one at a time: `claim` moves a job to `sending` and counts the attempt; the adapter call runs outside any transaction; the result is written in one transaction (delivery row, channel message, channel counters). +- **Coalescing and supersede.** A job renders the notification's current message. When the channel message's `last_revision` already covers the job's revision the job is `superseded`; claiming a job supersedes older pending jobs of the same notification and channel. An edit is deferred (not an attempt) until 3 s after the message's last update. +- **Send, edit, delete.** `send` stores the ref in `notification_channel_messages` with `last_revision` and `expires_at` (from the channel's TTL for the category; never by default, D-35) and sets `expires_at = now` on a resolved notification when "delete when resolved" is on. `edit` addresses the stored ref; `message_gone` turns an alerting revision into a new `send` and marks the rest `superseded`. `delete` jobs are enqueued by the TTL sweep for expired, undeleted messages; a platform that cannot delete gives `suppressed: delete_unsupported`; `too_old` ends `dead` with reason `could_not_delete: too_old`; a delete more than a minute past its deadline is logged as late. +- **Retries.** Retryable failures back off exponentially (1 s · 2ⁿ, capped at 15 min) with ±20 % jitter, or wait the platform's `retry_after`; a job is `dead` after 8 attempts or when 24 h old. Non-retryable failures (`auth`, `rejected`) are `dead` at once. +- **Crash recovery.** At start every `sending` job returns to `retrying` due now. A `send` interrupted after the platform accepted it can therefore duplicate one message (at least once). +- **Breaker.** A success resets the channel's consecutive `failure_count` and sets `last_ok_at`; a failure increments it and records `last_error`. At 5 the channel becomes `broken`, its pending jobs are `suppressed: channel_paused`, `notification.channel.changed` is published and the `channel.broken` in-app notification is produced. No `system.degraded` is ever raised for a channel. Resuming a channel through the channel API sets it `active` and resets the count. +- **Backlog collapse.** When more than 20 `info` sends are pending for one channel, the newest is sent with a "you missed N" footer and the others are `superseded: collapsed`. +- **Telemetry.** `browserhive.notifications.deliveries{channel_kind,status}` counts each finished job; each platform call is a `notification.deliver` span (10 §6, §7). ## 10. Design notes - `POST /sessions/{id}/input` exists so takeover can be scripted without a WebSocket client; it shares the attention gate and audit path with the WS `input` command. diff --git a/specs/04-admin-frontend.md b/specs/04-admin-frontend.md index 9874afd..871e811 100644 --- a/specs/04-admin-frontend.md +++ b/specs/04-admin-frontend.md @@ -136,7 +136,7 @@ An injectable class (constructed in `SocketProvider`, replaceable in tests) expo Server-backed (D-16): `useQuery` for the page and the bell's unread count; the WS `notifications` topic patches both through the bridge (§5). Toast policy (`planToast`, pure and tested): -- Only `notification.created` raises a toast. `notification.updated` (a growing group, or read/dismiss elsewhere) updates a toast this tab raised, in place, or closes it once the row is read or dismissed; it never raises a new one. +- Only `notification.created` raises a toast. `notification.updated` (a growing group, a lifecycle revision, or read/dismiss elsewhere) updates a toast this tab raised, in place, or closes it once the row is read or dismissed or its `state` has left `open` (an attention request or vault confirmation resolved elsewhere); it never raises a new one. - Types that toast come from `/me/preferences` `notifications.types`; with none stored, `DEFAULT_TOAST_TYPES` applies, which excludes `error` (tool errors go to the bell only). `notifications.toasts === false` silences all. - An `error` for the session the operator is already viewing never toasts. The title is prefixed with `session_slug` unless it already contains it. - Only `attention` toasts persist; the rest auto-dismiss. "Open session" / "Review in vault" is omitted when the operator is already at the target (`isAlreadyAt`). @@ -203,7 +203,7 @@ Never: module-level mutable singletons holding server state. - Right: a search field-button ("Search sessions, pages…" + `Ctrl K`/`⌘K` by platform) that opens the palette (an icon button under 768 px), `HealthPill`, `NotificationBell`, `ThemeMenu`, `PrincipalMenu` (display name, Change password, Keyboard shortcuts, Log out). - `document.title` = `${title} · BrowserHive`. - `HealthPill` is a ghost `Button` opening a popover with Realtime (WS `connected | connecting | offline` + reason), REST, and daemon version. -- `NotificationBell`: badge anchored top-right, capped at "9+"; rows show the session slug (unless the title has it), `updated_at`, and "first …" for grouped rows; "View all" → `/notifications`. +- `NotificationBell`: badge anchored top-right, capped at "9+"; rows show the session slug (unless the title has it), `updated_at`, "first …" for grouped rows and the same outcome pill as the inbox (§12.11); "View all" → `/notifications`. - 404 (`NotFoundPage`) renders inside `AuthGate` and the shell and sets its title ("Page not found · BrowserHive"). ### 6.2 Route table as data @@ -562,7 +562,7 @@ Search: `read` (`all|unread|read`, default `all`), `type` (csv), `range` (`24h|7 - Header: accent "N unread" pill, description, Dismiss all (ghost, confirm), Mark all read. - `FilterBar`: labelled "Show" segmented control + "Period" range in the first row, type chips (no counts) in the second, "N matching" + Clear all. -- List: day groups ("Today", "Yesterday", `Mon D`) of `Panel`s with `LinkRow`s ordered and grouped by `updated_at`. **The whole row is a real link to its target and opening it (including middle/ctrl-click) marks it read.** Row: tinted type icon (type also in sr-only text), title (semibold + accent dot while unread), 2-line body, meta (session slug with a session icon unless the title leads with it; "first