Skip to content
Merged
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -29,6 +29,7 @@ Releases before 3.0.0 are not archived: those entries carry no prose to move.
- **PostgreSQL's server log is read by a classifier, not by two regexes** ([#3601]) - Darling now tails the same log the deadlock and plan readers tail — `pg_read_file` self-hosted, the RDS log API on Aurora/RDS — and classifies every line into families: `error` (WARNING or worse), `connection`, `lock_wait`, plus `temp_file`/`autovacuum`/`checkpoint` recognised for the structured tables #3602/#3603 add. Events land in `collect.pg_log_events` (V129, 30-day retention) with `message`/`detail`/`context` redacted by the plan parser's own patterns and the statement never stored, only fingerprinted; `get_pg_log_events` reads them with `family`/`min_severity` filters and the honest page contract. The three log readers now share one tailer (`PgServerLogTail`) instead of three copies.
- **The log said what the spill cost and what the vacuum cost; the store kept the sentence and threw away the number** ([#3602], [#3603]) - Two parser families on the log-event pipeline lift the figures out of the prose: a `temp_file` event carries the spilled file's exact bytes beside the fingerprint of the statement that spilled, and an `autovacuum` event carries the run's relation, duration, pages, tuples, buffers and WAL, on V130 columns of `pg_log_events` rather than sibling tables. `get_pg_log_events` publishes them where the line carried them; `get_pg_autovacuum_health` gains `recent_runs` per table, so "is it keeping up" and "what does each run cost" answer from one place. Line shapes pinned per PostgreSQL 16/17/18 from the server source.
- **Alert families route to their own channels** ([#3598]) - Every alert used to fan out to every configured channel from one install-time row, so self-monitor alerts, scheduled reports, job failures and performance pages interleaved in one stream and the reader did the routing in their head. A sparse `config_notification_routes` table (V131) layers over that row: a route names a family (`self-monitor`, `reports`, `agent-jobs`, `performance` — a closed taxonomy every alert maps into, census-pinned) or an exact metric, and carries a destination per channel; resolution runs once per firing after cooldown — exact name, then the firing a recovery pairs with, then family, then the parent's default — per channel, so an empty column inherits and a store with zero routes behaves byte-for-byte as before. The delivery ledger records which route matched, recoveries land where their firing did, the viewer gains a Routes grid, `get_notification_routes` / `set_notification_route` / `delete_notification_route` expose it to agents, and routes hot-reload through the same beacon mute rules use.
- **PostgreSQL targets enter the analysis pipeline** ([#3542] v1 plumbing) - The scheduled pass and every MCP analysis tool used to skip a PostgreSQL target behind an engine tombstone ("does not apply … use the get_pg_* reads"), because the pipeline read SQL Server tables and would have sat at "0 hours, still collecting" forever. The analysis service now holds two engine component sets and picks one PER CALL from the registry's `engine_kind` (two of its construction sites are singletons shared by every server, so the constructor could not be told); a PostgreSQL pass measures its 24-hour gate, its coverage witness and its baseline gate on `pg_database_stats`, scores a `pg_`-prefixed fact vocabulary declared once before any row persists, and clears the tombstone row on its first real pass. Every shared switch (scorer, advisory roots, graph, advice, anomaly reconciler, tool recommendations) gained exactly one delegating `pg_` arm so the nine content lanes build in parallel without touching shared files; `audit_config` answers `not_collected` for a PostgreSQL target instead of "the config collector may not have run yet". The SQL Server pass is byte-identical. No schema, no content. Rider (#3653/#3616): `HADR_SYNC_COMMIT` findings on Darling now carry next_tools (`get_wait_trend`, `get_ag_health`, `get_perfmon_trend`, `get_file_io_stats`), the pair of Lite's #3659.

### Fixed

Expand Down Expand Up @@ -99,7 +100,6 @@ Releases before 3.0.0 are not archived: those entries carry no prose to move.
- **The deprecated Dashboard mirrors five of the brains-review honesty fixes in its own idiom** ([#3653]) - The frozen Full-edition Dashboard (bug-fix support) said the same five untrue things the shipping SKUs fixed today: `get_blocking` / `get_deadlocks` / `get_alert_history` published a capped page count as a total (now `*_returned` + `truncated`, with the reader's fixed cap named as `limit_applied_by_reader`); the Poison Wait alert paged on one slow wait's average and slept through a THREADPOOL storm (now accumulated over the ten-minute window and graded on the shared `PoisonWaitEvaluator` bars, clearing only on an observed quiet window; the ms preference is retired but round-trips); `mute_analysis_finding` answered "muted" for any hash (now `registered` / `matched_now` / `muted_unmatched`); an online server with no collector banded was a green "OK" (now Unknown / `--`); and the Perfmon per-second rate was integer division (now fractional on the read). The anomaly detector's "calibrate on the lab box" markers now state what the fleet pass measured and that it never reached the Dashboard tier. The wait-fact nominal-window divisor stays declined per the issue.
- **Six small honesty riders from the brains-review residue** ([#3653]) - `poison_wait.threshold_ms` (both SKUs) now carries `threshold_ms_note` saying #3593 retired it, and Darling's `update_alert_settings` stores it with a `warnings[]` entry instead of pretending it tunes anything; `get_pg_server_config` publishes `captured_at` and joins the latest-snapshot census the `GetCurrent*Async` name had kept it out of; Lite's `get_latch_stats` / `get_spinlock_stats` bind their hidden `LIMIT 20` to the caller's `limit` and report `*_returned` / `truncated` (the grid keeps its 20 by name); a no-server census fails any reader on either SKU whose `ORDER BY` is an integer expression PostgreSQL/DuckDB would fold to a constant (zero hits today); V79/V122 rung docs name the PRs that superseded them; Lite's `analyze_server` recommends next reads for `HADR_SYNC_COMMIT`.
- **The PostgreSQL poison-wait host holds on silence like the SQL Server engine does, and three presence-flat alerts grade their severity from the bars the health bands already measured** ([#3653]) - The PG Poison Wait host announced "Cleared" the moment its read came back empty — collector silence read as recovery. Because `pg_wait_stats` skips idle events (#2694), the rows cannot witness their own absence, so the read now carries the collector's own `collection_log` run count over the same window and the host clears only on an observed window (the #3593 contract, ported). Deadlocks Detected grades Warning/Critical on the #3368 rate tiers (Darling's V120 knob honoured via `IAlertEngineSettings.DeadlockRateThresholds`; Lite the shipped pair), High CPU grades Critical at the CPU health band's 95% bar, tempdb Space fires an explicit Warning with no Critical tier (no measured bar exists to cite) — so history grids, `get_alert_history` and every channel render the tier the alert fired at instead of the colour its name implies.
- **PostgreSQL targets enter the analysis pipeline** ([#3542] v1 plumbing) - The scheduled pass and every MCP analysis tool used to skip a PostgreSQL target behind an engine tombstone ("does not apply … use the get_pg_* reads"), because the pipeline read SQL Server tables and would have sat at "0 hours, still collecting" forever. The analysis service now holds two engine component sets and picks one PER CALL from the registry's `engine_kind` (two of its construction sites are singletons shared by every server, so the constructor could not be told); a PostgreSQL pass measures its 24-hour gate, its coverage witness and its baseline gate on `pg_database_stats`, scores a `pg_`-prefixed fact vocabulary declared once before any row persists, and clears the tombstone row on its first real pass. Every shared switch (scorer, advisory roots, graph, advice, anomaly reconciler, tool recommendations) gained exactly one delegating `pg_` arm so the nine content lanes build in parallel without touching shared files; `audit_config` answers `not_collected` for a PostgreSQL target instead of "the config collector may not have run yet". The SQL Server pass is byte-identical. No schema, no content. Rider (#3653/#3616): `HADR_SYNC_COMMIT` findings on Darling now carry next_tools (`get_wait_trend`, `get_ag_health`, `get_perfmon_trend`, `get_file_io_stats`), the pair of Lite's #3659.
- **The Darling viewer's trend charts and calendar tell the same truth the MCP tools learned today** ([#3653]) - The Performance Trends tab read the raw tier only, so a 7-day chart on a TimescaleDB store plotted the 4 days raw still held under an axis that said 7; the query-duration, procedure-duration and execution-count charts now route by retention tier through the same ladder and hourly SQL the MCP trend tools use (moved to Storage as `DurationTrendRouting`, pinned equal on both sides) and each chart titles itself with the tier it served from and a truncation note when the store no longer holds the window's head. The viewer's trend SQL no longer fabricates a 0 for the first differenced point; an unrated point is skipped, not plotted or interpolated. The Performance Calendar judges every day against the store's retention horizon with the one `DailySummaryRetention` decision the MCP reader makes (`DailySummaryHorizon`, Storage; the purge's retention constants and fleet rule move down and the service aliases them), so a purged day is grey, not green, and its tooltip says why - on both SKUs. The PVS top-5 trend no longer plots an unmeasured size as 0 MB (Lite's `get_pvs_trend` carries it as null with `pvs_measured`).

## [3.8.0] - 2026-09-17
Expand Down
Loading