diff --git a/CHANGELOG.md b/CHANGELOG.md index ae6f03a23..14425eb4a 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -30,6 +30,40 @@ Releases before 3.0.0 are not archived: those entries carry no prose to move. - **The log said what the spill cost and what the vacuum cost; the store kept the sentence and threw away the number** ([#3602], [#3603]) - Two parser families on the log-event pipeline lift the figures out of the prose: a `temp_file` event carries the spilled file's exact bytes beside the fingerprint of the statement that spilled, and an `autovacuum` event carries the run's relation, duration, pages, tuples, buffers and WAL, on V130 columns of `pg_log_events` rather than sibling tables. `get_pg_log_events` publishes them where the line carried them; `get_pg_autovacuum_health` gains `recent_runs` per table, so "is it keeping up" and "what does each run cost" answer from one place. Line shapes pinned per PostgreSQL 16/17/18 from the server source. - **Alert families route to their own channels** ([#3598]) - Every alert used to fan out to every configured channel from one install-time row, so self-monitor alerts, scheduled reports, job failures and performance pages interleaved in one stream and the reader did the routing in their head. A sparse `config_notification_routes` table (V131) layers over that row: a route names a family (`self-monitor`, `reports`, `agent-jobs`, `performance` — a closed taxonomy every alert maps into, census-pinned) or an exact metric, and carries a destination per channel; resolution runs once per firing after cooldown — exact name, then the firing a recovery pairs with, then family, then the parent's default — per channel, so an empty column inherits and a store with zero routes behaves byte-for-byte as before. The delivery ledger records which route matched, recoveries land where their firing did, the viewer gains a Routes grid, `get_notification_routes` / `set_notification_route` / `delete_notification_route` expose it to agents, and routes hot-reload through the same beacon mute rules use. - **PostgreSQL targets enter the analysis pipeline** ([#3542] v1 plumbing) - The scheduled pass and every MCP analysis tool used to skip a PostgreSQL target behind an engine tombstone ("does not apply … use the get_pg_* reads"), because the pipeline read SQL Server tables and would have sat at "0 hours, still collecting" forever. The analysis service now holds two engine component sets and picks one PER CALL from the registry's `engine_kind` (two of its construction sites are singletons shared by every server, so the constructor could not be told); a PostgreSQL pass measures its 24-hour gate, its coverage witness and its baseline gate on `pg_database_stats`, scores a `pg_`-prefixed fact vocabulary declared once before any row persists, and clears the tombstone row on its first real pass. Every shared switch (scorer, advisory roots, graph, advice, anomaly reconciler, tool recommendations) gained exactly one delegating `pg_` arm so the nine content lanes build in parallel without touching shared files; `audit_config` answers `not_collected` for a PostgreSQL target instead of "the config collector may not have run yet". The SQL Server pass is byte-identical. No schema, no content. Rider (#3653/#3616): `HADR_SYNC_COMMIT` findings on Darling now carry next_tools (`get_wait_trend`, `get_ag_health`, `get_perfmon_trend`, `get_file_io_stats`), the pair of Lite's #3659. +- **A PostgreSQL target's durability posture is stated, not inferred** ([#3542] lane 8) - The analysis pass read nothing about `fsync`, `full_page_writes` or `synchronous_commit`, so a cluster running with fsync off analysed exactly like one running safe. Three `pg_posture` facts now come off the newest `pg_server_config` snapshot: fsync/full_page_writes off is a CRITICAL policy card naming the corruption trade, synchronous_commit off a 0.4 advisory naming the acknowledged-then-lost commit window, and Aurora's platform-managed pair is present, marked managed and ungraded. No amplifier, graph edge or optimisation word can reach a posture fact — pinned in source and at runtime (D6). +- **A PostgreSQL analysis pass now names the statement** ([#3542] lane 7) - A PostgreSQL target's stories ended at a server-level symptom and never said which statement. `PG_BAD_ACTOR_` facts from `pg_statement_stats` stored deltas now carry one shape's share of the window's total execution time (denominator over the whole window, never the page), its calls/sec over the span its deltas accrued and mean/max ms; the scorer grades the share behind an idle-window gate with every bar declared unmeasured (`threshold_lineage = 0`); the card states the figures, both levers with their counter-objectives and the queryid re-key caveat; the drill-down carries the normalised text, its hash and the per-database breakdown. +- **The two PostgreSQL knobs learn to speak** ([#3542] lane 2) - The analysis pass had nothing to say about `shared_buffers` and `max_wal_size`, the two settings the OtterTune field study found carry 75–95% of measured tuning value. Both are now cut from the latest `pg_server_config` snapshot: at their shipped defaults (128 MB, the line `initdb` writes; 1 GB) they root a 0.4 ADVISORY card and never an incident on their own, and reach the incident line only when the engine's own workload signal co-fires — `PG_CHECKPOINT_PRESSURE` (requested-dominant checkpoints from `pg_write_stats` deltas, with WAL bytes per interval) or `PG_BUFFER_CACHE_PRESSURE` (a composite of hit ratio, bgwriter and, on PostgreSQL 16+, `pg_io_stats` evictions, each arm carrying its own evidence and honesty flag: `evictions_tracked` below 16, `hit_ratio_suppressed` on Aurora where the storage layer answers reads). Every bar carries lineage — defaults engine-defined, ramps declared unmeasured with the calibrating read named — and the advice states the measured share, the value on the host, both levers and each lever's counter-objective. A `PgSettingValue` helper normalises PostgreSQL's unit zoo (8kB pages, kB/MB/GB, ms/s/min, -1 sentinels) once. +- **A PostgreSQL target's analysis pass gains its vacuum family, graded on the same bars the Tier-0 alerts page on** ([#3542]) - `PG_AUTOVACUUM_BACKLOG` names the worst table persistently past its own autovacuum trigger line with the ratio to that line, the dead-tuple slope over the run and whether autovacuum ran and lost (cumulative counts differenced, never summed); `PG_WRAPAROUND_TREND` grades XID and MultiXact separately through the alert evaluator's four arms and adds a time-to-wall estimate with an explicit not-computable branch; `PG_XMIN_HOLD` grades the alert's identity arm and names the fix per holder kind. The wraparound/xmin/slot constants move to one shared `PostgresOutagePredictorThresholds` that both the alert evaluator and the analysis scorer reference, pinned so the surface that pages and the surface that narrates cannot disagree; the three facts chain into one story with value-stated advice that never suggests disabling autovacuum. +- **A PostgreSQL pool near its connection ceiling is graded as the outage it is** ([#3542]) - A PostgreSQL target's analysis pass had no session fact: a pool at 96 of 97 usable slots — one deploy from `FATAL: too many clients already` — produced an empty verdict at full coverage. `PG_CONNECTION_SATURATION` now grades the window's peak session count from `pg_session_states` against the engine's own line, `max_connections − superuser_reserved_connections` (read off the config facts, never re-read), scoring zero below 0.8 and 1.0 at 0.9 with `threshold_lineage = 0`; the card states the division with all three numbers, the peak capture's active / idle-in-transaction / other breakdown, and both levers — a pooler or a higher ceiling — with what each costs. When the monitoring login lacks `pg_read_all_stats` and most stored rows are redacted, `PG_MONITORING_PERMISSIONS` says so instead of a ratio built on blank rows. +- **A PostgreSQL analysis pass now names temp-file spill and gates work_mem on it** ([#3542]) - PG_TEMP_SPILL grades spilled bytes per second of observed time from the reset-aware pg_database_stats difference (a pg_stat_reset() mid-window is clamped and reported), names the database that spilled the most, and leads a story to CONFIG_PG_WORK_MEM, which scores 0 without that evidence and reaches the incident line only beside it; the drill-down lists the statements that wrote the most temp blocks and the advice states the host's own work_mem × max_connections arithmetic. The same read emits PG_TPS (with a half-window trend) and PG_DEADLOCK_RATE from the engine's counter on the alert band's measured tiers, saying how many the log captured beside how many the engine counted. +- **A PostgreSQL analysis pass reads the wait profile — measured on Aurora, estimated on stock, and says which** ([#3542]) - Lock, LWLock, IO and IPC rollups with the Lock:relation, LWLock:WALWrite, IO:DataFileRead and IO:WALSync standouts, each a fraction of the wait source's own observed time. Aurora's pg_wait_stats deltas are summed under the three-state interval (a restart collection is not a sample); a stock target's pg_wait_sampling counts × profile_period are read only when the Aurora table is empty and every such fact carries is_sampled and its resolution, with advice that says "estimated from sampling" and never "spent". Bars are unmeasured and say so; a rollup yields to its fired standout so one wait is graded once; WAL waits chain into checkpoint pressure, data-file reads into buffer-cache pressure, and Lock waits mesh with connection saturation. Lock advice names PostgreSQL's levers, not the SQL Server isolation remedy. +- **A PostgreSQL target learns its own normal** ([#3542] lane 9) - A PostgreSQL analysis pass had no baseline and no anomaly detector, so a 6× throughput surge, a tripled connection pool or a serverless instance pinned at its configured ceiling was indistinguishable from a quiet Tuesday - and had the anomalies existed, they would have scored 0 and never left the 1.49 cap (#3584's defect, again). Five baselines over the pg_* raw hypertables (TPS and deadlocks/hour off the reset-aware per-database difference, the capture's total_sessions, the Aurora all-types wait rate with CPU excluded, percent of the configured ACU ceiling - never raw cpu_percent) and five detectors through the shared AnomalyGate writing the metadata the shared ramp and extremity escape read; a PostgreSQL ratio ramp with no sentinel for a first occurrence; load-family co-fires confirmed only by a measured capacity reading; PG_CPU_PERCENT graded only when capacity was measured; advice in PostgreSQL nouns. Every unmeasured bar carries its lineage and threshold_lineage = 0. Stock PostgreSQL's sampled waits get no anomaly baseline in v1. +- **The measurement contract had ten rules and two tests** ([#3653]) - The contract #3540 stated in one sentence is now a numbered list in `Lite.Tests/MeasurementContractCensusTests` (4 and 7 keep their tests' numbers; 9–10 are left empty because the review's enumeration is not on the record), and rules 1, 2, 3 and 6 are censuses over both SKUs' sources with positive controls and population floors: no reader divides by a stored interval except through NULLIF, no rate assumes a cadence, exactly two carriers observe an epoch and every forget site is named, every per-second alias is a quotient. Rule 5 is rostered against the real continuous-aggregate definitions — the #3698 successors carry the interval predicate; the superseded pair and three hourly rollups do not — and rule 8 waits, pinned, on the `cntr_type` rung. +- **v2 plumbing for the PostgreSQL-target engine, and the wait-profile anomaly now folds onto its wait and can page** ([#3691]) - The I/O, replication and bloat families exist as stubs behind one delegating arm per shared switch, with the v2 vocabulary declared once so lanes 11–15 never touch a shared file. `ANOMALY_PG_WAIT_PROFILE` stopped sitting one hundredth under the page line: it folds onto the wait card its dominant contributor names and leaves the 1.49 cap on the same three readings as its SQL Server twin, anchored on its own firing multiple. +- **Schema 133: the PostgreSQL saturation numerator and the sampler's duty cycle are stored** ([#3691]) - `pg_database_stats` gains `numbackends` (the client backends connected per database at the instant of the read, a level, never differenced) and `pg_wait_sampling` gains `sampled_ms` (the milliseconds each five-minute collection actually observed - 30,000 from the service-side sampler, NULL from the extension arm that watches the whole interval), because "count / max_connections" had only an exception-capture count to divide and "samples × period / interval" understated the sampler arm ~10× with nothing in the row saying so. Written by the collectors from this rung; read by nothing yet. Lite stores no `pg_*` table, so no DuckDB bump. +- **The pull-request review learns the repository's traps and stops answering LGTM by default** ([#3710]) - The bot review had read several hundred pull requests and posted LGTM on very nearly all of them; every defect it missed was later found by behaviour, not by reading. Three structural causes, three fixes: a `.github/REVIEW-TRAPS.md` of the repository's recurring defect shapes, each citing the pull request that bit, which the review must consult and name; a blind first phase that reads the diff and its callers before it is allowed the author's account, then reports where the two diverge; and a structured verdict — every claim in the body VERIFIED, UNVERIFIED or REFUTED, the three riskiest lines and what pins each, a machine-readable ledger line — without which the run fails, so a bare LGTM is no longer a valid review. A paths filter sends analysis, collectors, storage, service, alerting and notification changes to a full-tier Opus pass with the 1M context window and everything else to a lighter Sonnet pass; the verdict contract is identical either way. Lands dark under the workflow-drift rule; the first live run is the first pull request after the release sync. +- **A server configuration change now has a consequence the engine states — `CONFIG_CHANGED` compares the ±4 h around it and says what moved, or that nothing did** ([#3653]) - The analysis engine graded the current value of seven `sp_configure` settings and never asked whether a change had an effect; that join lived in the operator's head between `get_server_config_changes` and `compare_analysis`. Both SKUs now emit one Information-level `CONFIG_CHANGED` finding per change event observed inside the pass window, running `ComparePeriodsAsync` over the four hours before the observation and the (honestly clamped) four hours after, banded by the same dispersion rules as `compare_analysis`, with the setting, old → new, and the moved metrics — or the sentence that none moved — frozen into the finding. Because `server_config` is captured on connect, the finding says "first observed", states the span the real change landed in, and never claims to be a causal test; database-config and trace-flag changes are the next slice. +- **PostgreSQL deadlock cards carry the captured exemplars, grouped by shape and ranked by recurrence** ([#3691]) - The card said "the engine counted N; M were captured from the log" and stopped. The drill-down now reads pg_deadlocks for the window, groups reports into shapes (participants, lock modes, resources — deadlock_hash identifies one report, not a shape), carries the top three with bounded statement and graph text, reuses the counter the card was graded on, and re-freezes the advice with the value-stated sentence and the ordering lever. Zero capture beside a non-zero counter names the logging-posture settings instead of showing an empty list. +- **The PostgreSQL bloat family grades fourteen-day GROWTH behind a size floor, never a spot percentage** ([#3691]) - `bloat_pct_estimate` read 71–99.8 % on most of the measured fleet's small heaps while their fourteen-day growth was zero bytes, so a percentage-graded finding would have paged everywhere and argued for VACUUM FULL on a Tuesday. `PG_BLOAT_TREND` (hourly `pg_table_bloat_stats`) and `PG_INDEX_BLOAT_TREND` (daily `pg_index_bloat`) now grade the estimate's growth across a fourteen-day lookback — 256 MiB and 25 % of the earlier value, 1 GiB critical, 64 MiB floor, all measured on the dogfood fleet — exclude `estimate_unavailable` and skipped rows outright, name the top three objects, state their own sample counts, and chain onto a same-table `PG_AUTOVACUUM_BACKLOG` as its damage. Withheld estimates are facts that say why, never zero bloat. +- **A chain that fired every Tuesday at the same hour was rated a fresh incident each week and nobody was told; the pass now labels it "recurring at this hour" at unchanged severity, and a weekly Agent job whose slot slid gets "maintenance window moved"** ([#3653]) - Ruling Q3: label, not discount. After the anomaly fold and before the finding is persisted, one store read of the prior three weeks (on the target's clock via `server_properties.utc_offset_minutes`, UTC disclosed when absent) lets `RecurrenceLabeler` append one sentence to the frozen advice of every chain that fired in the same hour×weekday slot for three consecutive weeks, and one naming the job and both slots to every story tied to a fired `RUNNING_JOBS` whose slot moved since last week. Severity is bit-identical with and without the label; both `get_analysis_findings` twins publish `recurring_at_this_hour`, `recurrence_weeks` and `maintenance_window_moved`. +- **PostgreSQL read latency is now a graded fact, and the pass says why when it cannot know** ([#3691]) - `PG_IO_READ_LATENCY_MS` states milliseconds per data-file read from `pg_stat_io`, graded against bars measured on the dogfood fleet's Aurora storage (10 / 30 ms, population named), with `ANOMALY_PG_IO_LATENCY` when an hour departs this server's own routine. With `track_io_timing` off the pass reports the reads it counted and that none were timed — never 0.000 ms — and on a pre-16 major it says `pg_stat_io` is absent; write latency is stated, not graded. +- **Replication lag and slot retention stories for the PostgreSQL-target engine** ([#3691]) - A PostgreSQL operator's agent now hears whether a standby is behind or FALLING behind and at which stage (sent / write / flush / replay), whether a replication slot is filling the disk or has already lost the WAL its consumer needs, and whether a slot's horizon is what holds vacuum back — three facts, one chain into PG_XMIN_HOLD, and the replay-lag anomaly against the server's own hour-of-week baseline. Every shared bar is the Tier-0 slot alert's own by reference; bytes are graded and NULL replay_lag_ms is tolerated; the drift multiple is unmeasured (n = 1). +- **The CPU sample's UTC instant and the server's time-zone id are now stored, so UTC-window readers stop deriving an offset that is an hour wrong across DST** ([#3653]) - `cpu_utilization_stats.sample_time` is the monitored server's LOCAL wall clock, and every reader that aligned it against UTC derived the offset per batch or from the single collected `utc_offset_minutes`, placing the samples nearest a DST transition an hour wrong. Darling V134 / Lite v63 add `sample_time_utc` beside the unchanged local stamp (the same instant off `SYSUTCDATETIME()`), which the viewer, MCP `get_cpu_utilization` and the Lite CPU window now prefer, and `server_properties.time_zone_id` (`CURRENT_TIMEZONE_ID()`, SQL Server 2022+ / Azure SQL; NULL elsewhere) beside the offset that could not say which side of a transition an instant fell on. Nullable, no backfill: the offset a server had at a past sample's instant is exactly what the store never recorded. +- **PostgreSQL sessions: an operator now learns which application left a transaction open and walked away** ([#3691]) - The analysis pass said nothing about idle-in-transaction sessions beyond a share of the pool; `PG_IDLE_IN_TRANSACTION` now names the longest holder (application, role, database) from `pg_session_states`, graded on duration bars the fleet calibration measured (60 s / 10 min — zero rows reached 60 s over 7 days × 50 clusters), a band higher when it pins the xmin horizon (never on the collector's "pins nothing" sentinel), amplified when the same holder recurs, joined to the saturation story only when parked sessions are a quarter of the pool, and honest when the monitoring login cannot see state (`idle_in_transaction_unobservable = 1`). The lever is the application; both server-side backstops state that they roll the work back. +- **The Long-Running Query opt-out knob was read-only on Darling — V135 gives it a store home with the production read's seeds as the column default, editable in the Viewer and through `update_alert_settings` on both SKUs** ([#3653]) - Two `text[]` columns on `config_alert_settings` default to the seeded lists, so a pre-rung store evaluates exactly what it did before and a list an operator clears stays cleared; `get_alert_settings` publishes `long_running_query.excluded_program_name_prefixes` / `excluded_logins` on Lite and Darling alike, and the viewer's connect-time probe maps a current store to V135. +- **PostgreSQL blocking: the pass names the head of the chain, how long the sessions behind it really waited, and whether the log agrees** ([#3691]) - `PG_BLOCKING_CHAIN` reconstructs each capture's chains from the sampled `pg_blocking_edges` (a head is a blocker not itself blocked; cycles are counted, never attributed), names the head with the most sessions behind it and grades on the blocked side's own duration — never captures × cadence — routing to the idle-in-transaction holder or the long-running query behind it. `PG_LOCK_WAIT_EVENTS` reads the event-grain `log_lock_waits` lines (count, longest, relation) and says `log_lock_waits` is off rather than reporting zero; `PG_LONG_RUNNING_QUERY` grades client backends' own query duration; `ANOMALY_PG_BLOCKING` baselines blocked sessions per minute with `collection_log` supplying the "looked and found nothing" zeros a healthy instance never writes to the edges table, gated on peak AND mean so one mid-flight handoff cannot fire it. Every bar unmeasured and says so; `deadlock_timeout` is the one engine-defined line, stated as context. + +### Changed + +- **The review guard tells a pending workflow fix from hostile drift** - When `dev`'s `claude-review.yml` differs from `main`'s because a reviewed dev PR changed it and the release has not synced yet, the guard now warns "review pending release" and names the merged PR instead of failing every PR in the repo as if the file had been tampered with; unexplained drift still fails, and an origin the guard cannot determine fails closed. +- **The tempdb Space alert stops paging on a single collected sample** ([#3653]) - it fired on the first `tempdb_stats` row at or above the threshold and resolved on the next one under it, which on one production store class was 22 page/resolve pairs in 14 days about runs of 1-4 samples lasting at most 206 s. The arm now sits behind the shared persistence gate High CPU adopted in #3282: 3 consecutive breaching COLLECTIONS to fire (identified by the row's `collection_time`, so a 30 s sweep re-reading the last row does not count twice), the first collected sample under the bar to resolve, the open-incident bit persisted under its own metric name in the existing state table so a restart neither re-announces nor forgets it, and a tempdb reading that stops arriving freezes the gate instead of announcing "back to N/A". Blocking Wait Time and Long-Running Query await their rulings. +- **The analysis names the Agent job that is running long instead of counting it** ([#3653]) - the RUNNING_JOBS fact on both SKUs now carries the name of the job furthest past its own history (chosen among the running-long rows by percent of average, then duration), and the job card's headline, first sentence and remediation say which job — and "and N others" when several overran — where they used to send the operator to the Running Jobs view to find out. Findings persisted before this, and windows where jobs ran but none ran long, read exactly as before. +- **MCP tool failures were two wire shapes — a bare sentence on 214 tools and a JSON envelope on the PostgreSQL reads — and a `status`-keyed client read the sentences as successful text; now every tool on both SKUs answers a caught exception with one envelope through `McpHelpers.FormatError`** ([#3653]) - **WIRE CHANGE.** A tool that throws returns `{"status":"error","message":"Error during : …","hints":{"operation":""}}` — the same envelope and serializer the four miss words use, the message text unchanged — including the frozen Dashboard twin, which links the same helper. The 30 PostgreSQL catches drop their "Reading X failed" dialect for the one grammar; the web surface maps the envelope to HTTP 500 (so PostgreSQL failures that passed through as 200 are 500 now) and keeps its `{"error": sentence}` body; the payload-contract census retires its two-shape inventory and fails by name on any catch that bypasses the helper. Validation refusals are unchanged and still bare sentences — a separate ruling. +- **PostgreSQL bars the fleet measured now say so** ([#3691]) - Every v1 bar the 2026-09-19 calibration of 50 Aurora PostgreSQL clusters validated (temp 1/10 MiB/s, backlog 3/10×/5, deadlocks 5/20, bad-actor busy floor 0.05, wait rollups 0.20/1.0 and io 0.40/2.0, CPU 80/95, TPS floor 50 / fallback 500, CPU floor 40 / fallback 95, deadlock floor 1/h, wait fallback 500 ms/s) carries a measured comment naming its percentile, population and date, and the facts graded on them show `threshold_lineage = 1` in get_analysis_facts instead of calling the bar a judgment; sessions 0.8/0.9 stay engine-defined with the measured note (fleet max 10.3 % of ceiling). Still unmeasured and still saying so: wait standout bars and sampling floor, session COUNT floors, the ratio families' firing multiple, every co-fire boost, checkpoint pressure (not applicable on Aurora) and buffer cache. No bar value changed. +- **Anomaly detectors fired on one hot sample and fired more the longer the window was; the gate now judges the window peak AND the window mean, and I/O reads the pair like its siblings** ([#3653]) - Every z-score detector tested the window MAX against a per-sample hour×dow distribution, so a 24-hour anchored pass mechanically reported more anomalies than the 4-hour scheduled pass on identical behaviour. The shared AnomalyGate fires only when both the peak and the window mean clear the existing cutoffs (z on both when the baseline is trustworthy; the fallback bar on the peak and the magnitude floor on the mean when it is not) — no cutoff moves, the reported sigma stays the peak's, and the mean's sigma rides beside it in the story. I/O latency, the one family reading AVG alone, now reads MAX and AVG on both SKUs and reports the peak. +- **Confidence chooses the channel** ([#3712]) - a scheduled-analysis finding now earns a page by corroboration, never by severity alone: two or more facts in its chain, a matched co-fire check on its root, or one of the two by-construction stories. A lone uncorroborated fact at or above the notify floor is persisted, visible on the web and MCP surfaces the instant it fires, recorded in the alert history as `notification_type: digest` with the gate's reason (`routing` / `routing_reason` on `get_alert_history`), marked *Not paged* in Lite's Recommendations, and named once a day in Darling's new Analysis Singles Digest - but delivered to no paging channel; the day it gains corroboration it pages as a new firing. Measured on the first day the recalibrated engine ran on a large production fleet: forty-plus anomaly pages in seven hours at confidence 0.20-0.35, mostly true and redundant. One knob, `analysis.uncorroborated_route` (`digest` default, `page` restores the old behaviour): Lite's Settings → Alerts, Darling's file-level `analysis.uncorroboratedRoute`. +- **Blocking Wait Time paged on one snapshot and cleared on the next, and Long-Running Query had no way to stop reporting the same permanent background sessions forever; blocking now fires on one snapshot at 3× the bar or on 3 consecutive collections through the shared persistence gate, and long-running queries gain an opt-out knob seeded from the production read** ([#3653]) - Measured on one production store class, 97 of 102 Blocking Wait Time episodes were a single snapshot, so a plain consecutive gate would have dropped exactly the severe pile-ups; the fire names which arm admitted it (`single_snapshot_3x` / `consecutive_k3`), a quiet collector cycle starts a new episode, and a stale snapshot still clears. The Long-Running Query knob is two editable lists — program-name PREFIXES seeded with `SQLAgent - TSQL JobStep` and exact logins seeded with the two `NT AUTHORITY` service accounts, the classes a 7-day read of one large production store showed to be permanent background; the application's admin login is deliberately not excluded and named humans never are — applied inside the read ahead of the row cap, with the card counting the sessions each list removed. Darling's store column and both SKUs' `get_alert_settings` twins follow as rung V135. +- **PostgreSQL saturation now tells queueing from load** ([#3691]) - the connection-saturation finding is amplified when the session-count anomaly fired against the server's own hour-of-week baseline while the transaction-rate anomaly did not — more connections than this hour usually carries, doing no more work — and says so in the advice with the measured multiple. The parked raw TPS-trend co-fire is retired: the fleet calibration showed routine 20–50× in-window TPS bursts on every cluster, so that trend was noise. ### Fixed @@ -101,6 +135,38 @@ Releases before 3.0.0 are not archived: those entries carry no prose to move. - **Six small honesty riders from the brains-review residue** ([#3653]) - `poison_wait.threshold_ms` (both SKUs) now carries `threshold_ms_note` saying #3593 retired it, and Darling's `update_alert_settings` stores it with a `warnings[]` entry instead of pretending it tunes anything; `get_pg_server_config` publishes `captured_at` and joins the latest-snapshot census the `GetCurrent*Async` name had kept it out of; Lite's `get_latch_stats` / `get_spinlock_stats` bind their hidden `LIMIT 20` to the caller's `limit` and report `*_returned` / `truncated` (the grid keeps its 20 by name); a no-server census fails any reader on either SKU whose `ORDER BY` is an integer expression PostgreSQL/DuckDB would fold to a constant (zero hits today); V79/V122 rung docs name the PRs that superseded them; Lite's `analyze_server` recommends next reads for `HADR_SYNC_COMMIT`. - **The PostgreSQL poison-wait host holds on silence like the SQL Server engine does, and three presence-flat alerts grade their severity from the bars the health bands already measured** ([#3653]) - The PG Poison Wait host announced "Cleared" the moment its read came back empty — collector silence read as recovery. Because `pg_wait_stats` skips idle events (#2694), the rows cannot witness their own absence, so the read now carries the collector's own `collection_log` run count over the same window and the host clears only on an observed window (the #3593 contract, ported). Deadlocks Detected grades Warning/Critical on the #3368 rate tiers (Darling's V120 knob honoured via `IAlertEngineSettings.DeadlockRateThresholds`; Lite the shipped pair), High CPU grades Critical at the CPU health band's 95% bar, tempdb Space fires an explicit Warning with no Critical tier (no measured bar exists to cite) — so history grids, `get_alert_history` and every channel render the tier the alert fired at instead of the colour its name implies. - **The Darling viewer's trend charts and calendar tell the same truth the MCP tools learned today** ([#3653]) - The Performance Trends tab read the raw tier only, so a 7-day chart on a TimescaleDB store plotted the 4 days raw still held under an axis that said 7; the query-duration, procedure-duration and execution-count charts now route by retention tier through the same ladder and hourly SQL the MCP trend tools use (moved to Storage as `DurationTrendRouting`, pinned equal on both sides) and each chart titles itself with the tier it served from and a truncation note when the store no longer holds the window's head. The viewer's trend SQL no longer fabricates a 0 for the first differenced point; an unrated point is skipped, not plotted or interpolated. The Performance Calendar judges every day against the store's retention horizon with the one `DailySummaryRetention` decision the MCP reader makes (`DailySummaryHorizon`, Storage; the purge's retention constants and fleet rule move down and the service aliases them), so a purged day is grey, not green, and its tooltip says why - on both SKUs. The PVS top-5 trend no longer plots an unmeasured size as 0 MB (Lite's `get_pvs_trend` carries it as null with `pvs_measured`). +- **`get_pg_server_config` counted the page and called it the server** ([#3653]) - `non_default_count` was a count over the rows fetched, so a limit below the number of chosen settings reported the limit as a fact about the snapshot, and `truncated` was inferred from a full page — always up on the default view. The count is now the snapshot's, computed on the row statement above the cap; the page's own count is `non_default_returned`; the `include_defaults` filter is in the statement so `truncated` is about the population asked for; and the three sibling pages in the same file observe truncation off a `limit + 1` fetch. +- **The four PostgreSQL host alerts fire with the tier they earned instead of no tier at all** ([#3653]) - High CPU, Deadlocks Detected, Blocking Detected and Long-Running Query on a PostgreSQL target passed no severity, so the history grids, get_alert_history and every channel coloured them by name after #3635. CPU now grades on the CPU health band's own classifier (Critical at 95% of the configured ACU ceiling), deadlocks on the shared rate grader against the store's deadlock tiers, blocking and long-running query as an explicit Warning with the missing measured bar named. Rows written before this keep their by-name colour. +- **Darling's `get_pvs_stats` trend drew an unmeasured pass as 0 MB** ([#3653]) - a collection on which the DMV reported no PVS size for a database was published as a real zero in the top-5 series, a cliff in a series that had none. The point is now null with `pvs_measured` beside it, the same two keys Lite's twin has carried since #3666; a measured 0 MB is still a zero. +- **The Query Store duration chart now says what served it and where the rollup's floor cut the window** ([#3653]) - The viewer's fourth Performance Trends chart plotted what the corrected Query Store rollup had materialized under an axis spanning the requested window, with no title; it now carries the sibling charts' disclosure — the route word the MCP tool publishes, the grain seam, and a "data begins" clause only when the rollup's measured floor sits above the requested start. The MCP trend reader's tier ladder and hourly SQL became aliases of the Storage definitions rather than pinned-equal copies, so the tool and the viewer cannot drift. +- **The analysis pass no longer says "collection appears to have stopped … NOT an all-clear" for a window the coverage witness proves the collector observed** ([#3653]) - The empty-window rule in both analysis services tested `facts.Count == 0` before it tested the coverage witness, so an observed window that simply produced no fact wore the dead-collector envelope: `unavailable` with a pointer at collection health from `analyze_server`, and the Viewer's dead-collector marker from the scheduled pass. The rule now gates on `ObservedDurationMs <= 0` alone; an observed window with no facts runs the pass and reads `empty` at its stated coverage, with one Information line naming that coverage for the scheduled pass. +- **The PostgreSQL analysis vocabulary absorbs what six content lanes reported back** ([#3542]) - Two `pg_config` facts were inert by construction: `autovacuum = off` and `maintenance_work_mem` had no scorer arm, so they scored 0 and no co-fire could lift them; `autovacuum = off` is now a 0.9 posture fact and `maintenance_work_mem` is evidence-gated on the backlog the way `work_mem` is on the spill. The connection ceiling subtracts PostgreSQL 16's `reserved_connections`, edges can name "the bad actor" through a stable `PG_BAD_ACTOR` alias the graph resolves to the window's heaviest statement, and a bad actor's next reads lead with `get_pg_query_duration_trend`. +- **A target that restarted, failed over or was re-pointed no longer keeps subtracting from the old instance's counters** ([#3653]) - Every delta family subtracted from a baseline cached under the server_id with nothing checking the readings came from the same instance; a failover to a busier replica stored `new total − old baseline` as one interval's work. Both hosts now detect an identity epoch on a read they already make — `sqlserver_start_time`/`@@SERVERNAME` for SQL Server, `pg_stat_statements_info.stats_reset` for the PostgreSQL statements family — compared against the pair persisted in collector_state, forget the affected baselines before subtracting, mark the run's collection_log note (`identity_epoch_changes=1`) and log old→new at Information; Darling's same-id reconnect on a definition edit forgets its baselines too. +- **A quiet hour no longer halves the next hour's Query Store rate** ([#3653]) - the rollup-routed Query Store duration trend rated every point over the gap to the previous EMITTED point, so a corrected-hourly bucket with no rows made its neighbour's denominator 7,200 seconds and published half its true rate; each point class now carries its own denominator - a rollup bucket its bucket width (and is therefore always rated, first bucket included), a raw Query Store interval its spacing (the store holds no interval length for it, said on the CTE). Both the MCP tool and the viewer read one builder. +- **The query-stats trends read the interval the store has, on both SKUs** ([#3653]) - `query_stats` has carried `sample_interval_seconds` from the start, yet the query-duration and execution-count trends (Darling viewer, Lite) LAG-recomputed it from row spacing and divided a restart row's fabricated 0 into a confident 0.00 ms/sec; they now read it three-state like the procedure trends have since V128 - stored interval, 0 = unrated, LAG only for a pre-upgrade collection - through one Storage builder. The MCP reader's own raw query-stats const is the named one-line follow-up. +- **Compose delta aggregates exclude the restart marker** ([#3653]) - every raw-tier SUM/AVG/MIN/MAX over a delta column on one of the ten interval-carrying families now carries its own `FILTER (WHERE sample_interval_seconds IS DISTINCT FROM 0)`, so a restart's (0, 0) marker is no longer averaged in as a measured zero or reported as the window's minimum; gauges, overlays, ratios, COUNT(*), Query Store and the rollup route are untouched. +- **`mute_analysis_finding` stops writing the hash as the path and stops registering the same mute twice, `remove_server` can remove a server that never connected, and eight MCP descriptions stop saying what the code does not do** ([#3653]) - The mute registry write is one guarded `INSERT … WHERE NOT EXISTS … RETURNING` on both SKUs: idempotent per (scope, hash), the `story_path` column names the chain resolved from the retained findings (the hash only as a disclosed placeholder), and the payload says `registered` / `already_muted` / `story_path` truthfully. `remove_server` resolves against `config_monitored_servers` — the table it deletes from — so a definition that never connected is removable, with `ever_connected` and `matched_in` disclosed. Descriptions fixed: `audit_config` (no edition branch exists; PostgreSQL targets are refused and redirected), Lite `get_cpu_utilization` (one ring-buffer record per minute; 15 s is Azure's `dm_db_resource_stats`), the plan tools (`top_operators` publishes its cap/count/truncation and its ranking basis; `missing_indexes` carries a labelled statement-scoped impact and column lists, no paste-ready DDL), `get_ag_health` (lag not banded, lag 0 while suspended, stale last snapshot), the duration-trend catalogue line, and `get_resource_semaphore` row order now matches across SKUs. +- **The web server page labelled page sums as totals, drew one instant as a trend, and four tools called their capture time by another name** ([#3653]) - `get_pg_database_stats` now computes `total_temp_files` / `total_temp_bytes` / `total_deadlocks`, `database_count`, a cluster-wide `cache_hit_pct` and the reset flag as window aggregates on the same statement as its rows, publishes the page's own figures as `returned_*` / `databases_returned` / `cache_hit_pct_of_returned`, observes `truncated` off a `limit + 1` fetch instead of inferring `limit_reached`, and the A7 census's one stated allowance is gone; `get_pg_io_stats` spells its page count `combinations_returned` like every other paged tool. The Memory Grants line chart drew `grants[]` — one instant, every row at the same stamp — under a window caption; it and the Resource Semaphore table are now two tables per read, the `window[]` aggregates (peak waiters and when, peak granted, floor of available, timeouts and forced grants across the window) first and the newest snapshot second, labelled as a moment. `get_database_sizes`, `get_running_jobs`, `get_server_properties` and `get_session_stats` stamp `captured_at` on both SKUs and join the latest-is-a-time roster; the `StampedUnderCollectionTime` allowance is deleted. +- **The wait and perfmon anomaly baselines stop counting every restart's fabricated zero as a quiet sample** ([#3653]) - Interval-honest successor aggregates (`sample_interval_seconds IS DISTINCT FROM 0` baked in, the measured interval carried) replace the legacy pair in place; the provider reads whichever supply covers its window, divides by the stored interval and applies the `LAG > N` heuristic only to pre-column rows, and the legacy pair retires by a self-executing coverage condition about five days after the release. The startup baseline backfill now fires on the first start (its gate read coverage through the real-time view), and a Daily-routed Performance Calendar answers the days its rollup has not reached from raw instead of printing `unique_queries = 0`. +- **Three PostgreSQL pages stop guessing truncation from a full page, and four of the eleven payload rules stop being sentences** ([#3653]) - `get_pg_deadlocks`, `get_pg_plan_capture_readiness` and `get_pg_index_bloat` published `truncated = rows.Count >= limit`, and the two that withhold their summaries on that flag withheld them for pages that were the whole set; `get_pg_plans` cut at `limit` and said nothing. All four now fetch `limit + 1` and observe through one shared `McpHelpers.BoundPage`, the page census sweeps the inference across every tool body on both SKUs instead of a roster (which is why it missed these), and `McpPayloadContractCensusTests` pins errors-one-shape, refuse-what-you-cannot-honor and name-is-truth as far as a census honestly can, with the one-vocabulary inventories (14 truncation dialects, 8 severity spellings, two error shapes) as exact rosters for the lane that collapses them. +- **The MCP query-duration trend still divided a restart's zero into 0.00 after the viewer stopped, and two sentences that #3695/#3696 made false** ([#3653]) - `get_query_duration_trend`'s raw route now reads `query_stats.sample_interval_seconds` three-state through the same Storage builder the viewer runs (`DarlingTrendReader.QueryDurationTrendSql` and both `ProcedureDurationTrendSql` consts are aliases of `DurationTrendRouting`), so a restart collection is unrated instead of a confident 0.00 ms/sec and a measured pass is no longer halved by a LAG spanning it. The trend trio's descriptions say which points are rated over what (rollup buckets over their width, first bucket included; raw points over the stored interval or, lacking one, the LAG). `get_cpu_utilization`'s Darling note ports Lite's corrected cadences (one ring-buffer record per minute on-prem; 15 s is Azure's `sys.dm_db_resource_stats`), pinned byte-identical across SKUs. +- **The perfmon chart plotted deltas under a "Value" label with the divisor sitting unused on the row, and a latch/spinlock restart read as zero** ([#3653]) - both viewers now shape every perfmon series through the stored `sample_interval_seconds` via one shared `DeltaSeriesShaping`: a counter whose name ends in `/sec` plots per second (a name proxy until the `cntr_type` rung lands, its mis-classes named), every other counter plots its raw delta under a legend and axis that say so, and a restart's (0, 0) is a line break, not a trough. Both latch/spinlock snapshot grids render the (0, 0) marker as "—" with an Interval (sec) column saying "restart / first sample", and Lite's `get_latch_stats` / `get_spinlock_stats` publish null, not 0, for it. The gap clause was already landed by #1944's cadence rule; this adds the restart-at-normal-cadence break it could not see. +- **One payload said NoData and No Data, and the page cut had five names** ([#3653]) - get_daily_summary / get_daily_summary_range publish overall_health and health_band from one enum token on both SKUs; four PostgreSQL tools, the autovacuum run history and get_analysis_findings stop inferring a cut from a full page and fetch one past their cap through McpHelpers.BoundPage, publishing truncated and *_returned; McpPayloadContractCensusTests classifies every cut key and severity-word literal by what it is (six of #3699's eight "severity spellings" were other vocabularies) and fails a retired or unclassified spelling by name. McpHelpers.ValidateDaysBack adopts the five inline days_back refusals and list_servers goes through FormatError; FormatError's wire shape awaits ruling Q11. +- **Five families subtracted from the old instance once before anyone noticed it had changed, and Aurora's wait counters had no one watching for a restart at all** ([#3653]) - #3694's only SQL Server identity carrier was `cpu_utilization`, tenth in the schedule, so on the pass where a restart, failover or re-point first became visible wait_stats, latch_stats, spinlock_stats, query_stats and procedure_stats each subtracted the new instance's counters from the dead instance's baseline once and then ate a second (0, 0) re-baseline pass. `WaitStatsCollector`, first in the order on both hosts, now carries the identity pair (`sqlserver_start_time`, `@@SERVERNAME`) as a second result set observed after the rows are read and before anything subtracts, so every SQL Server delta family re-baselines against nothing on the epoch pass; the CPU carrier stays as the fallback for an operator who disables wait_stats, and `ServerEpoch` remembers per calculator and per server the identity it last forgot to, so two carriers seeing one epoch forget once. `PgWaitStatsCollector` carries `pg_postmaster_start_time()` as its own epoch (a clean restart zeroes `aurora_stat_system_waits()` while leaving `pg_stat_statements_info.stats_reset` continuous) and forgets only the wait groups. Both census rosters (`FamiliesThatSubtractOnceBeforeTheCarrier`, `FamiliesWithoutAnEpochCarrier`) are empty and asserted empty. No Azure SQL DB epoch (#3694's ruling stands); the marker is still the run's `collection_log` note plus `collector_state`, not rendered. +- **Lite's plan-cache trend descriptions rated every point over the gap since the previous one, the shared unrated_note named one of two unrated reasons, and the latch/spinlock restart row was spelled two ways across SKUs** ([#3653]) - Lite's `get_query_duration_trend` / `get_procedure_duration_trend` descriptions and instructions row now carry the three-state rule (stored interval; a restart's 0 is unrated; LAG only where no interval was stored). The trio's `unrated_note` names both the stored-0 restart and the first-in-window LAG in one sentence, byte-identical on both SKUs. `get_latch_stats` / `get_spinlock_stats` publish the unknowable row one way on both SKUs — per-second rates null, `interval_seconds` null beside them, the latest delta null — with Lite gaining `interval_seconds` and the per-second pair, Darling's latch nulling `severity` and the banded-from delta instead of a LOW from a zero nobody measured, and Darling's spinlock query returning the interval its null rates lacked. +- **Fifteen rate arms spelled "unknowable" as 0 behind a guard that happened to hide it** ([#3653]) - Every `CASE WHEN interval > 0 THEN delta / interval` that fell back to `ELSE 0` — the file-I/O throughput trend on both SKUs, nine LAG-differenced PostgreSQL trend rates, and the two wait-ms/sec baseline arms — now ends at `END`, so a collection with no interval to rate over yields NULL rather than a measured 0.00; the PostgreSQL trend readers and `get_pg_*_trend` carry that null instead of reading it back as 0, and `evictions_per_second` follows. Two of the fifteen were live: two collections landing in one second on a pre-column store rated a 0 ms/sec sample into the wait baseline (mean 66.7 where the rated collections say 100), so the baseline's row filter also takes Lite's `interval_sec > 0` before the restart LAG. Both census rosters are empty and asserted empty. +- **A falling gauge read as a counter reset because the store did not know it was a gauge** ([#3653]) - `perfmon_stats` gains `cntr_type` (Darling V132 / Lite v62), so the collector writes a gauge such as Total Server Memory (KB) as its level with no delta instead of differencing it like Batch Requests/sec and storing the (0, 0) "counter reset" marker when the level fell. Both viewers' perfmon charts plot a gauge as its value, a rate per second and everything else as its per-interval delta, classified by the stored type with the `/sec` name proxy kept only for rows written before the rung; both SKUs' `get_perfmon_stats` / `get_perfmon_trend` publish `cntr_type` and `counter_kind` and null a gauge's `delta_value`. The ten Wait Statistics "Average wait time (ms)" instances are PERF_AVERAGE_BULK numerators whose average needs a base row this store does not join — stated, not fixed. +- **A maintenance job's own sub-threshold anomalies fold onto the job's incident and name the job** ([#3704]) - #3632's three maintenance edges (SCH_M, IO_WRITE_LATENCY_MS, WRITELOG → RUNNING_JOBS, gated on the job having fired) were keyed on the regular symptom facts, and AnomalyIncidentReconciler folded an ANOMALY_* story only onto a regular story carrying its family — so four one-fact anomalies during one server's 116-minute index-maintenance job paged separately and never said the job's name. A maintenance-family anomaly's fold-target set now includes any story carrying RUNNING_JOBS (the same fired-gate, seen through the story set; no new graph edge), and its frozen Investigation gains one sentence naming the job from the fact's ObjectName (#3693). Non-maintenance anomalies stay solo, as #3632 ruled. +- **PostgreSQL baselines survive a shortened retention, and pg_cpu carries the CPU dispersion floor** ([#3691]) - The four `pg_*` raw hypertables the PostgreSQL-target baselines read directly were user-editable and unfloored, so a 7-day retention would have silently switched every PostgreSQL anomaly detector off; the purge now floors them at the 30-day baseline window, the detector's own minimum-history gate, exactly as #1757 floored `cpu_utilization` / `file_io_stats`. `BaselineMath.AbsStdDevFloorFor` gives `pg_cpu` (percent of the capacity ceiling) the 5-point floor SQL Server CPU carries, by unit parity, so a flat month of ACU utilisation no longer manufactures a display-capped z-score on a one-point wobble. +- **`compare_analysis` sigma-bands the PostgreSQL baselined metrics** ([#3691]) - On a PostgreSQL target the tool banded a 10 → 60 tps move "stable" (absolute rule; `PG_TPS` sits at ladder 0 by design) in the window the detector called 25σ, because only SQL Server keys were mapped to baselines. `PG_TPS`, `PG_CONNECTION_SATURATION` (on its peak session count, never the fraction), `PG_DEADLOCK_RATE` and `PG_CPU_PERCENT` (capacity unit only) now band against their own `pg_*` hour-of-week buckets; the `plan_cache_churn` note speaks the engine of the facts. SQL Server output is byte-identical. +- **`get_analysis_facts` now runs the anomaly detector on both engines** ([#3691]) - the read the tool sold as "every observation the engine sees" was collector + scorer only on Lite and Darling alike, so it never showed an ANOMALY_* / ANOMALY_PG_* fact and the detector's gate metadata (deviation_sigma, fire_threshold, baseline_samples, threshold_lineage) was reachable only through a finding that had already crossed the severity floor. The resolved engine's detector now runs between collector and scorer over the same window, gated on the coverage witness like the pass, so anomaly facts arrive scored with their metadata — including the ones that fired but stayed under the finding floor. Descriptions and instructions say so on both SKUs. +- **Full-edition `perfmon_stats.cntr_value_per_second` divided as an integer, so every counter under one event per second read 0/sec** ([#3653]) - The computed column is now `cntr_value_delta * 1.0 / NULLIF(sample_interval_seconds, 0)` (numeric(33,12)) in `install/02` and `06`, and `06` converges an existing install idempotently, guarded on `sys.computed_columns` by the column's integer type rather than its rewritten definition text. `report.daily_summary_v2`'s throughput rows stop averaging truncated zeros; the Dashboard's read was already fixed in #3658, this is the schema half. +- **The write family tells Aurora the truth** ([#3691]) - On aurora-postgres, PG_CHECKPOINT_PRESSURE and CONFIG_PG_MAX_WAL_SIZE were a permanent measured 0 and a forever-0.4 advisory about a knob the engine never consults (fifty clusters: sixty synthetic timed checkpoints an hour, requested share 0). Both are now emitted `not_applicable` with the shape and reason in metadata, score 0, root nothing, and say what Aurora does instead. Where the engine reports WAL, PG_WAL_VOLUME_SHIFT states the window's mean/peak bytes/s (or why it cannot) and ANOMALY_PG_WAL_VOLUME grades the peak against the server's own hour-of-week baseline, lifting and folding onto the checkpoint incident it leads. +- **Three hourly rollups counted every restart's fabricated zero as a sample, and a service outage longer than a day left a two-hour hole under a floor that said covered** ([#3653]) - `query_stats_interval_hourly`, `procedure_stats_interval_hourly` and `query_stats_db_interval_hourly` carry the collector's knowability verdict in their WHERE and the measured interval summed, every hourly-tier reader takes them where they reach as far as the legacy, and the refresh phase grid re-derived itself around them (heaviest refresh at :18, watch line 900 s). At service start every continuous aggregate is scanned for bucket ranges its source holds rows for and it never materialized, and each is closed with one refresh over exactly its bounds — no policy window widened. +- **The PostgreSQL engine absorbs what the v2 lanes reported back** ([#3691]) - Shortening PostgreSQL retention no longer starves the I/O-latency and WAL-volume detectors (both tables floored at the 30-day window). The I/O-latency baseline is built at a quarter-hour grain with a 250-read floor, so it is judged against this hour of the week rather than collapsing to the hour of the day. A slot-retention or replication-lag finding is now lifted by a fired WAL-volume anomaly (the co-fire had read a fact that could never fire), a Lock wait walks to the idle-in-transaction holder behind it, and the idle finding's next reads include the xmin-horizon and blocking tools. compare_analysis says which PostgreSQL keys are baseline-banded. Wave-3 stubs for the blocking family (lane 17). +- **`get_fleet_overview` runs its collection-health rollup once per minute per host, however many overview calls race** ([#3735]) - the 7-day per-collector health aggregate behind every fleet card is the one read in the overview that does not depend on `hours_back`, so a caller racing three calls with three windows had the store run three identical copies of it concurrently — photographed on the largest production store inside the collectors' flush band, the slowest crossing the `mcp` role's 15 s `statement_timeout` while the plan runs in 680 ms alone. Concurrent callers now share one in-flight scan and a result under 60 seconds old (the fastest collector cadence, so nothing the rollup bands can have moved) is served from memory, per host; a caller that gives up releases only itself, a failed scan is never memoized, and the payload's new trailing `collection_health_age_seconds` says how old that half of the roll-up is. The SQL, the command deadline and the server-side cap are unchanged - the cap did its job. +- **`compare_analysis` sigma-bands the v2 PostgreSQL baselined metrics, and the write family has one definition of "WAL tracked"** ([#3691]) - The v2 lanes stored three more hour-of-week buckets (I/O read latency, replay lag bytes, WAL bytes per second) and their detectors judged against them, but the compare tool still banded those keys by the absolute ladder — a 16× latency move read as a ladder step in the same window its detector called 25σ. `PG_IO_READ_LATENCY_MS`, `PG_REPLICATION_LAG` and `PG_WAL_VOLUME_SHIFT` now band in their bucket's sigma (a fact that is `unavailable`, under its ops floor or holding a mean rather than the peak is withheld from sigma, never banded in the wrong unit); the checkpoint and WAL-volume reads apply the same `WalIsTracked` predicate, so a 0-not-NULL WAL series can no longer be "tracked" to one fact and "not reported" to the other. SQL Server output unchanged. +- **Hour-of-week baselines key on the target's local clock, not UTC** ([#3653]) - Both SKUs keyed the hour×dow baseline on the collector's UTC `collection_time`, so one bucket pooled two local hours across a DST change and a finding's `baseline_bucket` named an hour nobody on the server keeps. The scaffold and the two event arms now key on `collection_time` shifted by the target's offset step function (from `server_properties.time_zone_id`, falling back to `utc_offset_minutes`, then UTC), the lookup uses the same numbers, and nothing keyed is stored, so existing baselines re-bucket at the next compute with no migration. ## [3.8.0] - 2026-09-17 @@ -1367,6 +1433,11 @@ Full entries: [docs/changelog/3.0.md](docs/changelog/3.0.md) [#3650]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/3650 [#3652]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/3652 [#3653]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/3653 +[#3691]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/3691 +[#3704]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/3704 +[#3710]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/3710 +[#3712]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/3712 +[#3735]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/3735 [#3514]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/3514 [#3477]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/3477 [#3495]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/3495