A Claude review that posted nothing no longer finishes green: read-only tools end the denial churn, the job fails when no verdict was submitted, and the transcript survives the runner (main copy, lands first) - #3656
Closed
erikdarlingdata wants to merge 69 commits into
Conversation
get_top_queries_by_cpu and get_top_procedures_by_cpu ordered by summed delta_elapsed_time in both SKUs, so on a wait-bound server the real CPU consumers could be missing from the page entirely and attributed_cpu_ratio read as "hidden CPU" when it meant "wrong sort key". All eight ranking sites now key on CPU: Darling's TopQueriesSql, TopQueriesByHostObjectSql, and TopProceduresSql (CTE cut + post-WAITFOR-trim outer sort) and Lite's twin reads. The viewer's Duration grids keep their elapsed ranking by design. Tests: SQL-text ordering pins for all three Darling consts (plus the rollup const joins the Postgres-dialect theory), and behavior tests in both SKUs seeding a CPU king whose elapsed time ranks last with more groups than the over-fetch admits — under the old key it was cut, not merely mis-sorted. Red-watched: the Lite cut test fails against the old ordering. Fixes #3523 Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…-clear (#3524) (#3553) The analysis pipeline's data-span gate measures LIFETIME history, so a server whose collection died (broken credential, unreachable target) still passes it, collects zero facts over the empty window, and the MCP tool rendered that as status "empty" with "All metrics are within normal ranges" — an affirmative all-clear for a dead collector. Both SKUs' analysis services now set WindowEmptyMessage (beside InsufficientDataMessage) at the zero-facts branch, and both analyze tools return the same "unavailable" envelope get_analysis_facts already uses for this state, pointing the caller at get_collection_health. The true-negative all-clear is unchanged for runs that collected facts and found nothing. The unavailable envelope keeps the #2506 persistence hints, so anchored empty-window runs still disclose non-persistence. Tests, both SKUs, on the REAL 24h gate: 25h of history whose newest row is five hours old -> unavailable; the same server with one benign wait inside the window -> the existing all-clear. Darling's is gated (DARLING_TEST_PG) and ran green against a live PG 18.4 + TimescaleDB store; Lite's runs on the shared DuckDB fixture. The viewer Recommendations tabs render the same false all-clear from the same state and only branch on InsufficientDataMessage; that half is #3551 (out of this fix's file scope). Fixes #3524 Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…k worst slot by severity, and stop spelling unknown growth as measured zero (#3554) get_pg_autovacuum_health (#3534): the reader ranks by GREATEST(dead ratio, insert ratio) but severity came from the dead ratio alone, so an append-only worst_table ten times past its INSERT threshold read "ok". Severity now comes from the same GREATEST the ranking uses, the insert-side ratio is published as insert_threshold_ratio (mirroring the SQL CASE's -1 sentinel handling), a one-sample window nulls dead_tuples_growing and dead_tuple_change instead of claiming "flat" (first_seen_at published so the caller can see the window), and the page-scoped summary counts gain the sibling limit_reached discriminator with past_threshold_count reading both axes. get_pg_replication_slots (#3535): worst_slot was the fattest slot, not the worst-classified one — an active 45 GB keeping-pace slot ("ok") headlined over an inactive 2->8 GB grower ("critical_orphan_filling_disk"). Slots now order by a severity Rank (the wraparound tool's pattern) with size only as the tiebreak; growth is null when either endpoint is the -1 sentinel or the window holds one sample, with unknown-growth Classify arms (warning_retaining_wal_growth_unknown / info_inactive_growth_unknown) so a sentinel can never read as stable; and raw retained_wal_bytes nulls the sentinel as its _gb sibling always has. Tests: Classify unknown-growth arms and the Rank ladder pinned (including that every emittable severity holds a rung above ok); two new gated live classes drive both tools end to end against a real store — the severity-vs-size worst pick, the insert-driven worst_table's severity, sentinel and one-sample growth reading unknown, and limit_reached. Fixes #3534, fixes #3535 Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…s column, caveat-free RCSI advice (#3550) * Fix three honesty defects: null the fake granted-memory field, read the dropped physical-reads column, mirror the RCSI caveats #3529: get_memory_trend shipped a hardcoded total_granted_mb = 0.0 on both SKUs — an agent investigating memory-grant pressure read "granted was 0 all window" and ruled out the real culprit. The field is now an explicit null, the envelope carries a granted_note naming get_memory_grants as the grants series' source, and both tool descriptions stop promising granted memory. Pinned by payload tests (Lite real-DuckDB, Darling gated-live) and per-SKU description pins. #3530: Lite's QueryStore time-slice reader mapped TotalReads AND TotalLogicalReads to ordinal 4 (the logical aggregate — a deliberate alias, matching the QueryStats slicer) but never read ordinal 6, so the SELECT's total_physical_reads was computed and dropped. TotalPhysicalReads now maps ordinal 6, pinned by a fixture whose I/O columns carry pairwise-distinct values — equal values are exactly how the slip stayed invisible. #3531: the LCK_M_IS advice handed out the RCSI ALTER calling it "strictly better" with none of the caveats its LCK_M_S twin carried. Both twins now name the brief exclusive lock the ALTER takes, the tempdb version-store cost (phrased after FactRiskDisclosure's RCSI risk items), and the test-on-a-copy warning, pinned by a shared content test. Fixes #3529, fixes #3530, fixes #3531 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx * Update the two doc lines that still described total_granted_mb as an always-0 parity placeholder Boundary extension approved by the wave lead: doc lines going stale because of the #3529 change belong in the same PR. Darling/README.md's tool-catalog line and DarlingTrendReader's MemoryTrendPoint doc comment now describe the shipped behavior — explicit null plus a granted_note naming get_memory_grants. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx * QueryStore slicer: physical-reads sort plots the physical series, bars and overlay both #3547, folded in per the wave lead. The grid's sorting handler mapped "AvgPhysicalReads" to the LOGICAL series under a physical-reads label — a forced workaround while the reader dropped the physical column, which the #3530 fix ended. The mapping now targets TotalPhysicalReads with the bucket.Value case to match, and the selected-row overlay follows: the timeline read carries avg_physical_io_reads through the same deduped CTE (no new dedup site), the point record gains PhysicalReads, and the overlay selector plots it — otherwise honest bars would have drawn under an elapsed-ms overlay, a new mismatch replacing the old label lie. The distinct-per-column test now covers the timeline's read path too, so a logical/physical swap on either the bars' or the overlay's source goes red. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…tion floor (#3537) (#3555) Two edges escaped the identity-keyed persistence gate in opposite directions. False fire: the identity fraction's denominator counts only holder-bearing collections (the collector emits no rows when the horizon is unheld), so the first holder after quiet hours read as 1 of 1 - 100%, "chronic", off a single sample. False quiet: a horizon continuously pinned past the age threshold by a parade of DISTINCT holders never accumulates any single holder's fraction, so the alert never fired while its own claim ("vacuum is reclaiming nothing cluster-wide") was true the whole time. The identity arm stays as designed - it is the arm that names the thing to kill - and gains a 5-observation floor, the smallest denominator whose majority test cannot be satisfied by fewer than three sightings. The new horizon arm fires when the WINNING xmin_age sat at or above the threshold in a majority of the window's real captures, holder identity ignored. Its denominator comes from collection_log (this collector's own SUCCESS runs, so quiet zero-row captures count) rather than from the holder table - the cheapest honest "collections in the window", counting the exact collector being fractioned instead of inferring cadence from a cadence-mate table. The threshold reaches the SQL as a bind from the evaluator's own constant so read and evaluation cannot drift. A horizon-arm fire without a stable identity names the rotating-holder pattern, still carries the latest holder's remedy, and subjects the stable "rotating holders" sentinel so the host's per-subject cooldown holds across rotations instead of paging once per parade member. Tests: evaluator boundaries for both arms (1/1 and sub-floor totals no longer fire, rotating holders fire under the sentinel subject, the classic chronic holder still fires the identity arm with its original wording, below-threshold never fires), plus a live-Postgres read test that pins the new SQL's counting - ERROR runs, other collectors and other servers stay out of the capture denominator; a below-threshold winner stays out of the numerator while still counting as an observation. Fixes #3537 Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…io_timing (fixes #3533, fixes #3536) (#3558) #3533: the queryid filter ran client-side over a fetched top-duration page (limit * 10), so a plan ranked below the page was unfindable at any window size - and the filtered-empty branch then told the caller capture was working, no plan was captured, and it was "not the query to look at" while the plan sat in the store. The predicate now travels into DarlingPgPlanCaptureReader's SQL (($4::bigint IS NULL OR query_id = $4); null leaves the top-duration page unchanged), the over-fetch and the C# filter are gone (the Viewer's existing call keeps its overload), and the miss text says what was actually searched - the whole window, server-side - and what a miss can mean: never ran (get_pg_top_queries confirms), never crossed auto_explain.log_min_duration, capture not working when it ran (get_pg_plan_capture_readiness), or aged out of retention. #3536: track_io_timing is OFF by default in PostgreSQL, and this read rendered that as 0.000 ms latencies - an impossibly fast disk instead of "not measured". It now mirrors the trend sibling's contract: the setting is read from pg_server_config bounded by the window end (inferred from the window's data when the config was never collected, and io_timing_source says which), io_timing_tracked and timing_note are published, every time-derived field (read_time_ms, avg_read_ms, write_time_ms, extend_time_ms, the read-time shares, total_read_time_ms) is null when untracked, and busiest_basis states what the ranking used - the reader's secondary ORDER BY key (reads) is the entire ordering over a store of zeros, and is now documented as load-bearing. busiest_by_read_time keeps its key because the web tile reads it by name. Tests: SQL predicate pins; both miss-text arms; BuildIoJson wire shape for tracked, untracked, inferred, setting-overrides-observation, and the Aurora null-write interplay; and a gated live fixture seeding 20 expensive shapes plus a 21st cheap one - exactly the limit * 10 page the old path fetched - proving the pinned read finds it, the top page does not contain it, and an absent queryid gets the honest miss. Live parse analysis of every shipped PG read green on PostgreSQL 18. Fixes #3533, fixes #3536 Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…interval (#3527) (#3560) The PERFMON_*_SEC facts, the batch-request anomaly window, and both SKUs' batch-request baselines consumed delta_cntr_value — the per-COLLECTION-INTERVAL delta — as if it were already per-second: 60x truth at a 60s cadence, 300x at 5min. All three consuming reads now divide by the interval consistently, in both SKUs, so the window statistic, the baseline population, and the requests/sec floors (BatchRequestFloor 500, BatchRequestFallback 5000) finally meet in one unit: - Fact collectors (Pg + DuckDb): select the row's measured sample_interval_seconds (#2234) alongside the delta, emit delta / interval as the fact value, and keep cntr_value/delta_cntr_value raw in the metadata with the divisor added. Rows with interval <= 0 (no delta was knowable: first sighting, reset, gap past the delta policy) are filtered so rn = 1 lands on the newest USABLE row — an unknowable rate is never emitted as 0 or as the raw delta. - Anomaly windows (PgAnomalyDetector + Lite AnomalyDetector, Lite-verbatim SQL): AVG/MAX over delta * 1.0 / NULLIF(sample_interval_seconds, 0) with interval <= 0 rows excluded from the sample count. - Baselines: Lite divides by the stored measured interval on the raw rows; Darling's arm reads the perfmon_baseline continuous aggregate, which materializes no interval column and cannot grow one without forfeiting the 31 of 35 days of history the 4-day raw tier can't refill — so it derives interval_sec from LAG(collection_time) over the collapsed series, the established WaitMsPerSec idiom, with the io-arm's DOUBLE PRECISION cast. The restart-exclusion signature stays on the RAW delta in both SKUs (its > 1000 bar predates the division). No stored baseline state needs migration: both providers compute baselines on read (1-hour in-process cache only, cleared by the upgrade's restart), and the CAGG stores per-collection raw deltas — unit-neutral facts — not statistics, so all 35 days of materialized history serve the new unit immediately. Tests, both SKUs: the interval-60/delta-6000 fixture produces a fact of 100; interval-0 rows are skipped at every read (never rates of 0); the anomaly floors demonstrably compare per-second values (a raw delta past the floor whose rate is under it stays quiet); live-Postgres end-to-end proofs for the fact collector and the detector+baseline pair, including the LAG-derived baseline mean and a window that skips an interval-0 row. Fixes #3527 Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…and count measured metrics in the fleet Healthy label (#3528) (#3562) Three alert-semantics honesty fixes from the 2026-09 brains-review campaign: 1. DeadlockCountThreshold and BlockingCountThreshold now floor at PostgresAlertEvaluator.CountThresholdFloor on read, mirroring their V122 PostgreSQL twins, so a store row hand-edited to 0 can no longer make count >= threshold fire on a quiet server. The MCP write bounds name the same constant instead of a literal 1, and the PG twins' doc comments stop calling the gap pre-existing. 2. Store Disk Pressure gains a GB floor (V126: self_disk_free_warn_gb, default 50): the percent trigger additionally requires free space below the floor before firing, the pvs_floor_gb AND-composition, so 400 GB free on a 4 TB store volume stops paging CRITICAL; 0 removes the floor. Full knob plumbing: darling.json field, clamped settings property, seed/read, get_alert_settings/update_alert_settings key parity, and the viewer schema-probe arm. The Settings window box is deliberately deferred to the viewer pass and pinned as an abstinence. 3. Fleet cards carry measured_metric_count/metric_count beside the band (the worst-of fold skips Unknown, so an online server with five of six metrics structurally Unknown still bands Healthy); the web fleet page renders the qualifier ("1 of 6 measured") on partially-measured cards. Rank-neutrality of Unknown is unchanged. Fixes #3528 Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…licy (#3532) (#3559) A delta-family collector scheduled above the CollectorDeltaCalculator gap policy (DefaultMaxGapSeconds = 3600) takes the reset branch every cycle: every delta is (0, 0), facts vanish, baselines fill with zeros, and the product reads green precisely because it stopped measuring. The bound lives once, in PerformanceMonitor.Collectors: the delta family set (the calculator's exact caller census - 8 SQL Server + 2 PostgreSQL collectors, pinned by a source-scan test) and MaxDeltaFrequencyMinutes, half the policy so even an entirely missed cycle still yields a real delta (a cadence AT the policy already fails every cycle, since the gap is the cadence plus scheduling latency and the reset comparison is strict). Enforcement, per surface: - Lite ScheduleManager: UpdateSchedule / SetScheduleForServer refuse (before mutating), and the load path clamps a hand-edited or pre-fix collection_schedule.json with a warning, since there is no user to bounce it back to. - Lite schedule editor: validates before save (mirroring the Darling editor, which Lite previously had no counterpart to), commits pending cell edits first, and shows the cap in a footer line built from the shared constants. - Darling viewer editor: ValidateSchedule moved to the pure CollectorScheduleOverlay (now testable) and extended with the delta bound; footer hint appended from the same constants. - Darling service: StoreConfigProvider.ResolveSchedule's ValidFrequency now knows the collector, so a poisoned store row (hand-written or pre-fix) degrades to "no override" and falls through, exactly like a negative frequency. The web/MCP command path only writes enabled flags, so it needed nothing. Snapshot collectors keep long cadences - the bound is per-collector-kind. Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…he false all-clear (#3551) (#3565) The tabs only branched on InsufficientDataMessage, so an analysis window that collected zero facts (dead collector, unreachable target) rendered "All clear - no current recommendations." Both viewers now consume the WindowEmptyMessage state #3553 added, one rung over from insufficient-data: - Lite's Generate now reads AnalysisService.WindowEmptyMessage directly and renders a distinct notice pointing at the Collection Health tab. - The Darling viewer never calls the engine, so the worker persists the window-empty determination into the V19 analysis_state marker as insufficient_data = false + the engine's message (a shape nothing else writes, so no schema change), and AnalysisStateMarker.WindowEmpty derives it back for the Recommendations tab on both the read and Generate now doors. The genuine all-clear (facts measured, zero findings) is unchanged, and the marker self-heals on the next facts-bearing pass - pinned through the real viewer read against live Postgres. Fixes #3551 Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…real data (#3566) * get_memory_trend joins the grants series: total_granted_mb carries real data Fixes #3548. #3529 shipped the honest placeholder — an explicit null with a note naming get_memory_grants — because the memory_stats source has no grant data. The complete fix joins the memory-grant series both stores already collect (the same v_memory_grant_stats sum the viewers' Memory Overview overlays plot) into the trend payload on both SKUs. The join is nearest-match within 30 seconds, because each collector stamps its own DateTime.UtcNow per run: same-cycle rows sit seconds apart, so an equality join returns nothing, while a wider match would smear a slower grants cadence across points it never measured. Zero-vs-unknown discipline holds on both sides of the match: a snapshot measuring nothing granted is a genuine 0.0, an uncovered point is null, and the granted_note survives only in payloads that have a null to explain — a fully covered window is not captioned with an apology. Darling's read is a new DarlingTrendReader const pinned byte-identical to the viewer's proven MemoryGrantTrendSql; Lite reuses its overlay read with an asOfUtc pass-through so the grants window matches the memory window's. Twin tests pin the pool-summed match, the genuine zero, the uncovered null, note presence/absence, and the 30s boundary (in at 30, out at 31); the same claims run live against Postgres, plus the byte-identical SQL pin. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx * Reword a doc comment the as_of census reads as code AsOfWindowAnchorTests string-scans anchored tool bodies for the literal DateTime.UtcNow, and the tolerance constant's doc comment named it while describing the collectors' per-run stamps. Both tools are genuinely anchored (the joined grants read takes the same windowEnd as the memory read); the comment now says "capture clock" and the census passes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…s tiers, not any-deadlock-is-Critical (#3564) * The daily classifier bands deadlocks as a rate through the card band's tiers, not any-deadlock-is-Critical (#3525) The shared DailyHealthBandCalculator still shipped #3368's un-fixed twin: signals.Deadlocks > 0 banded the whole day Critical. On the measured 43-server fleet that read 87.9% of 24-hour windows Critical, so ~7 of 8 Performance Calendar day cells painted red from deadlocks alone, and the fleet sweep's variable 15-1440min span scaled its Critical rate ~6x from the cadence knob. DailyHealthSignals now carries the Window its counts cover (the ServerHealthMetrics.DeadlockWindow discipline: default(TimeSpan) is unusable, never a divisor), and Classify routes the Deadlocks signal through ServerHealthClassifier.DeadlockSeverity with the store-backed DeadlockRateThresholds (V120) via DailyHealthThresholds.DeadlockRates — Critical/Warning fold into the day's tiers, and the sub-1h/undeclared window arm falls to Warning-not-rate (zero stays out of the trigger). DeadlockRatePerHour/DeadlockSeverity widen int->long for the day-scale counts; every int caller converts implicitly. Producers declare their windows: the calendar-day projections (Darling health reader, viewer, Lite) pass 24h, the fleet sweep passes its own span. The Darling day reads hoist one store-tiers read per range/sweep (the fleet reader's own hot-swap argument), the sweep's Compose takes the tiers as a required parameter, and its would-have-paged deadlock family now fires only when the rate crossed the Critical tier, with rate + tier + count as evidence. The shared deadlock reason/tooltip line reports the rate beside the count ("120 deadlocks (5.0/hr)"), the card reason's own disclosure rule, and the sweep verdict JSON carries deadlock_rate_per_hour + window_minutes additively. Tests: two-window proofs on the day classifier (same rate bands the same over 1h and 24h; same count bands differently), the single-deadlock day control, the sub-hour arms, tier-override pins, sweep ledger pins, a whole-repo census that every production DailyHealthSignals bundle declares its Window (with the calendar projections pinned to 24h and the sweep pinned to its span), and the Lite live-DuckDB calendar now proves the rate path end-to-end with a 480-deadlock storm day. Fixes #3525 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx * Two live health-tool tests catch up to #3525's rate banding DarlingMcpHealthToolsLivePostgresTests planted one deadlock and asserted the day reads Critical - the any-deadlock-is-Critical semantics #3525 removes. One deadlock across a 24h day is 0.04/hr, far below the measured tiers, so the boundary test now pins the honest Healthy band plus the deadlock_count evidence (its real subject is date-boundary visibility), and the planted-rows test pins Warning (its 85% CPU sample trips the HighCpuWarningSamples tier) plus the same evidence. The Critical-band rate path is covered by the PR's own storm-day tests. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx * The still-forming day clamps its window to elapsed time (#3525 review) The calendar-day projections stamped Window = 24h on every row including today's half-open bucket, so an active storm diluted against hours that had not happened yet: 60 deadlocks in the last hour read 2.5/hr (Healthy) instead of 60/hr (Critical) - a false-negative the pre-#3525 any-deadlock trigger could not produce. The shared CalendarDayWindow helper now clamps the current day to its elapsed portion against the read's own clock: the anchored range reads (both SKUs) hand their resolved as_of window end so a backdated read clamps against itself, the live calendar reads use the wall clock, and a day minutes old falls to the sub-hour Warning-not-rate arm rather than a fabricated multiplied rate. Finished days band over their full 24 hours exactly as before, and the sweep path is untouched. The window census now pins all three projections to the clamp helper so a hand-rolled FromDays(1) cannot reintroduce the dilution, and the twin band suites carry the storm proof at the review's own numbers. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
… Store slice-total labels say Total (#3567) * Procedures slicer: physical sort plots the physical series, and slice totals stop calling themselves averages Fixes #3556 — the #3547 bug's verified siblings, found during that fix and kept out of its file boundary. The procedures grid's physical-reads sort mapped to the LOGICAL aggregate under a "Total Physical Reads" label, with no physical case in the value switch — while the slicer's SELECT computed total_physical_reads at ordinal 6 and the reader dropped it on the floor, the #3530 shape exactly. The reader now maps TotalPhysicalReads, and the handler mirrors the Query Store fix: metric TotalPhysReads, the value-switch case, honest label. The overlay side was already wired — ComputeProcOverlayPoints has carried a dormant TotalPhysReads -> DeltaPhysicalReads arm all along, and the handler re-triggers SelectionChanged on metric change, so bars and overlay go physical together with no further threading. The Query Store switch's two remaining dishonest labels — "Avg Reads" and "Avg Writes" over execution-weighted slice TOTALS — now say Total, matching the physical arm beside them and the query-stats grid's vocabulary. Pinned three ways: a distinct-per-column reader test (110/30/50 — equal fixture values are how the drop stayed invisible), and a source-text pin over both sorting handlers' switch arms in the AvailabilityGroupsGridSortTests style, since the handlers are private WPF event handlers. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx * Darling viewer: port the #3547/#3556 grid fixes the review found unported PR #3567's claude-review flagged that ViewerServerTab.Queries.cs — the viewer's copy of Lite's sort-driven slicer metric mapping — still carried both bug classes this PR fixes in Lite. The procedures grid's physical sort mapped to the LOGICAL series with no physical value case, and the Query Store grid had both the "Avg"-over-totals labels AND the un-ported #3547 physical/logical swap. Both handlers now match Lite: TotalPhysReads mapping, value case, honest Total labels. No reader threading was needed anywhere — the shared slicer reader (ReadQueryStatsSlicerAsync) has always mapped ordinal 6's total_physical_reads, all three slicer SELECTs compute it, and the three item-timeline reads already carry physical_reads into ItemTimelinePoint, with ComputeOverlayPoints' TotalPhysReads arm sitting dormant — so only the handlers were lying. One genuine port gap beyond the switch arms: Lite's sorting handlers re-run SelectionChanged after a metric change so a selected row's overlay re-projects onto the new metric; the viewer's handlers never did, which would have drawn the newly-honest physical bars under a stale-metric overlay — the exact mismatch the #3550 commit warned about. All three handlers gain the re-trigger. Pins mirror Lite's: a source-text pin file over both fixed handlers plus the re-trigger (the omission was per handler, so pinned per handler), the slicer SQL shape theory now asserts the reads/writes/physical alias order the shared reader's ordinals depend on, and the live slicer + dedup tests gain distinct-per-column fixture values (Lite's 110/30/50 signature; 41x100/41x10 for the Query Store bucket and 40x100/40x10 for the timeline point) — equal fixture values are exactly how a swap stays invisible. Verified against a throwaway PG 18.4 + TimescaleDB cluster: 11/11 live, 48/48 ungated. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* CHANGELOG: the brains-review wave, one splice (17 entries) Wave-1 of the 2026-09 brains-review campaign landed sixteen PRs on dev with a buffered changelog protocol - agents reported their entries to the coordinator instead of touching this file, so sixteen PRs merged without a single CHANGELOG conflict. This is the one post-wave splice: seventeen Fixed entries covering #3523-#3537, #3547, #3548, #3551, #3556, plus their reference links. #3549 was documentation-only and carries no entry. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx * CHANGELOG: fold the #3561 Dashboard divisor entry into the wave splice PR #3569 landed after the splice opened; its entry joins the same PR to keep the one-splice-per-wave discipline (18 entries now). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx * CHANGELOG: fold the #3563 viewer-pass entry into the wave splice (19 entries) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…) (#3569) Mirror of #3560 (Lite/Darling) into the deprecated Dashboard's frozen analysis twins: cntr_value_delta spans one collection interval, not one second, so the raw reads overstated by the cadence (60x at 60s, 300x at 5min) against thresholds defined in requests/sec. All three reads move together, or the z-score compares across units: - SqlServerFactCollector PERFMON_*_SEC facts divide the delta by the row's measured sample_interval_seconds; interval <= 0 rows (unknowable delta) are filtered so rn = 1 lands on the newest usable row - a counter with only interval-0 rows emits no fact, never 0. The raw delta and the divisor ride the metadata. - SqlServerAnomalyDetector's batch-request window AVG/MAX divide by NULLIF(sample_interval_seconds, 0) with interval <= 0 rows filtered, keeping the window statistic in the same requests/sec unit as the BatchRequestFloor/Fallback bars. - SqlServerBaselineProvider's batch_requests arm computes the baseline population as the per-second rate; the restart signature stays on the RAW delta (its > 1000 bar predates the division). The window and perfmon SQL move to public consts (the Darling twin's tested shape) and GetBaselineQuery goes internal so Dashboard.Tests can pin the text - the Dashboard has no test database. Semantics verified live against SQL Server via a temp-table clone of collect.perfmon_stats. Fixes #3561 Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…oor, and the measured-metric qualifier on the WPF fleet card (#3571) The V126 knob (self_disk_free_warn_gb) lands in the viewer: the Settings window gets its box one knob over from the percent sibling (prefill, >= 0 save gate matching the MCP write bound and read-side clamp, Restore Defaults at the shipped 50, master-switch follow), and ViewerDataService.AlertSettings.cs appends the column to the select / upsert / bind / reader at ordinal 66 ($67). The SelfDiskWarnGbFloorRungTests abstinence pin flips to pin the wired state, with the bind-order roundtrip and the Settings-window pins; V124's end-anchored bind pin hands its top-of-bind claim on, the V122->V124 precedent. The WPF fleet card also picks up the web fleet page's measured-metric qualifier: ServerSummaryItem carries MeasuredMetricCount / MetricCount through the shared classifier fold, and the card tooltip's band label says "Healthy - 1 of 6 measured" for a band folded over unmeasured metrics instead of claiming every metric is inside its threshold. Fixes #3563 Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…cached roots from prior rotations can never break Windows chain building (#3557) (#3572) Windows caches every root a TLS client processes into that user's intermediate-CA store, one entry per store rotation, forever. With every rotation minting a root under the same subject, the cached pile grows until CryptoAPI's subject-matched issuer walk fails outright ("unknown chain building error") - measured at ~50 cached roots on a dev box, killing SslStreamCertificateContext.Create server-side in the #2117 handshake test and the default-trust build a viewer's SslStream runs during connect. A random mint tag in the root CN makes every rotation a genuinely distinct CA, so cached copies of other rotations are never issuer candidates and the pile never forms, on any machine. SKI/AKI does NOT prevent the engine error (tested, matrix-controlled); subject uniqueness does (6/6 fresh-process runs green against a ~55-cert pile left in place). The test harness now returns the server-side exception it used to swallow - "handshake completed = false" alone cannot distinguish an Npgsql rejection from the server never reaching TLS at all, which is exactly the misread that kept this failure looking like a certificate-validation verdict. Fixes #3557 Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…3578) Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
… same table name in two schemas no longer reads as one object (#3576) (#3583) * The Locking & Contention grid names the table as schema.table, so the same table name in two schemas no longer reads as one object (#3576) An Azure SQL DB user reported the FinOps Locking & Contention grid showing the bare table name. On Azure SQL DB — and anywhere per-tenant or per-environment schemas are the pattern — the same table name routinely lives in several schemas, so dbo.Orders and archive.Orders rendered as one indistinguishable "Orders". The store always carried schema_name (the index drill's Identity panel shows it); only the grid dropped it. The row model gains a computed FullName (schema.table, bare name when the schema is empty — the ProcedureStatsRow.FullName idiom) and the Table column binds it, so the qualified name sorts, filters, and exports as one string. The header's filter Tag moves with the binding: the popup filter resolves Tag to a property by reflection, and a Tag naming a property the row lacks filters nothing, silently. Ported to the Darling viewer's copy of the grid and to the deprecated Dashboard's frozen twin (smallest faithful diff). Pinned in Lite.Tests (property, filter mechanism through the real matcher, and both SKUs' grid XAML from source) and Darling.Tests (the viewer's own row model). * The Postgres Index Usage grid names the table as schema.table too (#3576, sibling surface) Same defect class one tab over, Darling-only: the viewer's Postgres Index Usage grid showed the bare table name while the five other table-naming grids on that tab show Schema as a column. The store's reader carried SchemaName from the start; it was dropped in two places — the display row had no SchemaName member for the mapper to fill, and the grid bound TableName. Both halves fixed: the display IndexUsageRow gains SchemaName (mapper copies it) and a computed FullName (schema.table, bare on empty schema), and the Table column binds FullName at width 200. No filter buttons on this grid, so no Tag to move. Pinned through the real PgDisplay.IndexUsage mapper — a pin on the property alone would pass with the mapper still dropping the schema, which is the exact shape that shipped — plus the grid binding from source. Lite has no Postgres tab; the deprecated Dashboard has no such grid.
…, so the baseline engine is no longer notification-inert at shipped settings (Fixes #3526) (#3584) * Extreme, corroborated anomaly findings can cross the notify floor (#3526) Every ANOMALY_* ramp saturates its base at 1.0, the Layer-3 tuning-class cap holds the final at 1.49, no amplifier arm existed, and the shipped notify floor is 1.5 — so the baseline engine could never page at default settings. FactScorer: - IsExtremeAnomaly: a per-fact escape from the tuning-class cap for ANOMALY_* only, at 3x the cutoff the fact fired at (10.5σ robust, 15σ heavy-tail, 6σ classical; 3x the absolute bar on the low-quality fallback path, never for is_new wait profiles; ratio/count/delta families stay capped). - AnomalyAmplifiers: a co-fire arm — sibling anomaly families and measured absolute facts corroborate at +0.2/+0.3, so an extreme anomaly with two corroborators scores 1.6 and a lone or routine one cannot reach 1.5. - The impact-peer escape and CXPACKET's cap are unchanged. Lite.Tests/FactScorerTests: 12 new pins on the escape bar arithmetic per path, the co-fire arm, the routine/lone/single-corroborator negatives, the low-quality sigma blindness, CXPACKET staying capped beside extreme anomalies. * Bound the extremity escape bar at SigmaDisplayCap (review on #3584) AnomalyGate clamps the stored deviation_sigma at 25σ before the scorer sees it, so an operator-scaled anchor above ~8.3 made 3x unreachable and the escape silently dead for that metric — the defect this PR fixes, for a tuned store. IsExtremeAnomaly now takes min(3x anchor, SigmaDisplayCap); two Theory rows pin anchor 10.0 at 25.0 (released) and 24.99 (capped).
… are visible to, and proves rows exist where it can (#3574) (#3585) The block reported recording: true from the GUC (#3175) and stopped there. The reader's real question is "will I see rows", and timescaledb_information.job_history is a security_barrier view that shows a row only to a member of the job's owner role or of the database owner (2.28.1 views.sql: pg_has_role(current_user, <datdba>, 'MEMBER') OR pg_has_role(current_user, owner, 'MEMBER')), while the jobs and job_stats views a reader checks first are not filtered. A least-privilege role can list all 110 jobs and read an empty history on a store recording perfectly; that cost a real postmortem. The reader gains a second, failure-isolated statement (JobHistoryEvidenceSql) that evaluates the view's own two pg_has_role tests for the connection doing the reading, counts the rows it sees over a fixed 24-hour window, takes the population half from the unfiltered job_stats, and publishes reader_role, visibility (All/Partial/None), job counts, rows_observed, newest_row_at, jobs_run_in_window, newest_run_started_at and a contradiction flag. The note names the rule with the predicate, the role it read as, and what that role may see; recording on + reader admitted + jobs ran + zero rows is the new CONTRADICTION arm, with its one benign cause and how to settle it. Evaluating the predicate is load-bearing rather than decorative: in managed mode the MCP host connects as the mcp role, which the view filters OUT, so a bare count would have manufactured the very false contradiction the issue is about on every managed store. The block now says "None — the filter, not the table" there and names the role that can see; a bring-your-own owner connection is the self-proving case. Bind is an explicit timestamptz with Kind=Utc — the inverse of the store's naive discipline, because these are TimescaleDB's own TIMESTAMPTZ catalog columns.
…er will actually take, instead of walking the whole fleet's two-hour slice per server (#3573) (#3586) The read's (server_id, collection_time) composite was on the production store all along — V1 generates it — and the planner priced it out: server_id's physical correlation is 0.022, so the cost model charged one random page per row and preferred streaming the entire fleet's two-hour slice through the time index (Rows Removed by Filter: 691,058 to keep 37,878; 57,307 buffers; 422 ms warm, 10.3 s cold against a 10 s deadline). Forced under random_page_cost = 1.1 the same statement took the existing composite at 5,063 buffers, so a second plain composite would have been priced and ignored identically. Covering removes the heap term the model got wrong: (server_id, collection_time DESC) INCLUDE every other column the read touches, so it plans as an Index Only Scan over one server's rows (rig: 50 buffers vs 1,514, Heap Fetches: 0). It lives in PgTableTuning, not the ladder — results-invariant perf, autocommit, failure-isolated — as a plain CREATE INDEX: on TimescaleDB 2.28.1 compressed chunks get an 8 KB shell so only the live chunks carry it; transaction_per_chunk was measured leaving an invalid parent index after a mid-build cancel that IF NOT EXISTS then skips forever; CONCURRENTLY is refused on hypertables. Tests pin the index's column list against the read's own column references, the statement text against that list, the two rejected shapes, and (DARLING_TEST_PG) EXPLAIN the shipped statement on a seeded store for the Index Only Scan. The 1,744.9 ms deadline anchors keep their history and gain the arc; the 10 s deadline does not move.
…he toasts on the next poll instead of waiting on the service's reload (#3570) (#3587) The tray toast is the viewer's own alert channel, and until now the viewer applied no mute rule to it: a polled alert row toasted unless the SERVICE had stamped it muted. A Snooze wrote the rule to the store and bumped the reload beacon, and then the toast's suppression depended on a four-link chain in another process - beacon observed on the next 15 s tick, a monolithic config re-read succeeding in full, the mute cache refreshed, the alert re-firing THROUGH that cache - none of which the viewer observed. When a link was slow or broken the rule sat in Manage Mute Rules while the toasts kept coming. AlertToastCoordinator.SelectToasts now also takes the viewer's own rule set and skips any row a rule covers, judged with the shared MuteRule.MatchesAt over the row's own ToMuteContext (server + metric as the row spells them, plus the detail-text dimensions the Mute This Alert pre-fill already read). The set is the per-poll mute-rule read that already drives the sidebar bell, plus any rule this viewer just wrote (tray Snooze, server Silence) added persist-then-cache. Keep- last-known on a failed read, and a rule written mid-read is carried over so the guarantee has no one-tick hole. The status line says what happens on each clock, and a failed snooze is reported rather than only logged. Lite never had the gap: its snooze lands in the same in-process MuteRuleService its deliverer consults before the balloon.
…, and samples off the policies' run instant (#3575) (#3588) TimescaleDB's job_stats view assembles job_status from pg_stat_activity and next_start from the bgw_job_stat row, so at both edges of every healthy run one SELECT reads -infinity AND Scheduled - the predicate's dead-job arm. A production store paged on exactly that, 53 ms into a 63 ms run that succeeded. - ReadStuckCompressionJobsAsync re-reads StuckCompressionJobsSql five seconds after a -infinity trip and reports the job only if the arm still trips; a failed confirm defers judgement to the next hour rather than paging. The stuck-Running arm is reported from the first read as before. Injectable seam + pure merge (ClassifyCompressionJob / ConfirmStuckCompressionJobs) so the decision table pins without sleeping. - The worker snaps each hourly due time to :30 past the minute (TimescaleSupport.NextCompressionCheckUtc) instead of UtcNow + 1 h, which slipped a few seconds an hour across the :MM:00 instants the fixed-schedule policies fire on. - Both edges captured on a PG18 + TimescaleDB 2.28.1 rig; the crashed-worker state (the persistent -infinity row) reproduced and still reported. Tests: CompressionStuckConfirmReadTests (new), TimescaleSupportTests pins annotated. Census counts unchanged.
…nOps row marks become theme brushes Dark can actually see (#3577) (#3589) Two contrast defects from a community report, both measured rather than eyeballed (WCAG relative-luminance contrast; AA wants 4.5:1 for text, 3:1 for a non-text marker). The selected tab on Cool Breeze read 2.71:1 - and not because the theme chose that. Every TabItem style set a light Foreground for the selected state, but a string header becomes a TextBlock, and the themes' app-level implicit TextBlock style out-ranks the Foreground that TextBlock inherits from the TabItem, so the setter never reached the text: the header always rendered ForegroundBrush on the accent. The same dead setter meant Dark's selected tab was 1.99:1 (visible in the README's plan-viewer screenshot) and Light's would have been light-on-cyan had it ever applied. An empty implicit TextBlock style inside the header presenter shadows the app-level one, so the header inherits the TabItem's Foreground the way stock WPF intends, and a new per-theme AccentForegroundColor/Brush is the ink for text on any accent fill: white on Cool Breeze (5.39:1), the darkest surface tone on Dark (7.52:1), the page text on Light (6.79:1, no change). The same ForegroundBrush-on-accent pair recurred on the highlighted combo item, the accent button, the selected calendar day and Dark's grid selection (all 1.99:1 on Dark, 2.71:1 on Cool Breeze); they take the same ink through the same key. Cool Breeze's AccentHoverColor moves one step toward the accent (#2B87C8 -> #267BB8) because no ink passed on the old shade (white 3.89:1); white is 4.56:1 on the new one. TabCloseButton drops its hard-coded White so the x inherits the header ink - it was invisible on the light themes' unselected tabs. The FinOps "Mark Done / To Do / Do Not Do" tints were three fixed 20% overlays in DataGridRowMarks, which composite to 1.2-1.4:1 against Dark's near-black rows - the operator could not find the rows they had marked. They are theme brushes now (RowMarkDoneBrush / RowMarkToDoBrush / RowMarkDoNotBrush, painted by resource reference so a marked row follows a theme switch). Light and Cool Breeze keep the shipped tints; Dark builds its own from the theme's Success / Warning / Error colors at the opacity that pushes the mark as far from the row as it can go while the row text holds >= 4.5:1 on both row backgrounds (marks 2.9-3.1:1, text 4.5-5.1:1). The literals stay as the fallback for a host without the keys, and a test pins the keys in every theme file of both apps. Darling viewer theme copies ported 1:1; the deprecated Dashboard twin carries the identical tab/accent defect and gets the same fix (no marks there - it has none). Arm (B) of the report - user-maintainable colors with reset - is not in this change.
… truncation is detected not inferred, and no page count is called a total (#3541 A3) (#3594) * MCP pages say what bounded them: caps bind to the caller's limit, truncation is observed, no page count is a total (#3541 A3) Six MCP tool groups on both SKUs carried a HIDDEN reader cap -- LIMIT 200 / 50 / 500 under a tool that advertised `limit`, a Take(limit) on top, and the capped count published under a total_* name. An agent that cannot see the code read a 200-row page of a 5,000-event window as the window, and the "widen hours_back" advice could never help because the cap was on rows, not time. get_blocking / get_blocked_process_reports (200), get_deadlocks (50), get_alert_history (+ a hidden dismissed = FALSE filter), get_long_query_completions (200; Darling kept the SLOWEST, Lite kept the NEWEST then re-ranked, so Lite could omit the window's slowest run), get_plan_corrections (200 newest per-cycle re-captures ~ 16h reach regardless of hours_back), get_waiting_tasks (500 / unbounded, bare envelope) and get_wait_stats (50 under a limit up to 1,000). ONE pattern, the one #3287 established for get_collection_log: the cap is the caller's limit bound as a SQL parameter; the tool fetches limit + 1 and reads the extra row as `truncated` rather than inferring it from count >= limit; the page publishes oldest_returned_* / newest_returned_* and names its `order`; the page count is *_returned, never total_*. get_deadlock_detail and get_blocked_process_xml moved their graph/XML predicate into the SQL so limit counts graphs, not rows they would have discarded. get_blocking / get_deadlocks honour #2159's whole-window promise for dedup_key: the fingerprint scan is bounded by a stated 2,000-row ceiling, the payload carries rows_examined / scan_truncated, and a no-match answer on a scan that ran out says so. get_alert_history states and MEASURES its filter: dismissed_excluded and dismissed_excluded_count on every page, include_dismissed to lift it (each row then labelled `dismissed`), and an all-dismissed window names the filter instead of calling itself quiet. Lite's completions read is now duration-ranked in SQL, so both SKUs serve the same population. Descriptions on every touched tool say what bounds the page. Grids keep their caps (the readers default to them). Census tests pin the dialect across both SKUs; Lite's run against a real DuckDB, Darling's live half is gated on DARLING_TEST_PG. * Pass the as_of anchor by name on Lite's slowest-completions read; declare the new LF-reading pin Two census pins from the first CI run. AsOfWindowAnchorTests requires every LocalDataService call that CAN take asOfUtc to be given it by name, so the anchor's presence is visible to the scan rather than inferred from position; the new GetSlowestLongQueryCompletionsAsync call passed it positionally. RepoFileAdoptionTests holds the exact set of pins that read LF-normalised source; McpPageContractTests is one (its Lite-description anchor spans the line break on get_plan_corrections), so it is declared.
…urce like get_query_trend already does, so a 7-day request no longer returns 4 days labelled quiet (#3541 A2) (#3590) * The duration-trend trio routes by retention tier and discloses its source like get_query_trend already does (#3541 A2) get_query_duration_trend and get_procedure_duration_trend read raw query_stats / procedure_stats only, whose rows a TimescaleDB store drops at four days, while hours_back accepts 168 — so a 7-day request returned 4 days under a label saying 7, and when nothing survived the empty branch called the window "genuinely quiet — widen hours_back", which is false (dropped, not absent) and harmful (widening cannot help). Both now route through the ladder get_query_trend already uses (age → availability → coverage, DarlingTrendReader.ResolveTier) onto the hourly rollups, dividing by the bucket width rather than a LAG, and every sibling — the Query Store one included — publishes source / effective_start / effective_hours_back / truncated / bucket / aggregate_note on the data path and the empty one. get_query_trend gains the availability + coverage rungs (it answered 42P01 on plain PostgreSQL past four days). Lite's twins publish the same block with Lite's truth (raw, per-collection). value gains its named twin elapsed_ms_per_second on both SKUs. Partial: completes checklist item A2 of #3541 only. * The retention clock lives in the reader, not the anchored tool bodies AsOfWindowAnchorTests pins, as an absolute, that a tool advertising as_of never names DateTime.UtcNow — its only "now" is the anchor it resolved. The trio's route resolution measured raw age against the wall clock IN the tool body, which is the right clock for retention (the anchor decides the window, not how old its rows are) named in the wrong place. ResolveQueryDurationTrendRoute / ResolveProcedureDurationTrendRoute now default nowUtc to the real clock the way GetQueryHistoryAsync already does, and the route carries ResolvedAtUtc so the empty branch's horizon comparison reads it from the route. * RawRetentionApplies is scoped to the grain, because the arming gate is Review finding: the flag was rollups != None, a store-wide test, while the #1680 gate arms each raw table's purge only once that table's OWN rollup covers it. On #1664's failure-isolated partial build, a grain whose rollup is missing has its raw rows intact — and the store-wide flag would have told that grain's caller the rows were dropped and widening cannot help. Now hourlyAvailable, the per-grain flag already threaded through the resolver; the past-horizon raw branch's dead "rollup not present" arm goes with it (a grain with no rollup can no longer reach that branch), and the pin asserts the per-grain split. * RawRetentionApplies: say that existence-not-armed is the deliberate upper bound, and why (review question)
… over, so a restart's fabricated zero reads as unknowable instead of 0.00 ms/sec (#3540, V127 / Lite v60) (#3595) * Interval columns for the four naked delta families (V127 / Lite v60), so a restart's fabricated zero reads as unknowable instead of 0.00 (#3540) * Rung test counts the view statements by prefix; the lock-wait live test asserts the first collection is absent, not 0.00 * Lite pins: the wait_stats DDL carries the trailing interval; the lock-wait tool tests assert the first collection is absent, not 0.00
…n, tiered from 14 days of fleet data, so a week of blocking and an hour of blocking stop getting the same colour (#3539 A2/A3/A8d) (#3596) * Blocking and CPU bands are rates over the window they were measured in (#3539 A2/A3/A8d) BlockingSeverity bands blocked-process reports per hour over a declared window (Warning 5/hr, Critical 20/hr, the 60 s / 10 s wait arms rate-independent), with the deadlock band's too-short-window rule; tiers off 14 days of dogfood-fleet blocking. The daily classifier delegates blocking to that band, scales its high-CPU bar with the window (excursion-scale minimum 6, sustained-heat rate 1.25/hr = 30/day), and bands collection errors as a share of the window's runs against the collector-health classifier's own 20% bar at its tier (Warning, never Critical). CollectorSeverity grades the failing count as a share of banded collectors (Critical past the same bar). The daily SQL projects collection_runs (trailing column, both SKUs); every producer declares the new denominators; the sweep ledger, cards, calendar tooltips and MCP day tools disclose the rates and shares they banded on. * Re-key the viewer collector-health helper pin on its widened tuple (#3539 A8d) ViewerFleetRollupTests anchored on the helper's declaration text; the tuple gained the Total the collector share divides by, so the anchor names the new shape.
… ten-minute window like its PostgreSQL twin, so one slow wait no longer pages and a THREADPOOL storm no longer sleeps (#3539 A4) (#3593) * SQL Server's Poison Wait alert measures accumulated starvation over a ten-minute window like its PostgreSQL twin (#3539 A4) The SQL Server Poison Wait alert judged ONE collector delta row's avg-ms-per-wait against PoisonWaitThresholdMs (500), presence-flat CRITICAL, no volume floor: one 600 ms wait paged and a THREADPOOL storm of thousands of 8 ms waits slept. The PostgreSQL twin (#2711) already fired on accumulated wait over a ten-minute window, graded Warning/Critical, under the same alert name, switch and mute key. Port the PostgreSQL shape to SQL Server. The read (both SKUs) SUMs delta_wait_time_ms / delta_waiting_tasks per poison wait type over the window with no task filter, no LIMIT and no threshold; the new shared PoisonWaitEvaluator grades the sum (Warning at 1.0 avg tasks stuck = 600 s / 10 min, Critical at 10) and both engines' constants are one definition (PostgresAlertEvaluator's are aliases). Fleet margin on the SQL side: the worst ten-minute bucket across 43 servers / 4 days was 0.0097 avg waiters, ~100x under the bar. The clear arm now requires an observed window (an empty read holds the alert; the retired shape announced Cleared on collector silence). PoisonWaitThresholdMs stays readable and persisted but is no longer consulted; docs and settings comments say so. * Settings windows stop offering the retired poison-wait ms bar; live E2E compares at the store's precision Review finding on #3593: both SKUs' Settings windows still labelled the poison checkbox "avg wait >=" beside an editable ms box, and the save-preview line read "poison waits >= Nms avg" - a knob the engine no longer consults. Label now says the alert measures accumulated wait over a 10-minute window and that the ms bar is retired; the box is disabled with a tooltip and kept so saved settings round-trip; the preview states the live bar from the shared constants. Identical edit on Lite and the Darling Viewer. CI: the live read-adapter E2E compared a seeded DateTime (100 ns ticks) to the value PostgreSQL stored at microsecond precision; compare at the store's precision and assert the in-window row's clock won separately. * PoisonWaitsSql: state the read's cost bound by shape (predicate-identical to the retired read, ~30 rows per server in the head chunk), not by index * Settings windows: the retired poison-wait ms box stays disabled through the master-switch loop, both SKUs (review round 3)
…: the adapter returns the collection that produced the rise and the engine remembers which one it already reported (#3579) (#3628) * Forced Plan Failing fires once per observation, not once per cooldown: the adapter returns the collection that produced the rise and the engine remembers which one it already reported (#3579) * The live access-path seed is floored to whole microseconds so the #3579 stamp round-trips PostgreSQL's timestamp to the tick (#3579)
…very surface that shows it, so a 10 GB bar stops meaning 10 GB per five minutes on one store and per day on another (#3539 A8c) (#3631) * The File Growth threshold means one thing — megabytes per hour — at every surface that shows it (#3539 A8c) FileGrowthRiseMb was compared against the raw growth inside the lookback window while the lookback was a second knob clamped 5..1440 minutes, so the same 10240 meant 10 GB per five minutes on one store and 10 GB per day on another. The knob is now a rate: GetBreachedFiles holds the in-window growth to FileGrowthRiseBarMb (rate × lookback/60), which is byte-identical on the shipped 60-minute lookback and the same sustained rate on every other. Scaled to the CONFIGURED window rather than read off the measured span, because the measured span on a freshly-collecting server is one collection interval and one autogrowth in it would extrapolate to twelve an hour. Both Settings rows, both preview lines, the alert's threshold line (rate + window + the bar in MB), the card's Growth field, IAlertEngineSettings, both MCP descriptions and the settings carriers now say MB/hr averaged over the lookback, in one shared phrase (AlertContextBuilders.FileGrowthRiseUnit) that FileGrowthRiseUnitCensusTests holds across nine surfaces. Key, column and integer are unchanged. * DarlingConfig.FileGrowthLookbackMinutes carries the same averaging-window note as its twins (#3539 A8c, review nit)
…he nightly rebuild's three cards become one incident that names the job (#3538 A6/A9) (#3632) * Story confidence measures corroboration instead of path length, and the nightly rebuild's three cards become one incident that names the job (#3538 A6/A9) * StoryConfidence in its own file; the exhaustive pins run to the engine's real 11-node maximum path (review)
… asks 'was it delivered': the once-a-day gate now stamps delivered-today in the store, so a restart re-announces nothing and a failed delivery still retries (#3580) (#3633) * The daily digest and sweep rollup gate on delivered-today in the store, not fired-today in process memory: a restart no longer re-announces them, and a failed delivery still retries (#3580) Both once-a-day gates were a ConcurrentDictionary under one fixed key, so a fresh process delivered both documents again whatever the previous process had delivered an hour earlier — six re-announcements among ~23 channel posts on the v3.8.0 install night. The gate now reads a delivery stamp from collect.collector_state (V44, no rung; server_id 0 = the fleet sentinel both existing fleet-scope writers use, collector_name 'self_alert') and writes it only when the deliverer reports a disposition other than `failed`, so the install night's recovery case — a failed pair re-attempted after the restart and landing — is preserved by construction. - IAlertDeliverer.DeliverAndReportAsync: a default-implemented reporting twin of DeliverAsync (null = unreported); DarlingAlertDeliverer reports the AlertDelivery its history row was written with. - DarlingSelfAlertEvaluator: FireAsync returns the report; the two documents stamp AFTER the fire; the store seeds memory after a start (#981's cooldown seed shape), memory serves the process; a stamp read fault falls open to memory, warns and is counted (#3013); a stamp write fault warns, uncounted. - Census pins re-justified: AlertMasterSwitchSurfaceTests (new delivery seam name, generic FireAsync return), AlertReadFailureSurfaceTests (9th counted evaluator site, 12th exempt), Lite.Tests inventory (8th fleet-scoped read), AlertReadFailureCounter.FleetScopedReads names the new read. - Tests: SelfAlertDeliveryStampTests (both documents, simulated restarts, failed/muted/unreported dispositions, read/write faults) and a DARLING_TEST_PG-gated live round-trip through the real table. * The stamp store is asked once per document per process: an empty or failed answer is cached as a sentinel, so the apply half's gate does not re-read, re-warn or re-count a fault the evaluate half already met on the same tick (#3580 review) Review catch: the "one store read between the two halves" cost model only held when the read succeeded — a throw (or no row) returned without seeding memory, so the apply half re-asked, doubling the warning and the #3013 count per tick for one logical fault. Exact under the single-writer fact: a store with no row on this process's first tick cannot gain one except through this process, which seeds memory directly. New pin drives two consults on one tick with both the read and the delivery failing — one read, one warning, one count.
…t-history grids show the severity the alert actually fired at instead of the colour its name implies (#3539 A6/A8e) (#3635) * A server nothing has banded yet is Unknown, not Healthy, and the alert-history grids show the severity the alert fired at instead of the colour its name implies (#3539 A6/A8e) * Keep dismissed the last selected column (the pin anchors on it), and paint the nothing-measured border with the cached Warning brush
…nd query_stats stores the statement offsets its delta key is made of, so no restart zero reads as a measurement anywhere and the query seed can finally find its keys (#3540, V128 / Lite v61) (#3630) * Every delta family stores the interval its deltas accrued over, and query_stats stores the statement offsets its delta key is made of (#3540, V128 / Lite v61) The completion of V127: sample_interval_seconds on procedure_stats, memory_grant_stats, pg_wait_stats and pg_statement_stats, written as the minimum over each row's delta groups; statement_start_offset / statement_end_offset on query_stats, stored raw (-1 included) so both hosts' restart seeds rebuild the collector's delta key from the store. The census's still-naked list is empty; the seeding census's pass-window-only set is empty. Readers of the four families that divide by time take the stored interval first, LAG only for pre-V128 rows, and never ELSE 0. get_resource_semaphore emits sample_interval_seconds again on both hosts. * CI round 1: the PostgreSQL pair's CREATE TABLE rungs carry the interval for the fresh population (the V101 two-places rule); the resolving-view definer literal moves to 128; the backlog fallback pin records why it stays hash-keyed; the pre-V128 first snapshot is absent in the trend read test; two of my own pins fixed (V4 passthrough skipped, CTE-first FROM) * CI round 2: the two Lite procedure-trend tool tests expect the pre-v61 first snapshot to be absent, the same correction v60 made for the wait trends
…t once per cooldown: the adapter returns the collection that produced the growth and the engine remembers which one it already reported per file (#3636) (#3639) The rise gate's input is a stored delta between two collections of database_size_stats, which lands hourly; the engine re-asked every ~30 s and re-fired on every 5-minute cooldown against the same two rows — up to twelve cards per growth event. #3579's mechanism, one condition over, at the worse ratio. Both SKUs' reads project the newest sample's collection_time as observed_at, appended last so the fourteen bound ordinals do not move. DatabaseFileGrowthInfo gains a nullable ObservedAtUtc. The engine keeps, per (server, database, file), the stamp it last fired on beside the per-server cooldown clock and active flag, and declines to send the card when no breached file is news: a level breach is always news (standing level, re-fired by design); a rise-only file is news only with a stamp newer than the one it last fired on. Stamped on fire only, including muted fires; pruned as files leave the breached set; cleared on recovery. In-memory like the forced-plan memory — a restart may re-fire once, which the emptied cooldown clock already did.
…defaulted one: CONTRIBUTING's Two-Store Parity rule names this interface, so Lite's deliverer and every test fake now state their (null) answer by hand (#3580 review) (#3640) Review note taken: a default body compiled and left LiteAlertDeliverer and fourteen fakes quietly inheriting an answer nobody wrote down — the exact pattern CONTRIBUTING forbids for this seam. Lite's implementation returns null and says why (its send seam returns no disposition; it hosts neither daily document); the DarlingAlertingTests wrapper forwards the inner deliverer's report rather than swallowing it. Also: an xUnit2013 analyzer warning in the new two-consults pin (Assert.Single over the Regex matches).
…folds one cause into one row, so same-hour-yesterday noise stops reading as a verdict (#3538 A3) (#3634) * compare_analysis bands each delta by the server's own dispersion and folds one cause into one row, so same-hour-yesterday noise stops reading as a verdict (#3538 A3) The tool compared ONE window against ONE window and banded every key by its severity delta on a flat +/-0.1 dead-band. Severity is a threshold-formula artifact: a saturating ladder read a doubling from 30% to 60% of observed time as "stable", a trace of a 1%-bar wait read a few seconds an hour as "worse", one I/O stall surfaced as four to six worse keys, and BAD_ACTOR_<hash> keys churned into new_issues/resolved_issues on plan-cache eviction alone. Shared ComparisonBanding (PerformanceMonitor.Analysis) now decides every verdict for both SKUs: a key with a same-unit per-server baseline (CPU %, read latency, connections) is banded in that bucket's robust sigma (delta_sigma, band_source "baseline", trustworthy buckets only); every other key changes status only when the value moved >= 25% of the larger side AND the larger side sits >= 0.25 up its own Layer-1 ladder (the scorer's base severity, so every concerning bar is reused with no number copied); one-sided keys are new/resolved only when they register on their ladder; BAD_ACTOR_* appearances are plan_cache_churn; rows carry a physical-cause family and the payload carries one family row per cause; the rules are stated in band_rules; every verdict row carries coverage_caveat when either side was partly observed (#3592 composes). ComparePeriodsAsync on both SKUs returns the dispersion map, fenced so a failed baseline read degrades to the absolute rule. * Verdict order is direction before magnitude, so a mixed-direction family's worst member is its regression (#3538 A3 review) The row sort partitioned only stable from non-stable; worse and better shared a bucket ordered by relative move, so a family whose BLOCKING_EVENTS fell 60% while its LCK_M_S rose 30% took the improvement as its worst member and left families_worse. Rows and families now rank worse > better > stable, then move; a mixed-direction family pin covers it.
…rs, their Long-Running Query alert skips maintenance like SQL Server's does, and every one-sided alert says why it is one-sided (#3539 parity) (#3638) The fleet card's deadlock dot on a PostgreSQL target read Unknown while pg_database_stats held the server's own deadlock counter, because the fleet reader took v_deadlocks (the SQL Server extended-event capture) and ServerMetricSources nulled the structural zero. Both fleet surfaces (the service's get_fleet_overview / /api/fleet reader and the viewer's Overview card and totals) now difference pg_stat_database.deadlocks per (server, database) series between consecutive samples, clamp at zero across a statistics reset, sum across databases, and band the result through the SAME deadlock_warn_per_hour / deadlock_critical_per_hour tiers as the SQL Server graph count. Never SUM(deadlocks): the column is a lifetime counter repeated in every sample. FleetDeadlockSource.PostgresTarget now means "counted from the server counter" and is a covered arm; servers_read counts it and postgres_servers is a sub-count. A PostgreSQL target whose pg_database_stats collector is silent or denied reads those arms, exactly as a SQL Server whose deadlocks collector is. The PostgreSQL Long-Running Query read gains the noise opt-outs SQL Server's has had: non-client backends (mirrors session_id > 50), the VACUUM/ANALYZE/REINDEX/CLUSTER statements by command tag (no text is stored, so the tag is the only handle), the dump/restore utilities by application_name on the shared longRunningQueryExcludeBackups switch, and excludedDatabases after the read. CREATE stays reported with its tag. The Darling README gains an engine-coverage table for every built-in alert and band, stating why Blocking Wait Time and Database File Growth are SQL-only, why the PostgreSQL blocking BAND stays Unknown (sampled evidence against tiers measured on engine-recorded reports), that Poison Wait is one shape since #3593, and that slot retention's SQL analogue (log_reuse_wait_desc) is collected only as a load-time snapshot today.
…cet by facet from the stored config, so a target with everything off stops looking identical to an instrumented one (#3607) (#3643) * get_pg_logging_audit judges a PostgreSQL target's logging settings facet by facet from the stored config, so a target with everything off stops looking identical to an instrumented one (#3607) Plan-capture readiness established the shape - a facet per setting, the remedy per finding, the hosting flavour's own syntax - and covered only auto_explain's three preconditions. The rest of the logging surface (log_lock_waits, log_temp_files, log_autovacuum_min_duration, log_checkpoints, log_connections / log_disconnections, log_min_duration_statement) had no audit at all, and a target with all of them off answered every counter-based read exactly like a fully instrumented one. This is a READ over pg_server_config's newest snapshot, not a collector: every judged setting is a core GUC the config collector already stores hourly with its value, source, unit and context, so the judgment is a pure function computed on request and the snapshot's collection_time is the audit's as_of. Verdicts are instrumented / partial / off / unknown; partial names what a threshold hides and is the recommended posture for log_min_duration_statement. Consumers are named honestly as PLANNED (#3601/#3602/#3603) beside the counter read that exists today. Hosting flavour comes from rds.* parameters in the same snapshot, which is the one fact the registry's engine token cannot supply. Plan capture's own settings are listed with the readiness facet that owns each, never judged twice. Registered beside the plan tools, dispatched on the web with a Configuration-tab-adjacent panel under readiness, censused in the instructions / README / llms.txt / runbook, ratcheted in the cross-app inventory pin. * The audit speaks #3541 A10's stamp dialect: captured_at on the wire, LATEST IS A TIME in the description, and a Darling-only allowance in the stamp roster dev moved under the lane: #3637 landed McpLatestSnapshotStampTests, whose reader-call sweep caught GetNewestSnapshotAsync and asked for a Shape. The roster is a roster of PAIRS (every entry names its Lite twin's file), so a Darling-only tool gets a DarlingOnlyStamped list held to the same three Stamped assertions, and the new reader's SQL joins the stamp-column theory. Threshold() became a block body so TsqlConventionGuardTests' member scan reads it whole rather than growing KnownTruncatedRanges. * Review round 2: pending_restart reaches the row and the summary, the catch re-checks the engine gate, the ten-minute default is PostgreSQL's own verdict, and the two unknowns are told apart pending_restart was fetched, documented as load-bearing, and dropped before the Facet - on exactly the row whose remedy needs qualifying, because the value judged is the RUNNING one and the file already holds another. It is now on every facet with a restart_note, named at the top as pending_restart_settings the way get_pg_server_config names them, and seeded in the live test. The catch block asks the engine gate before answering error, as the plan tools do. atDefault reads source = 'default' rather than comparing setting text to boot_val - an administrator who writes 600000 into the file has made a choice, which is the anti-pattern DarlingPgServerConfigReader's own comment names. An unparseable value now says it IS in the snapshot rather than that it is not.
…ilter is part of the query: parallel_only/min_dop/blocking_only cut before the page, a negative window is refused, an unknown source names the accepted set (#3541 A9/A13) (#3641) * The daily summary stops painting purged months green, and every MCP filter is part of the query (#3541 A9/A13) A9: the daily aggregate's day spine outlives its signals (collection log 60 d, alert log 90 d, signals 30 d) and COALESCEs each purged signal to a measured zero, so a year-long range banded the purged months Healthy. Both SKUs' readers now compute the store's retention horizon (Darling: the shortest EFFECTIVE fleet retention among the seven signal collectors and the two log constants; Lite: RetentionService.ArchiveRetentionMonths) on the reader's wall clock, the SQL projects how many signal sources still hold rows for the day, and the shared DailySummaryRetention.StateFor judges each day: purged (past horizon, no signal rows -> No Data, never Healthy), past_horizon (rows survive, verdict withheld), no_run_record, collected. The range tool publishes retention_horizon / days_before_horizon / purged_day_count / collected_day_count and each row its data_state + data_note; the single-day tool answers a purged day with the unavailable envelope. summary_date is TryParseExact yyyy-MM-dd on both SKUs (McpHelpers.ParseSummaryDate). A13: parallel_only / min_dop are a HAVING floor on the grouped population before the CPU ranking and the cap (both SKUs, both Darling grouping reads), with filter_applied on the payload and a filtered-miss sentence instead of "no query stats". get_active_queries pushes database_name and blocking_only into the SQL, counts the filtered population with COUNT(*) OVER (), pages at limit + 1 (snapshots_returned / truncated / bounds / order), never strips a head blocker some row in the same capture names (is_head_blocker), and names why a victim's blocker is absent (blocker_not_shown: not_captured / filtered / past_page). The three uncapped reads refuse a non-positive hours_back through McpHelpers.ValidateUncappedWindow instead of Math.Abs. The get_analysis_facts source filter is validated against FactScorer.KnownSources (15, pinned to every collector literal) and refused with the whole set. Census: get_active_queries joins McpPageContractTests' paged dialect on both SKUs; new A13 pins (no .Where after the read, predicates before ORDER BY/LIMIT, no Math.Abs on a parameter anywhere, the three uncapped reads on the shared validator); live PostgreSQL round-trips for the filter, head-blocker, purged/past-horizon/override arms; Lite DuckDB twins. * xUnit2009: Assert.EndsWith for the daily SQL's trailing ORDER BY pin * A day inside retention with no run record keeps its band; the displaced range-read summary goes back to its member; the two fixed-date Lite fixtures follow the horizon CI's first execution caught three things the Mac harness could not. (1) PerformanceCalendarDataTests has always pinned an alert-only day (no collection-log run) as Warning — the alert is real and inside retention every zero is a measurement — so no_run_record is now a DISCLOSURE (error share has no denominator) that keeps the band, and only the two past-horizon states withhold it; both SKUs' ToSignals, the descriptions, the instructions rows and the pins say so. (2) The same fixture's fixed July 2026 month would have drifted past the three-month horizon within weeks; it is now the calendar month two months before the current one. (3) DocCommentHygiene: the DailySummaryRangeReadResult insertion had displaced GetDailySummaryRangeAsync's summary block. Plus the FindingsRetentionHorizonPinTests literal pin follows the promoted RetentionService.ArchiveRetentionMonths.
…y hour and paid twelve index inserts per row for eleven indexes nothing reads; the capture-down alert read decompressed a server's whole collection_log to learn two statuses (#3597, partial) (#3647) * The interval-hourly refresh stops paying twelve index inserts per re-materialized row for eleven indexes nothing reads, and the capture-down alert read stops decompressing a server's whole collection_log to learn two collectors' latest status (#3597, partial) Rig-measured on PostgreSQL 18.4 / TimescaleDB 2.28.1 at one tenth of the largest production store's scale: a refresh of query_store_stats_interval_hourly recomputes exactly the hour-buckets the invalidation log marks dirty — the newly closed one every hour, plus one whole fleet-wide bucket per distinct past hour any backdated row landed in (a ONE-row backdated insert re-materialized 36,120 rows; twelve backdated rows in one transaction across twelve hours re-materialized twelve buckets, clean hours included). The window's width is not the cost; the dirty-bucket count times the per-bucket cost is. The only backdated writer is QueryStoreBackfill, by design (collection_time = slice ceiling). The per-bucket cost carried TimescaleDB's default create_group_indexes: eleven (column, bucket DESC) btrees on the L1 materialization that no reader uses — its three child aggregates read it by bucket range, the coverage probe and arming gate read min(bucket), retention drops chunks. EXPLAIN (ANALYZE, BUFFERS, WAL) of one bucket's materialization INSERT: 45.5 MB WAL / 483,689 records / 1.64 M buffer touches / 6,040 dirtied with them; 10.7 MB / 72,735 / 447 K / 56 with the bucket index alone. L1 is now created without them and EnsureIntervalDedupMaterializationIndexesAsync drops them on an existing store at startup, per index under a 10 s lock_timeout so a refresh in flight is yielded to rather than convoyed. The capture-down read was the #3496 shape #3496 named and left: ROW_NUMBER() OVER (ORDER BY log_id DESC) over every collection_log row the server had for two collectors — 9,573 buffers across 61 chunks (59 decompressed) per alert pass per server on the rig. It is now one chunk-orderable LIMIT 1 per collector: 14 buffers, the newest chunk only, 120 of 122 chunk scans never executed. Partial: the dirty-bucket multiplier is the backfill's designed behaviour and is not changed here; the production dirty-bucket count per hour is the read that decides how much of the 417 s average this removes, and the exact statements for it are on the issue. * IntervalDedupMaterializationIndexesTests records why it is not serialized against the live collection: its gated arm mints a scratch database, so it cannot race the shared store (#1776 own-store, the census's own guidance)
…y at all and its empty chart reads as a broken product: the collector grows a service-side pg_stat_activity sampler arm as the floor, the three wait tiers become one connect-time decision, and every wait read discloses its instrument (#3604) (#3645) * A stock PostgreSQL target without pg_wait_sampling gets no wait accumulation at all, so its empty wait chart reads as a broken product: pg_wait_sampling grows a service-side pg_stat_activity sampler arm as the floor under stock targets, the three wait tiers are one connect-time decision, and every wait read discloses which instrument fed it (#3604) The pg_wait_sampling collector forks on a new connect-time fact, CollectorTargetInfo.HasPgWaitSamplingExtension: with the extension it reads the 10 ms profile as before; without it, it runs one multi-statement command of thirty one-second pg_stat_activity snapshots (each behind pg_stat_clear_snapshot, because the view is cached per transaction) and accumulates the tallies across cycles in its own collector_state, so the rows it writes are cumulative and the existing newest-minus-oldest read differences them unchanged. Which arm ran is recorded in state; get_pg_wait_sampling, get_pg_wait_stats and the Viewer panel disclose it as instrument: engine_cumulative | extension_sampled | service_sampled, with the floor caveat on the service tier. Aurora is gated off pg_wait_sampling (it cannot preload the module; the hourly EXTENSION_MISSING there was a permanent gap dressed as a precondition), and CoveredInsteadBy now points an Aurora caller at pg_wait_stats. The collector moves from hourly to five minutes and is detached from the sequential sweep body beside query_store and plan_correction, because a deliberate 30 s window inline would delay every other collector on the target; the window is 30 s and not the cycle because SweepPressureClassifier sums single-run costs against a 60 s body budget. No migration: same table, same columns, same collector name. The sampler needs nothing beyond pg_monitor. * Three pins encoded 'Aurora is a strict superset of what the PostgreSQL collectors read'; pg_wait_sampling is now the one documented exception (the module cannot be preloaded there), so the kind-axis split, the pre-flight probe line and the compose measure census each name it rather than assume it away; the empty-arm helper moves above the JSON builder's doc block it had split (#3604) * The sampler's pg_sleep argument is formatted under InvariantCulture, as every other number this file puts into SQL or state is: harmless at a whole-second period, a comma-decimal host's pg_sleep(0,5) the day it is not (#3645 review) * The extension arm clears the sampler's tally every cycle it runs: the host persists only PendingState keys, so a tally left alone would sit in collector_state for as long as the extension stayed installed and be resumed as a months-old baseline the day the target fell back to the sampler arm — splicing two eras into one series when the extension's own profile had just been reset (#3645 review) * The service-tier note claimed a series resets when the service restarts; the tally is persisted collector state reloaded every cycle, so it survives one — the note now names what actually starts a series over (cap eviction, arm reversion) and a pin keeps the false claim from returning (#3645 review) * Two Lite pins catch up: the state-declaring set is two collectors now, and get_pg_wait_stats' instrument_note naming aurora_stat_system_waits() is accounted for as PROSE in the Aurora-surface allow-list (#3604) --------- Co-authored-by: erikdarlingdata <erik@erikdarling.com>
…plan maximum kept as a dated history, so a MAXDOP-1 instance stops reporting 16 from a plan compiled before the pin (#3648) (#3651) * Top Cpu Queries' Max Dop is the newest plan's reading with the cross-plan maximum kept as a dated history, so a MAXDOP-1 instance stops reporting 16 from a plan compiled before the pin (#3648) * Parity guard moves its cross-SKU comparison to Darling.Tests (the filter that reaches both trees), the live test cleans up through LiveStoreCleanup, and the DuckDB fixture reads UtcNow once so a round-tripped timestamp compares equal (#3648)
…er seen, a regression with no baseline stays null, and the first point of a differenced trend is no longer a fabricated 0 (#3541 A12) (#3642) * Zero is a measurement: health parsers say whether their source was ever seen, a regression with no baseline stays null, and the first point of a differenced trend is no longer a fabricated 0 (#3541 A12) Six MCP sites published a 0, an empty, or a nominal-window label where nothing had been measured. Eight of the nine get_health_parser_* tools answered a dead system_health session with the same empty a healthy hour earns; now every one carries source_observed / last_captured_at and climbs the four-rung ladder significant_waits alone had, with nothing-ever-captured answering unavailable. get_query_store_regressions kept a NULL percent NULL (a 0 baseline has no ratio) with the reason under undefined_percents, and nulls a severity banded from a ratio that does not exist. The LAG-differenced duration trends (three raw reads, the Query Store rollup builder, Lite's four) leave the window's first collection unrated - null, counted in unrated_points, explained in unrated_note - instead of 0.0. get_pg_xmin_horizon divides a holder's wins by every capture in the window (collection_log SUCCESS rows, the alert evaluator's own denominator) rather than by the source's own rows, and answers unavailable when the collector never looked. get_pvs_stats divides any MEASURED size, so 0 MB is 0.00% and only an unmeasured numerator or absent denominator yields null, with pvs_measured and the reason. get_table_index_sizes projects its baselines raw and derives each growth figure from exactly the baseline it names, publishing history_days_available and growth_over_available_history_* where the nominal windows are out of reach, plus tables_returned / truncated. * Compose with #3630 (V128 / v61): the procedure trend keeps its stored-interval derivation and the MCP reader keeps the NULL-rate row as an unrated point rather than dropping it; the SQL-mirror pin's comment lines move to the C# doc; the #3540 pins that expected the first snapshot absent now expect it unrated * CI round 1: rung 3's closing sentence no longer contains the word the sibling rungs are pinned NOT to say; the Lite trend parity test's lone in-window collection is asserted unrated rather than compared as two zeros; the eight expression-bodied growth derivations join KnownTruncatedRanges (they strand no literal) * Re-cut head: the Claude review posted nothing on 1deff83 across five attempts (the #2229 swallowed-output class); a fresh head gives the review and the guard fresh runs. No code change.
…ly tools end the denial churn, the job fails when no verdict was submitted, and the transcript survives the runner (#3650) The prompt told the reviewer the branch was checked out and asked for a correctness, parity and security review, while --allowedTools permitted only four gh verbs and the inline-comment tool. Every Read, Grep, Glob and git call was a permission denial; the swallowed runs' result blocks read 50 turns / 18 denials, 39 / 21, 18 / 17, each ending subtype=success with nothing posted. Three PRs, about thirteen paid runs in one night. - --allowedTools gains Read, Grep, Glob, Bash(git diff:*), Bash(git log:*), Bash(git show:*). Nothing that writes. - A step after the action fails the job when claude[bot] submitted no non-empty-bodied review since the run's own start stamp, and prints the transcript's result block. A refusal-to-run (this file differing from the default branch's copy) is named separately. - claude-execution-output.json is uploaded as the claude-review-transcript artifact, 7 days. Lands on main and dev together: claude-code-action refuses to run whenever this file differs from the default branch's copy, so a dev-only edit would disable review on every PR until the next release.
This was referenced Sep 18, 2026
erikdarlingdata
marked this pull request as draft
September 18, 2026 22:49
erikdarlingdata
marked this pull request as ready for review
September 18, 2026 23:18
erikdarlingdata
enabled auto-merge (squash)
September 18, 2026 23:18
Owner
Author
|
Closing unmerged: this pointed the dev-cut lane branch at main, so it carried 69 commits — all of dev's unreleased work — where the pair strategy wants exactly one. Replaced by #3662, a main-cut branch cherry-picking the same workflow commit (c39de4c), byte-identical to #3657's dev half. Merge #3662 + #3657 together to keep the drift-arm window closed. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The lie
A Claude review run that posted nothing — no comment, no inline note, no verdict — finished green. Tonight that happened on three PRs across roughly thirteen paid runs (#3618, #3642 ×6, #3646); the guard (#2229/#3492) caught each one after the fact, and each cost a run, a guard cycle and a human's time. One PR was drafted to clear the board instead of merged on its earlier clean LGTM.
The mechanism (verified from the swallowed runs' own result blocks)
50 / 18, 39 / 21, 18 / 17 — denials ≈ turns. The prompt says "The PR branch is already checked out in the working directory" and asks for a correctness / parity / security review, an invitation to read code, while
--allowedToolspermitted onlymcp__github_inline_comment__create_inline_commentand fourgh prverbs. NoRead,Grep,Glob, nogit. Every attempt to open a file the diff touches, find a symbol's other callers, or check a parity twin was refused. After enough refusals the session ends —subtype: success— without ever reaching the VERDICT PROTOCOL. Stochastic: a small diff can get lucky fromgh pr diffalone; a 37-file diff did not, five times on one head.The change — one file,
.github/workflows/claude-review.yml--allowedToolsgains read-only tools:Read,Grep,Glob,Bash(git diff:*),Bash(git log:*),Bash(git show:*). Nothing that writes; the reviewer still cannot edit, commit, push, or post outside the gh verbs already listed. The comment aboveclaude_argscarries the denial-ratio evidence.date -u +%FT%TZ; a step after it (if: always() && token present) countsclaude[bot]reviews on the PR withsubmitted_at >= stampand a non-empty body, viagh api repos/…/pulls/N/reviews. Zero →exit 1, naming Claude review runs finish green having posted nothing: the allowlist forbids every read tool the prompt invites, denial churn ends sessions without a verdict, and the job never checks that it posted #3650 and [CI] Review workflow can execute fully (real token spend) and post nothing - green check with vanished output #2229, with the transcript's result block (num_turns,permission_denials_count, denied tool names) in the log. The body test is load-bearing and came out of the dry run: on The interval-hourly refresh recomputes one fleet-wide bucket per dirty hour and paid twelve index inserts per row for eleven indexes nothing reads; the capture-down alert read decompressed a server's whole collection_log to learn two statuses (#3597, partial) #3647 the bot has four reviews, two of them the empty-bodied carrierscreate_inline_commentfiles each note inside — a verdict (gh pr review --comment|--request-changes --body) always has a body (gh and the API both refuse a bodiless one), so "posted inline notes but no verdict" (the Give fleet sweeps a place to remember: the sweep-state store (#3466, lane 1 of 4) #3470 shape) fails here too. Because the prompt already mandates anLGTMreview on a clean run, zero verdicts is never legitimate on this repo — a clean review is never silent — so there is no false-red on a clean one-liner.execution_fileonly after running): "ran and posted no verdict" (Claude review runs finish green having posted nothing: the allowlist forbids every read tool the prompt invites, denial churn ends sessions without a verdict, and the job never checks that it posted #3650) vs "never ran — this file differs from the default branch's copy" (cause 1 in the guard's header). Bothexit 1; both are defects here. The step stands down (::notice) when the review step's own outcome is notsuccess(a failed step already reds the job; cancelled is a superseded push), and on agh apifailure it warns UNCONFIRMED and exits 0 per [CI] Review guard hard-fails on an unguarded lookup 404 — a lookup failure becoming the verdict its own comments forbid #2309 — a transient API error must not force a paid re-run of the whole job; the guard's own tally is a separate, free-to-rerun workflow and stays the enforcement.claude-execution-output.jsonuploaded asclaude-review-transcript,retention-days: 7,if-no-files-found: ignore. The next swallowed run is diagnosable in one click instead of by inference from step-log result blocks.Not changed: the prompt, the verdict protocol, the model, the step name
Claude review(the guard'sREVIEW_STEPcoupling),fetch-depth, permissions (pull-requests: writealready covers the read;GH_TOKEN: ${{ github.token }}).Why this PR targets
main, and why a twin targetsdevclaude-code-actionrefuses to run — server-side, during the OIDC token exchange, exiting success in ~4 s with no output — whenever the branch'sclaude-review.ymldiffers from the default branch's (main). So adev-only edit of this file disables review on every PR until the next release, and the guard's drift arm (#2229, kept by #3492) then hard-fails every non-editing PR — correctly. Lifting the guard would not help; the refusal is upstream of it.mainanddevhold byte-identical copies of this file today. So: this PR lands onmainFIRST, so thatdev's twin does not drift; thedevPR is #3657 (same commit, same bytes). Sequenced by hand, no auto-merge on either. The window in which otherdevPRs' reviews refuse runs from this merge to the twin's merge — minutes — and is stamped on #3650.Expected oddities on this PR, all non-required on
main(build+Darling PostgreSQL testsare the required checks):Check pull request target branchis red by design (it permitsmainPRs only fromdev— this is the sanctioned exception for the file that cannot land any other way),Check version bumpskips (head_ref != dev), and the review itself refuses to run on this very PR — cause 1 in action on the PR fixing it — so the new post-step will red the (non-required) review job with "Claude never ran", and the guard reports "Expected: this PR edits the workflow, a human must review it." A human has: the diff is 150 lines, all in one workflow, all of it read-only tools, a count, and an upload.Verified
actionlint(with shellcheck) clean on the file;yaml.safe_loadok;ALWAYS POSTsentinel andClaude reviewstep name unchanged.failure→ notice/exit 0; past stamp on The interval-hourly refresh recomputes one fleet-wide bucket per dirty hour and paid twelve index inserts per row for eleven indexes nothing reads; the capture-down alert read decompressed a server's whole collection_log to learn two statuses (#3597, partial) #3647 → 2 verdicts (the two bodiless carriers correctly excluded)/exit 0; stamp between the carrier and the final verdict → 1/exit 0; future stamp + transcript → ran and posted no verdict, result block printed, exit 1; future stamp, no execution file, no transcript → Claude never ran, exit 1; nonexistent PR → 404 → UNCONFIRMED warning, exit 0.Ref #3650 (closed by the
devtwin).