Skip to content

A Claude review that posted nothing no longer finishes green: read-only tools end the denial churn, the job fails when no verdict was submitted, and the transcript survives the runner (main copy, lands first) - #3656

Closed
erikdarlingdata wants to merge 69 commits into
mainfrom
fix/3650-lane
Closed

erikdarlingdata wants to merge 69 commits into
mainfrom
fix/3650-lane

Conversation

@erikdarlingdata

@erikdarlingdata erikdarlingdata commented Sep 18, 2026 •

Copy link
Copy Markdown
Owner

The lie

A Claude review run that posted nothing — no comment, no inline note, no verdict — finished green. Tonight that happened on three PRs across roughly thirteen paid runs (#3618, #3642 ×6, #3646); the guard (#2229/#3492) caught each one after the fact, and each cost a run, a guard cycle and a human's time. One PR was drafted to clear the board instead of merged on its earlier clean LGTM.

The mechanism (verified from the swallowed runs' own result blocks)

"type": "result", "subtype": "success", "is_error": false,
"num_turns": 50, "permission_denials_count": 18

50 / 18, 39 / 21, 18 / 17 — denials ≈ turns. The prompt says "The PR branch is already checked out in the working directory" and asks for a correctness / parity / security review, an invitation to read code, while --allowedTools permitted only mcp__github_inline_comment__create_inline_comment and four gh pr verbs. No Read, Grep, Glob, no git. Every attempt to open a file the diff touches, find a symbol's other callers, or check a parity twin was refused. After enough refusals the session ends — subtype: success — without ever reaching the VERDICT PROTOCOL. Stochastic: a small diff can get lucky from gh pr diff alone; a 37-file diff did not, five times on one head.

The change — one file, .github/workflows/claude-review.yml

  1. --allowedTools gains read-only tools: Read, Grep, Glob, Bash(git diff:*), Bash(git log:*), Bash(git show:*). Nothing that writes; the reviewer still cannot edit, commit, push, or post outside the gh verbs already listed. The comment above claude_args carries the denial-ratio evidence.
  2. The job fails when it posted no verdict. A step before the action stamps date -u +%FT%TZ; a step after it (if: always() && token present) counts claude[bot] reviews on the PR with submitted_at >= stamp and a non-empty body, via gh api repos/…/pulls/N/reviews. Zero → exit 1, naming Claude review runs finish green having posted nothing: the allowlist forbids every read tool the prompt invites, denial churn ends sessions without a verdict, and the job never checks that it posted #3650 and [CI] Review workflow can execute fully (real token spend) and post nothing - green check with vanished output #2229, with the transcript's result block (num_turns, permission_denials_count, denied tool names) in the log. The body test is load-bearing and came out of the dry run: on The interval-hourly refresh recomputes one fleet-wide bucket per dirty hour and paid twelve index inserts per row for eleven indexes nothing reads; the capture-down alert read decompressed a server's whole collection_log to learn two statuses (#3597, partial) #3647 the bot has four reviews, two of them the empty-bodied carriers create_inline_comment files each note inside — a verdict (gh pr review --comment|--request-changes --body) always has a body (gh and the API both refuse a bodiless one), so "posted inline notes but no verdict" (the Give fleet sweeps a place to remember: the sweep-state store (#3466, lane 1 of 4) #3470 shape) fails here too. Because the prompt already mandates an LGTM review on a clean run, zero verdicts is never legitimate on this repo — a clean review is never silent — so there is no false-red on a clean one-liner.
  3. The transcript survives the runner: claude-execution-output.json uploaded as claude-review-transcript, retention-days: 7, if-no-files-found: ignore. The next swallowed run is diagnosable in one click instead of by inference from step-log result blocks.

Not changed: the prompt, the verdict protocol, the model, the step name Claude review (the guard's REVIEW_STEP coupling), fetch-depth, permissions (pull-requests: write already covers the read; GH_TOKEN: ${{ github.token }}).

Why this PR targets main, and why a twin targets dev

claude-code-action refuses to run — server-side, during the OIDC token exchange, exiting success in ~4 s with no output — whenever the branch's claude-review.yml differs from the default branch's (main). So a dev-only edit of this file disables review on every PR until the next release, and the guard's drift arm (#2229, kept by #3492) then hard-fails every non-editing PR — correctly. Lifting the guard would not help; the refusal is upstream of it.

main and dev hold byte-identical copies of this file today. So: this PR lands on main FIRST, so that dev's twin does not drift; the dev PR is #3657 (same commit, same bytes). Sequenced by hand, no auto-merge on either. The window in which other dev PRs' reviews refuse runs from this merge to the twin's merge — minutes — and is stamped on #3650.

Expected oddities on this PR, all non-required on main (build + Darling PostgreSQL tests are the required checks): Check pull request target branch is red by design (it permits main PRs only from dev — this is the sanctioned exception for the file that cannot land any other way), Check version bump skips (head_ref != dev), and the review itself refuses to run on this very PR — cause 1 in action on the PR fixing it — so the new post-step will red the (non-required) review job with "Claude never ran", and the guard reports "Expected: this PR edits the workflow, a human must review it." A human has: the diff is 150 lines, all in one workflow, all of it read-only tools, a count, and an upload.

Verified

Ref #3650 (closed by the dev twin).

erikdarlingdata and others added 30 commits September 18, 2026 03:49
get_top_queries_by_cpu and get_top_procedures_by_cpu ordered by summed
delta_elapsed_time in both SKUs, so on a wait-bound server the real CPU
consumers could be missing from the page entirely and attributed_cpu_ratio
read as "hidden CPU" when it meant "wrong sort key".

All eight ranking sites now key on CPU: Darling's TopQueriesSql,
TopQueriesByHostObjectSql, and TopProceduresSql (CTE cut + post-WAITFOR-trim
outer sort) and Lite's twin reads. The viewer's Duration grids keep their
elapsed ranking by design.

Tests: SQL-text ordering pins for all three Darling consts (plus the rollup
const joins the Postgres-dialect theory), and behavior tests in both SKUs
seeding a CPU king whose elapsed time ranks last with more groups than the
over-fetch admits — under the old key it was cut, not merely mis-sorted.
Red-watched: the Lite cut test fails against the old ordering.

Fixes #3523


Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…-clear (#3524) (#3553)

The analysis pipeline's data-span gate measures LIFETIME history, so a
server whose collection died (broken credential, unreachable target)
still passes it, collects zero facts over the empty window, and the MCP
tool rendered that as status "empty" with "All metrics are within
normal ranges" — an affirmative all-clear for a dead collector.

Both SKUs' analysis services now set WindowEmptyMessage (beside
InsufficientDataMessage) at the zero-facts branch, and both analyze
tools return the same "unavailable" envelope get_analysis_facts already
uses for this state, pointing the caller at get_collection_health. The
true-negative all-clear is unchanged for runs that collected facts and
found nothing. The unavailable envelope keeps the #2506 persistence
hints, so anchored empty-window runs still disclose non-persistence.

Tests, both SKUs, on the REAL 24h gate: 25h of history whose newest row
is five hours old -> unavailable; the same server with one benign wait
inside the window -> the existing all-clear. Darling's is gated
(DARLING_TEST_PG) and ran green against a live PG 18.4 + TimescaleDB
store; Lite's runs on the shared DuckDB fixture.

The viewer Recommendations tabs render the same false all-clear from
the same state and only branch on InsufficientDataMessage; that half is
#3551 (out of this fix's file scope).

Fixes #3524


Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…k worst slot by severity, and stop spelling unknown growth as measured zero (#3554)

get_pg_autovacuum_health (#3534): the reader ranks by GREATEST(dead ratio,
insert ratio) but severity came from the dead ratio alone, so an
append-only worst_table ten times past its INSERT threshold read "ok".
Severity now comes from the same GREATEST the ranking uses, the
insert-side ratio is published as insert_threshold_ratio (mirroring the
SQL CASE's -1 sentinel handling), a one-sample window nulls
dead_tuples_growing and dead_tuple_change instead of claiming "flat"
(first_seen_at published so the caller can see the window), and the
page-scoped summary counts gain the sibling limit_reached discriminator
with past_threshold_count reading both axes.

get_pg_replication_slots (#3535): worst_slot was the fattest slot, not
the worst-classified one — an active 45 GB keeping-pace slot ("ok")
headlined over an inactive 2->8 GB grower ("critical_orphan_filling_disk").
Slots now order by a severity Rank (the wraparound tool's pattern) with
size only as the tiebreak; growth is null when either endpoint is the -1
sentinel or the window holds one sample, with unknown-growth Classify arms
(warning_retaining_wal_growth_unknown / info_inactive_growth_unknown) so
a sentinel can never read as stable; and raw retained_wal_bytes nulls the
sentinel as its _gb sibling always has.

Tests: Classify unknown-growth arms and the Rank ladder pinned (including
that every emittable severity holds a rung above ok); two new gated live
classes drive both tools end to end against a real store — the
severity-vs-size worst pick, the insert-driven worst_table's severity,
sentinel and one-sample growth reading unknown, and limit_reached.

Fixes #3534, fixes #3535


Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…s column, caveat-free RCSI advice (#3550)

* Fix three honesty defects: null the fake granted-memory field, read the dropped physical-reads column, mirror the RCSI caveats

#3529: get_memory_trend shipped a hardcoded total_granted_mb = 0.0 on both
SKUs — an agent investigating memory-grant pressure read "granted was 0 all
window" and ruled out the real culprit. The field is now an explicit null,
the envelope carries a granted_note naming get_memory_grants as the grants
series' source, and both tool descriptions stop promising granted memory.
Pinned by payload tests (Lite real-DuckDB, Darling gated-live) and
per-SKU description pins.

#3530: Lite's QueryStore time-slice reader mapped TotalReads AND
TotalLogicalReads to ordinal 4 (the logical aggregate — a deliberate alias,
matching the QueryStats slicer) but never read ordinal 6, so the SELECT's
total_physical_reads was computed and dropped. TotalPhysicalReads now maps
ordinal 6, pinned by a fixture whose I/O columns carry pairwise-distinct
values — equal values are exactly how the slip stayed invisible.

#3531: the LCK_M_IS advice handed out the RCSI ALTER calling it "strictly
better" with none of the caveats its LCK_M_S twin carried. Both twins now
name the brief exclusive lock the ALTER takes, the tempdb version-store
cost (phrased after FactRiskDisclosure's RCSI risk items), and the
test-on-a-copy warning, pinned by a shared content test.

Fixes #3529, fixes #3530, fixes #3531

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx

* Update the two doc lines that still described total_granted_mb as an always-0 parity placeholder

Boundary extension approved by the wave lead: doc lines going stale because
of the #3529 change belong in the same PR. Darling/README.md's tool-catalog
line and DarlingTrendReader's MemoryTrendPoint doc comment now describe the
shipped behavior — explicit null plus a granted_note naming
get_memory_grants.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx

* QueryStore slicer: physical-reads sort plots the physical series, bars and overlay both

#3547, folded in per the wave lead. The grid's sorting handler mapped
"AvgPhysicalReads" to the LOGICAL series under a physical-reads label — a
forced workaround while the reader dropped the physical column, which the
#3530 fix ended. The mapping now targets TotalPhysicalReads with the
bucket.Value case to match, and the selected-row overlay follows: the
timeline read carries avg_physical_io_reads through the same deduped CTE
(no new dedup site), the point record gains PhysicalReads, and the overlay
selector plots it — otherwise honest bars would have drawn under an
elapsed-ms overlay, a new mismatch replacing the old label lie.

The distinct-per-column test now covers the timeline's read path too, so a
logical/physical swap on either the bars' or the overlay's source goes red.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…tion floor (#3537) (#3555)

Two edges escaped the identity-keyed persistence gate in opposite
directions. False fire: the identity fraction's denominator counts only
holder-bearing collections (the collector emits no rows when the horizon
is unheld), so the first holder after quiet hours read as 1 of 1 - 100%,
"chronic", off a single sample. False quiet: a horizon continuously
pinned past the age threshold by a parade of DISTINCT holders never
accumulates any single holder's fraction, so the alert never fired while
its own claim ("vacuum is reclaiming nothing cluster-wide") was true the
whole time.

The identity arm stays as designed - it is the arm that names the thing
to kill - and gains a 5-observation floor, the smallest denominator
whose majority test cannot be satisfied by fewer than three sightings.

The new horizon arm fires when the WINNING xmin_age sat at or above the
threshold in a majority of the window's real captures, holder identity
ignored. Its denominator comes from collection_log (this collector's own
SUCCESS runs, so quiet zero-row captures count) rather than from the
holder table - the cheapest honest "collections in the window", counting
the exact collector being fractioned instead of inferring cadence from a
cadence-mate table. The threshold reaches the SQL as a bind from the
evaluator's own constant so read and evaluation cannot drift.

A horizon-arm fire without a stable identity names the rotating-holder
pattern, still carries the latest holder's remedy, and subjects the
stable "rotating holders" sentinel so the host's per-subject cooldown
holds across rotations instead of paging once per parade member.

Tests: evaluator boundaries for both arms (1/1 and sub-floor totals no
longer fire, rotating holders fire under the sentinel subject, the
classic chronic holder still fires the identity arm with its original
wording, below-threshold never fires), plus a live-Postgres read test
that pins the new SQL's counting - ERROR runs, other collectors and
other servers stay out of the capture denominator; a below-threshold
winner stays out of the numerator while still counting as an
observation.

Fixes #3537


Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…io_timing (fixes #3533, fixes #3536) (#3558)

#3533: the queryid filter ran client-side over a fetched top-duration page
(limit * 10), so a plan ranked below the page was unfindable at any window
size - and the filtered-empty branch then told the caller capture was
working, no plan was captured, and it was "not the query to look at" while
the plan sat in the store. The predicate now travels into
DarlingPgPlanCaptureReader's SQL (($4::bigint IS NULL OR query_id = $4);
null leaves the top-duration page unchanged), the over-fetch and the C#
filter are gone (the Viewer's existing call keeps its overload), and the
miss text says what was actually searched - the whole window, server-side -
and what a miss can mean: never ran (get_pg_top_queries confirms), never
crossed auto_explain.log_min_duration, capture not working when it ran
(get_pg_plan_capture_readiness), or aged out of retention.

#3536: track_io_timing is OFF by default in PostgreSQL, and this read
rendered that as 0.000 ms latencies - an impossibly fast disk instead of
"not measured". It now mirrors the trend sibling's contract: the setting is
read from pg_server_config bounded by the window end (inferred from the
window's data when the config was never collected, and io_timing_source
says which), io_timing_tracked and timing_note are published, every
time-derived field (read_time_ms, avg_read_ms, write_time_ms,
extend_time_ms, the read-time shares, total_read_time_ms) is null when
untracked, and busiest_basis states what the ranking used - the reader's
secondary ORDER BY key (reads) is the entire ordering over a store of
zeros, and is now documented as load-bearing. busiest_by_read_time keeps
its key because the web tile reads it by name.

Tests: SQL predicate pins; both miss-text arms; BuildIoJson wire shape for
tracked, untracked, inferred, setting-overrides-observation, and the Aurora
null-write interplay; and a gated live fixture seeding 20 expensive shapes
plus a 21st cheap one - exactly the limit * 10 page the old path fetched -
proving the pinned read finds it, the top page does not contain it, and an
absent queryid gets the honest miss. Live parse analysis of every shipped
PG read green on PostgreSQL 18.

Fixes #3533, fixes #3536


Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…interval (#3527) (#3560)

The PERFMON_*_SEC facts, the batch-request anomaly window, and both SKUs'
batch-request baselines consumed delta_cntr_value — the per-COLLECTION-INTERVAL
delta — as if it were already per-second: 60x truth at a 60s cadence, 300x at
5min. All three consuming reads now divide by the interval consistently, in
both SKUs, so the window statistic, the baseline population, and the
requests/sec floors (BatchRequestFloor 500, BatchRequestFallback 5000) finally
meet in one unit:

- Fact collectors (Pg + DuckDb): select the row's measured
  sample_interval_seconds (#2234) alongside the delta, emit
  delta / interval as the fact value, and keep cntr_value/delta_cntr_value
  raw in the metadata with the divisor added. Rows with interval <= 0 (no
  delta was knowable: first sighting, reset, gap past the delta policy) are
  filtered so rn = 1 lands on the newest USABLE row — an unknowable rate is
  never emitted as 0 or as the raw delta.
- Anomaly windows (PgAnomalyDetector + Lite AnomalyDetector, Lite-verbatim
  SQL): AVG/MAX over delta * 1.0 / NULLIF(sample_interval_seconds, 0) with
  interval <= 0 rows excluded from the sample count.
- Baselines: Lite divides by the stored measured interval on the raw rows;
  Darling's arm reads the perfmon_baseline continuous aggregate, which
  materializes no interval column and cannot grow one without forfeiting the
  31 of 35 days of history the 4-day raw tier can't refill — so it derives
  interval_sec from LAG(collection_time) over the collapsed series, the
  established WaitMsPerSec idiom, with the io-arm's DOUBLE PRECISION cast.
  The restart-exclusion signature stays on the RAW delta in both SKUs (its
  > 1000 bar predates the division).

No stored baseline state needs migration: both providers compute baselines on
read (1-hour in-process cache only, cleared by the upgrade's restart), and the
CAGG stores per-collection raw deltas — unit-neutral facts — not statistics,
so all 35 days of materialized history serve the new unit immediately.

Tests, both SKUs: the interval-60/delta-6000 fixture produces a fact of 100;
interval-0 rows are skipped at every read (never rates of 0); the anomaly
floors demonstrably compare per-second values (a raw delta past the floor
whose rate is under it stays quiet); live-Postgres end-to-end proofs for the
fact collector and the detector+baseline pair, including the LAG-derived
baseline mean and a window that skips an interval-0 row.

Fixes #3527


Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…and count measured metrics in the fleet Healthy label (#3528) (#3562)

Three alert-semantics honesty fixes from the 2026-09 brains-review campaign:

1. DeadlockCountThreshold and BlockingCountThreshold now floor at
   PostgresAlertEvaluator.CountThresholdFloor on read, mirroring their
   V122 PostgreSQL twins, so a store row hand-edited to 0 can no longer
   make count >= threshold fire on a quiet server. The MCP write bounds
   name the same constant instead of a literal 1, and the PG twins' doc
   comments stop calling the gap pre-existing.

2. Store Disk Pressure gains a GB floor (V126: self_disk_free_warn_gb,
   default 50): the percent trigger additionally requires free space
   below the floor before firing, the pvs_floor_gb AND-composition, so
   400 GB free on a 4 TB store volume stops paging CRITICAL; 0 removes
   the floor. Full knob plumbing: darling.json field, clamped settings
   property, seed/read, get_alert_settings/update_alert_settings key
   parity, and the viewer schema-probe arm. The Settings window box is
   deliberately deferred to the viewer pass and pinned as an abstinence.

3. Fleet cards carry measured_metric_count/metric_count beside the band
   (the worst-of fold skips Unknown, so an online server with five of
   six metrics structurally Unknown still bands Healthy); the web fleet
   page renders the qualifier ("1 of 6 measured") on partially-measured
   cards. Rank-neutrality of Unknown is unchanged.

Fixes #3528


Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…licy (#3532) (#3559)

A delta-family collector scheduled above the CollectorDeltaCalculator gap
policy (DefaultMaxGapSeconds = 3600) takes the reset branch every cycle:
every delta is (0, 0), facts vanish, baselines fill with zeros, and the
product reads green precisely because it stopped measuring.

The bound lives once, in PerformanceMonitor.Collectors: the delta family
set (the calculator's exact caller census - 8 SQL Server + 2 PostgreSQL
collectors, pinned by a source-scan test) and MaxDeltaFrequencyMinutes,
half the policy so even an entirely missed cycle still yields a real
delta (a cadence AT the policy already fails every cycle, since the gap
is the cadence plus scheduling latency and the reset comparison is
strict).

Enforcement, per surface:
- Lite ScheduleManager: UpdateSchedule / SetScheduleForServer refuse
  (before mutating), and the load path clamps a hand-edited or pre-fix
  collection_schedule.json with a warning, since there is no user to
  bounce it back to.
- Lite schedule editor: validates before save (mirroring the Darling
  editor, which Lite previously had no counterpart to), commits pending
  cell edits first, and shows the cap in a footer line built from the
  shared constants.
- Darling viewer editor: ValidateSchedule moved to the pure
  CollectorScheduleOverlay (now testable) and extended with the delta
  bound; footer hint appended from the same constants.
- Darling service: StoreConfigProvider.ResolveSchedule's ValidFrequency
  now knows the collector, so a poisoned store row (hand-written or
  pre-fix) degrades to "no override" and falls through, exactly like a
  negative frequency. The web/MCP command path only writes enabled
  flags, so it needed nothing.

Snapshot collectors keep long cadences - the bound is per-collector-kind.


Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…he false all-clear (#3551) (#3565)

The tabs only branched on InsufficientDataMessage, so an analysis window that
collected zero facts (dead collector, unreachable target) rendered "All clear -
no current recommendations." Both viewers now consume the WindowEmptyMessage
state #3553 added, one rung over from insufficient-data:

- Lite's Generate now reads AnalysisService.WindowEmptyMessage directly and
  renders a distinct notice pointing at the Collection Health tab.
- The Darling viewer never calls the engine, so the worker persists the
  window-empty determination into the V19 analysis_state marker as
  insufficient_data = false + the engine's message (a shape nothing else
  writes, so no schema change), and AnalysisStateMarker.WindowEmpty derives
  it back for the Recommendations tab on both the read and Generate now doors.

The genuine all-clear (facts measured, zero findings) is unchanged, and the
marker self-heals on the next facts-bearing pass - pinned through the real
viewer read against live Postgres.

Fixes #3551


Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…real data (#3566)

* get_memory_trend joins the grants series: total_granted_mb carries real data

Fixes #3548. #3529 shipped the honest placeholder — an explicit null with a
note naming get_memory_grants — because the memory_stats source has no grant
data. The complete fix joins the memory-grant series both stores already
collect (the same v_memory_grant_stats sum the viewers' Memory Overview
overlays plot) into the trend payload on both SKUs.

The join is nearest-match within 30 seconds, because each collector stamps
its own DateTime.UtcNow per run: same-cycle rows sit seconds apart, so an
equality join returns nothing, while a wider match would smear a slower
grants cadence across points it never measured. Zero-vs-unknown discipline
holds on both sides of the match: a snapshot measuring nothing granted is a
genuine 0.0, an uncovered point is null, and the granted_note survives only
in payloads that have a null to explain — a fully covered window is not
captioned with an apology.

Darling's read is a new DarlingTrendReader const pinned byte-identical to
the viewer's proven MemoryGrantTrendSql; Lite reuses its overlay read with
an asOfUtc pass-through so the grants window matches the memory window's.
Twin tests pin the pool-summed match, the genuine zero, the uncovered null,
note presence/absence, and the 30s boundary (in at 30, out at 31); the same
claims run live against Postgres, plus the byte-identical SQL pin.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx

* Reword a doc comment the as_of census reads as code

AsOfWindowAnchorTests string-scans anchored tool bodies for the literal
DateTime.UtcNow, and the tolerance constant's doc comment named it while
describing the collectors' per-run stamps. Both tools are genuinely
anchored (the joined grants read takes the same windowEnd as the memory
read); the comment now says "capture clock" and the census passes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…s tiers, not any-deadlock-is-Critical (#3564)

* The daily classifier bands deadlocks as a rate through the card band's tiers, not any-deadlock-is-Critical (#3525)

The shared DailyHealthBandCalculator still shipped #3368's un-fixed twin:
signals.Deadlocks > 0 banded the whole day Critical. On the measured
43-server fleet that read 87.9% of 24-hour windows Critical, so ~7 of 8
Performance Calendar day cells painted red from deadlocks alone, and the
fleet sweep's variable 15-1440min span scaled its Critical rate ~6x from
the cadence knob.

DailyHealthSignals now carries the Window its counts cover (the
ServerHealthMetrics.DeadlockWindow discipline: default(TimeSpan) is
unusable, never a divisor), and Classify routes the Deadlocks signal
through ServerHealthClassifier.DeadlockSeverity with the store-backed
DeadlockRateThresholds (V120) via DailyHealthThresholds.DeadlockRates —
Critical/Warning fold into the day's tiers, and the sub-1h/undeclared
window arm falls to Warning-not-rate (zero stays out of the trigger).
DeadlockRatePerHour/DeadlockSeverity widen int->long for the day-scale
counts; every int caller converts implicitly.

Producers declare their windows: the calendar-day projections (Darling
health reader, viewer, Lite) pass 24h, the fleet sweep passes its own
span. The Darling day reads hoist one store-tiers read per range/sweep
(the fleet reader's own hot-swap argument), the sweep's Compose takes
the tiers as a required parameter, and its would-have-paged deadlock
family now fires only when the rate crossed the Critical tier, with
rate + tier + count as evidence. The shared deadlock reason/tooltip
line reports the rate beside the count ("120 deadlocks (5.0/hr)"), the
card reason's own disclosure rule, and the sweep verdict JSON carries
deadlock_rate_per_hour + window_minutes additively.

Tests: two-window proofs on the day classifier (same rate bands the same
over 1h and 24h; same count bands differently), the single-deadlock day
control, the sub-hour arms, tier-override pins, sweep ledger pins, a
whole-repo census that every production DailyHealthSignals bundle
declares its Window (with the calendar projections pinned to 24h and
the sweep pinned to its span), and the Lite live-DuckDB calendar now
proves the rate path end-to-end with a 480-deadlock storm day.

Fixes #3525

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx

* Two live health-tool tests catch up to #3525's rate banding

DarlingMcpHealthToolsLivePostgresTests planted one deadlock and asserted the
day reads Critical - the any-deadlock-is-Critical semantics #3525 removes.
One deadlock across a 24h day is 0.04/hr, far below the measured tiers, so
the boundary test now pins the honest Healthy band plus the deadlock_count
evidence (its real subject is date-boundary visibility), and the planted-rows
test pins Warning (its 85% CPU sample trips the HighCpuWarningSamples tier)
plus the same evidence. The Critical-band rate path is covered by the PR's
own storm-day tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx

* The still-forming day clamps its window to elapsed time (#3525 review)

The calendar-day projections stamped Window = 24h on every row including
today's half-open bucket, so an active storm diluted against hours that
had not happened yet: 60 deadlocks in the last hour read 2.5/hr (Healthy)
instead of 60/hr (Critical) - a false-negative the pre-#3525 any-deadlock
trigger could not produce. The shared CalendarDayWindow helper now clamps
the current day to its elapsed portion against the read's own clock: the
anchored range reads (both SKUs) hand their resolved as_of window end so
a backdated read clamps against itself, the live calendar reads use the
wall clock, and a day minutes old falls to the sub-hour Warning-not-rate
arm rather than a fabricated multiplied rate. Finished days band over
their full 24 hours exactly as before, and the sweep path is untouched.
The window census now pins all three projections to the clamp helper so
a hand-rolled FromDays(1) cannot reintroduce the dilution, and the twin
band suites carry the storm proof at the review's own numbers.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
… Store slice-total labels say Total (#3567)

* Procedures slicer: physical sort plots the physical series, and slice totals stop calling themselves averages

Fixes #3556 — the #3547 bug's verified siblings, found during that fix and
kept out of its file boundary.

The procedures grid's physical-reads sort mapped to the LOGICAL aggregate
under a "Total Physical Reads" label, with no physical case in the value
switch — while the slicer's SELECT computed total_physical_reads at ordinal
6 and the reader dropped it on the floor, the #3530 shape exactly. The
reader now maps TotalPhysicalReads, and the handler mirrors the Query Store
fix: metric TotalPhysReads, the value-switch case, honest label. The
overlay side was already wired — ComputeProcOverlayPoints has carried a
dormant TotalPhysReads -> DeltaPhysicalReads arm all along, and the handler
re-triggers SelectionChanged on metric change, so bars and overlay go
physical together with no further threading.

The Query Store switch's two remaining dishonest labels — "Avg Reads" and
"Avg Writes" over execution-weighted slice TOTALS — now say Total, matching
the physical arm beside them and the query-stats grid's vocabulary.

Pinned three ways: a distinct-per-column reader test (110/30/50 — equal
fixture values are how the drop stayed invisible), and a source-text pin
over both sorting handlers' switch arms in the AvailabilityGroupsGridSortTests
style, since the handlers are private WPF event handlers.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx

* Darling viewer: port the #3547/#3556 grid fixes the review found unported

PR #3567's claude-review flagged that ViewerServerTab.Queries.cs — the
viewer's copy of Lite's sort-driven slicer metric mapping — still carried
both bug classes this PR fixes in Lite. The procedures grid's physical
sort mapped to the LOGICAL series with no physical value case, and the
Query Store grid had both the "Avg"-over-totals labels AND the un-ported
#3547 physical/logical swap. Both handlers now match Lite: TotalPhysReads
mapping, value case, honest Total labels.

No reader threading was needed anywhere — the shared slicer reader
(ReadQueryStatsSlicerAsync) has always mapped ordinal 6's
total_physical_reads, all three slicer SELECTs compute it, and the three
item-timeline reads already carry physical_reads into ItemTimelinePoint,
with ComputeOverlayPoints' TotalPhysReads arm sitting dormant — so only
the handlers were lying.

One genuine port gap beyond the switch arms: Lite's sorting handlers
re-run SelectionChanged after a metric change so a selected row's overlay
re-projects onto the new metric; the viewer's handlers never did, which
would have drawn the newly-honest physical bars under a stale-metric
overlay — the exact mismatch the #3550 commit warned about. All three
handlers gain the re-trigger.

Pins mirror Lite's: a source-text pin file over both fixed handlers plus
the re-trigger (the omission was per handler, so pinned per handler), the
slicer SQL shape theory now asserts the reads/writes/physical alias order
the shared reader's ordinals depend on, and the live slicer + dedup tests
gain distinct-per-column fixture values (Lite's 110/30/50 signature;
41x100/41x10 for the Query Store bucket and 40x100/40x10 for the timeline
point) — equal fixture values are exactly how a swap stays invisible.
Verified against a throwaway PG 18.4 + TimescaleDB cluster: 11/11 live,
48/48 ungated.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* CHANGELOG: the brains-review wave, one splice (17 entries)

Wave-1 of the 2026-09 brains-review campaign landed sixteen PRs on dev
with a buffered changelog protocol - agents reported their entries to the
coordinator instead of touching this file, so sixteen PRs merged without a
single CHANGELOG conflict. This is the one post-wave splice: seventeen
Fixed entries covering #3523-#3537, #3547, #3548, #3551, #3556, plus their
reference links. #3549 was documentation-only and carries no entry.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx

* CHANGELOG: fold the #3561 Dashboard divisor entry into the wave splice

PR #3569 landed after the splice opened; its entry joins the same PR to
keep the one-splice-per-wave discipline (18 entries now).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx

* CHANGELOG: fold the #3563 viewer-pass entry into the wave splice (19 entries)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…) (#3569)

Mirror of #3560 (Lite/Darling) into the deprecated Dashboard's frozen
analysis twins: cntr_value_delta spans one collection interval, not one
second, so the raw reads overstated by the cadence (60x at 60s, 300x at
5min) against thresholds defined in requests/sec.

All three reads move together, or the z-score compares across units:

- SqlServerFactCollector PERFMON_*_SEC facts divide the delta by the
  row's measured sample_interval_seconds; interval <= 0 rows (unknowable
  delta) are filtered so rn = 1 lands on the newest usable row - a
  counter with only interval-0 rows emits no fact, never 0. The raw
  delta and the divisor ride the metadata.
- SqlServerAnomalyDetector's batch-request window AVG/MAX divide by
  NULLIF(sample_interval_seconds, 0) with interval <= 0 rows filtered,
  keeping the window statistic in the same requests/sec unit as the
  BatchRequestFloor/Fallback bars.
- SqlServerBaselineProvider's batch_requests arm computes the baseline
  population as the per-second rate; the restart signature stays on the
  RAW delta (its > 1000 bar predates the division).

The window and perfmon SQL move to public consts (the Darling twin's
tested shape) and GetBaselineQuery goes internal so Dashboard.Tests can
pin the text - the Dashboard has no test database. Semantics verified
live against SQL Server via a temp-table clone of collect.perfmon_stats.

Fixes #3561


Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…oor, and the measured-metric qualifier on the WPF fleet card (#3571)

The V126 knob (self_disk_free_warn_gb) lands in the viewer: the Settings
window gets its box one knob over from the percent sibling (prefill, >= 0
save gate matching the MCP write bound and read-side clamp, Restore
Defaults at the shipped 50, master-switch follow), and
ViewerDataService.AlertSettings.cs appends the column to the select /
upsert / bind / reader at ordinal 66 ($67). The
SelfDiskWarnGbFloorRungTests abstinence pin flips to pin the wired state,
with the bind-order roundtrip and the Settings-window pins; V124's
end-anchored bind pin hands its top-of-bind claim on, the V122->V124
precedent.

The WPF fleet card also picks up the web fleet page's measured-metric
qualifier: ServerSummaryItem carries MeasuredMetricCount / MetricCount
through the shared classifier fold, and the card tooltip's band label says
"Healthy - 1 of 6 measured" for a band folded over unmeasured metrics
instead of claiming every metric is inside its threshold.

Fixes #3563


Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…cached roots from prior rotations can never break Windows chain building (#3557) (#3572)

Windows caches every root a TLS client processes into that user's
intermediate-CA store, one entry per store rotation, forever. With every
rotation minting a root under the same subject, the cached pile grows until
CryptoAPI's subject-matched issuer walk fails outright ("unknown chain
building error") - measured at ~50 cached roots on a dev box, killing
SslStreamCertificateContext.Create server-side in the #2117 handshake test
and the default-trust build a viewer's SslStream runs during connect. A
random mint tag in the root CN makes every rotation a genuinely distinct CA,
so cached copies of other rotations are never issuer candidates and the pile
never forms, on any machine. SKI/AKI does NOT prevent the engine error
(tested, matrix-controlled); subject uniqueness does (6/6 fresh-process runs
green against a ~55-cert pile left in place).

The test harness now returns the server-side exception it used to swallow -
"handshake completed = false" alone cannot distinguish an Npgsql rejection
from the server never reaching TLS at all, which is exactly the misread that
kept this failure looking like a certificate-validation verdict.

Fixes #3557


Claude-Session: https://claude.ai/code/session_01QMYFB4qv5tqGwMmdPsa4Zx

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
… same table name in two schemas no longer reads as one object (#3576) (#3583)

* The Locking & Contention grid names the table as schema.table, so the same table name in two schemas no longer reads as one object (#3576)

An Azure SQL DB user reported the FinOps Locking & Contention grid showing
the bare table name. On Azure SQL DB — and anywhere per-tenant or
per-environment schemas are the pattern — the same table name routinely
lives in several schemas, so dbo.Orders and archive.Orders rendered as
one indistinguishable "Orders". The store always carried schema_name (the
index drill's Identity panel shows it); only the grid dropped it.

The row model gains a computed FullName (schema.table, bare name when the
schema is empty — the ProcedureStatsRow.FullName idiom) and the Table
column binds it, so the qualified name sorts, filters, and exports as one
string. The header's filter Tag moves with the binding: the popup filter
resolves Tag to a property by reflection, and a Tag naming a property the
row lacks filters nothing, silently.

Ported to the Darling viewer's copy of the grid and to the deprecated
Dashboard's frozen twin (smallest faithful diff). Pinned in Lite.Tests
(property, filter mechanism through the real matcher, and both SKUs' grid
XAML from source) and Darling.Tests (the viewer's own row model).

* The Postgres Index Usage grid names the table as schema.table too (#3576, sibling surface)

Same defect class one tab over, Darling-only: the viewer's Postgres Index
Usage grid showed the bare table name while the five other table-naming
grids on that tab show Schema as a column. The store's reader carried
SchemaName from the start; it was dropped in two places — the display row
had no SchemaName member for the mapper to fill, and the grid bound
TableName. Both halves fixed: the display IndexUsageRow gains SchemaName
(mapper copies it) and a computed FullName (schema.table, bare on empty
schema), and the Table column binds FullName at width 200. No filter
buttons on this grid, so no Tag to move.

Pinned through the real PgDisplay.IndexUsage mapper — a pin on the
property alone would pass with the mapper still dropping the schema,
which is the exact shape that shipped — plus the grid binding from source.
Lite has no Postgres tab; the deprecated Dashboard has no such grid.
…, so the baseline engine is no longer notification-inert at shipped settings (Fixes #3526) (#3584)

* Extreme, corroborated anomaly findings can cross the notify floor (#3526)

Every ANOMALY_* ramp saturates its base at 1.0, the Layer-3 tuning-class cap
holds the final at 1.49, no amplifier arm existed, and the shipped notify floor
is 1.5 — so the baseline engine could never page at default settings.

FactScorer:
- IsExtremeAnomaly: a per-fact escape from the tuning-class cap for ANOMALY_*
  only, at 3x the cutoff the fact fired at (10.5σ robust, 15σ heavy-tail, 6σ
  classical; 3x the absolute bar on the low-quality fallback path, never for
  is_new wait profiles; ratio/count/delta families stay capped).
- AnomalyAmplifiers: a co-fire arm — sibling anomaly families and measured
  absolute facts corroborate at +0.2/+0.3, so an extreme anomaly with two
  corroborators scores 1.6 and a lone or routine one cannot reach 1.5.
- The impact-peer escape and CXPACKET's cap are unchanged.

Lite.Tests/FactScorerTests: 12 new pins on the escape bar arithmetic per path,
the co-fire arm, the routine/lone/single-corroborator negatives, the
low-quality sigma blindness, CXPACKET staying capped beside extreme anomalies.

* Bound the extremity escape bar at SigmaDisplayCap (review on #3584)

AnomalyGate clamps the stored deviation_sigma at 25σ before the scorer sees
it, so an operator-scaled anchor above ~8.3 made 3x unreachable and the escape
silently dead for that metric — the defect this PR fixes, for a tuned store.
IsExtremeAnomaly now takes min(3x anchor, SigmaDisplayCap); two Theory rows
pin anchor 10.0 at 25.0 (released) and 24.99 (capped).
… are visible to, and proves rows exist where it can (#3574) (#3585)

The block reported recording: true from the GUC (#3175) and stopped there. The
reader's real question is "will I see rows", and timescaledb_information.job_history
is a security_barrier view that shows a row only to a member of the job's owner role
or of the database owner (2.28.1 views.sql: pg_has_role(current_user, <datdba>,
'MEMBER') OR pg_has_role(current_user, owner, 'MEMBER')), while the jobs and job_stats
views a reader checks first are not filtered. A least-privilege role can list all 110
jobs and read an empty history on a store recording perfectly; that cost a real
postmortem.

The reader gains a second, failure-isolated statement (JobHistoryEvidenceSql) that
evaluates the view's own two pg_has_role tests for the connection doing the reading,
counts the rows it sees over a fixed 24-hour window, takes the population half from
the unfiltered job_stats, and publishes reader_role, visibility (All/Partial/None),
job counts, rows_observed, newest_row_at, jobs_run_in_window, newest_run_started_at
and a contradiction flag. The note names the rule with the predicate, the role it read
as, and what that role may see; recording on + reader admitted + jobs ran + zero rows
is the new CONTRADICTION arm, with its one benign cause and how to settle it.

Evaluating the predicate is load-bearing rather than decorative: in managed mode the
MCP host connects as the mcp role, which the view filters OUT, so a bare count would
have manufactured the very false contradiction the issue is about on every managed
store. The block now says "None — the filter, not the table" there and names the role
that can see; a bring-your-own owner connection is the self-proving case.

Bind is an explicit timestamptz with Kind=Utc — the inverse of the store's naive
discipline, because these are TimescaleDB's own TIMESTAMPTZ catalog columns.
…er will actually take, instead of walking the whole fleet's two-hour slice per server (#3573) (#3586)

The read's (server_id, collection_time) composite was on the production
store all along — V1 generates it — and the planner priced it out:
server_id's physical correlation is 0.022, so the cost model charged one
random page per row and preferred streaming the entire fleet's two-hour
slice through the time index (Rows Removed by Filter: 691,058 to keep
37,878; 57,307 buffers; 422 ms warm, 10.3 s cold against a 10 s deadline).
Forced under random_page_cost = 1.1 the same statement took the existing
composite at 5,063 buffers, so a second plain composite would have been
priced and ignored identically.

Covering removes the heap term the model got wrong: (server_id,
collection_time DESC) INCLUDE every other column the read touches, so it
plans as an Index Only Scan over one server's rows (rig: 50 buffers vs
1,514, Heap Fetches: 0). It lives in PgTableTuning, not the ladder —
results-invariant perf, autocommit, failure-isolated — as a plain CREATE
INDEX: on TimescaleDB 2.28.1 compressed chunks get an 8 KB shell so only
the live chunks carry it; transaction_per_chunk was measured leaving an
invalid parent index after a mid-build cancel that IF NOT EXISTS then
skips forever; CONCURRENTLY is refused on hypertables.

Tests pin the index's column list against the read's own column
references, the statement text against that list, the two rejected
shapes, and (DARLING_TEST_PG) EXPLAIN the shipped statement on a seeded
store for the Index Only Scan. The 1,744.9 ms deadline anchors keep
their history and gain the arc; the 10 s deadline does not move.
…he toasts on the next poll instead of waiting on the service's reload (#3570) (#3587)

The tray toast is the viewer's own alert channel, and until now the
viewer applied no mute rule to it: a polled alert row toasted unless
the SERVICE had stamped it muted. A Snooze wrote the rule to the
store and bumped the reload beacon, and then the toast's suppression
depended on a four-link chain in another process - beacon observed on
the next 15 s tick, a monolithic config re-read succeeding in full,
the mute cache refreshed, the alert re-firing THROUGH that cache -
none of which the viewer observed. When a link was slow or broken the
rule sat in Manage Mute Rules while the toasts kept coming.

AlertToastCoordinator.SelectToasts now also takes the viewer's own
rule set and skips any row a rule covers, judged with the shared
MuteRule.MatchesAt over the row's own ToMuteContext (server + metric
as the row spells them, plus the detail-text dimensions the Mute This
Alert pre-fill already read). The set is the per-poll mute-rule read
that already drives the sidebar bell, plus any rule this viewer just
wrote (tray Snooze, server Silence) added persist-then-cache. Keep-
last-known on a failed read, and a rule written mid-read is carried
over so the guarantee has no one-tick hole. The status line says what
happens on each clock, and a failed snooze is reported rather than
only logged. Lite never had the gap: its snooze lands in the same
in-process MuteRuleService its deliverer consults before the balloon.
…, and samples off the policies' run instant (#3575) (#3588)

TimescaleDB's job_stats view assembles job_status from pg_stat_activity and
next_start from the bgw_job_stat row, so at both edges of every healthy run
one SELECT reads -infinity AND Scheduled - the predicate's dead-job arm. A
production store paged on exactly that, 53 ms into a 63 ms run that succeeded.

- ReadStuckCompressionJobsAsync re-reads StuckCompressionJobsSql five seconds
  after a -infinity trip and reports the job only if the arm still trips; a
  failed confirm defers judgement to the next hour rather than paging. The
  stuck-Running arm is reported from the first read as before. Injectable
  seam + pure merge (ClassifyCompressionJob / ConfirmStuckCompressionJobs) so
  the decision table pins without sleeping.
- The worker snaps each hourly due time to :30 past the minute
  (TimescaleSupport.NextCompressionCheckUtc) instead of UtcNow + 1 h, which
  slipped a few seconds an hour across the :MM:00 instants the fixed-schedule
  policies fire on.
- Both edges captured on a PG18 + TimescaleDB 2.28.1 rig; the crashed-worker
  state (the persistent -infinity row) reproduced and still reported.

Tests: CompressionStuckConfirmReadTests (new), TimescaleSupportTests pins
annotated. Census counts unchanged.
…nOps row marks become theme brushes Dark can actually see (#3577) (#3589)

Two contrast defects from a community report, both measured rather than
eyeballed (WCAG relative-luminance contrast; AA wants 4.5:1 for text, 3:1
for a non-text marker).

The selected tab on Cool Breeze read 2.71:1 - and not because the theme
chose that. Every TabItem style set a light Foreground for the selected
state, but a string header becomes a TextBlock, and the themes' app-level
implicit TextBlock style out-ranks the Foreground that TextBlock inherits
from the TabItem, so the setter never reached the text: the header always
rendered ForegroundBrush on the accent. The same dead setter meant Dark's
selected tab was 1.99:1 (visible in the README's plan-viewer screenshot)
and Light's would have been light-on-cyan had it ever applied. An empty
implicit TextBlock style inside the header presenter shadows the app-level
one, so the header inherits the TabItem's Foreground the way stock WPF
intends, and a new per-theme AccentForegroundColor/Brush is the ink for
text on any accent fill: white on Cool Breeze (5.39:1), the darkest
surface tone on Dark (7.52:1), the page text on Light (6.79:1, no change).
The same ForegroundBrush-on-accent pair recurred on the highlighted combo
item, the accent button, the selected calendar day and Dark's grid
selection (all 1.99:1 on Dark, 2.71:1 on Cool Breeze); they take the same
ink through the same key. Cool Breeze's AccentHoverColor moves one step
toward the accent (#2B87C8 -> #267BB8) because no ink passed on the old
shade (white 3.89:1); white is 4.56:1 on the new one. TabCloseButton drops
its hard-coded White so the x inherits the header ink - it was invisible on
the light themes' unselected tabs.

The FinOps "Mark Done / To Do / Do Not Do" tints were three fixed 20%
overlays in DataGridRowMarks, which composite to 1.2-1.4:1 against Dark's
near-black rows - the operator could not find the rows they had marked.
They are theme brushes now (RowMarkDoneBrush / RowMarkToDoBrush /
RowMarkDoNotBrush, painted by resource reference so a marked row follows a
theme switch). Light and Cool Breeze keep the shipped tints; Dark builds
its own from the theme's Success / Warning / Error colors at the opacity
that pushes the mark as far from the row as it can go while the row text
holds >= 4.5:1 on both row backgrounds (marks 2.9-3.1:1, text 4.5-5.1:1).
The literals stay as the fallback for a host without the keys, and a test
pins the keys in every theme file of both apps.

Darling viewer theme copies ported 1:1; the deprecated Dashboard twin
carries the identical tab/accent defect and gets the same fix (no marks
there - it has none). Arm (B) of the report - user-maintainable colors
with reset - is not in this change.
… truncation is detected not inferred, and no page count is called a total (#3541 A3) (#3594)

* MCP pages say what bounded them: caps bind to the caller's limit, truncation is observed, no page count is a total (#3541 A3)

Six MCP tool groups on both SKUs carried a HIDDEN reader cap -- LIMIT 200 / 50 /
500 under a tool that advertised `limit`, a Take(limit) on top, and the capped
count published under a total_* name. An agent that cannot see the code read a
200-row page of a 5,000-event window as the window, and the "widen hours_back"
advice could never help because the cap was on rows, not time.

get_blocking / get_blocked_process_reports (200), get_deadlocks (50),
get_alert_history (+ a hidden dismissed = FALSE filter), get_long_query_completions
(200; Darling kept the SLOWEST, Lite kept the NEWEST then re-ranked, so Lite could
omit the window's slowest run), get_plan_corrections (200 newest per-cycle
re-captures ~ 16h reach regardless of hours_back), get_waiting_tasks (500 / unbounded,
bare envelope) and get_wait_stats (50 under a limit up to 1,000).

ONE pattern, the one #3287 established for get_collection_log: the cap is the
caller's limit bound as a SQL parameter; the tool fetches limit + 1 and reads the
extra row as `truncated` rather than inferring it from count >= limit; the page
publishes oldest_returned_* / newest_returned_* and names its `order`; the page
count is *_returned, never total_*. get_deadlock_detail and get_blocked_process_xml
moved their graph/XML predicate into the SQL so limit counts graphs, not rows they
would have discarded.

get_blocking / get_deadlocks honour #2159's whole-window promise for dedup_key:
the fingerprint scan is bounded by a stated 2,000-row ceiling, the payload carries
rows_examined / scan_truncated, and a no-match answer on a scan that ran out says so.

get_alert_history states and MEASURES its filter: dismissed_excluded and
dismissed_excluded_count on every page, include_dismissed to lift it (each row then
labelled `dismissed`), and an all-dismissed window names the filter instead of
calling itself quiet. Lite's completions read is now duration-ranked in SQL, so
both SKUs serve the same population.

Descriptions on every touched tool say what bounds the page. Grids keep their caps
(the readers default to them). Census tests pin the dialect across both SKUs;
Lite's run against a real DuckDB, Darling's live half is gated on DARLING_TEST_PG.

* Pass the as_of anchor by name on Lite's slowest-completions read; declare the new LF-reading pin

Two census pins from the first CI run. AsOfWindowAnchorTests requires every LocalDataService call that
CAN take asOfUtc to be given it by name, so the anchor's presence is visible to the scan rather than
inferred from position; the new GetSlowestLongQueryCompletionsAsync call passed it positionally.
RepoFileAdoptionTests holds the exact set of pins that read LF-normalised source; McpPageContractTests
is one (its Lite-description anchor spans the line break on get_plan_corrections), so it is declared.
…urce like get_query_trend already does, so a 7-day request no longer returns 4 days labelled quiet (#3541 A2) (#3590)

* The duration-trend trio routes by retention tier and discloses its source like get_query_trend already does (#3541 A2)

get_query_duration_trend and get_procedure_duration_trend read raw query_stats /
procedure_stats only, whose rows a TimescaleDB store drops at four days, while
hours_back accepts 168 — so a 7-day request returned 4 days under a label saying
7, and when nothing survived the empty branch called the window "genuinely quiet
— widen hours_back", which is false (dropped, not absent) and harmful (widening
cannot help). Both now route through the ladder get_query_trend already uses
(age → availability → coverage, DarlingTrendReader.ResolveTier) onto the hourly
rollups, dividing by the bucket width rather than a LAG, and every sibling —
the Query Store one included — publishes source / effective_start /
effective_hours_back / truncated / bucket / aggregate_note on the data path and
the empty one. get_query_trend gains the availability + coverage rungs (it
answered 42P01 on plain PostgreSQL past four days). Lite's twins publish the same
block with Lite's truth (raw, per-collection). value gains its named twin
elapsed_ms_per_second on both SKUs.

Partial: completes checklist item A2 of #3541 only.

* The retention clock lives in the reader, not the anchored tool bodies

AsOfWindowAnchorTests pins, as an absolute, that a tool advertising as_of never
names DateTime.UtcNow — its only "now" is the anchor it resolved. The trio's
route resolution measured raw age against the wall clock IN the tool body,
which is the right clock for retention (the anchor decides the window, not how
old its rows are) named in the wrong place. ResolveQueryDurationTrendRoute /
ResolveProcedureDurationTrendRoute now default nowUtc to the real clock the way
GetQueryHistoryAsync already does, and the route carries ResolvedAtUtc so the
empty branch's horizon comparison reads it from the route.

* RawRetentionApplies is scoped to the grain, because the arming gate is

Review finding: the flag was rollups != None, a store-wide test, while the #1680
gate arms each raw table's purge only once that table's OWN rollup covers it.
On #1664's failure-isolated partial build, a grain whose rollup is missing has
its raw rows intact — and the store-wide flag would have told that grain's
caller the rows were dropped and widening cannot help. Now hourlyAvailable, the
per-grain flag already threaded through the resolver; the past-horizon raw
branch's dead "rollup not present" arm goes with it (a grain with no rollup can
no longer reach that branch), and the pin asserts the per-grain split.

* RawRetentionApplies: say that existence-not-armed is the deliberate upper bound, and why (review question)
… over, so a restart's fabricated zero reads as unknowable instead of 0.00 ms/sec (#3540, V127 / Lite v60) (#3595)

* Interval columns for the four naked delta families (V127 / Lite v60), so a restart's fabricated zero reads as unknowable instead of 0.00 (#3540)

* Rung test counts the view statements by prefix; the lock-wait live test asserts the first collection is absent, not 0.00

* Lite pins: the wait_stats DDL carries the trailing interval; the lock-wait tool tests assert the first collection is absent, not 0.00
…n, tiered from 14 days of fleet data, so a week of blocking and an hour of blocking stop getting the same colour (#3539 A2/A3/A8d) (#3596)

* Blocking and CPU bands are rates over the window they were measured in (#3539 A2/A3/A8d)

BlockingSeverity bands blocked-process reports per hour over a declared window
(Warning 5/hr, Critical 20/hr, the 60 s / 10 s wait arms rate-independent), with the
deadlock band's too-short-window rule; tiers off 14 days of dogfood-fleet blocking.
The daily classifier delegates blocking to that band, scales its high-CPU bar with
the window (excursion-scale minimum 6, sustained-heat rate 1.25/hr = 30/day), and
bands collection errors as a share of the window's runs against the collector-health
classifier's own 20% bar at its tier (Warning, never Critical). CollectorSeverity
grades the failing count as a share of banded collectors (Critical past the same
bar). The daily SQL projects collection_runs (trailing column, both SKUs); every
producer declares the new denominators; the sweep ledger, cards, calendar tooltips
and MCP day tools disclose the rates and shares they banded on.

* Re-key the viewer collector-health helper pin on its widened tuple (#3539 A8d)

ViewerFleetRollupTests anchored on the helper's declaration text; the tuple gained
the Total the collector share divides by, so the anchor names the new shape.
… ten-minute window like its PostgreSQL twin, so one slow wait no longer pages and a THREADPOOL storm no longer sleeps (#3539 A4) (#3593)

* SQL Server's Poison Wait alert measures accumulated starvation over a ten-minute window like its PostgreSQL twin (#3539 A4)

The SQL Server Poison Wait alert judged ONE collector delta row's avg-ms-per-wait
against PoisonWaitThresholdMs (500), presence-flat CRITICAL, no volume floor: one
600 ms wait paged and a THREADPOOL storm of thousands of 8 ms waits slept. The
PostgreSQL twin (#2711) already fired on accumulated wait over a ten-minute
window, graded Warning/Critical, under the same alert name, switch and mute key.

Port the PostgreSQL shape to SQL Server. The read (both SKUs) SUMs
delta_wait_time_ms / delta_waiting_tasks per poison wait type over the window
with no task filter, no LIMIT and no threshold; the new shared PoisonWaitEvaluator
grades the sum (Warning at 1.0 avg tasks stuck = 600 s / 10 min, Critical at 10)
and both engines' constants are one definition (PostgresAlertEvaluator's are
aliases). Fleet margin on the SQL side: the worst ten-minute bucket across 43
servers / 4 days was 0.0097 avg waiters, ~100x under the bar. The clear arm now
requires an observed window (an empty read holds the alert; the retired shape
announced Cleared on collector silence). PoisonWaitThresholdMs stays readable and
persisted but is no longer consulted; docs and settings comments say so.

* Settings windows stop offering the retired poison-wait ms bar; live E2E compares at the store's precision

Review finding on #3593: both SKUs' Settings windows still labelled the poison
checkbox "avg wait >=" beside an editable ms box, and the save-preview line read
"poison waits >= Nms avg" - a knob the engine no longer consults. Label now says
the alert measures accumulated wait over a 10-minute window and that the ms bar
is retired; the box is disabled with a tooltip and kept so saved settings
round-trip; the preview states the live bar from the shared constants. Identical
edit on Lite and the Darling Viewer.

CI: the live read-adapter E2E compared a seeded DateTime (100 ns ticks) to the
value PostgreSQL stored at microsecond precision; compare at the store's
precision and assert the in-window row's clock won separately.

* PoisonWaitsSql: state the read's cost bound by shape (predicate-identical to the retired read, ~30 rows per server in the head chunk), not by index

* Settings windows: the retired poison-wait ms box stays disabled through the master-switch loop, both SKUs (review round 3)
erikdarlingdata and others added 19 commits September 18, 2026 18:50
…: the adapter returns the collection that produced the rise and the engine remembers which one it already reported (#3579) (#3628)

* Forced Plan Failing fires once per observation, not once per cooldown: the adapter returns the collection that produced the rise and the engine remembers which one it already reported (#3579)

* The live access-path seed is floored to whole microseconds so the #3579 stamp round-trips PostgreSQL's timestamp to the tick (#3579)
…very surface that shows it, so a 10 GB bar stops meaning 10 GB per five minutes on one store and per day on another (#3539 A8c) (#3631)

* The File Growth threshold means one thing — megabytes per hour — at every surface that shows it (#3539 A8c)

FileGrowthRiseMb was compared against the raw growth inside the lookback window while the lookback was a
second knob clamped 5..1440 minutes, so the same 10240 meant 10 GB per five minutes on one store and 10 GB
per day on another. The knob is now a rate: GetBreachedFiles holds the in-window growth to
FileGrowthRiseBarMb (rate × lookback/60), which is byte-identical on the shipped 60-minute lookback and the
same sustained rate on every other. Scaled to the CONFIGURED window rather than read off the measured span,
because the measured span on a freshly-collecting server is one collection interval and one autogrowth in
it would extrapolate to twelve an hour.

Both Settings rows, both preview lines, the alert's threshold line (rate + window + the bar in MB), the
card's Growth field, IAlertEngineSettings, both MCP descriptions and the settings carriers now say MB/hr
averaged over the lookback, in one shared phrase (AlertContextBuilders.FileGrowthRiseUnit) that
FileGrowthRiseUnitCensusTests holds across nine surfaces. Key, column and integer are unchanged.

* DarlingConfig.FileGrowthLookbackMinutes carries the same averaging-window note as its twins (#3539 A8c, review nit)
…he nightly rebuild's three cards become one incident that names the job (#3538 A6/A9) (#3632)

* Story confidence measures corroboration instead of path length, and the nightly rebuild's three cards become one incident that names the job (#3538 A6/A9)

* StoryConfidence in its own file; the exhaustive pins run to the engine's real 11-node maximum path (review)
… asks 'was it delivered': the once-a-day gate now stamps delivered-today in the store, so a restart re-announces nothing and a failed delivery still retries (#3580) (#3633)

* The daily digest and sweep rollup gate on delivered-today in the store, not fired-today in process memory: a restart no longer re-announces them, and a failed delivery still retries (#3580)

Both once-a-day gates were a ConcurrentDictionary under one fixed key, so a
fresh process delivered both documents again whatever the previous process had
delivered an hour earlier — six re-announcements among ~23 channel posts on the
v3.8.0 install night. The gate now reads a delivery stamp from
collect.collector_state (V44, no rung; server_id 0 = the fleet sentinel both
existing fleet-scope writers use, collector_name 'self_alert') and writes it
only when the deliverer reports a disposition other than `failed`, so the
install night's recovery case — a failed pair re-attempted after the restart
and landing — is preserved by construction.

- IAlertDeliverer.DeliverAndReportAsync: a default-implemented reporting twin
  of DeliverAsync (null = unreported); DarlingAlertDeliverer reports the
  AlertDelivery its history row was written with.
- DarlingSelfAlertEvaluator: FireAsync returns the report; the two documents
  stamp AFTER the fire; the store seeds memory after a start (#981's cooldown
  seed shape), memory serves the process; a stamp read fault falls open to
  memory, warns and is counted (#3013); a stamp write fault warns, uncounted.
- Census pins re-justified: AlertMasterSwitchSurfaceTests (new delivery seam
  name, generic FireAsync return), AlertReadFailureSurfaceTests (9th counted
  evaluator site, 12th exempt), Lite.Tests inventory (8th fleet-scoped read),
  AlertReadFailureCounter.FleetScopedReads names the new read.
- Tests: SelfAlertDeliveryStampTests (both documents, simulated restarts,
  failed/muted/unreported dispositions, read/write faults) and a
  DARLING_TEST_PG-gated live round-trip through the real table.

* The stamp store is asked once per document per process: an empty or failed answer is cached as a sentinel, so the apply half's gate does not re-read, re-warn or re-count a fault the evaluate half already met on the same tick (#3580 review)

Review catch: the "one store read between the two halves" cost model only
held when the read succeeded — a throw (or no row) returned without seeding
memory, so the apply half re-asked, doubling the warning and the #3013 count
per tick for one logical fault. Exact under the single-writer fact: a store
with no row on this process's first tick cannot gain one except through this
process, which seeds memory directly. New pin drives two consults on one tick
with both the read and the delivery failing — one read, one warning, one count.
…t-history grids show the severity the alert actually fired at instead of the colour its name implies (#3539 A6/A8e) (#3635)

* A server nothing has banded yet is Unknown, not Healthy, and the alert-history grids show the severity the alert fired at instead of the colour its name implies (#3539 A6/A8e)

* Keep dismissed the last selected column (the pin anchors on it), and paint the nothing-measured border with the cached Warning brush
…nd query_stats stores the statement offsets its delta key is made of, so no restart zero reads as a measurement anywhere and the query seed can finally find its keys (#3540, V128 / Lite v61) (#3630)

* Every delta family stores the interval its deltas accrued over, and query_stats stores the statement offsets its delta key is made of (#3540, V128 / Lite v61)

The completion of V127: sample_interval_seconds on procedure_stats, memory_grant_stats,
pg_wait_stats and pg_statement_stats, written as the minimum over each row's delta groups;
statement_start_offset / statement_end_offset on query_stats, stored raw (-1 included) so both
hosts' restart seeds rebuild the collector's delta key from the store. The census's still-naked
list is empty; the seeding census's pass-window-only set is empty. Readers of the four families
that divide by time take the stored interval first, LAG only for pre-V128 rows, and never ELSE 0.
get_resource_semaphore emits sample_interval_seconds again on both hosts.

* CI round 1: the PostgreSQL pair's CREATE TABLE rungs carry the interval for the fresh population (the V101 two-places rule); the resolving-view definer literal moves to 128; the backlog fallback pin records why it stays hash-keyed; the pre-V128 first snapshot is absent in the trend read test; two of my own pins fixed (V4 passthrough skipped, CTE-first FROM)

* CI round 2: the two Lite procedure-trend tool tests expect the pre-v61 first snapshot to be absent, the same correction v60 made for the wait trends
…t once per cooldown: the adapter returns the collection that produced the growth and the engine remembers which one it already reported per file (#3636) (#3639)

The rise gate's input is a stored delta between two collections of database_size_stats, which lands hourly; the engine re-asked every ~30 s and re-fired on every 5-minute cooldown against the same two rows — up to twelve cards per growth event. #3579's mechanism, one condition over, at the worse ratio.

Both SKUs' reads project the newest sample's collection_time as observed_at, appended last so the fourteen bound ordinals do not move. DatabaseFileGrowthInfo gains a nullable ObservedAtUtc. The engine keeps, per (server, database, file), the stamp it last fired on beside the per-server cooldown clock and active flag, and declines to send the card when no breached file is news: a level breach is always news (standing level, re-fired by design); a rise-only file is news only with a stamp newer than the one it last fired on. Stamped on fire only, including muted fires; pruned as files leave the breached set; cleared on recovery. In-memory like the forced-plan memory — a restart may re-fire once, which the emptied cooldown clock already did.
…defaulted one: CONTRIBUTING's Two-Store Parity rule names this interface, so Lite's deliverer and every test fake now state their (null) answer by hand (#3580 review) (#3640)

Review note taken: a default body compiled and left LiteAlertDeliverer and
fourteen fakes quietly inheriting an answer nobody wrote down — the exact
pattern CONTRIBUTING forbids for this seam. Lite's implementation returns
null and says why (its send seam returns no disposition; it hosts neither
daily document); the DarlingAlertingTests wrapper forwards the inner
deliverer's report rather than swallowing it. Also: an xUnit2013 analyzer
warning in the new two-consults pin (Assert.Single over the Regex matches).
…folds one cause into one row, so same-hour-yesterday noise stops reading as a verdict (#3538 A3) (#3634)

* compare_analysis bands each delta by the server's own dispersion and folds one cause into one row, so same-hour-yesterday noise stops reading as a verdict (#3538 A3)

The tool compared ONE window against ONE window and banded every key by its
severity delta on a flat +/-0.1 dead-band. Severity is a threshold-formula
artifact: a saturating ladder read a doubling from 30% to 60% of observed time
as "stable", a trace of a 1%-bar wait read a few seconds an hour as "worse",
one I/O stall surfaced as four to six worse keys, and BAD_ACTOR_<hash> keys
churned into new_issues/resolved_issues on plan-cache eviction alone.

Shared ComparisonBanding (PerformanceMonitor.Analysis) now decides every
verdict for both SKUs: a key with a same-unit per-server baseline (CPU %,
read latency, connections) is banded in that bucket's robust sigma
(delta_sigma, band_source "baseline", trustworthy buckets only); every other
key changes status only when the value moved >= 25% of the larger side AND the
larger side sits >= 0.25 up its own Layer-1 ladder (the scorer's base severity,
so every concerning bar is reused with no number copied); one-sided keys are
new/resolved only when they register on their ladder; BAD_ACTOR_* appearances
are plan_cache_churn; rows carry a physical-cause family and the payload
carries one family row per cause; the rules are stated in band_rules; every
verdict row carries coverage_caveat when either side was partly observed
(#3592 composes). ComparePeriodsAsync on both SKUs returns the dispersion map,
fenced so a failed baseline read degrades to the absolute rule.

* Verdict order is direction before magnitude, so a mixed-direction family's worst member is its regression (#3538 A3 review)

The row sort partitioned only stable from non-stable; worse and better shared
a bucket ordered by relative move, so a family whose BLOCKING_EVENTS fell 60%
while its LCK_M_S rose 30% took the improvement as its worst member and left
families_worse. Rows and families now rank worse > better > stable, then move;
a mixed-direction family pin covers it.
…rs, their Long-Running Query alert skips maintenance like SQL Server's does, and every one-sided alert says why it is one-sided (#3539 parity) (#3638)

The fleet card's deadlock dot on a PostgreSQL target read Unknown while
pg_database_stats held the server's own deadlock counter, because the fleet
reader took v_deadlocks (the SQL Server extended-event capture) and
ServerMetricSources nulled the structural zero. Both fleet surfaces (the
service's get_fleet_overview / /api/fleet reader and the viewer's Overview
card and totals) now difference pg_stat_database.deadlocks per
(server, database) series between consecutive samples, clamp at zero across
a statistics reset, sum across databases, and band the result through the
SAME deadlock_warn_per_hour / deadlock_critical_per_hour tiers as the SQL
Server graph count. Never SUM(deadlocks): the column is a lifetime counter
repeated in every sample. FleetDeadlockSource.PostgresTarget now means
"counted from the server counter" and is a covered arm; servers_read counts
it and postgres_servers is a sub-count. A PostgreSQL target whose
pg_database_stats collector is silent or denied reads those arms, exactly as
a SQL Server whose deadlocks collector is.

The PostgreSQL Long-Running Query read gains the noise opt-outs SQL Server's
has had: non-client backends (mirrors session_id > 50), the
VACUUM/ANALYZE/REINDEX/CLUSTER statements by command tag (no text is stored,
so the tag is the only handle), the dump/restore utilities by
application_name on the shared longRunningQueryExcludeBackups switch, and
excludedDatabases after the read. CREATE stays reported with its tag.

The Darling README gains an engine-coverage table for every built-in alert
and band, stating why Blocking Wait Time and Database File Growth are
SQL-only, why the PostgreSQL blocking BAND stays Unknown (sampled evidence
against tiers measured on engine-recorded reports), that Poison Wait is one
shape since #3593, and that slot retention's SQL analogue
(log_reuse_wait_desc) is collected only as a load-time snapshot today.
…cet by facet from the stored config, so a target with everything off stops looking identical to an instrumented one (#3607) (#3643)

* get_pg_logging_audit judges a PostgreSQL target's logging settings facet by facet from the stored config, so a target with everything off stops looking identical to an instrumented one (#3607)

Plan-capture readiness established the shape - a facet per setting, the remedy per finding, the hosting flavour's own syntax - and covered only auto_explain's three preconditions. The rest of the logging surface (log_lock_waits, log_temp_files, log_autovacuum_min_duration, log_checkpoints, log_connections / log_disconnections, log_min_duration_statement) had no audit at all, and a target with all of them off answered every counter-based read exactly like a fully instrumented one.

This is a READ over pg_server_config's newest snapshot, not a collector: every judged setting is a core GUC the config collector already stores hourly with its value, source, unit and context, so the judgment is a pure function computed on request and the snapshot's collection_time is the audit's as_of. Verdicts are instrumented / partial / off / unknown; partial names what a threshold hides and is the recommended posture for log_min_duration_statement. Consumers are named honestly as PLANNED (#3601/#3602/#3603) beside the counter read that exists today. Hosting flavour comes from rds.* parameters in the same snapshot, which is the one fact the registry's engine token cannot supply. Plan capture's own settings are listed with the readiness facet that owns each, never judged twice.

Registered beside the plan tools, dispatched on the web with a Configuration-tab-adjacent panel under readiness, censused in the instructions / README / llms.txt / runbook, ratcheted in the cross-app inventory pin.

* The audit speaks #3541 A10's stamp dialect: captured_at on the wire, LATEST IS A TIME in the description, and a Darling-only allowance in the stamp roster

dev moved under the lane: #3637 landed McpLatestSnapshotStampTests, whose reader-call sweep caught GetNewestSnapshotAsync and asked for a Shape. The roster is a roster of PAIRS (every entry names its Lite twin's file), so a Darling-only tool gets a DarlingOnlyStamped list held to the same three Stamped assertions, and the new reader's SQL joins the stamp-column theory. Threshold() became a block body so TsqlConventionGuardTests' member scan reads it whole rather than growing KnownTruncatedRanges.

* Review round 2: pending_restart reaches the row and the summary, the catch re-checks the engine gate, the ten-minute default is PostgreSQL's own verdict, and the two unknowns are told apart

pending_restart was fetched, documented as load-bearing, and dropped before the Facet - on exactly the row whose remedy needs qualifying, because the value judged is the RUNNING one and the file already holds another. It is now on every facet with a restart_note, named at the top as pending_restart_settings the way get_pg_server_config names them, and seeded in the live test. The catch block asks the engine gate before answering error, as the plan tools do. atDefault reads source = 'default' rather than comparing setting text to boot_val - an administrator who writes 600000 into the file has made a choice, which is the anti-pattern DarlingPgServerConfigReader's own comment names. An unparseable value now says it IS in the snapshot rather than that it is not.
…ilter is part of the query: parallel_only/min_dop/blocking_only cut before the page, a negative window is refused, an unknown source names the accepted set (#3541 A9/A13) (#3641)

* The daily summary stops painting purged months green, and every MCP filter is part of the query (#3541 A9/A13)

A9: the daily aggregate's day spine outlives its signals (collection log 60 d, alert log 90 d,
signals 30 d) and COALESCEs each purged signal to a measured zero, so a year-long range banded the
purged months Healthy. Both SKUs' readers now compute the store's retention horizon (Darling: the
shortest EFFECTIVE fleet retention among the seven signal collectors and the two log constants;
Lite: RetentionService.ArchiveRetentionMonths) on the reader's wall clock, the SQL projects how
many signal sources still hold rows for the day, and the shared DailySummaryRetention.StateFor
judges each day: purged (past horizon, no signal rows -> No Data, never Healthy), past_horizon
(rows survive, verdict withheld), no_run_record, collected. The range tool publishes
retention_horizon / days_before_horizon / purged_day_count / collected_day_count and each row its
data_state + data_note; the single-day tool answers a purged day with the unavailable envelope.
summary_date is TryParseExact yyyy-MM-dd on both SKUs (McpHelpers.ParseSummaryDate).

A13: parallel_only / min_dop are a HAVING floor on the grouped population before the CPU ranking
and the cap (both SKUs, both Darling grouping reads), with filter_applied on the payload and a
filtered-miss sentence instead of "no query stats". get_active_queries pushes database_name and
blocking_only into the SQL, counts the filtered population with COUNT(*) OVER (), pages at
limit + 1 (snapshots_returned / truncated / bounds / order), never strips a head blocker some row
in the same capture names (is_head_blocker), and names why a victim's blocker is absent
(blocker_not_shown: not_captured / filtered / past_page). The three uncapped reads refuse a
non-positive hours_back through McpHelpers.ValidateUncappedWindow instead of Math.Abs. The
get_analysis_facts source filter is validated against FactScorer.KnownSources (15, pinned to every
collector literal) and refused with the whole set.

Census: get_active_queries joins McpPageContractTests' paged dialect on both SKUs; new A13 pins
(no .Where after the read, predicates before ORDER BY/LIMIT, no Math.Abs on a parameter anywhere,
the three uncapped reads on the shared validator); live PostgreSQL round-trips for the filter,
head-blocker, purged/past-horizon/override arms; Lite DuckDB twins.

* xUnit2009: Assert.EndsWith for the daily SQL's trailing ORDER BY pin

* A day inside retention with no run record keeps its band; the displaced range-read summary goes back to its member; the two fixed-date Lite fixtures follow the horizon

CI's first execution caught three things the Mac harness could not. (1) PerformanceCalendarDataTests has always
pinned an alert-only day (no collection-log run) as Warning — the alert is real and inside retention every zero
is a measurement — so no_run_record is now a DISCLOSURE (error share has no denominator) that keeps the band,
and only the two past-horizon states withhold it; both SKUs' ToSignals, the descriptions, the instructions
rows and the pins say so. (2) The same fixture's fixed July 2026 month would have drifted past the three-month
horizon within weeks; it is now the calendar month two months before the current one. (3) DocCommentHygiene:
the DailySummaryRangeReadResult insertion had displaced GetDailySummaryRangeAsync's summary block. Plus the
FindingsRetentionHorizonPinTests literal pin follows the promoted RetentionService.ArchiveRetentionMonths.
…y hour and paid twelve index inserts per row for eleven indexes nothing reads; the capture-down alert read decompressed a server's whole collection_log to learn two statuses (#3597, partial) (#3647)

* The interval-hourly refresh stops paying twelve index inserts per re-materialized row for eleven indexes nothing reads, and the capture-down alert read stops decompressing a server's whole collection_log to learn two collectors' latest status (#3597, partial)

Rig-measured on PostgreSQL 18.4 / TimescaleDB 2.28.1 at one tenth of the largest production store's scale: a refresh of query_store_stats_interval_hourly recomputes exactly the hour-buckets the invalidation log marks dirty — the newly closed one every hour, plus one whole fleet-wide bucket per distinct past hour any backdated row landed in (a ONE-row backdated insert re-materialized 36,120 rows; twelve backdated rows in one transaction across twelve hours re-materialized twelve buckets, clean hours included). The window's width is not the cost; the dirty-bucket count times the per-bucket cost is. The only backdated writer is QueryStoreBackfill, by design (collection_time = slice ceiling).

The per-bucket cost carried TimescaleDB's default create_group_indexes: eleven (column, bucket DESC) btrees on the L1 materialization that no reader uses — its three child aggregates read it by bucket range, the coverage probe and arming gate read min(bucket), retention drops chunks. EXPLAIN (ANALYZE, BUFFERS, WAL) of one bucket's materialization INSERT: 45.5 MB WAL / 483,689 records / 1.64 M buffer touches / 6,040 dirtied with them; 10.7 MB / 72,735 / 447 K / 56 with the bucket index alone. L1 is now created without them and EnsureIntervalDedupMaterializationIndexesAsync drops them on an existing store at startup, per index under a 10 s lock_timeout so a refresh in flight is yielded to rather than convoyed.

The capture-down read was the #3496 shape #3496 named and left: ROW_NUMBER() OVER (ORDER BY log_id DESC) over every collection_log row the server had for two collectors — 9,573 buffers across 61 chunks (59 decompressed) per alert pass per server on the rig. It is now one chunk-orderable LIMIT 1 per collector: 14 buffers, the newest chunk only, 120 of 122 chunk scans never executed.

Partial: the dirty-bucket multiplier is the backfill's designed behaviour and is not changed here; the production dirty-bucket count per hour is the read that decides how much of the 417 s average this removes, and the exact statements for it are on the issue.

* IntervalDedupMaterializationIndexesTests records why it is not serialized against the live collection: its gated arm mints a scratch database, so it cannot race the shared store (#1776 own-store, the census's own guidance)
…nterleaved fields: a drill-down's rows travel as records beside their flat fields, and a five-fact page's fixed items cost 20 blocks instead of 33 (#3644) (#3649)
…y at all and its empty chart reads as a broken product: the collector grows a service-side pg_stat_activity sampler arm as the floor, the three wait tiers become one connect-time decision, and every wait read discloses its instrument (#3604) (#3645)

* A stock PostgreSQL target without pg_wait_sampling gets no wait accumulation at all, so its empty wait chart reads as a broken product: pg_wait_sampling grows a service-side pg_stat_activity sampler arm as the floor under stock targets, the three wait tiers are one connect-time decision, and every wait read discloses which instrument fed it (#3604)

The pg_wait_sampling collector forks on a new connect-time fact,
CollectorTargetInfo.HasPgWaitSamplingExtension: with the extension it reads the
10 ms profile as before; without it, it runs one multi-statement command of
thirty one-second pg_stat_activity snapshots (each behind pg_stat_clear_snapshot,
because the view is cached per transaction) and accumulates the tallies across
cycles in its own collector_state, so the rows it writes are cumulative and the
existing newest-minus-oldest read differences them unchanged. Which arm ran is
recorded in state; get_pg_wait_sampling, get_pg_wait_stats and the Viewer panel
disclose it as instrument: engine_cumulative | extension_sampled |
service_sampled, with the floor caveat on the service tier.

Aurora is gated off pg_wait_sampling (it cannot preload the module; the hourly
EXTENSION_MISSING there was a permanent gap dressed as a precondition), and
CoveredInsteadBy now points an Aurora caller at pg_wait_stats. The collector
moves from hourly to five minutes and is detached from the sequential sweep body
beside query_store and plan_correction, because a deliberate 30 s window inline
would delay every other collector on the target; the window is 30 s and not the
cycle because SweepPressureClassifier sums single-run costs against a 60 s body
budget.

No migration: same table, same columns, same collector name. The sampler needs
nothing beyond pg_monitor.

* Three pins encoded 'Aurora is a strict superset of what the PostgreSQL collectors read'; pg_wait_sampling is now the one documented exception (the module cannot be preloaded there), so the kind-axis split, the pre-flight probe line and the compose measure census each name it rather than assume it away; the empty-arm helper moves above the JSON builder's doc block it had split (#3604)

* The sampler's pg_sleep argument is formatted under InvariantCulture, as every other number this file puts into SQL or state is: harmless at a whole-second period, a comma-decimal host's pg_sleep(0,5) the day it is not (#3645 review)

* The extension arm clears the sampler's tally every cycle it runs: the host persists only PendingState keys, so a tally left alone would sit in collector_state for as long as the extension stayed installed and be resumed as a months-old baseline the day the target fell back to the sampler arm — splicing two eras into one series when the extension's own profile had just been reset (#3645 review)

* The service-tier note claimed a series resets when the service restarts; the tally is persisted collector state reloaded every cycle, so it survives one — the note now names what actually starts a series over (cap eviction, arm reversion) and a pin keeps the false claim from returning (#3645 review)

* Two Lite pins catch up: the state-declaring set is two collectors now, and get_pg_wait_stats' instrument_note naming aurora_stat_system_waits() is accounted for as PROSE in the Aurora-surface allow-list (#3604)

---------

Co-authored-by: erikdarlingdata <erik@erikdarling.com>
…plan maximum kept as a dated history, so a MAXDOP-1 instance stops reporting 16 from a plan compiled before the pin (#3648) (#3651)

* Top Cpu Queries' Max Dop is the newest plan's reading with the cross-plan maximum kept as a dated history, so a MAXDOP-1 instance stops reporting 16 from a plan compiled before the pin (#3648)

* Parity guard moves its cross-SKU comparison to Darling.Tests (the filter that reaches both trees), the live test cleans up through LiveStoreCleanup, and the DuckDB fixture reads UtcNow once so a round-tripped timestamp compares equal (#3648)
…er seen, a regression with no baseline stays null, and the first point of a differenced trend is no longer a fabricated 0 (#3541 A12) (#3642)

* Zero is a measurement: health parsers say whether their source was ever seen, a regression with no baseline stays null, and the first point of a differenced trend is no longer a fabricated 0 (#3541 A12)

Six MCP sites published a 0, an empty, or a nominal-window label where nothing had been measured. Eight of the nine get_health_parser_* tools answered a dead system_health session with the same empty a healthy hour earns; now every one carries source_observed / last_captured_at and climbs the four-rung ladder significant_waits alone had, with nothing-ever-captured answering unavailable. get_query_store_regressions kept a NULL percent NULL (a 0 baseline has no ratio) with the reason under undefined_percents, and nulls a severity banded from a ratio that does not exist. The LAG-differenced duration trends (three raw reads, the Query Store rollup builder, Lite's four) leave the window's first collection unrated - null, counted in unrated_points, explained in unrated_note - instead of 0.0. get_pg_xmin_horizon divides a holder's wins by every capture in the window (collection_log SUCCESS rows, the alert evaluator's own denominator) rather than by the source's own rows, and answers unavailable when the collector never looked. get_pvs_stats divides any MEASURED size, so 0 MB is 0.00% and only an unmeasured numerator or absent denominator yields null, with pvs_measured and the reason. get_table_index_sizes projects its baselines raw and derives each growth figure from exactly the baseline it names, publishing history_days_available and growth_over_available_history_* where the nominal windows are out of reach, plus tables_returned / truncated.

* Compose with #3630 (V128 / v61): the procedure trend keeps its stored-interval derivation and the MCP reader keeps the NULL-rate row as an unrated point rather than dropping it; the SQL-mirror pin's comment lines move to the C# doc; the #3540 pins that expected the first snapshot absent now expect it unrated

* CI round 1: rung 3's closing sentence no longer contains the word the sibling rungs are pinned NOT to say; the Lite trend parity test's lone in-window collection is asserted unrated rather than compared as two zeros; the eight expression-bodied growth derivations join KnownTruncatedRanges (they strand no literal)

* Re-cut head: the Claude review posted nothing on 1deff83 across five attempts (the #2229 swallowed-output class); a fresh head gives the review and the guard fresh runs. No code change.
…ly tools end the denial churn, the job fails when no verdict was submitted, and the transcript survives the runner (#3650)

The prompt told the reviewer the branch was checked out and asked for a correctness,
parity and security review, while --allowedTools permitted only four gh verbs and the
inline-comment tool. Every Read, Grep, Glob and git call was a permission denial; the
swallowed runs' result blocks read 50 turns / 18 denials, 39 / 21, 18 / 17, each ending
subtype=success with nothing posted. Three PRs, about thirteen paid runs in one night.

- --allowedTools gains Read, Grep, Glob, Bash(git diff:*), Bash(git log:*), Bash(git show:*).
  Nothing that writes.
- A step after the action fails the job when claude[bot] submitted no non-empty-bodied
  review since the run's own start stamp, and prints the transcript's result block. A
  refusal-to-run (this file differing from the default branch's copy) is named separately.
- claude-execution-output.json is uploaded as the claude-review-transcript artifact, 7 days.

Lands on main and dev together: claude-code-action refuses to run whenever this file
differs from the default branch's copy, so a dev-only edit would disable review on every PR
until the next release.
@erikdarlingdata

Copy link
Copy Markdown
Owner Author

Closing unmerged: this pointed the dev-cut lane branch at main, so it carried 69 commits — all of dev's unreleased work — where the pair strategy wants exactly one. Replaced by #3662, a main-cut branch cherry-picking the same workflow commit (c39de4c), byte-identical to #3657's dev half. Merge #3662 + #3657 together to keep the drift-arm window closed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant