diff --git a/CHANGELOG.md b/CHANGELOG.md index 5a4e219e3..ee21ba5a8 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -8,6 +8,12 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 ## [Unreleased] ### Added +- **The member-range scan now reports ranges that stop short of their own content, which it previously read as healthy** ([#3230]) - `ShapeOf` named two failing shapes and both were over-reads; a range ending early is closed, is under its successor's start, and overlaps nothing, so it satisfied every arm and reported `WholeMember`. `DeclaredRange` now carries a second, independently derived content boundary and `RangeShape.Truncated` reports the disagreement. Thirty members whose ranges stop short are carried as a set-asserted inventory rather than rewritten, because no merged consumer of the map under-reads because of them. +- **The clock-frame census now sees a renderer reached through a one-hop wrapper** ([#3207]) - four server-local `plan_correction` stamps in the Darling viewer were in the wrong frame and invisible to it, because the render scan keys on the renderer's name and a wrapper is a different name reaching the same renderer. The wrapper set is derived and pinned at set equality, so a new one must declare the frame it renders, and no wrapper may be handed a column every table declaring it frames the other way. +- **`.github/dependabot.yml` now tells whoever reviews a `Microsoft.Data.SqlClient` or `Azure.Identity` bump to check whether `ExcludeBrokerCredential`'s default or the broker's packaging changed** ([#3219]) - Lite's `EntraDefaultCredential` mode is broker-free only because `Azure.Identity.Broker` is absent from the graph, `Azure.Identity` is unpinned and arrives transitively through SqlClient's range, and the two pins guarding that absence assert the absence of the package rather than the outcome. +- **Lite now records which Azure credential `DefaultAzureCredential` selected for an Existing Sign-In (`az login`) connection** ([#3224], a follow-up to #3218, and see [#3219]) - the mode signs in as whichever identity the machine already has, in an order the driver owns and the app cannot narrow, which the connection dialog and the README warned about and nothing recorded. A narrowly-scoped `EventListener` enables `Azure.Identity`'s `Azure-Identity` event source for exactly the duration of one connection open, forwards event 13 and nothing else, and writes the selected credential's type name - `AzureCliCredential`, `EnvironmentCredential`, `ManagedIdentityCredential` and so on - through `AppLogger`. The filter is an allowlist on the event id rather than a level or a keyword because neither can narrow this: event 13 is itself `Informational`, so that is the lowest level which delivers it, `EnableEvents` admits everything at or above the level requested (28 of that source's 29 events), and not one event in it declares `Keywords` - so the callback filter is the only barrier between the app's log and the siblings carrying tenant ids, account details and raw exceptions. `Payload[0]` is the only value read and only if it has the shape of a CLR type name, which is what survives a future payload reorder. Expect one line per run of the app: `Azure.Identity` reports the selection once per credential instance and the driver caches that instance in a process-wide static, so later connections have nothing new to report and say so at `Debug` rather than silently. Observability only - the connection path, the connection string and the credential chain are unchanged, and no other authentication mode enables the event source or pays anything for this. +- **A guard over the boundary three clock-frame defects shipped through** ([#3208]) - a stored column holding the monitored server's local wall clock, handed to an MCP payload or a WPF render site that assumes naive UTC. The census is derived from the collector catalog — 66 timestamp columns, 49 of them SQL Server — and each column's frame is read out of its own collector's query text, per (table, column), with the nine columns whose T-SQL provenance cannot be read answered by a declaration carrying the C#-side evidence rather than defaulting to UTC. Whether a column NAME may stand in for its column is itself derived and cross-checked against `StoreSqlClockDisciplineTests`' hand-written ambiguous-name register. 33 MCP payload sites and 22 desktop render sites are carried as a labelled inventory pinned at set equality, so a new offender and a fixed one both fail the build. +- **Added a sixth authentication mode, Azure — Existing Sign-In (`az login`), for Azure SQL Database where Microsoft Entra MFA cannot work** ([#3214], reported via [#3196]) - `Microsoft.Data.SqlClient` forces the Windows account broker on for any caller using the driver's own Entra application id, `UseWamBroker` is inert in that configuration, and MSAL falls back to a browser only when the broker is *unavailable* rather than when an invoked broker refuses - so a broker refusal has no in-process remedy and Entra MFA cannot be repaired in place. The new mode maps to `SqlAuthenticationMethod.ActiveDirectoryDefault`, whose arm in the driver returns before any MSAL public-client application is constructed, which is the only thing `WithBroker` is ever attached to. It signs in as whichever Azure identity is already established on the machine and **never prompts** - the driver hard-codes `ExcludeInteractiveBrowserCredential`, so a machine with no Azure sign-in on it gets a failure rather than a sign-in window. That failure is classified into "nothing found, run `az login`" versus "one found and broken", because a mode that fails as opaquely as the one it replaces is not a fix. The credential search order is the driver's and cannot be narrowed from a connection string, so on a machine with several Azure identities configured this connects as whichever comes first - the connection dialog says so, and Service Principal remains the mode that names an identity. Darling does not offer it: the Darling service's connect path builds Windows-integrated or SQL-login connections only and acquires no tokens, so its viewer's auth whitelist rejects it with no change needed. **Unverified against a live Entra tenant**, which is the only place this class can be confirmed. - **A collector's PostgreSQL extension dependency is declared rather than described** ([#3187]) - `Darling/README.md`'s permissions paragraph named the collectors that need an extension installed, and nothing derived that list - a commit in #3184 named **four of the six** and all eight checks passed on it, corrected only because a reviewer read the paragraph. An enumeration in prose is only better than a count if something breaks when it is wrong. `ICollectorSchemaInfo.RequiredPgExtensions` now carries the extension each collector cannot run without plus what installing it costs - `CreateExtension` for a statement, `SharedPreloadLibraries` for the ones needing a server restart first - and the paragraph is pinned to it in BOTH directions, so a dependency the product has and the documentation omits fails the build, and so does a paragraph claiming one nothing declares. **The nearby list that looks like the authority is not it**: `PgExtensionAvailabilityCollector`'s roster carries eight entries because it exists so ABSENCE is reportable, three of which no collector reads, and it deliberately omits `pg_wait_sampling`, which one does. Two axes stay derived rather than declared twice - per-database is `RunsPerDatabase`, and `pg_stat_kcache` sitting on `pg_stat_statements` is a property of that extension. A second pin closes the case the README pin cannot see, where a new collector reads an extension and never declares it: a PostgreSQL definition whose query text touches an extension-owned object has to declare it, and the verification rig has to be able to load and create everything declared. The paragraph's three word-numerals over the same set are deleted rather than pinned, the #3072 alternative. **Two corrections found by reading rather than by a pin**: `pg_stat_kcache` and `pg_qualstats` need `shared_preload_libraries` too - which the MCP tools already told operators and the rig's own preload line confirms - so following the paragraph meant creating the extension and watching the collector store nothing. And the paragraph explained `pg_extension_availability`'s exclusion of `pg_wait_sampling` by saying preload-only modules never appear in `pg_available_extensions`, which is false for that one case: the rig's `seed.sql` creates it with `CREATE EXTENSION` and the collector reads a function only that statement creates. The exclusion now carries the set-membership reason instead - reportable-absence and collector-dependency are different sets - which is the same reason the new type states. - **A refresh-ceiling staleness finding, reported separately from where the reading sits against its slot** ([#3182]) - a live run above a constant recorded as a maximum means the constant is wrong, which is a different fact with a different remedy from the slot being exceeded, and it is reported from every band including the routine one. Rate-limited by a high-water mark per constant. `OtherHourlyRefreshObservedCeilingSeconds` gets a live per-view feed for the first time, and the guard's derivation gets the upper bound the hour imposes on it, so a re-derivation that does not fit is red at the constant. - **A TimescaleDB compression run is watched against the clearance its own minute of the hour has before the next continuous-aggregate refresh starts** ([#3112]) - a daily chunk-close run that overruns into a refresh is reported instead of reading as 15% of its hourly cadence; both the refresh window's watch line and the new compression one are derived over a single named lead-time fraction. @@ -155,10 +161,22 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 - **Lite's portable ZIP is self-contained, which HALVED it** ([#2501]) - `Publish Lite` is now `-r win-x64 --self-contained` in both `build.yml` and `nightly.yml`, so neither Lite artifact has a .NET prerequisite any more and the failure [#2489] documented stops existing: a tester who unzips onto a stock Windows Server no longer meets the .NET host's bare `You must install .NET to run this application` before a line of our code runs. **The size went the opposite way from what bundling a runtime suggests.** The old publish was RID-agnostic, so it copied every platform its packages ship - **537 MB of `runtimes\` on a 565 MB tree** (osx 130, linux-x64 116, linux-arm64 70, win-arm64 56, then win-x86, musl, loongarch64 and riscv64), of which only the **52 MB `win-x64`** folder could ever load on Windows. `DuckDB.NET.Bindings.Full` is most of it, SkiaSharp and SqlClient behind it. Dropping ~485 MB of unloadable native payload beats the cost of bundling .NET, WPF and ASP.NET Core by roughly two to one: measured on one commit and one SDK, **565 MB tree / 212.7 MB zipped becomes 277 MB / 114.2 MB**. It matters most for the **nightly** ZIP, which is the UAT download and is not offered as a `Setup.exe` at all. **A RID-specific publish needed two more files than the flag.** `Lite/packages.lock.json` had only a `net10.0-windows7.0` target, and a RID restore adds `net10.0-windows7.0/win-x64` to it - after which the `dotnet restore --locked-mode` that BOTH workflows run before the publish fails `NU1004: the project's runtime identifiers have changed`, because locked mode compares the PROJECT's RID set (empty) against the lock file's (win-x64). Reproduced locally; that is a red CI run on every PR, not the future `--no-restore` trap it was filed as. The fix is `win-x64` in `PerformanceMonitorLite.csproj`, so the project itself asks for that graph and one committed lock file satisfies the RID-less locked-mode restore and the RID publish alike; `RuntimeIdentifiers` (plural) sets no RID on the build, so a plain `dotnet build` stays RID-agnostic and `Lite.Tests` is untouched. **SignPath needed nothing** - the `Lite` artifact-configuration slug already receives both shapes today, and the signed re-zip reads `signed/Lite/*`, inheriting whatever shape `publish/Lite` has. Auto-update is unaffected; the ZIP is not a Velopack channel. `LiteRuntimePrerequisiteDocsTests` went red on the flag alone (3 of its 7 facts) and was rewritten to state every claim BOTH ways round: [#2499]'s version asserted only that the docs DID name the runtimes, so two of its facts stayed green while the prose went stale. It now also derives the lock file's RID coverage from the `-r` flags in the workflows, and every new assertion was proven red with its fix reverted. ### Fixed +- **The de-skew's accuracy claim is now scoped where a column can outlive a DST transition** ([#3231]) - `server_properties.utc_offset_minutes` is `DATEDIFF(MINUTE, GETUTCDATE(), GETDATE())`, the offset in force at collection time, so subtracting it from a stored timestamp is exact only while the timestamp and the offset sit on the same side of a transition, and #3212's sixteen de-skewed fields shipped without saying so. Thirteen of the sixteen describe current state - a running job, an open transaction, a cleaner that ran seconds ago - and are always in the same DST period as the collected offset. `get_index_usage`'s `last_user_access` is not: it is the `GREATEST` of four `sys.dm_db_index_usage_stats` columns that persist since the instance restarted, routinely months on a stable box, so a large share of those values predate the most recent transition and come back 60 minutes early - silently, and in the plausible direction, which is the bad one. `get_pvs_stats` gets the same note one tier down, since `aborted_version_cleaner_*` can reach back weeks on a quiet database. Documented rather than fixed, deliberately: the proper fix is a zone instead of an offset - `CURRENT_TIMEZONE_ID()` collected alongside `utc_offset_minutes`, then `AT TIME ZONE` at the read boundary - which is a collected column, a migration rung and five readers, and the offset path has to be *kept* rather than replaced, because `CURRENT_TIMEZONE_ID()` is SQL Server 2019+ and the product supports servers older than that. **There is no backfill**: the collector ships the column verbatim and the offset is subtracted in the projection, so nothing stored is wrong today, and a single current zone id is legitimately valid retroactively where a single current offset is not - which is the whole defect. +- **Desktop timestamp columns were rendered in the wrong clock frame in both SKUs, in opposite directions, with each SKU correct exactly where the other was wrong** ([#3207], [#3221]) - a column holding the monitored server's own wall clock — `sys.dm_exec_*` creation / cache / last-execution stamps, `plan_correction`'s recommendation and action stamps, the blocked-process report's transaction and batch attributes, `query_snapshots.tran_start_time`, msdb Agent's job start time and the deadlock graph's `lasttranstarted` — was sent through the renderer that adds the collected UTC offset, displaying it four hours early on a fleet at UTC-4; and Query Store's first- and last-execution stamps, which the collector normalises to UTC, were rendered raw and displayed four hours late. Each of twenty-nine render sites now takes the renderer matching its column's frame, and Lite gains `ServerTimeHelper.FormatServerClock` for the frame it had no renderer for, so a server-clock value converts exactly once and still honours the Server / Local / UTC display preference. +- **Four doc comments stated a column's clock frame incorrectly, one of them appealing to a sibling tab as precedent and one asserting a cross-SKU parity that did not hold** ([#3207]) - they are the reason one wrong site became several, and each now states the frame the code uses. +- **Composed panel queries now assert that every compiled statement binds exactly the parameters its plan predicts and cites every one of them** ([#3211]) - `ComposeCompiler` allocates positional parameters by hand at each site, so a site binding a value it already bound gets a fresh ordinal and shifts every later one — a class that survived the whole suite. Measured correction to the proposed fix: when the second bind's placeholder is the one the new site uses, every ordinal still appears in the SQL, so placeholder coverage and any max-based comparison are both satisfied; only a count derived from the plan sees it. Coverage is kept alongside for the spelling where the placeholder never reaches the text. Both sweep a catalog-derived corpus whose reach is asserted as membership rather than as a total, and the prediction's own reach is pinned against the compiler's `ParamList` call sites. +- **`ServiceCommandDeadlines.SerialLoopSeconds` keeps its value of 5 and loses the chain bound that justified it** ([#3204]) - `DarlingWorker.SweepWatchdogSeconds` clocks per-server collection bodies from their own launch and only selects a log level, so it never observed the serial loop; and the site census that formed the product's other operand covers only the commands carrying this constant, while the same thread awaits others under `CollectionSweepSeconds`, `AlertPassCommandTimeoutSeconds`, `AnalysisCommandTimeoutSeconds` and two 300 s budgets. Nothing bounds the chain in aggregate. The pin now enforces the relationship that does hold — the constant must stay under the collection loop's own 15 s tick, whose delay is finish-to-start, so every second spent pushes the next tick's launches out one for one — and records that nothing in the service forbids 6. +- **The SQL-text pins over the object-stats, PVS and plan-correction reads are now insensitive to alias qualification and whitespace** ([#3217]) - so a semantically neutral rewrite no longer reds four untouched files, while a removed filter, ordering key or cap still does. Schema qualifiers and case are deliberately not normalised. `TheCountAndTheRowsShareTheirFilter` is untouched: it asserts textual identity across two statements, which is what the invariant is, and under the rewrite that broke the others it is the one pin that correctly fires. The database filter and the unused-first ordering are additionally asserted by running the query against a real PostgreSQL over a fixture planting two databases in one capture, since both claims are properties of the returned rows. +- **`CONTRIBUTING.md`'s build instructions name the deprecated Dashboard and CLI Installer projects at their real paths under `deprecated/`, matching what `README.md` already said** ([#3226]) - the two documents no longer disagree on the first command a new contributor runs. The three root scripts carrying the same stale paths — `build-dashboard.cmd`, `build-all.cmd` and `package-release.cmd` — are deleted rather than repaired: nothing outside themselves referenced them, no workflow invoked them, CI publishes no Dashboard or Installer artifact, and none of the three built Darling at all, so repairing the paths would have preserved a "release" script that omits the flagship product. `build-lite.cmd` is unaffected. +- **Sixteen fields across five MCP tools returned the monitored server's local wall clock, unmarked, beside naive-UTC fields on the same JSON object** ([#3206]) - `ToString("o")` on a `DateTimeKind.Unspecified` value renders no offset suffix, so the two frames serialize identically and nothing in the payload distinguishes them, and `server_properties.utc_offset_minutes` is exposed by no MCP tool, so a caller could not correct for it either. The de-skewed fields are `get_running_jobs`' `start_time` (the msdb Agent clock, against which `current_duration_seconds` is computed with `GETDATE()` on the next line), the six `blocked_`/`blocking_` transaction and batch stamps on `get_blocking` and Lite's `get_blocked_process_reports`, `get_index_usage`' `last_user_access`, `get_pvs_stats`' four ADR cleaner start and end times, and `get_plan_corrections`' `valid_since`, `last_refresh` and two action-initiated stamps. Payload field names are unchanged and only the frame is, and the stored frame stays local, because collector dedup and duration arithmetic are local-vs-local and converting at the read boundary is what the repo has settled on: Darling de-skews in SQL from a single-row `COALESCE` offset CTE, so a server that has not completed its first `server_properties` cycle still gets its rows, while Lite de-skews the row in C# at the MCP projection because DuckDB has no `make_interval` and the underlying reads are shared with WPF grids wrong on a *different* subset of columns. **No row selection moves** - every affected read windows, snapshots and orders on a naive-UTC column, which is asserted rather than assumed, so there is no stored-data migration and no schema change. The offset is exactly −240 minutes and fleet-wide, measured at 4 h behind the same collector run's `collection_time` on 42 of 42 servers, and confirmed end-to-end on live `get_pvs_stats` payloads across both production stores: a continuously-running off-row version cleaner read as four hours stale and resolved to under a minute after the correction, which is what rules out the cleaner really having last run four hours ago. `McpPayloadClockFrameDisciplineTests` pins the payload boundary none of the three existing clock guards could reach - the first has the right discriminator over the wrong roots, the second the right roots but hunts bare clock functions and disclaims projections, the third the right per-column reasoning but reads collector source only and never asks who consumes it - scoping every fact to a (read, column) pair with its collector evidence named, because the frame genuinely cannot be keyed on the column name. `AlreadyUtcReads` names the four sibling reads that must NOT acquire the de-skew, since applying it to a value already in UTC is the same defect with the sign flipped. +- **The product version is declared in exactly one place** ([#3222]) - six csproj files each carried their own ``, two of them three releases behind, and the release gate read a different file from the one the nightly named artifacts with, so bumping the gate's file shipped artifacts carrying the old version while bumping the nightly's file failed the gate on a correct build. Nothing compared them, and no test had ever read a workflow file. One `` now lives in `Directory.Build.props` and all 22 projects inherit it. The three deprecated projects were also still hand-setting `AssemblyVersion` / `FileVersion` / `InformationalVersion`, which #2113 deleted from the shipped three because a hand-set copy wins the stamp - that is why `deprecated/Dashboard/Dashboard.csproj` declared 3.6.0 while its binary stamped 3.3.0.0, in the very file the release gate took its judgement from. All nine version reads across the workflows and the root build scripts now name the one declaring file, two of which had been pointing at a path that stopped existing in #1612. Each of those reads now fails its own step when it reads nothing, which a pull request could not previously reveal because every consumer of the value was gated on the release event. Two existing guards needed following: the `darling` path filter now reaches Lite's project file, since the new pin enumerates it, and `CrossAppGuardCiGateTests` had pinned the ABSENCE of `Directory.Build.props` and now asserts all three imported build files by value. A release bump is now one line, and `ProductVersionDeclarationTests` fails the build on a second declaration, a re-added derived property, a reader naming another path, an unguarded reader, or a comment that pushes the text-scraping readers off the element. +- **`StoreSqlClockDisciplineTests` scanned store SQL one string literal at a time, so a predicate assembled by concatenation was split before the discriminator could see it** ([#3223]) - the pin flags store SQL comparing a naive collector timestamp against a bare clock function, because PostgreSQL resolves the mixed comparison at the store session's TimeZone — a documented one-hour predicate spanning five under `America/New_York`. Iterating literal bodies one at a time put the column in one body and the clock in another, so the comparison was in neither and the second body was not even SQL-shaped. The scan now reads maximal runs of literals welded by C# `+` concatenation, rendered as the string the runtime builds: nothing between a bare `+`, since inserting a space there would invent a token boundary in a statement merely split across source lines, and a single space where the glue carries an expression whose value is unknown. Reachability is measured rather than argued - 815 runs in the corpus weld two or more literals and **8 are units no member body was SQL-shaped enough to reach the discriminator**, four of them the `config_alert_log` reads this pin's own remarks name as its reason for existing, none of which had ever been read. Offenders are 1 before and 1 after. What is deliberately not welded is separate statements, separate arguments or elements, separate blocks and the two arms of a conditional, each held by a test whose fixture is modelled on a real corpus gap: a `+` at **both** ends of the gap, and a separator at expression **depth zero**. +- **`DarlingPgReadSqlParsesLiveTests` had a floor of 10 against a population of 49, and one shipped read was outside it** ([#3223]) - this is the suite that `PREPARE`s every shipped PostgreSQL read against a real server, and the defect it exists for - an ambiguous `LEFT JOIN` column throwing 42702 on every call for months while a dozen text assertions passed - is invisible to every other instrument, so a read that drops out of the population loses its only parse check. Its doc claimed twelve constants across nine readers; measured, **28 reader types ship 49 reads**, so a floor of 10 tolerated a 39-read drop. Discovery also filtered on `IsLiteral`, and a `static readonly string` composed by concatenation is not a literal, so `DarlingPgColumnStatsReader.CoverageEvidenceSql` - a shipped read executed at every column-stats coverage call - had never been parse-checked. The single floor is replaced by three clauses that fail differently: a ratchet at the live population, so growth never trips it while a removed read reds; a per-type clause, since one reader ships 9 and no total with slack can see a type emptying; and a source census over every `const string` and `static readonly string` declared in the reader files, the only clause whose denominator comes from outside reflection and so the only one that sees a read leave reflection's reach while its declaration sits in the file. The census runs **without a server**, where the old floor sat inside the container-gated test and checked nothing on the Windows `build` job. It keys by declaring type rather than file stem, because `CONTRIBUTING.md` endorses partial classes and a split reader file would fail *loudly* for a field both reflection and the parse check cover; it is anchored to a declaration rather than pattern-matched, because C# allows a local `const string` that reflection can never see; and it resolves the type through `CSharpMemberMap`, which now exposes `EnclosingType` beside `EnclosingMember` rather than carrying a private copy of the containment scan - a copy that had already diverged, treating an unterminated body as running to EOF where `EnclosingMember` escalates it to `Unknown`. +- **The Query Store liveness touch guard is one constant, six hours wide** ([#3189]) - the `last_seen` touch that keeps Query Store plan XML alive against the dimension GC was guarded by `interval '1 hour'` written THREE times, twice in `QueryStorePlanMap.TouchAndProbeSql` (the `touched` CTE and `dim_touch`) and once in `QueryStoreTextStore.TouchAndProbeSql` - so the map row and the dimension row were protected by copies that happened to agree, in a statement whose own summary calls their divergence structurally impossible. `QueryStoreLivenessTouchGuard` is now the only place the width is written, and it is DERIVED: the smaller of the two tables' prune margins (`QueryStorePlanMap.PruneMarginDays` = 1 day against `QueryStoreTextStore.PruneMarginDays` = 2) divided by a stated `MarginShareDivisor` of 4. Derived from the MARGIN and not from the retention horizon, deliberately - the margins are what `DarlingRetention`'s cutoffs reserve for a trailing stamp, while the retention term is the operator's and moves with `plan_content_retention_days`, so a guard sized as a fraction of the horizon would widen when retention widened while its actual budget did not move. The divisor is CHOSEN, not measured, and its doc says so: safety picks nothing here, since anything from 1 to 24 hours stays inside the margin and the hard bound is 2 days in the tightest reachable configuration. What it buys is WRITES, not examinations - the verdict SELECT returns a verdict for every reference in the batch regardless of the guard - and `last_seen` is indexed on all three tables, so every touch is a full non-HOT update; measured at the one-hour width, 22.1 M of them a day across the three tables, whose per-row ceiling the six-hour width divides by six. `ComputeMapCutoff`'s margin rationale is corrected with it: it justified the one-day margin by `TouchAndProbeSql` refreshing the map's stamp eagerly while the dim's was hourly-guarded, and the shipped SQL guards both, now at one width by construction. The `min`'s MEMBERSHIP is derived rather than stated, because an enumeration with no inclusion criterion is the next defect: a table contributes a term exactly when this touch writes its `last_seen`, and `PgStatementText.PruneMarginDays` is deliberately excluded - it is 2, so folding it in could not change the value, while `PgStatementText` has no guard site at all (its upsert advances `last_seen` on every conflict under a monotonicity guard) and its margin is a fact-outliving one wearing the same constant name. `query_plan_dim`, the third table the touch writes, needs no term because `MarginOrderingHolds` already pins the map's margin strictly inside the dimension's. The pin is the single-sourcing and the membership rather than the value - a source scan over both statement files that reds on any hard-coded interval, plus a project-wide guard-site sweep whose derived store types must equal the terms in the constant's own initializer. Both are mutation-proved on changes that move NO value: a correct-looking literal in the text store reds only the literal scan, folding in `PgStatementText`'s margin reds only the membership equality, and crippling either control makes its own mutation pass silently. - **`get_default_trace_events` returned `event_time` in the monitored server's local wall clock** ([#3198]) - every neighbouring timestamp — `collection_time`, `last_collection`, the XE `event_time` columns and the tool's own `as_of` — is UTC, so an event at 04:28 UTC on a server at UTC-4 rendered as 00:28 and read as having preceded the incident it coincided with. The read now de-skews the stored `StartTime` by the collected `server_properties.utc_offset_minutes` and both returns and windows the UTC value, matching the viewer's read of the same column and Lite's read of the same tool. - **Composed-panel event annotations from `default_trace_events` were both selected and plotted in the monitored server's local clock** ([#3198]) - on an axis bucketed in UTC and beside four annotation sources that are UTC — two frames in one chart. Each annotation source now declares its clock frame, and the compiler de-skews the server-local ones per server. - **Adding an Entra MFA server could fail with a Windows broker error that was recorded nowhere** ([#3196]) - `AddServerDialog`'s connection test kept `ex.Message` and never called `AppLogger`, so a failed Entra MFA sign-in left a dialog and no log entry; the reporter of #3196 could not send anything but a screenshot. The failure is now classified by which stage of the WAM broker handshake it died at - the application not supplying a window handle, the broker runtime not loading, or the broker running and refusing - and the whole inner exception chain is logged, since the driver wraps MSAL's exception which wraps the broker's and `Message` is only the outermost layer. The dialog names the resolved log directory, because that path is not guessable from a dialog that does not print it. `ServerManager`'s `SqlException` arm passes the exception rather than its message, matching the generic arm beside it - that is the path a saved Entra MFA server takes on every collection cycle. **The authentication failure itself is not fixed and cannot be from here**: `Microsoft.Data.SqlClient` forces the WAM broker on for any caller using the driver's own Entra application id, `UseWamBroker` is inert in that configuration, MSAL falls back to a browser only when the broker is unavailable rather than when an invoked broker refuses, and the broker's stated reason is redacted to `Context: (pii)` with no seam to enable it. Unverified against a live Entra-MFA tenant, which is the only place this class can be confirmed. -- **The store disk-pressure check stopped walking the whole store every five minutes to decorate an alert message** ([#3199]) - `SELECT pg_database_size(current_database())` ran on the collection loop's serial thread under `ServiceCommandDeadlines.SerialLoopSeconds` = 5 s, a bound floored on **6.2 ms measured against a 4.05 GB store**. That function stats every file in the database directory, so its cost follows the store rather than the single row it returns: on a **225 GiB** production store — 56x the fixture — thirteen samples spanned **2,090-3,745 ms**, leaving 1.3-2.4x of the ~806x the derivation claimed. **None of the thirteen crossed 5 s, and that is the finding.** Cancels for the statement ran mean **5.8/day** against a nominal ~288 iterations (**2.0%**), so the eight cancelled statements that found this were the visible tail and the other 98% cost ~2.5 s each while leaving no trace anywhere - under the deadline so no cancel, not a collector run so no `collection_log` row - which is **~11.9 min/day** of a single-threaded loop's wall time spent computing a number the store already held. `collect.store_metrics` has carried the whole-store size hourly since V53, so the check now reads the newest recorded row: **0.101 ms** cold, three buffers, `Index Scan Backward` on the existing V53 index, **no new index**. **The walk is relocated rather than eliminated** and the honest claim is about which path pays it - `StoreSelfMetrics.StoreInsertSql` runs `pg_database_size` itself, so 312 executions a day become 24, all of them under the 300 s budget #2317 sized against this same store's sizing queries; the ~31,000x is the single read's latency and never the change's overall effect. Raising the deadline was arithmetically closed rather than merely unattractive: the shared constant is pinned at `10 sites × N < 60 s`, so `N < 6`, and a private constant for this one call caps at 14 s with nothing measured saying 14 is enough. The `object_kind` literal became one const with six consumers, because a reader filtering on a kind the writer stopped writing returns **zero rows rather than erroring** and every consumer renders that as the same null a never-swept store produces. `SerialLoopStoreSizeSourceTests` pins the **category** - no command on the serial loop may run a read whose cost scales with the store - over the population `StartupCommandTimeoutTests` already owns, projected rather than re-listed so it cannot cover nine members while claiming ten: that pin counts DEADLINES, which is why it stayed green across the whole life of this defect and would go green again the moment a size-scaling read came back. **Two of the new pins could not fail and were found by mutating the instrument rather than the subject** - an `Assert.DoesNotContain` read over source whose transform had already removed the pattern it asserted absent, and a negative census whose loop body never runs on a population that correctly contains nothing - so both now share their scan with a positive control over a planted body. +- **The store disk-pressure check stopped walking the whole store every five minutes to decorate an alert message** ([#3199]) - `SELECT pg_database_size(current_database())` ran on the collection loop's serial thread under `ServiceCommandDeadlines.SerialLoopSeconds` = 5 s, a bound floored on **6.2 ms measured against a 4.05 GB store**. That function stats every file in the database directory, so its cost follows the store rather than the single row it returns: on a **225 GiB** production store — 56x the fixture — thirteen samples spanned **2,090-3,745 ms**, leaving 1.3-2.4x of the ~806x the derivation claimed. **None of the thirteen crossed 5 s, and that is the finding.** Cancels for the statement ran mean **5.8/day** against a nominal ~288 iterations (**2.0%**), so the eight cancelled statements that found this were the visible tail and the other 98% cost ~2.5 s each while leaving no trace anywhere - under the deadline so no cancel, not a collector run so no `collection_log` row - which is **~11.9 min/day** of a single-threaded loop's wall time spent computing a number the store already held. `collect.store_metrics` has carried the whole-store size hourly since V53, so the check now reads the newest recorded row: **0.101 ms** cold, three buffers, `Index Scan Backward` on the existing V53 index, **no new index**. **The walk is relocated rather than eliminated** and the honest claim is about which path pays it - `StoreSelfMetrics.StoreInsertSql` runs `pg_database_size` itself, so 312 executions a day become 24, all of them under the 300 s budget #2317 sized against this same store's sizing queries; the ~31,000x is the single read's latency and never the change's overall effect. A private constant for this one call caps at 14 s with nothing measured saying 14 is enough. The `object_kind` literal became one const with six consumers, because a reader filtering on a kind the writer stopped writing returns **zero rows rather than erroring** and every consumer renders that as the same null a never-swept store produces. `SerialLoopStoreSizeSourceTests` pins the **category** - no command on the serial loop may run a read whose cost scales with the store - over the population `StartupCommandTimeoutTests` already owns, projected rather than re-listed so it cannot cover nine members while claiming ten: that pin counts DEADLINES, which is why it stayed green across the whole life of this defect and would go green again the moment a size-scaling read came back. **Two of the new pins could not fail and were found by mutating the instrument rather than the subject** - an `Assert.DoesNotContain` read over source whose transform had already removed the pattern it asserted absent, and a negative census whose loop body never runs on a population that correctly contains nothing - so both now share their scan with a positive control over a planted body. - **`CompressionPhaseGuardMinutes` is a DECLARED four-minute width rather than the light-refresh ceiling rounded up** ([#3188]) - That member IS UnboundedLightRefreshSeparationMinutes, so the width decided the adjacency the twelve light refreshes ran under and their runtimes were a measurement of that adjacency - the grid's shape was a function of a measurement the shape produced, in both directions. The measurement is now checked against the width, so a light refresh past 240 s goes red instead of silently widening a band subtracted from a fixed hour. Nothing moves. - **The two refresh-ceiling constants state one estimator and one closure rule, word for word, and name themselves PREFIX MAXIMA rather than maxima over closed populations** ([#3188]) - A read instant closes the read and not the series, so a later run of the same regime joins the population and can exceed the value - which is what took one of them from 896 s to 1,134 s in an evening with no code change. The refresh-slot Error band and both constants now state the same proposition from its two sides, and the staleness Warning stopped saying that a closed population had been overtaken. No constant value moves and no job moves. - **`query_store`'s `sql_duration_ms` was mostly store time, under a column documented as the monitored server's** ([#3192]) - the enumerated driver's per-item stopwatch wraps the whole `readItem` closure, and for `query_store` that closure round-trips the store to decide what plan XML and statement text are already held before writing back what came off the target. One production run: 124,972 ms of `sql_duration_ms` of which 107,334 ms - **86%** - was the two store probes, against 6,494 ms of plan-plus-text target time; fleet-wide the probe is 55.4% of `plan_fetch` and 80.6% of `text_fetch`. So `get_collector_cost`'s tens of millions of ms/day of "target-side query DURATION" (**64.7 M ms/day** measured over 7 days on one production store, where #3192's ~45.8 M does not reproduce - a different store or window; the ratio is the claim, not the absolute), the `get_collection_log` projection comment stating that a collector slow because the target is slow needs work on that server, and the web grid's **"On Server"** column header were all pointing at the monitored servers for time spent in the monitoring store - the ~6.5 : 1 anti-target bias V110 declined a rollup shape over, arriving through the parent column instead. **The column is NOT re-based**, and the reason is that the past cannot follow it: `CollectorRunResult.SqlMs` also feeds `collect.collector_cost`, nine columns with no phase split, flushed hourly from an in-memory accumulator rather than aggregated from `collection_log` - so there is nothing there to subtract, no source to re-derive from, and a re-based column would leave 90 days meaning one thing and every row after meaning another under a self-alert whose baseline window is 14 days, in which a real regression is measured against an inflated baseline. Routing the probe into `store_duration_ms` is worse rather than kinder: that column is the binary COPY and nothing else, and it is the measurement `ServiceCommandDeadlines` derives the 10 s COPY deadline from ("worst of 200 runs, 1.53 s"), so absorbing a 54 s probe would widen a deadline the sweep watchdog depends on. Instead **`sql_store_ms`** derives the store share from the V110 columns already on the row, so nothing is stored, nothing vanishes, a 124,972 ms run still sums to 124,972 ms, and the attribution applies **retroactively to every row since V110**. It is a **FLOOR**, said out loud on the property and on both tool descriptions: the per-item watermark refresh is also a store read inside the same stopwatch, the enumerated branch never raises V108's measured flag so `watermark_ms` is NULL on exactly these rows, and that component is recorded nowhere - which makes `sql_duration_ms - sql_store_ms` an *upper* bound on target time. `get_collector_cost` can only carry the caveat and point at the per-run tool, and that bound is stated as the limit of what any fix could reach there. Scope checked rather than assumed: the server-scoped and per-database paths both keep every store round trip *outside* their sql stopwatch, deliberately and already pinned, so the enumerated path is the only one. Found in the same doc block: the driver's budget doc claimed the per-item budget is null for "every collector but `query_store`" in **three** places, one at the runner's own call site - four definitions declare one and two of those enumerate, so `plan_correction` arrives non-null too; corrected and pinned by deriving the count from `CollectorCatalog.All`. Proven red fourteen ways with a green control and the baseline printed, including the option-1 change applied at the cost surface, and including one pin that had to be rebuilt after its first mutation build-failed - the shared return statement passes a variable, so asserting the absence of a literal there was a check nothing could turn red. @@ -3549,9 +3567,25 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 [#3185]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/3185 [#3187]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/3187 [#3188]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/3188 +[#3189]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/3189 [#3192]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/3192 [#3193]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/3193 [#3196]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/3196 [#3198]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/3198 [#3199]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/3199 [#3200]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/3200 +[#3204]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/3204 +[#3206]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/3206 +[#3207]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/3207 +[#3208]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/3208 +[#3211]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/3211 +[#3214]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/3214 +[#3217]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/3217 +[#3219]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/3219 +[#3221]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/3221 +[#3222]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/3222 +[#3223]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/3223 +[#3224]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/3224 +[#3226]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/3226 +[#3230]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/3230 +[#3231]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/3231