Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -45,7 +45,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

- **The database-state heal tests no longer assume a best-effort maintenance cycle always runs** ([#2266] item 2, root-caused from source) - `RebaselinedByHandDuringAnOutage_HealsOnceTheDatabaseRecovers` failed on a PR whose diff could not reach it and passed on a re-run of the same commit, reporting `ExpectedState = SUSPECT` alongside `StateDesc = ONLINE`. That pair is only reachable one way: `GetDatabaseStateDeviationsAsync` performs its seeding, the [#2189] heal, the [#2203] forget and the prune inside a block that opens the write connection with a **5-second** lock acquisition and, on `TimeoutException`, skips the entire maintenance block while still running the deviation read - which is deliberate and documented, because skipping is the only lossless option when archival holds the lock. That write lock is **static, shared by the whole process**, and xunit runs test classes in parallel, so another class can hold it long enough for a cycle to skip its maintenance. The test was therefore asserting that the heal lands in ONE cycle, which the design does not promise; it now sweeps until the expectation settles, bounded, which is the actual contract. It cannot mask a regression: a genuinely broken heal never settles, every cycle runs, and the caller's own assertion fails on the final result with its own message exactly as before. Ruled out first: the fixture mints a unique temp directory per instance so no two classes share a store file, and the server id is used by no other class - the shared resource is the LOCK, not the data. **Not fixed here, and worth knowing**: that `catch (TimeoutException)` logs nothing, so in production a sustained contention window means baselines quietly stop being seeded and healed with no evidence anywhere, which matters because [#2189] exists precisely because an unhealed baseline inverts the alert permanently. Item 1 of that issue (the sub-second scale-test comparison) is deliberately untouched: every candidate fix needs the jitter distribution, and guessing a threshold is how an intermittent test stops looking broken without becoming correct.

- **`get_top_queries_by_cpu` can rank a procedure's dynamic SQL as ONE statement: `group_by: "host_object"`** ([#2235]) - `query_hash` is a SHAPE hash, so dynamic SQL built with per-value literals fragments one logical statement across as many hashes as there are literal sets. Measured on `prod-pos-use2-apex-01`: **21 hashes** for a single `API.GetInventoryWithLabsV5` `insert #result` statement, whose fragments were **58-65% of the instance's worker_time** in every window sampled - while the hash never entered the 168-hour top 20, and the per-query ranking as a whole accounted for roughly a **tenth** of the box's CPU. A top-N-by-hash list cannot surface that no matter how large N is, and nothing in the output said so. Setting `group_by: "host_object"` collapses every statement of a hosting procedure or function into one row, which is what makes the real consumer rank first; `distinct_query_hashes` reports how many hashes the row rolled up (21, in the reported case) and is the number that explains why the default ranking missed it, with a `rollup_note` saying so in words. `query_hash` and `query_text` in a rolled-up row are one representative fragment, exactly as `query_text` already is when `distinct_texts > 1`. **Ad-hoc statements keep their per-hash grouping in BOTH modes, and that is the load-bearing part**: ad-hoc rows carry `host_object_name = NULL`, so a bare `GROUP BY host_object_name` would pool every unrelated ad-hoc statement in a database into one meaningless row - a worse attribution bug than the one being fixed - and the representative-text lookup would start serving an unrelated statement's text. The grouping key keys those rows on their own `query_hash` instead, identical to the default read. The per-hash grouping ([#2012] stage 2) stays the DEFAULT and is unchanged: two procedures sharing a hash genuinely are different work, which is why that split exists, so this is an additional lens rather than a replacement. Implemented as a sibling SQL const rather than a built clause because Postgres cannot parameterize `GROUP BY` and every read here is a public const so the suite can pin its dialect without a live store; an unrecognised `group_by` is rejected rather than silently falling back, since a caller who asked for a rollup and got a per-hash ranking would read it as "this procedure is not hot", which is the exact wrong conclusion. `group_by` is trailing and optional, so Lite-shaped calls are unaffected, and it is deliberately absent from `get_top_procedures_by_cpu` (already keyed on the object) and `get_query_store_top` (keys on `query_id`, which does not fragment). Pinned by a live-Postgres test that asserts the collapse, the ad-hoc non-pooling, per-row text correctness, and that total CPU and executions are CONSERVED across both groupings - a rollup must redistribute attribution, never invent or lose it.
- **`get_top_queries_by_cpu` can rank a procedure's dynamic SQL as ONE statement: `group_by: "host_object"`** ([#2235]) - `query_hash` is a SHAPE hash, so dynamic SQL built with per-value literals fragments one logical statement across as many hashes as there are literal sets. Measured on `prod-sql-use2-alpha-01`: **21 hashes** for a single `API.GetInventoryWithLabsV5` `insert #result` statement, whose fragments were **58-65% of the instance's worker_time** in every window sampled - while the hash never entered the 168-hour top 20, and the per-query ranking as a whole accounted for roughly a **tenth** of the box's CPU. A top-N-by-hash list cannot surface that no matter how large N is, and nothing in the output said so. Setting `group_by: "host_object"` collapses every statement of a hosting procedure or function into one row, which is what makes the real consumer rank first; `distinct_query_hashes` reports how many hashes the row rolled up (21, in the reported case) and is the number that explains why the default ranking missed it, with a `rollup_note` saying so in words. `query_hash` and `query_text` in a rolled-up row are one representative fragment, exactly as `query_text` already is when `distinct_texts > 1`. **Ad-hoc statements keep their per-hash grouping in BOTH modes, and that is the load-bearing part**: ad-hoc rows carry `host_object_name = NULL`, so a bare `GROUP BY host_object_name` would pool every unrelated ad-hoc statement in a database into one meaningless row - a worse attribution bug than the one being fixed - and the representative-text lookup would start serving an unrelated statement's text. The grouping key keys those rows on their own `query_hash` instead, identical to the default read. The per-hash grouping ([#2012] stage 2) stays the DEFAULT and is unchanged: two procedures sharing a hash genuinely are different work, which is why that split exists, so this is an additional lens rather than a replacement. Implemented as a sibling SQL const rather than a built clause because Postgres cannot parameterize `GROUP BY` and every read here is a public const so the suite can pin its dialect without a live store; an unrecognised `group_by` is rejected rather than silently falling back, since a caller who asked for a rollup and got a per-hash ranking would read it as "this procedure is not hot", which is the exact wrong conclusion. `group_by` is trailing and optional, so Lite-shaped calls are unaffected, and it is deliberately absent from `get_top_procedures_by_cpu` (already keyed on the object) and `get_query_store_top` (keys on `query_id`, which does not fragment). Pinned by a live-Postgres test that asserts the collapse, the ad-hoc non-pooling, per-row text correctness, and that total CPU and executions are CONSERVED across both groupings - a rollup must redistribute attribution, never invent or lose it.

- **The Query Store tick and the Query Store backfill no longer run against one server at the same time** ([#2165]) - the per-tick `query_store` collection and the [#2058] first-contact backfill were independent loops with no per-server coordination at all, and both do heavy Query Store text extraction. Dogfood evidence from a 4-core multi-tenant box mid-consolidation: a 64 MB backfill slice for a freshly restored database ran concurrently with the tick's collection of a SIBLING database - a 12:50:58 backfill ship overlapping a 12:51:09 tick completion - so roughly 128 MB of extraction was in flight at once on the box least able to afford it. That overlap is not bad luck: a big catalog arriving is exactly what triggers BOTH the backfill and budget-bound tick passes, so the two loops collide precisely when the server is already drowning. A per-server gate now excludes them, in both apps - Darling's `DarlingWorker` keyed by server id, Lite's `RemoteCollectorService` keyed by server - sharing one `QueryStoreServerGate` primitive beside `AbandonableStep` (that one bounds how long a step may hold a loop; this one bounds what may run beside it). **Nothing ever waits, deliberately.** Both sides try-acquire with a zero timeout and SKIP on failure, because these are shared fleet loops: an in-flight slice runs to a 180-300 second abandonment deadline, so a blocking acquire would let one slow server stall collection for the entire fleet - the [#2148] wedge arriving through a lock instead of a hang. **Skipping is safe for this collector specifically** because its window is a watermark ([#1960]): the next pass resumes from the same boundary, so a skipped pass defers rows rather than dropping them - which is also why the gate must not be reused for a collector whose window is wall-clock derived. The "tick wins, backfill defers" bias is realized by CADENCE rather than preemption (stopping a statement already running on the monitored server would mean killing it): the tick retries on its own ~1-minute interval against the backfill's 5, so it recovers five times faster from a collision, and a slice is byte-budgeted so it is short in the healthy case. The backfill takes the gate OUTSIDE its `AbandonableStep`, so an abandoned-but-still-wedged slice keeps the gate closed - the statement is genuinely still running on the server and the tick must keep yielding to it. Built on an interlocked flag rather than a `SemaphoreSlim`: a gate that never waits needs none of what a semaphore provides, and would otherwise own one undisposed kernel object per monitored server forever. Leases are idempotent on dispose, because a stray second `Dispose()` would otherwise clear a flag the OTHER loop had since taken and let both run at once - the exact condition being prevented, reached from the wrong direction.

Expand Down
Loading
Loading