The Query Store liveness touch updates rows in place: drop the last_seen indexes, fillfactor 90 (#4250) - #4472
Merged
Merged
Conversation
…en indexes, fillfactor 90 (#4250) - V149: drops idx_query_store_plan_map_last_seen and idx_query_store_text_last_seen, sets fillfactor = 90 on both tables. - Removes the CREATE INDEX ... last_seen lines from both tables' CreateTableSql so startup convergence does not recreate them. - Adds the viewer probe sentinel (negative: index absent + fillfactor=90 present), StorageVersion bump, and the full rung recipe. - New live pin file QueryStoreLivenessHotTouchLiveTests.cs: schema pin (RED verified against pre-fix dev), convergence pin (RED verified via a re-add mutation), HOT pin (n_tup_hot_upd increments on the second touch of a settled row). - Updates QueryStoreTextStoreTests and ReadLatencyFlushLiveTests for the new top rung. Refs #4250.
…delete (#4250 item 3) The Query Store plan-map and statement-text prunes each ran three sequential scans per statement without their last_seen index (two min() subqueries plus the DELETE) — on the busiest measured store, the text table's slice alone read on the order of tens of GB per call. A new unordered, row-capped builder (a sibling of the plan dimension's ordered one) replaces all three scans with one: a fixed cutoff means every matched row is already expired regardless of which ones a batch happens to take, so dropping the ORDER BY cannot strand a floor the way it could against a moving cutoff. The cap (300,000) clears the busiest single day observed in the field's last_seen age histogram for either table in one batch, so a normal run issues one statement, and a multi-day backlog still drains in bounded per-batch transactions.
Merge origin/dev (#4467), and add a live loop test for #4250 item 3's row-capped prune through DarlingRetention.PurgeAsync (not the builder or drain loop called directly): seeds both liveness-touched tables past 2x a small test cap plus in-cutoff rows, asserts expired rows gone / in-cutoff rows kept / batch count matches the drain contract, and verifies a stop-after-one-batch mutation turns it RED. Adds a test-only cap-override seam to PurgeAsync defaulting to the shipped constant, so production is unchanged.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Refs #4250.
Why
The Query Store liveness touch (
QueryStorePlanMap.TouchAndProbeSql,QueryStoreTextStore.TouchAndProbeSql) refresheslast_seenon every referenced plan and statement, on every collector cycle. On a busy store the field measured 1.43 M touches/day, 0 HOT updates onquery_store_plan_map, and roughly 7-8.5 KB of WAL per touch — non-HOT because the only column the touch ever writes besideslast_seenis an unindexed hash column, butlast_seenitself carries a btree index, so every touch also rewrites that index entry and typically moves the tuple to a new page.A production read (no store named) showed
query_store_plan_map'slast_seenindex at 881 MB — bigger than the table's own heap (709 MB) — with 178 M updates and 0 HOT.query_store_text's index was 1.2 GB against a 9.4 GB table (mostly TOAST), 320 M updates, 91k HOT.Reader inventory
git grep -n "last_seen"over the storage and service code turns up exactly two live readers of either index:QueryStorePlanMap.TouchAndProbeSql/QueryStoreTextStore.TouchAndProbeSql— the liveness touch itself, the thing this change fixes.QueryStorePlanMap.PruneSql/QueryStoreTextStore.PruneSql— the retention purge's time-slicedDELETE ... WHERE last_seen < $1 AND last_seen >= (SELECT min(last_seen) ...) AND last_seen < (SELECT min(last_seen) ...) + INTERVAL 'N days', run twice a day by the retention sweep.No reader orders, filters, or range-scans
last_seenfor anything else — every other hit across the service, storage, viewer and MCP code is a join on the primary key or an unrelated table.What changes
Migration V149:
Plus the removal of the
CREATE INDEX IF NOT EXISTS ... last_seenlines from both tables'CreateTableSql, so startup convergence doesn't recreate what the migration drops.The retention prune for both tables is also rewritten. The old
PruneSqlran three sequential scans per statement without the dropped index (twomin()subqueries plus theDELETE). It's replaced byDarlingRetention.UnorderedRowCappedDeleteSql, a sibling of the plan dimension's row-capped, ordered builder but with noORDER BY:DELETE FROM t WHERE ctid IN (SELECT ctid FROM t WHERE last_seen < $1 LIMIT n). Dropping the ordering is safe here specifically because the cutoff is fixed per run rather than recomputed from the table's own remaining minimum: every row the predicate matches is already expired regardless of which ones a given batch happens to take, and each batch strictly shrinks the matching set, so the drain still terminates — it just can't strand a floor the way an unordered cap against a moving cutoff could. The cap (300,000 rows) is sized from a field read of thelast_seenage histogram so a normal run clears the whole day's retirement in one batch for both tables, while a multi-day backlog still drains in bounded per-batch transactions via the existing drain loop.Lock / transaction
PgMigrations.MigrateAsyncwraps every rung in one transaction, soCREATE/DROP INDEX CONCURRENTLYis not available here (Postgres refuses it inside a transaction) — same constraint the V142 rung's doc comment already documents for a different table. Both statements in V149 are metadata-only:DROP INDEXon a btree does not touch the heap, andALTER TABLE ... SET (fillfactor = ...)only changes a catalog reloption, so the shortACCESS EXCLUSIVEeach one takes is a catalog update rather than a rewrite. Nothing else in the migrate session holds a competing lock on either table at that point.Measurements
Measured on a rig (
timescale/timescaledb:2.30.1-pg18), two databases —before(btreelast_seenindex, default fillfactor) andafter(V149applied AFTER a 1.1 M-row population and
VACUUM ANALYZE, so existing pagesstart full, as on an upgraded store). Four touch cycles per table
(
CHECKPOINT, an UPDATE over the due half of the rows,VACUUMbetweencycles, rows aged back 13h to simulate 2 touches/row/day).
Touch WAL/FPI/HOT%, steady state (cycle 4):
query_store_plan_mapquery_store_textquery_store_textshows the larger win: HOT climbs to 65% by cycle 4 and thefillfactor-90 + no-index change is not only a
plan_mapfix.Prune cost without the index (field-scale numbers from a real store, read
only, no store named):
query_store_plan_mapmain fork 709 MB / 4.04 M rows;query_store_textmain fork 4,992 MB + TOAST 4,355 MB / 6.86 M rows. Oneslice of the current
PruneSql(2min()subqueries + the DELETE), warmcache: plan_map 1.31 s indexed vs 1.68 s without (544,248 buffers, ~4.25 GB);
text 10.24 s indexed vs 7.46 s without, but the no-index text slice touches
~3.83 M buffers ≈ 29 GB logical (2.27 M read from outside shared_buffers).
The drain loop runs ~2 slices/table, 2 runs/day, so
query_store_textalone is on the order of tens of GB/day up to roughly 60-120 GB/day of extra
reads without its index, depending on cache residency. No measured slice
anywhere near the ~10-minute stop condition (worst case 10.24 s).
A rig-measured one-scan prune shape (the
UnorderedRowCappedDeleteSqlnow shipped, dropping the two
min()subqueries and the upper-boundre-scan) reads roughly 2.5-3x fewer buffers than the old 3-scan shape at
every size tested on the rig — field-scale copies of both tables, no index,
cutoff selecting ~20% of rows:
query_store_plan_map290,476 buffers(current 3-scan) vs 117,448 (one-scan);
query_store_text1,089,958 buffers(current) vs ~497,812-632,414 (one-scan, two runs on the same rig gave a
noisy 3,528 ms/460 ms wall-time split but a stable buffer count). Scaled to
the field's
query_store_textfigures above (tens of GB/day, up to roughly60-120 GB/day), the one-scan shape cuts that to roughly 20-40 GB/day —
real, but well under half.
The field saving is not projected here. The rig's per-touch WAL (0.4–2.3 KB) is far below the field's 7–8.5 KB, because the rig's tables fit in cache and many touched rows share a page, so the rig can't reproduce the field's full-page-image-heavy shape. On the field, a non-HOT touch also writes into the primary key and the
last_seenindex (881 MB and 1.2 GB on the busiest store), and a HOT touch writes into no index at all. So the field saving could be larger or smaller than the rig's 23% / 35%.Trade shipped: both
last_seenindexes are dropped (the touch's real,large WAL/HOT win on both tables) AND the prune moves to the one-scan
shape, so the no-index prune's own read cost — real, but bounded to twice a
day and never near the 10-minute stop condition at any measured scale —
is itself cut by more than half rather than left at the 3-scan number.
The field saving is unmeasured until the 24-hour
pg_stat_statementsdeltaafter install.
Pins
New file
Darling/Darling.Tests/QueryStoreLivenessHotTouchLiveTests.cs, five facts, all run against a rig container (timescale/timescaledb:2.30.1-pg18) and all green on the branch:TheRungIsRegisteredAtTheTopOfADenseLadder— the ladder shape invariant.TheProbeMapsAFullyMigratedStoreToThisTopRung— the viewer probe/map/arm three-way agreement.AfterMigrate_NeitherLastSeenIndexExists_AndBothTablesCarryFillfactor90— the LIVE schema pin.StartupConvergence_RunTwice_DoesNotRecreateEitherIndex— running each table'sCreateTableSqltwice after migrate must not bring the index back.TwoTouchesOfTheSameRow_TheSecondTouchReportsAHotUpdate— the payoff:n_tup_hot_updincrements on the settled-row touch, viapg_stat_force_next_flush()plus a short settle.RED on origin/dev: ran the same file (verbatim) against a detached worktree at dev's tip (
6c3b026ed) with the same rig. All 5 facts fail — 2 by design failure of the assertions (index present, no fillfactor reloption), 3 by reflection failures (nohasHotLivenessTouchparameter, noquery-store-liveness-hot-touchrung).Mutations (both reverted after verifying RED, confirmed green again after revert):
CREATE INDEX ... idx_query_store_plan_map_last_seenline toQueryStorePlanMap.CreateTableSqlturnsStartupConvergence_RunTwice_DoesNotRecreateEitherIndexRED.ALTER TABLE collect.query_store_plan_map SET (fillfactor = 90);line from V149's SQL turnsAfterMigrate_NeitherLastSeenIndexExists_AndBothTablesCarryFillfactor90RED.Totals:
dotnet buildforDarling.TestsandLite.Testsboth 0 warnings / 0 errors. Ran (all green):QueryStoreLivenessHotTouchLiveTests(5),ReadLatencyFlushLiveTests(5, updated for the new top rung),QueryStoreTextStoreTests(updated for the dropped index),MigrationLadderPins,DocCommentHygieneTests,ViewerDataServiceTests— 103 total, 0 failed.Gated tests (test-only nightly)
QueryStoreLivenessHotTouchLiveTests(new, this PR)ReadLatencyFlushLiveTests(updated top-rung claim)QueryStoreTextStoreTestsMigrationLadderPinsDocCommentHygieneTestsViewerDataServiceTestsDarling.TestsandLite.Testsboth targetnet10.0-windows: they build here but cannot run on this machine; CI decides them.Prune pins (this pass)
Added to
DarlingRetentionTests: the new builder's exact SQL shape (noORDER BY, nomin(), theLIMITandctid INpresent, right table/column),and a source-anchored check that both the map and text purge call sites use
the new builder with the shared cap as
batchSizeand leaveadaptiveRowCapTimeColumnunset (that parameter is the plan dimension's#4130 retry-at-half-cap behavior, not measured as needed here — both tables
clear their whole steady-state backlog in one batch at the sized cap).
Moved the old
QueryStoreTextStore.PruneSqlshape pin out ofQueryStoreTextStoreTestssince that method no longer exists; bothQueryStorePlanMap.PruneSqlandQueryStoreTextStore.PruneSqlare retired(no other caller — verified with
git grep).Mutations (reverted after verifying RED, confirmed green again after
revert): restoring
ORDER BY {timeColumn}inUnorderedRowCappedDeleteSqlturns the shape pin RED (
UnorderedRowCappedDelete_HasNoOrderByAndNoMinSubqueryfailed, Total: 47, Failed: 1); dropping the
LIMITclause turns it RED.Live loop test (added this pass)
New file
Darling/Darling.Tests/QueryStoreLivenessTablePruneLiveTests.cs, run against atimescale/timescaledb:2.30.1-pg18rig via the same#1776 own-store/ScratchPostgrespattern as the other live pins.
A test-only seam was added to
DarlingRetention.PurgeAsync: an optionallivenessTouchedTablePruneRowCapparameter that defaults to the shippedLivenessTouchedTablePruneRowCapconstant (300,000). Both production callers (the daily sweepand the on-demand
purge_nowcommand) omit the argument, so production always runs the shippedcap unchanged; the test passes a small cap (1,000) so a multi-batch drain can be exercised
without seeding ~650k rows per table. The default is pinned twice: once in the new file, once in
DarlingRetentionTests.LivenessTouchedTablePruneRowCapSeam_DefaultsToTheShippedConstant.Seed: 2,137 expired rows per table (
1,000 * 2 + 137, two full batches plus one short one atthe test cap) 400 days back, plus 5 rows 1 hour back that must survive.
Run:
DarlingRetention.PurgeAsync(postgres, timescaleAvailable: false, purgeLog, ct, livenessTouchedTablePruneRowCap: 1000)— the real product entry point.Assert: every expired row gone, every in-cutoff row kept, and each table's batch count is
exactly 3 (
ceil(2137 / 1000)), read fromPurgeOneAsync's own "Retention purge drained ... in{Batches} batch(es)" log line via a
CapturingTestLogger.Mutation (reverted after verifying RED, confirmed green again after revert): temporarily
made
DrainBatchesAsyncbreakunconditionally after its firstexecuteBatchcall — turns thenew test RED (only 1,000 of 2,137 expired rows removed per table, batch count 1 not 3). Reverted;
151 tests green across the classes run in this pass.
RED on origin/dev: not run as a separate check — the test's
livenessTouchedTablePruneRowCapargument does not exist on dev's
PurgeAsyncsignature, so the file fails to compile there(compile-failure RED, the same shape already used for other #4250 pins in this PR).
Prune at the shipped cap
Measured on a 2.5M-row rig copy of both tables in this PR's V149 shape (no
last_seenbtree,
fillfactor = 90), withquery_store_textcarrying a distinct ~700-byte value perrow, against the exact
DarlingRetention.UnorderedRowCappedDeleteSqlstatement at theshipped cap of 300,000. Buffers are shared hit + shared read; 1 buffer = 8 KB.
(a) Steady state (~110,000 rows expired):
One batch clears the whole steady-state backlog on both tables. It deletes fewer rows than the cap, so the drain loop stops after it (the loop's contract: a short batch ends it; the live loop test pins this). There is no second, empty-result call in the product.
(b) Backlog (~2,000,000 rows expired, 7 batches, no VACUUM between batches):
plan_map buffers rise slightly batch over batch (≈4.82 GB → ≈5.00 GB, +4%) as earlier
batches' dead tuples accumulate; per-batch time stays in the 390–430 ms range. text's
buffers stay flatter (≈5.2–5.6 GB) but per-batch time is noisier (642–1,283 ms), driven by
TOAST reads for the payload column rather than the row cap. No batch anywhere near a
command timeout.
(c) Old shape (pre-#4250 3-scan
PruneSql) on the same steady-state slice:On plan_map the two shapes read about the same (≈1.8 GB vs ≈2.0 GB), and the new one is 4× faster. text's old shape reads ~1.9x this PR's shape for the same expired slice (992,223 vs.
510,033 buffers), from the two extra
min()sequential scans the oldPruneSqlran beforeits
DELETE. On the production store's full-size text table (6.86 M rows), the old shape measured roughly 29 GB of logical buffers per slice.Scale to field size: field main forks are
query_store_plan_map709 MB andquery_store_text4,992 MB main + 4,355 MB TOAST (9,347 MB total), giving size ratios of≈2.36x (plan_map) and ≈4.03x (text) over this rig's tables. Scaling each shape's worst
observed single-batch time by that ratio: plan_map steady-state ≈0.45 s, backlog ≈0.93 s;
text steady-state ≈1.95 s, backlog ≈5.17 s. Every projected batch stays roughly two orders
of magnitude under the 300-second command timeout.
CHANGELOG entry
SECTION: Changed
ENTRY:
last_seenindexes are dropped (store migration V149), and their retention prune reads each table once per run.REF:
[The Query Store liveness touch updates rows in place: drop the last_seen indexes, fillfactor 90 (#4250) #4472]: The Query Store liveness touch updates rows in place: drop the last_seen indexes, fillfactor 90 (#4250) #4472