Skip to content

The Query Store liveness touch updates rows in place: drop the last_seen indexes, fillfactor 90 (#4250) - #4472

Merged
erikdarlingdata merged 7 commits into
devfrom
fix/4250-hot-liveness-touch
Sep 27, 2026
Merged

erikdarlingdata merged 7 commits into
devfrom
fix/4250-hot-liveness-touch

Conversation

@erikdarlingdata

@erikdarlingdata erikdarlingdata commented Sep 27, 2026 •

Copy link
Copy Markdown
Owner

Refs #4250.

Why

The Query Store liveness touch (QueryStorePlanMap.TouchAndProbeSql, QueryStoreTextStore.TouchAndProbeSql) refreshes last_seen on every referenced plan and statement, on every collector cycle. On a busy store the field measured 1.43 M touches/day, 0 HOT updates on query_store_plan_map, and roughly 7-8.5 KB of WAL per touch — non-HOT because the only column the touch ever writes besides last_seen is an unindexed hash column, but last_seen itself carries a btree index, so every touch also rewrites that index entry and typically moves the tuple to a new page.

A production read (no store named) showed query_store_plan_map's last_seen index at 881 MB — bigger than the table's own heap (709 MB) — with 178 M updates and 0 HOT. query_store_text's index was 1.2 GB against a 9.4 GB table (mostly TOAST), 320 M updates, 91k HOT.

Reader inventory

git grep -n "last_seen" over the storage and service code turns up exactly two live readers of either index:

  • QueryStorePlanMap.TouchAndProbeSql / QueryStoreTextStore.TouchAndProbeSql — the liveness touch itself, the thing this change fixes.
  • QueryStorePlanMap.PruneSql / QueryStoreTextStore.PruneSql — the retention purge's time-sliced DELETE ... WHERE last_seen < $1 AND last_seen >= (SELECT min(last_seen) ...) AND last_seen < (SELECT min(last_seen) ...) + INTERVAL 'N days', run twice a day by the retention sweep.

No reader orders, filters, or range-scans last_seen for anything else — every other hit across the service, storage, viewer and MCP code is a join on the primary key or an unrelated table.

What changes

Migration V149:

DROP INDEX IF EXISTS collect.idx_query_store_plan_map_last_seen;
DROP INDEX IF EXISTS collect.idx_query_store_text_last_seen;
ALTER TABLE collect.query_store_plan_map SET (fillfactor = 90);
ALTER TABLE collect.query_store_text SET (fillfactor = 90);

Plus the removal of the CREATE INDEX IF NOT EXISTS ... last_seen lines from both tables' CreateTableSql, so startup convergence doesn't recreate what the migration drops.

The retention prune for both tables is also rewritten. The old PruneSql ran three sequential scans per statement without the dropped index (two min() subqueries plus the DELETE). It's replaced by DarlingRetention.UnorderedRowCappedDeleteSql, a sibling of the plan dimension's row-capped, ordered builder but with no ORDER BY: DELETE FROM t WHERE ctid IN (SELECT ctid FROM t WHERE last_seen < $1 LIMIT n). Dropping the ordering is safe here specifically because the cutoff is fixed per run rather than recomputed from the table's own remaining minimum: every row the predicate matches is already expired regardless of which ones a given batch happens to take, and each batch strictly shrinks the matching set, so the drain still terminates — it just can't strand a floor the way an unordered cap against a moving cutoff could. The cap (300,000 rows) is sized from a field read of the last_seen age histogram so a normal run clears the whole day's retirement in one batch for both tables, while a multi-day backlog still drains in bounded per-batch transactions via the existing drain loop.

Lock / transaction

PgMigrations.MigrateAsync wraps every rung in one transaction, so CREATE/DROP INDEX CONCURRENTLY is not available here (Postgres refuses it inside a transaction) — same constraint the V142 rung's doc comment already documents for a different table. Both statements in V149 are metadata-only: DROP INDEX on a btree does not touch the heap, and ALTER TABLE ... SET (fillfactor = ...) only changes a catalog reloption, so the short ACCESS EXCLUSIVE each one takes is a catalog update rather than a rewrite. Nothing else in the migrate session holds a competing lock on either table at that point.

Measurements

Measured on a rig (timescale/timescaledb:2.30.1-pg18), two databases —
before (btree last_seen index, default fillfactor) and after (V149
applied AFTER a 1.1 M-row population and VACUUM ANALYZE, so existing pages
start full, as on an upgraded store). Four touch cycles per table
(CHECKPOINT, an UPDATE over the due half of the rows, VACUUM between
cycles, rows aged back 13h to simulate 2 touches/row/day).

Touch WAL/FPI/HOT%, steady state (cycle 4):

Table before WAL/touch before HOT% after WAL/touch after HOT% WAL reduction
query_store_plan_map 558 B 0.0% 430 B 25.8% -23.0%
query_store_text 2262 B 0.0% 1480 B 64.9% -34.6%

query_store_text shows the larger win: HOT climbs to 65% by cycle 4 and the
fillfactor-90 + no-index change is not only a plan_map fix.

Prune cost without the index (field-scale numbers from a real store, read
only, no store named): query_store_plan_map main fork 709 MB / 4.04 M rows;
query_store_text main fork 4,992 MB + TOAST 4,355 MB / 6.86 M rows. One
slice of the current PruneSql (2 min() subqueries + the DELETE), warm
cache: plan_map 1.31 s indexed vs 1.68 s without (544,248 buffers, ~4.25 GB);
text 10.24 s indexed vs 7.46 s without, but the no-index text slice touches
~3.83 M buffers ≈ 29 GB logical (2.27 M read from outside shared_buffers).
The drain loop runs ~2 slices/table, 2 runs/day, so query_store_text
alone is on the order of tens of GB/day up to roughly 60-120 GB/day of extra
reads without its index
, depending on cache residency. No measured slice
anywhere near the ~10-minute stop condition (worst case 10.24 s).

A rig-measured one-scan prune shape (the UnorderedRowCappedDeleteSql
now shipped, dropping the two min() subqueries and the upper-bound
re-scan) reads roughly 2.5-3x fewer buffers than the old 3-scan shape at
every size tested on the rig — field-scale copies of both tables, no index,
cutoff selecting ~20% of rows: query_store_plan_map 290,476 buffers
(current 3-scan) vs 117,448 (one-scan); query_store_text 1,089,958 buffers
(current) vs ~497,812-632,414 (one-scan, two runs on the same rig gave a
noisy 3,528 ms/460 ms wall-time split but a stable buffer count). Scaled to
the field's query_store_text figures above (tens of GB/day, up to roughly
60-120 GB/day), the one-scan shape cuts that to roughly 20-40 GB/day —
real, but well under half.

The field saving is not projected here. The rig's per-touch WAL (0.4–2.3 KB) is far below the field's 7–8.5 KB, because the rig's tables fit in cache and many touched rows share a page, so the rig can't reproduce the field's full-page-image-heavy shape. On the field, a non-HOT touch also writes into the primary key and the last_seen index (881 MB and 1.2 GB on the busiest store), and a HOT touch writes into no index at all. So the field saving could be larger or smaller than the rig's 23% / 35%.

Trade shipped: both last_seen indexes are dropped (the touch's real,
large WAL/HOT win on both tables) AND the prune moves to the one-scan
shape, so the no-index prune's own read cost — real, but bounded to twice a
day and never near the 10-minute stop condition at any measured scale —
is itself cut by more than half rather than left at the 3-scan number.

The field saving is unmeasured until the 24-hour pg_stat_statements delta
after install.

Pins

New file Darling/Darling.Tests/QueryStoreLivenessHotTouchLiveTests.cs, five facts, all run against a rig container (timescale/timescaledb:2.30.1-pg18) and all green on the branch:

  • TheRungIsRegisteredAtTheTopOfADenseLadder — the ladder shape invariant.
  • TheProbeMapsAFullyMigratedStoreToThisTopRung — the viewer probe/map/arm three-way agreement.
  • AfterMigrate_NeitherLastSeenIndexExists_AndBothTablesCarryFillfactor90 — the LIVE schema pin.
  • StartupConvergence_RunTwice_DoesNotRecreateEitherIndex — running each table's CreateTableSql twice after migrate must not bring the index back.
  • TwoTouchesOfTheSameRow_TheSecondTouchReportsAHotUpdate — the payoff: n_tup_hot_upd increments on the settled-row touch, via pg_stat_force_next_flush() plus a short settle.

RED on origin/dev: ran the same file (verbatim) against a detached worktree at dev's tip (6c3b026ed) with the same rig. All 5 facts fail — 2 by design failure of the assertions (index present, no fillfactor reloption), 3 by reflection failures (no hasHotLivenessTouch parameter, no query-store-liveness-hot-touch rung).

Mutations (both reverted after verifying RED, confirmed green again after revert):

  • Re-adding the CREATE INDEX ... idx_query_store_plan_map_last_seen line to QueryStorePlanMap.CreateTableSql turns StartupConvergence_RunTwice_DoesNotRecreateEitherIndex RED.
  • Dropping the ALTER TABLE collect.query_store_plan_map SET (fillfactor = 90); line from V149's SQL turns AfterMigrate_NeitherLastSeenIndexExists_AndBothTablesCarryFillfactor90 RED.

Totals: dotnet build for Darling.Tests and Lite.Tests both 0 warnings / 0 errors. Ran (all green): QueryStoreLivenessHotTouchLiveTests (5), ReadLatencyFlushLiveTests (5, updated for the new top rung), QueryStoreTextStoreTests (updated for the dropped index), MigrationLadderPins, DocCommentHygieneTests, ViewerDataServiceTests — 103 total, 0 failed.

Gated tests (test-only nightly)

  • QueryStoreLivenessHotTouchLiveTests (new, this PR)
  • ReadLatencyFlushLiveTests (updated top-rung claim)
  • QueryStoreTextStoreTests
  • MigrationLadderPins
  • DocCommentHygieneTests
  • ViewerDataServiceTests

Darling.Tests and Lite.Tests both target net10.0-windows: they build here but cannot run on this machine; CI decides them.

Prune pins (this pass)

Added to DarlingRetentionTests: the new builder's exact SQL shape (no
ORDER BY, no min(), the LIMIT and ctid IN present, right table/column),
and a source-anchored check that both the map and text purge call sites use
the new builder with the shared cap as batchSize and leave
adaptiveRowCapTimeColumn unset (that parameter is the plan dimension's
#4130 retry-at-half-cap behavior, not measured as needed here — both tables
clear their whole steady-state backlog in one batch at the sized cap).
Moved the old QueryStoreTextStore.PruneSql shape pin out of
QueryStoreTextStoreTests since that method no longer exists; both
QueryStorePlanMap.PruneSql and QueryStoreTextStore.PruneSql are retired
(no other caller — verified with git grep).

Mutations (reverted after verifying RED, confirmed green again after
revert):
restoring ORDER BY {timeColumn} in UnorderedRowCappedDeleteSql
turns the shape pin RED (UnorderedRowCappedDelete_HasNoOrderByAndNoMinSubquery
failed, Total: 47, Failed: 1); dropping the LIMIT clause turns it RED.

Live loop test (added this pass)

New file Darling/Darling.Tests/QueryStoreLivenessTablePruneLiveTests.cs, run against a
timescale/timescaledb:2.30.1-pg18 rig via the same #1776 own-store / ScratchPostgres
pattern as the other live pins.

A test-only seam was added to DarlingRetention.PurgeAsync: an optional
livenessTouchedTablePruneRowCap parameter that defaults to the shipped
LivenessTouchedTablePruneRowCap constant (300,000). Both production callers (the daily sweep
and the on-demand purge_now command) omit the argument, so production always runs the shipped
cap unchanged; the test passes a small cap (1,000) so a multi-batch drain can be exercised
without seeding ~650k rows per table. The default is pinned twice: once in the new file, once in
DarlingRetentionTests.LivenessTouchedTablePruneRowCapSeam_DefaultsToTheShippedConstant.

Seed: 2,137 expired rows per table (1,000 * 2 + 137, two full batches plus one short one at
the test cap) 400 days back, plus 5 rows 1 hour back that must survive.

Run: DarlingRetention.PurgeAsync(postgres, timescaleAvailable: false, purgeLog, ct, livenessTouchedTablePruneRowCap: 1000) — the real product entry point.

Assert: every expired row gone, every in-cutoff row kept, and each table's batch count is
exactly 3 (ceil(2137 / 1000)), read from PurgeOneAsync's own "Retention purge drained ... in
{Batches} batch(es)" log line via a CapturingTestLogger.

Mutation (reverted after verifying RED, confirmed green again after revert): temporarily
made DrainBatchesAsync break unconditionally after its first executeBatch call — turns the
new test RED (only 1,000 of 2,137 expired rows removed per table, batch count 1 not 3). Reverted;
151 tests green across the classes run in this pass.

RED on origin/dev: not run as a separate check — the test's livenessTouchedTablePruneRowCap
argument does not exist on dev's PurgeAsync signature, so the file fails to compile there
(compile-failure RED, the same shape already used for other #4250 pins in this PR).

Prune at the shipped cap

Measured on a 2.5M-row rig copy of both tables in this PR's V149 shape (no last_seen
btree, fillfactor = 90), with query_store_text carrying a distinct ~700-byte value per
row, against the exact DarlingRetention.UnorderedRowCappedDeleteSql statement at the
shipped cap of 300,000. Buffers are shared hit + shared read; 1 buffer = 8 KB.

(a) Steady state (~110,000 rows expired):

table buffers time WAL
plan_map ≈2.0 GB 190 ms 18 MB
text ≈3.9 GB 483 ms 6.1 MB

One batch clears the whole steady-state backlog on both tables. It deletes fewer rows than the cap, so the drain loop stops after it (the loop's contract: a short batch ends it; the live loop test pins this). There is no second, empty-result call in the product.

(b) Backlog (~2,000,000 rows expired, 7 batches, no VACUUM between batches):

plan_map buffers rise slightly batch over batch (≈4.82 GB → ≈5.00 GB, +4%) as earlier
batches' dead tuples accumulate; per-batch time stays in the 390–430 ms range. text's
buffers stay flatter (≈5.2–5.6 GB) but per-batch time is noisier (642–1,283 ms), driven by
TOAST reads for the payload column rather than the row cap. No batch anywhere near a
command timeout.

(c) Old shape (pre-#4250 3-scan PruneSql) on the same steady-state slice:

table buffers time
plan_map ≈1.81 GB 819 ms
text ≈7.75 GB 1,888 ms

On plan_map the two shapes read about the same (≈1.8 GB vs ≈2.0 GB), and the new one is 4× faster. text's old shape reads ~1.9x this PR's shape for the same expired slice (992,223 vs.
510,033 buffers), from the two extra min() sequential scans the old PruneSql ran before
its DELETE. On the production store's full-size text table (6.86 M rows), the old shape measured roughly 29 GB of logical buffers per slice.

Scale to field size: field main forks are query_store_plan_map 709 MB and
query_store_text 4,992 MB main + 4,355 MB TOAST (9,347 MB total), giving size ratios of
≈2.36x (plan_map) and ≈4.03x (text) over this rig's tables. Scaling each shape's worst
observed single-batch time by that ratio: plan_map steady-state ≈0.45 s, backlog ≈0.93 s;
text steady-state ≈1.95 s, backlog ≈5.17 s. Every projected batch stays roughly two orders
of magnitude under the 300-second command timeout.

CHANGELOG entry

SECTION: Changed
ENTRY:

…en indexes, fillfactor 90 (#4250)

- V149: drops idx_query_store_plan_map_last_seen and idx_query_store_text_last_seen,
  sets fillfactor = 90 on both tables.
- Removes the CREATE INDEX ... last_seen lines from both tables' CreateTableSql so
  startup convergence does not recreate them.
- Adds the viewer probe sentinel (negative: index absent + fillfactor=90 present),
  StorageVersion bump, and the full rung recipe.
- New live pin file QueryStoreLivenessHotTouchLiveTests.cs: schema pin (RED verified
  against pre-fix dev), convergence pin (RED verified via a re-add mutation), HOT pin
  (n_tup_hot_upd increments on the second touch of a settled row).
- Updates QueryStoreTextStoreTests and ReadLatencyFlushLiveTests for the new top rung.

Refs #4250.
…delete (#4250 item 3)

The Query Store plan-map and statement-text prunes each ran three
sequential scans per statement without their last_seen index (two
min() subqueries plus the DELETE) — on the busiest measured store,
the text table's slice alone read on the order of tens of GB per
call. A new unordered, row-capped builder (a sibling of the plan
dimension's ordered one) replaces all three scans with one: a fixed
cutoff means every matched row is already expired regardless of
which ones a batch happens to take, so dropping the ORDER BY cannot
strand a floor the way it could against a moving cutoff. The cap
(300,000) clears the busiest single day observed in the field's
last_seen age histogram for either table in one batch, so a normal
run issues one statement, and a multi-day backlog still drains in
bounded per-batch transactions.
Merge origin/dev (#4467), and add a live loop test for #4250 item 3's row-capped prune through DarlingRetention.PurgeAsync (not the builder or drain loop called directly): seeds both liveness-touched tables past 2x a small test cap plus in-cutoff rows, asserts expired rows gone / in-cutoff rows kept / batch count matches the drain contract, and verifies a stop-after-one-batch mutation turns it RED. Adds a test-only cap-override seam to PurgeAsync defaulting to the shipped constant, so production is unchanged.
@erikdarlingdata
erikdarlingdata marked this pull request as ready for review September 27, 2026 16:43
@erikdarlingdata
erikdarlingdata merged commit 4fb8962 into dev Sep 27, 2026
23 of 24 checks passed
@erikdarlingdata
erikdarlingdata deleted the fix/4250-hot-liveness-touch branch September 27, 2026 16:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant