Skip to content

Store rung V150: a per-collector watermark index and a per-server job-history index (#4469, #4477) - #4489

Merged
erikdarlingdata merged 2 commits into
devfrom
feat/v150-watermark-and-job-history-indexes
Sep 27, 2026
Merged

erikdarlingdata merged 2 commits into
devfrom
feat/v150-watermark-and-job-history-indexes

Conversation

@erikdarlingdata

@erikdarlingdata erikdarlingdata commented Sep 27, 2026 •

Copy link
Copy Markdown
Owner

Fixes #4469.
Refs #4477.

Why

The startup watermark read (DarlingWorker.ReadCollectorWatermarksAsync) seeds every collector's next-due time from its last recorded run. #4480 already bounded the read to a 2-day floor, which cut out the compressed chunks beyond that floor — about 11% of a field-measured 4,775 ms cold read. The other 89% was a bitmap heap scan of the two newest, uncompressed chunks: TimescaleDB had no index shaped to answer "the newest row per collector" directly, so it still had to read every row for the server across those chunks before it could group by collector name.

The Job History tab's read (ViewerDataService.BuildJobHistorySql) selects the newest N rows fleet-wide by a de-skewed run_datetime, which today reads every row in the window per server before sorting and limiting.

What changes

Store rung V150 adds two indexes:

  • idx_collection_log_watermark on collection_log (server_id, collector_name, collection_time DESC).
  • idx_job_history_server_run on job_history (server_id, run_datetime DESC, instance_id DESC) — the trailing column matches the Job History read's own tie-break (run_datetime DESC, instance_id DESC) so a per-server top-N lookup against the index needs no extra sort.

DarlingWorker.ReadCollectorWatermarksSql changes from a single GROUP BY statement to a per-collector lookup: SELECT c.name, (SELECT l.collection_time FROM collection_log l WHERE l.server_id = $1 AND l.collector_name = c.name AND l.collection_time >= $2 ORDER BY l.collection_time DESC LIMIT 1) FROM unnest($3::text[]) AS c(name). $3 binds the same CollectorScheduleDefaults.All.Keys list both call sites already hold in memory — no new store read. The absent-when-never-run-within-the-floor contract is unchanged.

The Job History read itself is not changed in this PR — only the index it will use ships here.

Measurements

Rig: timescale/timescaledb:2.30.1-pg18, 43 servers, real collector cadences, chunks compressed the product way.

Watermark read (~40 collectors, 15M collection_log rows, 9 of 11 chunks compressed):

cold warm
old GROUP BY ~24 ms ~12 ms
new per-collector lookup ~0.65 ms ~0.6 ms

A >30x floor reduction, cold and warm alike — one index-only descent per collector name instead of a scan of every row in the newest chunks.

Job History base row selection (43 servers, ~2.7M job_history rows over 4 days, 3 of 5 chunks compressed), cold buffers:

buffers (cold)
unindexed scan ~6,883
per-server LATERAL top-N against the new index ~1,506

About 4.6x fewer buffers. The full read (base selection plus the job_stats per-job average/max aggregate, which this index does not cover) runs ~721 ms cold against the unindexed ~1,362 ms — about 1.9x, smaller than the watermark win because the aggregate half still scans every matching row in the window.

Costs

  • Build time: ~0.94 s for the watermark index on 2.4M uncompressed rows (well inside the migration command timeout); the job-history index built in well under a second on this rig's smaller uncompressed set. A store many times today's size would still build both in low seconds.
  • Lock: each CREATE INDEX takes a ShareLock for its build duration, blocking concurrent writes to that table until it finishes. A store whose uncompressed chunks have grown unusually large (a compression-policy gap, or a raised CompressAfterDays) would make the lock window grow with the uncompressed row count.
  • Size: ~115 MB for the watermark index at the field's estimated row count (2.4M uncompressed rows, 30 compressed chunks at ~8 KB each).
  • Write overhead: +38% wall time on a 100,000-row bulk COPY into the newest chunk (0.353 s → 0.487 s median of 3) — one more btree every future collection_log write maintains. The service's real write path inserts one row per collector run, so the absolute per-row cost is far smaller.
  • Both indexes are plain CREATE INDEX IF NOT EXISTS; WITH (timescaledb.transaction_per_chunk) cannot run inside MigrateAsync's per-rung transaction (confirmed: it raises an error), and the measured build times don't need it.

IMPORTANT

Store rung V150 builds two indexes on first start after upgrade; collection_log and job_history writes wait for the build (about a second per 2.4 million uncompressed rows on the test rig).

Test plan

  • MigrationDataMovingRungCensusPins — declared V150 in the census (metadata-only CREATE INDEX over populated hypertables, costed rather than assumed, same shape as V104/V142).
  • MigrationLadderPins, PgSchemaGeneratorTests, DocCommentHygieneTests — green, no fresh-vs-upgraded divergence (neither index exists on a fresh install today, so this rung is the real create on both paths).
  • DarlingWatermarkFloorPlanShapeLiveTests — updated to assert the new plan shape (idx_collection_log_watermark, Function Scan on unnest, no HashAggregate/Bitmap Heap Scan/Seq Scan). RED on dev (old code has no such index and plans a Finalize HashAggregate). A first cut of this pin that only checked "an Index Only Scan is present" passed against a mutated build that still had the old GROUP BY text, because the old statement can also use the new index — the pin was strengthened to require the per-collector plan shape specifically, and the mutation (GROUP BY text restored, index still present) was re-verified RED against the strengthened pin, then reverted.
  • DarlingWatermarkIndexEquivalenceLiveTests (new) — the new read returns the identical per-collector answer the bounded GROUP BY oracle does, run through ReadCollectorWatermarksAsync itself (the product's own call path), including the absent-when-outside-the-floor and absent-when-never-run cases.
  • DarlingWatermarkFloorScanBoundLiveTests, DarlingWatermarkSeedLiveTests — unaffected, both still green (they call the async method directly, not the SQL text).
  • Totals: 137/137 green on the touched classes plus the migration census/ladder/doc-comment/schema classes.

CHANGELOG

SECTION: Changed
ENTRY:

IMPORTANT: Store rung V150 builds two indexes on first start after upgrade; collection_log and job_history writes wait for the build (about a second per 2.4 million uncompressed rows on the test rig).

…-history index (#4469, #4477)

The startup watermark read looks up each collector's newest collection_log row
through a new index instead of scanning the newest chunks: idx_collection_log_watermark
on collection_log (server_id, collector_name, collection_time DESC), plus a per-collector
unnest()/LATERAL LIMIT 1 lookup replacing the GROUP BY. Also adds
idx_job_history_server_run on job_history (server_id, run_datetime DESC, instance_id DESC)
for the Viewer's Job History read.

Refs #4469
Refs #4477
…rm, the top-rung pins (#4469)

StorageVersion.SchemaVersion moves to 150. The Viewer schema probe gains a V150 sentinel (both
idx_collection_log_watermark and idx_job_history_server_run present) and its newest-first arm above
V149's, so a fully-migrated store maps to exactly 150 instead of falling through to V149 and showing a
spurious upgrade banner. QueryStoreLivenessHotTouchLiveTests' 'I am the top rung' facts move to a new
CollectionLogWatermarkAndJobHistoryIndexesRungTests class; V149's and V148's own facts are rewritten to
assert what stays true forever (registered, ladder-dense, gated below the current top's arm) rather than
'is exactly the top'.

Refs #4469, #4477
@erikdarlingdata
erikdarlingdata marked this pull request as ready for review September 27, 2026 18:56
@erikdarlingdata
erikdarlingdata merged commit 67ba074 into dev Sep 27, 2026
21 of 22 checks passed
@erikdarlingdata
erikdarlingdata deleted the feat/v150-watermark-and-job-history-indexes branch September 27, 2026 18:56
erikdarlingdata added a commit that referenced this pull request Sep 28, 2026
…00Z UTC (#4621)

Moves the CHANGELOG entries carried in merged pull-request descriptions into [Unreleased]. The cut is PRs merged after 2026-09-26T17:37:33Z and at or before 2026-09-28T17:40:00Z; the next splice starts after it.

- 76 PRs are spliced: Fixed 36, Changed 28 and Added 12, each counted once under its first section. That is 79 bullets: #4481's Fixed entry holds three, and #4548 adds a second group under Added.
- 21 PRs have no user-visible entry (None, test-only, CI-only, or deferred to a parent).
- The [#4509] link definition, which #4538 also carries, is defined once.
- The IMPORTANT upgrade notes from #4489, #4495, #4501, #4506 and #4541 are held for the release cut and are not in this change.
- Only CHANGELOG.md changes.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant