Skip to content

MCP top queries and procedures over long windows: no dropped filters, null for what the rollup lacks, per-server coverage - #4715

Merged
erikdarlingdata merged 14 commits into
devfrom
fix/4711-mcp-topn-hourly-honesty
Sep 29, 2026
Merged

erikdarlingdata merged 14 commits into
devfrom
fix/4711-mcp-topn-hourly-honesty

Conversation

@erikdarlingdata

@erikdarlingdata erikdarlingdata commented Sep 29, 2026 •

Copy link
Copy Markdown
Owner

Fixes #4711.

What changed

  1. Raw-only filters keep the read on raw. parallel_only, min_dop and group_by=host_object used to be dropped on the hourly rollup while filter_applied claimed them. The reader now forces raw when any is set, so filter_applied, the group_by echo and precision_note are true. precision_note says the read stayed on raw; the raw floor probe reports window_truncated / effective_start. The empty min_dop status carries effective_start, effective_hours_back and window_truncated; when raw no longer holds the whole window it says which part it read, and otherwise it keeps the sentence Lite shares word for word.
  2. Null for what the rollup lacks. On tier_used=hourly, plan hash, plan handle, DOP, reads/writes/physical reads/rows/spills, avg_reads, is_parallel, distinct_texts and min/max cpu/elapsed are JSON null (procedures: object_type, handles, I/O totals and all four extremes). The rollup's min/max are per-collection sums, so none are selected. sql_handle is now returned (exact, it is a rollup group column). precision_note says so, and so does each tool's guide (get_tool_guide); the served tools/list heads are unchanged. No new payload keys.
  3. Text lookup inside the ranked read. A LEFT JOIN LATERAL over v_query_stats replaces one query per row; the read over-fetches by 5 and drops WAITFOR shells as raw does. A miss gives a null query_text and a text_note.
  4. Per-server first-bucket probe, behind one method a cache can wrap. window_truncated, effective_start and truncation_note on hourly now come from it. When the window crosses the stitch floor F (RollupCoverage.StitchFloor, the same F the stitched relation splits at), the probe is least() of two ordered first-row probes: the legacy view below F, and the successor from F (HourlyFirstBucketSql). That is exact, because the legacy half only holds rows below F. With nothing to split, it probes the single spliced relation (HourlyFirstBucketSingleRelationSql, which throws if the splice is ever a UNION). A first row can't be taken cheaply from a stitched UNION ALL subquery, because the planner sorts every row of the server in the window to return one.
  5. Hour edges disclosed in precision_note by the new HourlyWindowEdges.Note: the start (no bucket before the first served one), and the end, against the rollup's materialization ceiling (see below).

Exactness fixes, second pass

  • End edge and the materialization ceiling. When an hourly window ends past the rollup's materialization ceiling, the note no longer claims the end hour was counted whole. It says "the hourly rollup is materialized only to {ceiling}; nothing after it was read". When the start edge also moved, it says "served from {start} to {ceiling}". This lives in the sentence only: window_truncated and effective_start still describe the start edge.
  • Forced-raw read over a window raw no longer holds. parallel_only, min_dop and group_by force the raw tier. With an aged as_of raw holds nothing there. The tool now returns status empty ("raw query_stats holds nothing in this window; the hourly rollup, which does, cannot apply parallel_only/min_dop/group_by=host_object"), with window_truncated = true and effective_start = null, instead of "the window's answer".
  • WAITFOR shells on the hourly tier. A row whose text has aged out of raw carries a text_note that also says it "may be a WAITFOR shell, which raw query_stats filters out".
  • sql_handle on the hourly tier. precision_note states that the handle is the rollup's MAX(sql_handle), which can name a handle seen only on a zero-interval collection that raw excludes. Totals are unaffected.
  • Start wording. The start note reads "hourly buckets start on the hour: no bucket before {served}: the data from {requested start} to {served} is not included". It no longer calls a multi-hour gap "the partial hour".
  • cpu_attribution on the hourly tier. The share divides by the span the rollup served (first bucket to min(ceiling, end of the as_of hour), on hour edges), and the existing note states that span. Raw reads keep the requested window.
  • Cleanup. The first-bucket probe task is observed if the ranked read throws, and the reader is disposed once.

Exactness fixes, third pass

  • Both hourly ranked reads and both first-bucket probes now carry AND f.bucket < $N, where $N is the relation's materialization ceiling bound as a naive-UTC timestamp parameter (a placeholder swapped for the clause, or for nothing when the ceiling is unknown). A bucket materialized after the cached coverage snapshot is neither read nor counted, so "nothing after it was read" is true by construction. The split probe keeps its two ordered LIMIT 1s inside least().
  • The CPU samples for cpu_attribution are read over the span the hourly rollup served (first bucket to the ceiling), on both tools, so the measured denominator and the ranked numerator cover the same span.
  • A null ceiling now says "materialization ceiling unknown; the end edge is not verified" (window note and attribution note). "served no bucket span" appears only when a ceiling exists but no span was served.
  • Wording: "the hourly rollup, which covers this window, cannot apply ...".

Cost

The probe is one extra indexed query per hourly-tier call, uncached. Measured read-only on a large production monitoring store, with 43 servers and 1-day rollup chunks, most of them compressed:

  • Successor only (the window starts after F): the busiest server at 168 h took 4.6 ms. Per server, the worst pass (72 h, cold) had a p50 of 1.8 ms, a p95 of 2.5 ms and a max of 4.1 ms; for procedures, the max was 8.8 ms. There was no seq scan: each probe is an ordered chunk append that stops at the first chunk.
  • Stitched (the window crosses F; 720 h):
    • A first draft probed through the stitched UNION ALL and took 146 ms on the busiest server, sorting about 207 K rows. That's why the probe splits at F.
    • The split form took 0.39 ms on the same server. Per server, the p50 was 10.3 ms cold and 2.6 ms warm, and the max 17.2 ms; the cold cost is mostly planning over about 20 chunks.

Tests

  • New/rewritten live facts (TopQueriesHourlyRoutingLiveTests): forced-raw with floor disclosure and empty status; hourly nulls plus sql_handle plus null text and text_note; young server reports window_truncated. Procedures live test asserts the nulls. New source pins in TopQueriesHourlyRoutingTests, new HourlyWindowEdgesTests.
  • Ran on a local TimescaleDB rig: TopQueriesHourlyRoutingLiveTests 4/4, TopProceduresHourlyRoutingLiveTests 2/2, TopQueriesHourlyRoutingTests 4/4, HourlyWindowEdgesTests 4/4, McpPayloadContractCensusTests 68/68, McpPageContractTests 43/43, DocCommentHygieneTests 77/77, StorageCommandTimeoutTests, StartupCommandTimeoutTests, McpReadCommandTimeoutTests, CommentFilterAdoptionTests, RepoFileAdoptionTests, StitchedRelationSqlTests all green. Darling.Tests builds 0 errors, no CA/CS/IDE warnings.
  • Pin changed: McpPayloadContractCensusTests.WindowFloorBlocks 4 -> 5.
  • Split probe:
    • The live fact HourlyRouted_StitchedWindow_CoverageProbeSplitsAtTheFloor: a window crossing F reports the legacy half's first bucket, and a successor-only server reports the successor's first bucket. On the rig, 5/5 in its class.
    • The shape pins: least(, exactly two ordered probes, no UNION in either constant, and the probe method calls StitchFloor. On the previous tip they're a compile RED (HourlyFirstBucketSingleRelationSql doesn't exist there), and the old single stitched constant has no least(.
  • Mutation check: disabling the force-raw clause (if (false && tier == Hourly && ...)) makes the forced-raw fact fail: Expected: "raw" Actual: "hourly". Restored; green.
  • RED on old code: not run in a detached worktree. HourlyWindowEdgesTests is a compile RED on dev (type absent); the other pins assert shapes dev lacks. The forced-raw fact's RED is the mutation above.
  • Mutations of the split probe, on the same rig: swapping the two halves' floor comparisons, or dropping the legacy half, each makes HourlyRouted_StitchedWindow_CoverageProbeSplitsAtTheFloor fail (expected the legacy half's 01:00 first bucket, got the successor's 10:00); restored, the classes pass 5/5 and 2/2.
  • The tools/list budget, the guide-head and SKU-parity pins, and Lite's shared miss-sentence pin pass unchanged (no ceiling raised).
  • The full suite: run by CI.
  • Web: pages render through fmtInt/fmtNum/fmtMs, which return "—" for null; no JS change.

Exactness fixes, second pass: tests

RED = the pre-change head (e43d9b27c) with the new test files copied in. GREEN = this branch.

Pin RED (old code) GREEN
HourlyRouted_EndBeyondTheMaterializationCeiling_SaysNothingAfterItWasRead (live) FAIL: Assert.Contains sub-string not found (old note says "included whole") pass
ForcedRaw_OverAWindowRawNoLongerHolds_SaysNothingWasRead (live, aged as_of) FAIL: Assert.Contains sub-string not found (old text: "the window's answer") pass
HourlyRouted_MissingColumnsAreNull_AndSqlHandleIsCarried (live, WAITFOR text_note) FAIL: Assert.Contains sub-string not found pass
HourlyWindowEdgesTests start wording FAIL: Assert.Contains sub-string not found (old sentence says "the partial hour") pass
HourlyWindowEdgesTests ceiling, combined-edge and ServedSpan cases compile RED: Note takes 3 arguments, ServedSpan does not exist pass
HourlyRouted_StitchedWindow_CoverageProbeSplitsAtTheFloor (live, legacy-only server D) passes on old code by design: it guards the < $4 bound, see the mutation row pass
HourlyAttributionSpanTests (3 unit facts) new file; with the helper returning the requested window, 2 of 3 FAIL (Assert.Equal values differ: start, and the 0.75 share) 3 pass
  • < $4 mutation: removing f.bucket < $4 AND from the legacy half of HourlyFirstBucketSql makes the stitched fact FAIL on server D: expected effective_start 13:00, actual 11:00 (the legacy-only bucket wins). Restored: pass.
  • Finding 8 RED: changing HourlyAttributionSpan to return the requested window for hourly reads fails Hourly_LateFirstBucketAndLowCeiling_UsesTheServedSpan and Hourly_ShareDividesByTheServedSpan_NotTheRequestedWindow. Restored: 3 of 3 pass.
  • Sweep: all *Census*|*Ratchet*|*Inventory*|*Adoption* classes plus the MCP budget, guide, heads-data, page-contract and payload-contract classes, both hourly routing live classes, both *HourlyRoutingTests, StitchedRelationSqlTests, CpuAttributionTests, DocCommentHygieneTests, the store command timeout classes: 43 classes, 586 tests, 0 failed, 0 skipped (live rig attached). Lite.Tests builds with 0 warnings and 0 errors (build-only; it cannot run here).
  • Shared sentence: Darling's untruncated raw empty-page message contains, byte for byte, the sentence in Lite/Mcp/McpQueryTools.cs: "The filter was applied in SQL over the whole window, so this is the window's answer rather than a page artefact — drop parallel_only / min_dop to see the unfiltered ranking, or confirm current parallelism with analyze_query_plan."

Exactness fixes, third pass: tests

  • HourlyRouted_EndBeyondTheMaterializationCeiling_SaysNothingAfterItWasRead (live, extended): a bucket materialized past the ceiling is not read (RED without the clause: Assert.Single sees 2 rows; GREEN: 7/7).
  • Same test, sql_cpu_seconds_in_window against seeded samples over a known span (RED with the CPU call site reverted to the requested window: expected 7200, actual 10800; GREEN: 7/7).
  • HourlyWindowEdgesTests.NullCeiling_SaysTheCeilingIsUnknown and HourlyAttributionSpanTests.Hourly_NullCeiling_SaysUnknown_NotThatNoSpanWasServed: green after the wording change.
  • Also green: the procedures live and routing classes, McpToolsListBudgetTests, McpToolGuideTests, McpToolGuideHeadsDataTests, McpPayloadContractCensusTests, McpPageContractTests, StitchedRelationSqlTests, StorageCommandTimeoutTests, McpReadCommandTimeoutTests, StartupCommandTimeoutTests, DocCommentHygieneTests, the Adoption classes.

CHANGELOG

None: every defect fixed here is in the hourly routing and the window disclosure that are new in 3.9 (#4396, #4413, #4278 and #4279; none is in v3.8.0), so nothing a 3.8 user has seen changes. The 3.9 entries for those changes will describe this final behavior.

erikdarlingdata and others added 14 commits September 29, 2026 01:14
…ters, null for missing columns, in-statement text, per-server coverage, edge disclosure
A stitched UNION ALL cannot give an ordered first row, so ORDER BY bucket LIMIT 1 over it sorted every row the server has. The probe is now least() of two ordered first-row probes, one per relation, split at the same floor the stitch uses; a single relation keeps the one-probe form.
…uide marker, and the raw empty page keeps the shared sentence (#4711)
@erikdarlingdata
erikdarlingdata marked this pull request as ready for review September 29, 2026 13:44
@erikdarlingdata
erikdarlingdata merged commit dceb4f1 into dev Sep 29, 2026
16 of 18 checks passed
@erikdarlingdata
erikdarlingdata deleted the fix/4711-mcp-topn-hourly-honesty branch September 29, 2026 13:44
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant