MCP budget: get_collection_health compacts healthy collectors under the size budget (#4198) - #4268
Merged
Merged
Conversation
…4198 budget Measured 41,669 bytes at default arguments with every field on every collector (42-collector fleet) against the 32,768-byte budget. There is no row to cut here: every collector is one row, and a health read must never hide one that is failing, stale, disabled or erroring by leaving it off the page. Cuts fields per-row instead - a collector that is HEALTHY with zero errors, session/extension-missing runs, denials or abandoned runs this window, and either stored rows or is a known event collector resting at zero, compacts to 7 fields (collector/status/compact/total_runs/rows_stored/avg_duration_ms/ last_success). Everything else always keeps every field, including three HEALTHY-banded edge cases the band alone would miss (a partial error rate, an older permission denial behind a fresh success, and a non-event collector's unexplained zero). New full_detail=true opt-in restores every field on every row; collector_detail_note says how many rows this call compacted. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
…t updated, Lite parity test DarlingWebEndpoints.cs's /api/read row now passes full_detail: true explicitly so the web viewer and Custom Views keep today's full per-collector payload; a no-rig source-scan pin holds it. Custom Views' Collection Health template also asks for full_detail so its errors/avg_duration_ms columns stay populated. McpToolsListBudget/DarlingMcpDataTools.txt and TotalCeilingBytes updated for the new full_detail parameter (+240 bytes measured). Adds the Lite.Tests twin of the Darling live test: same seeding shape, same never-compact assertions, against local DuckDB (no rig needed). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
- Darling: heatmap (+147) from dev + TK full_detail (+240) = 172,607 - Lite: heatmap (+147) from dev + TK full_detail (+229) = 90,647 - Add Lite pin: param get_collection_health.full_detail 191 (TK forgot it) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ
erikdarlingdata
marked this pull request as ready for review
September 25, 2026 10:35
erikdarlingdata
added a commit
that referenced
this pull request
Sep 25, 2026
Measured with McpToolsListBudgetTests after merging origin/dev (#4261, #4258, #4265, #4267, #4264, #4266 and #4268) plus this PR's catalog changes. Budget, census and tool-guide classes: 219/219. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ
erikdarlingdata
added a commit
that referenced
this pull request
Sep 25, 2026
…nder the response budget (#4198) (#4272) * MCP budget: describe_custom_view_catalog default groups measures by source (#4198) Default default-argument call was 98,173 bytes (#4198's own measurement), three times the tool's 32 KB budget. It is pure static reference data (no server/store read), so the cut groups the 179 measures by source and keeps only key/displayName/kind/unitFamily/validAggregates per measure; source=<name> drills into one source's full detail, full_detail=true returns the original shape unfiltered. /api/catalog (the web Custom Views editor) calls the underlying builder directly, never this MCP method, so it is unaffected - pinned in DarlingComposeTests. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3 * Trigger CI (draft PR skipped the Build workflow) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ * Add comment to DarlingMcpCustomViewCatalogSizeTests.cs to trigger Build CI GitHub did not fire pull_request events for this draft-opened PR. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ * fix(#4272): reword validAggregates description - compact drops allowedDimensions, source=<name> returns it Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3 * #4272: set TotalCeilingBytes to the measured 174,236 after merging dev Measured with McpToolsListBudgetTests after merging origin/dev (#4261, #4258, #4265, #4267, #4264, #4266 and #4268) plus this PR's catalog changes. Budget, census and tool-guide classes: 219/219. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
erikdarlingdata
added a commit
that referenced
this pull request
Sep 25, 2026
…as merged After merging dev, the per-tool fixes for get_blocking (#4267), get_collection_log (#4265), get_query_store_regressions (#4264), get_collection_health (#4268), describe_custom_view_catalog (#4272) and get_query_store_top (#4273) are all in, and get_fleet_overview already fit (CI's stale-exemption message). Every row goes; CI's live run decides. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ
erikdarlingdata
added a commit
that referenced
this pull request
Sep 25, 2026
CI on 08e6e81 measured get_collection_health at 35,145 B, over the 32 KB default. #4268 compacts only healthy collectors with nothing to report, and this fixture's collectors are not healthy, so nothing compacts. The row goes back, under #4198, which stays open for it. Every other exempt tool fits. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ
3 tasks done
erikdarlingdata
added a commit
that referenced
this pull request
Sep 25, 2026
* Add the #4198 MCP read-tool budget pin (part of #4198) McpReadToolBudgetLiveTests reflects over every [McpServerTool] method on every [McpServerToolType] class in the Darling service assembly, excludes write tools/analyze_*/compare_*/audit_config, binds each tool's DI services and server_name generically, and asserts the reply stays under McpResponseBudget.DefaultBytes unless the tool is named in an explicit, byte-stamped exemption roster. A fix lands by deleting its row. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3 * MCP budget pin: empty ExemptOffenders now every #4198 per-tool lane has merged After merging dev, the per-tool fixes for get_blocking (#4267), get_collection_log (#4265), get_query_store_regressions (#4264), get_collection_health (#4268), describe_custom_view_catalog (#4272) and get_query_store_top (#4273) are all in, and get_fleet_overview already fit (CI's stale-exemption message). Every row goes; CI's live run decides. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ * MCP budget pin: get_collection_health is still over on this fixture CI on 08e6e81 measured get_collection_health at 35,145 B, over the 32 KB default. #4268 compacts only healthy collectors with nothing to report, and this fixture's collectors are not healthy, so nothing compacts. The row goes back, under #4198, which stays open for it. Every other exempt tool fits. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
5 tasks done
erikdarlingdata
added a commit
that referenced
this pull request
Sep 25, 2026
The exempt row's comment said "#4268, still open", but #4268 merged. Word it like Darling's twin row: #4268 compacts only healthy collectors, this fixture's are not, so the tool stays over and open under #4198. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ
erikdarlingdata
added a commit
that referenced
this pull request
Sep 25, 2026
) * Pin Lite's MCP read tools under the #4198 response budget Part of #4198. Mirrors Darling's McpReadToolBudgetLiveTests: reflects over every [McpServerTool] in the Lite service assembly, excludes analyze_*/ compare_*/audit_config/mute_analysis_finding, binds each remaining tool's parameters generically and asserts the reply stays under McpResponseBudget.DefaultBytes on one busy DuckDB-only fixture. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3 * Lite budget pin: #4268 has merged; the row stays open under #4198 The exempt row's comment said "#4268, still open", but #4268 merged. Word it like Darling's twin row: #4268 compacts only healthy collectors, this fixture's are not, so the tool stays over and open under #4198. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
11 of 12 tasks
erikdarlingdata
added a commit
that referenced
this pull request
Sep 25, 2026
…4319) * Lite: get_collection_health's default fits the MCP response budget (#4198) A row that fails IsCollectionHealthCompactEligible used to keep the full ~30-field shape regardless of full_detail, and measured on both #4198 fixtures that alone could not clear the response budget (30+ small fields times every row that needs a look, which previewing free text cannot touch). That row now gets a new leaner partial_detail shape: enough to say what is wrong and since when, never reusing the compact marker so a caller scanning for compact != true still never skips it. CollectionHealthPayloadBudgetToolTests (already added by #4268, not new as #4198's design notes assumed) gets the ruling's 80% ceiling and the new marker assertions. McpReadToolBudgetTests' ExemptOffenders drops get_collection_health, its last row. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3 * Darling: get_collection_health's default fits the MCP response budget (#4198) Field-for-field mirror of Lite's fix: a row that fails IsCollectionHealthCompactEligible now gets a new leaner partial_detail shape instead of the full ~30-field fallback, with its own marker so a caller scanning for compact != true still never skips a row that needs a look. CollectionHealthPayloadBudgetLiveTests gets the ruling's 80% ceiling and the new marker assertions; McpReadToolBudgetLiveTests' ExemptOffenders drops get_collection_health, its last row. Live-test numbers to follow once the rig confirms them. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3 * Fix full-suite fallout: stacked doc summary and census gaps for #4198 DocCommentHygieneTests: my ExemptOffenders edit left the original <summary> block stacked above the replacement one instead of overwriting it. McpPayloadContractCensusTests: last_error_truncated and output_finding_truncated are new *_truncated keys get_collection_health's partial_detail shape introduces; classified as FieldPreviewCutKeys beside error_message_truncated, the existing entry for the same two files. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3 * Name why a HEALTHY row needs a look in get_collection_health's partial shape (#4198) PartialCollectionHealthRow dropped session_missing, abandoned and last_denied_at, so a row failing IsCollectionHealthCompactEligible only on one of those three carried partial_detail: true with nothing in the shape saying why. Add all three, same names/formats as the full row, on both SKUs, and update both get_collection_health guide texts' field list to match (still byte-identical, CrossSkuSurfaceSourceTests). ExtensionMissingCount is in the predicate but surfaces nowhere in the full row, so nothing was added for it. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3 --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part of #4198.
Why
get_collection_healthhas no row to cut. Every collector on the server is one row, and a cut must never hide a failing, stale, disabled, or erroring collector. Measured with a realistic 42-collector SQL Server fleet (everyCollectorCatalogname whoseTargetEngineisSqlServer), all fields on all rows total 41,669 bytes. That exceeds the shared 32,768-byteMcpResponseBudget.DefaultBytestarget.What changes
DarlingMcpDataTools.cs) and Lite (McpHealthTools.cs) compact a collector row to 7 fields by default:collector, status, compact, total_runs, rows_stored, avg_duration_ms, last_success. A row is boring whenHealthStatus == HEALTHYand it has no errors. It must also have zero missed runs of any kind and either stored rows or be a known event collector at zero.full_detailor not. The new live tests seed three edge cases the HEALTHY band alone misses: sub-threshold errors, stale denials, and unexplained persistent zeros. Newfull_detailbool argument (default false) restores every field on every row. The envelope's newcollector_detail_notesays how many rows this call compacted and how to get the rest.DarlingWebEndpoints.cs's/api/readrow forget_collection_healthnow passesfull_detail: trueexplicitly, so the dashboard keeps today's full payload unchanged. A no-rig source-scan test (WebViewerRow_PassesFullDetailTrue) pins that row's text. The Custom Views "Collection Health" template also passesfull_detail: true. Its columns (errors,avg_duration_ms) need the full row. A grep confirms/api/readis the only call site ofGetCollectionHealthin the whole service, so that fix covers every web path including Custom Views. The template change is belt-and-suspenders.full_detailparameter, +240 bytes Darling, +229 bytes Lite (McpToolsListBudget/DarlingMcpDataTools.txt,McpToolsListBudget/McpHealthTools.txt). After merging dev (which also added +147 bytes for the heatmap lane), DarlingTotalCeilingBytesis 172,607 and Lite is 90,647. The head description is unchanged.ExemptOffendersroster in the still-unmergedMcpReadToolBudgetLiveTests. Per the coordinator, that file is not ondevand this PR does not create or edit it. Whichever of A budget pin over every MCP read tool (part of #4198) #4224 or this PR merges second must removeget_collection_healthfrom that roster, since this PR makes it pass the budget on its own.get_query_store_topin the same Darling file tonight. This PR touches onlyGetCollectionHealthand its two new private helpers.Test plan
Darling.Testsbuild: 0 warnings, 0 errors.Lite.Testsbuild: 0 warnings, 0 errors.CollectionHealthPayloadBudgetLiveTests(own file, own seeding, rig on port 55984): default 20,707 bytes (under 32,768) vs. 41,669 bytes atfull_detail=true. Asserts the four not-boring collectors (wait_stats,memory_grant_stats,query_store_health,database_scoped_config) never compact, and thatdeadlocks(an event collector resting at zero) does compact. Passed.WebViewerRow_PassesFullDetailTrue(no rig): passed.Lite.TeststwinCollectionHealthPayloadBudgetToolTestsagainst local DuckDB: same shape, same assertions. Passed.McpToolsListBudgetTests,DarlingMcpDataToolsSurfaceAndSqlTests,DarlingMcpDataToolsLivePostgresTests,McpPayloadContractCensusTests, and the two new classes above: 139 tests, 0 failed.git merge origin/dev: conflict inMcpToolsListBudgetTests.csresolved. Darling ceiling raised to 172,607 (dev 172,367 + TK's +240). Lite ceiling raised to 90,647 (dev 90,418 + TK's +229). Added missing Lite pinparam get_collection_health.full_detail 191.Darling.Testssuite: not run this pass. Stopped short to stay inside the context budget after two watchdog warnings. The targeted run above covers every file this PR touches plus the shared budget and census gates.Lite.Testssuite: not run this pass, same reason.CHANGELOG entry
SECTION: Fixed
ENTRY:
get_collection_healthcaps its response size ([MCP read tools have no default response-size budget: at default arguments 12 tools return >50 KB and 3 return >100 KB for one server, more than an agent client's per-result cap #4198]): A default call on a busy SQL Server fleet reached 41,669 bytes, over the 32 KB limit. Healthy collectors with nothing to report now compact to a shorter row by default. Any failing, stale, stopped, erroring, denied, or regressed collector keeps every field. A newfull_detailargument restores every field on request. The web dashboard and Custom Views are unaffected.REF:
[MCP read tools have no default response-size budget: at default arguments 12 tools return >50 KB and 3 return >100 KB for one server, more than an agent client's per-result cap #4198]: MCP read tools have no default response-size budget: at default arguments 12 tools return >50 KB and 3 return >100 KB for one server, more than an agent client's per-result cap #4198
Generated with Claude Code
https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ