Problem
#3897 (delivered in #3968) put the trend reads on a point budget sized for a model's context. The other read tools have row limits (limit, top), but no budget on what one call returns. Rows vary in width (plan XML, query text, deadlock graphs, per-object stats), so a default call can return far more than an MCP client will put in context. When that happens, the client either spills the result to disk (the agent then has to page through a file) or truncates it. Either way the tool's carefully written envelope (truncated, oldest_returned_*, the notes) may never reach the model.
This is invisible on a lab store with few databases and quiet servers, and shows up on a large production store.
Measured
Every read tool at its default arguments, one server each, on two production stores, build 3.8.0-nightly.20260924.467, 2026-09-25. Write tools, analyze_*/compare_* and audit_config were excluded; 126 tools ran on each store, sequentially.
SQL Server store, one busy server. Median 1.5 KB, p90 47 KB; 12 tools over 50 KB, 3 over 100 KB:
| tool |
bytes |
ms |
get_query_store_regressions |
211,326 |
3,725 |
get_query_heatmap |
144,757 |
1,763 |
get_deadlock_detail |
120,454 |
247 |
describe_custom_view_catalog |
98,173 |
291 |
get_plan_corrections |
95,428 |
224 |
get_collection_log |
87,816 |
239 |
get_object_locking |
71,332 |
460 |
get_index_usage |
69,290 |
848 |
get_fleet_overview |
67,157 |
687 |
get_analysis_findings |
66,838 |
687 |
get_active_queries |
56,586 |
186 |
get_blocking |
52,524 |
160 |
PostgreSQL-target store, one cluster: get_pg_cpu_utilization 103,415 B (filed separately, #4193), describe_custom_view_catalog 98,173, get_collection_log 85,206, get_fleet_overview 80,807, get_pg_io_trend 55,628.
Client-side: Claude Code refused three get_fleet_overview results inline (67,351 / 65,860 / 80,816 characters), each with "result exceeds maximum allowed tokens", and wrote them to files. The fleet overview is the natural first call of any session. Another agent seat had to build a scripting wrapper to use these reads at all.
For scale, tools/list is 172,969 B for 159 tools (down from 336 KB before #3898).
Where
Fix shape
- A shared default response budget for MCP reads, e.g. 32 KB (~8 k tokens), stated in one place like
TrendBuckets.McpPointBudget. Each tool's defaults are sized so that a default call on a large store stays under it:
- Row limits: default
limit sized to the tool's measured bytes per row (a regression row with query text ≈ 10 KB, so 20 rows → ~5; a deadlock graph → ~3).
- Wide fields: truncate long text fields (query text, plan fragments, deadlock XML) at default to a preview with
*_truncated: true, and add a full_text opt-in, as get_store_query_stats already does.
get_fleet_overview: a detail parameter, summary by default: the rollup, worst_servers and per-band counts. cards returns what it does today. It's also worth adding worst_only / band filters.
describe_custom_view_catalog: a family/read filter, with the default listing names and one-line purposes only.
- Keep the envelope honest:
truncated / *_returned already exist, so the only change is that defaults stop sitting past the context limit.
- Pin: a live test in the style of
TrendPayloadBudgetLiveTests that calls every read tool at default arguments on a seeded multi-database store and asserts each payload is under the budget. New tools inherit the test.
Problem
#3897 (delivered in #3968) put the trend reads on a point budget sized for a model's context. The other read tools have row limits (
limit,top), but no budget on what one call returns. Rows vary in width (plan XML, query text, deadlock graphs, per-object stats), so a default call can return far more than an MCP client will put in context. When that happens, the client either spills the result to disk (the agent then has to page through a file) or truncates it. Either way the tool's carefully written envelope (truncated,oldest_returned_*, the notes) may never reach the model.This is invisible on a lab store with few databases and quiet servers, and shows up on a large production store.
Measured
Every read tool at its default arguments, one server each, on two production stores, build 3.8.0-nightly.20260924.467, 2026-09-25. Write tools,
analyze_*/compare_*andaudit_configwere excluded; 126 tools ran on each store, sequentially.SQL Server store, one busy server. Median 1.5 KB, p90 47 KB; 12 tools over 50 KB, 3 over 100 KB:
get_query_store_regressionsget_query_heatmapget_deadlock_detaildescribe_custom_view_catalogget_plan_correctionsget_collection_logget_object_lockingget_index_usageget_fleet_overviewget_analysis_findingsget_active_queriesget_blockingPostgreSQL-target store, one cluster:
get_pg_cpu_utilization103,415 B (filed separately, #4193),describe_custom_view_catalog98,173,get_collection_log85,206,get_fleet_overview80,807,get_pg_io_trend55,628.Client-side: Claude Code refused three
get_fleet_overviewresults inline (67,351 / 65,860 / 80,816 characters), each with "result exceeds maximum allowed tokens", and wrote them to files. The fleet overview is the natural first call of any session. Another agent seat had to build a scripting wrapper to use these reads at all.For scale,
tools/listis 172,969 B for 159 tools (down from 336 KB before #3898).Where
Darling/PerformanceMonitor.Darling.Service/Mcp/and their Lite twins. Defaults are per tool (limit/topdefaults and full-text fields).PerformanceMonitor.Common/Mcp/TrendBuckets.cs: the budget contract The trend tools answer in time buckets sized to the window: a day of get_file_io_trend is 34 KB, not 1.4 MB (#3897) #3968 built for trends, the model to extend.get_fleet_overviewtakes onlyhours_back. It always returns every card (43–50 cards, ~1.5 KB each) with no summary option.Fix shape
TrendBuckets.McpPointBudget. Each tool's defaults are sized so that a default call on a large store stays under it:limitsized to the tool's measured bytes per row (a regression row with query text ≈ 10 KB, so 20 rows → ~5; a deadlock graph → ~3).*_truncated: true, and add afull_textopt-in, asget_store_query_statsalready does.get_fleet_overview: adetailparameter,summaryby default: the rollup,worst_serversand per-band counts.cardsreturns what it does today. It's also worth addingworst_only/bandfilters.describe_custom_view_catalog: afamily/readfilter, with the default listing names and one-line purposes only.truncated/*_returnedalready exist, so the only change is that defaults stop sitting past the context limit.TrendPayloadBudgetLiveTeststhat calls every read tool at default arguments on a seeded multi-database store and asserts each payload is under the budget. New tools inherit the test.