Skip to content

Pin Lite's MCP read tools under the #4198 response budget (#4198) - #4290

Merged
erikdarlingdata merged 2 commits into
devfrom
test/4198-lite-read-tool-budget-pin
Sep 25, 2026
Merged

erikdarlingdata merged 2 commits into
devfrom
test/4198-lite-read-tool-budget-pin

Conversation

@erikdarlingdata

@erikdarlingdata erikdarlingdata commented Sep 25, 2026 •

Copy link
Copy Markdown
Owner

Part of #4198.

Why

#4224 added a test that measures every Darling MCP read tool against the shared 32 KB response budget. If a change makes a tool's default answer too big, that test fails. Lite has 89 MCP tools and had no such test. A tool that grew past the budget on Lite alone went unnoticed.

What changes

This PR adds Lite.Tests/McpReadToolBudgetTests.cs, Lite's version of Darling's test. It runs on DuckDB and needs no PostgreSQL rig, like the other *BudgetTests classes in Lite.

  • It finds every [McpServerTool] method in the Lite assembly by reflection. A new tool is measured with no change here.
  • It skips analyze_* (5 tools), compare_analysis and audit_config by name. It skips write tools through an exact list. The lane checked all 89 tool names for add, create, update, delete, remove, set and mute verbs. Lite has no MCP tools for adding servers, Custom Views, custom alert rules or notification routes, so mute_analysis_finding is its only write tool. As in Darling's test, a new write tool must be added to the list on purpose.
  • It fills each tool's parameters the way Darling's test does. LocalDataService, ServerManager, AnalysisService and McpToolGuideCatalog come by type, and server_name is the seeded server. Every other parameter keeps its default.
  • It seeds one busy server in the same seven tables as Darling's fixture. It reuses column lists that McpPageContractTests, QueryStoreRegressionsBudgetTests and QueryStoreTopBudgetTests already use against this DuckDB schema. The seed holds:
    • Query Store regressions with about 3.8 KB of query text each
    • 30 deadlock graphs and 30 blocked-process reports
    • 10 DMV blocking snapshots
    • collection-log history for every SQL Server collector, with SUCCESS and ERROR rows mixed
    • 30 plan-correction recommendations and 20 long query completions
  • It measures each reply's UTF-8 bytes against McpResponseBudget.DefaultBytes. The test fails when a tool is over the budget and not in ExemptOffenders. It also fails when an exempt tool now fits, or when an exempt name was not measured at all.

The exempt list

Tool Bytes Tracked in
get_collection_health 43,115 #4198

#4268 made get_collection_health compact only healthy collectors with nothing to report. This fixture's collectors are not healthy on purpose, so nothing compacts and the answer stays over the budget. Darling's test exempts the same tool for the same reason. #4268 is merged, and the tool stays open under #4198. No other Lite tool measured over the budget.

The largest answers on this fixture

Tool Bytes
get_collection_health 43,115 (exempt)
get_query_store_regressions 24,252
get_plan_corrections 19,696
get_long_query_completions 17,269
get_query_store_top 16,106
get_collection_log 15,938
get_blocked_process_reports 15,879
get_blocked_process_xml 13,203
get_deadlock_detail 12,876
get_deadlocks 5,972
get_blocking_stats 3,660

The test measured 81 read tools: 89 tools, minus the write tool and the 7 skipped by name. The other 70 answered 2,373 bytes or less.

Known gaps

Some tools pass without a real test:

  • A required parameter with no usable default. The test passes a placeholder, so these tools answer empty or invalid. They are get_mute_rules (its MuteRuleService is not built by the test), get_perfmon_trend (counter_name), get_plan_xml (query_hash), get_query_trend (query_hash, database_name) and get_wait_trend (wait_type).
  • Not seeded. As in Darling's test, the fixture has no rows for get_index_usage and get_object_locking (index_object_stats), get_query_heatmap, get_top_queries_by_cpu and get_top_procedures_by_cpu (query_stats), or get_analysis_findings.

Test plan

CHANGELOG entry

None. This PR changes tests only.

🤖 Generated with Claude Code

https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ

erikdarlingdata and others added 2 commits September 25, 2026 10:32
Part of #4198. Mirrors Darling's McpReadToolBudgetLiveTests: reflects over
every [McpServerTool] in the Lite service assembly, excludes analyze_*/
compare_*/audit_config/mute_analysis_finding, binds each remaining tool's
parameters generically and asserts the reply stays under
McpResponseBudget.DefaultBytes on one busy DuckDB-only fixture.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
The exempt row's comment said "#4268, still open", but #4268 merged. Word it
like Darling's twin row: #4268 compacts only healthy collectors, this
fixture's are not, so the tool stays over and open under #4198.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant