You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Agentic QE's MCP server accepts typed inputs but does not publish or enforce typed output contracts. On protected main at release v3.14.3 SHA ffc0c6a5173874070bb2d07b29e9ae2a7383ed55:
ToolDefinition has parameters and annotations but no outputSchema.
handleToolsList() advertises only name, description, annotations, and inputSchema.
Successful calls, including cache hits, are serialized only into a text content block (lines 647-665, lines 673-729). There is no structuredContent, server-side result validation, or output-contract version.
The caught execution-error path returns a JSON-looking text block but does not set isError: true (lines 730-767). A client can therefore receive an execution failure in a protocol-success-shaped result.
ToolResult<T> gives internal TypeScript callers a generic envelope, but that compile-time type neither proves a concrete tool's payload at runtime nor survives as an advertised wire contract.
This permits two silent-failure classes:
A handler, middleware, bridge, or cache can omit, rename, truncate, or mistype fields while still returning a normal text result.
A caught tool failure can look successful to MCP clients that use isError rather than reparsing arbitrary JSON text.
The issue is not “make every result huge” or “trust schemas as truth.” It is: define the minimum successful payload for each stable tool, validate it before emission and cache admission, and preserve explicit error/incomplete dispositions at the wire boundary.
Why now
Agent-tool completeness evidence
Silent Failures in Agent–Tool Interaction audited 15 scientific tools (104 wrappers) after sampling 63 ToolUniverse resources. It manually validated 91 failures; 51 were attributed to API-layer gaps and 25 to wrapper-layer gaps. Missing information was the most common issue family, and four cases involved error handling or false-success declarations. The study is useful as a failure taxonomy, not a prevalence estimate: it is ToolUniverse-specific, selected 15 tools after candidate discovery, used LLM-assisted discovery/adjudication, and manually rechecked only a subset.
The MCP tools specification permits outputSchema and structuredContent; when an output schema is provided, servers must return conforming structured results and clients should validate them. The official TypeScript SDK validates structuredContent against outputSchema before a result leaves the server.
The new clean-control backtest benchmark shows why contract tests need matched clean controls rather than malformed-only cases. Across 1,440 audits, one open-prompt condition flagged 93.8% of clean code controls; reporting recall alone would have hidden a 79-point separation among models. The dataset is small and synthetic, but the design principle transfers: output validation must reject damaging mutations without flagging valid optional/forward-compatible payloads.
Start with a narrow, high-value slice such as fleet_status, task_status, pipeline_status, quality_assess, security_scan, and one QE slash-named tool. Do not bulk-generate permissive {} schemas.
For a contracted successful call:
validate the post-middleware result against the declared JSON Schema before caching or emission;
return it in structuredContent and retain equivalent serialized JSON in content for compatibility;
fail closed with a typed output_contract_violation tool error if required fields, types, discriminants, or completeness metadata are invalid;
never cache a result that failed output validation;
bind cache entries to tool-output contract version/hash so a schema change cannot replay an older shape as current evidence.
P0 — truthful MCP error semantics
Every caught tool execution error must return isError: true with a stable, sanitized machine-readable error code/disposition.
Preserve the human-readable content block, but do not require clients to parse free-form JSON text to learn that a call failed.
Distinguish protocol errors (unknown tool/malformed call), tool execution errors, output-contract violations, partial/incomplete success where explicitly supported, and unavailable evidence.
A middleware exception or serializer failure must not be emitted as an apparently successful empty/partial payload.
P1 — contract derivation, parity, and compatibility
Prefer one source of truth from concrete result schemas/types; generated schemas must be checked in or deterministically reproducible and reviewed.
Preserve output contracts through hard-coded tools, registry/QE bridge tools, lazy/dynamic registration, tools/list, stdio and HTTP paths, cache hits, post-result middleware, and packaged artifacts.
Unknown/plugin tools remain uncontracted and visibly unqualified; absence of a schema must never imply validated output.
Allow additive optional fields when the contract says additive, while required/discriminated fields remain strict.
For external/open-world tools whose structural schema cannot establish semantic completeness, add a separate optional receipt rather than overclaiming what JSON Schema proves:
This receipt complements #651's downstream verification-reach contract and #695's bounded-buffer completeness; it does not claim that structural validation proves factual correctness.
Successful contracted calls return conforming structuredContent plus equivalent text content for compatibility.
Server-side validation occurs after result middleware and before cache admission or wire emission.
Missing required fields, wrong types, invalid discriminants, and forbidden false-success shapes return isError: true with output_contract_violation.
Caught handler, middleware, and serialization errors always set isError: true and expose no stack trace, secret, or raw sensitive payload.
Cache entries carry output-contract identity; stale-shape cache hits are rejected or invalidated.
Hard-coded, registry-backed, slash-named, bridge, stdio, HTTP, and packaged-artifact paths preserve the same contract.
Unknown/plugin tools are explicitly unqualified rather than assigned a permissive schema or treated as validated.
Additive optional fields pass under an additive contract, while removal/rename/type mutation of required fields fails.
Structurally valid but partial external results carry a distinct completeness disposition where the tool can determine one.
tools/list and tools/call remain compatible with clients that ignore outputSchema/structuredContent.
The contract inventory and tests cover every tool claimed as qualified; no generated schema can silently become {} or omit required evidence fields.
Validation experiments
Required-field mutation: delete one required field from each contracted handler result; every call must fail before cache/write-back and emit isError: true.
Type/discriminant mutation: change a numeric count to a string and a terminal disposition to an unknown value; both must fail deterministically.
Clean controls: add valid optional fields, reorder object keys, and use allowed empty arrays/nullables; valid payloads must not be over-rejected.
False-success error path: inject handler, middleware, and JSON serialization failures; clients must observe isError: true without reparsing content text.
Cache poisoning: cache a v1-shaped result, advance the output contract, and verify the old entry cannot satisfy v2.
Middleware corruption: remove a field in a post-result hook; validation after middleware must catch it.
Transport parity: compare stdio, HTTP, direct registry, and QE bridge results for schema, structured content, error flag, and contract identity.
Package smoke: install the packed tarball, call the selected vertical slice, and compare its live inventory/results with exact source.
Partial-result control: simulate pagination truncation or provider capping; the payload can be structurally valid but must not claim complete.
Backward compatibility: use a client that ignores structured content and a validating current client; both receive usable content, while only conforming payloads leave the server.
Fuzz/property test: generate conforming and single-fault payloads from each schema; measure false accepts and false rejects separately.
Sensitive error test: throw values containing paths, tokens, and payload fragments; the typed error stays useful while redacting sensitive material.
Duplicate review and scope
Targeted live searches for outputSchema, structuredContent, MCP isError, tool-result schemas, response contracts, silent failure, and completeness found no issue owning this wire-level output contract.
feat(mcp): publish explicit tool safety annotations #681 owns MCP safety annotations (readOnlyHint, destructiveHint, idempotentHint, openWorldHint); annotations describe side-effect expectations, not result shape or error truth.
This issue does not claim JSON Schema proves factual correctness, require contracts for every dynamic plugin immediately, replace domain-specific oracles, or make a successful MCP call sufficient evidence for release/deployment.
Suggested priority
P0 for truthful error signaling; P1 for the first contracted tool slice. The implementation should be incremental: land the shared validator/wire semantics first, then qualify stable tools with mutation and clean-control tests.
Problem
Agentic QE's MCP server accepts typed inputs but does not publish or enforce typed output contracts. On protected
mainat release v3.14.3 SHAffc0c6a5173874070bb2d07b29e9ae2a7383ed55:ToolDefinitionhas parameters and annotations but nooutputSchema.handleToolsList()advertises onlyname,description,annotations, andinputSchema.structuredContent, server-side result validation, or output-contract version.isError: true(lines 730-767). A client can therefore receive an execution failure in a protocol-success-shaped result.ToolResult<T>gives internal TypeScript callers a generic envelope, but that compile-time type neither proves a concrete tool's payload at runtime nor survives as an advertised wire contract.This permits two silent-failure classes:
isErrorrather than reparsing arbitrary JSON text.The issue is not “make every result huge” or “trust schemas as truth.” It is: define the minimum successful payload for each stable tool, validate it before emission and cache admission, and preserve explicit error/incomplete dispositions at the wire boundary.
Why now
Agent-tool completeness evidence
Silent Failures in Agent–Tool Interaction audited 15 scientific tools (104 wrappers) after sampling 63 ToolUniverse resources. It manually validated 91 failures; 51 were attributed to API-layer gaps and 25 to wrapper-layer gaps. Missing information was the most common issue family, and four cases involved error handling or false-success declarations. The study is useful as a failure taxonomy, not a prevalence estimate: it is ToolUniverse-specific, selected 15 tools after candidate discovery, used LLM-assisted discovery/adjudication, and manually rechecked only a subset.
Protocol evidence
The MCP tools specification permits
outputSchemaandstructuredContent; when an output schema is provided, servers must return conforming structured results and clients should validate them. The official TypeScript SDK validatesstructuredContentagainstoutputSchemabefore a result leaves the server.Calibration evidence
The new clean-control backtest benchmark shows why contract tests need matched clean controls rather than malformed-only cases. Across 1,440 audits, one open-prompt condition flagged 93.8% of clean code controls; reporting recall alone would have hidden a 79-point separation among models. The dataset is small and synthetic, but the design principle transfers: output validation must reject damaging mutations without flagging valid optional/forward-compatible payloads.
Proposed capability
P0 — versioned output contracts for stable built-ins
Extend tool definitions with an explicit successful-result schema and contract identity:
Start with a narrow, high-value slice such as
fleet_status,task_status,pipeline_status,quality_assess,security_scan, and one QE slash-named tool. Do not bulk-generate permissive{}schemas.For a contracted successful call:
structuredContentand retain equivalent serialized JSON incontentfor compatibility;output_contract_violationtool error if required fields, types, discriminants, or completeness metadata are invalid;P0 — truthful MCP error semantics
isError: truewith a stable, sanitized machine-readable error code/disposition.P1 — contract derivation, parity, and compatibility
tools/list, stdio and HTTP paths, cache hits, post-result middleware, and packaged artifacts.unqualified; absence of a schema must never imply validated output.additive, while required/discriminated fields remain strict.P2 — contextual completeness receipts
For external/open-world tools whose structural schema cannot establish semantic completeness, add a separate optional receipt rather than overclaiming what JSON Schema proves:
This receipt complements #651's downstream verification-reach contract and #695's bounded-buffer completeness; it does not claim that structural validation proves factual correctness.
Acceptance criteria
outputSchemavalues intools/list.structuredContentplus equivalent text content for compatibility.isError: truewithoutput_contract_violation.isError: trueand expose no stack trace, secret, or raw sensitive payload.tools/listandtools/callremain compatible with clients that ignoreoutputSchema/structuredContent.{}or omit required evidence fields.Validation experiments
isError: true.isError: truewithout reparsing content text.complete.Duplicate review and scope
Targeted live searches for
outputSchema,structuredContent, MCPisError, tool-result schemas, response contracts, silent failure, and completeness found no issue owning this wire-level output contract.readOnlyHint,destructiveHint,idempotentHint,openWorldHint); annotations describe side-effect expectations, not result shape or error truth.This issue does not claim JSON Schema proves factual correctness, require contracts for every dynamic plugin immediately, replace domain-specific oracles, or make a successful MCP call sufficient evidence for release/deployment.
Suggested priority
P0 for truthful error signaling; P1 for the first contracted tool slice. The implementation should be incremental: land the shared validator/wire semantics first, then qualify stable tools with mutation and clean-control tests.