Skip to content

feat(mcp): validate structured tool outputs end to end #729

Description

@proffesor-for-testing

Problem

Agentic QE's MCP server accepts typed inputs but does not publish or enforce typed output contracts. On protected main at release v3.14.3 SHA ffc0c6a5173874070bb2d07b29e9ae2a7383ed55:

  • ToolDefinition has parameters and annotations but no outputSchema.
  • handleToolsList() advertises only name, description, annotations, and inputSchema.
  • Successful calls, including cache hits, are serialized only into a text content block (lines 647-665, lines 673-729). There is no structuredContent, server-side result validation, or output-contract version.
  • The caught execution-error path returns a JSON-looking text block but does not set isError: true (lines 730-767). A client can therefore receive an execution failure in a protocol-success-shaped result.
  • ToolResult<T> gives internal TypeScript callers a generic envelope, but that compile-time type neither proves a concrete tool's payload at runtime nor survives as an advertised wire contract.

This permits two silent-failure classes:

  1. A handler, middleware, bridge, or cache can omit, rename, truncate, or mistype fields while still returning a normal text result.
  2. A caught tool failure can look successful to MCP clients that use isError rather than reparsing arbitrary JSON text.

The issue is not “make every result huge” or “trust schemas as truth.” It is: define the minimum successful payload for each stable tool, validate it before emission and cache admission, and preserve explicit error/incomplete dispositions at the wire boundary.

Why now

Agent-tool completeness evidence

Silent Failures in Agent–Tool Interaction audited 15 scientific tools (104 wrappers) after sampling 63 ToolUniverse resources. It manually validated 91 failures; 51 were attributed to API-layer gaps and 25 to wrapper-layer gaps. Missing information was the most common issue family, and four cases involved error handling or false-success declarations. The study is useful as a failure taxonomy, not a prevalence estimate: it is ToolUniverse-specific, selected 15 tools after candidate discovery, used LLM-assisted discovery/adjudication, and manually rechecked only a subset.

Protocol evidence

The MCP tools specification permits outputSchema and structuredContent; when an output schema is provided, servers must return conforming structured results and clients should validate them. The official TypeScript SDK validates structuredContent against outputSchema before a result leaves the server.

Calibration evidence

The new clean-control backtest benchmark shows why contract tests need matched clean controls rather than malformed-only cases. Across 1,440 audits, one open-prompt condition flagged 93.8% of clean code controls; reporting recall alone would have hidden a 79-point separation among models. The dataset is small and synthetic, but the design principle transfers: output validation must reject damaging mutations without flagging valid optional/forward-compatible payloads.

Proposed capability

P0 — versioned output contracts for stable built-ins

Extend tool definitions with an explicit successful-result schema and contract identity:

interface ToolOutputContract {
  schemaVersion: string;
  outputSchema: JsonSchema2020_12;
  compatibility: 'exact' | 'additive';
  sensitivePaths?: string[];
}

interface ToolDefinition {
  // existing fields...
  output?: ToolOutputContract;
}

Start with a narrow, high-value slice such as fleet_status, task_status, pipeline_status, quality_assess, security_scan, and one QE slash-named tool. Do not bulk-generate permissive {} schemas.

For a contracted successful call:

  1. validate the post-middleware result against the declared JSON Schema before caching or emission;
  2. return it in structuredContent and retain equivalent serialized JSON in content for compatibility;
  3. fail closed with a typed output_contract_violation tool error if required fields, types, discriminants, or completeness metadata are invalid;
  4. never cache a result that failed output validation;
  5. bind cache entries to tool-output contract version/hash so a schema change cannot replay an older shape as current evidence.

P0 — truthful MCP error semantics

  • Every caught tool execution error must return isError: true with a stable, sanitized machine-readable error code/disposition.
  • Preserve the human-readable content block, but do not require clients to parse free-form JSON text to learn that a call failed.
  • Distinguish protocol errors (unknown tool/malformed call), tool execution errors, output-contract violations, partial/incomplete success where explicitly supported, and unavailable evidence.
  • A middleware exception or serializer failure must not be emitted as an apparently successful empty/partial payload.

P1 — contract derivation, parity, and compatibility

  • Prefer one source of truth from concrete result schemas/types; generated schemas must be checked in or deterministically reproducible and reviewed.
  • Preserve output contracts through hard-coded tools, registry/QE bridge tools, lazy/dynamic registration, tools/list, stdio and HTTP paths, cache hits, post-result middleware, and packaged artifacts.
  • Unknown/plugin tools remain uncontracted and visibly unqualified; absence of a schema must never imply validated output.
  • Allow additive optional fields when the contract says additive, while required/discriminated fields remain strict.
  • Include contract identity in protocol parity tests and machine-readable inventory alongside feat(mcp): publish explicit tool safety annotations #681's safety annotation disposition.

P2 — contextual completeness receipts

For external/open-world tools whose structural schema cannot establish semantic completeness, add a separate optional receipt rather than overclaiming what JSON Schema proves:

interface ToolEvidenceDisposition {
  sourceRevision?: string;
  retrievedAt?: string;
  requestedScope?: string;
  returnedCount?: number;
  totalKnown?: number;
  paginationComplete?: boolean;
  truncation?: 'none' | 'client' | 'server' | 'unknown';
  limitations: string[];
  disposition: 'complete' | 'partial' | 'unknown' | 'unavailable';
}

This receipt complements #651's downstream verification-reach contract and #695's bounded-buffer completeness; it does not claim that structural validation proves factual correctness.

Acceptance criteria

  • Contracted tools advertise deterministic JSON Schema 2020-12 outputSchema values in tools/list.
  • Successful contracted calls return conforming structuredContent plus equivalent text content for compatibility.
  • Server-side validation occurs after result middleware and before cache admission or wire emission.
  • Missing required fields, wrong types, invalid discriminants, and forbidden false-success shapes return isError: true with output_contract_violation.
  • Caught handler, middleware, and serialization errors always set isError: true and expose no stack trace, secret, or raw sensitive payload.
  • Cache entries carry output-contract identity; stale-shape cache hits are rejected or invalidated.
  • Hard-coded, registry-backed, slash-named, bridge, stdio, HTTP, and packaged-artifact paths preserve the same contract.
  • Unknown/plugin tools are explicitly unqualified rather than assigned a permissive schema or treated as validated.
  • Additive optional fields pass under an additive contract, while removal/rename/type mutation of required fields fails.
  • Structurally valid but partial external results carry a distinct completeness disposition where the tool can determine one.
  • tools/list and tools/call remain compatible with clients that ignore outputSchema/structuredContent.
  • The contract inventory and tests cover every tool claimed as qualified; no generated schema can silently become {} or omit required evidence fields.

Validation experiments

  1. Required-field mutation: delete one required field from each contracted handler result; every call must fail before cache/write-back and emit isError: true.
  2. Type/discriminant mutation: change a numeric count to a string and a terminal disposition to an unknown value; both must fail deterministically.
  3. Clean controls: add valid optional fields, reorder object keys, and use allowed empty arrays/nullables; valid payloads must not be over-rejected.
  4. False-success error path: inject handler, middleware, and JSON serialization failures; clients must observe isError: true without reparsing content text.
  5. Cache poisoning: cache a v1-shaped result, advance the output contract, and verify the old entry cannot satisfy v2.
  6. Middleware corruption: remove a field in a post-result hook; validation after middleware must catch it.
  7. Transport parity: compare stdio, HTTP, direct registry, and QE bridge results for schema, structured content, error flag, and contract identity.
  8. Package smoke: install the packed tarball, call the selected vertical slice, and compare its live inventory/results with exact source.
  9. Partial-result control: simulate pagination truncation or provider capping; the payload can be structurally valid but must not claim complete.
  10. Backward compatibility: use a client that ignores structured content and a validating current client; both receive usable content, while only conforming payloads leave the server.
  11. Fuzz/property test: generate conforming and single-fault payloads from each schema; measure false accepts and false rejects separately.
  12. Sensitive error test: throw values containing paths, tokens, and payload fragments; the typed error stays useful while redacting sensitive material.

Duplicate review and scope

Targeted live searches for outputSchema, structuredContent, MCP isError, tool-result schemas, response contracts, silent failure, and completeness found no issue owning this wire-level output contract.

This issue does not claim JSON Schema proves factual correctness, require contracts for every dynamic plugin immediately, replace domain-specific oracles, or make a successful MCP call sufficient evidence for release/deployment.

Suggested priority

P0 for truthful error signaling; P1 for the first contracted tool slice. The implementation should be incremental: land the shared validator/wire semantics first, then qualify stable tools with mutation and clean-control tests.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions