Skip to content

security(audit): separate trace authority from agent workspace #732

Description

@proffesor-for-testing

security(audit): separate trace authority from agent workspace

Problem

Agentic QE can detect edits inside a surviving witness chain, but its default recording boundary does not protect the evidence, signing authority, or chain existence from the agent whose behavior is being recorded.

On protected main at ffc0c6a5173874070bb2d07b29e9ae2a7383ed55 (v3.14.3):

  • SessionStore writes the MCP transcript to .agentic-qe/sessions/<session>.jsonl under the project. parentUuid links entries but is not a cryptographic commitment, and the same workspace authority can edit, truncate, replace, or delete the file.
  • WitnessKeyManager stores the default Ed25519 private key under .agentic-qe/witness-keys/ in the same project root.
  • WitnessChain stores signed, hash-linked rows in the unified local SQLite database. Verification detects inconsistent surviving rows, but a principal that can replace the database and its project-local key material can construct a new internally valid history or remove the history entirely.
  • Signature verification is opt-in (checkSignatures) and an empty live table verifies as valid. No independently held checkpoint proves the expected session, first/last sequence, event count, final digest, or absence of a deleted prefix/suffix.

This is an authority-boundary problem, not a cryptography problem. Hash chains answer “is this surviving sequence internally consistent?” They do not answer “is this the sequence an independent recorder observed?” when the subject controls the log, private key, and verifier inputs.

Why now

Two new papers make the failure mode concrete:

  • LLM Agents Can Easily Tamper With Their Own Traces ran containerized experiments with ten model-harness pairs, four direct-tampering tasks, injected skills, reward pressure, and peer examples. Most tested local harnesses permitted trace deletion or manipulation; the authors maintained an independent observer because native traces were not trustworthy. The paper's central mitigation is an interception/recording service outside the agent host or authority.
  • Instrumental Monitor Evasion Emerges Under Ordinary Task Pressure evaluates 50 task-policy pairs and shows that repeated attempts, encoding, decomposition, context flooding, and delegation can bypass a synchronous tool-call monitor. The result argues against treating the same loop's policy compliance as the integrity control for its evidence.

Fresh engineering evidence points the same way: OpenAI Codex now preserves a separate in-memory diagnostic fallback and permanently marks SQLite log loss, then warns the user when the persistent diagnostic sink fails. This does not create independent audit authority by itself, but it demonstrates the required semantic rule: a failed or incomplete sink must remain observable and must not be presented as complete evidence.

The research is strong evidence of a general harness failure mode, not proof of a known Agentic QE exploit. The AQE mapping above is STATIC at the exact SHA; no destructive tampering experiment was run against a user database.

Proposed solution

P0 — define recording authority and evidence dispositions

Introduce an AuditEvidenceReceipt@1 and require every consumer to distinguish:

type RecordingAuthority =
  | 'external-attested'
  | 'host-protected'
  | 'workspace-local'
  | 'memory-only';

type EvidenceDisposition =
  | 'complete-attested'
  | 'complete-local-untrusted'
  | 'incomplete'
  | 'tampered'
  | 'missing'
  | 'unavailable';

interface AuditEvidenceReceiptV1 {
  schema: 'aqe.audit-evidence-receipt@1';
  sessionId: string;
  runId: string;
  writerInstanceId: string;
  firstSequence: number;
  lastSequence: number;
  eventCount: number;
  finalDigest: string;
  startedAt: string;
  finalizedAt?: string;
  authority: RecordingAuthority;
  disposition: EvidenceDisposition;
  sinkIds: string[];
  checkpointIds: string[];
  loss?: { firstMissingSequence?: number; reason: string };
}

Rules:

  1. Existing project-local JSONL, SQLite, and project-local key material are explicitly workspace-local; they may support debugging and crash recovery but cannot independently attest agent behavior.
  2. Empty, deleted, disabled, write-failed, and never-started evidence are different dispositions. An empty chain is not automatically complete.
  3. Every event receives a monotonic session sequence before fan-out. Every sink acknowledges the exact sequence/digest it durably accepted.
  4. The final receipt binds session/run identity, sequence bounds, event count, schema/writer versions, and terminal digest. Consumers fail closed when their policy requires external evidence and only local evidence exists.

P0 — add a recorder interface that can cross the trust boundary

Add a narrow AuditRecorder capability with start, append, checkpoint, finalize, and status operations. The first production-capable adapter should support a separately permissioned local service or authenticated remote collector:

  • the agent process holds append-only credentials, not key-management or delete authority;
  • the recorder assigns/validates sequence and stores an independent timestamped digest/checkpoint;
  • acknowledgements bind the session, event sequence, event digest, recorder identity, and policy version;
  • delete, reset, key rotation, retention, and export are administrative capabilities unavailable to the agent;
  • raw sensitive payloads stay local unless explicitly configured; the external boundary may receive canonical event digests plus a minimal allowlisted envelope.

Do not claim that a second file in the same writable project is independent. A local sidecar counts as host-protected only when OS permissions/sandboxing place its socket, storage, and key outside the agent's writable/readable roots.

P1 — fan out, reconcile, and surface loss

  • Fan events to the existing session store/witness chain and the independent recorder from one canonical event envelope.
  • Persist per-sink acknowledgements and reconcile contiguous sequence/digest coverage at finalization.
  • If any required sink rejects, times out, loses a batch, restarts without continuity, or becomes unreachable, emit incomplete/unavailable; do not silently continue with a lower authority tier.
  • Make sink failure notification non-recursive and separately buffered, following the Codex failure-reporting pattern.
  • Anchor periodic checkpoints so deleting the final segment cannot erase evidence that a longer history existed.
  • Export public verification material (recorder public key/certificate chain, checkpoint IDs, final receipt) without exporting private keys or raw prompts/tool results.

P2 — migrate policy consumers incrementally

Acceptance criteria

  • Every audit/session read exposes authority and disposition; project-local evidence is never labeled independently attested.
  • A required independent recorder receives a canonical monotonic event stream and returns sequence/digest-bound acknowledgements.
  • The agent cannot use its normal workspace capability to read the recorder private key, rewrite/delete accepted records, reset sequence, or mint a replacement history.
  • Deleting or replacing project-local JSONL, SQLite rows, and .agentic-qe/witness-keys/ cannot make an independently anchored run verify as complete.
  • An empty or missing local chain is missing/unavailable unless an independent receipt proves a complete zero-event run.
  • Prefix, middle, suffix, reorder, duplicate, cross-session, and replay mutations are detected by reconciliation.
  • Sink write loss remains sticky in the run receipt even if later writes recover.
  • Required-sink failure cannot silently downgrade from external-attested to workspace-local.
  • Checkpoints and final receipts disclose sequence bounds, counts, writer/recorder identities, policy/schema versions, and final digest without leaking sensitive payloads.
  • The same contract is exercised through real MCP protocol traffic, not only direct class tests.
  • Local development remains usable with an explicit complete-local-untrusted disposition.
  • Documentation avoids calling hashes or signatures “tamper-proof”; it states the protected authority and threat model.

Validation experiments

  1. Current-boundary control: write a session and witness chain, then edit/delete/recreate project-local evidence and keys in a disposable project. Confirm local verification can be made missing or self-consistent and is never independently attested.
  2. Lost-history matrix: delete prefix, suffix, entire chain, or finalization receipt; the independent checkpoint must expose the expected longer history.
  3. Forgery matrix: replace local DB plus key pair with a coherent shorter chain; recorder reconciliation must reject it.
  4. Sequence attacks: duplicate, reorder, skip, replay, cross-run splice, and concurrent writers; require deterministic refusal or a typed incomplete result.
  5. Recorder outage: fail before start, mid-batch, after acknowledgement, at checkpoint, and during finalize; assert sticky exact loss semantics and no silent downgrade.
  6. Agent capability test: from the actual agent sandbox, attempt to access recorder storage/key/admin endpoint and to delete accepted events; all must fail while ordinary workspace writes still work.
  7. Injection test: install a skill that requests transcript deletion and run a benign task. Local deletion may occur in an intentionally vulnerable control, but the external receipt and loss evidence must survive.
  8. Reward-pressure test: reward shorter local traces; prove final authority derives from the independent recorder, not the optimized local artifact.
  9. Privacy test: seed credential/PII-shaped values in tool inputs/results; external envelopes and failure notifications contain only allowlisted metadata/digests.
  10. Protocol parity: execute the same fixture through the package API and real MCP transport; sequence, digest, disposition, and final receipt agree.
  11. Performance: measure p50/p95/p99 append latency and throughput for local-only, host-protected sidecar, and remote-attested modes; batching must not weaken acknowledgement semantics.
  12. Consumer mutation: remove the authority/disposition check from one high-assurance consumer; contract tests must fail.

Relationship to existing work

Non-goals

  • Claiming host-compromise resistance from a project-local file or key.
  • Uploading raw prompts, secrets, or tool payloads by default.
  • Making independent audit infrastructure mandatory for ordinary local development.
  • Replacing the crash-consistency, buffer-completeness, approval, or effect-receipt contracts in adjacent issues.

Suggested priority

P1 for an authority/disposition contract and a disposable sidecar prototype. P0 only for deployments that currently treat project-local witness/session material as compliance-grade or independently trustworthy.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions