Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

PatchWitness

PatchWitness is a local app that checks whether an AI-made code change actually fixes the requested behavior. It runs the same checks before and after the change, then shows the evidence in a report you can inspect and replay.

PatchWitness adds no cloud storage: its worktrees, settings, and reports stay on your computer, and it stores credential references rather than secret values. During a live run, the agent/provider you configure may process repository context under that provider’s own policies. Open sample evidence calls no provider.

Start here — no programming experience required

What does PatchWitness do?

Imagine asking an AI to repair something in a program. A confident answer such as “fixed” is not proof. PatchWitness gives the work to two separate agents:

  1. The Builder makes the change in a temporary copy of your project.
  2. A fresh Verifier decides what should be tested without seeing the Builder’s private reasoning.
  3. PatchWitness runs the checks on both the original and changed copies.
  4. You get a report showing what failed before, what passed afterward, and whether anything else broke.

PatchWitness never gives the Builder permission to approve its own work.

Try it without setting up an AI

This is the easiest first step:

  1. Start PatchWitness using the installation steps below.
  2. Open the local web address shown in the terminal, normally http://localhost:3000.
  3. Press Open sample evidence.

The sample is already prepared and calls no AI provider. It lets you explore a finished report, including the before/after test evidence, in under a minute. Sample reports are clearly marked SAMPLE / DETERMINISTIC so they cannot be confused with live AI verification.

Install and open PatchWitness

You need:

  • Node.js version 20 or newer. The current LTS version is the easiest choice.
  • Git, which keeps the project history PatchWitness compares.
  • macOS, Windows, or Linux. On Linux, the Choose file/folder window uses Zenity; if it is unavailable, paste the path instead.

Download this project, then open Terminal on macOS/Linux or PowerShell on Windows. Move into the downloaded PatchWitness folder and run:

npm install
npm run dev

Open the address printed after “Local” in your browser. Keep the terminal window open while using PatchWitness. Press Ctrl+C in that terminal when you want to stop it.

If you prefer to download with Git:

git clone https://github.com/GameAssasin/PatchWitness.git
cd PatchWitness
npm install
npm run dev

Check one of your own projects

A “Git project” is a project folder whose changes have been saved with Git. PatchWitness needs at least one saved commit because it creates clean before-and-after copies from that history.

  1. On New Verification, press Choose folder or Choose file. You can choose the project root, a folder inside it, or one committed file.
  2. Leave Auto-fill project defaults on and press Validate. PatchWitness finds the project root, current branch or commit, nearby JavaScript test command, dependency lockfile, and a safe path scope.
  3. Read the Requested behavior draft. It is only a draft; PatchWitness has not discovered a bug for you. Edit it so it says exactly what should change, then press Confirm requested behavior.
  4. Review any warning or blocker shown in Project preflight.
  5. Press Start verified change.

A live run needs a working Codex installation or another compatible agent/provider configured under Settings. If you only want to explore the product, use Open sample evidence instead.

The ? button turns on Control Explainer. While it is on, select any highlighted button or setting to see a plain explanation and an example. Press Escape to close it.

Which files and folders can I choose?

PatchWitness is extension-agnostic at intake: it does not require a file to end in .js, .ts, or any other particular extension.

Selection Supported? What happens
The Git project root Yes The run may change anything allowed by your scope settings.
A committed folder inside the project Yes PatchWitness keeps that nested folder as the suggested scope.
A committed regular file Yes The exact file becomes the suggested scope.
Names containing spaces, commas, Unicode, or dots Yes The name is preserved exactly through validation and scope checks.
Hidden files and hidden folders Yes They work when your operating-system picker shows them, or when you paste the path.
A regular binary file such as an image Path selection: yes The file can be scoped, but the selected agent and project tests must understand its format. The generic text-tool agent only reads and writes UTF-8 text.
main, master, trunk, a feature branch, tag, commit, or detached HEAD Yes Auto-fill uses the real current branch or exact commit instead of assuming main.
A modified file already present in the selected commit Yes PatchWitness warns that the unsaved Git changes are ignored; the run starts from the committed version.
A new file/folder not present in the selected commit Not yet Commit it first. A temporary worktree cannot contain an uncommitted selection.
A symbolic link No Select the real target inside the repository. Direct links are rejected to avoid path confusion and escapes.
A socket, named pipe, or device file No These special operating-system objects are rejected safely.
A folder that is not a Git repository, or a Git repository with no commits No Initialize Git and create the first commit, then validate again.

This is the supported “all normal files and folders” boundary. No app can safely promise every filesystem object on every operating system; PatchWitness rejects ambiguous or non-replayable objects with a specific correction instead of pretending they worked.

If an earlier parent directory is an operating-system alias or symbolic link, PatchWitness resolves and displays the canonical real repository path. The selected file or folder itself still cannot be a symbolic link.

Understand the result

  • Original FAIL → Patched PASS: the requested behavior changed in the right direction.
  • Original PASS → Patched PASS: an existing behavior was preserved.
  • Original PASS → Patched FAIL: the change caused a regression.
  • NONRUNNABLE, timeout, blocked command, or missing tool: PatchWitness could not collect reliable evidence. The Overview explains the primary cause and what to do next.

The report is evidence for review, not a promise that every possible behavior was tested.

Common problems

  • Start button is disabled: read the sentence beside the button. Usually you still need to validate, confirm the requested-behavior draft, enter a test command, or fix a preflight blocker.
  • “Base branch or commit does not exist”: turn on Auto-fill, or enter HEAD.
  • “Selected file or folder is not part of the current commit”: commit that item in Git, then select it again.
  • “No test command was detected”: enter the command your project normally uses, such as npm test.
  • The folder window does not open: paste the full path into Project file or folder, then press Validate.
  • Provider or model error: open Settings → Providers, test the connection, then run Test capabilities for the model. PatchWitness shows whether it can be a Builder, Verifier, both, or neither.
  • A run failed: open its Overview and read Why this run needs attention. Use Start a new run with this configuration after correcting the problem; PatchWitness will not silently reuse an old temporary worktree.

A short safety note

PatchWitness runs code and tests from the selected repository. Only use projects you trust, or run untrusted projects inside a proper virtual machine or operating-system sandbox. The app’s network and command policies reduce mistakes but are not a complete security boundary.


Developer reference

Evidence contract

PatchWitness is a local, replayable verification harness for AI-authored code changes. A Builder implements a task, a fresh Verifier receives the task, original repository, final diff, and observable baseline results, and the same behavioral claims run against detached original and patched worktrees.

The key counterfactual classifications are:

  • FAIL → PASS (BEHAVIOR_CHANGED) is causal evidence that the patch changed the exercised behavior.
  • PASS → PASS (PRESERVED) is preservation evidence, not proof of the requested fix.
  • PASS → FAIL (REGRESSION), non-runnable commands, scope violations, and missing executions block or qualify the verdict.

PatchWitness intentionally emits no numerical confidence score. A score would hide which claim ran, on which revision, and whether execution completed. Recorded execution outranks Builder/Verifier prose; agent summaries remain operational context and never certify a patch.

Architecture

flowchart LR
  UI[Local Next.js UI / API] --> O[RunOrchestrator]
  O --> V[Repository validation\nbase revision]
  V --> W1[Original worktree]
  V --> W2[Patched worktree]
  W2 --> B[Selected Builder runtime\nworkspace-write]
  B --> P[Patch + scope capture]
  P --> I[Fresh Verifier runtime\nread-only original worktree]
  I --> T[Generated tests + selected\nexisting commands]
  T --> E1[Execute on original]
  T --> E2[Execute on patched]
  E1 --> C[Counterfactual classifier]
  E2 --> C
  C --> R[HTML + JSON report\nreplay command]
  O --> D[(Local .patchwitness data)]
Loading

Builder and Verifier are independently assigned versioned profiles containing runtime, provider, model, supported parameters, permissions, timeout, and approved fallback chain. Codex is the ready default. The generic tool-agent adapter supports OpenAI Responses-compatible and local endpoints; the deterministic runtime remains restricted to explicit tests and sample evidence.

Each run snapshots the resolved profiles, capabilities, command and verification policies, budgets, permissions, fallbacks, and verifier-isolation decision. Later settings edits do not rewrite historical run configuration. The Verifier never receives Builder reasoning, confidence, self-assessment, or conversation state.

Technical requirements and commands

  • Node.js 20 or newer (engines.node is >=20).
  • Git on PATH.
  • macOS, Linux, or Windows. POSIX-like environments have the broadest runtime coverage; native picker and process-group handling are platform-specific.
  • A supported Codex runtime/authentication or configured compatible provider for live runs. PatchWitness stores no provider credential values.
npm install
npm run dev

Useful project checks:

npm run typecheck
npm run lint
npm test
npm run build
npm run test:e2e

Repository intake and preflight

The picker and validation API accept an absolute path to a committed regular file or directory inside a local Git repository. Validation canonicalizes the repository root while retaining the original selected path and scope. It uses argument arrays rather than a shell, and Git’s NUL-delimited status/numstat formats preserve spaces, commas, quotes, and Unicode in changed paths.

With Auto-fill project defaults enabled, /api/repositories/validate returns the existing repository fields plus an additive preflight object containing:

  • resolved revision and inference source;
  • detected test command;
  • package manager and lockfile;
  • dependency-preparation strategy;
  • selected allowed scope;
  • editable requested-behavior draft;
  • warnings and blocking issues.

Auto-fill never claims a defect was discovered. The user must review or edit its generic task draft. A manually entered revision is resolved and reported exactly when Auto-fill is off. Only concrete blockers disable Start, and the correction is displayed beside the button.

Built-in command detection and locked dependency preparation currently target JavaScript projects using npm, pnpm, Yarn, or Bun. Other language repositories can still be selected because intake is file-extension agnostic, but the operator must supply the appropriate safe test command and prepare any required tooling. This release does not claim universal language-runner support.

Provider configuration

  • Codex: use the built-in runtime, Codex authentication provider, and runtime-default model. Authentication stays with the installed runtime.
  • OpenAI or OpenAI-compatible: add a provider, supply its base URL when required, and store only an environment-variable name such as OPENAI_API_KEY. Set that variable in the process launching PatchWitness.
  • Local endpoint (LM Studio, compatible local servers): choose Local, use the loopback/base URL, select no credential when supported, test the connection, then discover or register a model.

Connection success proves only reachability. Test capabilities separately probes a plain response, sequential list/read/write tool calls, and Verifier-style JSON without touching repository files. The UI reports named results and whether a model is eligible for Builder, Verifier, both, or neither. Incompatible profiles cannot be assigned, while manually configured OpenAI-compatible providers remain supported.

When writing a task, describe the repository outcome in ordinary language. Do not paste tool-call JSON, request direct read_file/write_file calls, or use absolute paths. Malformed or multiple tool calls stop before execution and include a model/template correction instead of exposing only a raw HTTP error.

Provider credentials are references only. Raw tokens and custom-header values are rejected from settings, effective-run snapshots, and reports. Every actual fallback is recorded with its reason and policy effect.

Existing flat settings migrate atomically to schema version 2. Legacy model, reasoning, timeout, network, scope, and verification choices map to separate Builder and Verifier profiles. A legacy executable allowlist becomes a visibly marked compatibility policy for review.

Live demo and deterministic sample

Prepare the live fixture:

npm run demo:setup
npm run dev

Choose Use demo repository, review its task and src, tests scope, then start the run. This is a real provider-backed run and requires the selected runtime/authentication.

Choose Open sample evidence for the zero-provider path. It creates the same fixture’s evidence with deterministic in-process adapters, calls no provider, marks every surface SAMPLE / DETERMINISTIC, and remains filterable separately from live records. It validates the UI, orchestration, counterfactual classifier, and report flow—not live model behavior.

Terminal equivalents:

npm run demo:setup
npm run demo:run                 # live selected runtime
npm run demo:run -- --test-adapter  # explicit deterministic test path

Replay uses the command recorded in a completed report, for example npm run demo:run -- --replay <run-id>. It reuses recorded input and current settings but creates fresh worktrees.

How Codex and GPT-5.6 were used

PatchWitness was developed through iterative Codex sessions during OpenAI Build Week. Codex and GPT-5.6 were used to trace the repository workflow, implement the isolated Builder/Verifier architecture, harden local and OpenAI-compatible provider handling, classify truthful failure diagnostics, build the deterministic sample, repair responsive and accessible UI states, and audit the final path-compatibility matrix and documentation.

Important product decisions remained explicit and reviewable: execution evidence outranks agent prose; the Builder cannot certify itself; sample evidence is labeled separately from live evidence; credentials never enter reports; and unsupported filesystem objects fail closed. Automated changes were checked with focused regression tests followed by typecheck, lint, build, integration, and browser validation rather than accepted from agent summaries.

Security and containment

Each run creates detached original and patched Git worktrees in a private OS-temporary directory outside both the source checkout and PatchWitness data root. Only records and reports persist under the data root. Generated tests are limited to .patchwitness/generated/ in each isolated worktree.

Commands launch with shell: false. Structured policies match executable, arguments, role, and network policy server-side. Shell metacharacters and escaping paths are rejected, working directories must be absolute existing paths, and output/errors are sanitized and bounded. Builder package installation is policy-gated; Verifier installation is blocked. Data/worktree parents use mode 0700, and persisted JSON, HTML, patches, and generated tests use 0600 where supported. Existing symlink segments are rejected on controlled tool and test paths.

Network access is disabled by default for Codex threads, and package managers receive offline flags when disabled. Node child processes do not provide OS-level network isolation. Repository commands and generated tests execute with the host user’s authority, so untrusted repositories require an external sandbox or VM.

Local data layout

The default data root is <PatchWitness project>/.patchwitness; override it with an absolute PATCHWITNESS_DATA_DIR.

.patchwitness/
├── settings.json                 # versioned policies; credential references only
├── runs/<run-id>/
│   ├── run.json                  # persisted run record
│   ├── report.html               # standalone report
│   ├── report.json               # sanitized structured evidence
│   └── submitted.patch           # tests-only input, when used

<OS temporary directory>/patchwitness-<run-id>-*/
├── original/                     # detached base revision
└── patched/                      # Builder or supplied patch

Failed worktrees are preserved by default. After confirming a run is inactive, use Cleanup or POST /api/runs/:id/cleanup with { "confirm": "<run-id>" }. Cleanup removes only temporary worktrees, not evidence records. Removing .patchwitness deletes local evidence and should be an explicit operator action.

Reports, replay, and diagnostics

Every completed or failed run attempts to write report.html and report.json. Reports include the primary structured failure diagnosis; base commit and task; patch, changed files, and scope findings; truthful blocked/failed stage timeline; Builder/Verifier identities and disclosure boundary; immutable runtime/profile/policy snapshots; fallback events; acceptance claims; exact commands and working directories; sanitized output; original/patched outcomes; counterfactual classifications; final verdict; recommended action; and replay command.

Optional diagnostics: FailureDiagnostic[] records category, stage, command, explanation, and recommended action. Older runs without stored diagnostics remain readable and derive diagnostics from recorded commands/errors when opened. Policy blocks, missing tools, dependency failures, timeouts, provider failures, invalid agent output, baseline failures, and ordinary test failures remain distinct.

Developer troubleshooting

  • No test command detected: provide a safe override or change the baseline-test policy. Baseline execution is required by default.
  • Base revision missing: use Auto-fill or HEAD; feature branches, tags, commits, and detached checkouts are supported.
  • Selection not in current commit: commit the file/folder. Dirty tracked changes are recorded as ignored source state because worktrees use the resolved commit.
  • Baseline exits nonzero: the run may continue so an existing failure can become valid FAIL → PASS evidence. The counterfactual comparison determines the verdict.
  • Codex authentication/runtime failure: authenticate the Codex runtime in the environment that launches Next.js or demo:run.
  • Network/proxy URL renders but does not respond: development allows current machine interface addresses. For a custom proxy hostname, set PATCHWITNESS_DEV_ORIGINS=your-hostname and restart.
  • Builder produced no patch: inspect the task and preserved worktree, then start a fresh run.
  • NONRUNNABLE, timeout, or cancellation: inspect the structured diagnostic, command, revision, and stderr. These never become passes.
  • Scope violation or regression: inspect exact changed paths and PASS → FAIL rows, correct the task/scope, and rerun.

PatchWitness’s own test suite proves the product code paths it exercises. It does not replace one live repository verification with the provider/runtime named in a submission or release claim.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages