validrig is the engine of DearAuditor Eval (CLI: rig); formerly "Harness Factory".
Last updated: 2026-09-04
The engine (validrig/) is use-case-agnostic; all use-case content lives in
packs/. Everything below runs fully offline and deterministically via a
first-class fake model and fake judge.
python3.12 -m venv .venv && .venv/bin/pip install -e ".[dev]"
.venv/bin/pytest -q # 225 tests, ~6s
.venv/bin/rig new packs/my-pack --id my-pack # scaffold a new pack
.venv/bin/rig lint packs/demo-tumor-board # check a pack for authoring gaps
.venv/bin/rig run packs/demo-tumor-board --battery smoke --out ./runs --seed 1
.venv/bin/rig run packs/demo-tumor-board --battery regression --out ./runs --seed 1
.venv/bin/rig diff --out ./runs --baseline <run_a> --candidate <run_b>
.venv/bin/rig qms packs/demo-tumor-board --run <run_id> --out ./runs
.venv/bin/rig monitor packs/demo-tumor-board --run <run_id> --events events.jsonl --out ./runs # M5 monitoring
.venv/bin/rig dossier packs/demo-tumor-board --run <run_id> --out ./runs # printable HTML validation dossier
.venv/bin/rig ui packs/demo-tumor-board --out ./runs # calibration review UI (needs [ui] extra)
.venv/bin/rig publish packs/demo-tumor-board --runs ./runs --run <run_id> --format ts --out content.ts # site-ready content from pinned runs (see README)| Milestone | Status | Notes |
|---|---|---|
| M1 engine core | ✅ complete | pack loader, casebank, LLM-call SUT adapter (+ OpenAI-compatible, mock-tested), ablation+format axes, judge grading, append-only SQLite+parquet store, bootstrap stats, InputContract + ValidationReport, demo-tumor-board demo, end-to-end + determinism + zero-leak gates |
| M2 tumor-board tooling | 🟡 authoring + ingest boundary | ✅ DE language axis + battery axis-scoping; adjudication ingestion; rig lint; rig new scaffold; and now the pseudonymization boundary (validrig/ingest/, [deid] extra) — a Pseudonymizer contract + a lightweight Presidio backend (small spaCy model + regex patterns, no HF/transformers) with reversible encrypt/deanonymize (AES key from env, hospital-side, never stored). ❌ next (measured, if needed): higher-recall domain NER (OpenMed); PDF/OCR ingestion; judge-vs-gold metric. Lightweight recall on clinical free-text is not yet validated — established the boundary, not adequacy |
| M3 regression discipline | 🟢 done (core) | ✅ battery pinning, RegressionDiff, acceptance gating, CLI diff, native G-Eval LLMJudge (Gap 1), and now the judge-calibration loop (Gap 2): deterministic sampling, append-only human grades, Cohen's κ + % agreement, and a standalone calibration gate (advisory when underpowered). Surfaced in the review UI. Gate is kept out of the synchronous run/report path by design (calibration is asynchronous) |
| Review UI (M2/M3) | 🟢 both jobs | ✅ FastAPI + Jinja2, behind the [ui] extra, localhost-bound, never mutates immutable grades. Calibration: grade sampled generations, agreement/κ + gate. Adjudication: blind per-case gold authoring, writes pack rubric/adjudication/*.json (now consumed by the loader). Data path tested via TestClient + live uvicorn smoke. Clinician-oriented improvements (2026-09-04): grading instructions visible, ground-truth evidence display, case elements with labels, pre-fill existing values, progress indicator, critical-item styling. ❌ deferred: clinical UX review with a real clinician, multi-rater, SSO |
| M4 agent SUTs | 🟢 done | ✅ deterministic tool mocks (keyed by case/tool/args-hash, pack content → pack_hash), a fake agent (kind=agent) emitting a Trace, process rubrics (RubricItem.target=trace, N/A for non-agent SUTs), proven by right-answer-wrong-process. Plus the agent perturbation axes: tool_availability (remove a tool) and tool_response (error/empty) threaded to the agent via SUTContext; the agent_robustness battery shows the tool removed/degraded → process fails while output survives (agent compensates from the note, doesn't hallucinate). ❌ remaining: wrapping a real (non-fake) agent framework via the trace protocol |
| M5 monitoring | 🟢 done (core) | ✅ ingest pseudonymized production events (element-presence booleans + override flag, extra="forbid"), MonitoringSnapshot (override rate + three-state completeness vs the validated contract), and drift on two separate baselines — absolute (vs thresholds) and trend (vs prior snapshot), degradation-only, advisory when underpowered. rig monitor emits snapshot + drift + PMS + AIMS. ❌ deferred: production-side capture/pseudonymization of real encounters (hospital integration, not offline-verifiable) |
| QMS integration (§6) | 🟢 complete | ✅ maps runs → r05 (QMS-2026-07-09-R005): V&V plan, V&V report (baseline verdict; perturbations as characterization; calibration gate folded in async), change request, calibration status, package manifest, and now PMS periodic report (from a snapshot) + AIMS event (only on a real drift finding), plus a consolidated validation dossier rendered as self-contained printable HTML (rig dossier) with info-value bars + label-backed status (grayscale-safe). Attestation over pinned inputs; unsigned drafts; a prepared seam for immutable-release signing (docs/proposals/2026-07-17-release-attestation.md). Traceability map in docs/qms-traceability.md |
docs/proposals/2026-07-16-m3-regression-fixes.md— LLM judge ✅ shipped; calibration gate ✅ shippeddocs/proposals/2026-07-16-adjudication-calibration-ui.md— calibration review UI ✅ v1 shippeddocs/proposals/2026-07-16-m2-tooling-fixes.md— adjudication ingestion ✅ + rubric authoring CLI ✅ shipped; pseudonymization boundary + PDF/OCR not yet built
- Deterministic replay (
test_end_to_end.py,test_cli.py): two runs at the same seed produce byte-identical content (generations/grades/contract), in-process and across separate processes. Run id derives only from pins. - Zero use-case code in engine (
test_no_usecase_leak.py): the engine contains no use-case vocabulary. This gate already caught one real leak. - Contract distinguishes unknown from zero: an element ablated only in a
bundle (or never) is reported
measured: false, information_value: null, not0.0— load-bearing for RegressionDiff (a newly-dropped dependency vs. a never-measured one). - Acceptance gates on the baseline (intended-input) condition; ablation feeds the input contract and a separate robustness section, so a battery containing deliberate sabotage does not fail the baseline validation.
- Input contract (
smoke): pathology/molecular/imaging each carry information value ~0.33; prior_notes/meds are unmeasured (never isolated). - "What the new version broke" (
regression): a deliberately-regressed SUT that stopped reporting molecular findings. Both models still pass baseline acceptance (mean 0.67 > 0.60), but the diff pinpoints it: −0.27 mean score (significant), 36 item regressions, molecular_report information value −0.33. This is why acceptance thresholds alone are insufficient and the diff is the product. - Language sensitivity (
multilingual): in DE, diagnosis extraction fails (English-phrasing-dependent) while molecular/staging survive (language- invariant tokens).
- Pseudonymization boundary (§5): its safety-critical property is recall, which is unverifiable on synthetic data — needs a clinical decision, own turn.
- PDF/OCR ingestion (M2): needs external binaries and real documents.
- Review UI clinical UX review: the data path is tested; whether a clinician wants this flow is not.
- Real model endpoints: OpenAI-compatible SUT + LLM judge are built and mock-tested; no live call is made in tests (offline guarantee).
- Wrapping a real agent framework (M4) and production-side capture (M5): both need a real external system / clinical integration, not offline-verifiable.
- Plan:
docs/superpowers/plans/2026-07-16-harness-factory-m1.md - Engine:
validrig/—models/,packio/,perturb/,sut/,judge/,store/,stats/,artifacts/,calibration/,agent/,monitoring/,qms/,authoring/,ui/,execute.py,diff.py,cli.py,power.py,pathsafe.py - Demo packs:
packs/demo-tumor-board/(LLM),packs/demo-agent/(tool-using agent) - Tests:
tests/(one file per subsystem + integration gates) - Repo:
github.com/AliakseiT/validrig(private); commits authored as the AliakseiT noreply identity, no Claude attribution.