Skip to content

Latest commit

 

History

History
97 lines (80 loc) · 8.81 KB

File metadata and controls

97 lines (80 loc) · 8.81 KB

validrig — Build Status

validrig is the engine of DearAuditor Eval (CLI: rig); formerly "Harness Factory".

Last updated: 2026-09-04

What runs today

The engine (validrig/) is use-case-agnostic; all use-case content lives in packs/. Everything below runs fully offline and deterministically via a first-class fake model and fake judge.

python3.12 -m venv .venv && .venv/bin/pip install -e ".[dev]"
.venv/bin/pytest -q                                   # 225 tests, ~6s
.venv/bin/rig new packs/my-pack --id my-pack      # scaffold a new pack
.venv/bin/rig lint packs/demo-tumor-board         # check a pack for authoring gaps
.venv/bin/rig run packs/demo-tumor-board --battery smoke --out ./runs --seed 1
.venv/bin/rig run packs/demo-tumor-board --battery regression --out ./runs --seed 1
.venv/bin/rig diff --out ./runs --baseline <run_a> --candidate <run_b>
.venv/bin/rig qms packs/demo-tumor-board --run <run_id> --out ./runs
.venv/bin/rig monitor packs/demo-tumor-board --run <run_id> --events events.jsonl --out ./runs  # M5 monitoring
.venv/bin/rig dossier packs/demo-tumor-board --run <run_id> --out ./runs   # printable HTML validation dossier
.venv/bin/rig ui packs/demo-tumor-board --out ./runs   # calibration review UI (needs [ui] extra)
.venv/bin/rig publish packs/demo-tumor-board --runs ./runs --run <run_id> --format ts --out content.ts  # site-ready content from pinned runs (see README)

Milestone status

Milestone Status Notes
M1 engine core ✅ complete pack loader, casebank, LLM-call SUT adapter (+ OpenAI-compatible, mock-tested), ablation+format axes, judge grading, append-only SQLite+parquet store, bootstrap stats, InputContract + ValidationReport, demo-tumor-board demo, end-to-end + determinism + zero-leak gates
M2 tumor-board tooling 🟡 authoring + ingest boundary ✅ DE language axis + battery axis-scoping; adjudication ingestion; rig lint; rig new scaffold; and now the pseudonymization boundary (validrig/ingest/, [deid] extra) — a Pseudonymizer contract + a lightweight Presidio backend (small spaCy model + regex patterns, no HF/transformers) with reversible encrypt/deanonymize (AES key from env, hospital-side, never stored). ❌ next (measured, if needed): higher-recall domain NER (OpenMed); PDF/OCR ingestion; judge-vs-gold metric. Lightweight recall on clinical free-text is not yet validated — established the boundary, not adequacy
M3 regression discipline 🟢 done (core) ✅ battery pinning, RegressionDiff, acceptance gating, CLI diff, native G-Eval LLMJudge (Gap 1), and now the judge-calibration loop (Gap 2): deterministic sampling, append-only human grades, Cohen's κ + % agreement, and a standalone calibration gate (advisory when underpowered). Surfaced in the review UI. Gate is kept out of the synchronous run/report path by design (calibration is asynchronous)
Review UI (M2/M3) 🟢 both jobs ✅ FastAPI + Jinja2, behind the [ui] extra, localhost-bound, never mutates immutable grades. Calibration: grade sampled generations, agreement/κ + gate. Adjudication: blind per-case gold authoring, writes pack rubric/adjudication/*.json (now consumed by the loader). Data path tested via TestClient + live uvicorn smoke. Clinician-oriented improvements (2026-09-04): grading instructions visible, ground-truth evidence display, case elements with labels, pre-fill existing values, progress indicator, critical-item styling. ❌ deferred: clinical UX review with a real clinician, multi-rater, SSO
M4 agent SUTs 🟢 done ✅ deterministic tool mocks (keyed by case/tool/args-hash, pack content → pack_hash), a fake agent (kind=agent) emitting a Trace, process rubrics (RubricItem.target=trace, N/A for non-agent SUTs), proven by right-answer-wrong-process. Plus the agent perturbation axes: tool_availability (remove a tool) and tool_response (error/empty) threaded to the agent via SUTContext; the agent_robustness battery shows the tool removed/degraded → process fails while output survives (agent compensates from the note, doesn't hallucinate). ❌ remaining: wrapping a real (non-fake) agent framework via the trace protocol
M5 monitoring 🟢 done (core) ✅ ingest pseudonymized production events (element-presence booleans + override flag, extra="forbid"), MonitoringSnapshot (override rate + three-state completeness vs the validated contract), and drift on two separate baselines — absolute (vs thresholds) and trend (vs prior snapshot), degradation-only, advisory when underpowered. rig monitor emits snapshot + drift + PMS + AIMS. ❌ deferred: production-side capture/pseudonymization of real encounters (hospital integration, not offline-verifiable)
QMS integration (§6) 🟢 complete ✅ maps runs → r05 (QMS-2026-07-09-R005): V&V plan, V&V report (baseline verdict; perturbations as characterization; calibration gate folded in async), change request, calibration status, package manifest, and now PMS periodic report (from a snapshot) + AIMS event (only on a real drift finding), plus a consolidated validation dossier rendered as self-contained printable HTML (rig dossier) with info-value bars + label-backed status (grayscale-safe). Attestation over pinned inputs; unsigned drafts; a prepared seam for immutable-release signing (docs/proposals/2026-07-17-release-attestation.md). Traceability map in docs/qms-traceability.md

Proposals

  • docs/proposals/2026-07-16-m3-regression-fixes.md — LLM judge ✅ shipped; calibration gate ✅ shipped
  • docs/proposals/2026-07-16-adjudication-calibration-ui.md — calibration review UI ✅ v1 shipped
  • docs/proposals/2026-07-16-m2-tooling-fixes.md — adjudication ingestion ✅ + rubric authoring CLI ✅ shipped; pseudonymization boundary + PDF/OCR not yet built

Design guarantees proven by tests

  • Deterministic replay (test_end_to_end.py, test_cli.py): two runs at the same seed produce byte-identical content (generations/grades/contract), in-process and across separate processes. Run id derives only from pins.
  • Zero use-case code in engine (test_no_usecase_leak.py): the engine contains no use-case vocabulary. This gate already caught one real leak.
  • Contract distinguishes unknown from zero: an element ablated only in a bundle (or never) is reported measured: false, information_value: null, not 0.0 — load-bearing for RegressionDiff (a newly-dropped dependency vs. a never-measured one).
  • Acceptance gates on the baseline (intended-input) condition; ablation feeds the input contract and a separate robustness section, so a battery containing deliberate sabotage does not fail the baseline validation.

Key demonstrations

  • Input contract (smoke): pathology/molecular/imaging each carry information value ~0.33; prior_notes/meds are unmeasured (never isolated).
  • "What the new version broke" (regression): a deliberately-regressed SUT that stopped reporting molecular findings. Both models still pass baseline acceptance (mean 0.67 > 0.60), but the diff pinpoints it: −0.27 mean score (significant), 36 item regressions, molecular_report information value −0.33. This is why acceptance thresholds alone are insufficient and the diff is the product.
  • Language sensitivity (multilingual): in DE, diagnosis extraction fails (English-phrasing-dependent) while molecular/staging survive (language- invariant tokens).

Still deferred (need real data or a clinical decision)

  • Pseudonymization boundary (§5): its safety-critical property is recall, which is unverifiable on synthetic data — needs a clinical decision, own turn.
  • PDF/OCR ingestion (M2): needs external binaries and real documents.
  • Review UI clinical UX review: the data path is tested; whether a clinician wants this flow is not.
  • Real model endpoints: OpenAI-compatible SUT + LLM judge are built and mock-tested; no live call is made in tests (offline guarantee).
  • Wrapping a real agent framework (M4) and production-side capture (M5): both need a real external system / clinical integration, not offline-verifiable.

Where the code is

  • Plan: docs/superpowers/plans/2026-07-16-harness-factory-m1.md
  • Engine: validrig/ — models/, packio/, perturb/, sut/, judge/, store/, stats/, artifacts/, calibration/, agent/, monitoring/, qms/, authoring/, ui/, execute.py, diff.py, cli.py, power.py, pathsafe.py
  • Demo packs: packs/demo-tumor-board/ (LLM), packs/demo-agent/ (tool-using agent)
  • Tests: tests/ (one file per subsystem + integration gates)
  • Repo: github.com/AliakseiT/validrig (private); commits authored as the AliakseiT noreply identity, no Claude attribution.