Skip to content

Repository files navigation

Release Equivalence & Recovery (A4)

Engineering profile: RESEARCH Toolchain: CPython 3.14 + uv Quality gate: make check Declared deviations: see "Contract deviation ledger" below — not none. An independent forensic review of the initial candidate (feb40de) found this line was false; it is corrected here rather than silently restated.

Purpose and boundary

A4 asks: when a validated model-system crosses a release boundary, under what conditions does it remain behaviorally equivalent to the system that was evaluated, and when can rollback or recovery actually restore the declared decision behavior? The full frozen scientific contract — release identity model, equivalence-level definitions and tolerances, the ten release-skew mechanisms, decision policies, rollback semantics, controls, metrics, negative-result targets, threats to validity, and prohibited claims — is docs/RESEARCH-CONTRACT.md. Nothing in this README overrides it; where the two disagree, the contract is authoritative.

This is a measurement study of release-boundary behavioral equivalence. It is not a deployment platform, an MLOps framework, a Kubernetes study, a model-registry comparison, or A5's delayed-feedback intervention question (contract section 2).

Status

Corrected candidate, builder-complete, following an independent forensic review of the initial candidate (feb40de, docs/reviews/2026-09-03T173215Z-feb40de-forensic-review.md, verdict REVISE). That review found two blocking defects: A4-001, the release-identity manifest computed CODE/SCHEMA/PREPROCESSING_LOGIC from a free-form label string that several mechanisms did not update to match their real behavior, so the live manifest diff did not match the frozen contract's mechanism table for 12 of 18 Arm A cells, and no test caught it; and A4-002, Arm C's reported latency table and RUN_MANIFEST_VALID claim did not describe files a reader could actually regenerate (a second data fetch's volatile timestamp had invalidated every manifest, and latency was measured on the wrong batch size). Both are corrected in this candidate — see "Contract deviation ledger" below and docs/RESEARCH-CONTRACT.md's dated amendment for the exact before/after.

make check passes (lock, format, lint, type, 99 tests including property-based tests and a comprehensive per-mechanism manifest-diff regression test, bounded reproduction, manifest validation). The canonical Arm A mechanism sweep (3,600 replications), Arm A boundary-sensitive sweep (1,000 replications), Arm B rollback grid (600 replications) and Arm C frontier were all re-run against the real, fetched German Credit data after the corrections; every number in docs/EVIDENCE.md traces to results/*.csv via scripts/summarize_evidence.py or scripts/make_figures.py, and all four canonical run manifests are now RUN_MANIFEST_VALID (previously invalid against committed provenance).

Not yet done, and out of this candidate's scope: independent confirmation of this corrected candidate (routed externally per the review-handoff protocol, not performed by this repository's builder), publication, A5, and anything past a local, unpushed git history (this study has no configured remote, matching sibling studies A1/A2/A3).

Architecture

src/release_equivalence/
  config.py             enums + frozen configs: PerturbationType, Magnitude, ReleaseComponent,
                         RollbackStage, PerturbationSpec, DecisionConfig, Tolerances
  data.py                German Credit fetch/load/split/bootstrap-sample + verify_source_identity
  preprocessing.py        numeric standardization + one-hot encoding; two variance algorithms
                          (two_pass / welford_online); name- vs position-based column reading
  model.py                logistic regression fit + safe-JSON serialization (no pickle)
  release.py              the eight-component ReleaseManifest, computed structurally, + diff_manifests
  perturbation.py         the single dispatch point for all ten mechanisms + score() + manifest_for()
  decision.py             threshold and exact-budget top-K decision policies (contract section 7)
  equivalence.py           the L0-L4 comparison harness (contract section 5)
  rollback.py              recovery_report: mechanically-computed rollback completeness
  metrics.py               Wilson intervals, Spearman rank correlation, Jaccard overlap
  experiment.py            Arm A: one replication -> one result row; boundary-sensitive sampling
  grid.py                  Arm A: JSON grid spec -> replications -> rows
  rollback_experiment.py   Arm B: the bad-release + three-stage rollback grid
  operational.py           Arm C: Configuration LR vs. GB latency/size/quality measurement
  manifest.py              MyWorld run-manifest v1 builder (shared by every script)
  errors.py                ReleaseIncompatibilityError: explicit failure, never a silent fallback
  bounded_fixture.py       the bounded, network-free, synthetic-shaped smoke path
scripts/
  reproduce.py                  bounded, deterministic, network-free — part of `make check`
  fetch_data.py                 downloads credit-g from OpenML, records provenance
  run_experiment.py             canonical Arm A (main-grid.json or boundary-grid.json)
  run_rollback_experiment.py    canonical Arm B
  run_frontier.py                canonical Arm C
  summarize_evidence.py          prints every docs/EVIDENCE.md table from results/*.csv
  make_figures.py                 regenerates figures/ from results/ only
  validate_manifest.py            validates a MyWorld run-manifest v1

Commands

Command What it does
make setup Exact sync from uv.lock.
make check Lock, format, lint, type, test, bounded reproduction and manifest gates.
make test Unit + property-based tests (99 tests).
make reproduce Bounded, deterministic, network-free evidence path.
make data Fetches credit-g from OpenML and records provenance. Requires network.
make experiment Canonical Arm A (configs/main-grid.json + configs/boundary-grid.json). Requires make data first. Not part of make check.
make rollback Canonical Arm B (configs/rollback-grid.json). Requires make data first.
make frontier Canonical Arm C (configs/frontier.json). Requires make data first.
make figures Regenerates all figures from canonical results/ only.
make evidence Prints every docs/EVIDENCE.md table from results/*.csv.
make clean Removes generated bounded output and local caches only.

Data provenance

German Credit (credit-g, OpenML dataset id 31): 1,000 applicants, 20 raw features (7 numeric, 13 categorical), binary good/bad credit-risk label. Fetched by scripts/fetch_data.py from https://www.openml.org/data/get_csv/31/dataset_31_credit-g.arff. The raw CSV is gitignored (reproducible from source); data/provenance/ commits the source URL, retrieval timestamp and SHA-256 digest instead — see data/README.md. 700 rows are held out as the model-fitting split (produces the reference MODEL_ARTIFACT/PREPROCESSING_STATE exactly once); the remaining 300 rows are the evaluation pool every experiment bootstrap-resamples from — never used for fitting.

Methodology summary

A release is represented as an explicit eight-component manifest (CODE, MODEL_ARTIFACT, PREPROCESSING_LOGIC, PREPROCESSING_STATE, SCHEMA, DECISION_CONFIG, DEPENDENCY_TOOLCHAIN, INPUT_DATA — contract section 4, ADR-001), computed structurally from each release's real behavioral fields rather than a free-form label. Five equivalence levels (L0 artifact, L1 numerical, L2 predictive, L3 decision, L4 operational — contract section 5) are computed by one comparison harness; L0-L3 are computed for every Arm A/B replication, and L4 is computed only when operational measurements (latency/size ratios) are supplied — which happens at the configuration level in Arm C, not per-replication in Arm A/B. Ten typed release-skew mechanisms (contract section 6) perturb one or more named components at a time, including a deliberately "silent-failure" pair (SCHEMA_DEFAULT_FALLBACK, FEATURE_ORDER_SKEW) and a "config-only" pair (CONFIG_THRESHOLD_DRIFT, and the frozen-model side of RETRAIN_RELEASE) that changes decisions without changing the model file. Two bounded decision policies (threshold approval, finite-capacity top-K — contract section 7) translate scores into consequential outcomes. Arm B constructs a "bad release" that changes three components at once and measures whether no / partial (model-only) / full rollback actually restores decision equivalence. Arm C measures a real quality/latency/cost frontier between a logistic and a gradient-boosted release configuration on this host.

Evidence and headline findings

See docs/EVIDENCE.md for the full run-by-run account (canonical commands, manifest digests, and every headline number traced to its results/*.csv row). Figures in figures/ are regenerated from that evidence only (make figures); nothing below is manually transcribed. docs/hypothesis-disposition.json binds each of contract section 16's five claims to the exact results/*.csv cell IDs and figure that support it.

Reproduction

Bounded (make check / make reproduce): deterministic, network-free, a fraction of a second, exercises one small mechanism/rollback grid on a synthetic-shaped fixture batch (bounded_fixture.py) — proves the reference -> perturbation -> equivalence -> recovery -> evidence path works, not a scientific result.

Canonical: make data fetches the real German Credit CSV and records provenance; make experiment runs the full contract-section-9 Arm A grid (3,600 + 1,000 replications, a few seconds of wall time); make rollback runs Arm B (600 replications); make frontier runs Arm C. None of the three is part of make check or CI (network-required for the first, and not appropriate for every commit).

Evidence lineage

Every canonical run writes a MyWorld run-manifest v1 alongside its CSV (manifest.py / scripts/validate_manifest.py), recording config identity + SHA-256, the git SHA, toolchain versions, seeds, duration and output digests. The manifest's inputs list binds to the fetched raw CSV's own content hash directly (data.verify_source_identity checks this against the committed provenance record before any canonical script runs) — not to the whole provenance JSON file, whose retrieved_at_utc transport-metadata field changes on every make data call independent of the dataset's actual content. This separation (canonical content identity vs. transport metadata) is a corrected-candidate fix: forensic review A4-002 found a second fetch silently invalidated every existing manifest under the previous design.

Contract deviation ledger

Every remaining departure from the frozen docs/RESEARCH-CONTRACT.md, as of this corrected candidate. This replaces the initial candidate's incorrect "Declared deviations: none" (forensic review finding).

Frozen requirement Actual implementation Reason Scientific consequence Claims narrowed?
Contract §6: FLOAT_PRECISION_SKEW changes PREPROCESSING_LOGIC, MODEL_ARTIFACT Changes PREPROCESSING_LOGIC only A serving-time precision cast never rewrites the stored model bytes; there is no second artifact to hash None — corrected by dated contract amendment (2026-09-03), not by implementation change No
Contract §6: DEPENDENCY_NUMERIC_VARIANT changes DEPENDENCY_TOOLCHAIN, PREPROCESSING_LOGIC Changes PREPROCESSING_LOGIC, PREPROCESSING_STATE ADR-002 already declared this mechanism never varies the live dependency stack; the original table row contradicted its own ADR None — corrected by dated contract amendment, not by implementation change No
Contract §6: CATEGORICAL_ENCODING_SKEW changes PREPROCESSING_LOGIC, PREPROCESSING_STATE Changes PREPROCESSING_STATE only Rotating a fitted vocabulary's order is a state change, not an algorithm change None — corrected by dated contract amendment, not by implementation change No
Contract §4: SCHEMA identity covers "ordered column names + dtypes" Hashes column order + which declared columns are absent; does not hash dtypes No mechanism in this study perturbs a dtype independently of presence/order None known; disclosed rather than silently claimed as full compliance No
Contract §4: DEPENDENCY_TOOLCHAIN is a release identity component Recorded as interpreter/library version strings; does not include a uv.lock digest No mechanism varies this component at all (see above); narrow fix scope for this correction cycle None — nothing in this study depends on lockfile-level granularity No
Contract §9: Arm C latency uses "a fixed batch size (100)" Fixed in this candidate. Was: measured on the full 300-row evaluation split. Now: a fixed, deterministically-resampled 100-row batch, matching the contract Original implementation error (forensic review A4-002), now corrected Latency ratio changed from ~55-58× (uncorrected, batch-300, superseded) to 44.5× and 39.9× across two corrected (batch-100) process invocations — 39.9× is canonical (results/arm_c_frontier.csv) Yes — reported as "order-of-magnitude ~40×," not a fixed constant, given observed run-to-run variance
Contract's implicit expectation that a frozen research contract predates its results in git history The candidate commit (4c5c0f9) contains the frozen contract, full implementation, tests and evidence together; no independent pre-results commit exists Single-session build; no scaffolding-only freeze commit was made before implementation began Pre-registration of hypotheses/predictions (e.g. Arm C's directional prediction) is a builder-process claim, not independently git-verifiable Yes — stated explicitly in docs/EVIDENCE.md
Arm B "full rollback...restores behavior" (contract §9) The FULL stage substitutes the original reference ReleaseInstance object directly, not a reconstruction from declared component values Simplest correct control; contract §9 itself already called this "definitionally the reference release again" 200/200 full-rollback recovery is an internal-consistency control, not independent empirical evidence; the partial-rollback 0/200 result carries the actual recovery finding Yes — stated explicitly in docs/EVIDENCE.md and rollback_experiment.py

If a future correction closes one of the disclosed-but-not-required rows above (dtype hashing, uv.lock digest, a scaffolding-only pre-registration commit for a future study), update this ledger rather than deleting the row.

Assumptions and threats to validity

See docs/RESEARCH-CONTRACT.md section 12 for the full, frozen account: DEPENDENCY_NUMERIC_VARIANT is a modeled proxy for library-version numeric skew, not a literal two-installed-version test (ADR-002); findings are scoped to one dataset (German Credit, a 1994-era benchmark — not a contemporary lending-policy claim) and one model family per arm; latency and size measurements are single-host (macOS, Apple Silicon), single process invocation, and are an operational-contract-level comparison, never a low-level CPU/cache/PMU claim; magnitude levels are author-chosen for interpretable separation, not sampled from a real severity distribution; boundary-sensitive sampling ratios are, for the CONFIG_THRESHOLD_DRIFT mechanism specifically, largely a property of the sampling design itself rather than an independent empirical discovery (see docs/EVIDENCE.md).

Limitations and non-claims

A4 does not claim a production release-safety certification for any real system, that its ten mechanisms are exhaustive of real release-skew causes, that the dependency-numeric-variant proxy is equivalent to a real library version-pair test, a causal account of any real production incident, any low-level CPU/cache/PMU performance claim, or anything about contemporary consumer credit risk or lending fairness (contract section 16; German Credit is a 1994-era benchmark used only for its realistic tabular structure). Full rollback's 200/200 recovery result is an internal- consistency control (it substitutes the reference release object directly), not independent evidence that rollback restores behavior in general — the partial-rollback 0/200 result is the finding that carries recovery-question weight. Pre-registration of Arm C's directional prediction is a documented builder-process claim, not independently verifiable from this repository's git history. Where evidence shows a safeguard failing — an identical model artifact with divergent decisions, a numerically-untouched batch whose decisions still flip under config drift, a partial rollback that never restores decision equivalence — that is reported as a finding, not smoothed over (contract section 11).

License and data rights

Repository code license. The code, tests, configuration and documentation in this repository are licensed under the MIT License (copyright Joshua de Freitas). This license covers everything this study authored: src/, scripts/, tests/, configs/, and the documentation in docs//README.md.

Third-party dataset rights — distinct from the code license. credit-g (German Credit / Statlog, OpenML dataset id 31) is fetched from OpenML at run time (scripts/fetch_data.py, source https://www.openml.org/data/get_csv/31/dataset_31_credit-g.arff) and is not covered by the MIT license above. Per OpenML's own dataset record (https://www.openml.org/api/v1/json/data/31, inspected 2026-09-03): author Dr. Hans Hofmann, original source UCI Machine Learning Repository (1994), OpenML's own licence field is "Public", and OpenML/UCI ask that reuse cite the UCI citation policy (https://archive.ics.uci.edu/ml/citation_policy.html). This repository does not restate, relicense, or modify those third-party terms; it only records what it found and defers to the source for anything this note does not cover.

The raw dataset is deliberately not committed to this repository. data/raw/ is gitignored; only the SHA-256/URL/retrieval-timestamp provenance record (data/provenance/credit-g.provenance.json) is tracked in git. This is a deliberate design choice, preserved through the A4 correction cycle: redistribution rights for republishing the raw dataset file itself were not independently confirmed, so this repository fetches it fresh from the authoritative source at reproduction time (make data) rather than bundling a copy. Reproduction of the canonical arms requires that fetch to succeed.

About

Release identity, behavioral equivalence, and rollback recovery across model-system boundaries.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages