Engineering profile: RESEARCH
Toolchain: CPython 3.14 + uv
Quality gate: make check
Declared deviations: see "Contract deviation ledger" below — not none.
An independent forensic review of the initial candidate (feb40de) found
this line was false; it is corrected here rather than silently restated.
A4 asks: when a validated model-system crosses a release boundary, under
what conditions does it remain behaviorally equivalent to the system that
was evaluated, and when can rollback or recovery actually restore the
declared decision behavior? The full frozen scientific contract — release
identity model, equivalence-level definitions and tolerances, the ten
release-skew mechanisms, decision policies, rollback semantics, controls,
metrics, negative-result targets, threats to validity, and prohibited
claims — is docs/RESEARCH-CONTRACT.md. Nothing in this README overrides
it; where the two disagree, the contract is authoritative.
This is a measurement study of release-boundary behavioral equivalence. It is not a deployment platform, an MLOps framework, a Kubernetes study, a model-registry comparison, or A5's delayed-feedback intervention question (contract section 2).
Corrected candidate, builder-complete, following an independent
forensic review of the initial candidate (feb40de,
docs/reviews/2026-09-03T173215Z-feb40de-forensic-review.md, verdict
REVISE). That review found two blocking defects: A4-001, the
release-identity manifest computed CODE/SCHEMA/PREPROCESSING_LOGIC
from a free-form label string that several mechanisms did not update to
match their real behavior, so the live manifest diff did not match the
frozen contract's mechanism table for 12 of 18 Arm A cells, and no test
caught it; and A4-002, Arm C's reported latency table and
RUN_MANIFEST_VALID claim did not describe files a reader could actually
regenerate (a second data fetch's volatile timestamp had invalidated every
manifest, and latency was measured on the wrong batch size). Both are
corrected in this candidate — see "Contract deviation ledger" below and
docs/RESEARCH-CONTRACT.md's dated amendment for the exact before/after.
make check passes (lock, format, lint, type, 99 tests including
property-based tests and a comprehensive per-mechanism manifest-diff
regression test, bounded reproduction, manifest validation). The canonical
Arm A mechanism sweep (3,600 replications), Arm A boundary-sensitive sweep
(1,000 replications), Arm B rollback grid (600 replications) and Arm C
frontier were all re-run against the real, fetched German Credit data after
the corrections; every number in docs/EVIDENCE.md traces to
results/*.csv via scripts/summarize_evidence.py or
scripts/make_figures.py, and all four canonical run manifests are now
RUN_MANIFEST_VALID (previously invalid against committed provenance).
Not yet done, and out of this candidate's scope: independent confirmation of this corrected candidate (routed externally per the review-handoff protocol, not performed by this repository's builder), publication, A5, and anything past a local, unpushed git history (this study has no configured remote, matching sibling studies A1/A2/A3).
src/release_equivalence/
config.py enums + frozen configs: PerturbationType, Magnitude, ReleaseComponent,
RollbackStage, PerturbationSpec, DecisionConfig, Tolerances
data.py German Credit fetch/load/split/bootstrap-sample + verify_source_identity
preprocessing.py numeric standardization + one-hot encoding; two variance algorithms
(two_pass / welford_online); name- vs position-based column reading
model.py logistic regression fit + safe-JSON serialization (no pickle)
release.py the eight-component ReleaseManifest, computed structurally, + diff_manifests
perturbation.py the single dispatch point for all ten mechanisms + score() + manifest_for()
decision.py threshold and exact-budget top-K decision policies (contract section 7)
equivalence.py the L0-L4 comparison harness (contract section 5)
rollback.py recovery_report: mechanically-computed rollback completeness
metrics.py Wilson intervals, Spearman rank correlation, Jaccard overlap
experiment.py Arm A: one replication -> one result row; boundary-sensitive sampling
grid.py Arm A: JSON grid spec -> replications -> rows
rollback_experiment.py Arm B: the bad-release + three-stage rollback grid
operational.py Arm C: Configuration LR vs. GB latency/size/quality measurement
manifest.py MyWorld run-manifest v1 builder (shared by every script)
errors.py ReleaseIncompatibilityError: explicit failure, never a silent fallback
bounded_fixture.py the bounded, network-free, synthetic-shaped smoke path
scripts/
reproduce.py bounded, deterministic, network-free — part of `make check`
fetch_data.py downloads credit-g from OpenML, records provenance
run_experiment.py canonical Arm A (main-grid.json or boundary-grid.json)
run_rollback_experiment.py canonical Arm B
run_frontier.py canonical Arm C
summarize_evidence.py prints every docs/EVIDENCE.md table from results/*.csv
make_figures.py regenerates figures/ from results/ only
validate_manifest.py validates a MyWorld run-manifest v1
| Command | What it does |
|---|---|
make setup |
Exact sync from uv.lock. |
make check |
Lock, format, lint, type, test, bounded reproduction and manifest gates. |
make test |
Unit + property-based tests (99 tests). |
make reproduce |
Bounded, deterministic, network-free evidence path. |
make data |
Fetches credit-g from OpenML and records provenance. Requires network. |
make experiment |
Canonical Arm A (configs/main-grid.json + configs/boundary-grid.json). Requires make data first. Not part of make check. |
make rollback |
Canonical Arm B (configs/rollback-grid.json). Requires make data first. |
make frontier |
Canonical Arm C (configs/frontier.json). Requires make data first. |
make figures |
Regenerates all figures from canonical results/ only. |
make evidence |
Prints every docs/EVIDENCE.md table from results/*.csv. |
make clean |
Removes generated bounded output and local caches only. |
German Credit (credit-g, OpenML dataset id 31): 1,000 applicants, 20 raw
features (7 numeric, 13 categorical), binary good/bad credit-risk label.
Fetched by scripts/fetch_data.py from
https://www.openml.org/data/get_csv/31/dataset_31_credit-g.arff. The raw
CSV is gitignored (reproducible from source); data/provenance/ commits the
source URL, retrieval timestamp and SHA-256 digest instead — see
data/README.md. 700 rows are held out as the model-fitting split (produces
the reference MODEL_ARTIFACT/PREPROCESSING_STATE exactly once); the
remaining 300 rows are the evaluation pool every experiment bootstrap-resamples
from — never used for fitting.
A release is represented as an explicit eight-component manifest (CODE,
MODEL_ARTIFACT, PREPROCESSING_LOGIC, PREPROCESSING_STATE, SCHEMA,
DECISION_CONFIG, DEPENDENCY_TOOLCHAIN, INPUT_DATA — contract section
4, ADR-001), computed structurally from each release's real behavioral
fields rather than a free-form label. Five equivalence levels (L0 artifact,
L1 numerical, L2 predictive, L3 decision, L4 operational — contract section
5) are computed by one comparison harness; L0-L3 are computed for every Arm
A/B replication, and L4 is computed only when operational measurements
(latency/size ratios) are supplied — which happens at the configuration
level in Arm C, not per-replication in Arm A/B. Ten typed release-skew mechanisms
(contract section 6) perturb one or more named components at a time,
including a deliberately "silent-failure" pair (SCHEMA_DEFAULT_FALLBACK,
FEATURE_ORDER_SKEW) and a "config-only" pair
(CONFIG_THRESHOLD_DRIFT, and the frozen-model side of RETRAIN_RELEASE)
that changes decisions without changing the model file. Two bounded decision
policies (threshold approval, finite-capacity top-K — contract section 7)
translate scores into consequential outcomes. Arm B constructs a "bad
release" that changes three components at once and measures whether no /
partial (model-only) / full rollback actually restores decision equivalence.
Arm C measures a real quality/latency/cost frontier between a logistic and a
gradient-boosted release configuration on this host.
See docs/EVIDENCE.md for the full run-by-run account (canonical commands,
manifest digests, and every headline number traced to its results/*.csv
row). Figures in figures/ are regenerated from that evidence only
(make figures); nothing below is manually transcribed.
docs/hypothesis-disposition.json binds each of contract section 16's five
claims to the exact results/*.csv cell IDs and figure that support it.
Bounded (make check / make reproduce): deterministic, network-free, a
fraction of a second, exercises one small mechanism/rollback grid on a
synthetic-shaped fixture batch (bounded_fixture.py) — proves the reference
-> perturbation -> equivalence -> recovery -> evidence path works, not a
scientific result.
Canonical: make data fetches the real German Credit CSV and records
provenance; make experiment runs the full contract-section-9 Arm A grid
(3,600 + 1,000 replications, a few seconds of wall time); make rollback
runs Arm B (600 replications); make frontier runs Arm C. None of the three
is part of make check or CI (network-required for the first, and not
appropriate for every commit).
Every canonical run writes a MyWorld run-manifest v1 alongside its CSV
(manifest.py / scripts/validate_manifest.py), recording config identity +
SHA-256, the git SHA, toolchain versions, seeds, duration and output
digests. The manifest's inputs list binds to the fetched raw CSV's own
content hash directly (data.verify_source_identity checks this against
the committed provenance record before any canonical script runs) —
not to the whole provenance JSON file, whose retrieved_at_utc
transport-metadata field changes on every make data call independent of
the dataset's actual content. This separation (canonical content identity
vs. transport metadata) is a corrected-candidate fix: forensic review
A4-002 found a second fetch silently invalidated every existing manifest
under the previous design.
Every remaining departure from the frozen docs/RESEARCH-CONTRACT.md, as of
this corrected candidate. This replaces the initial candidate's incorrect
"Declared deviations: none" (forensic review finding).
| Frozen requirement | Actual implementation | Reason | Scientific consequence | Claims narrowed? |
|---|---|---|---|---|
Contract §6: FLOAT_PRECISION_SKEW changes PREPROCESSING_LOGIC, MODEL_ARTIFACT |
Changes PREPROCESSING_LOGIC only |
A serving-time precision cast never rewrites the stored model bytes; there is no second artifact to hash | None — corrected by dated contract amendment (2026-09-03), not by implementation change | No |
Contract §6: DEPENDENCY_NUMERIC_VARIANT changes DEPENDENCY_TOOLCHAIN, PREPROCESSING_LOGIC |
Changes PREPROCESSING_LOGIC, PREPROCESSING_STATE |
ADR-002 already declared this mechanism never varies the live dependency stack; the original table row contradicted its own ADR | None — corrected by dated contract amendment, not by implementation change | No |
Contract §6: CATEGORICAL_ENCODING_SKEW changes PREPROCESSING_LOGIC, PREPROCESSING_STATE |
Changes PREPROCESSING_STATE only |
Rotating a fitted vocabulary's order is a state change, not an algorithm change | None — corrected by dated contract amendment, not by implementation change | No |
Contract §4: SCHEMA identity covers "ordered column names + dtypes" |
Hashes column order + which declared columns are absent; does not hash dtypes | No mechanism in this study perturbs a dtype independently of presence/order | None known; disclosed rather than silently claimed as full compliance | No |
Contract §4: DEPENDENCY_TOOLCHAIN is a release identity component |
Recorded as interpreter/library version strings; does not include a uv.lock digest |
No mechanism varies this component at all (see above); narrow fix scope for this correction cycle | None — nothing in this study depends on lockfile-level granularity | No |
| Contract §9: Arm C latency uses "a fixed batch size (100)" | Fixed in this candidate. Was: measured on the full 300-row evaluation split. Now: a fixed, deterministically-resampled 100-row batch, matching the contract | Original implementation error (forensic review A4-002), now corrected | Latency ratio changed from ~55-58× (uncorrected, batch-300, superseded) to 44.5× and 39.9× across two corrected (batch-100) process invocations — 39.9× is canonical (results/arm_c_frontier.csv) |
Yes — reported as "order-of-magnitude ~40×," not a fixed constant, given observed run-to-run variance |
| Contract's implicit expectation that a frozen research contract predates its results in git history | The candidate commit (4c5c0f9) contains the frozen contract, full implementation, tests and evidence together; no independent pre-results commit exists |
Single-session build; no scaffolding-only freeze commit was made before implementation began | Pre-registration of hypotheses/predictions (e.g. Arm C's directional prediction) is a builder-process claim, not independently git-verifiable | Yes — stated explicitly in docs/EVIDENCE.md |
| Arm B "full rollback...restores behavior" (contract §9) | The FULL stage substitutes the original reference ReleaseInstance object directly, not a reconstruction from declared component values |
Simplest correct control; contract §9 itself already called this "definitionally the reference release again" | 200/200 full-rollback recovery is an internal-consistency control, not independent empirical evidence; the partial-rollback 0/200 result carries the actual recovery finding | Yes — stated explicitly in docs/EVIDENCE.md and rollback_experiment.py |
If a future correction closes one of the disclosed-but-not-required rows
above (dtype hashing, uv.lock digest, a scaffolding-only pre-registration
commit for a future study), update this ledger rather than deleting the row.
See docs/RESEARCH-CONTRACT.md section 12 for the full, frozen account:
DEPENDENCY_NUMERIC_VARIANT is a modeled proxy for library-version numeric
skew, not a literal two-installed-version test (ADR-002); findings are
scoped to one dataset (German Credit, a 1994-era benchmark — not a
contemporary lending-policy claim) and one model family per arm; latency
and size measurements are single-host (macOS, Apple Silicon), single
process invocation, and are an operational-contract-level comparison, never
a low-level CPU/cache/PMU claim; magnitude levels are author-chosen for
interpretable separation, not sampled from a real severity distribution;
boundary-sensitive sampling ratios are, for the CONFIG_THRESHOLD_DRIFT
mechanism specifically, largely a property of the sampling design itself
rather than an independent empirical discovery (see docs/EVIDENCE.md).
A4 does not claim a production release-safety certification for any real system, that its ten mechanisms are exhaustive of real release-skew causes, that the dependency-numeric-variant proxy is equivalent to a real library version-pair test, a causal account of any real production incident, any low-level CPU/cache/PMU performance claim, or anything about contemporary consumer credit risk or lending fairness (contract section 16; German Credit is a 1994-era benchmark used only for its realistic tabular structure). Full rollback's 200/200 recovery result is an internal- consistency control (it substitutes the reference release object directly), not independent evidence that rollback restores behavior in general — the partial-rollback 0/200 result is the finding that carries recovery-question weight. Pre-registration of Arm C's directional prediction is a documented builder-process claim, not independently verifiable from this repository's git history. Where evidence shows a safeguard failing — an identical model artifact with divergent decisions, a numerically-untouched batch whose decisions still flip under config drift, a partial rollback that never restores decision equivalence — that is reported as a finding, not smoothed over (contract section 11).
Repository code license. The code, tests, configuration and
documentation in this repository are licensed under the
MIT License (copyright Joshua de Freitas). This license covers
everything this study authored: src/, scripts/, tests/, configs/,
and the documentation in docs//README.md.
Third-party dataset rights — distinct from the code license.
credit-g (German Credit / Statlog, OpenML dataset id 31) is fetched from
OpenML at run time (scripts/fetch_data.py, source
https://www.openml.org/data/get_csv/31/dataset_31_credit-g.arff) and is
not covered by the MIT license above. Per OpenML's own dataset record
(https://www.openml.org/api/v1/json/data/31, inspected 2026-09-03): author
Dr. Hans Hofmann, original source UCI Machine Learning Repository (1994),
OpenML's own licence field is "Public", and OpenML/UCI ask that reuse
cite the UCI citation policy
(https://archive.ics.uci.edu/ml/citation_policy.html). This repository
does not restate, relicense, or modify those third-party terms; it only
records what it found and defers to the source for anything this note does
not cover.
The raw dataset is deliberately not committed to this repository.
data/raw/ is gitignored; only the SHA-256/URL/retrieval-timestamp
provenance record (data/provenance/credit-g.provenance.json) is tracked
in git. This is a deliberate design choice, preserved through the A4
correction cycle: redistribution rights for republishing the raw dataset
file itself were not independently confirmed, so this repository fetches it
fresh from the authoritative source at reproduction time
(make data) rather than bundling a copy. Reproduction of the canonical
arms requires that fetch to succeed.