Engineering profile: RESEARCH
Toolchain: CPython 3.14 + uv 0.12.5
Quality gate: make check
Declared deviations: none to the frozen contract's design; two UCI
data-handling clarifications not fully pinned by the contract's prose are
recorded below and in docs/references/README.md.
Contract: docs/RESEARCH-CONTRACT.md — read
this first. It is frozen; anything below implements it, and any deviation
discovered during implementation is recorded here, never silently made in
the contract itself.
Correction record: an independent forensic review
(docs/reviews/2026-09-02-6fd57f3-forensic-review.md,
verdict REVISE) found three blocking defects in the candidate this document
used to describe. All three are corrected in this version — see
"Correction record" below before trusting any number that predates it.
Under what conditions is ranking/prediction quality (discrimination) alone insufficient for a reliable finite-capacity decision, and how do calibration and uncertainty-aware decision rules change realized decision value relative to a transparent baseline and to plain score ranking? Second foundation of Track A (Model & Data Systems) in the Dual-Foundation Public Technical Evidence Programme. Self-contained: its own controlled DGP plus one independently sourced real dataset (UCI Bank Marketing); it consumes no A1 artifacts or results.
This is not a claim about any real deployed system. The synthetic arm is illustrative (a designed mechanism, not fit to real data); the real-data arm carries no causal claim about telemarketing effectiveness (contract section 15).
An independent forensic review of candidate 6fd57f3 found three blocking
defects (docs/reviews/2026-09-02-6fd57f3-forensic-review.md, verdict
REVISE). All three are corrected in this candidate; every canonical CSV,
figure and number below is from the corrected code, not the reviewed one.
- DA-001 (evidence claim). The frozen candidate's Figure 3 and README
captioned per-segment rows as a test of H1's "pooling amplifies the
calibration effect" mechanism. They were not: capacity in this candidate
is always allocated once, over the pooled population, so a segment's row
is an arithmetic slice of that one selection —
gap(all) = gap(A) + gap(B)always, by construction, regardless of any pooling mechanism. Correction: the claim is retracted, not "fixed" — see "Hypothesis disposition" below and the corrected Figure 3 caption (scripts/make_figures.py). No experiment was expanded to manufacture a within-stratum test; the honest statement is that this candidate does not contain one. - DA-002 (correctness).
SplitConformalClassifier.lower_boundcomputedproba_one - qhaton conformal-ambiguous cases, not the frozen contract'sp_lo(x) = min{p_hat(y|x) : y in C(x)}(contract section 9) — at the reviewer's example (qhat=0.95, ambiguousproba_one=0.05) the old code gave-0.90, an unbounded, non-contract value. Correction: the formula now matches the contract exactly (min(p, 1-p)on ambiguous cases), withtests/test_conformal.pybinding the implementation to that formula, including the reviewer's own numeric example. - DA-003 (correctness).
selective_policycould over- or under-fill capacity: it copied the primary top-K mask, then independently overwrote deferred cases with the baseline rule's own unconstrained top-K decision, which is not guaranteed to have the same size as the freed slots. On the frozen candidate's own canonical grid this showed up as underfill (e.g. ~583 of a K=600 budget spent at f=0.10). Correction:selective_policynow certainly-selects everything outside the deferred set, then fills exactly the freed slots from the deferred pool by the baseline rule's preference, guaranteeing|S| = min(K, N)always.tests/test_policies.pyreproduces the reviewer's exact n=5, K=2 counterexample as a regression, plus ahypothesisproperty test over random inputs. The capacity invariant was re-verified directly against the regenerated canonical CSVs (every policy, every cell): zero mismatches.
Material (non-blocking) findings were also actioned: a calibration-set-size
shrinkage test now exists for H2's stopping condition (DA-004); the Brier
score's reliability/resolution/uncertainty decomposition required by
contract section 10 is now computed and written to every results row
(DA-006); the UCI non-stationarity narrative below and in Figure 5 now
names the unseen-month-dummy confound the review identified and does not
claim to have isolated it from genuine covariate shift (DA-007); every
Monte Carlo "gap" claim below states its normalized SE and, per contract
section 13's 15%-of-mean rule, is labelled inconclusive at this seed
count where that threshold is crossed (DA-005). np.quantile's edge
behaviour (DA-009) and the fetch script's non-fail-closed digest check
(DA-010) are unchanged — non-blocking, not in scope for this correction,
and not expected to change any reported conclusion.
| Command | What it does |
|---|---|
make setup |
Exact sync from uv.lock. |
make check |
Lock, format, lint, type, test, bounded reproduction and content-verified manifest gates. Offline. |
make test |
Unit + property tests (pytest + hypothesis). |
make reproduce |
Bounded, deterministic, offline smoke run of the real pipeline at tiny scale; validates its manifest. |
make experiment |
Canonical synthetic-DGP main grid + noise-sensitivity sweep. Offline, not cheap — see wall time below. |
make external |
Canonical real-data arm: fetches UCI Bank Marketing (network required), runs the policy comparison, validates its manifest. |
make figures |
Regenerates figures/*.png from results/*.csv. |
make clean |
Removes generated bounded output and local caches only. |
make check never touches the network and never runs experiment/external
(those are canonical reproduction, not the fast local gate).
Lineage: public claim → figure/table → machine-readable CSV → run manifest
(MyWorld run-manifest v1) → source/config/input/environment identity. Every
canonical run writes results/<name>-run-manifest.json next to its CSV,
validated by scripts/validate_manifest.py (byte-level digest checks against
the actual config/input/output files, not just schema shape).
Manifest git_sha lineage note. The canonical manifests record
git_sha=8032699 (the preserved-REVISE commit), not the correction commit
6bf5872 that the currently-committed code actually is. This is because the
post-correction canonical CSVs were regenerated from a working tree whose
HEAD was still 8032699 at generation time, immediately before 6bf5872
was committed — an honest lineage gap, not a fabrication. The r01
confirmation review (docs/reviews/2026-09-03-6bf5872-r01-confirmation.md)
independently verified this by replaying the corrected code at 6bf5872
and matching the leftover CSVs to floating-point noise (max |ΔV| = 2.84e-14),
establishing that the committed 6bf5872 code is in fact what produced the
evidence, even though the manifest's own git_sha field predates that
commit by one step. The manifests are left exactly as generated (preserving
the reviewed evidence chain intact); this note is the closure documentation
that resolves the apparent mismatch without rewriting any hash-checked file.
Run in this order, in this repository, from a clean checkout:
uv sync --locked --group dev
make check # fast, offline gate
make experiment # synthetic main grid (600 reps) + noise sweep (60 reps)
make external # network: fetches UCI Bank Marketing, runs the real arm
make figures| Stage | Command | Result | Wall time |
|---|---|---|---|
make check |
see above | 79 tests passed; bounded manifest valid | ~10s |
| Synthetic main grid | run_experiment.py --config configs/main-grid.json |
results/main_grid.csv, 600 replications × 33 rows = 19,800 rows |
293.4s (~4.9 min) |
| Noise sensitivity | run_experiment.py --config configs/noise-sensitivity.json |
results/noise_sensitivity.csv, 60 × 33 = 1,980 rows |
34.1s |
| Real-data arm | fetch_external_data.py then run_external_experiment.py |
results/external_bank_marketing.csv, 1,100 rows |
35.2s (run) + a few seconds to fetch |
| Figures | make_figures.py |
5 PNGs in figures/ |
~5s |
Timings above are from the post-correction regeneration (7 new tests added
by the correction — 72 to 79 — and 3 new Brier-decomposition columns per
row; row counts and grid sizes are unchanged from the original candidate).
All canonical runs completed on the first attempt at their contract-frozen
configuration both times (no seed count, grid size, or metric was adjusted
after inspecting results — contract section 13). Each wrote a MyWorld
run-manifest v1 validated by scripts/validate_manifest.py
(RUN_MANIFEST_VALID for all four, both before and after correction).
Configuration span actually run for this candidate's evidence (AGENTS.md
protocol convention): 40 seeds × 5 capacity fractions × 3 shift
magnitudes for the main grid; 20 seeds × 3 noise temperatures for the
sensitivity sweep; 10 model seeds × 2 split modes × 5 capacity fractions
for the real-data arm, evaluated once on the one available real dataset (not
resampled — contract section 13). Not varied: the benefit/cost ratio
(fixed at 5:1 by contract), the DGP's structural form (weights, segment
count, feature count), the conformal target alpha (fixed at 0.10), or the
ensemble/logistic hyperparameters (fixed, never tuned per replication).
All numbers below are Observed (Reference Standard claim labels) from
this candidate's own canonical, post-correction CSVs —
results/main_grid.csv (19,800 rows), results/noise_sensitivity.csv
(1,980 rows) and results/external_bank_marketing.csv (1,100 rows) —
produced by the runs in the table above, using the corrected
selective_policy and SplitConformalClassifier.lower_bound. Figures are
in figures/.
Notation (normalized per DA-005): for a stochastic quantity, mean ± CI95 is always the mean and a normal 95% confidence interval on the mean
(1.96 × sd/√n); sd is stated separately, in parentheses, only when the
spread itself (not the precision of the mean) is the point. Every
between-condition "gap" is computed as a paired per-seed difference
(common random numbers), and its normalized Monte Carlo SE
(|SE/mean|) is checked against contract section 13's rule: if that
ratio exceeds 15%, the gap is labelled inconclusive at this seed count
and is not asserted as a finding, whatever its sign. A bit-identical
(zero-variance) result is exact, not a Monte Carlo estimate, and is never
labelled inconclusive.
H2 — calibration as a no-op within one stratum: confirmed exactly for Platt, small but real net effect for isotonic
For the logistic scorer, raw_ranking and calibrated_ranking_platt
produced bit-identical decision value in all 1,800 compared cells
(max |raw − Platt| = 0.0). Platt scaling on a linear score is a strictly
monotonic transform of that same score, so H2's identity holds to
floating-point precision.
Isotonic calibration, only weakly monotonic and fit on a finite 2,000-row
calibration window, does not hold the identity exactly. Pooled over every
main-grid cell (segment all, all scorers/capacities/shifts, n=1,200
seed-cells): mean gap calibrated_isotonic − raw = −2.59 ± 0.34
(sd 6.01), SE/mean = 6.7% — not inconclusive, a real small net-negative
effect. Logistic only (n=600): −2.98 ± 0.50 (sd 6.26), SE/mean = 8.6% —
also not inconclusive. Every individual per-capacity, per-segment cell
shown in Figure 3, however, is individually inconclusive at n=40 seeds
(SE/mean ranges 25%–253% across the 15 (segment × capacity) cells at
shift=0) — Figure 3 is illustrative of where the pooled net effect lands,
not 15 independent findings. tests/test_calibration.py additionally
confirms directly (contract section 12's stopping condition, previously
untested — DA-004) that this raw-vs-isotonic top-K disagreement shrinks as
the calibration set grows, so the mechanism producing the small pooled
effect is a genuine finite-sample property, not an implementation defect.
H1 — pooling a shared capacity budget across heterogeneous segments: retracted; not tested by this candidate (DA-001)
An earlier version of this document, and Figure 3's caption, claimed the
per-segment rows above showed H1's "pooling amplifies the calibration
effect" mechanism, with the sign reversed from what was hypothesized. That
claim is retracted, not merely re-signed: capacity in this candidate is
always allocated once, per period, over the pooled two-segment
population — there is no per-segment top-K allocation anywhere in this
candidate's code. Decision value sums independently over any partition of a
selection mask, so gap(all) = gap(A) + gap(B) by construction, for
any scorer, any capacity, any run — that is arithmetic, not evidence about
a pooling mechanism (verified: max deviation from exact additivity across
the grid is 2×10⁻¹³, floating-point noise). Whether a true within-stratum
allocation arm (never implemented here) would show calibration helping,
hurting, or doing nothing is genuinely unknown and not addressed by
this candidate. See docs/hypothesis-disposition.json and Figure 3's
corrected caption. Per the correction brief's own guidance, the fix here is
retracting the claim, not building a new experimental arm to rescue it.
H3 — does the ensemble's complexity earn its place? Yes at the default cell; the noise-sweep's edge cases are inconclusive at this seed count
At capacity=0.10, shift=0 (n=40 paired seeds, main grid): ensemble−logistic decision-value gap = +17.55 ± 3.20 (sd 10.31), SE/mean = 9.3% — not inconclusive, a real advantage. Regret against the synthetic Bayes oracle at the same cell: ensemble 43.95 ± 3.13 (≈11% of its own achievable value), logistic 61.50 ± 3.33 (≈15.5%) — the ensemble leaves less value on the table but neither candidate is close to oracle.
The separate noise-sensitivity sweep (Figure 4, n=20 paired seeds, its own smaller seed count per contract section 17's targeted-sweep design) shows the same gap shrinking as the signal weakens: +37.20 ± 4.33 (sd 9.88, SE/mean 5.9%) at τ=0.6 — not inconclusive; +20.05 ± 6.00 (sd 13.70, SE/mean 15.3%) at τ=1.0 (default) — inconclusive at this seed count; +5.05 ± 4.40 (sd 10.04, SE/mean 44.5%) at τ=1.6 (noisiest tested) — inconclusive at this seed count. Notably, the same nominal condition (τ=1.0, f=0.10, shift=0) is confirmed on the 40-seed main grid but inconclusive on the 20-seed noise sweep — both numbers are reported honestly rather than picking the one that reads as a stronger finding; the difference is seed count and independent sample draw, not a contradiction. This candidate does not claim the ensemble's advantage would stay positive beyond τ=1.6, and does not claim the τ=1.0/1.6 sweep cells are confirmed findings.
H4 — conformal coverage requires exchangeability: confirmed synthetically (modest), confirmed far more dramatically — and only partially explained — on real data
Synthetic (Figure 2, capacity=0.10, n=40): coverage falls from 0.910 ± 0.004 (no shift) to 0.891 ± 0.008 (Δ=2.5, logistic) and 0.854 ± 0.009 (Δ=2.5, ensemble) — a real but modest decline (SE/mean well under 1% throughout; every coverage cell here is precisely estimated).
Real data (Figure 5, UCI Bank Marketing, n=10 model seeds — see the logistic-solver-determinism limitation below): the chronological split (fit on the earliest ~1/3 of the 2008–2010 campaign by row order, calibrate on the next ~1/6, deploy on the rest) collapses coverage to a flat 0.219 for the logistic scorer (zero variance — a deterministic solver on fixed data, not a Monte Carlo estimate) and 0.806 ± 0.012 for the ensemble — both far below the random-split control (0.905 logistic, 0.908 ± 0.002 ensemble, both exchangeable by construction).
This candidate does not isolate why the collapse is that large (DA-007).
Two mechanisms are entangled and were not separated: (a) genuine 2008–2010
economic covariate shift (emp.var.rate mean 1.23 in the training window
vs. −1.12 in deploy), and (b) a real encoding artifact — pd.get_dummies
is fit on the full frame before splitting, so the chronological deploy
window contains one-hot month indicators the model never saw an activated
training example of (train sees only {may, jun, jul}; deploy includes
seven months absent from training). Both are real properties of this
candidate's pipeline on this dataset; this study reports the coverage
collapse as evidence that exchangeability was violated, and explicitly does
not attribute a specific fraction of it to "drift" versus "unseen
categories" — that would be manufacturing evidence this candidate did not
collect. (Interpretive, still not isolated: the tree ensemble's smaller
degradation is consistent with, but not proven to be caused by, its
predictions staying bounded by leaf values seen in training, unlike the
logistic model's unbounded linear extrapolation.)
H5 — selective abstention and the uncertainty-aware policy: corrected implementations, then measured (DA-002, DA-003)
The numbers below are from the corrected selective_policy (exact
min(K,N) capacity, DA-003) and the corrected lower_bound (contract's
min(p, 1-p) formula, DA-002). The pre-correction numbers in the reviewed
candidate are superseded and not reported as evidence — they measured a
different, non-contract policy and an under-filling allocation bug, not
the mechanisms named below.
Sign convention: every paired gap in this section is
calibrated_ranking_isotonic minus the named policy (mirroring H3's
ensemble − logistic convention above). A positive gap means calibrated
ranking scored higher, i.e. the named policy performed worse; a negative
gap means the named policy scored higher.
selective_abstain vs. calibrated_ranking_isotonic, logistic, shift=0,
paired per-seed gap (n=40) at every tested capacity: f=0.02: +0.03 ±
0.05 (inconclusive); f=0.05: −0.10 ± 0.24 (inconclusive); f=0.10:
−0.05 ± 0.60 (inconclusive); f=0.20: −0.18 ± 1.19 (inconclusive);
f=0.40: −0.45 ± 0.87 (inconclusive). Every capacity fraction is
inconclusive at this seed count — selective abstention is statistically
indistinguishable from plain calibrated ranking here. This is expected once
capacity is respected exactly: abstention_rate (the fraction of cases
actually deferred) is tiny — 0.026% at f=0.02 rising to only 1.5%
at f=0.40 — so few decisions are ever actually reassigned. (The pre-
correction candidate's apparent "small tracked-together" result was itself
partly an artifact of the DA-003 underfill bug quietly removing slots
rather than reassigning them; this corrected, mostly-inconclusive-at-zero
result is the honest one.)
uncertainty_aware vs. calibrated_ranking_isotonic, logistic, shift=0,
paired gap (n=40): f=0.02: +0.23 ± 0.28 (inconclusive); f=0.05:
+2.30 ± 2.94 (inconclusive); f=0.10: +17.08 ± 9.19 (inconclusive);
f=0.20: +73.08 ± 13.71 (sd 44.24, SE/mean 9.6% — not inconclusive);
f=0.40: +63.15 ± 14.17 (sd 45.73, SE/mean 11.5% — not inconclusive).
So: at small capacity fractions the corrected uncertainty-aware policy is
not distinguishably worse than calibrated ranking at this seed count; at
f≥0.20 it is really and distinguishably worse — a genuine, still-
reported negative result, just a smaller and more precisely bounded one
than the pre-correction, wrong-formula candidate claimed (previously
~599 units at f=0.40 under the buggy discount; now 63 units, using the
contract's own min(p, 1-p) formula). Roughly 49% of deployment cases
fall in the conformal-ambiguous band at α=0.10 in this DGP (mean
prediction-set size ≈1.49, 0% empty sets), so re-ranking by the
contract-specified conservative bound still measurably reshuffles a large
fraction of the population at larger capacities — a real, disclosed
limitation of "rank by the conformal lower bound" as a policy, not of
conformal prediction itself.
Real-data arm: on the chronological split, selective_abstain and
uncertainty_aware produced decision values identical to
calibrated_ranking_isotonic at every capacity fraction and both scorers
(e.g. ensemble, f=0.40: all three = −143.1; 100/100 compared cells
identical) — no case in that deployment window was both conformal-ambiguous
and capacity-boundary-proximate in a way that changed the outcome. On the
random split, the three policies are not identical: 40 of 100 compared
policy-pairs differ, by up to 4.0 decision-value units. The confirmed
finding is therefore split-specific: on the chronological split (the
split that matches how this system would actually be deployed), neither
extra mechanism earned its keep over plain calibrated ranking; the random
split shows small, real differences the chronological identity does not
generalize to a claim of "no case anywhere changed the outcome" (correction
recorded per the r01 confirmation review, finding DA-C01).
At the real 11.3% base rate and the contract's fixed (b=1, c=0.2) utility,
the break-even selection probability is 0.2 — above the population base
rate. Realized decision value on the real arm is accordingly negative at
every capacity fraction ≥ 0.10 for every policy (e.g. fixed_threshold at
f=0.40: −214.8; calibrated_ranking_isotonic/ensemble at f=0.40:
−143.1 ± 8.8), because filling a large capacity necessarily pulls in
many below-break-even cases. This is not a bug: it is a genuine illustration
of why "more capacity" is not free once a realistic cost is attached, and it
matches the finite-capacity framing directly (contract section 3).
Main grid: 40 generator seeds × 5 capacity fractions × 3 shift
magnitudes, one DGP structural form, one architecture pair, one benefit/cost
setting. Noise sweep: 20 seeds × 3 temperatures, same fixed structure.
Real arm: 10 model seeds × 2 split modes × 5 capacity fractions on
one real dataset (not resampled). Not varied: the DGP's structural
form (weights, segment count/definition, feature count), the benefit/cost
ratio, the conformal target α, the ensemble/logistic hyperparameters, or the
number of segments/strata. A material limitation discovered while
analysing results, not anticipated in the contract: RegularisedLogistic's
lbfgs solver is deterministic given its inputs, so its real-arm results
have zero variance across the 10 model_seed values — the real arm's
seed sweep gives genuine Monte Carlo information only for the ensemble
scorer, not the logistic one. The logistic real-arm numbers above are
therefore single deterministic point estimates, not seed-averaged means,
and are reported as such.
Machine-readable summary: docs/hypothesis-disposition.json.
| Hypothesis | Disposition |
|---|---|
| H1 (pooling amplifies calibration's effect) | Retracted — untested. No within-stratum allocation exists in this candidate; the figure formerly cited as evidence is an arithmetic identity (DA-001). |
| H2 (monotonic recalibration ≈ no-op within a stratum) | Confirmed exactly for Platt (bit-identical); small real net effect for isotonic, with the required calibration-size shrinkage now directly tested. |
| H3 (complexity earns its place) | Confirmed at the default cell (40 seeds); inconclusive at this seed count at the noise sweep's default and noisiest cells (20 seeds, per contract §13). |
| H4 (coverage needs exchangeability) | Confirmed, both arms; the real arm's much larger effect is not isolated from a named, disclosed confound (unseen categorical levels vs. covariate shift). |
| H5 (selective abstention value) | Corrected implementations, then measured. Abstention: inconclusive at every tested capacity. Uncertainty-aware ranking: inconclusive at small capacity, really worse at f≥0.20 — smaller and more precisely bounded than the pre-correction, wrong-formula result. |
None to the frozen design (question, hypotheses, DGP, candidates, policies, metrics, robustness arms, stopping criteria). Two data-handling decisions for the real-data arm were not fully pinned by the contract's prose and are recorded here rather than by editing the frozen contract:
durationis dropped as a feature in the UCI Bank Marketing arm — it is a documented post-outcome leak (seedocs/references/README.md).month/day_of_weekare kept as one-hot features, not only as the temporal-split axis, since they are known at contact time. (This choice is also the source of the unseen-category confound named in H4 above and DA-007 — it is not reversed here, since doing so would be new experimental design, not a claim correction; it is disclosed instead.)
SplitConformalClassifier.lower_bound and selective_policy are not
listed as deviations: both now implement exactly what the frozen contract
specifies (sections 9 and 3 respectively). Before this correction cycle
they were undeclared deviations — see "Correction record" above (DA-002,
DA-003) for what changed and why that is a defect fix, not a documented
design choice.
Restated from the contract (section 15), with what this candidate's results made concrete added inline, and updated for the correction cycle:
- The synthetic DGP is illustrative of a mechanism, not fit to real data; generalization is bounded by how much a real deployment resembles the pooled-heterogeneous-capacity structure it encodes. Its shift arm, in particular, turned out to be gentler than the real UCI arm's own non-stationarity (coverage fell to 0.85-0.89 synthetically vs. 0.22-0.81 really) — a limitation of the synthetic calibration, not a claim that real drift is always this severe.
- No regret/oracle claim exists for the real-data arm (no known p(x)); only an ex-post/illustrative ceiling is reported there, never an achievable policy.
- No interference/spillover assumption is tested.
- Conformal marginal coverage does not imply subgroup-conditional coverage; subgroup/period coverage is diagnostic only.
- The benefit/cost ratio (5:1) is fixed by contract and not swept; the real arm's negative decision values at larger capacity are a direct, expected consequence of that fixed ratio against an 11.3% base rate, not a calibration or policy defect.
- The real dataset's "non-stationarity" is contact order within one campaign (2008-2010, spanning the financial crisis in its economic features), entangled with an unseen-categorical-level encoding artifact this candidate did not isolate from it (H4 above, DA-007) — real, and considerably more severe in its effect on coverage than this study anticipated when the contract was frozen, but its size should not be read as a clean "drift" number, and it remains one dataset's one campaign, not a general claim about drift severity.
- The real arm's
model_seedsweep varies only the ensemble meaningfully;LogisticRegression's deterministic solver means its real-arm numbers are point estimates, not seed-averaged means (see Findings, "Configuration span actually varied"). - Every per-cell H1 illustration and every H3/H5 gap not explicitly marked
"not inconclusive" above should be read as
inconclusive at this seed count(contract section 13) — this study does not silently upgrade an inconclusive Monte Carlo estimate into a confirmed finding anywhere. - Contract section 10 requires "reliability-diagram data (bin means)
written to results for plotting."
metrics.reliability_bins()computes this and is unit-tested, and its aggregate (the Brier reliability/resolution/uncertainty decomposition) is now a column in every results row; the per-bin data itself is not yet written to a separate results artifact. This is a disclosed gap (A2 forensic review DA-006), not silently omitted: ECE plus the Brier decomposition are the calibration diagnostics actually available inresults/*.csvtoday. brier/ece/the Brier decomposition columns are scorer-level diagnostics of the calibrated probability, computed once per scorer and repeated identically across every policy row that shares that scorer (includingraw_ranking, whose own ranking score is not itself a probability) — they are not a distinct calibration measurement per policy. This was always true of the implementation; it is now stated explicitly (DA-006).np.quantile(..., method="higher")'s edge behaviour at small n can deviate from the textbook order statistic (A2 forensic review DA-009, non-blocking); no-shift coverage on the canonical grid is still ≈0.91, slightly conservative, and this is not treated as breaking H4.scripts/fetch_external_data.pyprints digests for the operator to record rather than failing closed against the pins indocs/references/README.md(DA-010, non-blocking); a future fetch of a mutated upstream file would not be automatically rejected.
docs/RESEARCH-CONTRACT.md frozen contract (read this first)
docs/reviews/ independent forensic review this correction responds to
docs/hypothesis-disposition.json machine-readable H1-H5 status summary
docs/references/ source/provenance record (Reference and Evidence Standard)
src/decision_assurance/ dgp.py, models.py, calibration.py, conformal.py,
allocation.py, policies.py, metrics.py,
experiment.py, grid.py, external.py, manifest.py
scripts/ reproduce.py, run_experiment.py, fetch_external_data.py,
run_external_experiment.py, make_figures.py, validate_manifest.py
configs/ bounded-fixture.json, main-grid.json,
noise-sensitivity.json, external-experiment.json
tests/ unit + property tests; tests/fixtures/hand_authored/
holds the only hand-authored (non-generated, non-downloaded) fixture
results/ generated CSVs + run manifests (gitignored; regenerate via scripts/)
figures/ generated PNGs (committed, so they render on GitHub)
data/external/ UCI Bank Marketing, fetched on demand (gitignored, digest-pinned)
A model can be excellent at telling who is more likely to have a positive outcome and still make a worse real decision than a simpler, honestly- calibrated rule, the moment a fixed budget has to be split across groups whose raw scores are not on the same scale, or the world changes between tuning and use. This study measures that gap directly with a distribution- free uncertainty method that is honest about when it can and cannot promise coverage. What fails if this is wrong: a high-discrimination but poorly calibrated model looks safe on a leaderboard metric while quietly misallocating a real, bounded budget — and, as this study's own correction cycle showed, a figure caption can make the same mistake at one remove: an earlier version of Figure 3 told a reader that pooling capacity across groups was tested and found to reverse-sign calibration's effect, when what was actually plotted was an arithmetic slice of one selection that cannot show that. What remains uncertain: how far the synthetic mechanism generalizes beyond this study's own DGP and one real dataset; whether a true within-stratum allocation arm (never built here) would show pooling mattering at all; and how sensitive the conclusions are to the fixed benefit/cost ratio this contract deliberately does not sweep.
The code, tests, configuration and documentation in this repository are licensed under the MIT License (copyright Joshua de Freitas).
This study's real-data arm fetches the UCI Machine Learning Repository's
Bank Marketing dataset at run time (scripts/fetch_external_data.py); the
raw dataset is not committed to this repository (data/external/ is
gitignored) and is not covered by the MIT license above. It remains subject
to the UCI Machine Learning Repository's own terms
(https://archive.ics.uci.edu/dataset/222/bank+marketing), which this
repository does not restate or modify. Reproduction of the real-data arm
requires fetching that dataset directly from its source under those terms.