Skip to content

Repository files navigation

Calibrated Finite-Capacity Decision Assurance (A2)

Engineering profile: RESEARCH Toolchain: CPython 3.14 + uv 0.12.5 Quality gate: make check Declared deviations: none to the frozen contract's design; two UCI data-handling clarifications not fully pinned by the contract's prose are recorded below and in docs/references/README.md. Contract: docs/RESEARCH-CONTRACT.md — read this first. It is frozen; anything below implements it, and any deviation discovered during implementation is recorded here, never silently made in the contract itself. Correction record: an independent forensic review (docs/reviews/2026-09-02-6fd57f3-forensic-review.md, verdict REVISE) found three blocking defects in the candidate this document used to describe. All three are corrected in this version — see "Correction record" below before trusting any number that predates it.

Purpose and boundary

Under what conditions is ranking/prediction quality (discrimination) alone insufficient for a reliable finite-capacity decision, and how do calibration and uncertainty-aware decision rules change realized decision value relative to a transparent baseline and to plain score ranking? Second foundation of Track A (Model & Data Systems) in the Dual-Foundation Public Technical Evidence Programme. Self-contained: its own controlled DGP plus one independently sourced real dataset (UCI Bank Marketing); it consumes no A1 artifacts or results.

This is not a claim about any real deployed system. The synthetic arm is illustrative (a designed mechanism, not fit to real data); the real-data arm carries no causal claim about telemarketing effectiveness (contract section 15).

Correction record (A2 forensic review)

An independent forensic review of candidate 6fd57f3 found three blocking defects (docs/reviews/2026-09-02-6fd57f3-forensic-review.md, verdict REVISE). All three are corrected in this candidate; every canonical CSV, figure and number below is from the corrected code, not the reviewed one.

  • DA-001 (evidence claim). The frozen candidate's Figure 3 and README captioned per-segment rows as a test of H1's "pooling amplifies the calibration effect" mechanism. They were not: capacity in this candidate is always allocated once, over the pooled population, so a segment's row is an arithmetic slice of that one selection — gap(all) = gap(A) + gap(B) always, by construction, regardless of any pooling mechanism. Correction: the claim is retracted, not "fixed" — see "Hypothesis disposition" below and the corrected Figure 3 caption (scripts/make_figures.py). No experiment was expanded to manufacture a within-stratum test; the honest statement is that this candidate does not contain one.
  • DA-002 (correctness). SplitConformalClassifier.lower_bound computed proba_one - qhat on conformal-ambiguous cases, not the frozen contract's p_lo(x) = min{p_hat(y|x) : y in C(x)} (contract section 9) — at the reviewer's example (qhat=0.95, ambiguous proba_one=0.05) the old code gave -0.90, an unbounded, non-contract value. Correction: the formula now matches the contract exactly (min(p, 1-p) on ambiguous cases), with tests/test_conformal.py binding the implementation to that formula, including the reviewer's own numeric example.
  • DA-003 (correctness). selective_policy could over- or under-fill capacity: it copied the primary top-K mask, then independently overwrote deferred cases with the baseline rule's own unconstrained top-K decision, which is not guaranteed to have the same size as the freed slots. On the frozen candidate's own canonical grid this showed up as underfill (e.g. ~583 of a K=600 budget spent at f=0.10). Correction: selective_policy now certainly-selects everything outside the deferred set, then fills exactly the freed slots from the deferred pool by the baseline rule's preference, guaranteeing |S| = min(K, N) always. tests/test_policies.py reproduces the reviewer's exact n=5, K=2 counterexample as a regression, plus a hypothesis property test over random inputs. The capacity invariant was re-verified directly against the regenerated canonical CSVs (every policy, every cell): zero mismatches.

Material (non-blocking) findings were also actioned: a calibration-set-size shrinkage test now exists for H2's stopping condition (DA-004); the Brier score's reliability/resolution/uncertainty decomposition required by contract section 10 is now computed and written to every results row (DA-006); the UCI non-stationarity narrative below and in Figure 5 now names the unseen-month-dummy confound the review identified and does not claim to have isolated it from genuine covariate shift (DA-007); every Monte Carlo "gap" claim below states its normalized SE and, per contract section 13's 15%-of-mean rule, is labelled inconclusive at this seed count where that threshold is crossed (DA-005). np.quantile's edge behaviour (DA-009) and the fetch script's non-fail-closed digest check (DA-010) are unchanged — non-blocking, not in scope for this correction, and not expected to change any reported conclusion.

Commands

Command What it does
make setup Exact sync from uv.lock.
make check Lock, format, lint, type, test, bounded reproduction and content-verified manifest gates. Offline.
make test Unit + property tests (pytest + hypothesis).
make reproduce Bounded, deterministic, offline smoke run of the real pipeline at tiny scale; validates its manifest.
make experiment Canonical synthetic-DGP main grid + noise-sensitivity sweep. Offline, not cheap — see wall time below.
make external Canonical real-data arm: fetches UCI Bank Marketing (network required), runs the policy comparison, validates its manifest.
make figures Regenerates figures/*.png from results/*.csv.
make clean Removes generated bounded output and local caches only.

make check never touches the network and never runs experiment/external (those are canonical reproduction, not the fast local gate).

Evidence and reproduction

Lineage: public claim → figure/table → machine-readable CSV → run manifest (MyWorld run-manifest v1) → source/config/input/environment identity. Every canonical run writes results/<name>-run-manifest.json next to its CSV, validated by scripts/validate_manifest.py (byte-level digest checks against the actual config/input/output files, not just schema shape).

Manifest git_sha lineage note. The canonical manifests record git_sha=8032699 (the preserved-REVISE commit), not the correction commit 6bf5872 that the currently-committed code actually is. This is because the post-correction canonical CSVs were regenerated from a working tree whose HEAD was still 8032699 at generation time, immediately before 6bf5872 was committed — an honest lineage gap, not a fabrication. The r01 confirmation review (docs/reviews/2026-09-03-6bf5872-r01-confirmation.md) independently verified this by replaying the corrected code at 6bf5872 and matching the leftover CSVs to floating-point noise (max |ΔV| = 2.84e-14), establishing that the committed 6bf5872 code is in fact what produced the evidence, even though the manifest's own git_sha field predates that commit by one step. The manifests are left exactly as generated (preserving the reviewed evidence chain intact); this note is the closure documentation that resolves the apparent mismatch without rewriting any hash-checked file.

Canonical reproduction — exact commands and observed wall time

Run in this order, in this repository, from a clean checkout:

uv sync --locked --group dev
make check                 # fast, offline gate
make experiment            # synthetic main grid (600 reps) + noise sweep (60 reps)
make external               # network: fetches UCI Bank Marketing, runs the real arm
make figures
Stage Command Result Wall time
make check see above 79 tests passed; bounded manifest valid ~10s
Synthetic main grid run_experiment.py --config configs/main-grid.json results/main_grid.csv, 600 replications × 33 rows = 19,800 rows 293.4s (~4.9 min)
Noise sensitivity run_experiment.py --config configs/noise-sensitivity.json results/noise_sensitivity.csv, 60 × 33 = 1,980 rows 34.1s
Real-data arm fetch_external_data.py then run_external_experiment.py results/external_bank_marketing.csv, 1,100 rows 35.2s (run) + a few seconds to fetch
Figures make_figures.py 5 PNGs in figures/ ~5s

Timings above are from the post-correction regeneration (7 new tests added by the correction — 72 to 79 — and 3 new Brier-decomposition columns per row; row counts and grid sizes are unchanged from the original candidate). All canonical runs completed on the first attempt at their contract-frozen configuration both times (no seed count, grid size, or metric was adjusted after inspecting results — contract section 13). Each wrote a MyWorld run-manifest v1 validated by scripts/validate_manifest.py (RUN_MANIFEST_VALID for all four, both before and after correction).

Configuration span actually run for this candidate's evidence (AGENTS.md protocol convention): 40 seeds × 5 capacity fractions × 3 shift magnitudes for the main grid; 20 seeds × 3 noise temperatures for the sensitivity sweep; 10 model seeds × 2 split modes × 5 capacity fractions for the real-data arm, evaluated once on the one available real dataset (not resampled — contract section 13). Not varied: the benefit/cost ratio (fixed at 5:1 by contract), the DGP's structural form (weights, segment count, feature count), the conformal target alpha (fixed at 0.10), or the ensemble/logistic hyperparameters (fixed, never tuned per replication).

Findings

All numbers below are Observed (Reference Standard claim labels) from this candidate's own canonical, post-correction CSVs — results/main_grid.csv (19,800 rows), results/noise_sensitivity.csv (1,980 rows) and results/external_bank_marketing.csv (1,100 rows) — produced by the runs in the table above, using the corrected selective_policy and SplitConformalClassifier.lower_bound. Figures are in figures/.

Notation (normalized per DA-005): for a stochastic quantity, mean ± CI95 is always the mean and a normal 95% confidence interval on the mean (1.96 × sd/√n); sd is stated separately, in parentheses, only when the spread itself (not the precision of the mean) is the point. Every between-condition "gap" is computed as a paired per-seed difference (common random numbers), and its normalized Monte Carlo SE (|SE/mean|) is checked against contract section 13's rule: if that ratio exceeds 15%, the gap is labelled inconclusive at this seed count and is not asserted as a finding, whatever its sign. A bit-identical (zero-variance) result is exact, not a Monte Carlo estimate, and is never labelled inconclusive.

H2 — calibration as a no-op within one stratum: confirmed exactly for Platt, small but real net effect for isotonic

For the logistic scorer, raw_ranking and calibrated_ranking_platt produced bit-identical decision value in all 1,800 compared cells (max |raw − Platt| = 0.0). Platt scaling on a linear score is a strictly monotonic transform of that same score, so H2's identity holds to floating-point precision.

Isotonic calibration, only weakly monotonic and fit on a finite 2,000-row calibration window, does not hold the identity exactly. Pooled over every main-grid cell (segment all, all scorers/capacities/shifts, n=1,200 seed-cells): mean gap calibrated_isotonic − raw = −2.59 ± 0.34 (sd 6.01), SE/mean = 6.7% — not inconclusive, a real small net-negative effect. Logistic only (n=600): −2.98 ± 0.50 (sd 6.26), SE/mean = 8.6% — also not inconclusive. Every individual per-capacity, per-segment cell shown in Figure 3, however, is individually inconclusive at n=40 seeds (SE/mean ranges 25%–253% across the 15 (segment × capacity) cells at shift=0) — Figure 3 is illustrative of where the pooled net effect lands, not 15 independent findings. tests/test_calibration.py additionally confirms directly (contract section 12's stopping condition, previously untested — DA-004) that this raw-vs-isotonic top-K disagreement shrinks as the calibration set grows, so the mechanism producing the small pooled effect is a genuine finite-sample property, not an implementation defect.

H1 — pooling a shared capacity budget across heterogeneous segments: retracted; not tested by this candidate (DA-001)

An earlier version of this document, and Figure 3's caption, claimed the per-segment rows above showed H1's "pooling amplifies the calibration effect" mechanism, with the sign reversed from what was hypothesized. That claim is retracted, not merely re-signed: capacity in this candidate is always allocated once, per period, over the pooled two-segment population — there is no per-segment top-K allocation anywhere in this candidate's code. Decision value sums independently over any partition of a selection mask, so gap(all) = gap(A) + gap(B) by construction, for any scorer, any capacity, any run — that is arithmetic, not evidence about a pooling mechanism (verified: max deviation from exact additivity across the grid is 2×10⁻¹³, floating-point noise). Whether a true within-stratum allocation arm (never implemented here) would show calibration helping, hurting, or doing nothing is genuinely unknown and not addressed by this candidate. See docs/hypothesis-disposition.json and Figure 3's corrected caption. Per the correction brief's own guidance, the fix here is retracting the claim, not building a new experimental arm to rescue it.

H3 — does the ensemble's complexity earn its place? Yes at the default cell; the noise-sweep's edge cases are inconclusive at this seed count

At capacity=0.10, shift=0 (n=40 paired seeds, main grid): ensemble−logistic decision-value gap = +17.55 ± 3.20 (sd 10.31), SE/mean = 9.3% — not inconclusive, a real advantage. Regret against the synthetic Bayes oracle at the same cell: ensemble 43.95 ± 3.13 (≈11% of its own achievable value), logistic 61.50 ± 3.33 (≈15.5%) — the ensemble leaves less value on the table but neither candidate is close to oracle.

The separate noise-sensitivity sweep (Figure 4, n=20 paired seeds, its own smaller seed count per contract section 17's targeted-sweep design) shows the same gap shrinking as the signal weakens: +37.20 ± 4.33 (sd 9.88, SE/mean 5.9%) at τ=0.6 — not inconclusive; +20.05 ± 6.00 (sd 13.70, SE/mean 15.3%) at τ=1.0 (default) — inconclusive at this seed count; +5.05 ± 4.40 (sd 10.04, SE/mean 44.5%) at τ=1.6 (noisiest tested) — inconclusive at this seed count. Notably, the same nominal condition (τ=1.0, f=0.10, shift=0) is confirmed on the 40-seed main grid but inconclusive on the 20-seed noise sweep — both numbers are reported honestly rather than picking the one that reads as a stronger finding; the difference is seed count and independent sample draw, not a contradiction. This candidate does not claim the ensemble's advantage would stay positive beyond τ=1.6, and does not claim the τ=1.0/1.6 sweep cells are confirmed findings.

H4 — conformal coverage requires exchangeability: confirmed synthetically (modest), confirmed far more dramatically — and only partially explained — on real data

Synthetic (Figure 2, capacity=0.10, n=40): coverage falls from 0.910 ± 0.004 (no shift) to 0.891 ± 0.008 (Δ=2.5, logistic) and 0.854 ± 0.009 (Δ=2.5, ensemble) — a real but modest decline (SE/mean well under 1% throughout; every coverage cell here is precisely estimated).

Real data (Figure 5, UCI Bank Marketing, n=10 model seeds — see the logistic-solver-determinism limitation below): the chronological split (fit on the earliest ~1/3 of the 2008–2010 campaign by row order, calibrate on the next ~1/6, deploy on the rest) collapses coverage to a flat 0.219 for the logistic scorer (zero variance — a deterministic solver on fixed data, not a Monte Carlo estimate) and 0.806 ± 0.012 for the ensemble — both far below the random-split control (0.905 logistic, 0.908 ± 0.002 ensemble, both exchangeable by construction).

This candidate does not isolate why the collapse is that large (DA-007). Two mechanisms are entangled and were not separated: (a) genuine 2008–2010 economic covariate shift (emp.var.rate mean 1.23 in the training window vs. −1.12 in deploy), and (b) a real encoding artifact — pd.get_dummies is fit on the full frame before splitting, so the chronological deploy window contains one-hot month indicators the model never saw an activated training example of (train sees only {may, jun, jul}; deploy includes seven months absent from training). Both are real properties of this candidate's pipeline on this dataset; this study reports the coverage collapse as evidence that exchangeability was violated, and explicitly does not attribute a specific fraction of it to "drift" versus "unseen categories" — that would be manufacturing evidence this candidate did not collect. (Interpretive, still not isolated: the tree ensemble's smaller degradation is consistent with, but not proven to be caused by, its predictions staying bounded by leaf values seen in training, unlike the logistic model's unbounded linear extrapolation.)

H5 — selective abstention and the uncertainty-aware policy: corrected implementations, then measured (DA-002, DA-003)

The numbers below are from the corrected selective_policy (exact min(K,N) capacity, DA-003) and the corrected lower_bound (contract's min(p, 1-p) formula, DA-002). The pre-correction numbers in the reviewed candidate are superseded and not reported as evidence — they measured a different, non-contract policy and an under-filling allocation bug, not the mechanisms named below.

Sign convention: every paired gap in this section is calibrated_ranking_isotonic minus the named policy (mirroring H3's ensemble − logistic convention above). A positive gap means calibrated ranking scored higher, i.e. the named policy performed worse; a negative gap means the named policy scored higher.

selective_abstain vs. calibrated_ranking_isotonic, logistic, shift=0, paired per-seed gap (n=40) at every tested capacity: f=0.02: +0.03 ± 0.05 (inconclusive); f=0.05: −0.10 ± 0.24 (inconclusive); f=0.10: −0.05 ± 0.60 (inconclusive); f=0.20: −0.18 ± 1.19 (inconclusive); f=0.40: −0.45 ± 0.87 (inconclusive). Every capacity fraction is inconclusive at this seed count — selective abstention is statistically indistinguishable from plain calibrated ranking here. This is expected once capacity is respected exactly: abstention_rate (the fraction of cases actually deferred) is tiny — 0.026% at f=0.02 rising to only 1.5% at f=0.40 — so few decisions are ever actually reassigned. (The pre- correction candidate's apparent "small tracked-together" result was itself partly an artifact of the DA-003 underfill bug quietly removing slots rather than reassigning them; this corrected, mostly-inconclusive-at-zero result is the honest one.)

uncertainty_aware vs. calibrated_ranking_isotonic, logistic, shift=0, paired gap (n=40): f=0.02: +0.23 ± 0.28 (inconclusive); f=0.05: +2.30 ± 2.94 (inconclusive); f=0.10: +17.08 ± 9.19 (inconclusive); f=0.20: +73.08 ± 13.71 (sd 44.24, SE/mean 9.6% — not inconclusive); f=0.40: +63.15 ± 14.17 (sd 45.73, SE/mean 11.5% — not inconclusive). So: at small capacity fractions the corrected uncertainty-aware policy is not distinguishably worse than calibrated ranking at this seed count; at f≥0.20 it is really and distinguishably worse — a genuine, still- reported negative result, just a smaller and more precisely bounded one than the pre-correction, wrong-formula candidate claimed (previously ~599 units at f=0.40 under the buggy discount; now 63 units, using the contract's own min(p, 1-p) formula). Roughly 49% of deployment cases fall in the conformal-ambiguous band at α=0.10 in this DGP (mean prediction-set size ≈1.49, 0% empty sets), so re-ranking by the contract-specified conservative bound still measurably reshuffles a large fraction of the population at larger capacities — a real, disclosed limitation of "rank by the conformal lower bound" as a policy, not of conformal prediction itself.

Real-data arm: on the chronological split, selective_abstain and uncertainty_aware produced decision values identical to calibrated_ranking_isotonic at every capacity fraction and both scorers (e.g. ensemble, f=0.40: all three = −143.1; 100/100 compared cells identical) — no case in that deployment window was both conformal-ambiguous and capacity-boundary-proximate in a way that changed the outcome. On the random split, the three policies are not identical: 40 of 100 compared policy-pairs differ, by up to 4.0 decision-value units. The confirmed finding is therefore split-specific: on the chronological split (the split that matches how this system would actually be deployed), neither extra mechanism earned its keep over plain calibrated ranking; the random split shows small, real differences the chronological identity does not generalize to a claim of "no case anywhere changed the outcome" (correction recorded per the r01 confirmation review, finding DA-C01).

External arm as a real capacity-economics check

At the real 11.3% base rate and the contract's fixed (b=1, c=0.2) utility, the break-even selection probability is 0.2 — above the population base rate. Realized decision value on the real arm is accordingly negative at every capacity fraction ≥ 0.10 for every policy (e.g. fixed_threshold at f=0.40: −214.8; calibrated_ranking_isotonic/ensemble at f=0.40: −143.1 ± 8.8), because filling a large capacity necessarily pulls in many below-break-even cases. This is not a bug: it is a genuine illustration of why "more capacity" is not free once a realistic cost is attached, and it matches the finite-capacity framing directly (contract section 3).

Configuration span actually varied (protocol convention)

Main grid: 40 generator seeds × 5 capacity fractions × 3 shift magnitudes, one DGP structural form, one architecture pair, one benefit/cost setting. Noise sweep: 20 seeds × 3 temperatures, same fixed structure. Real arm: 10 model seeds × 2 split modes × 5 capacity fractions on one real dataset (not resampled). Not varied: the DGP's structural form (weights, segment count/definition, feature count), the benefit/cost ratio, the conformal target α, the ensemble/logistic hyperparameters, or the number of segments/strata. A material limitation discovered while analysing results, not anticipated in the contract: RegularisedLogistic's lbfgs solver is deterministic given its inputs, so its real-arm results have zero variance across the 10 model_seed values — the real arm's seed sweep gives genuine Monte Carlo information only for the ensemble scorer, not the logistic one. The logistic real-arm numbers above are therefore single deterministic point estimates, not seed-averaged means, and are reported as such.

Hypothesis disposition

Machine-readable summary: docs/hypothesis-disposition.json.

Hypothesis Disposition
H1 (pooling amplifies calibration's effect) Retracted — untested. No within-stratum allocation exists in this candidate; the figure formerly cited as evidence is an arithmetic identity (DA-001).
H2 (monotonic recalibration ≈ no-op within a stratum) Confirmed exactly for Platt (bit-identical); small real net effect for isotonic, with the required calibration-size shrinkage now directly tested.
H3 (complexity earns its place) Confirmed at the default cell (40 seeds); inconclusive at this seed count at the noise sweep's default and noisiest cells (20 seeds, per contract §13).
H4 (coverage needs exchangeability) Confirmed, both arms; the real arm's much larger effect is not isolated from a named, disclosed confound (unseen categorical levels vs. covariate shift).
H5 (selective abstention value) Corrected implementations, then measured. Abstention: inconclusive at every tested capacity. Uncertainty-aware ranking: inconclusive at small capacity, really worse at f≥0.20 — smaller and more precisely bounded than the pre-correction, wrong-formula result.

Deviations from the contract

None to the frozen design (question, hypotheses, DGP, candidates, policies, metrics, robustness arms, stopping criteria). Two data-handling decisions for the real-data arm were not fully pinned by the contract's prose and are recorded here rather than by editing the frozen contract:

  • duration is dropped as a feature in the UCI Bank Marketing arm — it is a documented post-outcome leak (see docs/references/README.md).
  • month/day_of_week are kept as one-hot features, not only as the temporal-split axis, since they are known at contact time. (This choice is also the source of the unseen-category confound named in H4 above and DA-007 — it is not reversed here, since doing so would be new experimental design, not a claim correction; it is disclosed instead.)

SplitConformalClassifier.lower_bound and selective_policy are not listed as deviations: both now implement exactly what the frozen contract specifies (sections 9 and 3 respectively). Before this correction cycle they were undeclared deviations — see "Correction record" above (DA-002, DA-003) for what changed and why that is a defect fix, not a documented design choice.

Limitations

Restated from the contract (section 15), with what this candidate's results made concrete added inline, and updated for the correction cycle:

  • The synthetic DGP is illustrative of a mechanism, not fit to real data; generalization is bounded by how much a real deployment resembles the pooled-heterogeneous-capacity structure it encodes. Its shift arm, in particular, turned out to be gentler than the real UCI arm's own non-stationarity (coverage fell to 0.85-0.89 synthetically vs. 0.22-0.81 really) — a limitation of the synthetic calibration, not a claim that real drift is always this severe.
  • No regret/oracle claim exists for the real-data arm (no known p(x)); only an ex-post/illustrative ceiling is reported there, never an achievable policy.
  • No interference/spillover assumption is tested.
  • Conformal marginal coverage does not imply subgroup-conditional coverage; subgroup/period coverage is diagnostic only.
  • The benefit/cost ratio (5:1) is fixed by contract and not swept; the real arm's negative decision values at larger capacity are a direct, expected consequence of that fixed ratio against an 11.3% base rate, not a calibration or policy defect.
  • The real dataset's "non-stationarity" is contact order within one campaign (2008-2010, spanning the financial crisis in its economic features), entangled with an unseen-categorical-level encoding artifact this candidate did not isolate from it (H4 above, DA-007) — real, and considerably more severe in its effect on coverage than this study anticipated when the contract was frozen, but its size should not be read as a clean "drift" number, and it remains one dataset's one campaign, not a general claim about drift severity.
  • The real arm's model_seed sweep varies only the ensemble meaningfully; LogisticRegression's deterministic solver means its real-arm numbers are point estimates, not seed-averaged means (see Findings, "Configuration span actually varied").
  • Every per-cell H1 illustration and every H3/H5 gap not explicitly marked "not inconclusive" above should be read as inconclusive at this seed count (contract section 13) — this study does not silently upgrade an inconclusive Monte Carlo estimate into a confirmed finding anywhere.
  • Contract section 10 requires "reliability-diagram data (bin means) written to results for plotting." metrics.reliability_bins() computes this and is unit-tested, and its aggregate (the Brier reliability/resolution/uncertainty decomposition) is now a column in every results row; the per-bin data itself is not yet written to a separate results artifact. This is a disclosed gap (A2 forensic review DA-006), not silently omitted: ECE plus the Brier decomposition are the calibration diagnostics actually available in results/*.csv today.
  • brier/ece/the Brier decomposition columns are scorer-level diagnostics of the calibrated probability, computed once per scorer and repeated identically across every policy row that shares that scorer (including raw_ranking, whose own ranking score is not itself a probability) — they are not a distinct calibration measurement per policy. This was always true of the implementation; it is now stated explicitly (DA-006).
  • np.quantile(..., method="higher")'s edge behaviour at small n can deviate from the textbook order statistic (A2 forensic review DA-009, non-blocking); no-shift coverage on the canonical grid is still ≈0.91, slightly conservative, and this is not treated as breaking H4.
  • scripts/fetch_external_data.py prints digests for the operator to record rather than failing closed against the pins in docs/references/README.md (DA-010, non-blocking); a future fetch of a mutated upstream file would not be automatically rejected.

Repository layout

docs/RESEARCH-CONTRACT.md   frozen contract (read this first)
docs/reviews/                independent forensic review this correction responds to
docs/hypothesis-disposition.json  machine-readable H1-H5 status summary
docs/references/            source/provenance record (Reference and Evidence Standard)
src/decision_assurance/     dgp.py, models.py, calibration.py, conformal.py,
                             allocation.py, policies.py, metrics.py,
                             experiment.py, grid.py, external.py, manifest.py
scripts/                    reproduce.py, run_experiment.py, fetch_external_data.py,
                             run_external_experiment.py, make_figures.py, validate_manifest.py
configs/                    bounded-fixture.json, main-grid.json,
                             noise-sensitivity.json, external-experiment.json
tests/                      unit + property tests; tests/fixtures/hand_authored/
                             holds the only hand-authored (non-generated, non-downloaded) fixture
results/                    generated CSVs + run manifests (gitignored; regenerate via scripts/)
figures/                    generated PNGs (committed, so they render on GitHub)
data/external/               UCI Bank Marketing, fetched on demand (gitignored, digest-pinned)

Plain-language model

A model can be excellent at telling who is more likely to have a positive outcome and still make a worse real decision than a simpler, honestly- calibrated rule, the moment a fixed budget has to be split across groups whose raw scores are not on the same scale, or the world changes between tuning and use. This study measures that gap directly with a distribution- free uncertainty method that is honest about when it can and cannot promise coverage. What fails if this is wrong: a high-discrimination but poorly calibrated model looks safe on a leaderboard metric while quietly misallocating a real, bounded budget — and, as this study's own correction cycle showed, a figure caption can make the same mistake at one remove: an earlier version of Figure 3 told a reader that pooling capacity across groups was tested and found to reverse-sign calibration's effect, when what was actually plotted was an arithmetic slice of one selection that cannot show that. What remains uncertain: how far the synthetic mechanism generalizes beyond this study's own DGP and one real dataset; whether a true within-stratum allocation arm (never built here) would show pooling mattering at all; and how sensitive the conclusions are to the fixed benefit/cost ratio this contract deliberately does not sweep.

License and data rights

The code, tests, configuration and documentation in this repository are licensed under the MIT License (copyright Joshua de Freitas).

This study's real-data arm fetches the UCI Machine Learning Repository's Bank Marketing dataset at run time (scripts/fetch_external_data.py); the raw dataset is not committed to this repository (data/external/ is gitignored) and is not covered by the MIT license above. It remains subject to the UCI Machine Learning Repository's own terms (https://archive.ics.uci.edu/dataset/222/bank+marketing), which this repository does not restate or modify. Reproduction of the real-data arm requires fetching that dataset directly from its source under those terms.

About

Calibrated finite-capacity decision assurance under predictive uncertainty.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages