Skip to content

Restore modern NumPy/SciPy compatibility, fix the long-range LD filter, compile the sampler, modernise packaging/CI - #164

Merged
bvilhjal merged 8 commits into
masterfrom
claude/organization-efficiency-qv29rl
Jul 29, 2026
Merged

bvilhjal merged 8 commits into
masterfrom
claude/organization-efficiency-qv29rl

Conversation

@bvilhjal

@bvilhjal bvilhjal commented Jul 29, 2026 •

Copy link
Copy Markdown
Owner

Follow-up to a review of organisation and computational efficiency. Every item here was reproduced before being changed; claims that did not survive measurement were dropped rather than implemented (see §8).

8 commits, 26 files.

1. LDpred did not run at all on a current scientific Python stack

from scipy import isfinite raised ImportError, and 230 further call sites used NumPy aliases in the scipy namespace (sp.dot, sp.array, sp.clip, …) that SciPy has since removed. Verified on scipy 1.17.1 / numpy 2.4.6, where all 32 distinct attributes used are gone.

  • Import numpy directly and rename every sp. to np. — all 33 map 1:1, so it is a pure rename. from scipy import stats / linalg are genuine SciPy imports and stay.
  • dtype='bool8' (removed in NumPy 2.0) → dtype=bool.

Two latent bugs on the same path that NumPy 2 turns into hard errors:

  • calc_risk_scores reshaped pval_derived_effects_prs in place to feed the covariate regressions, mutating the array shared with its caller; the score writer then formatted length-1 arrays with %0.6e. The regressions now use a separate column view and the vector stays 1-d.
  • lstsq returns the residual sum of squares as a length-1 array for a column right-hand side, so every derived pred_r2 was an array rather than a float. Squeezed through a small _scalar helper.

2. Long-range LD handling — two bugs, one behaviour change, one new option

The exclusion never fired. load_lrld_dict initialised each chromosome as {'reg_dict': {}} but stored regions as siblings of reg_dict, so is_in_lrld read an always-empty dict and returned False for everything. Sweeping a 1 Mb grid genome-wide flagged 0 positions; chr6:30 Mb (the MHC) came back False. The loop body underneath was dead code that would have raised TypeError on a string key had it ever run.

Fixing that alone would have traded a silent no-op for a KeyError: callers passed a counter over util.chromosomes_list, which runs to 24 entries, while the region table is keyed 1–23. Lookups now key off the group name, and get_lrld_chrom_key resolves 'chrom_X' to 23. get_snp_lrld_status is vectorised over the region list (171× on 500k positions, identical mask).

Skipped variants were reported as zero. Once the exclusion actually ran, a second bug became reachable. A held-out variant stays frozen at whatever it starts with, so the starting value decides both what it contributes to its neighbours' residuals and what weight it is reported with. curr_post_means is only written inside the loop body the skip bypasses, so it kept its initial 0 and the output weight was zero — contradicting the help text and silently dropping ~5% of the genome from the score.

New --long-range-ld {include,inf,exclude}, default inf, replacing the boolean:

value behaviour
inf (default) hold out of the sampler, keep in the model at the LDpred-inf effect, report that weight
exclude hold out with a zero effect, so they leave the model entirely, report a zero weight
include sample them like any other SNP

--incl-long-range-ld still works as a deprecated spelling of --long-range-ld include and warns; combining it with an explicit --long-range-ld is an error rather than a silent precedence rule. The run summary reports which policy was used.

Behaviour change on real data: with ldpred gibbs at the default, variants in the 34 Price et al. (2008) regions are now held out of the sampler and reported with their LDpred-inf estimates. Scores over those regions will differ from previous releases. Recorded in CHANGES.txt; the policy choice is evidenced in §5.

3. Compiled Gibbs sampler (Numba)

Profiling put ldpred_gibbs at 66.7 s cumulative against 18.0 s for everything in ldpred_inf — the sampler is where a genome-wide run spends its time, and it ran as a per-SNP Python loop.

The sweep is now a standalone kernel compiled by Numba, an optional dependency (pip install ldpred[fast]). Without it the sampler uses an equivalent NumPy implementation. Median of 5 runs, m=20000, ld_radius=100, 60 iterations, with the provenance of each build verified:

median vs pre-numba
pre-numba 5.883 s —
Numba off (fallback) 4.822 s 1.22× — no regression
Numba on 0.500 s 11.8×

Both backends are kept deliberately: running the kernel's scalar loop in the interpreter would be ~12× slower than the code it replaces, so the fallback expresses the inner product as one np.dot instead. LDPRED_DISABLE_NUMBA=1 forces it, which is how the two were compared. The active backend is reported in the run summary.

Random draws are still generated outside the sweep, so the random stream is byte-for-byte what the previous implementation consumed. The backends differ only in summation order: largest relative difference 2.0e-16 across the fixed-N, varying-N and long-range-LD paths.

4. Other measured efficiency fixes

Change Speed-up Agreement
calc_auc: O(n²) threshold sweep → rank/Mann-Whitney 504× at n=20000 1e-14 (tie-free)
ldpred_inf: pinv (SVD) → Cholesky solve 3–9.5× 1e-14
get_snp_lrld_status: Python loop → vectorised mask 171× identical
Gibbs inner loop: dropped two tautological isreal() calls per SNP per sweep — bit-identical

calc_auc also fixes a correctness bug: on tied scores the old sweep returned 0.889 where brute force over all pairs gives 0.722.

Scope note: on the integration suite these are worth ~30% end-to-end, but that workload over-weights ldpred_inf. On a genome-wide run with a --f grid, §4 alone is worth ~13%; §3 is what moves the total.

5. Evidence for the long-range LD policy

The exclusion had never actually run before this PR, so there was no evidence for how it should behave — and the goldens cannot supply any, since every simulated test SNP sits below 5000 bp while the earliest real region starts at 25 Mb.

simulations/lrld_policy.py compares five arms; three are reachable from the command line. Mean test R², 10 replicates per cell (simulations/README.md, per-replicate rows in lrld_policy_results.csv):

arm loading 0.0 (control) 0.5 1.0 2.0
include 0.435 0.014 0.029 0.080
zero (previous behaviour, no longer reachable) 0.324 0.328 0.329 0.330
inf (default) 0.444 0.374 0.300 0.288
exclude 0.326 0.333 0.334 0.333
drop (coordination-step only, not exposed) 0.347 0.332 0.336 0.332
  • include does not merely do worse — the sampler diverges, with Σβ² reaching 250–860× the true h², concentrated in the region. That is exactly what LDpred's existing convergence warning reports.
  • Excluding is not free. In the control column — a region on the exclusion list that behaves normally — zero costs −0.111 R², exclude −0.109 and drop −0.087.
  • inf is the only arm that is never bad, which is why it is the default.
  • exclude tracks the full panel rebuild (drop) to within noise at every loading ≥ 0.5 and beats inf there by ~0.03–0.04, which is why it is the one exposed rather than drop. Rebuilding LD on a reduced panel is a coordination-step operation the sampler cannot perform.

Every arm sets its region weights explicitly rather than inheriting the sampler's default, so the arms stay independent of whatever the default happens to be.

Caveat: idealised Gaussian genotypes, one synthetic region, one h², one polygenicity. The relative ordering of inf vs exclude is exactly what could shift on real data, and the default depends on it. The README says plainly this should be repeated on real genotypes with the actual Price regions.

6. Organisation

  • setup.cgf was a typo for setup.cfg, so its only content had never been read by anything. Removed.
  • setup.py → pyproject.toml. Same distribution: version still single-sourced from ldpred.__version__, same three console scripts, same package data. Classifiers no longer advertise Python 2.7.
  • requirements.txt carries lower bounds instead of bare names; new [fast] extra for Numba.
  • CircleCI (pinned to a python:3.7.4 image, running only SimpleTests plus one ComplexTests case) → GitHub Actions matrix over 3.9/3.11/3.13 running the unit tests and both golden suites in full, each a second time with the NumPy fallback forced, plus a wheel/sdist build-and-import smoke job.
  • Added a .gitignore; the repo had none, so 16 __pycache__/*.pyc files were tracked and every test run dirtied the tree.
  • get_LDpred_sample_size existed byte-identically in LDpred_gibbs and LDpred_fast, with the LDpred_fast copy never called. Now lives once in util.

7. Tests

All 16 golden-file tests pass byte-identically on numpy 2.4.6 / scipy 1.17.1, with Numba enabled and with it disabled.

38 unit tests added — the repo had none.

file count covers
tests/test_util.py 12 long-range LD lookup, mask, calc_auc incl. ties
tests/test_sampler.py 19 packing helpers, backend equivalence, both sample-size paths, all three long-range LD policies, deprecated-flag resolution
tests/test_lrld_simulation.py 7 simulation stays runnable and reproduces the divergence result

The goldens are unaffected by the §2 behaviour change because none of their 31000 simulated SNPs falls in a real region — which is precisely why the unit tests exist. All three policies were additionally verified end to end through the CLI.

8. Claims that did not survive verification, and were therefore not implemented

  • Tiled BLAS-3 LD table. Measured 2.0–3.3×, not the order of magnitude expected. It is sharply tile-sensitive (2.26× at tile 128, 0.49× at tile 2048) because a dense GEMM over a banded problem computes (tile+2r)/(2r+1) times more correlations than needed, and it perturbs D at ~6e-07. Decisively, the LD table is built once and cached: on a 7-value --f grid it is 3.0% of runtime, so even a 3.3× is worth ~2% overall.
  • stats.norm.rvs → standard_normal. Only ~1.5× on modern SciPy.

9. Noted, not changed

  • The tests package ships ~43 MB of golden fixtures into site-packages because the ldpred-unittest / ldpred-inttest entry points import it. Dropping those two scripts would let the wheel shed all of it.
  • coord_genotypes.py:532 reads and parses the entire interpolated genetic-map file inside a loop whose body is entirely commented out, so genetic_map always comes back empty. A missing feature rather than a regression.

🤖 Generated with Claude Code

https://claude.ai/code/session_01Fc3Ljp4v5ANQA7bqVXdGqR

claude added 6 commits July 29, 2026 08:37
LDpred failed at import on any current SciPy: `from scipy import isfinite`
raised ImportError, and 230 further call sites used NumPy aliases in the
scipy namespace (sp.dot, sp.array, sp.clip, ...) that SciPy has since
removed. Verified against scipy 1.17.1 / numpy 2.4.6, where all 32 distinct
attributes used are gone.

- Import numpy directly and rename every sp.<attr> use to np.<attr>. All 33
  attributes map 1:1 onto numpy, so this is a pure rename; `from scipy import
  stats` and `from scipy import linalg` are genuine SciPy imports and stay.
- Replace dtype='bool8' (removed in NumPy 2.0) with dtype=bool.
- Drop the now-unused bare `import scipy` from run.py.

Two latent bugs that NumPy 2 turns into hard errors are fixed alongside,
since they sit on the same code path:

- calc_risk_scores reshaped pval_derived_effects_prs in place to feed the
  covariate regressions, mutating the array it shares with its caller. The
  score writer then formatted length-1 arrays with %0.6e, which NumPy 2
  rejects. The regressions now use a separate column view and the vector
  stays 1-d.
- lstsq returns the residual sum of squares as a length-1 array for a column
  right-hand side, so every derived pred_r2 was a length-1 array rather than
  a float. They are squeezed through a small _scalar helper before being
  printed or stored.

Verified: the full golden-file suite (14 SimpleTests + 2 ComplexTests) passes
byte-identically on numpy 2.4.6 / scipy 1.17.1, so the migration preserves
numerics exactly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fc3Ljp4v5ANQA7bqVXdGqR
… shrink

Long-range LD exclusion never fired. load_lrld_dict initialised each
chromosome as {'reg_dict': {}} but then stored regions as siblings of
reg_dict, so is_in_lrld read an always-empty dict and returned False for
every variant. Sweeping a 1 Mb grid across the genome flagged 0 positions;
chr6:30 Mb (the MHC) came back False. The loop body underneath was dead code
that would have raised TypeError on a string key had it ever run.

Fixing that alone would have swapped a silent no-op for a KeyError: callers
passed a counter over util.chromosomes_list, which runs to 24 entries, while
the region table is keyed 1-23. Lookups now key off the group name, and
get_lrld_chrom_key resolves 'chrom_X' to 23 the way coordinated data stores
it. get_snp_lrld_status is also vectorised over the region list rather than
looping in Python per position (171x on 500k positions, same mask).

Note this changes results on real data: variants in long-range LD regions
are now actually skipped when --incl-long-range-ld is off, which is what the
flag always claimed to do.

Also in this change:

- calc_auc computed the AUC by sweeping every observed score as a threshold
  and rescanning the whole prediction vector at each one, which is O(n^2):
  1.7 s at n=20000. It now uses the rank/Mann-Whitney form, which is exact
  and O(n log n) - 504x faster at n=20000, agreeing to 1e-14 on tie-free
  input. It also fixes tie handling: on tied scores the sweep returned 0.889
  where the correct value, confirmed by brute force over all pairs, is 0.722.
- ldpred_inf inverted the shrink matrix with pinv (an SVD) and then
  multiplied. It now solves A x = n*beta_hat directly. A is (m/h2) I + n D,
  symmetric and normally positive definite, so a Cholesky solve applies, with
  fallbacks to the symmetric solver and then pinv because shrinking D can
  cost positive definiteness. 3-9.5x faster, agreeing to 1e-14.
- Removed the isreal() guards in the Gibbs inner loop. Both operands are
  products of real floats with exp() of a real, so the checks were tautologies
  and their else-branches unreachable, at two array calls per SNP per sweep.
- Dropped a shadowed chromosomes_list definition in util.py, and hoisted
  positions out of the `if out_file_prefix` block in ldpred_genomewide, where
  it was read unconditionally a few lines later.

Tests: the 16 golden-file tests still pass byte-identically, and the
integration suite got ~32% faster (222 s -> 149 s for ComplexTests). Those
goldens do not cover the long-range LD path at all - all 31000 simulated test
SNPs sit at positions below 5000 bp - so tests/test_util.py adds 12 unit
tests over the region lookup, the mask, and calc_auc, including the tie case.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fc3Ljp4v5ANQA7bqVXdGqR
The repository had no .gitignore, so 16 __pycache__/*.pyc files were tracked
and every run of the test suite showed up as a dirty tree with binary diffs.
Untrack them and ignore the usual Python build, test and environment output.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fc3Ljp4v5ANQA7bqVXdGqR
Packaging:
- setup.cgf was a typo for setup.cfg, so its only content (the py2/py3
  universal wheel flag, itself obsolete) had never been read by anything.
  Removed.
- Replaced setup.py with pyproject.toml. Same distribution: version still
  single-sourced from ldpred.__version__, same three console scripts, same
  package data. Classifiers no longer advertise Python 2.7, and
  requires-python is >=3.9 to match the numpy floor.
- requirements.txt now carries lower bounds instead of bare package names,
  and says it mirrors pyproject.

CI:
- Replaced the CircleCI config, pinned to a python:3.7.4 image and running
  only SimpleTests plus a single ComplexTests case, with a GitHub Actions
  matrix over 3.9/3.11/3.13 that runs the unit tests and both golden suites
  in full, plus a wheel/sdist build-and-import smoke job. BLAS threads are
  pinned so the golden comparisons stay reproducible.

Deduplication:
- get_LDpred_sample_size existed byte-identically in LDpred_gibbs and
  LDpred_fast, and the LDpred_fast copy was never called. It now lives once
  in util alongside the other shared helpers.
- Dropped an unused `import random` in run.py.

Tests: 12 unit tests and all 16 golden-file tests pass; `pip install .` from
the new pyproject produces the same version, entry points and dependency set,
and the installed `ldpred` CLI runs.

Note for follow-up, not changed here: the tests package ships ~43 MB of
golden fixtures into site-packages because the ldpred-unittest and
ldpred-inttest entry points import it. Dropping those two scripts would let
the wheel shed all of it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fc3Ljp4v5ANQA7bqVXdGqR
The Gibbs sampler is where a genome-wide run spends its time: profiling the
integration suite puts ldpred_gibbs at 66.7 s cumulative against 18.0 s for
everything in ldpred_inf. It ran as a per-SNP Python loop.

The sweep is now a standalone kernel compiled by Numba. Measured on
m=20000, ld_radius=100, 60 iterations (median of 5 runs, provenance of each
build confirmed):

  pre-numba (920cc9f)      5.883 s
  new, Numba disabled      4.822 s     1.22x  (no regression)
  new, Numba enabled       0.500 s    11.8x

Numba is an optional dependency, installed with the new [fast] extra. Without
it the sampler uses _gibbs_sweep_numpy, which is the same arithmetic with the
inner product expressed as one np.dot. Both backends are kept because running
the kernel's scalar loop in the interpreter would be ~12x slower than the code
this replaces; the fallback is instead slightly faster than before.
LDPRED_DISABLE_NUMBA=1 forces the fallback, which is how the two paths were
compared. The active backend is reported in the run summary.

Supporting changes:
- constants_as_arrays flattens prepare_constants' scalars-or-dict-of-dicts
  into plain arrays, so the fixed-N and varying-N paths share one kernel.
- flatten_ld_dict packs the per-SNP LD vectors into one contiguous array,
  deriving each window's left edge exactly as the Python loop did, from
  ld_boundaries when a genetic map was used and ld_radius otherwise. Packing
  happens per chromosome, so peak memory is unaffected at genome scale.
- beta_hats and curr_betas are pinned to float64. The Python loop got this
  implicitly because np.dot upcast a float32 LD vector against float64
  effects; the kernel is explicitly typed, so the dtype should not depend on
  what the caller passed.
- The random draws are still generated outside the sweep, so the random
  stream is byte-for-byte what the previous implementation consumed.

Numerics: the two backends differ only in summation order (scalar loop vs
BLAS). Across the fixed-N, varying-N and long-range-LD paths the largest
relative difference is 2.0e-16, i.e. last-bit floating point. All 16
golden-file tests pass byte-identically with Numba enabled and with it
disabled, and ComplexTests drops from 174 s to 134 s.

tests/test_sampler.py adds 10 tests covering the packing helpers, the
equivalence of the two backends including the skipped-variant path, and both
sample-size paths. CI installs [test,fast] and runs the unit tests and
SimpleTests a second time with the fallback forced, so it cannot rot behind
the compiled path.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fc3Ljp4v5ANQA7bqVXdGqR
The long-range LD exclusion had never actually run before it was fixed, so
there was no evidence for how it should behave -- and the golden tests cannot
supply any, since every simulated test SNP sits below 5000 bp while the
earliest real region starts at 25 Mb.

simulations/lrld_policy.py compares the four possible policies: sampling
those variants normally (--incl-long-range-ld), skipping them and emitting
zero (what the code does), skipping them and emitting the LDpred-inf weight
(what --help claims), and dropping them from the panel entirely. Background
LD is AR(1) along the variant index; the region adds a shared latent factor
so every pair inside it is correlated regardless of distance. --loading 0 is
the control: a region on the exclusion list that behaves normally.

Results over 10 replicates (simulations/README.md, per-replicate rows in
lrld_policy_results.csv):

  mean test R2       loading 0.0    0.5    1.0    2.0
    include                0.435  0.014  0.029  0.080
    zero                   0.324  0.328  0.329  0.330
    inf                    0.444  0.374  0.300  0.288
    drop                   0.347  0.332  0.336  0.332

Three findings:

- include does not merely do worse, the sampler diverges. Sum of squared
  effects reaches 250-860x the true h2, concentrated in the region, which is
  what LDpred's own convergence warning reports.
- Excluding is not free. In the control column zero costs -0.111 R2 and drop
  -0.087 against including, which matters because the Price list is broad and
  may be benign for a given trait or ancestry.
- inf is the only policy safe in both regimes: it never diverges and pays no
  penalty when the region is benign. Under strong long-range LD zero and drop
  edge ahead of it by ~0.03-0.04.

This argues for making the code honour its help text rather than emitting
zero, and possibly for a three-way option instead of a boolean. No behaviour
is changed here -- that is a scientific call, and the README says plainly that
this should be repeated on real genotypes with the actual Price regions first.

tests/test_lrld_simulation.py adds 7 fast tests (3.5 s) so the simulation
stays runnable: the panel has the intended LD structure, replicates are
deterministic, and the divergence result reproduces.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fc3Ljp4v5ANQA7bqVXdGqR
@bvilhjal bvilhjal changed the title Restore modern NumPy/SciPy compatibility, fix the long-range LD filter, and modernise packaging/CI Restore modern NumPy/SciPy compatibility, fix the long-range LD filter, compile the sampler, modernise packaging/CI Jul 29, 2026
claude added 2 commits July 29, 2026 14:36
Variants in long-range LD regions are held out of the Gibbs sweep because the
LD band is badly misspecified there -- sampling them makes the chain diverge,
with the sum of squared effects reaching 250-860x the true h2 in simulation.
Holding them out is right. Reporting them as zero was not.

The skipped variants stayed frozen at their LDpred-inf effect internally, so
they still absorbed LD from their neighbours, but curr_post_means was only
ever written inside the loop body the skip bypassed. It kept its initial 0,
and averaging zeros produced a zero weight in the output. That contradicted
the --incl-long-range-ld help text, which has always said the LDpred-inf
estimates are used for these, and it silently dropped ~5% of the genome from
the score. It went unnoticed because the exclusion never fired at all until
the preceding commit fixed the region lookup.

Their posterior mean is now seeded from the starting effects once, before the
iterations begin. The sweep never writes those entries, so the value persists
and averaging it over the retained iterations returns it unchanged, for any
num_iter / burn_in.

Chosen on the evidence in simulations/lrld_policy.py, which compares the four
possible policies over 10 replicates. Emitting the LDpred-inf weight is the
only policy that is safe in both regimes: it never diverges, and unlike
zeroing it pays no penalty when a region on the exclusion list turns out to
behave normally, where zeroing costs 0.111 R2. Under strong long-range LD,
dropping the variants outright is better by ~0.03-0.04; a three-way
include/inf/drop option would express that trade-off but is not implemented
here.

The simulation now sets each arm's region weights explicitly instead of
inheriting the sampler's default, so the 'zero' arm stays a real comparison
rather than silently becoming a duplicate of 'inf'. Its numbers are unchanged.

--incl-long-range-ld's help text is expanded to say what the default does and
why, and CHANGES.txt records the behaviour change: scores over those regions
will differ from previous releases.

Tests: two new cases pin the output weight to the LDpred-inf estimate and
check it is invariant to num_iter/burn_in. 31 unit tests and all 16
golden-file tests pass on both sampler backends; the goldens are unaffected
because none of their 31000 simulated SNPs falls in a real region.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fc3Ljp4v5ANQA7bqVXdGqR
The boolean --incl-long-range-ld could only say "sample them" or "hold them
out", and holding out has two sensible meanings that the simulation separates
clearly. It becomes a three-way choice:

  include   sample them like any other SNP
  inf       hold out of the sampler, keep in the model at the LDpred-inf
            effect, report that weight                          (default)
  exclude   hold out with a zero effect, so they leave the model
            entirely, and report a zero weight

A held-out variant stays frozen at whatever it starts with, so the starting
value decides both what it contributes to its neighbours' residuals and what
weight it is reported with. That is the whole mechanism: 'inf' seeds the
posterior mean from the LDpred-inf effect, 'exclude' zeroes the effect first.

--incl-long-range-ld still works as a deprecated spelling of
--long-range-ld include, and warns. Combining it with an explicit
--long-range-ld is an error rather than a silent precedence rule.

Why 'exclude' rather than the simulation's 'drop': dropping properly means
re-deriving every retained variant's LD window on the reduced panel, which is
a coordination-step operation, not something the sampler can do. Zeroing the
effect achieves the same thing for the sampler -- a zero contributes nothing
to any neighbour -- and a new simulation arm confirms it tracks a full panel
rebuild to within noise wherever it matters:

  mean test R2     loading 0.0    0.5    1.0    2.0
    include              0.435  0.014  0.029  0.080
    zero                 0.324  0.328  0.329  0.330
    inf                  0.444  0.374  0.300  0.288
    exclude              0.326  0.333  0.334  0.333
    drop                 0.347  0.332  0.336  0.332

'exclude' matches 'drop' at every loading >= 0.5 and beats 'inf' there by
~0.03-0.04. It trails 'drop' only in the control column, which is the regime
where 'inf' is the right choice anyway. That is why 'inf' stays the default:
it is the only arm that is never bad, and a region on the Price list may well
behave normally in a given dataset.

The run summary now reports which policy was used.

Tests: 7 new cases covering the exclude path, the difference it makes to
retained neighbours, rejection of unknown policies, and the deprecated-flag
resolution including the conflict case. 38 unit tests and all 16 golden-file
tests pass on both sampler backends. Verified end to end through the CLI for
all three policies plus the deprecated flag.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fc3Ljp4v5ANQA7bqVXdGqR
@bvilhjal
bvilhjal marked this pull request as ready for review July 29, 2026 22:23
@bvilhjal
bvilhjal merged commit 4b62f46 into master Jul 29, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants