Symptom
--mode flowmc-phipsimarg --angle-marg-scheme laplace dies with:
jax.errors.JaxRuntimeError: RESOURCE_EXHAUSTED: Out of memory while trying to allocate 833.83GiB
on a 24 GiB card. It happens after [jax_ile] compiling JAX kernel … done, during [flowMC] sampling (inv_T=0.02) — not in the pilot. Reproduced on four independent condor jobs at SNR 40 and 80.
Cause
angle_marg_eval_chunk bounds the anglemarg schemes' per-sample buffer, and it works — but it is called from exactly three places, all of them batched-eval helpers:
samplers.py:364 eval_lnL
samplers.py:1084 eval_lnL_4
samplers.py:1181 eval_lnL_3
The flowMC path does not go through them. It builds a logpdf closure and hands it to flowMC's Sampler, which calls it at whatever batch its own configuration implies (chains × local steps, vmapped). Nothing on that route consults the cap, so the accurate schemes' large per-sample-point buffer is multiplied by an unbounded sample axis.
The August OOM that motivated the cap (c5b81dd61, 36.41 GiB) was on the eval path, so the fix landed there and the sampling path was never covered.
Why it matters
This is not a tuning problem: it makes exact, laplace and peak-local unusable under flowMC at any amplitude, which is the mode the differentiable driver defaults to for phase- and polarization-marginalized sampling. grid survives only because it is exempt from the cap by having a small enough buffer to begin with.
It also means a larger buffer target does not help — see #250, which makes the cap device-aware. A bigger allowance on an uncapped path changes nothing.
Suggested direction
Either chunk the flowMC logpdf the way eval_lnL chunks its input, or refuse the combination up front with a message naming the cap — the second is cheap and would have turned four silent 6-hour failures into an immediate, readable error.
Found while re-measuring the flowMC negative result in ap:flow of the RIFT_O4d paper set; that measurement is blocked on this.
Symptom
--mode flowmc-phipsimarg --angle-marg-scheme laplacedies with:on a 24 GiB card. It happens after
[jax_ile] compiling JAX kernel … done, during[flowMC] sampling (inv_T=0.02)— not in the pilot. Reproduced on four independent condor jobs at SNR 40 and 80.Cause
angle_marg_eval_chunkbounds the anglemarg schemes' per-sample buffer, and it works — but it is called from exactly three places, all of them batched-eval helpers:The flowMC path does not go through them. It builds a
logpdfclosure and hands it to flowMC'sSampler, which calls it at whatever batch its own configuration implies (chains × local steps, vmapped). Nothing on that route consults the cap, so the accurate schemes' large per-sample-point buffer is multiplied by an unbounded sample axis.The August OOM that motivated the cap (
c5b81dd61, 36.41 GiB) was on the eval path, so the fix landed there and the sampling path was never covered.Why it matters
This is not a tuning problem: it makes
exact,laplaceandpeak-localunusable under flowMC at any amplitude, which is the mode the differentiable driver defaults to for phase- and polarization-marginalized sampling.gridsurvives only because it is exempt from the cap by having a small enough buffer to begin with.It also means a larger buffer target does not help — see #250, which makes the cap device-aware. A bigger allowance on an uncapped path changes nothing.
Suggested direction
Either chunk the flowMC
logpdfthe wayeval_lnLchunks its input, or refuse the combination up front with a message naming the cap — the second is cheap and would have turned four silent 6-hour failures into an immediate, readable error.Found while re-measuring the flowMC negative result in
ap:flowof the RIFT_O4d paper set; that measurement is blocked on this.