portfolio and adaptive_cartesian could not run with a GPU visible - #368
Merged
oshaughnessy-junior merged 8 commits intoSep 18, 2026
Merged
Conversation
F1 The defaults reader returned on the FIRST ast.Dict it walked into, so a second
construction site with wrong defaults was never looked at. Pin the construction-site
count exactly, and check every site's dict.
F3 Normalize both sides of that comparison through ast.unparse instead of deleting every
space from one side, and refuse an unresolvable dict key with a named assertion rather
than an AttributeError on k.value.
F4 The drop-the-factor test pins a partition of six parameter TUPLES, not of the eight
definitions; say so, since the claim belongs to it and _EXPECTED_SIGNATURES together.
Also assert the calling and dropping sets are disjoint.
F5 make_e2e_calibration: refuse a bad --lane before mkdtemp and build_event, which cost a
fixture build and leaked a directory on /.
F6 Tag lanes by enumerate rather than LANES.index, a first-match lookup; correct the
field-count comment above LANES.
F2 is NOT a defect. xpy.asarray(x, dtype=float) does land on the device AND take object
dtype: cupy refuses object dtype only when no dtype is passed, and the reviewer's suggested
xpy.asarray(np.asarray(x, dtype=float)) would raise on every cupy input. Measured on cupy
12.0.0 on two hosts and recorded on the line; pinned by a new 5 s test on the dtype keyword.
Each new guard was mutation-checked against the mutation it exists for, with the
pre-change file as the control arm.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The F2 conclusion rested on a probe I ran by hand, and the test I shipped for it covered only the numpy path; its own docstring said the device half was "an argument in the comment, not a test". That is the read-and-trust shape. Fixed: New section 4 runs the shipped example and the driver's own stand-in factory against REAL cupy -- both input kinds (device float64, host object dtype), both lnL conventions, a fully-sampled signature and one falling back to a scalar default -- and checks the result is ON the device and equals the numpy arm. _FakeXpy stays: it runs everywhere, and a fake can be more permissive than what it stands for, so the two cover each other. _cupy_or_skip separates the two failures that look alike. `import cupy` fails on a host with no CUDA runtime; it SUCCEEDS on the CIT GPU head nodes while the visible device is one this cupy cannot build a kernel for (pcdev13 slots 0-2, all of pcdev11: Blackwell cc 12.0, cupy 12.0.0 -> "invalid value for --gpu-architecture"). So it runs a kernel rather than trusting the import, probes at dispatch because the slot map moves, and says in the skip reason that a skip is not a pass. Measured, ldas-pcdev2 slot 0 (A100, cc 8.0) and ldas-pcdev13 slot 3 (RTX 2080 Ti, cc 7.5), cupy 12.0.0, same HEAD and same md5s as the local tree: 27 passed on both. Mutation on the A100: the rewrite the comment warns against, xpy.asarray(np.asarray(x, dtype=float)), now fails 3 tests instead of being refused by an argument; dropping dtype= fails 3; a factor that ignores xpy and returns numpy fails 3. Also corrects the module docstring and the CI comment, which both said this file needs no GPU. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… that cannot
The whole file pinned CUDA_VISIBLE_DEVICES="" and so said nothing about the device path,
which is the path production takes by default on a GPU node: RIFT binds its array module
at import from whether cupy imports, not from --gpu. Section 5 runs the prior-only and
closed-form answers with a device visible and asserts the child reached one.
WHAT IT COVERS: the extrinsic sampler on device, the --zero-likelihood stand-in building
its base array with xpy_default=cupy, and the supplementary factor evaluating on device.
NOT the GPU signal likelihood -- --zero-likelihood replaces likelihood_function outright,
so NoLoop never runs. Measured, not assumed: AV with and without --vectorized --gpu
--force-xpy returns bit-identical lnZ, and that inertness is now its own test.
THE BLOCKER DID NOT REPRODUCE. "CUDA_VISIBLE_DEVICES='' is required, the scalar AV path
raises Unsupported dtype float128" is true where it was written
(test_psi_marginalization.py) and false here: under --zero-likelihood the stand-in builds
lnL with xpy.zeros, float64 on device, and the scalar likelihood that would have built a
RiftFloat never runs. AV on an A100 completes and lands 1.5 sigma from the closed form.
TWO SAMPLERS CANNOT RUN ON A DEVICE, both pre-existing, both reproduced with no
supplementary factor, both `TypeError: Unsupported type <class 'numpy.ndarray'>` from a
cupy ufunc, and both invisible in production because the driver prints FAILED ANALYSIS and
exits 0:
portfolio RIFT/likelihood/vectorized_general_tools.py histogram(), where
xpy.maximum gets a device blank_array and numpy samples
adaptive_cartesian RIFT/integrators/mcsampler.py, fval*joint_p_prior/joint_p_s, device
integrand against numpy priors
They are RECORDED lanes, pinned by site and exception, not skips: skipping is what hid
them. The test reddens if one starts working and if the failure moves.
The GMM device lane is A=0.75 B=3, not A=8 B=2. This file already records that GMM at
A=8 with the inclination term is an n_eff lottery -- 2 of 8 seeds collapsed to sigma 0.028
and 0.073, over MAX_SIGMA -- which is why no CPU lane runs it. Eight clean GPU seeds do
not refute a lottery; the sweep that found it was also eight seeds.
Calibration, 8 seeds on ldas-pcdev2 slot 0 (A100, cc 8.0), cupy 12.0.0, recorded in the
table with its host twins. Worst |z| across the whole gate moves 2.15 -> 2.30, so
Z_TOLERANCE's margin is 2.2x rather than 2.3x; no constant changed.
Mutation-checked on the A100: a silent host fallback reddens all six device lanes; a
recorded failure that starts passing reddens; a failure that moves site reddens.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…pu-hostdevice-portfolio-ac
…at can move Both died in a cupy ufunc handed a numpy array, both reproduce with no supplementary likelihood factor, and both were invisible in production because the ILE driver catches the exception, prints FAILED ANALYSIS and exits 0. Recorded by site and exception in test_e2e_analytic_pipeline.py section 5 and fixed here. The two are NOT the same defect and do not convert the same way. portfolio (AV + AC). mcsamplerPortfolio pins self.xpy = numpy and aggregates on the host, so an AC member built on cupy was handed host samples through external_rvs, and vectorized_general_tools.histogram's xpy.maximum(samples, xpy.zeros(...)) raised. Fixed in mcsamplerGPU.compute_hist, which puts the samples and weights on the MEMBER'S backend: the histogram has to end up on the device anyway (histogram_edges/histogram_cdf are preallocated there and the draws read them), and nothing here carries precision worth keeping on the host -- the coordinates are float64 and the histogram is a proposal density, not an estimator. Fixed at the sampler's own boundary rather than in the shared helper; compute_hist is the only production caller of histogram(), and the symmetric conversion for member DRAWS is already at the same boundary in the portfolio. adaptive_cartesian. fval*joint_p_prior/joint_p_s with a device integrand and host priors. Fixed the other way, in mcsampler.integrate: everything below that line is host numpy, and joint_p_prior is deliberately RiftFloat, which cupy has no dtype for -- so the host side is the one that cannot move. The integrand is copied back once, right after the call, through a new to_host() that mirrors infer_array_module and keeps this file's numpy-only import list. A host input is returned unchanged, so the numpy path is untouched. Casting the ARGUMENT instead -- numpy.asarray(x) on a cupy array -- is what cupy refuses, and is the rewrite both comments warn against. MEASURED, ldas-pcdev2, CUDA slot 0 (A100 cc 8.0, identified by running a kernel: the slot map on that node has moved since it was last written down), cupy 12.0.0. Before: the two recorded lanes fail at their recorded sites. After: both complete and land on the closed form. Mutation-checked one fix at a time -- with only mcsamplerGPU reverted, portfolio fails again at vectorized_general_tools.py and adaptive_cartesian stays fixed; with only mcsampler reverted, the reverse. Neither fix is doing the other's work. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…es now They were recorded in _GPU_KNOWN_FAILING, which asserts the failure by site and exception and goes red when one starts passing. Both now pass, so the dict, its paragraph and the test that carried them are gone, and both samplers join _GPU_SAMPLERS with lane coefficients. Each device lane mirrors the host lane that sampler already runs, so its row sits on top of a host row: portfolio takes A=0.75 B=3 (its own host lane with a B term, and NOT A=8 B=2, which is the recorded GMM n_eff lottery), adaptive_cartesian takes A=8 B=2 with the --n-max its host lane needed, now named once as _AC_N_MAX and reached from the device lane through _GPU_LANE_KW. Every lane keeps a B term, so mis-routing inclination stays detectable. CALIBRATION, eight seeds (1000-1007) on ldas-pcdev2 CUDA slot 0 (A100 cc 8.0, probed with a kernel), cupy 12.0.0: GPU prior-only, portfolio 2.22 0.0104 2335 (host 1.68/0.0104/2316) GPU A=0.75 B=3, portfolio 2.32 0.0228 411 (host 1.54/0.0227/ 392) GPU prior-only, adaptive_cartesian 2.39 0.0091 3262 (no host twin) GPU A=8 B=2, adaptive_cartesian 1.35 0.0305 250 (host 1.35/0.0305/ 250) The last row is its host twin exactly, which is expected and now stated in the file as a tripwire: under --zero-likelihood the only device work is xpy.zeros, and the integrand is copied to the host before anything else touches it, so both arms run the same host arithmetic off the same seed. Drift there means real device arithmetic appeared. The four AV/GMM device rows were re-derived in the same sweep and came back identical to the values measured before the two fixes, so the fixes do not move the lanes that already worked. Worst |z| over all device lanes is now 2.39, leaving Z_TOLERANCE = 5 at 2.1x; MAX_SIGMA and MIN_NEFF are unmoved. Collected count 24 -> 26 in .travis/test-integrate.sh (four lanes added, two recorded failures removed), and the generator grows the four rows so the table stays re-derivable. RESULTS. ldas-pcdev2, slot 0, cupy 12.0.0: 26 passed, nothing skipped. ldas-grid, cupy absent: 17 passed, 9 skipped -- every device lane, with "a skip is NOT a pass" in the reason. test-ci-roster.py PASS; test-core-units.sh PASS (580 passed, 13 skipped) on ldas-grid. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
oshaughnessy-junior
deployed
to
private-review-dispatch-rift
September 18, 2026 12:48 — with
GitHub Actions
Active
…tion #365 landed with three review-response commits this branch was not built on, all in the same files: sections renumbered (5 -> 4), _run_ile/_invoke_ile made keyword-only past sampler_args, a _no_gpu() that FAILS instead of skipping under RIFT_CI_REQUIRE_GPU=1, skip guards in test-integrate.sh, and the device rows made droppable in the calibration generator. Resolved by taking rift_O4d's version of all four conflicting files outright and re-applying the promotion on top of it, rather than merging hunk by hunk. Re-applied against the new text: _GPU_SAMPLERS gains portfolio and adaptive_cartesian, _GPU_LANE_COEFFS/_GPU_LANE_KW, _AC/_AC_N_MAX named once and reused by the CPU lane and the generator, _GPU_KNOWN_FAILING and its test deleted, the four new calibration rows, and the device-lane count 7 -> 9 in _no_gpu's docstring and in both skip-guard rationales. _E2E_EXPECTED 24 -> 26. Two of their corrections change what this branch has to say. The device lanes are NOT CPU-only: .gitlab-ci.yml's gpu_integration job runs test-integrate.sh with RIFT_CI_REQUIRE_GPU=1 and CUDA_VISIBLE_DEVICES=0, so the four lanes added here will run in CI on that runner and any skip there is a failure. And they had already recorded that "slot 0" is CUDA's ordering rather than nvidia-smi's on ldas-pcdev2, so the duplicate note this branch carried is dropped in favour of theirs. _invoke_ile's docstring justified itself by the known-failing lanes that are now gone; it says what it is instead. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
oshaughnessy-junior
deployed
to
private-review-dispatch-rift
September 18, 2026 21:21 — with
GitHub Actions
Active
…gh, one stated gap WRONG, AND MINE. The CALIBRATION note said the GPU A=8 B=2 adaptive_cartesian row equals its host twin because "the only device work is xpy.zeros", so any drift would mean device arithmetic had appeared. The reviewer disproved the mechanism by reading the code and then measuring it: for a sampler with return_lnL False the --zero-likelihood stand-in builds xpy.ones(n) * xpy.exp(supp), so cos and exp run on the DEVICE already. cupy 12.0.0 and numpy 1.26.4 disagree bitwise on about a fifth of 200000 float64 draws, and at seed 1000 that lane returns 6.646959838539267 on the host against 6.646957839585277 on the device. The rows agree at the precision the table PRINTS, which is a different and much weaker claim, and the tripwire built on the wrong one would have sent a reader hunting for device arithmetic that was there from the start. Restated with the measurement. to_host PASSED A DEVICE ARRAY THROUGH for a backend it could not convert. infer_array_module recognizes any module exposing asarray/where/clip; to_host only converted when it also exposed asnumpy, and torch is in that gap. So the one function written to remove a silent host/device mix would have reintroduced one, one backend later. Now: asnumpy, else the array's own .get(), else a named TypeError. Not reachable today -- nothing in the ILE path returns a torch tensor -- which is why it had to be closed by reading rather than by a test going red. All four branches checked, including that the numpy and RiftFloat paths still return the SAME OBJECT. A STATED GAP, not a fix. --sampler-method adaptive_cartesian_gpu is the driver's own GPU default and reaches the same compute_hist conversion by a different route, and section 4 has no lane for it. Measured by hand on ldas-pcdev2 slot 0 (A100), cupy 12.0.0: it works and lands on the closed form, z = +1.0 prior-only and -0.40 at A=8 B=2. Recorded in the file with the numbers and labelled a to-do rather than coverage, because a hand measurement is not a lane. Three further findings -- the calibration generator refusing --lane portfolio outright, _invoke_ile's docstring still citing the deleted known-failing lanes, and the host adaptive_cartesian row left un-centralized while its device twin used _AC/_AC_N_MAX -- were already resolved by the rift_O4d merge in the previous commit. The reviewer also ran the base-arm control this branch's claim rests on: at 02d66cc with the new test file, the four promoted lanes give "4 failed ... Unsupported type <class 'numpy.ndarray'>"; at the fix, 10 passed. And confirmed the CPU path is untouched by byte-comparing five CPU lanes' .dat and status JSON across the two trees. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
oshaughnessy-junior
deployed
to
private-review-dispatch-rift
September 18, 2026 21:35 — with
GitHub Actions
Active
oshaughnessy-junior
marked this pull request as ready for review
September 18, 2026 22:11
oshaughnessy-junior
had a problem deploying
to
private-review-dispatch-rift
September 18, 2026 22:11 — with
GitHub Actions
Error
oshaughnessy-junior
deployed
to
private-review-dispatch-rift
September 18, 2026 22:12 — with
GitHub Actions
Active
oshaughnessy-junior
added a commit
that referenced
this pull request
Sep 18, 2026
…or the worst row Closes the two coverage gaps section 4 recorded rather than fixed, to keep #368 narrow. adaptive_cartesian_gpu is mcsamplerGPU standalone and the ILE driver's own default. It reaches compute_hist through its own integrate adaptation block, not through a portfolio member, and had no lane. It takes _AC_N_MAX, measured: at --n-max 20000 with the factor on it stops on the budget at n_eff 230. The device prior-only adaptive_cartesian row held the worst |z| in the calibration table and had no host twin. It has one now, and the twin returns the same three columns. Device table re-derived at 8 seeds on ldas-pcdev12 CUDA slot 0; all eight pre-existing rows identical to the ldas-pcdev2 numbers. Worst |z| is now 2.54. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes the two host/device defects PR #365 pinned but did not fix. Both raised
TypeError: Unsupported type <class 'numpy.ndarray'>from a cupy ufunc, and both wereinvisible because the ILE driver catches it, prints FAILED ANALYSIS and exits 0.
They convert in opposite directions, on purpose:
mcsamplerGPU.compute_histputs the aggregator's host samples onthe member's device. The histogram must end up there anyway, and it is a proposal density,
not an estimator.
mcsampler.integratecopies the integrand back to the host.joint_p_priorisRiftFloat, which cupy has no dtype for, so the host side cannot move.Both samplers move out of
_GPU_KNOWN_FAILINGinto_GPU_SAMPLERS, with lane coefficientsmirroring their host twins and a re-measured eight-seed calibration table.
_E2E_EXPECTED24 to 26. Note these lanes are not CPU-only:
gpu_integrationruns them for real.Measured. ldas-pcdev2, CUDA slot 0 (A100 cc 8.0, probed by running a kernel, since the
slot map had moved), cupy 12.0.0: 26 passed, nothing skipped. ldas-grid, cupy absent: 17
passed, 9 skipped. Each fix was mutation-checked alone: reverting one brings back its own
failure and not the other's.
Adversarial review found two defects, both fixed here. My note explaining why one device
row equals its host twin was wrong, disproved by measurement: that lane runs
cosandexpon the device, and the arms differ by 2e-6. And
to_hostsilently passed a device arraythrough for any backend exposing
asarray/where/clipbut notasnumpy, reintroducingthe mix it exists to remove; it now falls back to
.get()and otherwise raises.One gap stays open and is recorded in the file with its measurement:
adaptive_cartesian_gpureaches the same conversion by another route and has no lane. Itworks by hand (z = +1.0 and -0.40); a follow-up adds the lane.
🤖 Generated with Claude Code