Skip to content

portfolio and adaptive_cartesian could not run with a GPU visible - #368

Merged
oshaughnessy-junior merged 8 commits into
rift_O4dfrom
claude/gpu-hostdevice-portfolio-ac
Sep 18, 2026
Merged

oshaughnessy-junior merged 8 commits into
rift_O4dfrom
claude/gpu-hostdevice-portfolio-ac

Conversation

@oshaughnessy-junior

@oshaughnessy-junior oshaughnessy-junior commented Sep 18, 2026 •

Copy link
Copy Markdown
Owner

Fixes the two host/device defects PR #365 pinned but did not fix. Both raised
TypeError: Unsupported type <class 'numpy.ndarray'> from a cupy ufunc, and both were
invisible because the ILE driver catches it, prints FAILED ANALYSIS and exits 0.

They convert in opposite directions, on purpose:

  • portfolio (AV + AC): mcsamplerGPU.compute_hist puts the aggregator's host samples on
    the member's device. The histogram must end up there anyway, and it is a proposal density,
    not an estimator.
  • adaptive_cartesian: mcsampler.integrate copies the integrand back to the host.
    joint_p_prior is RiftFloat, which cupy has no dtype for, so the host side cannot move.

Both samplers move out of _GPU_KNOWN_FAILING into _GPU_SAMPLERS, with lane coefficients
mirroring their host twins and a re-measured eight-seed calibration table. _E2E_EXPECTED
24 to 26. Note these lanes are not CPU-only: gpu_integration runs them for real.

Measured. ldas-pcdev2, CUDA slot 0 (A100 cc 8.0, probed by running a kernel, since the
slot map had moved), cupy 12.0.0: 26 passed, nothing skipped. ldas-grid, cupy absent: 17
passed, 9 skipped. Each fix was mutation-checked alone: reverting one brings back its own
failure and not the other's.

Adversarial review found two defects, both fixed here. My note explaining why one device
row equals its host twin was wrong, disproved by measurement: that lane runs cos and exp
on the device, and the arms differ by 2e-6. And to_host silently passed a device array
through for any backend exposing asarray/where/clip but not asnumpy, reintroducing
the mix it exists to remove; it now falls back to .get() and otherwise raises.

One gap stays open and is recorded in the file with its measurement:
adaptive_cartesian_gpu reaches the same conversion by another route and has no lane. It
works by hand (z = +1.0 and -0.40); a follow-up adds the lane.

🤖 Generated with Claude Code

oshaughnessy-junior and others added 6 commits September 17, 2026 13:20
F1  The defaults reader returned on the FIRST ast.Dict it walked into, so a second
    construction site with wrong defaults was never looked at.  Pin the construction-site
    count exactly, and check every site's dict.
F3  Normalize both sides of that comparison through ast.unparse instead of deleting every
    space from one side, and refuse an unresolvable dict key with a named assertion rather
    than an AttributeError on k.value.
F4  The drop-the-factor test pins a partition of six parameter TUPLES, not of the eight
    definitions; say so, since the claim belongs to it and _EXPECTED_SIGNATURES together.
    Also assert the calling and dropping sets are disjoint.
F5  make_e2e_calibration: refuse a bad --lane before mkdtemp and build_event, which cost a
    fixture build and leaked a directory on /.
F6  Tag lanes by enumerate rather than LANES.index, a first-match lookup; correct the
    field-count comment above LANES.

F2 is NOT a defect.  xpy.asarray(x, dtype=float) does land on the device AND take object
dtype: cupy refuses object dtype only when no dtype is passed, and the reviewer's suggested
xpy.asarray(np.asarray(x, dtype=float)) would raise on every cupy input.  Measured on cupy
12.0.0 on two hosts and recorded on the line; pinned by a new 5 s test on the dtype keyword.

Each new guard was mutation-checked against the mutation it exists for, with the
pre-change file as the control arm.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The F2 conclusion rested on a probe I ran by hand, and the test I shipped for it covered
only the numpy path; its own docstring said the device half was "an argument in the
comment, not a test".  That is the read-and-trust shape.  Fixed:

New section 4 runs the shipped example and the driver's own stand-in factory against REAL
cupy -- both input kinds (device float64, host object dtype), both lnL conventions, a
fully-sampled signature and one falling back to a scalar default -- and checks the result
is ON the device and equals the numpy arm.  _FakeXpy stays: it runs everywhere, and a fake
can be more permissive than what it stands for, so the two cover each other.

_cupy_or_skip separates the two failures that look alike.  `import cupy` fails on a host
with no CUDA runtime; it SUCCEEDS on the CIT GPU head nodes while the visible device is one
this cupy cannot build a kernel for (pcdev13 slots 0-2, all of pcdev11: Blackwell cc 12.0,
cupy 12.0.0 -> "invalid value for --gpu-architecture").  So it runs a kernel rather than
trusting the import, probes at dispatch because the slot map moves, and says in the skip
reason that a skip is not a pass.

Measured, ldas-pcdev2 slot 0 (A100, cc 8.0) and ldas-pcdev13 slot 3 (RTX 2080 Ti, cc 7.5),
cupy 12.0.0, same HEAD and same md5s as the local tree: 27 passed on both.  Mutation on the
A100: the rewrite the comment warns against, xpy.asarray(np.asarray(x, dtype=float)), now
fails 3 tests instead of being refused by an argument; dropping dtype= fails 3; a factor
that ignores xpy and returns numpy fails 3.

Also corrects the module docstring and the CI comment, which both said this file needs no
GPU.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… that cannot

The whole file pinned CUDA_VISIBLE_DEVICES="" and so said nothing about the device path,
which is the path production takes by default on a GPU node: RIFT binds its array module
at import from whether cupy imports, not from --gpu.  Section 5 runs the prior-only and
closed-form answers with a device visible and asserts the child reached one.

WHAT IT COVERS: the extrinsic sampler on device, the --zero-likelihood stand-in building
its base array with xpy_default=cupy, and the supplementary factor evaluating on device.
NOT the GPU signal likelihood -- --zero-likelihood replaces likelihood_function outright,
so NoLoop never runs.  Measured, not assumed: AV with and without --vectorized --gpu
--force-xpy returns bit-identical lnZ, and that inertness is now its own test.

THE BLOCKER DID NOT REPRODUCE.  "CUDA_VISIBLE_DEVICES='' is required, the scalar AV path
raises Unsupported dtype float128" is true where it was written
(test_psi_marginalization.py) and false here: under --zero-likelihood the stand-in builds
lnL with xpy.zeros, float64 on device, and the scalar likelihood that would have built a
RiftFloat never runs.  AV on an A100 completes and lands 1.5 sigma from the closed form.

TWO SAMPLERS CANNOT RUN ON A DEVICE, both pre-existing, both reproduced with no
supplementary factor, both `TypeError: Unsupported type <class 'numpy.ndarray'>` from a
cupy ufunc, and both invisible in production because the driver prints FAILED ANALYSIS and
exits 0:
  portfolio           RIFT/likelihood/vectorized_general_tools.py histogram(), where
                      xpy.maximum gets a device blank_array and numpy samples
  adaptive_cartesian  RIFT/integrators/mcsampler.py, fval*joint_p_prior/joint_p_s, device
                      integrand against numpy priors
They are RECORDED lanes, pinned by site and exception, not skips: skipping is what hid
them.  The test reddens if one starts working and if the failure moves.

The GMM device lane is A=0.75 B=3, not A=8 B=2.  This file already records that GMM at
A=8 with the inclination term is an n_eff lottery -- 2 of 8 seeds collapsed to sigma 0.028
and 0.073, over MAX_SIGMA -- which is why no CPU lane runs it.  Eight clean GPU seeds do
not refute a lottery; the sweep that found it was also eight seeds.

Calibration, 8 seeds on ldas-pcdev2 slot 0 (A100, cc 8.0), cupy 12.0.0, recorded in the
table with its host twins.  Worst |z| across the whole gate moves 2.15 -> 2.30, so
Z_TOLERANCE's margin is 2.2x rather than 2.3x; no constant changed.

Mutation-checked on the A100: a silent host fallback reddens all six device lanes; a
recorded failure that starts passing reddens; a failure that moves site reddens.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…at can move

Both died in a cupy ufunc handed a numpy array, both reproduce with no supplementary
likelihood factor, and both were invisible in production because the ILE driver catches
the exception, prints FAILED ANALYSIS and exits 0.  Recorded by site and exception in
test_e2e_analytic_pipeline.py section 5 and fixed here.

The two are NOT the same defect and do not convert the same way.

portfolio (AV + AC).  mcsamplerPortfolio pins self.xpy = numpy and aggregates on the host,
so an AC member built on cupy was handed host samples through external_rvs, and
vectorized_general_tools.histogram's xpy.maximum(samples, xpy.zeros(...)) raised.  Fixed in
mcsamplerGPU.compute_hist, which puts the samples and weights on the MEMBER'S backend: the
histogram has to end up on the device anyway (histogram_edges/histogram_cdf are preallocated
there and the draws read them), and nothing here carries precision worth keeping on the host
-- the coordinates are float64 and the histogram is a proposal density, not an estimator.
Fixed at the sampler's own boundary rather than in the shared helper; compute_hist is the
only production caller of histogram(), and the symmetric conversion for member DRAWS is
already at the same boundary in the portfolio.

adaptive_cartesian.  fval*joint_p_prior/joint_p_s with a device integrand and host priors.
Fixed the other way, in mcsampler.integrate: everything below that line is host numpy, and
joint_p_prior is deliberately RiftFloat, which cupy has no dtype for -- so the host side is
the one that cannot move.  The integrand is copied back once, right after the call, through
a new to_host() that mirrors infer_array_module and keeps this file's numpy-only import
list.  A host input is returned unchanged, so the numpy path is untouched.

Casting the ARGUMENT instead -- numpy.asarray(x) on a cupy array -- is what cupy refuses,
and is the rewrite both comments warn against.

MEASURED, ldas-pcdev2, CUDA slot 0 (A100 cc 8.0, identified by running a kernel: the slot
map on that node has moved since it was last written down), cupy 12.0.0.  Before: the two
recorded lanes fail at their recorded sites.  After: both complete and land on the closed
form.  Mutation-checked one fix at a time -- with only mcsamplerGPU reverted, portfolio
fails again at vectorized_general_tools.py and adaptive_cartesian stays fixed; with only
mcsampler reverted, the reverse.  Neither fix is doing the other's work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…es now

They were recorded in _GPU_KNOWN_FAILING, which asserts the failure by site and exception
and goes red when one starts passing.  Both now pass, so the dict, its paragraph and the
test that carried them are gone, and both samplers join _GPU_SAMPLERS with lane
coefficients.

Each device lane mirrors the host lane that sampler already runs, so its row sits on top of
a host row: portfolio takes A=0.75 B=3 (its own host lane with a B term, and NOT A=8 B=2,
which is the recorded GMM n_eff lottery), adaptive_cartesian takes A=8 B=2 with the --n-max
its host lane needed, now named once as _AC_N_MAX and reached from the device lane through
_GPU_LANE_KW.  Every lane keeps a B term, so mis-routing inclination stays detectable.

CALIBRATION, eight seeds (1000-1007) on ldas-pcdev2 CUDA slot 0 (A100 cc 8.0, probed with a
kernel), cupy 12.0.0:

  GPU prior-only, portfolio            2.22   0.0104   2335   (host 1.68/0.0104/2316)
  GPU A=0.75 B=3, portfolio            2.32   0.0228    411   (host 1.54/0.0227/ 392)
  GPU prior-only, adaptive_cartesian   2.39   0.0091   3262   (no host twin)
  GPU A=8    B=2, adaptive_cartesian   1.35   0.0305    250   (host 1.35/0.0305/ 250)

The last row is its host twin exactly, which is expected and now stated in the file as a
tripwire: under --zero-likelihood the only device work is xpy.zeros, and the integrand is
copied to the host before anything else touches it, so both arms run the same host
arithmetic off the same seed.  Drift there means real device arithmetic appeared.

The four AV/GMM device rows were re-derived in the same sweep and came back identical to
the values measured before the two fixes, so the fixes do not move the lanes that already
worked.  Worst |z| over all device lanes is now 2.39, leaving Z_TOLERANCE = 5 at 2.1x;
MAX_SIGMA and MIN_NEFF are unmoved.

Collected count 24 -> 26 in .travis/test-integrate.sh (four lanes added, two recorded
failures removed), and the generator grows the four rows so the table stays re-derivable.

RESULTS.  ldas-pcdev2, slot 0, cupy 12.0.0: 26 passed, nothing skipped.  ldas-grid, cupy
absent: 17 passed, 9 skipped -- every device lane, with "a skip is NOT a pass" in the
reason.  test-ci-roster.py PASS; test-core-units.sh PASS (580 passed, 13 skipped) on
ldas-grid.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@oshaughnessy-junior
oshaughnessy-junior deployed to private-review-dispatch-rift September 18, 2026 12:48 — with GitHub Actions Active
…tion

#365 landed with three review-response commits this branch was not built on, all in the
same files: sections renumbered (5 -> 4), _run_ile/_invoke_ile made keyword-only past
sampler_args, a _no_gpu() that FAILS instead of skipping under RIFT_CI_REQUIRE_GPU=1, skip
guards in test-integrate.sh, and the device rows made droppable in the calibration
generator.  Resolved by taking rift_O4d's version of all four conflicting files outright
and re-applying the promotion on top of it, rather than merging hunk by hunk.

Re-applied against the new text: _GPU_SAMPLERS gains portfolio and adaptive_cartesian,
_GPU_LANE_COEFFS/_GPU_LANE_KW, _AC/_AC_N_MAX named once and reused by the CPU lane and the
generator, _GPU_KNOWN_FAILING and its test deleted, the four new calibration rows, and the
device-lane count 7 -> 9 in _no_gpu's docstring and in both skip-guard rationales.
_E2E_EXPECTED 24 -> 26.

Two of their corrections change what this branch has to say.  The device lanes are NOT
CPU-only: .gitlab-ci.yml's gpu_integration job runs test-integrate.sh with
RIFT_CI_REQUIRE_GPU=1 and CUDA_VISIBLE_DEVICES=0, so the four lanes added here will run in
CI on that runner and any skip there is a failure.  And they had already recorded that
"slot 0" is CUDA's ordering rather than nvidia-smi's on ldas-pcdev2, so the duplicate note
this branch carried is dropped in favour of theirs.

_invoke_ile's docstring justified itself by the known-failing lanes that are now gone; it
says what it is instead.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@oshaughnessy-junior
oshaughnessy-junior deployed to private-review-dispatch-rift September 18, 2026 21:21 — with GitHub Actions Active
…gh, one stated gap

WRONG, AND MINE.  The CALIBRATION note said the GPU A=8 B=2 adaptive_cartesian row equals
its host twin because "the only device work is xpy.zeros", so any drift would mean device
arithmetic had appeared.  The reviewer disproved the mechanism by reading the code and then
measuring it: for a sampler with return_lnL False the --zero-likelihood stand-in builds
xpy.ones(n) * xpy.exp(supp), so cos and exp run on the DEVICE already.  cupy 12.0.0 and
numpy 1.26.4 disagree bitwise on about a fifth of 200000 float64 draws, and at seed 1000
that lane returns 6.646959838539267 on the host against 6.646957839585277 on the device.
The rows agree at the precision the table PRINTS, which is a different and much weaker
claim, and the tripwire built on the wrong one would have sent a reader hunting for device
arithmetic that was there from the start.  Restated with the measurement.

to_host PASSED A DEVICE ARRAY THROUGH for a backend it could not convert.
infer_array_module recognizes any module exposing asarray/where/clip; to_host only
converted when it also exposed asnumpy, and torch is in that gap.  So the one function
written to remove a silent host/device mix would have reintroduced one, one backend later.
Now: asnumpy, else the array's own .get(), else a named TypeError.  Not reachable today --
nothing in the ILE path returns a torch tensor -- which is why it had to be closed by
reading rather than by a test going red.  All four branches checked, including that the
numpy and RiftFloat paths still return the SAME OBJECT.

A STATED GAP, not a fix.  --sampler-method adaptive_cartesian_gpu is the driver's own GPU
default and reaches the same compute_hist conversion by a different route, and section 4
has no lane for it.  Measured by hand on ldas-pcdev2 slot 0 (A100), cupy 12.0.0: it works
and lands on the closed form, z = +1.0 prior-only and -0.40 at A=8 B=2.  Recorded in the
file with the numbers and labelled a to-do rather than coverage, because a hand measurement
is not a lane.

Three further findings -- the calibration generator refusing --lane portfolio outright,
_invoke_ile's docstring still citing the deleted known-failing lanes, and the host
adaptive_cartesian row left un-centralized while its device twin used _AC/_AC_N_MAX -- were
already resolved by the rift_O4d merge in the previous commit.

The reviewer also ran the base-arm control this branch's claim rests on: at 02d66cc with
the new test file, the four promoted lanes give "4 failed ... Unsupported type
<class 'numpy.ndarray'>"; at the fix, 10 passed.  And confirmed the CPU path is untouched
by byte-comparing five CPU lanes' .dat and status JSON across the two trees.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@oshaughnessy-junior
oshaughnessy-junior deployed to private-review-dispatch-rift September 18, 2026 21:35 — with GitHub Actions Active
@oshaughnessy-junior
oshaughnessy-junior marked this pull request as ready for review September 18, 2026 22:11
@oshaughnessy-junior
oshaughnessy-junior merged commit f7f986a into rift_O4d Sep 18, 2026
35 checks passed
@oshaughnessy-junior
oshaughnessy-junior deployed to private-review-dispatch-rift September 18, 2026 22:12 — with GitHub Actions Active
oshaughnessy-junior added a commit that referenced this pull request Sep 18, 2026
…or the worst row

Closes the two coverage gaps section 4 recorded rather than fixed, to keep #368
narrow.

adaptive_cartesian_gpu is mcsamplerGPU standalone and the ILE driver's own
default. It reaches compute_hist through its own integrate adaptation block, not
through a portfolio member, and had no lane. It takes _AC_N_MAX, measured: at
--n-max 20000 with the factor on it stops on the budget at n_eff 230.

The device prior-only adaptive_cartesian row held the worst |z| in the
calibration table and had no host twin. It has one now, and the twin returns the
same three columns.

Device table re-derived at 8 seeds on ldas-pcdev12 CUDA slot 0; all eight
pre-existing rows identical to the ldas-pcdev2 numbers. Worst |z| is now 2.54.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

This branch was successfully deployed

1 active deployment
private-review-dispatch-rift — 6799e53d Deployed Sep 18, 2026 by oshaughnessy-junior via Dispatch exact RIFT PR generation #1412
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant