Skip to content

mcsamplerGPU reports a mismatched CUDA stack as "no cupy" and silently runs numpy under --gpu --force-xpy (162 of 171 runs in one campaign) #229

Description

@oshaughnessy-junior

mcsamplerGPU's availability probe is a bare except: around a trivial allocation. It cannot distinguish "cupy is not installed" from "cupy is installed and this host's CUDA stack is mismatched", and it reports both as the former:

junk_to_check_installed = cupy.array(5)   # a fill; the cheapest possible kernel
...
except:
    print(" no cupy (mcsamplerGPU)")

The driver then runs on numpy having been asked for --gpu. Two sessions hit this from opposite sides on the same week; the two failure modes together are the argument, because they show the probe is not predictive in either direction.

Side 1: silent degradation, at campaign scale

Measured by the paper-1 §V.A rebuild session across its whole ladder-2 campaign (their census, quoted with attribution):

A census over all 171 logs: 162 ran numpy with Override --gpu (not available), zero ran cupy.

Every one of those runs passed --gpu --force-xpy. 162 runs were labelled GPU in their own provenance and were not. That is not a corner case; it is the default outcome on a bare CIT host.

The root cause is a version mismatch, not a missing package, which is what makes the message actively misleading. On bare CIT, cupy 14.1.1 carries nvrtc 12.9. With CUDA_PATH=/usr/local/cuda-13.1 the probe kernel — a fill — compiles, and then every real kernel fails with a dozen errors out of cuda_fp8.hpp; CUDA_PATH=/usr/local/cuda-12.9 works. So even a "compile something" probe passes on a host where the run will die: the guard would have to compile a REDUCTION, not a fill, to be predictive.

Side 2: the same class of breakage escaping the guard entirely

From the PR #223 end-to-end work (this session), ldas-pcdev2, CVMFS IGWN python, cupy 12.0.0 outside a container, real data, --gpu --force-xpy:

CuPy Version       : 12.0.0
CUDA Build Version : 11020
CUDA Driver Version: 13030
CUDA Runtime Version: 11020
cuBLAS Version     : (available)

" no cupy" appears 0 times — the probe passed. The run then loaded frames and PSDs, printed its parameter table and its SNR guess, entered the integration, and died:

CUBLAS_STATUS_NOT_INITIALIZED
  File "cupy/cuda/device.pyx", line 69, in cupy.cuda.device.get_cublas_handle
Command exited with non-zero status 62

after 42 s of wall time and ~1.2 GB of file reads. Same class of defect (cupy built for CUDA 11.2, driver 13.0.3), opposite outcome: not a silent numpy run, a hard failure part-way through real work.

So: one mismatch is swallowed and mislabelled, another sails past the probe and kills the job mid-integration. What the probe currently guarantees is neither "cupy works" nor "cupy is absent".

Why it costs more than the runs

  • Provenance. A .dat from a silently-degraded run is indistinguishable downstream from a GPU one. The §V.A session found this only by writing a check_backend() gate that greps the log and raises when --gpu was requested and numpy ran; it fired on 162 of their own archived runs.
  • Diagnosis time. " no cupy (mcsamplerGPU)" sends you to check the package. Both of us checked the package. The package was fine both times.
  • An import-only probe is not evidence. python -c "import cupy; print(cupy.__version__)" succeeds on every host discussed here, including the two where the run does not work.

Suggested direction (not implemented here)

  1. Catch narrowly and say what happened. except ImportError for the absent-package case; let anything else print the actual exception. A one-line repr(exc) would have replaced two multi-probe debugging sessions.
  2. Make the probe representative. Compile and run a small reduction on the device, not cupy.array(5). A fill compiles under a mismatched toolkit; a reduction does not.
  3. Refuse rather than degrade under --force-xpy. A user who passed --gpu --force-xpy has said the GPU is the point. Falling back to numpy silently is the one behaviour that produces mislabelled science; exiting non-zero produces a rerun.
  4. Record the backend in the output. The status JSON already carries sampler_method; a backend: "cupy"|"numpy" field would make every downstream consumer able to check what the §V.A session had to grep for.

Point 3 changes behaviour and is RO'S's call, not ours. Points 1, 2 and 4 are diagnostics.

Related

Census and root cause: paper-1 §V.A rebuild session, 2026-09-02. Second failure mode and the #228 verification: PR #223 end-to-end work, same day; run logs under ~/pr223_e2e/ on CIT and the record in RIFT_roboto_paper analyses/limit_distance_e2e/.

🤖 Generated with Claude Code

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions