mcsamplerGPU's availability probe is a bare except: around a trivial allocation. It cannot distinguish "cupy is not installed" from "cupy is installed and this host's CUDA stack is mismatched", and it reports both as the former:
junk_to_check_installed = cupy.array(5) # a fill; the cheapest possible kernel
...
except:
print(" no cupy (mcsamplerGPU)")
The driver then runs on numpy having been asked for --gpu. Two sessions hit this from opposite sides on the same week; the two failure modes together are the argument, because they show the probe is not predictive in either direction.
Side 1: silent degradation, at campaign scale
Measured by the paper-1 §V.A rebuild session across its whole ladder-2 campaign (their census, quoted with attribution):
A census over all 171 logs: 162 ran numpy with Override --gpu (not available), zero ran cupy.
Every one of those runs passed --gpu --force-xpy. 162 runs were labelled GPU in their own provenance and were not. That is not a corner case; it is the default outcome on a bare CIT host.
The root cause is a version mismatch, not a missing package, which is what makes the message actively misleading. On bare CIT, cupy 14.1.1 carries nvrtc 12.9. With CUDA_PATH=/usr/local/cuda-13.1 the probe kernel — a fill — compiles, and then every real kernel fails with a dozen errors out of cuda_fp8.hpp; CUDA_PATH=/usr/local/cuda-12.9 works. So even a "compile something" probe passes on a host where the run will die: the guard would have to compile a REDUCTION, not a fill, to be predictive.
Side 2: the same class of breakage escaping the guard entirely
From the PR #223 end-to-end work (this session), ldas-pcdev2, CVMFS IGWN python, cupy 12.0.0 outside a container, real data, --gpu --force-xpy:
CuPy Version : 12.0.0
CUDA Build Version : 11020
CUDA Driver Version: 13030
CUDA Runtime Version: 11020
cuBLAS Version : (available)
" no cupy" appears 0 times — the probe passed. The run then loaded frames and PSDs, printed its parameter table and its SNR guess, entered the integration, and died:
CUBLAS_STATUS_NOT_INITIALIZED
File "cupy/cuda/device.pyx", line 69, in cupy.cuda.device.get_cublas_handle
Command exited with non-zero status 62
after 42 s of wall time and ~1.2 GB of file reads. Same class of defect (cupy built for CUDA 11.2, driver 13.0.3), opposite outcome: not a silent numpy run, a hard failure part-way through real work.
So: one mismatch is swallowed and mislabelled, another sails past the probe and kills the job mid-integration. What the probe currently guarantees is neither "cupy works" nor "cupy is absent".
Why it costs more than the runs
- Provenance. A
.dat from a silently-degraded run is indistinguishable downstream from a GPU one. The §V.A session found this only by writing a check_backend() gate that greps the log and raises when --gpu was requested and numpy ran; it fired on 162 of their own archived runs.
- Diagnosis time.
" no cupy (mcsamplerGPU)" sends you to check the package. Both of us checked the package. The package was fine both times.
- An import-only probe is not evidence.
python -c "import cupy; print(cupy.__version__)" succeeds on every host discussed here, including the two where the run does not work.
Suggested direction (not implemented here)
- Catch narrowly and say what happened.
except ImportError for the absent-package case; let anything else print the actual exception. A one-line repr(exc) would have replaced two multi-probe debugging sessions.
- Make the probe representative. Compile and run a small reduction on the device, not
cupy.array(5). A fill compiles under a mismatched toolkit; a reduction does not.
- Refuse rather than degrade under
--force-xpy. A user who passed --gpu --force-xpy has said the GPU is the point. Falling back to numpy silently is the one behaviour that produces mislabelled science; exiting non-zero produces a rerun.
- Record the backend in the output. The status JSON already carries
sampler_method; a backend: "cupy"|"numpy" field would make every downstream consumer able to check what the §V.A session had to grep for.
Point 3 changes behaviour and is RO'S's call, not ours. Points 1, 2 and 4 are diagnostics.
Related
Census and root cause: paper-1 §V.A rebuild session, 2026-09-02. Second failure mode and the #228 verification: PR #223 end-to-end work, same day; run logs under ~/pr223_e2e/ on CIT and the record in RIFT_roboto_paper analyses/limit_distance_e2e/.
🤖 Generated with Claude Code
mcsamplerGPU's availability probe is a bareexcept:around a trivial allocation. It cannot distinguish "cupy is not installed" from "cupy is installed and this host's CUDA stack is mismatched", and it reports both as the former:The driver then runs on numpy having been asked for
--gpu. Two sessions hit this from opposite sides on the same week; the two failure modes together are the argument, because they show the probe is not predictive in either direction.Side 1: silent degradation, at campaign scale
Measured by the paper-1 §V.A rebuild session across its whole ladder-2 campaign (their census, quoted with attribution):
Every one of those runs passed
--gpu --force-xpy. 162 runs were labelled GPU in their own provenance and were not. That is not a corner case; it is the default outcome on a bare CIT host.The root cause is a version mismatch, not a missing package, which is what makes the message actively misleading. On bare CIT, cupy 14.1.1 carries nvrtc 12.9. With
CUDA_PATH=/usr/local/cuda-13.1the probe kernel — a fill — compiles, and then every real kernel fails with a dozen errors out ofcuda_fp8.hpp;CUDA_PATH=/usr/local/cuda-12.9works. So even a "compile something" probe passes on a host where the run will die: the guard would have to compile a REDUCTION, not a fill, to be predictive.Side 2: the same class of breakage escaping the guard entirely
From the PR #223 end-to-end work (this session),
ldas-pcdev2, CVMFS IGWN python, cupy 12.0.0 outside a container, real data,--gpu --force-xpy:" no cupy"appears 0 times — the probe passed. The run then loaded frames and PSDs, printed its parameter table and its SNR guess, entered the integration, and died:after 42 s of wall time and ~1.2 GB of file reads. Same class of defect (cupy built for CUDA 11.2, driver 13.0.3), opposite outcome: not a silent numpy run, a hard failure part-way through real work.
So: one mismatch is swallowed and mislabelled, another sails past the probe and kills the job mid-integration. What the probe currently guarantees is neither "cupy works" nor "cupy is absent".
Why it costs more than the runs
.datfrom a silently-degraded run is indistinguishable downstream from a GPU one. The §V.A session found this only by writing acheck_backend()gate that greps the log and raises when--gpuwas requested and numpy ran; it fired on 162 of their own archived runs." no cupy (mcsamplerGPU)"sends you to check the package. Both of us checked the package. The package was fine both times.python -c "import cupy; print(cupy.__version__)"succeeds on every host discussed here, including the two where the run does not work.Suggested direction (not implemented here)
except ImportErrorfor the absent-package case; let anything else print the actual exception. A one-linerepr(exc)would have replaced two multi-probe debugging sessions.cupy.array(5). A fill compiles under a mismatched toolkit; a reduction does not.--force-xpy. A user who passed--gpu --force-xpyhas said the GPU is the point. Falling back to numpy silently is the one behaviour that produces mislabelled science; exiting non-zero produces a rerun.sampler_method; abackend: "cupy"|"numpy"field would make every downstream consumer able to check what the §V.A session had to grep for.Point 3 changes behaviour and is RO'S's call, not ours. Points 1, 2 and 4 are diagnostics.
Related
CuPy Version 14.1.1,cupy memory [total, available] 24026.7 23057.4 Mb,Seeding RNGs with N: cupy=seeded, zerono cupylines) rather than assumed. Had this probe been trusted, the control would have been numpy-vs-numpy and the finding wrong.Census and root cause: paper-1 §V.A rebuild session, 2026-09-02. Second failure mode and the #228 verification: PR #223 end-to-end work, same day; run logs under
~/pr223_e2e/on CIT and the record in RIFT_roboto_paperanalyses/limit_distance_e2e/.🤖 Generated with Claude Code