GPU parity legs: skip on a driverless host instead of failing - #370
Merged
Merged
Conversation
oshaughnessy-junior
deployed
to
private-review-dispatch-rift
September 18, 2026 23:39 — with
GitHub Actions
Active
`pytest.importorskip('cupy')` skips only on ModuleNotFoundError and re-raises
every other ImportError. The IGWN CVMFS environment installs the cupy package
on every CIT host, so on one with no CUDA driver the import reaches
`libcuda.so.1: cannot open shared object file` -- a plain ImportError -- and the
guard becomes a test FAILURE. Measured on ldas-grid at f7f986a (CVMFS igwn
python 3.11, cupy 12.0.0, pytest 9.1.1): both legs below failed, and read as a
broken GPU path. GitHub's runners have no cupy at all, get the clean
ModuleNotFoundError, and skip, so CI cannot see this.
Both now call gpu_slot_probe.cupy_or_skip(), a new sibling helper carrying the
approach test_e2e_analytic_pipeline.py settled on: a SUBPROCESS that imports
cupy and runs a kernel on each visible slot, and a one-word verdict line. That
also separates "no cupy here" from "cupy here but the visible device is one it
cannot compile for", which is the Blackwell case. The e2e gate's own copy is
left alone; it re-indexes slots for grandchildren and this one must not.
RIFT_CI_REQUIRE_GPU=1 keeps its meaning: the environment promised a device, so
a skip stays a failure.
Verified, two tests per arm:
ldas-grid (cupy installed, no driver) 2 skipped
ldas-grid, RIFT_CI_REQUIRE_GPU=1 2 failed
ldas-pcdev2, A100 at CUDA slot 0 2 passed
ldas-pcdev2, CUDA_VISIBLE_DEVICES=1 2 skipped (NOSLOT CompileException)
same, RIFT_CI_REQUIRE_GPU=1 2 failed
Slot 1 there is a Blackwell cc 12.0 that cupy 12.0.0 cannot build for;
importorskip imports it successfully and the test body then dies. CUDA's
ordering is not nvidia-smi's on that host -- the A100 is nvidia-smi 2 and CUDA
0 -- which is why the probe runs a kernel per slot rather than trusting an
index.
Collection counts are unchanged (peak-local 121, q-window-stencil 78/75/3);
.travis/test-q-window-stencil.sh and .travis/test-ci-roster.py pass on ldas-grid.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
oshaughnessy-junior
force-pushed
the
claude/gpu-guard-skip-not-fail
branch
from
September 19, 2026 00:03
e84e106 to
9195b2f
Compare
oshaughnessy-junior
deployed
to
private-review-dispatch-rift
September 19, 2026 00:03 — with
GitHub Actions
Active
…VMFS env
Six findings from an independent reviewer, all reproduced before acting.
THE PROVENANCE WAS WRONG, and it was the headline claim. Same interpreter on
ldas-grid, one test that does `pytest.importorskip('cupy')`:
pytest 9.1.1 from ~/.local/lib/python3.11/site-packages 1 failed
pytest 8.3.5 CVMFS's own, via PYTHONNOUSERSITE=1 1 skipped
8.3.5 defaults exc_type=ImportError and warns that it will stop doing so. The
user site-packages on the NFS home shadows CVMFS, so an ordinary interactive run
here gets 9.1.1 and does fail, but a container or a condor node gets 8.3.5 and
skips. So the claim is "under pytest >= 9.1", not "on CIT with CVMFS", and the
file and the q-window comment both said the latter.
Device(d).use() IS NOT ENOUGH when the FIRST visible slot is unusable. RIFT's
`try: import cupy` probes run at import, on the default device;
SphericalHarmonics_gpu runs cupy.array(5) there, and on a Blackwell it raises,
sets cupy_here=False and keys _coeffs by numpy alone. A later .use() cannot
move that. Measured on ldas-pcdev2 with CUDA_VISIBLE_DEVICES="1,0" and "1,2".
The two callers differ, so the new check is OPT-IN rather than blanket:
peak-local takes xpy as an argument, never touches SphericalHarmonics_gpu,
and PASSES at those pins. A blanket check would cost this lane.
noloop goes through the likelihood and died on KeyError <module cupy>.
It now skips, naming both flags and the remedy.
Also: no_gpu's claim that the shell enforces the same rule is false for these
two files (the peak-local block and test-q-window-stencil.sh score skips by
reason alone), so the pytest.fail is the only enforcement and says so; the
q-window comment counted two legitimate skips where the gate has three, leaving
MAX_SKIPS=3 with no headroom; `pragma: no cover` sat on a branch that fires
whenever a shared slot goes busy; and a transient probe outage was cached for
the session, so only real verdicts are cached now.
Arms, two tests each unless noted, at this commit:
ldas-grid 2 skipped
ldas-grid, RIFT_CI_REQUIRE_GPU=1 2 failed
pcdev2, A100 at CUDA slot 0 2 passed
pcdev2, CUDA_VISIBLE_DEVICES=1 2 skipped
same, RIFT_CI_REQUIRE_GPU=1 2 failed
pcdev2, CUDA_VISIBLE_DEVICES=1,0 noloop 1 skipped (was 1 failed, KeyError)
same, RIFT_CI_REQUIRE_GPU=1 noloop 1 failed
same, peak-local 1 passed
The cache fix was mutation-checked by forcing an OSError on the first probe
call: the second call retries and returns a real verdict.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
oshaughnessy-junior
had a problem deploying
to
private-review-dispatch-rift
September 19, 2026 01:18 — with
GitHub Actions
Error
oshaughnessy-junior
marked this pull request as ready for review
September 19, 2026 01:18
oshaughnessy-junior
deployed
to
private-review-dispatch-rift
September 19, 2026 01:18 — with
GitHub Actions
Active
oshaughnessy-junior
deployed
to
private-review-dispatch-rift
September 19, 2026 12:26 — with
GitHub Actions
Active
oshaughnessy-junior
added a commit
that referenced
this pull request
Sep 19, 2026
Found by sweeping for the defect class rather than for the spelling. PR #370 swept for `importorskip('cupy')` and fixed the two sites that matched. This is a third hand-rolled guard doing the same job, which that sweep could not see. It is NOT the importorskip defect -- it imports cupy in a try and already fails under RIFT_CI_REQUIRE_GPU=1. Its check is `getDeviceCount() >= 1`, which asks whether a device exists, not whether this cupy can build a kernel for it. Measured on ldas-pcdev2 2026-09-19, cupy 12.0.0, where CUDA slot 1 is a Blackwell cc 12.0: CUDA_VISIBLE_DEVICES before after 0 (A100) 1 passed 1 passed 1 (Blackwell) 1 failed, CompileException 1 skipped 1 + RIFT_CI_REQUIRE_GPU=1 1 failed 1,0 (bad first slot) 1 failed, CompileException 1 passed 1,2 (bad first slot) 1 failed, CompileException 1 passed The last two rows are the second thing the old guard could not do: it never chose a slot, so it ran on the default one and died with a usable card in the list. `require_rift_backend=True` is deliberately NOT passed, and the docstring now gives the measurement rather than an inference. An earlier draft argued from the exception TYPE -- CompileException rather than the KeyError that signals RIFT's import-time fallback -- which proves nothing, because the CompileException fires first and would mask a KeyError either way. What settles it: RIFT does fall back to numpy at "1,0" and "1,2" (xpy_default is numpy, cupy_here False, both measured), and this test passes regardless, because it takes xpy as an argument and never reaches SphericalHarmonics_gpu._coeffs. On ldas-grid (cupy present, no driver) the leg skips, and fails under RIFT_CI_REQUIRE_GPU=1, on pytest 9.1.1 and 8.3.5 alike. Collection floor unchanged at 171. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
From pytest 9.1,
pytest.importorskip('cupy')skips only onModuleNotFoundErrorandre-raises every other
ImportError. CIT hosts have the cupy package from CVMFS, so with noCUDA driver it reaches
libcuda.so.1: cannot open shared object fileand the guard failsthe test. Two GPU-parity tests failed that way on ldas-grid.
Both now call
gpu_slot_probe.cupy_or_skip(), a subprocess probe copied fromtest_e2e_analytic_pipeline.py.RIFT_CI_REQUIRE_GPU=1still makes a skip a failure.Which pytest
Same interpreter on ldas-grid. The NFS home shadows CVMFS, so an interactive run gets
9.1.1; a container or a condor node gets 8.3.5 and skips today.
~/.local/lib/python3.11/site-packagesPYTHONNOUSERSITE=1PytestDeprecationWarningVerified
Two tests per arm unless noted. The e2e gate's own probe is untouched.
RIFT_CI_REQUIRE_GPU=1CUDA_VISIBLE_DEVICES=1, Blackwell cc 12.0NOSLOT CompileExceptionRIFT_CI_REQUIRE_GPU=1CUDA_VISIBLE_DEVICES=1,0, noloopKeyError)RIFT_CI_REQUIRE_GPU=1, noloopLast three rows:
.use()is not enough when the FIRST visible slot is unusable, becauseRIFT's
try: import cupyprobes run at import on the default device. Opt-in, sincepeak-local takes
xpyas an argument and passes.Gates
ldas-grid, CVMFS igwn python 3.11, cupy 12.0.0.
.travis/test-q-window-stencil.sh.travis/test-ci-roster.pyKnown limit
Break the probe and both lanes skip naming the GPU, which the reason-matching gates accept.
.travis/test-integrate.sh:165already records this for the e2e gate.RIFT_CI_REQUIRE_GPU=1closes it andgpu_integrationsets it. Six other adversarial-reviewfindings are fixed in the second commit.
🤖 Generated with Claude Code