Skip to content

GPU parity legs: skip on a driverless host instead of failing - #370

Merged
oshaughnessy-junior merged 2 commits into
rift_O4dfrom
claude/gpu-guard-skip-not-fail
Sep 19, 2026
Merged

oshaughnessy-junior merged 2 commits into
rift_O4dfrom
claude/gpu-guard-skip-not-fail

Conversation

@oshaughnessy-junior

@oshaughnessy-junior oshaughnessy-junior commented Sep 18, 2026 •

Copy link
Copy Markdown
Owner

From pytest 9.1, pytest.importorskip('cupy') skips only on ModuleNotFoundError and
re-raises every other ImportError. CIT hosts have the cupy package from CVMFS, so with no
CUDA driver it reaches libcuda.so.1: cannot open shared object file and the guard fails
the test. Two GPU-parity tests failed that way on ldas-grid.

Both now call gpu_slot_probe.cupy_or_skip(), a subprocess probe copied from
test_e2e_analytic_pipeline.py. RIFT_CI_REQUIRE_GPU=1 still makes a skip a failure.

Which pytest

Same interpreter on ldas-grid. The NFS home shadows CVMFS, so an interactive run gets
9.1.1; a container or a condor node gets 8.3.5 and skips today.

pytest from result
9.1.1 ~/.local/lib/python3.11/site-packages 1 failed
8.3.5 CVMFS's own, via PYTHONNOUSERSITE=1 1 skipped, plus a PytestDeprecationWarning

Verified

Two tests per arm unless noted. The e2e gate's own probe is untouched.

Arm Result
ldas-grid, pristine base (control) 2 failed
ldas-grid 2 skipped
ldas-grid, RIFT_CI_REQUIRE_GPU=1 2 failed
ldas-pcdev2, A100 at CUDA slot 0 2 passed
ldas-pcdev2, CUDA_VISIBLE_DEVICES=1, Blackwell cc 12.0 2 skipped, NOSLOT CompileException
same, plus RIFT_CI_REQUIRE_GPU=1 2 failed
ldas-pcdev2, CUDA_VISIBLE_DEVICES=1,0, noloop 1 skipped (was 1 failed, KeyError)
same, plus RIFT_CI_REQUIRE_GPU=1, noloop 1 failed
same, peak-local 1 passed

Last three rows: .use() is not enough when the FIRST visible slot is unusable, because
RIFT's try: import cupy probes run at import on the default device. Opt-in, since
peak-local takes xpy as an argument and passes.

Gates

ldas-grid, CVMFS igwn python 3.11, cupy 12.0.0.

Gate Result
both files, full 124 passed, 2 skipped
both files, full, ldas-pcdev2 with the A100 126 passed
.travis/test-q-window-stencil.sh PASS, 78 collected, 75 passed, 3 skipped
.travis/test-ci-roster.py PASS
peak-local collection floor 121

Known limit

Break the probe and both lanes skip naming the GPU, which the reason-matching gates accept.
.travis/test-integrate.sh:165 already records this for the e2e gate.
RIFT_CI_REQUIRE_GPU=1 closes it and gpu_integration sets it. Six other adversarial-review
findings are fixed in the second commit.

🤖 Generated with Claude Code

@oshaughnessy-junior
oshaughnessy-junior deployed to private-review-dispatch-rift September 18, 2026 23:39 — with GitHub Actions Active
`pytest.importorskip('cupy')` skips only on ModuleNotFoundError and re-raises
every other ImportError.  The IGWN CVMFS environment installs the cupy package
on every CIT host, so on one with no CUDA driver the import reaches
`libcuda.so.1: cannot open shared object file` -- a plain ImportError -- and the
guard becomes a test FAILURE.  Measured on ldas-grid at f7f986a (CVMFS igwn
python 3.11, cupy 12.0.0, pytest 9.1.1): both legs below failed, and read as a
broken GPU path.  GitHub's runners have no cupy at all, get the clean
ModuleNotFoundError, and skip, so CI cannot see this.

Both now call gpu_slot_probe.cupy_or_skip(), a new sibling helper carrying the
approach test_e2e_analytic_pipeline.py settled on: a SUBPROCESS that imports
cupy and runs a kernel on each visible slot, and a one-word verdict line.  That
also separates "no cupy here" from "cupy here but the visible device is one it
cannot compile for", which is the Blackwell case.  The e2e gate's own copy is
left alone; it re-indexes slots for grandchildren and this one must not.

RIFT_CI_REQUIRE_GPU=1 keeps its meaning: the environment promised a device, so
a skip stays a failure.

Verified, two tests per arm:

  ldas-grid (cupy installed, no driver)   2 skipped
  ldas-grid, RIFT_CI_REQUIRE_GPU=1        2 failed
  ldas-pcdev2, A100 at CUDA slot 0        2 passed
  ldas-pcdev2, CUDA_VISIBLE_DEVICES=1     2 skipped (NOSLOT CompileException)
  same, RIFT_CI_REQUIRE_GPU=1             2 failed

Slot 1 there is a Blackwell cc 12.0 that cupy 12.0.0 cannot build for;
importorskip imports it successfully and the test body then dies.  CUDA's
ordering is not nvidia-smi's on that host -- the A100 is nvidia-smi 2 and CUDA
0 -- which is why the probe runs a kernel per slot rather than trusting an
index.

Collection counts are unchanged (peak-local 121, q-window-stencil 78/75/3);
.travis/test-q-window-stencil.sh and .travis/test-ci-roster.py pass on ldas-grid.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@oshaughnessy-junior
oshaughnessy-junior force-pushed the claude/gpu-guard-skip-not-fail branch from e84e106 to 9195b2f Compare September 19, 2026 00:03
@oshaughnessy-junior
oshaughnessy-junior deployed to private-review-dispatch-rift September 19, 2026 00:03 — with GitHub Actions Active
…VMFS env

Six findings from an independent reviewer, all reproduced before acting.

THE PROVENANCE WAS WRONG, and it was the headline claim.  Same interpreter on
ldas-grid, one test that does `pytest.importorskip('cupy')`:

    pytest 9.1.1  from ~/.local/lib/python3.11/site-packages   1 failed
    pytest 8.3.5  CVMFS's own, via PYTHONNOUSERSITE=1          1 skipped

8.3.5 defaults exc_type=ImportError and warns that it will stop doing so.  The
user site-packages on the NFS home shadows CVMFS, so an ordinary interactive run
here gets 9.1.1 and does fail, but a container or a condor node gets 8.3.5 and
skips.  So the claim is "under pytest >= 9.1", not "on CIT with CVMFS", and the
file and the q-window comment both said the latter.

Device(d).use() IS NOT ENOUGH when the FIRST visible slot is unusable.  RIFT's
`try: import cupy` probes run at import, on the default device;
SphericalHarmonics_gpu runs cupy.array(5) there, and on a Blackwell it raises,
sets cupy_here=False and keys _coeffs by numpy alone.  A later .use() cannot
move that.  Measured on ldas-pcdev2 with CUDA_VISIBLE_DEVICES="1,0" and "1,2".

The two callers differ, so the new check is OPT-IN rather than blanket:

  peak-local  takes xpy as an argument, never touches SphericalHarmonics_gpu,
              and PASSES at those pins.  A blanket check would cost this lane.
  noloop      goes through the likelihood and died on KeyError <module cupy>.
              It now skips, naming both flags and the remedy.

Also: no_gpu's claim that the shell enforces the same rule is false for these
two files (the peak-local block and test-q-window-stencil.sh score skips by
reason alone), so the pytest.fail is the only enforcement and says so; the
q-window comment counted two legitimate skips where the gate has three, leaving
MAX_SKIPS=3 with no headroom; `pragma: no cover` sat on a branch that fires
whenever a shared slot goes busy; and a transient probe outage was cached for
the session, so only real verdicts are cached now.

Arms, two tests each unless noted, at this commit:

  ldas-grid                                    2 skipped
  ldas-grid, RIFT_CI_REQUIRE_GPU=1             2 failed
  pcdev2, A100 at CUDA slot 0                  2 passed
  pcdev2, CUDA_VISIBLE_DEVICES=1               2 skipped
  same, RIFT_CI_REQUIRE_GPU=1                  2 failed
  pcdev2, CUDA_VISIBLE_DEVICES=1,0  noloop     1 skipped (was 1 failed, KeyError)
  same, RIFT_CI_REQUIRE_GPU=1       noloop     1 failed
  same,                             peak-local 1 passed

The cache fix was mutation-checked by forcing an OSError on the first probe
call: the second call retries and returns a real verdict.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@oshaughnessy-junior
oshaughnessy-junior marked this pull request as ready for review September 19, 2026 01:18
@oshaughnessy-junior
oshaughnessy-junior deployed to private-review-dispatch-rift September 19, 2026 01:18 — with GitHub Actions Active
@oshaughnessy-junior
oshaughnessy-junior merged commit 4adf965 into rift_O4d Sep 19, 2026
35 of 36 checks passed
@oshaughnessy-junior
oshaughnessy-junior deployed to private-review-dispatch-rift September 19, 2026 12:26 — with GitHub Actions Active
oshaughnessy-junior added a commit that referenced this pull request Sep 19, 2026
Found by sweeping for the defect class rather than for the spelling. PR #370
swept for `importorskip('cupy')` and fixed the two sites that matched. This is a
third hand-rolled guard doing the same job, which that sweep could not see.

It is NOT the importorskip defect -- it imports cupy in a try and already fails
under RIFT_CI_REQUIRE_GPU=1. Its check is `getDeviceCount() >= 1`, which asks
whether a device exists, not whether this cupy can build a kernel for it.

Measured on ldas-pcdev2 2026-09-19, cupy 12.0.0, where CUDA slot 1 is a
Blackwell cc 12.0:

  CUDA_VISIBLE_DEVICES   before                      after
  0    (A100)            1 passed                    1 passed
  1    (Blackwell)       1 failed, CompileException  1 skipped
  1  + RIFT_CI_REQUIRE_GPU=1                         1 failed
  1,0  (bad first slot)  1 failed, CompileException  1 passed
  1,2  (bad first slot)  1 failed, CompileException  1 passed

The last two rows are the second thing the old guard could not do: it never
chose a slot, so it ran on the default one and died with a usable card in the
list.

`require_rift_backend=True` is deliberately NOT passed, and the docstring now
gives the measurement rather than an inference. An earlier draft argued from the
exception TYPE -- CompileException rather than the KeyError that signals RIFT's
import-time fallback -- which proves nothing, because the CompileException fires
first and would mask a KeyError either way. What settles it: RIFT does fall back
to numpy at "1,0" and "1,2" (xpy_default is numpy, cupy_here False, both
measured), and this test passes regardless, because it takes xpy as an argument
and never reaches SphericalHarmonics_gpu._coeffs.

On ldas-grid (cupy present, no driver) the leg skips, and fails under
RIFT_CI_REQUIRE_GPU=1, on pytest 9.1.1 and 8.3.5 alike. Collection floor
unchanged at 171.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

This branch was successfully deployed

1 active deployment
private-review-dispatch-rift — c391d6a6 Deployed Sep 19, 2026 by oshaughnessy-junior via Dispatch exact RIFT PR generation #1421
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant