Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
90 changes: 83 additions & 7 deletions .travis/test-integrate.sh
Original file line number Diff line number Diff line change
Expand Up @@ -121,19 +121,65 @@ python -m pytest -q "$_PORTDENS_TESTS"
# against the driver's own supplemental_ln_likelihood call sites, its generated signature against
# every live likelihood_function signature, and its array module. This is the half the
# end-to-end gate below is structurally blind to -- right_ascension, phi_orb and psi are iid
# uniform on [0, 2pi), so NO marginal can tell a permutation of the three apart, and the runners
# have no cupy, so nothing there sees a host/device mistake. Pure AST + exec, ~4 s.
# uniform on [0, 2pi), so NO marginal can tell a permutation of the three apart, however and
# wherever that gate is run. Mostly AST + exec, ~5 s.
# Its section 4 runs the same device paths against REAL cupy. WHETHER THOSE RUN DEPENDS ON THE
# RUNNER, not on the code: they skip on a CPU-only runner, and they RUN under .gitlab-ci.yml's
# `gpu_integration` job, which invokes this script with RIFT_CI_REQUIRE_GPU=1 and
# CUDA_VISIBLE_DEVICES=0 inside the GPU container (the preflight at the top of this file then
# makes a working cupy a hard requirement). So a green run on a CPU-only runner has NOT
# exercised the device half; a green gpu_integration run has. Skipped tests are still
# collected, so the count below is the same on both. The SKIP GUARD after the run is what
# keeps a CPU-only pass honest.
# The collected count tracks the number of `def likelihood_function` signatures in the driver
# (one parametrized case each, currently 8); if a signature is added, look at the new one and
# update the number.
_ZLSTANDIN_TESTS=MonteCarloMarginalizeCode/Code/test/test_zero_likelihood_standin.py
_ZLSTANDIN_EXPECTED=23
_ZLSTANDIN_EXPECTED=27
_ZLSTANDIN_FOUND=$(python -m pytest -q --collect-only "$_ZLSTANDIN_TESTS" 2>/dev/null | grep -c '::' || true)
if [ "$_ZLSTANDIN_FOUND" -ne "$_ZLSTANDIN_EXPECTED" ]; then
echo "zero-likelihood stand-in gate: collected $_ZLSTANDIN_FOUND tests, expected $_ZLSTANDIN_EXPECTED" >&2
exit 1
fi
python -m pytest -q "$_ZLSTANDIN_TESTS"
# SKIP guard, same shape and same reason as the time-marginalization one below: `pytest -q`
# exits 0 with skips, the count guard above catches DESELECTION and not SKIPPING, and skipping
# is now this gate's NORMAL state on a CPU-only runner (3 of 27). Identify rather than count --
# the expected number of skips is a property of the runner, not of the code. Allow skips whose
# reason names cupy/GPU/CUDA; fail on any other, whatever the total.
#
# This is future-proofing, not a live hole, and the reason first written here was wrong: a
# missing ILE executable does NOT reach the module-level skipif, because the parametrize on
# test_the_factor_gets_the_raw_sampled_value_for_every_argument reads the driver at IMPORT, so
# it is a collection ERROR and the count guard above already catches it (collected drops to 0).
# Verified by moving the executable aside. The skipif is unreachable today.
# Streamed as well as captured, so a long gate is not silent on a runner -- but NOT via
# `tee /dev/stderr`. That opens /proc/self/fd/2 with O_TRUNC, so when stderr is a regular file
# (`bash .travis/test-integrate.sh > gate.log 2>&1`, the obvious way to run this) it truncates
# the log to zero and every earlier gate's output is gone, while the shell's own fd 2 keeps its
# offset and writes NULs into the hole. Measured; with stderr CLOSED it went on to overwrite
# the running script. `tee >(cat >&2)` writes through a pipe to a process that appends, which
# has none of that, keeps $_OUT intact for the grep below, and still propagates pytest's exit
# status under `set -o pipefail`.
_ZLSTANDIN_OUT=$(python -m pytest -q -rs "$_ZLSTANDIN_TESTS" 2>&1 | tee >(cat >&2)) \
|| { echo "zero-likelihood stand-in gate FAILED" >&2; exit 1; }
# On a runner that PROMISES a device, any skip is bad. The reason-matching below cannot tell
# "there is no GPU here" from "the GPU probe itself broke": break the probe and every device
# lane skips with a reason naming the GPU, which this would score as fine. Measured -- an
# unimportable module inside the e2e gate's _GPU_PROBE silently removed all 7 device lanes and
# scored 0 bad. RIFT_CI_REQUIRE_GPU=1 means the preflight at the top of this file already
# proved cupy works and a device computes, so there is nothing left for a skip to legitimately
# mean. Measured on ldas-pcdev2 slot 0: 0 skips.
if [[ "${RIFT_CI_REQUIRE_GPU:-0}" == "1" ]]; then
_ZLSTANDIN_BAD=$(echo "$_ZLSTANDIN_OUT" | grep -cE '^SKIPPED' || true)
else
_ZLSTANDIN_BAD=$(echo "$_ZLSTANDIN_OUT" | grep -E '^SKIPPED' | grep -vciE 'cupy|gpu|cuda' || true)
fi
_ZLSTANDIN_BAD=${_ZLSTANDIN_BAD:-0}
if [ "$_ZLSTANDIN_BAD" -ne 0 ]; then
echo "zero-likelihood stand-in gate: $_ZLSTANDIN_BAD unacceptable SKIPPED line(s) -- pytest -rs groups equal reasons, so this is not a test count (RIFT_CI_REQUIRE_GPU=${RIFT_CI_REQUIRE_GPU:-0}; with it set, ANY skip is unacceptable):" >&2
echo "$_ZLSTANDIN_OUT" | grep -E '^SKIPPED' >&2
exit 1
fi

# END-TO-END, ANALYTIC. ILE run to completion on a case whose answer is known in closed form:
# a zero-strain fixture, --zero-likelihood (exact ln Z = 0), and an analytic supplementary
Expand All @@ -146,16 +192,46 @@ python -m pytest -q "$_ZLSTANDIN_TESTS"
# driver's own default --sampler-method adaptive_cartesian. Both were invisible because the
# driver catches the exception, prints FAILED ANALYSIS and EXITS 0. It also caught ILE
# --sampler-method GMM returning an evidence ~30 nats wrong; #359 fixed that, and GMM is now a
# lane here rather than a recorded defect. Needs no network, no real event and no GPU; about
# lane here rather than a recorded defect. Needs no network and no real event; about
# 8-12 s per ILE arm plus ~15 s once for the distance-marginalization lookup table.
# Its section 4 runs the same answers with a GPU VISIBLE and asserts the child reached one.
# Those 7 lanes skip on a CPU-only runner and RUN under .gitlab-ci.yml's `gpu_integration` job
# (RIFT_CI_REQUIRE_GPU=1, CUDA_VISIBLE_DEVICES=0, GPU container). Skipped tests are still
# collected, so the count below is the same on both; the SKIP GUARD after the run is what stops
# a CPU-only pass from reading as device coverage.
_E2E_TESTS=MonteCarloMarginalizeCode/Code/test/test_e2e_analytic_pipeline.py
_E2E_EXPECTED=17
_E2E_EXPECTED=24
_E2E_FOUND=$(python -m pytest -q --collect-only "$_E2E_TESTS" 2>/dev/null | grep -c '::' || true)
if [ "$_E2E_FOUND" -ne "$_E2E_EXPECTED" ]; then
echo "e2e analytic gate: collected $_E2E_FOUND tests, expected $_E2E_EXPECTED" >&2
exit 1
fi
python -m pytest -q "$_E2E_TESTS"
# SKIP guard, as above. This gate skips 7 of 24 on a CPU-only runner, and it has a non-GPU
# skip path that matters: build_event returns None when lal_path2cache is missing, which skips
# the WHOLE module -- 24 silent skips under a green exit 0. So a skip whose reason does not
# name cupy/GPU/CUDA fails the gate.
# Streamed, not `tee /dev/stderr` -- see the note on the stand-in gate above. This is the gate
# that made streaming worth having: 3-9 minutes, and silent when captured.
_E2E_OUT=$(python -m pytest -q -rs "$_E2E_TESTS" 2>&1 | tee >(cat >&2)) \
|| { echo "e2e analytic gate FAILED" >&2; exit 1; }
# On a runner that PROMISES a device, any skip is bad. The reason-matching below cannot tell
# "there is no GPU here" from "the GPU probe itself broke": break the probe and every device
# lane skips with a reason naming the GPU, which this would score as fine. Measured -- an
# unimportable module inside the e2e gate's _GPU_PROBE silently removed all 7 device lanes and
# scored 0 bad. RIFT_CI_REQUIRE_GPU=1 means the preflight at the top of this file already
# proved cupy works and a device computes, so there is nothing left for a skip to legitimately
# mean. Measured on ldas-pcdev2 slot 0: 0 skips.
if [[ "${RIFT_CI_REQUIRE_GPU:-0}" == "1" ]]; then
_E2E_BAD=$(echo "$_E2E_OUT" | grep -cE '^SKIPPED' || true)
else
_E2E_BAD=$(echo "$_E2E_OUT" | grep -E '^SKIPPED' | grep -vciE 'cupy|gpu|cuda' || true)
fi
_E2E_BAD=${_E2E_BAD:-0}
if [ "$_E2E_BAD" -ne 0 ]; then
echo "e2e analytic gate: $_E2E_BAD unacceptable SKIPPED line(s) -- pytest -rs groups equal reasons, so this is not a test count (RIFT_CI_REQUIRE_GPU=${RIFT_CI_REQUIRE_GPU:-0}; with it set, ANY skip is unacceptable):" >&2
echo "$_E2E_OUT" | grep -E '^SKIPPED' >&2
exit 1
fi

# Supplementary-likelihood plugin hook: the NAL reader/evaluator (pure numpy, no data) and the
# static guard on the drivers' prepare-hook wiring, which is what makes the plugin receive the
Expand Down
27 changes: 20 additions & 7 deletions MonteCarloMarginalizeCode/Code/test/analytic_supplement_for_e2e.py
Original file line number Diff line number Diff line change
Expand Up @@ -69,14 +69,27 @@ def ln_analytic_factor(right_ascension, declination, phi_orb, inclination, psi,
#
# And it has to be USED, not just accepted. The three sites that pass it do
# `lnL += factor(...)` with lnL on the device, so a factor that always returns numpy raises
# there on a GPU host. Computing through xpy is what makes this example portable; the gate
# cannot check it, because every lane here pins CUDA_VISIBLE_DEVICES="" and the CI runners
# have no cupy.
# there on a GPU host. Computing through xpy is what makes this example portable, and it IS
# checked now: test_zero_likelihood_standin.py section 4 runs this function with xpy=cupy,
# and the e2e gate's section 4 runs the whole pipeline with a device visible. Both skip
# where there is no usable GPU.
#
# The CAST is not decoration either. mcsampler (--sampler-method adaptive_cartesian) hands
# its integrand object-dtype draws, on which np.cos raises "loop of ufunc does not support
# argument 0 of type float"; the driver's own non-vectorized likelihood casts for the same
# reason ("get rid of 'object'"). Any real supplementary factor needs it.
# The CAST is not decoration, and the EXPLICIT dtype is the whole of it. mcsampler
# (--sampler-method adaptive_cartesian) hands its integrand object-dtype draws, on which
# np.cos raises "loop of ufunc does not support argument 0 of type float"; the driver's own
# non-vectorized likelihood casts for the same reason ("get rid of 'object'").
#
# ONE line does BOTH jobs -- object dtype to float, and host to device -- and it does so
# only because `dtype=` is passed. Measured on cupy 12.0.0, 2026-09-17, on ldas-pcdev2
# (A100, cc 8.0) and ldas-pcdev13 (RTX 2080 Ti, cc 7.5):
# cupy.asarray(obj_arr) -> ValueError: Unsupported dtype object
# cupy.asarray(obj_arr, dtype=float) -> float64 cupy.ndarray, values preserved
# cupy.asarray(device_arr, dtype=float) -> float64 cupy.ndarray
# A review read the first of those three and concluded this line could not be doing both.
# Do NOT act on that by rewriting it as xpy.asarray(np.asarray(x, dtype=float)): np.asarray
# of a cupy array raises TypeError ("Implicit conversion to a NumPy array is not allowed"),
# so that form breaks every GPU run -- the exact path the line exists for. Re-check by
# running those three calls; it is a three-line probe, not an argument.
phi_orb = xpy.asarray(phi_orb, dtype=float)
out = A_COEFF * xpy.cos(phi_orb)
if B_COEFF:
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -31,7 +31,7 @@

_AV, _PORTFOLIO, _GMM = gate._AV, gate._PORTFOLIO, gate._GMM

# (label, sampler argv, A, B, exact-A, exact-B, extra kwargs for _run_ile)
# (label, sampler argv, A, B, extra kwargs for _run_ile). A is None for a prior-only lane.
LANES = [
("prior-only, AV", _AV, None, 0.0, {}),
("prior-only, portfolio", _PORTFOLIO, None, 0.0, {}),
Expand All @@ -53,6 +53,15 @@
("distance-marginalized", _AV, 8.0, 2.0, {"_needs_dmarg": True}),
("adaptive_cartesian, --n-max 60000",
["--sampler-method", "adaptive_cartesian"], 8.0, 2.0, dict(n_max=60000)),
# DEVICE lanes. They need --gpu-slot, and are refused rather than skipped without it, for
# the same reason as a --lane typo: a table that quietly stops covering four of its rows
# under a summary line that reads clean is worse than no table.
("GPU prior-only, AV", _AV, None, 0.0, {"_needs_gpu": True}),
("GPU prior-only, GMM", _GMM, None, 0.0, {"_needs_gpu": True}),
("GPU A=8 B=2, AV", _AV, 8.0, 2.0, {"_needs_gpu": True}),
# NOT A=8 B=2: that is the n_eff lottery recorded at the end of the gate's CALIBRATION
# section. Mirrors the CPU "A=0.75 B=3, GMM" row instead.
("GPU A=0.75 B=3, GMM", _GMM, 0.75, 3.0, {"_needs_gpu": True}),
]


Expand All @@ -65,10 +74,10 @@ def _dmarg_extra(event, out):
stderr=subprocess.STDOUT, timeout=1800)
if proc.returncode != 0 or not table.exists():
raise SystemExit("util_InitMargTable failed:\n%s" % proc.stdout.decode()[-1500:])
return dict(extra=("--distance-marginalization",
"--distance-marginalization-lookup-table", str(table),
"--d-min", "100", "--d-max", "1000",
"--time-marginalization", "--vectorized", "--gpu", "--force-xpy"))
return ("--distance-marginalization",
"--distance-marginalization-lookup-table", str(table),
"--d-min", "100", "--d-max", "1000",
"--time-marginalization", "--vectorized", "--gpu", "--force-xpy")


def main():
Expand All @@ -77,33 +86,80 @@ def main():
ap.add_argument("--seeds", type=int, default=8)
ap.add_argument("--first-seed", type=int, default=1000)
ap.add_argument("--lane", default=None, help="substring; measure only matching lanes")
ap.add_argument("--gpu-slot", default=None,
help="CUDA slot for the GPU lanes, as the CHILD should see it. Probe it "
"first: the CIT slot map moves and cupy 12.0.0 cannot build a kernel "
"for the Blackwell cards. Without it the GPU lanes are refused.")
args = ap.parse_args()

# Lane selection BEFORE anything expensive or on-disk. It depends on nothing but argv, and
# running it after mkdtemp/build_event meant a --lane typo cost a full fixture build and
# then leaked the directory: the refusal path below neither prints nor removes it, and on
# the CIT nodes that directory is on /, which is 20-22 GB.
#
# The index is carried alongside each lane, and it is the index into LANES rather than into
# this filtered list, so --lane does not renumber the output tags.
selected = [(i, L) for i, L in enumerate(LANES) if not args.lane or args.lane in L[0]]
# WITHOUT --gpu-slot the device rows are dropped, LOUDLY, and the rest are measured.
# Refusing outright broke the one invocation two comments tell you to run
# ("RE-DERIVE THIS TABLE, do not trust it" -> `make_e2e_calibration.py --seeds 8`): it
# exited 1 having measured nothing, and no single --lane selects the 17 host rows and
# excludes the 4 device ones, so the committed host table could not be regenerated at all.
# A comment demanding re-derivation, next to a command that cannot re-derive, is the exact
# unfalsifiable claim this script exists to remove.
#
# Naming them ON STDERR matters: the table below is built from stdout, so a dropped row
# cannot slip into it, and a person still sees which rows were not measured.
if args.gpu_slot is None:
dropped = [L[0].strip() for _i, L in selected if L[4].get("_needs_gpu")]
if dropped:
print("# NOT MEASURED (no --gpu-slot): %s" % ", ".join(dropped), file=sys.stderr)
selected = [(i, L) for i, L in selected if not L[4].get("_needs_gpu")]
if dropped and not selected:
# Dropping every selected row and then printing a clean summary over nothing is
# the shape this whole branch exists to remove, so that case is a refusal.
raise SystemExit("--gpu-slot is required: every lane selected%s needs a device:"
"\n %s\nProbe a slot this cupy can build a kernel for."
% ((" by --lane %r" % args.lane) if args.lane else "",
"\n ".join(dropped)))
if not selected:
# Exiting 0 having measured nothing, under a summary line that reads like a clean
# result, is the exact shape this whole branch exists to remove.
raise SystemExit("--lane %r matched none of:\n %s"
% (args.lane, "\n ".join(L[0].strip() for L in LANES)))

import pathlib
out = pathlib.Path(tempfile.mkdtemp(prefix="e2e_calibration_"))
event = gate.build_event(out)
if event is None:
raise SystemExit("could not build the fixture (lal_path2cache missing or failed)")
raise SystemExit(
"could not build the fixture (lal_path2cache missing or failed); the empty "
"directory is at %s" % out)

seeds = [args.first_seed + i for i in range(args.seeds)]
selected = [L for L in LANES if not args.lane or args.lane in L[0]]
if not selected:
# Exiting 0 having measured nothing, under a summary line that reads like a clean
# result, is the exact shape this whole branch exists to remove.
raise SystemExit("--lane %r matched none of:\n %s"
% (args.lane, "\n ".join(L[0].strip() for L in LANES)))
print("# lane max |z| max sigma min n_eff")
worst_z = worst_s = 0.0
least_n = float("inf")
for label, sampler, a, b, kw in selected:
kw0 = kw
for idx, (label, sampler, a, b, kw) in selected:
kw = dict(kw)
# Both of these EXTEND `extra` rather than assigning it, so a lane that carries its
# own options does not lose them, and a lane needing both keeps both.
if kw.pop("_needs_dmarg", False):
kw.update(_dmarg_extra(event, out))
kw["extra"] = tuple(kw.get("extra", ())) + _dmarg_extra(event, out)
if kw.pop("_needs_gpu", False):
# expect_device=True, so a lane that fell back to the host FAILS instead of
# contributing host numbers to a row labelled GPU.
kw.update(cuda=args.gpu_slot, expect_device=True,
extra=tuple(kw.get("extra", ())) + gate._GPU_FLAGS)
exact = 0.0 if a is None else gate._exact(a, b)
zs, sigs, neffs = [], [], []
for seed in seeds:
tag = "cal_%d_%d" % (LANES.index((label, sampler, a, b, kw0)), seed)
# `idx` from enumerate, NOT LANES.index(...): index is a first-match lookup, so two
# identical lane rows would silently share a tag and overwrite each other's ILE
# output directory. LANES.index happened to be correct for every row it ever had;
# it is wrong the moment one is duplicated, which is the kind of edit this table
# invites, and nothing would have said so.
tag = "cal_%d_%d" % (idx, seed)
lnL, sigma, neff = gate._run_ile(event, tag, sampler, a_coeff=a, b_coeff=b,
seed=seed, **kw)
zs.append(abs((lnL - exact) / sigma))
Expand All @@ -115,8 +171,8 @@ def main():
print("# %-44s%5.2f %6.4f %6.0f"
% (label, max(zs), max(sigs), min(neffs)))
print("#")
print("# across all lanes: worst |z| %.2f, worst sigma %.4f, least n_eff %.0f"
% (worst_z, worst_s, least_n))
print("# across the %d lane(s) MEASURED above: worst |z| %.2f, worst sigma %.4f, "
"least n_eff %.0f" % (len(selected), worst_z, worst_s, least_n))
# Printed, not deleted: a failing lane is worth inspecting. /tmp is small on the CIT
# nodes, so clean up when you are done.
print("# fixture kept at %s (rm -rf it when done)" % out)
Expand Down
Loading
Loading