diff --git a/.travis/test-integrate.sh b/.travis/test-integrate.sh index 6509772f7..ebc15119f 100755 --- a/.travis/test-integrate.sh +++ b/.travis/test-integrate.sh @@ -165,7 +165,7 @@ _ZLSTANDIN_OUT=$(python -m pytest -q -rs "$_ZLSTANDIN_TESTS" 2>&1 | tee >(cat >& # On a runner that PROMISES a device, any skip is bad. The reason-matching below cannot tell # "there is no GPU here" from "the GPU probe itself broke": break the probe and every device # lane skips with a reason naming the GPU, which this would score as fine. Measured -- an -# unimportable module inside the e2e gate's _GPU_PROBE silently removed all 7 device lanes and +# unimportable module inside the e2e gate's _GPU_PROBE silently removed all 9 device lanes and # scored 0 bad. RIFT_CI_REQUIRE_GPU=1 means the preflight at the top of this file already # proved cupy works and a device computes, so there is nothing left for a skip to legitimately # mean. Measured on ldas-pcdev2 slot 0: 0 skips. @@ -195,20 +195,20 @@ fi # lane here rather than a recorded defect. Needs no network and no real event; about # 8-12 s per ILE arm plus ~15 s once for the distance-marginalization lookup table. # Its section 4 runs the same answers with a GPU VISIBLE and asserts the child reached one. -# Those 7 lanes skip on a CPU-only runner and RUN under .gitlab-ci.yml's `gpu_integration` job +# Those 9 lanes skip on a CPU-only runner and RUN under .gitlab-ci.yml's `gpu_integration` job # (RIFT_CI_REQUIRE_GPU=1, CUDA_VISIBLE_DEVICES=0, GPU container). Skipped tests are still # collected, so the count below is the same on both; the SKIP GUARD after the run is what stops # a CPU-only pass from reading as device coverage. _E2E_TESTS=MonteCarloMarginalizeCode/Code/test/test_e2e_analytic_pipeline.py -_E2E_EXPECTED=24 +_E2E_EXPECTED=26 _E2E_FOUND=$(python -m pytest -q --collect-only "$_E2E_TESTS" 2>/dev/null | grep -c '::' || true) if [ "$_E2E_FOUND" -ne "$_E2E_EXPECTED" ]; then echo "e2e analytic gate: collected $_E2E_FOUND tests, expected $_E2E_EXPECTED" >&2 exit 1 fi -# SKIP guard, as above. This gate skips 7 of 24 on a CPU-only runner, and it has a non-GPU +# SKIP guard, as above. This gate skips 9 of 26 on a CPU-only runner, and it has a non-GPU # skip path that matters: build_event returns None when lal_path2cache is missing, which skips -# the WHOLE module -- 24 silent skips under a green exit 0. So a skip whose reason does not +# the WHOLE module -- 26 silent skips under a green exit 0. So a skip whose reason does not # name cupy/GPU/CUDA fails the gate. # Streamed, not `tee /dev/stderr` -- see the note on the stand-in gate above. This is the gate # that made streaming worth having: 3-9 minutes, and silent when captured. @@ -217,7 +217,7 @@ _E2E_OUT=$(python -m pytest -q -rs "$_E2E_TESTS" 2>&1 | tee >(cat >&2)) \ # On a runner that PROMISES a device, any skip is bad. The reason-matching below cannot tell # "there is no GPU here" from "the GPU probe itself broke": break the probe and every device # lane skips with a reason naming the GPU, which this would score as fine. Measured -- an -# unimportable module inside the e2e gate's _GPU_PROBE silently removed all 7 device lanes and +# unimportable module inside the e2e gate's _GPU_PROBE silently removed all 9 device lanes and # scored 0 bad. RIFT_CI_REQUIRE_GPU=1 means the preflight at the top of this file already # proved cupy works and a device computes, so there is nothing left for a skip to legitimately # mean. Measured on ldas-pcdev2 slot 0: 0 skips. diff --git a/MonteCarloMarginalizeCode/Code/RIFT/integrators/mcsampler.py b/MonteCarloMarginalizeCode/Code/RIFT/integrators/mcsampler.py index 98272e1fd..46bbfab5a 100644 --- a/MonteCarloMarginalizeCode/Code/RIFT/integrators/mcsampler.py +++ b/MonteCarloMarginalizeCode/Code/RIFT/integrators/mcsampler.py @@ -562,6 +562,16 @@ def integrate(self, func, *args, **kwargs): else: fval = func(**unpacked) # Chris' original plan: note this insures the function arguments are tied to the parameters, using a dictionary. + # THE INTEGRAND CONVERTS, NOT THE ACCUMULATORS. On a GPU node the ILE likelihood + # returns a cupy array (xpy_default binds at import from whether cupy imports, so + # a merely visible device is enough), while everything below here is host numpy: + # numpy.hstack, statutils update/finalize, math.isnan -- and joint_p_prior is + # deliberately RiftFloat, which cupy has no dtype for at all. So the host side is + # the one that cannot move. Without this, `fval*joint_p_prior/joint_p_s` raised + # "TypeError: Unsupported type " from a cupy ufunc and the + # driver swallowed it, printed FAILED ANALYSIS and exited 0. + fval = to_host(fval) + # # Check if there is any practical contribution to the integral # @@ -1035,6 +1045,36 @@ def infer_array_module(x, xpy=None): return numpy +def to_host(x): + """Return `x` as a host (numpy) array, copying it off a device backend if it is on one. + + numpy.asarray(cupy_array) RAISES -- cupy refuses implicit host conversion -- so the copy + has to go through the device module's own asnumpy. The module is looked up from the + value, the same way infer_array_module does it, so this file keeps its numpy-only import + list. A host input is returned UNCHANGED, not re-wrapped: the numpy path must stay bit + for bit what it was. + + A non-numpy backend it cannot convert RAISES rather than passing the value through. + infer_array_module recognizes any module exposing asarray/where/clip, which is a wider + set than the ones exposing asnumpy (torch is in the gap), so "return it unchanged" would + hand a device array to the host-only arithmetic below the call site -- the silent + host/device mix this function exists to remove, reintroduced one backend later. + """ + mod = infer_array_module(x) + if mod is numpy: + return x + if hasattr(mod, "asnumpy"): + return mod.asnumpy(x) + _get = getattr(x, "get", None) # cupy-like array, module without a top-level asnumpy + if callable(_get): + return numpy.asarray(_get()) + raise TypeError( + "mcsampler cannot bring a %s.%s back to the host: the module exposes neither asnumpy " + "nor a .get() on the array, and this sampler's accumulators are host RiftFloat, which " + "no device backend can hold. Use a sampler that runs on that backend." + % (mod.__name__, type(x).__name__)) + + def clip_angle_limits(lo, hi, kind): """Clip an angular range [lo,hi] (radians) to the physical domain of `kind` ('declination' -> [-pi/2,pi/2], 'inclination' -> [0,pi]). diff --git a/MonteCarloMarginalizeCode/Code/RIFT/integrators/mcsamplerGPU.py b/MonteCarloMarginalizeCode/Code/RIFT/integrators/mcsamplerGPU.py index 6750c57d0..fb1d112a9 100644 --- a/MonteCarloMarginalizeCode/Code/RIFT/integrators/mcsamplerGPU.py +++ b/MonteCarloMarginalizeCode/Code/RIFT/integrators/mcsamplerGPU.py @@ -326,6 +326,24 @@ def setup_hist_single_param(self, x_min, x_max, n_bins, param): def compute_hist(self, x_samples, param,weights=None,floor_level=0): + # PUT THE INPUTS ON THIS SAMPLER'S BACKEND FIRST. As a portfolio member the samples + # arrive from the aggregator's host _rvs (mcsamplerPortfolio pins self.xpy = numpy) + # while self.xpy here is cupy, and the first cupy ufunc downstream -- the + # xpy.maximum(samples, xpy.zeros(...)) clip in vectorized_general_tools.histogram -- + # raised "TypeError: Unsupported type ", killing every + # portfolio run with an AC member on a GPU node behind FAILED ANALYSIS and exit 0. + # + # DEVICE is the side to converge on here. The result has to end up there anyway + # (setup_hist_single_param preallocates histogram_edges/histogram_cdf with self.xpy, + # and cdf_inverse_from_hist is read by device draws), and nothing here carries + # precision worth keeping on the host: the coordinates are float64 and this histogram + # is a *proposal* density, not an estimator (see _bincount_weighted). Contrast + # mcsampler.py, whose accumulators are deliberately RiftFloat and so convert the + # other way. asarray is a no-op when the input is already on this backend, so the + # standalone GPU and pure-numpy paths are untouched. + x_samples = self.xpy.asarray(x_samples) + if weights is not None: + weights = self.xpy.asarray(weights) # Rescale the samples to [0, 1] y_samples = ( (x_samples - self.x_min[param]) / self.x_max_minus_min[param] diff --git a/MonteCarloMarginalizeCode/Code/test/expensive_before_merging/integrators/make_e2e_calibration.py b/MonteCarloMarginalizeCode/Code/test/expensive_before_merging/integrators/make_e2e_calibration.py index 49900187c..e8cd35d10 100644 --- a/MonteCarloMarginalizeCode/Code/test/expensive_before_merging/integrators/make_e2e_calibration.py +++ b/MonteCarloMarginalizeCode/Code/test/expensive_before_merging/integrators/make_e2e_calibration.py @@ -29,7 +29,7 @@ import test_e2e_analytic_pipeline as gate # noqa: E402 -_AV, _PORTFOLIO, _GMM = gate._AV, gate._PORTFOLIO, gate._GMM +_AV, _PORTFOLIO, _GMM, _AC = gate._AV, gate._PORTFOLIO, gate._GMM, gate._AC # (label, sampler argv, A, B, extra kwargs for _run_ile). A is None for a prior-only lane. LANES = [ @@ -52,9 +52,9 @@ dict(extra=("--time-marginalization",))), ("distance-marginalized", _AV, 8.0, 2.0, {"_needs_dmarg": True}), ("adaptive_cartesian, --n-max 60000", - ["--sampler-method", "adaptive_cartesian"], 8.0, 2.0, dict(n_max=60000)), + _AC, 8.0, 2.0, dict(n_max=gate._AC_N_MAX)), # DEVICE lanes. They need --gpu-slot, and are refused rather than skipped without it, for - # the same reason as a --lane typo: a table that quietly stops covering four of its rows + # the same reason as a --lane typo: a table that quietly stops covering eight of its rows # under a summary line that reads clean is worse than no table. ("GPU prior-only, AV", _AV, None, 0.0, {"_needs_gpu": True}), ("GPU prior-only, GMM", _GMM, None, 0.0, {"_needs_gpu": True}), @@ -62,6 +62,16 @@ # NOT A=8 B=2: that is the n_eff lottery recorded at the end of the gate's CALIBRATION # section. Mirrors the CPU "A=0.75 B=3, GMM" row instead. ("GPU A=0.75 B=3, GMM", _GMM, 0.75, 3.0, {"_needs_gpu": True}), + # portfolio and adaptive_cartesian could not run on a device at all until the host/device + # conversions in mcsamplerGPU.compute_hist and mcsampler.integrate. Each mirrors its host + # twin: portfolio the plain-portfolio lanes above, adaptive_cartesian its own --n-max, taken + # from the gate rather than repeated here. + ("GPU prior-only, portfolio", _PORTFOLIO, None, 0.0, {"_needs_gpu": True}), + ("GPU A=0.75 B=3, portfolio", _PORTFOLIO, 0.75, 3.0, {"_needs_gpu": True}), + ("GPU prior-only, adaptive_cartesian", _AC, None, 0.0, + dict(_needs_gpu=True, n_max=gate._AC_N_MAX)), + ("GPU A=8 B=2, adaptive_cartesian", _AC, 8.0, 2.0, + dict(_needs_gpu=True, n_max=gate._AC_N_MAX)), ] diff --git a/MonteCarloMarginalizeCode/Code/test/test_e2e_analytic_pipeline.py b/MonteCarloMarginalizeCode/Code/test/test_e2e_analytic_pipeline.py index fe585c61c..7d3ae0b5e 100644 --- a/MonteCarloMarginalizeCode/Code/test/test_e2e_analytic_pipeline.py +++ b/MonteCarloMarginalizeCode/Code/test/test_e2e_analytic_pipeline.py @@ -36,9 +36,8 @@ THE DEVICE. Section 4 runs the prior-only and factor answers again with a GPU VISIBLE, and asserts the child reached it. RIFT binds its array module at import from whether cupy imports, not from --gpu, so on a GPU node production is on the device path by default and the rest of -this file -- which pins CUDA_VISIBLE_DEVICES="" -- said nothing about it. Two samplers cannot -run there at all; they are recorded as known-failing lanes rather than skipped, so the hole is -visible. What section 4 does NOT cover is the GPU signal likelihood: --zero-likelihood +this file -- which pins CUDA_VISIBLE_DEVICES="" -- said nothing about it. All four samplers +run there. What section 4 does NOT cover is the GPU signal likelihood: --zero-likelihood replaces likelihood_function outright, so the NoLoop path never runs. WHAT EACH ARM COSTS. About 8-12 s of one core per ILE arm, plus ~15 s once for the @@ -82,7 +81,11 @@ _PORTFOLIO = ["--sampler-method", "portfolio", "--sampler-portfolio", "AV", "--sampler-portfolio", "AC"] _GMM = ["--sampler-method", "GMM"] -SAMPLER_ARGS = {"AV": _AV, "portfolio": _PORTFOLIO, "GMM": _GMM} +_AC = ["--sampler-method", "adaptive_cartesian"] +# adaptive_cartesian stops on the sample BUDGET rather than on --n-eff at the 20000 the other +# lanes use; see test_adaptive_cartesian for the eight-seed measurement behind this number. +_AC_N_MAX = 60000 +SAMPLER_ARGS = {"AV": _AV, "portfolio": _PORTFOLIO, "GMM": _GMM, "adaptive_cartesian": _AC} # --------------------------------------------------------------------------------------- @@ -191,7 +194,7 @@ def _no_gpu(reason): .travis/test-integrate.sh applies this rule too (with RIFT_CI_REQUIRE_GPU=1 any skip is fatal there), but a rule that lives only in the shell does not survive `pytest ` on the GPU runner -- which is what someone runs to reproduce a CI failure, and it would - report green with all 7 device lanes skipped. RIFT_CI_REQUIRE_GPU is read from the ambient + report green with all 9 device lanes skipped. RIFT_CI_REQUIRE_GPU is read from the ambient environment deliberately: _child_env strips RIFT_* from ILE CHILDREN, a different question from what this pytest process was promised.""" if os.environ.get("RIFT_CI_REQUIRE_GPU", "0") == "1": @@ -301,8 +304,9 @@ def _invoke_ile(event, tag, sampler_args, *, a_coeff=None, b_coeff=0.0, incl_is_ n_max=20000, n_eff=250, seed=1000, extra=(), cuda=""): """Run one ILE arm. Returns (dir, returncode, stdout text, whether the result row exists). - Split out of _run_ile because the recorded known-failing device lanes need the CHILD'S LOG, - not a pytest failure: what they pin is the shape of a failure that is real today.""" + Kept separate from _run_ile, which is its only caller today, because the raw child log is + what a lane pinning the SHAPE of a failure needs -- section 4 had two such lanes until the + defects they recorded were fixed, and the next one should not have to re-split this.""" d = event["dir"] / tag d.mkdir(exist_ok=True) env = _child_env(cuda=cuda) @@ -503,9 +507,9 @@ def test_adaptive_cartesian(event): # rather than on --n-eff: over eight seeds it reported n_eff 91 to 220 against a target of # 250, with sigma up to 0.0430 against a 0.06 cap. At 60000 it stops on n_eff instead, at # ntotal 30000, with n_eff ~300 and sigma ~0.03. One-and-a-half times the arm, and the - # lane stops sitting one bad draw away from its own informativeness floor. - lnL, sigma, neff = _run_ile(event, tag, ["--sampler-method", "adaptive_cartesian"], - a_coeff=a, b_coeff=b, n_max=60000) + # lane stops sitting one bad draw away from its own informativeness floor. The device + # lane in section 4 inherits it through _GPU_LANE_KW. + lnL, sigma, neff = _run_ile(event, tag, _AC, a_coeff=a, b_coeff=b, n_max=_AC_N_MAX) _assert_lnZ(tag, lnL, sigma, neff, _exact(a, b)) @@ -524,20 +528,45 @@ def test_adaptive_cartesian(event): # and without _GPU_FLAGS returns bit-identical lnZ, because those flags only choose a likelihood # that this configuration does not evaluate. -_GPU_SAMPLERS = ["AV", "GMM"] +# portfolio and adaptive_cartesian were RECORDED HERE AS BROKEN when this section was written: +# both mixed host and device arrays and died in a cupy ufunc with "TypeError: Unsupported type +# ", invisible in production because the driver catches it, prints FAILED +# ANALYSIS and exits 0. Fixed on OPPOSITE sides, because the two sides are not alike -- the +# samples go TO the device in mcsamplerGPU.compute_hist, the integrand comes BACK to the host in +# mcsampler.integrate, whose accumulators are deliberately RiftFloat and so cannot follow it. +# Read those two comments before moving either conversion. +_GPU_SAMPLERS = ["AV", "GMM", "portfolio", "adaptive_cartesian"] -# Coefficients per device lane, and they are NOT the same pair for both samplers on purpose. +# Coefficients per device lane, and they are NOT the same pair for every sampler on purpose. # Each mirrors the CPU lane that sampler already runs, so a device row is comparable with a host -# row in the table below, and neither lands on a configuration this file records as unstable. +# row in the table below, and none lands on a configuration this file records as unstable. # # GMM is A=0.75 B=3, not A=8 B=2. The note at the end of the CALIBRATION section records that # GMM at A=8 WITH the inclination term is an n_eff LOTTERY: over eight seeds, two collapsed to # n_eff 73.5 and 14.5 with sigma 0.028 and 0.073, and 0.073 is over MAX_SIGMA. The evidence was # right on every seed; the error bar was not. That is why no CPU lane runs it, and a device lane # that ran it would be a flake with a ~1-in-4 seed. Eight clean GPU seeds do not refute a -# lottery -- the CPU sweep that FOUND it was also eight seeds. Both lanes keep a B term, so -# mis-routing inclination is still detectable; that is what the B term is for. -_GPU_LANE_COEFFS = {"AV": (8.0, 2.0), "GMM": (0.75, 3.0)} +# lottery -- the CPU sweep that FOUND it was also eight seeds. portfolio takes the same pair as +# its own A=0.75 B=3 host lane. Every lane keeps a B term, so mis-routing inclination is still +# detectable; that is what the B term is for. +_GPU_LANE_COEFFS = {"AV": (8.0, 2.0), "GMM": (0.75, 3.0), + "portfolio": (0.75, 3.0), "adaptive_cartesian": (8.0, 2.0)} + +# WHAT SECTION 4 STILL DOES NOT COVER, stated because an unlisted sampler is an invisible gap +# rather than an absent one. `--sampler-method adaptive_cartesian_gpu` -- mcsamplerGPU +# standalone, and the ILE driver's own GPU default -- reaches the same +# mcsamplerGPU.compute_hist conversion these lanes exercise, by a different route (its +# integrate/integrate_log adaptation blocks rather than a portfolio member's external_rvs), +# and has no lane here. Measured by hand on ldas-pcdev2 slot 0 (A100), cupy 12.0.0, seed +# 1000: prior-only ln Z 0.00893 +- 0.00888 with n_eff 3364 (z = +1.0), and A=8 B=2 +# ln Z 6.64213 +- 0.02819 against 6.65332 (z = -0.40). It works; it is simply not gated. +# A hand measurement is not a lane -- see the note at the top of the CALIBRATION section +# about re-deriving rather than trusting -- so treat this as a to-do, not as coverage. + +# Per-lane _run_ile overrides, so a device lane runs its host twin's configuration and not a +# nearby one. Applied to the prior-only lane too: there the integrand is flat and the run stops +# on --n-eff long before any budget, so the larger cap costs nothing and keeps one definition. +_GPU_LANE_KW = {"adaptive_cartesian": dict(n_max=_AC_N_MAX)} @pytest.mark.parametrize("sampler", _GPU_SAMPLERS) @@ -545,7 +574,8 @@ def test_gpu_zero_likelihood_alone_gives_ln_Z_zero(event, gpu_slot, sampler): """The prior-only answer, on device. ln Z = 0 exactly, for a normalized extrinsic prior.""" tag = "gpu_zero_%s" % sampler lnL, sigma, neff = _run_ile(event, tag, SAMPLER_ARGS[sampler] + list(_GPU_FLAGS), - a_coeff=None, cuda=gpu_slot, expect_device=True) + a_coeff=None, cuda=gpu_slot, expect_device=True, + **_GPU_LANE_KW.get(sampler, {})) _assert_lnZ(tag, lnL, sigma, neff, 0.0) @@ -553,11 +583,12 @@ def test_gpu_zero_likelihood_alone_gives_ln_Z_zero(event, gpu_slot, sampler): def test_gpu_analytic_factor_marginal_is_exact(event, gpu_slot, sampler): """The factor's closed-form marginal, on device: ln Z = ln I0(A) + ln(sinh(B)/B). - See _GPU_LANE_COEFFS for why the two samplers get different coefficients.""" + See _GPU_LANE_COEFFS for why the samplers get different coefficients.""" a, b = _GPU_LANE_COEFFS[sampler] tag = "gpu_factor_%s" % sampler lnL, sigma, neff = _run_ile(event, tag, SAMPLER_ARGS[sampler] + list(_GPU_FLAGS), - a_coeff=a, b_coeff=b, cuda=gpu_slot, expect_device=True) + a_coeff=a, b_coeff=b, cuda=gpu_slot, expect_device=True, + **_GPU_LANE_KW.get(sampler, {})) _assert_lnZ(tag, lnL, sigma, neff, _exact(a, b)) @@ -579,52 +610,6 @@ def test_the_gpu_flags_do_not_decide_the_backend(event, gpu_slot): "something else now depends on them." % (bare, flagged)) -# Device lanes that DO NOT WORK today, recorded with the shape of the failure so the hole is -# visible instead of hidden behind CUDA_VISIBLE_DEVICES="". Both are host/device mixes, both -# reproduce with NO supplementary factor, and neither is caused by anything in this file: -# -# portfolio (AV + AC) RIFT/likelihood/vectorized_general_tools.py, histogram(): -# `xpy.maximum(samples, blank_array)` with blank_array built by -# xpy.zeros on device and `samples` arriving as a numpy array from -# mcsamplerGPU.compute_hist. -# adaptive_cartesian RIFT/integrators/mcsampler.py, `fval*joint_p_prior/joint_p_s`: the -# integrand is on device, the two prior arrays are numpy. -# -# Both surface as `TypeError: Unsupported type ` from a cupy ufunc, and -# both are invisible in production because the driver catches the exception, prints FAILED -# ANALYSIS and EXITS 0. Measured on ldas-pcdev2 (A100, cc 8.0), cupy 12.0.0, 2026-09-17. -_GPU_KNOWN_FAILING = { - "portfolio": (_PORTFOLIO, {}, "vectorized_general_tools.py"), - "adaptive_cartesian": (["--sampler-method", "adaptive_cartesian"], dict(n_max=60000), - "integrators/mcsampler.py"), -} - - -@pytest.mark.parametrize("sampler", sorted(_GPU_KNOWN_FAILING)) -def test_the_gpu_samplers_that_do_not_work_still_fail_the_known_way(event, gpu_slot, sampler): - """A RECORDED DEFECT, not an excused one. - - Two samplers cannot run on a device at all. Skipping them would leave the gate reporting - green over a broken configuration, which is what pinning CUDA_VISIBLE_DEVICES="" did for - the whole file. So the failure is pinned by its SITE and its EXCEPTION, and this test goes - red in both directions: if someone fixes one, promote it into _GPU_SAMPLERS and delete the - entry; if the failure moves somewhere else, that is a different defect and wants reading.""" - args, kw, site = _GPU_KNOWN_FAILING[sampler] - tag = "gpu_known_fail_%s" % sampler - _d, _rc, out, have_row = _invoke_ile(event, tag, args + list(_GPU_FLAGS), - a_coeff=8.0, b_coeff=2.0, cuda=gpu_slot, **kw) - assert _DEVICE_MARKER in out, ( - "%s never reached a device, so this says nothing about the GPU path" % tag) - assert not have_row, ( - "%s SUCCEEDED on a device. If that is a fix, move %r into _GPU_SAMPLERS, delete its " - "_GPU_KNOWN_FAILING entry and the paragraph above it, and re-measure the calibration " - "table for the new lane." % (tag, sampler)) - assert "Unsupported type " in out and site in out, ( - "%s still fails on a device, but not in the recorded way: expected a cupy ufunc " - "TypeError from %s. A different failure is a different defect; read the log.\n%s" - % (tag, site, out[-2500:])) - - # --------------------------------------------------------------------------------------- # CALIBRATION # @@ -676,28 +661,52 @@ def test_the_gpu_samplers_that_do_not_work_still_fail_the_known_way(event, gpu_s # calls the A100 index 2 and an RTX 3080 index 0, while CUDA's default FASTEST_FIRST ordering # puts the A100 at 0. Check with cupy's own getDeviceProperties, not with nvidia-smi. # -# The GMM row is A=0.75 B=3 and not A=8 B=2 deliberately; see _GPU_LANE_COEFFS. Each device row -# sits on top of its host twin, which is the point: the device path is not a different answer. +# The GMM and portfolio rows are A=0.75 B=3 and not A=8 B=2 deliberately; see _GPU_LANE_COEFFS. +# Each device row sits on top of its host twin, which is the point: the device path is not a +# different answer. # # lane max |z| max sigma min n_eff # GPU prior-only, AV 1.50 0.0134 1343 # GPU prior-only, GMM 0.93 0.0134 1352 # GPU A=8 B=2, AV 2.30 0.0325 257 # GPU A=0.75 B=3, GMM 1.49 0.0235 389 +# GPU prior-only, portfolio 2.22 0.0104 2335 +# GPU A=0.75 B=3, portfolio 2.32 0.0228 411 +# GPU prior-only, adaptive_cartesian 2.39 0.0091 3262 +# GPU A=8 B=2, adaptive_cartesian 1.35 0.0305 250 +# +# Host twins, in the same order, from the table above: 2.03/0.0134/1349, 1.88/0.0134/1335, +# 1.06/0.0332/247, 1.34/0.0235/381, 1.68/0.0104/2316, 1.54/0.0227/392, -- (no host prior-only +# adaptive_cartesian lane), 1.35/0.0305/250. Sigma is identical to all four decimals on five of +# the seven twinned rows; the other two differ by 0.0007 (A=8 B=2 AV) and 0.0001 (A=0.75 B=3 +# portfolio). The |z| values differ freely, because they are draws, not constants. +# +# THE LAST ROW PRINTS AS ITS HOST TWIN, and that is agreement at the table's precision, NOT a +# bitwise identity -- do not read it as one. An earlier version of this comment claimed the +# two arms run identical host arithmetic because "the only device work is xpy.zeros". That is +# wrong and was disproved by measurement: for a sampler with return_lnL False the +# --zero-likelihood stand-in builds xpy.ones(n) * xpy.exp(supp), so cos and exp run on the +# DEVICE. cupy 12.0.0 and numpy 1.26.4 disagree bitwise on about a fifth of 200000 float64 +# draws (max relative 2e-15), and at seed 1000 this lane returns ln Z 6.646959838539267 on the +# host against 6.646957839585277 on the device -- a 2e-6 difference, four orders below the +# 0.0305 error bar and well below what these three columns show. So a last-digit move in this +# row is a rounding boundary, not a finding; a move you can see in the SECOND digit is worth +# reading. # -# Host twins for those four, from the table above: 2.03/0.0134/1349, 1.88/0.0134/1335, -# 1.06/0.0332/247, 1.34/0.0235/381. Same sigma to three digits on three of the four; the |z| -# values differ because they are draws, not constants. +# The four AV/GMM rows above were RE-DERIVED in the sweep that added the last four, i.e. after +# the two host/device fixes, and came back identical -- so those fixes do not move the lanes +# that already worked. # -# Z_TOLERANCE = 5 is 2.2x the worst |z| seen, which is now 2.30 on the GPU A=8 B=2 AV lane; +# Z_TOLERANCE = 5 is 2.1x the worst |z| seen, which is now 2.39 on the GPU prior-only +# adaptive_cartesian lane; before that it was 2.30 on GPU A=8 B=2 AV; # the worst host lane is 2.15 (distance-marginalized), re-measured at eight # seeds after the stand-in started passing xpy= to factors that accept it # (max |z| 2.15, max sigma 0.0326, least n_eff 336, unchanged, seed 1000 # bit-identical). Adding the device lanes moved the margin from 2.3x to -# 2.2x and nothing else. +# 2.1x and nothing else. # MAX_SIGMA = 0.06 is 1.7x the worst sigma seen (0.0353, host). The worst device sigma is # 0.0325. 5 * MAX_SIGMA is a 0.30-nat band. -# MIN_NEFF = 30 is 5.4x below the worst n_eff seen (163, host; 257 on device). See its +# MIN_NEFF = 30 is 5.4x below the worst n_eff seen (163, host; 250 on device). See its # comment for why it is this loose. # # WHAT THE GATE HAS TO SEPARATE A CORRECT RUN FROM. Each row was run on this fixture, not