--distance-marginalization --sampler-method AV without --force-xpy raises a TypeError on every intrinsic point, and the driver exits 0 having produced no output files at all. To a DAG node, that is a success.
Reproduced independently on two events, two data types and two flag sets, by two sessions.
The exception
File ".../RIFT/integrators/mcsamplerAdaptiveVolume.py", line 1723, in integrate_log
lnL = _eval_integrand(rv)
File ".../RIFT/integrators/mcsamplerAdaptiveVolume.py", line 1718, in _eval_integrand
return lnF(**dict(list(zip(self.params_ordered, samples.T))))
TypeError: analyze_event.<locals>.likelihood_function() missing 1 required positional argument: 'distance'
params_ordered has dropped distance because the sampler marginalizes it, while the likelihood closure built on the non-xpy path still requires it as a keyword. The two halves disagree about whether distance is a sampled dimension.
Why it exits 0
integrate_likelihood_extrinsic_batchmode:5549-5554 — the per-point except classifies the error, prints, zeroes the extrinsic parameters and moves to the next point:
else:
print( " Cause not classified -- read the traceback above; it names the failing call.")
print( " Skipping the following binary! ")
P_list[indx].incl = P_list[indx].tref = P_list[indx].dist = ... = 0
That is reasonable when one point of many fails. It is not reasonable when it is every point: the loop ends, nothing is written, and the process returns 0. There is no "all points were skipped" check.
Two reproductions
|
data |
flags |
result |
| this session |
real O4 strain, S250114ax H1+L1, --d-prior Euclidean, --n-events-to-analyze 1 |
--distance-marginalization + lookup table, no --force-xpy |
3/3 seeds: TypeError x3, EXIT=0, output files NONE (only run.log) |
| paper-1 §V.A rebuild session |
ladder-2 zero-noise injection, different event |
--distance-marginalization, neither --gpu nor --force-xpy |
same traceback, EXIT CODE = 0, output files NONE |
So it is not specific to real data, to one event, or to my particular flag combination. The common factor is distance marginalization on the AV sampler without --force-xpy.
Why the exit code is the real defect
The TypeError is a bug worth fixing on its own, but a wrong answer that announces itself is cheap. A DAG node that exits 0 with no output is expensive: it is indistinguishable from success to condor_dagman, to a wrapper script, and to any completeness check that keys on the exit code. The §V.A rebuild session reports having already lost 38 of 87 runs to OOM kills that a wrapper reported as DONE, which is the same class through a different door; they now audit output completeness rather than exit codes, which is the workaround this defect forces on everyone.
Compounding it: nothing in the name --force-xpy suggests it is load-bearing for distance marginalization. A user who marginalizes distance and does not pass it gets a silent no-op run.
Suggested direction (not implemented here)
- Fix the disagreement. Build the likelihood closure's signature from the same
params_ordered the sampler will use, so the two cannot drift. Whichever path is right, they should not disagree about whether distance is sampled.
- Exit non-zero when every point was skipped. Count skips against points analyzed; if they are equal, fail. This is one counter and it converts a silent no-op into a rerun.
- Consider refusing the combination at parse time if
--force-xpy really is required for it — a SystemExit before the precompute costs seconds instead of minutes and names the fix.
(1) and (2) are diagnostics. (3) changes behaviour and is RO'S's call.
Credit where it is due
The handler's " Cause not classified -- read the traceback above; it names the failing call." is honest and is what made this diagnosable in one pass. An earlier version printed a hardcoded list of probable reasons that did not include the actual one, and that cost real time. Please keep the current wording.
Related
Run logs and status files: ~/pr223_e2e/S250114ax/out_DM_300{1,2,3}/ on CIT; record in RIFT_roboto_paper analyses/limit_distance_e2e/.
🤖 Generated with Claude Code
--distance-marginalization --sampler-method AVwithout--force-xpyraises aTypeErroron every intrinsic point, and the driver exits 0 having produced no output files at all. To a DAG node, that is a success.Reproduced independently on two events, two data types and two flag sets, by two sessions.
The exception
params_orderedhas droppeddistancebecause the sampler marginalizes it, while the likelihood closure built on the non-xpypath still requires it as a keyword. The two halves disagree about whether distance is a sampled dimension.Why it exits 0
integrate_likelihood_extrinsic_batchmode:5549-5554— the per-pointexceptclassifies the error, prints, zeroes the extrinsic parameters and moves to the next point:That is reasonable when one point of many fails. It is not reasonable when it is every point: the loop ends, nothing is written, and the process returns 0. There is no "all points were skipped" check.
Two reproductions
--d-prior Euclidean,--n-events-to-analyze 1--distance-marginalization+ lookup table, no--force-xpyTypeErrorx3,EXIT=0, output files NONE (onlyrun.log)--distance-marginalization, neither--gpunor--force-xpyEXIT CODE = 0, output files NONESo it is not specific to real data, to one event, or to my particular flag combination. The common factor is distance marginalization on the AV sampler without
--force-xpy.Why the exit code is the real defect
The
TypeErroris a bug worth fixing on its own, but a wrong answer that announces itself is cheap. A DAG node that exits 0 with no output is expensive: it is indistinguishable from success tocondor_dagman, to a wrapper script, and to any completeness check that keys on the exit code. The §V.A rebuild session reports having already lost 38 of 87 runs to OOM kills that a wrapper reported as DONE, which is the same class through a different door; they now audit output completeness rather than exit codes, which is the workaround this defect forces on everyone.Compounding it: nothing in the name
--force-xpysuggests it is load-bearing for distance marginalization. A user who marginalizes distance and does not pass it gets a silent no-op run.Suggested direction (not implemented here)
params_orderedthe sampler will use, so the two cannot drift. Whichever path is right, they should not disagree about whetherdistanceis sampled.--force-xpyreally is required for it — aSystemExitbefore the precompute costs seconds instead of minutes and names the fix.(1) and (2) are diagnostics. (3) changes behaviour and is RO'S's call.
Credit where it is due
The handler's
" Cause not classified -- read the traceback above; it names the failing call."is honest and is what made this diagnosable in one pass. An earlier version printed a hardcoded list of probable reasons that did not include the actual one, and that cost real time. Please keep the current wording.Related
--gpu --force-xpyand a sampled distance dimension). This is the neighbouring cell of that 2x2, and the only one that does not run.mcsamplerGPUreporting a broken CUDA stack as a missing package. Same theme: a failure mode that presents as something else, or as nothing.Run logs and status files:
~/pr223_e2e/S250114ax/out_DM_300{1,2,3}/on CIT; record in RIFT_roboto_paperanalyses/limit_distance_e2e/.🤖 Generated with Claude Code