Repository navigation
[drivers] master-only raise hangs all other ranks instead of failing the job #24
Description
Activity
We hit this in solid_dmft years ago and it keeps coming back, so let me argue for fixing it
one level down, in triqs core, rather than in dftkit — and for not writing the abort
machinery ourselves at all, because mpi4py already ships it.Why not in dftkit
The pattern isn't ours alone:
is_master_nodeappears ~58× in dftkit, ~29× in dft_tools and
~90× in solid_dmft. The ten explicitraises listed above are only the visible tip — any
exception in any of those blocks (aKeyError, an h5 read, a numpy error) hangs the job in
exactly the same way, and those can't be enumerated.dft_tools/sumk_dft_transport.pyalready
hand-rolls anMPI.COMM_WORLD.Abort(1), solid_dmft has four more, and dftkit carries a third
variant invasp/plovasp/sc_dmft.py. Four copies of the same idea, none shared.triqs.utility.mpiis also the only layer that knows whether we're running under mpi4py or
under the serialmpi_nompistub, so it's the only place the no-MPI fallback gets handled once.mpi4py already solved this
mpi4py.run.set_abort_status(1)(whatpython -m mpi4py script.pyuses internally) marks the
job so that mpi4py callsMPI_Abortinstead ofMPI_Finalizewhen the interpreter exits.That deferral is the whole point, and it is strictly better than calling
Abort()from inside
an excepthook the way all our current copies do: the traceback gets printed,finallyblocks
andatexithandlers run, stdio is flushed — then the job is torn down. Verified: an
atexithandler prints before OpenMPI'sMPI_ABORT was invoked...banner. So the
time.sleep(2)insolid_dmft/bin/solid_dmft.py, and the comment above it ("this sometimes
weirdly suppresses error output completely"), both become unnecessary.It also means the VASP-driver cleanup can move into a
finally/atexitand is then
guaranteed to run before the abort — instead of today's "remember to callDriver.kill()
before aborting", which is how you end up with an orphanedmpirun vaspstill burning nodes.Concrete proposal for
triqs.utility.mpiTwo functions in
mpi_mpi4py.py, with no-op/identity twins inmpi_nompi.py:def abort_on_exception(all_ranks_report=False): """Make an unhandled exception on *any* rank terminate the whole MPI job. Full traceback on master, one line on the other ranks, so an error that hits all ranks doesn't produce `size` interleaved tracebacks. Opt-in: applications call this from their entry point, libraries never do.""" from mpi4py.run import set_abort_status previous = sys.excepthook def _hook(typ, value, tb): if is_master_node() or all_ranks_report: previous(typ, value, tb) else: print(f'[rank {rank}] {typ.__name__}: {str(value).splitlines()[0]}', file=sys.stderr) set_abort_status(1) # abort at interpreter exit, after cleanup sys.excepthook = _hook @contextmanager def raise_on_all_ranks(comm=world): """Turn a failure inside a rank-local region into a raise on every rank.""" error = None try: yield except Exception: error = f'rank {comm.Get_rank()}:\n{traceback.format_exc()}' errors = [e for e in comm.allgather(error) if e is not None] if errors: raise MPIError('failed on a subset of ranks --\n' + '\n'.join(errors)) from None def on_master(fn, *args, **kwargs): """Function-shaped variant; broadcasts fn's return value.""" with raise_on_all_ranks(): result = fn(*args, **kwargs) if is_master_node() else None return bcast(result)
Note
allgather, notbcastfrom root 0: a failure on rank 2 is caught just as well as one on
the master, and the master's traceback then names the rank that actually failed.Tested
Prototyped against triqs unstable, 4 ranks, OpenMPI 5.0.10 under
mpirun, reproducing the
_run_plo_convertershape (master-only failure, other ranks inmpi.barrier(poll_msec=100)):case before with the proposal master-only raise, unwrapped hangs (killed at 25 s) exit 1 in <1 s, 1 traceback master-only raise, wrapped in raise_on_all_ranks— exit 1 in <1 s, master shows the failing rank's stack failure on rank 2 only, wrapped hangs exit 1 in ~1 s, rank 2's traceback shown on master raise on all ranks dies, 4 interleaved tracebacks exit 1, 1 traceback + 3 one-liners no failure at all — exit 0, no spurious abort serial run, no launcher ( mpi_nompi)— unchanged plain-Python traceback, exit 1 I also confirmed the current
solid_dmft/bin/solid_dmft.pyhook does not help here: it
aborts only on the non-master ranks, and in this scenario they never raise — they sit in the
barrier. The master prints its traceback and then blocks in mpi4py's atexitMPI_Finalize,
which is collective. Still a hang, measured.What dftkit then does
# entry point, once mpi.abort_on_exception()
and the ten sites keep their shape:
with mpi.raise_on_all_ranks(): if mpi.is_master_node(): ...
Three notes on the proposal above
with master_only():cannot work as written. A context manager cannot skip its own body
in Python —__enter__has no way to say "don't run the block" short ofsys.settrace
hacks. Hence the shape above: theif mpi.is_master_node():stays inside thewith, and the
context manager only handles the failure propagation. Same minimal diff at the call sites.- Don't stringify the exception.
f"{type(exc).__name__}: {exc}"discards the traceback,
which is the one thing you actually want from a remote rank.traceback.format_exc(). MPIHandleris the wrong home even after dedup: it's a dataclass about launching
subprocesses (mpi_exec, env vars), and wien2k uses neither copy. Dedupe it on its own
merits; MPI error policy belongs in triqs core.
And a helper alone won't close this: it only covers the sites someone remembered to wrap. The
excepthook is the backstop that catches the other ~170 master-only blocks;raise_on_all_ranks
is for the failures that should stay catchable by a caller (a notebook, dft_tools driving the
converter) — those never see an application's excepthook.Happy to open the triqs core issue/PR with the above if this direction looks right.
(mpi4py.run.set_abort_statushas been there since mpi4py 3.0; trivial to guard with a
fallback to a plainAbort(1)if we care about older.)🤖 Analysis, prototype and measurements above generated with Claude Code
All three of the DFT drivers raise on rank 0 inside a master-only block. The master
correctly throws while the other ranks are left behind. This causes the job to hang
instead of properly erroring. From
vasp/driver.py:186:wien2k/driver.py_check_inputs,_run_dmftproj,_run_scf,_write_qdmft,read_dft_energyvasp/driver.py_wait_for_vasp,_run_plo_converter,read_dft_energyqe/driver.py_run_qe_step,_run_w90_stepFix: a shared helper that runs the fallible part on the master, broadcasts
whether it failed, and raises on every rank.
A
with master_only():context manager would suit the block-shaped Wien2k sitesbetter. It needs a home:
MPIHandleris duplicated invasp/driver.py:34andqe/driver.py:27, and Wien2k uses neither.This is low priority. It takes a real error or ranks disagreeing about
shared filesystem state, and it hangs rather than corrupting anything.
This was brought up in review of #22 (point 3.2, @the-hampel). We opened
this issue to document.