Use default binding of tasks to cores in non-xstrid PE-layouts on Frontier - #8576
Conversation
There was a problem hiding this comment.
Pull request overview
This PR updates the Frontier machine configuration to avoid forcing Slurm’s
--distribution=*:block in non-fully-occupied PE layouts, restoring the prior
(default/cyclic) placement behavior at lower MPI-task counts while still
enabling *:block when the node is fully occupied (motivated by
#8557).
Changes:
- Make Frontier
GPU_BIND_ARGSconditionally include--distribution=*:block
based on the configuredMAX_MPITASKS_PER_NODE. - Preserve
--gpu-bind=closestbehavior for HIP builds.
|
@amametjanov , have you seen this error. When I create the test case of The above error in |
|
Mm, there is this at that line 1855: which is coming from PR #8532 . This branch is off latest master from today Jul-20. Is local cime submod pointing to |
grnydawn
left a comment
There was a problem hiding this comment.
Tested successfully with the reported case, SMS_P256.ne256pg2_ne256pg2.F2010-SCREAMv1. The e3sm_eamxx_v1_medres test suite also passed. Not all tests in the suite were completed due to a network disconnection, but all of the tests that ran passed. Approved.
Use default binding of tasks to cores in non-xstrid PE-layouts on Frontier.
Fixes #8557
[BFB]
Testing: this reverts to prior default (cyclic/round-robin) distribution at 8 mpi tasks per node; and sets the needed
--distribution=*:blockat fully occupied node with 56 mpi tasks per node.