Skip to content

Use default binding of tasks to cores in non-xstrid PE-layouts on Frontier - #8576

Merged
grnydawn merged 1 commit into
masterfrom
azamat/frontier-pes/dflt-task-distrib-non-xstrid
Jul 21, 2026
Merged

grnydawn merged 1 commit into
masterfrom
azamat/frontier-pes/dflt-task-distrib-non-xstrid

Conversation

@amametjanov

Copy link
Copy Markdown
Member

Use default binding of tasks to cores in non-xstrid PE-layouts on Frontier.

Fixes #8557

[BFB]


Testing: this reverts to prior default (cyclic/round-robin) distribution at 8 mpi tasks per node; and sets the needed --distribution=*:block at fully occupied node with 56 mpi tasks per node.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR updates the Frontier machine configuration to avoid forcing Slurm’s
--distribution=*:block in non-fully-occupied PE layouts, restoring the prior
(default/cyclic) placement behavior at lower MPI-task counts while still
enabling *:block when the node is fully occupied (motivated by
#8557).

Changes:

  • Make Frontier GPU_BIND_ARGS conditionally include --distribution=*:block
    based on the configured MAX_MPITASKS_PER_NODE.
  • Preserve --gpu-bind=closest behavior for HIP builds.

Comment thread cime_config/machines/config_machines.xml
@amametjanov
amametjanov requested a review from grnydawn July 20, 2026 21:39
@grnydawn

Copy link
Copy Markdown
Contributor

@amametjanov , have you seen this error. When I create the test case of SMS_P256.ne256pg2_ne256pg2.F2010-SCREAMv1 on your branch, I got this:

Output directory : /lustre/orion/cli115/scratch/grnydawn/e3sm_tests/AzFix
ERROR: Command: '/usr/bin/xmllint --xinclude --noout --schema /autofs/nccs-svm1_home1/grnydawn/repos/github/E3SM/cime/CIME/data/config/xml_schemas/config_machines.xsd /autofs/nccs-svm1_home1/grnydawn/repos/github/E3SM/cime_config/machines/config_machines.xml' failed with error '/autofs/nccs-svm1_home1/grnydawn/repos/github/E3SM/cime_config/machines/config_machines.xml:1855: element CMAKE_BACKEND: Schemas validity error : Element 'CMAKE_BACKEND': This element is not expected. Expected is one of ( TESTS, NTEST_PARALLEL_JOBS, BATCH_SYSTEM ).
/autofs/nccs-svm1_home1/grnydawn/repos/github/E3SM/cime_config/machines/config_machines.xml fails to validate' from dir '/autofs/nccs-svm1_home1/grnydawn/repos/github/workflow/e3sm'

The above error in config_machines.xml is not related to Frontier. However, it seems that your branch is quite far from the official master branch.

@amametjanov

Copy link
Copy Markdown
Member Author

Mm, there is this at that line 1855:

1838   <machine MACH="mappy">
...
1854     <BATCHED_BUILD>FALSE</BATCHED_BUILD>
1855     <CMAKE_BACKEND>ninja</CMAKE_BACKEND>

which is coming from PR #8532 . This branch is off latest master from today Jul-20. Is local cime submod pointing to ed580a836, and is there a local binary in externals/ninja/bin/ninja?

@grnydawn grnydawn left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tested successfully with the reported case, SMS_P256.ne256pg2_ne256pg2.F2010-SCREAMv1. The e3sm_eamxx_v1_medres test suite also passed. Not all tests in the suite were completed due to a network disconnection, but all of the tests that ran passed. Approved.

grnydawn added a commit that referenced this pull request Jul 21, 2026
…next (PR #8576)

Use default binding of tasks to cores in non-xstrid PE-layouts on Frontier.

Fixes #8557

[BFB]
@grnydawn
grnydawn merged commit 1dea29a into master Jul 21, 2026
1 check passed
@grnydawn
grnydawn deleted the azamat/frontier-pes/dflt-task-distrib-non-xstrid branch July 21, 2026 13:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Kokkos ERROR: HIP memory space failed to allocate error when running ne256 F2010-SCREAMv1 on Frontier

3 participants