Skip to content

Use one accuracy-compatible mode for model and optimizer numerics - #17

Merged
zrr1999 merged 3 commits into
PFCCLab:mainfrom
zrr1999:fix/glm45-reproducible-grad-norm
Sep 16, 2026
Merged

zrr1999 merged 3 commits into
PFCCLab:mainfrom
zrr1999:fix/glm45-reproducible-grad-norm

Conversation

@zrr1999

@zrr1999 zrr1999 commented Sep 15, 2026 •

Copy link
Copy Markdown
  • I, the PR author, have personally reviewed every line of this PR.

What does this PR do?

Use use_accuracy_compatible as the single mode for reference model numerics, AdamW selection and reproducible gradient clipping. Remove the separate norm, router, DSA, native-unfused-AdamW and reproducible-norm options. DSA-specific behavior follows the model architecture through a read-only derived property; ordinary training retains its existing fast implementations.

The reproducible L2 norm accumulates FP32 squares in integer bins using existing dense/expert ownership groups, then rounds once before the square root. Clipping remains disabled when the configured threshold is zero. Optimizer selection now happens once instead of choosing an implementation and overriding it later.

Validation:

  • Two GPU ranks: 14 norm/optimizer tests per rank, including ordinary-mode selection and an actual native AdamW update.
  • 83 focused model/configuration tests pass for DSA, normalization, routing, gradient ownership, MTP and the ordinary branches.
  • Final owning heads ran native Paddle and Torch entrypoints for 100 steps in formal-ci-followup-20260915-n100-r1: all 100 loss values are IEEE bit-identical, provenance passes, and both canonical safetensors checkpoints match. Each side also matches its prior accepted 100-step loss and checkpoint. GLM4.5 native EP2 10-step regression has 40/40 matching loss hashes and matches the previous baseline on both sides. Local receipts: experiments/ops/ci-followup-20260915/terminal.json and glm45-pair-r0/comparison.json. Frozen repository pins, existing completion-required items and the incomplete native-build snapshot remain unresolved; this is candidate acceptance, not complete + sealed or remote-CI acceptance.
  • Black, isort and Pylint pass on the changed files.

⚠️ For major changes (either in lines of code or in its impact), please make sure to first share a design doc with the team. If you're unsure what's the best way to do so, contact @NVIDIA/mcore-oncall.

Issue tracking

For PRs from open-source community contributors:

  • New features: a linked issue is required. Please open a feature request and reference it here before submitting the PR.
  • Small updates (bug fixes, minor improvements): a linked issue is recommended and will accelerate the PR review process.

Related failure: https://github.com/PaddlePaddle/PaddleFleet/actions/runs/34816312836/job/103892378045?pr=1961

Companions: Swift accuracy-mode integration and PaddleFleet GLM4.5 case. Reference wheels must include these changes before rerunning the accuracy job.

Contribution process

Pre-checks

  • I have added relevant unit tests
  • I have added relevant functional tests
  • I have added proper typing to my code Typing guidelines
  • I have added relevant documentation
  • I have run the autoformatter.sh on my PR

Code review

Feel free to message or comment @NVIDIA/mcore-oncall to help accelerate your merge into main. The less complex your PR is, the faster it will be approved and merged!

All PRs start as draft. If you open a non-draft PR, it will be automatically converted to draft.

Step 1: Mark PR as "Ready for Review"

  1. When your PR is ready, click Ready for Review.
  2. An oncall reviewer is auto-assigned and expert reviewers are notified based on your changes.
    • Some PRs may jump straight to step 2. This is determined by .github/CODEOWNERS.

⚠️ Only mark as ready once merge-conflicts are resolved and the CI is passing.
Final Review might get declined if these requirements are not fulfilled.

Step 2: Final Review

For PRs that change megatron/core, once all expert reviewers have approved, the Final Review label is applied automatically and final reviewers are assigned.

For PRs outside megatron/core, this step is skipped.

Step 3: Approved

Once all required reviewers have approved, the Approved label is applied automatically.

Merge

Any member of mcore-engineers will be able to merge your PR.

Preserve existing gradient ownership groups while making the FP32 norm independent of layout and partition order.

Signed-off-by: Zhan Rongrui <me@zrr.dev>
Signed-off-by: Zhan Rongrui <me@zrr.dev>
Signed-off-by: Zhan Rongrui <me@zrr.dev>
@zrr1999 zrr1999 changed the title Fix GLM4.5 clipping precision with an opt-in reproducible norm Use one accuracy-compatible mode for model and optimizer numerics Sep 15, 2026
@zrr1999
zrr1999 marked this pull request as ready for review September 16, 2026 02:00
@zrr1999
zrr1999 merged commit fd84191 into PFCCLab:main Sep 16, 2026
2 of 5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant