Skip to content

Add GLM-5.2 Megatron reference path for PaddleFleet alignment - #4

Merged
zrr1999 merged 29 commits into
PFCCLab:mainfrom
zrr1999:glm52-bit-exact-alignment
Sep 14, 2026
Merged

zrr1999 merged 29 commits into
PFCCLab:mainfrom
zrr1999:glm52-bit-exact-alignment

Conversation

@zrr1999

@zrr1999 zrr1999 commented Jul 27, 2026 •

Copy link
Copy Markdown
  • I, the PR author, have personally reviewed every line of this PR.

What does this PR do?

为 GLM-5.2 提供可配置的 DSA / MLA / MoE / MTP 参考计算路径,供配套 ms-swift 原生 SFT 与 PaddleFleet 对齐。

  • DSA 数值改动由默认关闭的 dsa_accuracy_compatible 控制;norm、router、非融合 AdamW 也通过显式配置选择。保留其他模型原有精度兼容分支与 FP64 loss 累加默认值。
  • 支持本地层规格、TP1 autograd、专家 padding、MTP 与对应梯度归约。
  • 移除新增的通用可复现梯度范数算法,恢复上游 clipping / optimizer 实现,减少公共训练路径的改动。
  • 增加必要的生产行为测试;未修改 CI workflow、构建脚本或依赖安装流程。

Validation

在官方权重抽取的缩层、缩 expert GLM-5.2 配置上,H800 / BF16、TP2 / PP2 / EP1 / ETP1、sequence parallel,双方经原生 CLI 独立训练 100 步。100 个原始 loss、187 个最终参数逐位一致,loss / provenance / checkpoint 三项严格检查通过;两侧结果也分别与改动前的 100 步基线一致。

验证提交:Megatron-LM 76dd11a1、ms-swift 76aa97fa、PaddleFleet f85e5d7a;invocation:formal-minimal-companions-20260914-n100-r0。

本地其他模型回归:GLM-4.5 与 MiniMax 各运行 10 步,两卡 EP2、BF16、精度兼容模式。每个模型的 40 条原始 loss 哈希与基线完全一致,导出的 checkpoint 文件也完全一致(分别 413 / 407 个 tensor)。MiniMax 使用未修改的上游基线;GLM-4.5 上游无法加载本地 dense MLP,因此基线仅加入与候选相同的 MLP norm 参数名加载修复。GLM-4.5 原始基线和收窄候选各三次运行逐位稳定;恢复上游原有 MoE clone 条件后,候选与基线恢复一致。

Megatron 定向测试 78 项、Swift 定向测试 28 项通过;恢复 MoE 条件后重验受影响的 4 项测试全部通过。关闭精度兼容模式的整图基线在 TE 前向阶段失败,未获得该模式的数值回归结论。上述证据不覆盖完整模型、其他拓扑或所有模型;远端 CI 与审批需分别满足。

本地证据:experiments/ops/minimal-companion-prs-20260914/terminal.json、regression-comparison-r1-glm45-minimax.json、glm45-determinism.json。

Dependencies and integration

建议按 Megatron-LM #4 → 对应 Megatron wheel → ms-swift #3 的顺序接入,再使用匹配的 Torch 参考环境验证 PaddleFleet #1961。本地测试直接使用上述 PR 候选提交,无需先合入。跨仓精度 CI 使用发布的 wheel,源码合入不表示 wheel 已更新。

Issue tracking

For PRs from open-source community contributors:

  • New features: a linked issue is required. Please open a feature request and reference it here before submitting the PR.
  • Small updates (bug fixes, minor improvements): a linked issue is recommended and will accelerate the PR review process.

Linked issue: none.

Contribution process

Pre-checks

  • I have added relevant unit tests
  • I have added relevant functional tests
  • I have added proper typing to my code Typing guidelines
  • I have added relevant documentation
  • I have run the autoformatter.sh on my PR

Code review

Feel free to message or comment @NVIDIA/mcore-oncall to help accelerate your merge into main. The less complex your PR is, the faster it will be approved and merged!

All PRs start as draft. If you open a non-draft PR, it will be automatically converted to draft.

Step 1: Mark PR as "Ready for Review"

  1. When your PR is ready, click Ready for Review.
  2. An oncall reviewer is auto-assigned and expert reviewers are notified based on your changes.
    • Some PRs may jump straight to step 2. This is determined by .github/CODEOWNERS.

⚠️ Only mark as ready once merge-conflicts are resolved and the CI is passing.
Final Review might get declined if these requirements are not fulfilled.

Step 2: Final Review

For PRs that change megatron/core, once all expert reviewers have approved, the Final Review label is applied automatically and final reviewers are assigned.

For PRs outside megatron/core, this step is skipped.

Step 3: Approved

Once all required reviewers have approved, the Approved label is applied automatically.

Merge

Any member of mcore-engineers will be able to merge your PR.

zrr1999 and others added 25 commits July 22, 2026 20:06
backend.linear() is required when DSA is built with the local (non-TE)
spec. Return TELinear in duplicated mode, distinct from column-parallel.
TESpecProvider.linear stays TELinear, so the TE-on default path is
unchanged. LocalSpecProvider.linear is the TE-off counterpart used when
accuracy-compatible forces LocalSpecProvider for DSA/MLA down-projections.
calculate_per_token_loss=true backprops the raw sum(loss*mask). Skipping
mcore's 1/num_tokens scale under UAC left every gradient ~44x too large
at step 1 (grad_norm 2511 vs 57) and drifted IEEE at step 2. Restore the
scale; keep the fp32 gate-wgrad DP all-reduce.

Signed-off-by: Zhan Rongrui <me@zrr.dev>
Keep E-172 1/num_tokens scaling. Drop GLM-4.5 UAC CE, expert-pad,
permute/unpermute, router-bias skip, and DistributedOptimizer AdamW
gate from main so the GLM-5.2 IEEE graph stays on 4e81432. Take CI
workflows and the dependence build script.

Signed-off-by: Zhan Rongrui <me@zrr.dev>
EP=1 / ETP=1 / TP=2 keeps a full expert replica on each tensor-parallel
rank while each rank only sees its sequence-parallel shard. The EP-only
allreduce flag left those partial sums in the dense dp_cp bucket (size 1).
Match the E-811 e495 helper so expert grads reduce over expt_dp.

Signed-off-by: Zhan Rongrui <me@zrr.dev>
Keep pull_request on the historical CodeSync/develop tarball and
develop/latest wheels. workflow_dispatch can opt into fail-closed
SHA checkout plus artifact digest checks; pairing remains unproven
unless this job built the wheels from the pin. MinimaxV2.5_EP2 and
GLM45Air_EP2 are unchanged.
workflow_dispatch left PR_ID empty so checkout stayed on the BOS main
tarball and the next exec defaulted to paddlefleet-0.0.0.whl before
setup_venvs. Checkout COMMIT_ID, export source trees from the pin
selector, and consume that env in the alignment step without the 0.0.0
fallback.
Do not let a leftover develop env overwrite the caller. Refuse unproven
develop/latest fallback, not a verified 0.0.0 artifact name. Handoff
fixture remains a path-pass, not a uv install.
Selector clone failure must fail the docker exec (set -e) instead of
continuing wget/build. Receipt checks live in a helper so the
single-quoted -c script has no nested quotes. Missing pin.env after an
error receipt is selector failure, not env-handoff.
stack-paired checkout_pin used a full default-branch git clone with no
retry. Swift 34038640242 failed Get Whl on curl 56 / early EOF before
ops, so Fleet 09bb4bd4 was not evaluated. Fetch origin $PIN_SHA at
--depth=1 --no-tags, detach-checkout, and keep exact HEAD. Transient
RPC/curl-56 retries up to 3 clean dests; HTTP 401/403, missing SHA,
and missing repo fail closed on attempt 1 with commit_verified=false.
Bound source submodule setup before environment preparation. Keep wheel installs on their existing path.

Signed-off-by: Zhan Rongrui <me@zrr.dev>
Keep FP32 accuracy routing, shared expert ordering and native expert accumulation behavior consistent with the validated candidates.

Signed-off-by: Zhan Rongrui <me@zrr.dev>
Retain embedding gradient accumulation and TP1 autograd structure in accuracy mode, with focused native behavior tests.

Signed-off-by: Zhan Rongrui <me@zrr.dev>
Preserve MTP loss attachment, rotary indexing, transformer exit and attention graph behavior in native accuracy mode.

Signed-off-by: Zhan Rongrui <me@zrr.dev>
Signed-off-by: Zhan Rongrui <me@zrr.dev>
…ions

Accumulate FP32 squares in exact integer bins on device, reduce over existing owner groups, and round once before FP32 clipping.

Signed-off-by: Zhan Rongrui <me@zrr.dev>
@zrr1999 zrr1999 changed the title Add gated GLM-5.2 bit-exact compatibility paths Add GLM-5.2 Megatron reference path for PaddleFleet alignment Sep 14, 2026
@zrr1999
zrr1999 marked this pull request as ready for review September 14, 2026 09:03
@zrr1999
zrr1999 merged commit dadaebb into PFCCLab:main Sep 14, 2026
2 of 5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant