Add GLM-5.2 Megatron reference path for PaddleFleet alignment - #4
Merged
Merged
Conversation
backend.linear() is required when DSA is built with the local (non-TE) spec. Return TELinear in duplicated mode, distinct from column-parallel.
TESpecProvider.linear stays TELinear, so the TE-on default path is unchanged. LocalSpecProvider.linear is the TE-off counterpart used when accuracy-compatible forces LocalSpecProvider for DSA/MLA down-projections.
calculate_per_token_loss=true backprops the raw sum(loss*mask). Skipping mcore's 1/num_tokens scale under UAC left every gradient ~44x too large at step 1 (grad_norm 2511 vs 57) and drifted IEEE at step 2. Restore the scale; keep the fp32 gate-wgrad DP all-reduce. Signed-off-by: Zhan Rongrui <me@zrr.dev>
Keep E-172 1/num_tokens scaling. Drop GLM-4.5 UAC CE, expert-pad, permute/unpermute, router-bias skip, and DistributedOptimizer AdamW gate from main so the GLM-5.2 IEEE graph stays on 4e81432. Take CI workflows and the dependence build script. Signed-off-by: Zhan Rongrui <me@zrr.dev>
EP=1 / ETP=1 / TP=2 keeps a full expert replica on each tensor-parallel rank while each rank only sees its sequence-parallel shard. The EP-only allreduce flag left those partial sums in the dense dp_cp bucket (size 1). Match the E-811 e495 helper so expert grads reduce over expt_dp. Signed-off-by: Zhan Rongrui <me@zrr.dev>
Keep pull_request on the historical CodeSync/develop tarball and develop/latest wheels. workflow_dispatch can opt into fail-closed SHA checkout plus artifact digest checks; pairing remains unproven unless this job built the wheels from the pin. MinimaxV2.5_EP2 and GLM45Air_EP2 are unchanged.
workflow_dispatch left PR_ID empty so checkout stayed on the BOS main tarball and the next exec defaulted to paddlefleet-0.0.0.whl before setup_venvs. Checkout COMMIT_ID, export source trees from the pin selector, and consume that env in the alignment step without the 0.0.0 fallback.
Do not let a leftover develop env overwrite the caller. Refuse unproven develop/latest fallback, not a verified 0.0.0 artifact name. Handoff fixture remains a path-pass, not a uv install.
Selector clone failure must fail the docker exec (set -e) instead of continuing wget/build. Receipt checks live in a helper so the single-quoted -c script has no nested quotes. Missing pin.env after an error receipt is selector failure, not env-handoff.
stack-paired checkout_pin used a full default-branch git clone with no retry. Swift 34038640242 failed Get Whl on curl 56 / early EOF before ops, so Fleet 09bb4bd4 was not evaluated. Fetch origin $PIN_SHA at --depth=1 --no-tags, detach-checkout, and keep exact HEAD. Transient RPC/curl-56 retries up to 3 clean dests; HTTP 401/403, missing SHA, and missing repo fail closed on attempt 1 with commit_verified=false.
Bound source submodule setup before environment preparation. Keep wheel installs on their existing path. Signed-off-by: Zhan Rongrui <me@zrr.dev>
Keep FP32 accuracy routing, shared expert ordering and native expert accumulation behavior consistent with the validated candidates. Signed-off-by: Zhan Rongrui <me@zrr.dev>
Retain embedding gradient accumulation and TP1 autograd structure in accuracy mode, with focused native behavior tests. Signed-off-by: Zhan Rongrui <me@zrr.dev>
Preserve MTP loss attachment, rotary indexing, transformer exit and attention graph behavior in native accuracy mode. Signed-off-by: Zhan Rongrui <me@zrr.dev>
Signed-off-by: Zhan Rongrui <me@zrr.dev>
Signed-off-by: Zhan Rongrui <me@zrr.dev>
This was referenced Sep 9, 2026
Draft
Closed
Signed-off-by: Zhan Rongrui <me@zrr.dev>
…ions Accumulate FP32 squares in exact integer bins on device, reduce over existing owner groups, and round once before FP32 clipping. Signed-off-by: Zhan Rongrui <me@zrr.dev>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
为 GLM-5.2 提供可配置的 DSA / MLA / MoE / MTP 参考计算路径,供配套 ms-swift 原生 SFT 与 PaddleFleet 对齐。
dsa_accuracy_compatible控制;norm、router、非融合 AdamW 也通过显式配置选择。保留其他模型原有精度兼容分支与 FP64 loss 累加默认值。Validation
在官方权重抽取的缩层、缩 expert GLM-5.2 配置上,H800 / BF16、TP2 / PP2 / EP1 / ETP1、sequence parallel,双方经原生 CLI 独立训练 100 步。100 个原始 loss、187 个最终参数逐位一致,loss / provenance / checkpoint 三项严格检查通过;两侧结果也分别与改动前的 100 步基线一致。
验证提交:Megatron-LM
76dd11a1、ms-swift76aa97fa、PaddleFleetf85e5d7a;invocation:formal-minimal-companions-20260914-n100-r0。本地其他模型回归:GLM-4.5 与 MiniMax 各运行 10 步,两卡 EP2、BF16、精度兼容模式。每个模型的 40 条原始 loss 哈希与基线完全一致,导出的 checkpoint 文件也完全一致(分别 413 / 407 个 tensor)。MiniMax 使用未修改的上游基线;GLM-4.5 上游无法加载本地 dense MLP,因此基线仅加入与候选相同的 MLP norm 参数名加载修复。GLM-4.5 原始基线和收窄候选各三次运行逐位稳定;恢复上游原有 MoE clone 条件后,候选与基线恢复一致。
Megatron 定向测试 78 项、Swift 定向测试 28 项通过;恢复 MoE 条件后重验受影响的 4 项测试全部通过。关闭精度兼容模式的整图基线在 TE 前向阶段失败,未获得该模式的数值回归结论。上述证据不覆盖完整模型、其他拓扑或所有模型;远端 CI 与审批需分别满足。
本地证据:
experiments/ops/minimal-companion-prs-20260914/terminal.json、regression-comparison-r1-glm45-minimax.json、glm45-determinism.json。Dependencies and integration
建议按 Megatron-LM #4 → 对应 Megatron wheel → ms-swift #3 的顺序接入,再使用匹配的 Torch 参考环境验证 PaddleFleet #1961。本地测试直接使用上述 PR 候选提交,无需先合入。跨仓精度 CI 使用发布的 wheel,源码合入不表示 wheel 已更新。
Issue tracking
For PRs from open-source community contributors:
Linked issue: none.
Contribution process
Pre-checks
Code review
Feel free to message or comment @NVIDIA/mcore-oncall to help accelerate your merge into main. The less complex your PR is, the faster it will be approved and merged!
All PRs start as draft. If you open a non-draft PR, it will be automatically converted to draft.
Step 1: Mark PR as "Ready for Review"
.github/CODEOWNERS.Final Review might get declined if these requirements are not fulfilled.
Step 2: Final Review
For PRs that change
megatron/core, once all expert reviewers have approved, theFinal Reviewlabel is applied automatically and final reviewers are assigned.For PRs outside
megatron/core, this step is skipped.Step 3: Approved
Once all required reviewers have approved, the
Approvedlabel is applied automatically.Merge
Any member of mcore-engineers will be able to merge your PR.