Repository navigation
Conversation
Signed-off-by: Zhan Rongrui <me@zrr.dev>
zrr1999
marked this pull request as ready for review
September 16, 2026 10:18
Draft
1 of 6 tasks
Apply the equivalent of PFCCLab/Megatron-LM commit 8aa1d8a30b1983debbae1bdea1a080d7eebe01d8 to the Swift accuracy workflow. Signed-off-by: Zhan Rongrui <me@zrr.dev>
morirun
reviewed
Sep 17, 2026
morirun
left a comment
There was a problem hiding this comment.
温和核对:
- 去掉 accuracy 模式下强制
clip_grad = 0的意图清楚;get_optimizer_and_scheduler对get_reproducible_grad_norm_bins的硬依赖需要配对的 Megatron-Core(见 PFCCLab/Megatron-LM#21)先落地或可探测,否则本地/CI 会直接ValueError。 - 当前 CI:
lint红(yapf / trailing-whitespace / end-of-file-fixer / double-quote-string-fixer);其中部分 hook 改到了本 PR 未触及的文件(如scripts/dependence/build.sh、swift/megatron/trainers/trainer.py),建议在分支上跑一遍pre-commit run --all-files再推,避免基线漂移干扰。 unittest日志里是 runnersudo/tty 基础设施失败,未必是本 diff 逻辑问题;alignment_model_accuracy镜像从 cu130 改到 cu129 请确认与 ops wheel 路径一致且为有意变更。
不代推;合入与否由维护者决定。
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
PR type
PR information
Preserve the caller's
clip_gradduring accuracy-mode SFT initialization. A requested positive threshold previously became zero before optimizer construction, silently disabling gradient clipping. Require companion Megatron reproducible-norm support when accuracy mode requests positive clipping, so an incompatible wheel fails explicitly.Retain all existing configuration names and model-specific switches, including
native_unfused_adamw. No switch renaming, model-path rewrite, or comparison-tolerance change is included. Related: #12, #13 and PaddleFleet #1961.Sync the accuracy CI image and PaddleFleet ops wheel download to CUDA 12.9, matching Megatron-LM commit 8aa1d8a. This changes exactly the image tag and the ops wheel CUDA directory. YAML parsing and all eight embedded shell blocks passed syntax checks. Torch installation settings remain controlled by the shared Fleet setup script. The new CI run must pass before merge.
Experiment results
Seven focused tests passed: real SFT construction preserves thresholds 0, 0.25 and 1 in both modes; an incompatible Megatron clipping implementation fails explicitly. Flake8, isort, YAPF and diff checks passed. The companion Megatron tests passed 14 cases on each of two GPU ranks. The previous head
dbf01b04passed GPU precision CI. The new CI-only head requires its own native results.Paired local validation uses Megatron
6b4771e50, Swiftd3d0b1a87, and Fleetd889f0a1: the original GLM52 CI profile passed 100 bitwise-identical loss steps and 187 identical canonical checkpoint tensors; the acceptance-profile native entrypoints separately passed 100 steps with strict loss, provenance and checkpoint oracles. Both sides also match their prior local baselines. The original two-rank GLM45 10-step regression withclip_grad=1.0passed all 40 raw per-token/final-loss hash records and both same-side baseline comparisons.These runs use local Paddle
3.4.0.post20260808+733f3454aa0and Torch2.12.1+cu129; they do not establish remote CI equivalence. Evidence:experiments/ops/restore-pair-20260916/{ci-terminal,terminal}.jsonandexperiments/ops/restore-pair-regression-20260916/glm45-pair-r0/{protocol,source-binding,comparison}.json.Companion PR: PFCCLab/Megatron-LM#21.
The GPU unit workflow now uses a checkout unique to the run and attempt. Docker receives a read-only source mount and runs the existing tests in a private copy; cleanup removes only the new checkout. Old root-owned workspace files no longer require a privileged recursive repair. CI-script edits trigger GPU tests, and hosted lint uses pinned Node 24 actions. The legacy GPU host uses native Git for the exact event commit because its glibc cannot run Node 24.
Four actual filesystem regressions, Actionlint, wrapper ShellCheck, shell syntax checks, full-repository pre-commit and actual commit hooks pass. Failures retain their exit codes and private copies are removed; host source and Git metadata remain unchanged. Only Git commit signing was disabled. Current native GPU unit/accuracy and NPU results remain required, with no test or comparison threshold reduced.
Git LFS is initialized only in the disposable checkout. If the binary is missing, the job verifies the official v3.8.0 archive SHA-256 before using a temporary binary, then fetches, materializes and checks the LFS objects. The current GPU unit job fails closed because the host has Git 1.8.3.1, while Git LFS requires Git >= 2.0. An earlier diagnostic run also established that the required
/mnt/modelscope/ci_env.shis absent. A compatible, correctly configured runner is required; these are infrastructure blockers, not passing unit tests. The NPU job remains required.Latest-head
a4736be3model-accuracy alignment passed (job111320729675): all recorded ranks and steps match for per-token and final loss MD5. The GPU unit runner and NPU requirements above remain open.