Skip to content

fix: support MiniMax-M3 pipeline parallelism in PD disaggregation - #36095

Draft
shiyang814-cpu wants to merge 2 commits into
sgl-project:mainfrom
shiyang814-cpu:fix/minimax-m3-pp-disagg
Draft

fix: support MiniMax-M3 pipeline parallelism in PD disaggregation#36095
shiyang814-cpu wants to merge 2 commits into
sgl-project:mainfrom
shiyang814-cpu:fix/minimax-m3-pp-disagg

Conversation

@shiyang814-cpu

@shiyang814-cpu shiyang814-cpu commented Aug 23, 2026

Copy link
Copy Markdown
Contributor
  • Preserve global sparse layer IDs for MiniMax index-K state registration.
  • Map sparse index transfer entries by global layer ID under pipeline parallelism.
  • Normalize scattered PP proxy tensors during Decode CUDA Graph capture.
  • Add regression tests for sparse state registration and PP proxy shapes.
  • Verified with TP4/PP2/EP4 on two 8×H20 nodes using Mooncake PD.
  • Decode CUDA Graph replay confirmed with batches [4, 8].

Checklist

Review and Merge Process

  1. Ping Merge Oncalls to start the process. See the PR Merge Process.
  2. Get approvals from CODEOWNERS and other reviewers.
  3. Trigger CI tests with comments or contact authorized users to do so.
    • Common commands include /tag-and-rerun-ci, /tag-run-ci-label, /rerun-failed-ci
  4. After green CI and required approvals, ask Merge Oncalls or people with Write permission to merge the PR.

CI States

Latest PR Test (Base): ❌ Run #32647732433
Latest PR Test (Extra): ❌ Run #32647732272
Latest PR Test (AMD ROCm 7.2): ❌ Run #32647732378

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant