(WIP) perf(kimi3): enable batched AttnRes runtime fusion - #1146
Draft
panditsa wants to merge 1 commit into
Draft
Conversation
This was referenced Aug 19, 2026
Signed-off-by: Sanket Pandit <sanket.pandit@amd.com>
panditsa
force-pushed
the
sanket/wip-k3-attnres-runtime
branch
from
August 19, 2026 20:22
0b55ce7 to
d5fee37
Compare
panditsa
force-pushed
the
sanket/wip-k3-sharded-finalize
branch
from
August 19, 2026 20:25
8ad030d to
ab93343
Compare
Contributor
|
The exact #1132 M=8/EP8 point differs from the published M=2/4 EP1 blocker. On current main (
So the activation is beneficial at M=8/EP8, but still regresses the already-reported M=2/4 EP1 points and only moves the #1132 round ratio from 2.157x to ~2.125x. This supports redesigning/gating the orchestration by shape/topology rather than landing the current activation globally. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Validation result
This WIP is not ready to land: activating the batched fusion regresses end-to-end decode even though the isolated kernel in #1142 is faster.
4K/1K TP8/EP1 CUDA graphs, median of three:
The parent figures are the exact #1144 head. #1145 only selects its sharded tail at M>=8, so M=2/4 are unchanged at the #1146 parent boundary.
The likely issue is orchestration rather than fused-kernel latency: joining the producer stream and consuming AttnRes scratch removes overlap that the prior split path retained. Runtime activation should remain gated until the dependency/stream schedule is redesigned and reprofiled.
Tests
Stack
WIP blocker