Repository navigation
Conversation
sherylll
commented
Oct 6, 2026
- Since u4 MMA only exists in a limited range of architectures, switch to native u8 MMA for better compatibility. Perf is not negatively affected because the kernel is latency bound.
- In MMA kernels, add warp guard on MMA instructions for elements past the useful range (new_size/old_size)
Skip warps whose output block is out of range; Remove redundant syncthreads in local_join_kernel_wmma Simplify BBQ staging function
22c0d30 to
3e6a7bb
Compare
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configuration
📒 Files selected for processing (2)
💤 Files with no reviewable changes (1)
Included review availability: This review used your included allowance. Your plan provides up to 4 included reviews per hour; 3 remain after this review. 📝 SummarySummary by CodeRabbit
WalkthroughNN-descent graph insertion now uses a fixed device-side degree. WMMA joins skip inactive tiles. The BBQ path derives self-join behavior from layouts and uses byte-fragment MMA to accumulate packed low and high nibbles. ChangesNN-descent kernels
Priority: ➖ Normal Estimated code review effort: 4 (Complex) | ~45 minutes Change: Refactor Suggested reviewers: Merge Risk: ⚪ Minimal · up to No specific issue has been established that would prevent merging after normal CUDA build and runtime checks. 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches 💡 1🛠️ Fix failing CI checks 💡
🧪 Generate unit tests (beta)
Comment |