Cache-aware DAG scheduling and memory optimization for SIMD/NPU compute graphs, motivated by LLM inference runtime and compiler problems.
graph-algorithms memory-allocation dag-scheduling llm-inference npu-scheduling spill-selection cache-aware-scheduling
-
Updated
Jul 2, 2026 - Python