Master's Student @ Soochow University Β· Efficient LLM Inference & KV-Cache Compression
ββββββββββ β¦ ββββββββββ
I am a Master's student at the Key Laboratory of Data Intelligence and Advanced Computing, Soochow University, supervised by Prof. Juntao Li.
My work focuses on efficient inference for large language models β from KV-cache compression to sparse attention mechanisms.
- π§ Efficient LLM Inference β pushing the limits of fast, economical LLM serving
- πΎ KV-Cache Compression β shrinking the memory footprint of long-context inference
- β‘ Sparse Attention β focusing computation on what matters most
- π¬ Model Optimization β keeping models lean without sacrificing quality
ββββββββββ β¦ ββββββββββ
"Advancing AI through innovative research and optimization."
Β© 2026 Quantong Qiu
