Skip to content
View qqtang-code's full-sized avatar
🎯
Focusing
🎯
Focusing

Highlights

  • Pro

Block or report qqtang-code

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
qqtang-code/README.md

Quantong Qiu

Master's Student @ Soochow University Β· Efficient LLM Inference & KV-Cache Compression

Website Google Scholar RΓ©sumΓ© Email

━━━━━━━━━━ ✦ ━━━━━━━━━━

About

I am a Master's student at the Key Laboratory of Data Intelligence and Advanced Computing, Soochow University, supervised by Prof. Juntao Li.

My work focuses on efficient inference for large language models β€” from KV-cache compression to sparse attention mechanisms.

Research

  • 🧠 Efficient LLM Inference β€” pushing the limits of fast, economical LLM serving
  • πŸ’Ύ KV-Cache Compression β€” shrinking the memory footprint of long-context inference
  • ⚑ Sparse Attention β€” focusing computation on what matters most
  • πŸ”¬ Model Optimization β€” keeping models lean without sacrificing quality

━━━━━━━━━━ ✦ ━━━━━━━━━━

"Advancing AI through innovative research and optimization."

Β© 2026 Quantong Qiu

Pinned Loading

  1. LCM-Lab/LOOM-Eval LCM-Lab/LOOM-Eval Public

    A comprehensive and efficient long-context model evaluation framework

    Python 31 4

  2. LCM-Lab/Elastic-Attention LCM-Lab/Elastic-Attention Public

    Elastic Attention: Test-time Adaptive Sparsity Ratios for Efficient Transformers

    Python 24 2

  3. FluxAttention FluxAttention Public

    Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference

    Python 7 1