C++ / CUDA / LLM inference systems · Maintainer of @open-infra-ai
- shenzhen
-
14:41
(UTC +08:00) - https://github.com/open-infra-ai
Pinned Loading
-
open-infra-ai/paged-infer
open-infra-ai/paged-infer PublicRust LLM Serving 控制面:Paged KV、continuous batching、OpenAI SSE、C ABI 后端与可信压测
Rust 1
-
open-infra-ai/cuflash-attn
open-infra-ai/cuflash-attn PublicCUDA C++ FlashAttention/FlashDecoding 专项:FP32/FP16/BF16、WMMA、差分测试与基准
Cuda
-
open-infra-ai/tiny-llm
open-infra-ai/tiny-llm PublicCUDA/C++ LLM 推理运行时:GGUF、W8A16、tokenizer、Paged KV、CUDA Graph 与可复现基准
C++
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.




