KV cache for vLLM and SGLang, backed by Rust. Reuse computed prefixes from pinned DRAM and optional SSD after GPU eviction or an engine restart. Both adapters use the same native Cache Manager API. Unencoded DRAM Restore runs in the engine process using shared payload arenas; SSD/codec Restore and Publish retain Manager workers and CUDA IPC tensor bindings.
Documentation · Quickstart · Architecture · Source
- DRAM/SSD prefix reuse with engine-owned GPU allocation.
- Automatic Rust SSD backend selection for both engines, with native cuFile writes/restores where available and io_uring otherwise; see the GPU storage configuration and qualification gates.
- Compiled recovery ranges for supported attention, window and checkpoint layouts.
- Optional GPU ANS lossless compression and experimental FP8/3-bit/4-bit TurboQuant storage in DRAM, SSD and peer transfers; see storage formats.
- Native query ownership, source grants, byte budgets and completion fences.
- Engine-local raw Restore with actual tensor ownership and an explicit engine readiness stream; both adapters use the whole-operation completion gate.
- Prometheus metrics and optional request timelines.
- Experimental peer cache sharing through Mooncake TENT.
- Experimental vLLM and SGLang P/D payload transfer through the same Rust TENT runtime; SGLang retains its native handoff control plane.
Pinned engine releases: vLLM 0.29.0 and SGLang 0.5.20. The new engine-local Restore path passes single-H20 Qwen3-8B DRAM correctness and engine-restart reuse in both pinned engines. The recorded vLLM end-to-end results compare native HBM, native CPU offload, OrbitKV and LMCache MP. Multiple-GPU, huge-page and broader serving qualification remain separate. See the measured Restore improvements and remaining overhead, model qualification and deployment support. Interfaces may change before 1.0.
register_context_batch(..., tensors=[...]) takes the real tensor/exporter
objects as well as their IPC metadata. Keep tensor order aligned with layer
registration. start_restore(..., ready_stream=...) takes the CUDA stream of
the engine's previous destination-page users; the native worker establishes
readiness before copying. Both bundled adapters provide these arguments.
The native worker owns accepted operations even if a Python handle is dropped
or wait_restore times out. Connectors must retain logical destination page IDs
until a terminal result. Local completion means DMA drained and does not wait
for the Manager's source-retirement ACK. Repeated registration of the same
binding is rejected; unregister and close drain accepted operations first.
Build the native client and Manager together: this cutover uses bootstrap 7, channel ABI 11, cache schema 9, lifecycle 4 and Restore grant schema 5, with no old-wire decoder. Fragmented raw Restore plans are partitioned into at most 1 MiB parts under one final drain fence, with 32 MiB operation and 64 MiB session metadata limits. Idle destination streams need no additional GPU event; busy streams are fenced with a reusable event. Layer overlap and graph replay dependencies remain future work.
| CUDA runtime | Distribution | Python import |
|---|---|---|
| CUDA 12 | orbitkv-llm |
orbitkv |
| CUDA 13 | orbitkv-llm-cu13 |
orbitkv |
Install only one distribution and one engine extra (vllm or sglang) per
environment. The extras pin the validated engine releases. The wheel includes
the native extension, Cache Manager and Mooncake runtime; the Manager also
needs compatible PyTorch. The base package does not install PyTorch.
The host supplies matching shared Python, CUDA driver/runtime and RDMA libraries
(libibverbs, librdmacm and the selected NIC provider). Other linked native
dependencies are repaired into the wheel with their redistribution notices;
its platform tag reflects all bundled binaries.
Version 0.1.0 is being prepared for release. Build a complete wheel from the repository, then install the produced file:
git submodule update --init --recursive third-party/mooncake
./scripts/build-wheel.sh --release --no-default-features --features cuda-13,mooncake
python -m pip install /absolute/path/to/the-built-wheel.whlUse ./scripts/build-wheel.sh --release for CUDA 12. See the
installation guide
for complete engine environments and the
release guide for artifact checks.
Start an independent Manager in a compatible PyTorch/CUDA environment:
orbitkv-cache-manager --addr 127.0.0.1:50055 --pool-size 8gbTo add SSD capacity, append
--ssd-cache-path /data/orbitkv/cache.bin --ssd-cache-capacity 100gb.
The Manager automatically tries native cuFile and falls back to io_uring;
normal deployments do not need --ssd-backend or a separate cuFile service.
Engines on the same host can connect to this Manager and share its external
capacity; see deployment requirements for runtime
compatibility, shared resources and current qualification limits.
Then enable the adapter on the same host:
# vLLM
vllm serve /path/to/immutable-model --enable-prefix-caching \
--kv-transfer-config '{"kv_connector":"OrbitKVConnector","kv_role":"kv_both","kv_connector_module_path":"orbitkv.vllm"}'
# SGLang, in its own environment
ORBITKV_SGLANG_ENDPOINT=unix:///tmp/orbitkv-50055.sock \
sglang serve --model-path /path/to/immutable-model --page-size 64 \
--enable-unified-cache-external-linker --radix-cache-backend orbitkvKeep the Manager alive, restart the engine and repeat a multi-block prompt to verify an external restore. SSD contents are recreated when the Manager restarts. The engine remains responsible for HBM and scheduling. Automatic request preparation is experimental and disabled by default.
Adapter reference · Test gates · Benchmarks · Roadmap
The runtime package contains only the native client and engine adapters. Tests
live in python/tests/; workloads and results live in benches/.