Website · Documentation · Quickstart · Architecture · Completion plan
OrbitKV extends the KV cache of vLLM and SGLang beyond GPU memory. Keep reusable prefixes in DRAM and SSD, then restore them when a matching request arrives. This helps workloads with repeated documents, shared system prompts, and conversations whose prefixes no longer fit in the engine's GPU cache.
Run an independent Cache Manager per host and connect the engines on that host to its shared cache. Engines own GPU memory and scheduling; OrbitKV manages external replicas and transfers. See deployment patterns for shared-instance budgets and container qualification limits. The pinned engine baselines are vLLM 0.29.0 and SGLang 0.5.20. The recorded engine-local raw Restore cutover passes single-H20 Qwen3-8B DRAM serving correctness and restart reuse in both engines. The recorded vLLM end-to-end comparison shows gains over HBM-eviction recomputation, while native CPU offload remains faster. LMCache comparisons and their limits are recorded in the same report. The matched communication measurements track the initial regression and the subsequent idle-stream/plan-compaction optimization. Dense transfers benefit; small-payload overhead and broader serving/deployment qualification remain open. Multi-node cache sharing is experimental. Interfaces may change before 1.0.
- DRAM and SSD caching. Reuse prefixes after GPU eviction or an engine restart while the Cache Manager remains alive.
- Optional reuse policies. Rust can protect reused pages within a byte cap and admit SSD writes selectively. See the policy controls and their cold-reuse tradeoff before enabling them.
- Native GPU transfers. Unencoded DRAM Restore executes inside the engine using independently imported shared payload arenas and retained tensors. Per-layer CUDA dependencies let consumers start before later copies finish; source and destination ownership lasts through the final drain. SSD/codec Restore and Publish retain Manager workers and CUDA IPC bindings; adapters supply the CUDA stream dependencies and Rust owns completion.
- Model-aware recovery. Cache identity includes model artifacts, computation settings and storage layout. Compiled recovery rules select the required attention pages, sliding windows and recurrent/conv checkpoints, including supported layouts that combine all three.
- Bounded resource use. Byte budgets cover pending reads, ready pages and active GPU transfers. Cancellation retains submitted I/O until completion.
- Observable behavior. Inspect Prometheus metrics and optional request timelines, and reproduce the published latency and throughput measurements. Opt-in cost observations compare matching copy/SSD-route evidence in shadow, with resource identity and uncertainty checks before suggesting a change.
- Experimental shared cache. Complete local global indexes locate peer replicas; etcd replicates locations and membership, and Mooncake TENT moves bytes. Source allocations remain budgeted through timeout; bounded completion records reconcile lost authorization replies and retry completion acknowledgements using reusable windows and generation-fenced tickets. Before the metadata cutover, both engines passed the recorded H20/A100 TCP natural-text recovery and restart gates, with cross-GPU numerical and RDMA limits documented separately. Peer SSD reads use exact-generation, bounded source-side io_uring staging before the same Mooncake transfer path; physical two-host qualification remains open.
- Experimental P/D handoff. vLLM uses OrbitKV's connector protocol;
SGLang
0.5.20keeps its native bootstrap/room protocol and can opt into the same Rust TENT payload owner withORBITKV_SGLANG_TENT=1. SGLang external H20 qualification remains open.
See supported deployments and model qualification before selecting a checkpoint and topology. Compiled recovery uses engine-declared state requirements; arbitrary model-graph analysis and future-token prediction are outside the current implementation.
Follow the installation guide to build and install a wheel for your Python and CUDA runtime. Use separate environments for vLLM and SGLang. The wheel includes the Cache Manager and Mooncake libraries; the Manager also requires compatible PyTorch. The first Python release is being prepared; see release preparation for package names and artifact checks.
Start a Cache Manager:
orbitkv-cache-manager --addr 127.0.0.1:50055 --http-addr 127.0.0.1:9091 --pool-size 8gbIn another terminal on the same host, start vLLM:
vllm serve /path/to/immutable-model \
--enable-prefix-caching \
--kv-transfer-config '{"kv_connector":"OrbitKVConnector","kv_role":"kv_both","kv_connector_module_path":"orbitkv.vllm"}'Or start SGLang:
ORBITKV_SGLANG_ENDPOINT=unix:///tmp/orbitkv-50055.sock \
sglang serve --model-path /path/to/immutable-model --page-size 64 \
--enable-unified-cache-external-linker --radix-cache-backend orbitkvTo enable SSD caching, add
--ssd-cache-path /data/orbitkv/cache.bin --ssd-cache-capacity 100gb to the
Manager command. The SSD cache is recreated when the Manager restarts.
No backend flag or engine-side storage setting is needed. The default
automatic SSD backend tries native cuFile on supported
mounts and uses io_uring when unavailable. cuFile is an optional library loaded
inside the Manager, not a separate service. It writes complete GPU state
groups and restores SSD demand hits through bounded GPU staging. Hardware
selection and native GDS performance qualification are separate. When cuFile
is selected, the Manager reserves the configured disk capacity before serving
and coalesces adjacent reads across cached blocks within each file. Rust submits
asynchronous I/O through two 4 MiB slots, keeping demand reads progressing
alongside bounded GPU writeback and event-tracked host copies.
SSD extent ownership is independent of its read route. For controlled comparisons,
--ssd-read-path uring|cufile selects demand restoration through pinned DRAM or
GPU staging over the same stored representation. --ssd-backend still controls
cuFile initialization and existing write behavior; leave both overrides unset for
the existing default policy. Prepare/warmup continues to target DRAM. See the
read-route contract.
GPU storage encoding supports nvCOMP ANS lossless
compression, FP8 and 3/4-bit TurboQuant with bounded batches and reusable GPU
workspace. Encoded pages stay compact in DRAM, SSD and peer transfers; cuFile can
write complete encoded groups and restore encoded SSD hits through GPU validation
and decode. GPU restore reconstructs the engine layout; FP8 CPU fallback selects
AVX-512F, AVX2 or scalar code at runtime. Set
--storage-codec ans|fp8|turboquant-4|turboquant-3 on the Manager. The default
is exact storage; lossy modes require model-quality qualification.
Keep the Manager alive, restart the engine, and repeat a multi-block prompt.
An increase in orbitkv_load_bytes_total confirms an external restore.
See the complete quickstart for identity, metrics and
container setup. Standalone caching requires neither etcd nor a gRPC listener.
The engine adapter identifies missing state and supplies GPU destinations. OrbitKV selects compatible cached ranges, reads them from the configured tiers, and retains page ownership until the GPU copy finishes. Newly computed KV is published for later reuse. The Manager owns cache placement and source grants; raw DRAM copies execute in the engine, while SSD/codec work remains with Manager workers. The same adapter API serves these routes and experimental remote fetches. The local executor partitions fragmented raw plans into at most 1 MiB parts under one whole-operation fence, with 32 MiB operation and 64 MiB session metadata limits. Raw per-layer CUDA dependencies allow consumption before later copies finish, with external-event graph replay and a final ownership fence. See execution scope and remaining gates.
The completion plan maps pinned LMCache, FlexKV and Mooncake mechanisms to deployment and validation work. The next milestone uses measured path costs and resource budgets across local tiers and Mooncake TENT transfers. Independent replicas, P/D handoff and TP/PP have separate completion and recovery contracts. Bounded Rust cost observations, raw-copy shadow predictions, independent local SSD read routes and a fixed-priority peer SSD route are implemented. SSD route shadow compares complete restoration to GPU readiness without changing execution. This is not yet a completed planner across all tiers. An opt-in selector can choose between equal-coverage owners of the same peer medium using fresh complete HostReady observations. Equal-coverage local SSD and single-owner peer routes also share a cross-medium shadow. A third, explicitly experimental opt-in can execute that choice, but it is not qualified until the external H20 TCP/RDMA matrix passes. All observations and execution selection remain off by default: the earlier overhead qualification has one open SGLang ANS SSD latency gate. The route changes require their own validation.
Read the architecture, hybrid recovery contract, and distributed design. Cross-engine byte conversion, and KV-aware routing remain planned work. The local global index and etcd metadata cutover are implemented; scale and multi-host failure qualification remain open. See the metadata design.
Performance depends on prefix reuse, cache capacity, storage and engine scheduling. The maintained guides link historical evidence and explain configurations, limits and reproduction:
| Report | Coverage |
|---|---|
| Single-node comparisons | Native HBM, engine CPU caches, OrbitKV, LMCache and FlexKV compatibility |
| SSD recovery | Restore readiness and sustained read/write pressure beyond DRAM capacity |
| Ordinary recovery | Qwen3-8B host reads, GPU transfers, notification delays and resource drain |
| Request preparation | Repeated preparation controls, DRAM recovery and read stopping policies |
| Shared-cache qualification | Independent replicas, remote GPU restoration, catalog replay and restart gates |
Request preparation remains off by default: the current Qwen3-8B controls
improve throughput in both engines, but SGLang P95 latency regresses. These
single-H20 measurements do not establish a universal advantage over other caches.
Benchmark programs live in benches/; generated results follow the external evidence policy.
- Get started: Installation · Adapter configuration · Manager options
- Operate: Metrics · Fault qualification · Deployment patterns
- Develop: Contributor guide · Python package · Test gates · Releases
- Development: Completion plan · Engine integration
Technical pages in docs/ are also published on the website. Contributions
should include the relevant checks and documentation changes.
OrbitKV is licensed under Apache-2.0.