Skip to content

Repository files navigation

OrbitKV — Compile the lifetime of state

Website · Documentation · Quickstart · Architecture · Completion plan

CI Python package v0.1.0 — not yet published Wheel targets: Python 3.10–3.14 Apache 2.0 license

About

OrbitKV extends the KV cache of vLLM and SGLang beyond GPU memory. Keep reusable prefixes in DRAM and SSD, then restore them when a matching request arrives. This helps workloads with repeated documents, shared system prompts, and conversations whose prefixes no longer fit in the engine's GPU cache.

Run an independent Cache Manager per host and connect the engines on that host to its shared cache. Engines own GPU memory and scheduling; OrbitKV manages external replicas and transfers. See deployment patterns for shared-instance budgets and container qualification limits. The pinned engine baselines are vLLM 0.29.0 and SGLang 0.5.20. The recorded engine-local raw Restore cutover passes single-H20 Qwen3-8B DRAM serving correctness and restart reuse in both engines. The recorded vLLM end-to-end comparison shows gains over HBM-eviction recomputation, while native CPU offload remains faster. LMCache comparisons and their limits are recorded in the same report. The matched communication measurements track the initial regression and the subsequent idle-stream/plan-compaction optimization. Dense transfers benefit; small-payload overhead and broader serving/deployment qualification remain open. Multi-node cache sharing is experimental. Interfaces may change before 1.0.

Key features

  • DRAM and SSD caching. Reuse prefixes after GPU eviction or an engine restart while the Cache Manager remains alive.
  • Optional reuse policies. Rust can protect reused pages within a byte cap and admit SSD writes selectively. See the policy controls and their cold-reuse tradeoff before enabling them.
  • Native GPU transfers. Unencoded DRAM Restore executes inside the engine using independently imported shared payload arenas and retained tensors. Per-layer CUDA dependencies let consumers start before later copies finish; source and destination ownership lasts through the final drain. SSD/codec Restore and Publish retain Manager workers and CUDA IPC bindings; adapters supply the CUDA stream dependencies and Rust owns completion.
  • Model-aware recovery. Cache identity includes model artifacts, computation settings and storage layout. Compiled recovery rules select the required attention pages, sliding windows and recurrent/conv checkpoints, including supported layouts that combine all three.
  • Bounded resource use. Byte budgets cover pending reads, ready pages and active GPU transfers. Cancellation retains submitted I/O until completion.
  • Observable behavior. Inspect Prometheus metrics and optional request timelines, and reproduce the published latency and throughput measurements. Opt-in cost observations compare matching copy/SSD-route evidence in shadow, with resource identity and uncertainty checks before suggesting a change.
  • Experimental shared cache. Complete local global indexes locate peer replicas; etcd replicates locations and membership, and Mooncake TENT moves bytes. Source allocations remain budgeted through timeout; bounded completion records reconcile lost authorization replies and retry completion acknowledgements using reusable windows and generation-fenced tickets. Before the metadata cutover, both engines passed the recorded H20/A100 TCP natural-text recovery and restart gates, with cross-GPU numerical and RDMA limits documented separately. Peer SSD reads use exact-generation, bounded source-side io_uring staging before the same Mooncake transfer path; physical two-host qualification remains open.
  • Experimental P/D handoff. vLLM uses OrbitKV's connector protocol; SGLang 0.5.20 keeps its native bootstrap/room protocol and can opt into the same Rust TENT payload owner with ORBITKV_SGLANG_TENT=1. SGLang external H20 qualification remains open.

See supported deployments and model qualification before selecting a checkpoint and topology. Compiled recovery uses engine-declared state requirements; arbitrary model-graph analysis and future-token prediction are outside the current implementation.

Quickstart

Follow the installation guide to build and install a wheel for your Python and CUDA runtime. Use separate environments for vLLM and SGLang. The wheel includes the Cache Manager and Mooncake libraries; the Manager also requires compatible PyTorch. The first Python release is being prepared; see release preparation for package names and artifact checks.

Start a Cache Manager:

orbitkv-cache-manager --addr 127.0.0.1:50055 --http-addr 127.0.0.1:9091 --pool-size 8gb

In another terminal on the same host, start vLLM:

vllm serve /path/to/immutable-model \
  --enable-prefix-caching \
  --kv-transfer-config '{"kv_connector":"OrbitKVConnector","kv_role":"kv_both","kv_connector_module_path":"orbitkv.vllm"}'

Or start SGLang:

ORBITKV_SGLANG_ENDPOINT=unix:///tmp/orbitkv-50055.sock \
  sglang serve --model-path /path/to/immutable-model --page-size 64 \
  --enable-unified-cache-external-linker --radix-cache-backend orbitkv

To enable SSD caching, add --ssd-cache-path /data/orbitkv/cache.bin --ssd-cache-capacity 100gb to the Manager command. The SSD cache is recreated when the Manager restarts. No backend flag or engine-side storage setting is needed. The default automatic SSD backend tries native cuFile on supported mounts and uses io_uring when unavailable. cuFile is an optional library loaded inside the Manager, not a separate service. It writes complete GPU state groups and restores SSD demand hits through bounded GPU staging. Hardware selection and native GDS performance qualification are separate. When cuFile is selected, the Manager reserves the configured disk capacity before serving and coalesces adjacent reads across cached blocks within each file. Rust submits asynchronous I/O through two 4 MiB slots, keeping demand reads progressing alongside bounded GPU writeback and event-tracked host copies.

SSD extent ownership is independent of its read route. For controlled comparisons, --ssd-read-path uring|cufile selects demand restoration through pinned DRAM or GPU staging over the same stored representation. --ssd-backend still controls cuFile initialization and existing write behavior; leave both overrides unset for the existing default policy. Prepare/warmup continues to target DRAM. See the read-route contract.

GPU storage encoding supports nvCOMP ANS lossless compression, FP8 and 3/4-bit TurboQuant with bounded batches and reusable GPU workspace. Encoded pages stay compact in DRAM, SSD and peer transfers; cuFile can write complete encoded groups and restore encoded SSD hits through GPU validation and decode. GPU restore reconstructs the engine layout; FP8 CPU fallback selects AVX-512F, AVX2 or scalar code at runtime. Set --storage-codec ans|fp8|turboquant-4|turboquant-3 on the Manager. The default is exact storage; lossy modes require model-quality qualification.

Keep the Manager alive, restart the engine, and repeat a multi-block prompt. An increase in orbitkv_load_bytes_total confirms an external restore. See the complete quickstart for identity, metrics and container setup. Standalone caching requires neither etcd nor a gRPC listener.

Architecture

OrbitKV architecture: local restore ownership, peer cache READ and P/D WRITE

The engine adapter identifies missing state and supplies GPU destinations. OrbitKV selects compatible cached ranges, reads them from the configured tiers, and retains page ownership until the GPU copy finishes. Newly computed KV is published for later reuse. The Manager owns cache placement and source grants; raw DRAM copies execute in the engine, while SSD/codec work remains with Manager workers. The same adapter API serves these routes and experimental remote fetches. The local executor partitions fragmented raw plans into at most 1 MiB parts under one whole-operation fence, with 32 MiB operation and 64 MiB session metadata limits. Raw per-layer CUDA dependencies allow consumption before later copies finish, with external-event graph replay and a final ownership fence. See execution scope and remaining gates.

The completion plan maps pinned LMCache, FlexKV and Mooncake mechanisms to deployment and validation work. The next milestone uses measured path costs and resource budgets across local tiers and Mooncake TENT transfers. Independent replicas, P/D handoff and TP/PP have separate completion and recovery contracts. Bounded Rust cost observations, raw-copy shadow predictions, independent local SSD read routes and a fixed-priority peer SSD route are implemented. SSD route shadow compares complete restoration to GPU readiness without changing execution. This is not yet a completed planner across all tiers. An opt-in selector can choose between equal-coverage owners of the same peer medium using fresh complete HostReady observations. Equal-coverage local SSD and single-owner peer routes also share a cross-medium shadow. A third, explicitly experimental opt-in can execute that choice, but it is not qualified until the external H20 TCP/RDMA matrix passes. All observations and execution selection remain off by default: the earlier overhead qualification has one open SGLang ANS SSD latency gate. The route changes require their own validation.

Read the architecture, hybrid recovery contract, and distributed design. Cross-engine byte conversion, and KV-aware routing remain planned work. The local global index and etcd metadata cutover are implemented; scale and multi-host failure qualification remain open. See the metadata design.

Performance

Performance depends on prefix reuse, cache capacity, storage and engine scheduling. The maintained guides link historical evidence and explain configurations, limits and reproduction:

Report Coverage
Single-node comparisons Native HBM, engine CPU caches, OrbitKV, LMCache and FlexKV compatibility
SSD recovery Restore readiness and sustained read/write pressure beyond DRAM capacity
Ordinary recovery Qwen3-8B host reads, GPU transfers, notification delays and resource drain
Request preparation Repeated preparation controls, DRAM recovery and read stopping policies
Shared-cache qualification Independent replicas, remote GPU restoration, catalog replay and restart gates

Request preparation remains off by default: the current Qwen3-8B controls improve throughput in both engines, but SGLang P95 latency regresses. These single-H20 measurements do not establish a universal advantage over other caches. Benchmark programs live in benches/; generated results follow the external evidence policy.

Documentation and contributing

Technical pages in docs/ are also published on the website. Contributions should include the relevant checks and documentation changes.

License

OrbitKV is licensed under Apache-2.0.

Releases

Packages

Used by

Contributors

Languages