Autoscaling for llm-d inference deployments. This repository holds KEDA manifest blueprints and the evaluation test bed used to validate and tune them.
WVA — the custom autoscaling controller this repository used to host — is
deprecated on main. It has not disappeared:
- The last supported code, manifests, and docs are on the
release-0.9branch, released asv0.9.0. Use that branch for anything WVA-related. - On
main, everything WVA has moved untouched intolegacy/. That directory is staging for removal — it is frozen, its CI is not wired up, and it will be deleted in a future release. Do not build on it.
Autoscaling for llm-d is driven by KEDA reading inference metrics (queue depth, KV-cache utilization, and other vLLM/EPP signals) straight from Prometheus and scaling model-server Deployments through the HPA it manages. No custom controller sits in that path.
This repository's role is therefore twofold:
Recommended, reviewed KEDA scaling strategies — ScaledObject min/max, scaling
behavior, and metric triggers per serving role (prefill/decode) — live with the
deployment topologies they belong to, under
benchmark/config/scenarios/:
scenarios/guides/— recommended blueprints, one per llm-d guide (e.g.pd-disaggregation.yaml). These are the configurations to copy from.scenarios/staging/— experiments and work in progress: trigger and threshold variants (baseline,queue-aggressive,kv-early,token-aware, …) staged for evaluation before being promoted.
Because scenarios are backend-agnostic, the same blueprint runs against
llm-d-inference-sim, a latency-simulating vLLM, or real GPU vLLM by swapping a
cluster-config overlay.
benchmark/ is an autoscaling test bed built on
llm-d-benchmark. It stands up a
scenario, drives load through the harness, and captures autoscaling behavior
(replicas, HPA/KEDA trigger values, latency, throughput) so blueprints are
compared on evidence rather than intuition.
# Optional: a local Kind cluster with emulated GPUs
make create-kind-cluster
# Stand up + run a scenario (see benchmark/README.md for the full lifecycle)
llmdbenchmark standup \
--spec benchmark/config/specification/guides/pd-disaggregation.yaml.j2 \
--cluster-config benchmark/config/cluster-configs/k8s/inference-sim.yaml \
--workspace benchmark/results -p <namespace>Results and reports: benchmark/docs/benchmark-report.md
and benchmark/docs/interactive-dashboard.md.
- Repository docs index
- Benchmark test bed
- llm-d autoscaling architecture
- llm-d KEDA autoscaling guide
We welcome contributions! See CONTRIBUTING.md.
Join the llm-d autoscaling community meetings to get involved.
Apache 2.0 — see LICENSE.