Skip to content
 
 

Repository files navigation

llm-d-autoscaling

License

Autoscaling for llm-d inference deployments. This repository holds KEDA manifest blueprints and the evaluation test bed used to validate and tune them.

1. The Workload-Variant-Autoscaler (WVA) is deprecated

WVA — the custom autoscaling controller this repository used to host — is deprecated on main. It has not disappeared:

  • The last supported code, manifests, and docs are on the release-0.9 branch, released as v0.9.0. Use that branch for anything WVA-related.
  • On main, everything WVA has moved untouched into legacy/. That directory is staging for removal — it is frozen, its CI is not wired up, and it will be deleted in a future release. Do not build on it.

2. KEDA is the autoscaling engine

Autoscaling for llm-d is driven by KEDA reading inference metrics (queue depth, KV-cache utilization, and other vLLM/EPP signals) straight from Prometheus and scaling model-server Deployments through the HPA it manages. No custom controller sits in that path.

This repository's role is therefore twofold:

KEDA manifest blueprints

Recommended, reviewed KEDA scaling strategies — ScaledObject min/max, scaling behavior, and metric triggers per serving role (prefill/decode) — live with the deployment topologies they belong to, under benchmark/config/scenarios/:

  • scenarios/guides/ — recommended blueprints, one per llm-d guide (e.g. pd-disaggregation.yaml). These are the configurations to copy from.
  • scenarios/staging/ — experiments and work in progress: trigger and threshold variants (baseline, queue-aggressive, kv-early, token-aware, …) staged for evaluation before being promoted.

Because scenarios are backend-agnostic, the same blueprint runs against llm-d-inference-sim, a latency-simulating vLLM, or real GPU vLLM by swapping a cluster-config overlay.

Evaluation

benchmark/ is an autoscaling test bed built on llm-d-benchmark. It stands up a scenario, drives load through the harness, and captures autoscaling behavior (replicas, HPA/KEDA trigger values, latency, throughput) so blueprints are compared on evidence rather than intuition.

# Optional: a local Kind cluster with emulated GPUs
make create-kind-cluster

# Stand up + run a scenario (see benchmark/README.md for the full lifecycle)
llmdbenchmark standup \
  --spec benchmark/config/specification/guides/pd-disaggregation.yaml.j2 \
  --cluster-config benchmark/config/cluster-configs/k8s/inference-sim.yaml \
  --workspace benchmark/results -p <namespace>

Results and reports: benchmark/docs/benchmark-report.md and benchmark/docs/interactive-dashboard.md.

Documentation

Contributing

We welcome contributions! See CONTRIBUTING.md.

Join the llm-d autoscaling community meetings to get involved.

License

Apache 2.0 — see LICENSE.

Related projects

About

Variant optimization autoscaler for distributed inference workloads

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages