Skip to content

Repository files navigation

Continuous Batching Improves Your P50 and Can Wreck Your P99: Measured Evidence from a GPU Droplet

A reproducible study of continuous vs static-style (gated) admission under mixed-length traffic on DigitalOcean GPU Droplets with vLLM — measuring first-token wait, mid-stream freezes, and what happens when chunked prefill is turned off.

  • Run date: August 4, 2026, 07:00–07:18 UTC (one suite window)
  • Hardware: DigitalOcean GPU Droplet gpu-h200x1-141gb (NVIDIA H200, driver 575.57.08), NYC2, 1-Click Inference Ready image
  • Engine: vllm/vllm-openai:v0.24.0 (sha256:251eba5cc7c12fed0b75da22a9240e582b1c9e39f6fbc064f86781b963bd814f)
  • Model: RedHatAI/Llama-3.1-8B-Instruct (BF16 Llama 3.1 8B Instruct redistributable)
  • Volume: 7 cells × 500 measured requests (+ 25 warmup each), zero failed requests in the published suite
  • Everything in this repository is what actually ran: the harness, the raw per-request JSON, the suite log, the plotting code, and the charts it produced.

This repository is the evidence base for the companion article "Continuous Batching Improves Your P50 and Can Wreck Your P99: The Measured Tradeoff" (DigitalOcean Community).

Related study (different question, same measurement discipline): serverless-inference-tail-latency-study.


Abstract

Continuous batching is usually sold on throughput and median latency. Anyscale’s widely cited result is up to 23× throughput while reducing p50 latency. This study asks a different question: under mixed-length traffic, where does the pain move?

We held one GPU, one engine build, one model, and one request mix fixed, then changed only how requests are admitted:

  1. Continuous admission — open-loop Poisson arrivals at 1, 5, 10, and 20 requests/second (modern vLLM defaults, chunked prefill on).
  2. Gated admission (B=16) — a disclosed stand-in for static batching: send 16, wait for all 16 to finish, then send the next 16.
  3. Continuous again with chunked prefill off — same rates 5 and 10, one honest config toggle.

The headline result is a tradeoff measured on live infrastructure, not a folklore cliff:

  • Gated / static-style admission loses at the door. Typical time to first token ~196 ms vs ~19–48 ms for continuous.
  • Continuous’s risk shows up mid-answer. At 10 req/s, the median worst inter-token freeze jumped to 129.6 ms even while first-token wait still looked fine.
  • Chunked prefill still helps, modestly. Turning it off at 10 req/s moved worst-gap p99 from 189.9 → 267.8 ms (1.41×) on modern vLLM V1 defaults.
  • Preemption never fired on 8B + 141 GB H200 for this trace (vllm:num_preemptions_total delta = 0 every cell).

The conclusion is practical: keep continuous batching and chunked prefill on for general API serving; measure mid-stream freezes (not only TTFT); reserve gated admission for niches that need stream smoothness more than admission speed.


Key findings

1. Gated admission makes users wait before the first word

Measured cell TTFT p50 TTFT p99 Worst-gap p99
Continuous @ 10 req/s (defaults) 24.1 ms 252.2 ms 189.9 ms
Gated B=16 195.8 ms 637.2 ms 203.6 ms

Same engine, same model, same GPU. Gated only changes the door policy. Users wait longer to see the first word; the stream itself is not the disaster mode.

p99 comparison across arms

Blue = wait before the first word. Orange = worst freeze while the answer is already typing. Green = time until the whole answer finishes. Continuous wins the blue bar. Gated pays at the door. “Chunked off” worsens the orange bar. Gated’s green bar can look smaller because gated only lets 16 requests in at a time — it never takes the same open rush of traffic.

2. Continuous starts fast, then stutters as traffic rises

Continuous defaults TTFT p50 Worst-gap p50 Worst-gap p99 Total p99
1 req/s 17.0 ms 5.9 ms 129.7 ms 3,434 ms
5 req/s 19.3 ms 18.3 ms 134.9 ms 4,409 ms
10 req/s 24.1 ms 129.6 ms 189.9 ms 5,723 ms
20 req/s 47.7 ms 159.9 ms 212.7 ms 12,510 ms

If you only watched median TTFT, rate 10 would still look healthy. The median freeze between tokens is already over 100 ms. That is the interactive failure mode continuous batching can hide from TTFT-only dashboards.

3. Turning chunked prefill off hurts the stream, modestly

Continuous @ 10 req/s TTFT p99 Worst-gap p99
Defaults (chunked on) 252.2 ms 189.9 ms
Chunked prefill off 249.4 ms 267.8 ms (1.41×)

TTFT and worst-gap p99 vs arrival rate

Left: wait for the first word as traffic rises — chunked on vs off barely differ. Right: mid-answer freezes — at 10 req/s, chunked-off jumps clearly worse. Keep the default on if users watch tokens appear on screen.

On this H200 + vLLM V1 pin, the measured delta was 1.41× on gap p99. Direction of the mechanism confirmed; the folklore magnitude of the chunked-off cliff is rejected for this setup.

4. Preemption was not the story of this run

Every published cell recorded preemptions_delta: 0.0 against vllm:num_preemptions_total. An 8B model on 141 GB of H200 simply had enough working memory for this mix. Preemption remains a real risk on tighter GPUs or longer contexts — absence here is headroom, not proof the mechanism is dead.

5. Full scorecard

Results summary table

All harness summaries in one image. Lower milliseconds are better. Last column stayed 0. If prose and this image disagree, trust results/*.json.

Rate Arm Config TTFT p50 TTFT p99 Gap p50 Gap p99 Total p99 Preempt
1 Continuous Defaults 17.0 154.0 5.9 129.7 3434.1 0
5 Continuous Defaults 19.3 161.2 18.3 134.9 4409.3 0
10 Continuous Defaults 24.1 252.2 129.6 189.9 5723.2 0
20 Continuous Defaults 47.7 394.8 159.9 212.7 12510.4 0
5 Continuous Chunked off 19.6 161.1 18.1 134.3 4482.3 0
10 Continuous Chunked off 24.2 249.4 129.4 267.8 5844.2 0
n/a Gated B=16 Defaults 195.8 637.2 52.8 203.6 4073.6 0

Methodology

What stayed fixed

  • One H200 Droplet, one vLLM image digest, one model ID
  • One mixed-length trace (seed 7): 70% short (≈200 prompt / 100 output), 20% medium (1000 / 300), 10% long (6000 / 600)
  • 500 measured requests + 25 discarded warmup per cell
  • Prefix caching disabled (--no-enable-prefix-caching) so identical filler prompts do not fake cheap prefills
  • Client: real OS threads (not a single asyncio loop), streaming completions, client-side TTFT and worst inter-token gap

What changed

  1. Continuous open-loop Poisson arrivals at rates 1, 5, 10, 20
  2. Gated admission with batch size 16 on the same default server
  3. Server restart with --no-enable-chunked-prefill; continuous rates 5 and 10 repeated

Measured per request: TTFT (first streamed content chunk), worst inter-token gap, total completion time, chunk count. Measured per run: delta of vllm:num_preemptions_total.

Known limitations, stated plainly

  • Gated is an admission-layer stand-in for static batching against a continuous engine. It does not reproduce static padding waste; treat it as mildly flattering to true static batching.
  • Filler-word prompts approximate token counts; both arms share the identical trace.
  • One model size (8B), one GPU class (H200), one day. Larger models or smaller GPUs will move preemption and gap magnitudes.
  • Open-loop total time at high rate includes queueing when arrivals outrun the GPU — do not read total p99 alone as “scheduler quality.”

Metadata: results/run_metadata.json. Suite stdout: results/suite.log.


Plain-language glossary

Term Plain meaning Simple example
Prefill / decode Read the prompt, then write the answer word by word Skim the email, then type the reply
Continuous batching Finished requests leave; new ones join after each small GPU step A revolving door
Static / gated (B=16) Only admit the next group of 16 after the current 16 all finish Tables of 16; no new seating until the party leaves
Chunked prefill Split a long prompt into pieces so one long read does not freeze everyone else’s stream Read a long book in short chapters between other answers
Chunked off That feature turned off on purpose Finish one giant order before touching anything else
TTFT Wait until the first word appears Spinner before the chat bubble starts
Worst gap Longest freeze while the answer is already typing Mid-sentence stutter
p50 / p99 Typical case / unlucky ~1-in-100 case Median vs tail
req/s How fast new requests arrive 1 = calm, 20 = rush

More detail: docs/GLOSSARY.md.


Repository layout

benchmarks/
  batching_bench.py      Stdlib-only harness that ran on the Droplet (continuous + gated).
  plot_results.py        Regenerates every chart from results/*.json.
results/
  cont_r*.json           Continuous defaults at each rate (raw per-request records + summary).
  nochunk_r*.json        Continuous with chunked prefill disabled.
  gated_b16.json         Gated admission, batch size 16.
  run_metadata.json      GPU, digest, flags, timestamps, summary table.
  suite.log              Verbatim harness stdout for the published suite.
figures/
  p99_vs_rate.png        First-word wait and mid-answer freezes vs traffic.
  p99_arm_comparison.png Continuous vs gated vs chunked-off at a glance.
  results_table.png      Full numeric scorecard image.
docs/
  RUNBOOK.md             Step-by-step reproduce instructions.
  THESIS.md              Claim structure.
  READING_RESULTS.md     Checklist for reading someone else's batching bench.
  GLOSSARY.md            Beginner terms.
continuous-vs-static-batching.md   Article draft with measured tables filled in.

Reproducing

Against the archived data (no GPU needed)

python3 -m venv .venv
source .venv/bin/activate
pip install matplotlib numpy
python3 benchmarks/plot_results.py --results-dir results --out-dir figures

This regenerates the three charts from the published JSON.

Against live infrastructure (your own numbers)

You need a GPU Droplet with Docker + NVIDIA drivers (the 1-Click Inference Ready image ships both).

docker pull vllm/vllm-openai:v0.24.0
docker images --digests | grep v0.24.0   # record the digest

docker run -d --name vllm-default --gpus all --ipc=host -p 8000:8000 \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  vllm/vllm-openai:v0.24.0 \
  --model RedHatAI/Llama-3.1-8B-Instruct \
  --no-enable-prefix-caching

# wait until ready
curl -s http://localhost:8000/v1/models | head -c 400

python3 benchmarks/batching_bench.py --arm continuous --rate 1  --n 500 \
  --model RedHatAI/Llama-3.1-8B-Instruct --out results/cont_r1.json
python3 benchmarks/batching_bench.py --arm continuous --rate 5  --n 500 \
  --model RedHatAI/Llama-3.1-8B-Instruct --out results/cont_r5.json
python3 benchmarks/batching_bench.py --arm continuous --rate 10 --n 500 \
  --model RedHatAI/Llama-3.1-8B-Instruct --out results/cont_r10.json
python3 benchmarks/batching_bench.py --arm continuous --rate 20 --n 500 \
  --model RedHatAI/Llama-3.1-8B-Instruct --out results/cont_r20.json

python3 benchmarks/batching_bench.py --arm gated --batch-size 16 --n 500 \
  --model RedHatAI/Llama-3.1-8B-Instruct --out results/gated_b16.json

# chunked-off arm (flag verified on v0.24.0)
docker rm -f vllm-default
docker run -d --name vllm-nochunk --gpus all --ipc=host -p 8000:8000 \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  vllm/vllm-openai:v0.24.0 \
  --model RedHatAI/Llama-3.1-8B-Instruct \
  --no-enable-prefix-caching \
  --no-enable-chunked-prefill

python3 benchmarks/batching_bench.py --arm continuous --rate 5  --n 500 \
  --model RedHatAI/Llama-3.1-8B-Instruct --out results/nochunk_r5.json
python3 benchmarks/batching_bench.py --arm continuous --rate 10 --n 500 \
  --model RedHatAI/Llama-3.1-8B-Instruct --out results/nochunk_r10.json

Full checklist: docs/RUNBOOK.md.

If your results differ from anything reported here — different GPU, model size, or traffic mix — that difference is the finding. Open an issue or PR with your run_metadata.json and summaries.


Citation

If you use this data or method, please cite:

Anish Singh Walia (2026). Continuous Batching Improves Your P50 and Can Wreck Your P99:
Measured Evidence from a GPU Droplet.
https://github.com/anishsingh20/continuous-vs-static-batching

License

MIT. See LICENSE. Model weights remain under their upstream licenses (Meta Llama / Red Hat redistributions).

About

Reproducible continuous vs gated (static-style) batching study on a DigitalOcean H200 GPU Droplet with vLLM: TTFT, mid-stream freezes, chunked-prefill toggle, raw JSON, harness, and charts (August 4, 2026 run)

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages