A reproducible study of continuous vs static-style (gated) admission under mixed-length traffic on DigitalOcean GPU Droplets with vLLM — measuring first-token wait, mid-stream freezes, and what happens when chunked prefill is turned off.
- Run date: August 4, 2026, 07:00–07:18 UTC (one suite window)
- Hardware: DigitalOcean GPU Droplet
gpu-h200x1-141gb(NVIDIA H200, driver 575.57.08), NYC2, 1-Click Inference Ready image - Engine:
vllm/vllm-openai:v0.24.0(sha256:251eba5cc7c12fed0b75da22a9240e582b1c9e39f6fbc064f86781b963bd814f) - Model:
RedHatAI/Llama-3.1-8B-Instruct(BF16 Llama 3.1 8B Instruct redistributable) - Volume: 7 cells × 500 measured requests (+ 25 warmup each), zero failed requests in the published suite
- Everything in this repository is what actually ran: the harness, the raw per-request JSON, the suite log, the plotting code, and the charts it produced.
This repository is the evidence base for the companion article "Continuous Batching Improves Your P50 and Can Wreck Your P99: The Measured Tradeoff" (DigitalOcean Community).
Related study (different question, same measurement discipline): serverless-inference-tail-latency-study.
Continuous batching is usually sold on throughput and median latency. Anyscale’s widely cited result is up to 23× throughput while reducing p50 latency. This study asks a different question: under mixed-length traffic, where does the pain move?
We held one GPU, one engine build, one model, and one request mix fixed, then changed only how requests are admitted:
- Continuous admission — open-loop Poisson arrivals at 1, 5, 10, and 20 requests/second (modern vLLM defaults, chunked prefill on).
- Gated admission (B=16) — a disclosed stand-in for static batching: send 16, wait for all 16 to finish, then send the next 16.
- Continuous again with chunked prefill off — same rates 5 and 10, one honest config toggle.
The headline result is a tradeoff measured on live infrastructure, not a folklore cliff:
- Gated / static-style admission loses at the door. Typical time to first token ~196 ms vs ~19–48 ms for continuous.
- Continuous’s risk shows up mid-answer. At 10 req/s, the median worst inter-token freeze jumped to 129.6 ms even while first-token wait still looked fine.
- Chunked prefill still helps, modestly. Turning it off at 10 req/s moved worst-gap p99 from 189.9 → 267.8 ms (1.41×) on modern vLLM V1 defaults.
- Preemption never fired on 8B + 141 GB H200 for this trace (
vllm:num_preemptions_totaldelta = 0 every cell).
The conclusion is practical: keep continuous batching and chunked prefill on for general API serving; measure mid-stream freezes (not only TTFT); reserve gated admission for niches that need stream smoothness more than admission speed.
| Measured cell | TTFT p50 | TTFT p99 | Worst-gap p99 |
|---|---|---|---|
| Continuous @ 10 req/s (defaults) | 24.1 ms | 252.2 ms | 189.9 ms |
| Gated B=16 | 195.8 ms | 637.2 ms | 203.6 ms |
Same engine, same model, same GPU. Gated only changes the door policy. Users wait longer to see the first word; the stream itself is not the disaster mode.
Blue = wait before the first word. Orange = worst freeze while the answer is already typing. Green = time until the whole answer finishes. Continuous wins the blue bar. Gated pays at the door. “Chunked off” worsens the orange bar. Gated’s green bar can look smaller because gated only lets 16 requests in at a time — it never takes the same open rush of traffic.
| Continuous defaults | TTFT p50 | Worst-gap p50 | Worst-gap p99 | Total p99 |
|---|---|---|---|---|
| 1 req/s | 17.0 ms | 5.9 ms | 129.7 ms | 3,434 ms |
| 5 req/s | 19.3 ms | 18.3 ms | 134.9 ms | 4,409 ms |
| 10 req/s | 24.1 ms | 129.6 ms | 189.9 ms | 5,723 ms |
| 20 req/s | 47.7 ms | 159.9 ms | 212.7 ms | 12,510 ms |
If you only watched median TTFT, rate 10 would still look healthy. The median freeze between tokens is already over 100 ms. That is the interactive failure mode continuous batching can hide from TTFT-only dashboards.
| Continuous @ 10 req/s | TTFT p99 | Worst-gap p99 |
|---|---|---|
| Defaults (chunked on) | 252.2 ms | 189.9 ms |
| Chunked prefill off | 249.4 ms | 267.8 ms (1.41×) |
Left: wait for the first word as traffic rises — chunked on vs off barely differ. Right: mid-answer freezes — at 10 req/s, chunked-off jumps clearly worse. Keep the default on if users watch tokens appear on screen.
On this H200 + vLLM V1 pin, the measured delta was 1.41× on gap p99. Direction of the mechanism confirmed; the folklore magnitude of the chunked-off cliff is rejected for this setup.
Every published cell recorded preemptions_delta: 0.0 against vllm:num_preemptions_total. An 8B model on 141 GB of H200 simply had enough working memory for this mix. Preemption remains a real risk on tighter GPUs or longer contexts — absence here is headroom, not proof the mechanism is dead.
All harness summaries in one image. Lower milliseconds are better. Last column stayed 0. If prose and this image disagree, trust results/*.json.
| Rate | Arm | Config | TTFT p50 | TTFT p99 | Gap p50 | Gap p99 | Total p99 | Preempt |
|---|---|---|---|---|---|---|---|---|
| 1 | Continuous | Defaults | 17.0 | 154.0 | 5.9 | 129.7 | 3434.1 | 0 |
| 5 | Continuous | Defaults | 19.3 | 161.2 | 18.3 | 134.9 | 4409.3 | 0 |
| 10 | Continuous | Defaults | 24.1 | 252.2 | 129.6 | 189.9 | 5723.2 | 0 |
| 20 | Continuous | Defaults | 47.7 | 394.8 | 159.9 | 212.7 | 12510.4 | 0 |
| 5 | Continuous | Chunked off | 19.6 | 161.1 | 18.1 | 134.3 | 4482.3 | 0 |
| 10 | Continuous | Chunked off | 24.2 | 249.4 | 129.4 | 267.8 | 5844.2 | 0 |
| n/a | Gated B=16 | Defaults | 195.8 | 637.2 | 52.8 | 203.6 | 4073.6 | 0 |
What stayed fixed
- One H200 Droplet, one vLLM image digest, one model ID
- One mixed-length trace (seed
7): 70% short (≈200 prompt / 100 output), 20% medium (1000 / 300), 10% long (6000 / 600) - 500 measured requests + 25 discarded warmup per cell
- Prefix caching disabled (
--no-enable-prefix-caching) so identical filler prompts do not fake cheap prefills - Client: real OS threads (not a single asyncio loop), streaming completions, client-side TTFT and worst inter-token gap
What changed
- Continuous open-loop Poisson arrivals at rates 1, 5, 10, 20
- Gated admission with batch size 16 on the same default server
- Server restart with
--no-enable-chunked-prefill; continuous rates 5 and 10 repeated
Measured per request: TTFT (first streamed content chunk), worst inter-token gap, total completion time, chunk count. Measured per run: delta of vllm:num_preemptions_total.
Known limitations, stated plainly
- Gated is an admission-layer stand-in for static batching against a continuous engine. It does not reproduce static padding waste; treat it as mildly flattering to true static batching.
- Filler-word prompts approximate token counts; both arms share the identical trace.
- One model size (8B), one GPU class (H200), one day. Larger models or smaller GPUs will move preemption and gap magnitudes.
- Open-loop total time at high rate includes queueing when arrivals outrun the GPU — do not read total p99 alone as “scheduler quality.”
Metadata: results/run_metadata.json. Suite stdout: results/suite.log.
| Term | Plain meaning | Simple example |
|---|---|---|
| Prefill / decode | Read the prompt, then write the answer word by word | Skim the email, then type the reply |
| Continuous batching | Finished requests leave; new ones join after each small GPU step | A revolving door |
| Static / gated (B=16) | Only admit the next group of 16 after the current 16 all finish | Tables of 16; no new seating until the party leaves |
| Chunked prefill | Split a long prompt into pieces so one long read does not freeze everyone else’s stream | Read a long book in short chapters between other answers |
| Chunked off | That feature turned off on purpose | Finish one giant order before touching anything else |
| TTFT | Wait until the first word appears | Spinner before the chat bubble starts |
| Worst gap | Longest freeze while the answer is already typing | Mid-sentence stutter |
| p50 / p99 | Typical case / unlucky ~1-in-100 case | Median vs tail |
| req/s | How fast new requests arrive | 1 = calm, 20 = rush |
More detail: docs/GLOSSARY.md.
benchmarks/
batching_bench.py Stdlib-only harness that ran on the Droplet (continuous + gated).
plot_results.py Regenerates every chart from results/*.json.
results/
cont_r*.json Continuous defaults at each rate (raw per-request records + summary).
nochunk_r*.json Continuous with chunked prefill disabled.
gated_b16.json Gated admission, batch size 16.
run_metadata.json GPU, digest, flags, timestamps, summary table.
suite.log Verbatim harness stdout for the published suite.
figures/
p99_vs_rate.png First-word wait and mid-answer freezes vs traffic.
p99_arm_comparison.png Continuous vs gated vs chunked-off at a glance.
results_table.png Full numeric scorecard image.
docs/
RUNBOOK.md Step-by-step reproduce instructions.
THESIS.md Claim structure.
READING_RESULTS.md Checklist for reading someone else's batching bench.
GLOSSARY.md Beginner terms.
continuous-vs-static-batching.md Article draft with measured tables filled in.
python3 -m venv .venv
source .venv/bin/activate
pip install matplotlib numpy
python3 benchmarks/plot_results.py --results-dir results --out-dir figuresThis regenerates the three charts from the published JSON.
You need a GPU Droplet with Docker + NVIDIA drivers (the 1-Click Inference Ready image ships both).
docker pull vllm/vllm-openai:v0.24.0
docker images --digests | grep v0.24.0 # record the digest
docker run -d --name vllm-default --gpus all --ipc=host -p 8000:8000 \
-v ~/.cache/huggingface:/root/.cache/huggingface \
vllm/vllm-openai:v0.24.0 \
--model RedHatAI/Llama-3.1-8B-Instruct \
--no-enable-prefix-caching
# wait until ready
curl -s http://localhost:8000/v1/models | head -c 400
python3 benchmarks/batching_bench.py --arm continuous --rate 1 --n 500 \
--model RedHatAI/Llama-3.1-8B-Instruct --out results/cont_r1.json
python3 benchmarks/batching_bench.py --arm continuous --rate 5 --n 500 \
--model RedHatAI/Llama-3.1-8B-Instruct --out results/cont_r5.json
python3 benchmarks/batching_bench.py --arm continuous --rate 10 --n 500 \
--model RedHatAI/Llama-3.1-8B-Instruct --out results/cont_r10.json
python3 benchmarks/batching_bench.py --arm continuous --rate 20 --n 500 \
--model RedHatAI/Llama-3.1-8B-Instruct --out results/cont_r20.json
python3 benchmarks/batching_bench.py --arm gated --batch-size 16 --n 500 \
--model RedHatAI/Llama-3.1-8B-Instruct --out results/gated_b16.json
# chunked-off arm (flag verified on v0.24.0)
docker rm -f vllm-default
docker run -d --name vllm-nochunk --gpus all --ipc=host -p 8000:8000 \
-v ~/.cache/huggingface:/root/.cache/huggingface \
vllm/vllm-openai:v0.24.0 \
--model RedHatAI/Llama-3.1-8B-Instruct \
--no-enable-prefix-caching \
--no-enable-chunked-prefill
python3 benchmarks/batching_bench.py --arm continuous --rate 5 --n 500 \
--model RedHatAI/Llama-3.1-8B-Instruct --out results/nochunk_r5.json
python3 benchmarks/batching_bench.py --arm continuous --rate 10 --n 500 \
--model RedHatAI/Llama-3.1-8B-Instruct --out results/nochunk_r10.jsonFull checklist: docs/RUNBOOK.md.
If your results differ from anything reported here — different GPU, model size, or traffic mix — that difference is the finding. Open an issue or PR with your run_metadata.json and summaries.
If you use this data or method, please cite:
Anish Singh Walia (2026). Continuous Batching Improves Your P50 and Can Wreck Your P99:
Measured Evidence from a GPU Droplet.
https://github.com/anishsingh20/continuous-vs-static-batching
MIT. See LICENSE. Model weights remain under their upstream licenses (Meta Llama / Red Hat redistributions).


