capacity planning for llm serving
One request contains a 20,000-token document and needs a 20-token answer. Another has a 100-token prompt and asks for 2,000 tokens back. Count them both as “one request” and your capacity dashboard has already hidden most of the work.
The first request puts a lot of work into prefill, which processes the input and builds attention state. The second spends more time decoding, generating tokens one at a time while reading the model and KV cache. They can occupy the same endpoint and still behave very differently under load. That's why requests per second, on its own, is a frustrating number to use for sizing an LLM service.
GPU utilization won't fill in all the missing detail either. Users may be waiting for a place in the queue or for cache space, even when a utilization graph doesn't suggest an obvious problem. Before deciding the fleet needs another replica, I want to see what those requests are waiting for.
put token lengths beside latency
Start with distributions of input and output lengths, arrivals and burst sizes. Split them by route or tenant wherever the workloads differ. Batch summarization and interactive chat shouldn't disappear into the same average if only one of them has a person waiting at a screen.
Time to first token, inter-token latency and completion time need separate targets. If the first token arrives late and the rest come steadily, the queue and prefill are reasonable places to investigate. If output arrives slowly despite little queueing, look at decode throughput, batch size and memory pressure. These aren't interchangeable symptoms, so one end-to-end latency chart leaves quite a bit to explain.
Little's law is a useful check on the measurements. At 12 arrivals per second and an average response time of five seconds, you should expect about 60 requests in flight in a stable system. If the numbers are wildly inconsistent, investigate before using them to size anything. Even when they agree, averages don't tell you much about a burst of long prompts. A small overload sustained for long enough will still fill a finite queue.
the cache has to fit
Each active sequence consumes KV-cache memory, with usage depending on token count, layer count, KV heads, head dimension and cache dtype. A limit of “ten concurrent requests” might work for ten short chats and fail for a handful of long documents. The requests are competing for memory in very different amounts.
For an unsharded model, a rough estimate is 2 × layers × kv_heads × head_dim × bytes_per_element × active_tokens. Keys and values account for the factor of two. It's only the cache estimate: weights, runtime buffers, allocator overhead and block fragmentation need room too. Parallelism changes the calculation per device. Still, this is enough to catch a proposed concurrency target that could never fit on the hardware.
Budget for the prompt already cached and for the output the request is permitted to generate. As the cache fills, the scheduler may have to queue, preempt or reject work. An admission policy that knows only the request count can't make a very informed decision about that pressure.
Continuous batching helps use the device by allowing requests to join while others progress. Context limits, output limits and per-tenant concurrency limits still matter. Long-document processing may justify a separate pool if it keeps occupying the interactive queue. Prefix caching can reduce repeated work where prompts share a stable beginning, but its value depends on the hit rate. I'd measure that on the actual traffic before subtracting anything from the capacity budget.
scaling takes time
A CPU metric from the API container is a poor proxy for this workload. Collect queue time, running requests, token throughput, first-token latency, inter-token latency and cache occupancy from the serving process. vLLM exposes queue and token-latency metrics, although the exact names depend on the version. An adapter can make custom or external metrics available to Kubernetes HPA.
Choosing the target is the consequential bit. A brief queue spike shouldn't provision a whole fleet, but waiting for some arbitrary GPU utilization threshold can leave you blind to cache pressure. A sustained queue signal, read alongside latency and available capacity, is more useful. It should also respond to the action the controller takes: if another replica doesn't improve that metric, the controller is chasing the wrong problem.
Then measure how long a replica takes to become useful. Provisioning the node, pulling the image, loading weights and warming the model all count. If those steps take minutes, autoscaling won't arrive in time for a burst that builds over seconds. You need enough warm capacity to cover it, or a gateway that rejects excess work within an agreed waiting limit.
HPA calculates desired replicas from the ratio between a metric and its target, subject to scaling rules. With a global queue, a per-pod average is understandable only when distribution is fair enough for it to represent the work. Try a step increase in load followed by a drop. Stale metrics, slow initialization and requests that outlive the scale-down window can leave the controller removing capacity just when it's still needed.
Rollouts complicate the arithmetic too. Starting new pods before stopping old ones needs spare GPU slots if each replica takes a whole device. Otherwise the replacements remain pending. MIG partitions supported NVIDIA hardware into isolated GPU instances, while time slicing shares a device with different isolation properties. Whichever arrangement you choose changes memory fit and interference. Measurements from a dedicated GPU don't establish how the same model behaves on a partition.
the benchmark i'd keep
Use a mix of short and long inputs, short and long outputs, and clustered arrivals in the proportions production receives. Identical prompts make comparison convenient, but they can flatter the server. Keep p50, p95 and p99 results for each traffic class, together with goodput: requests completed inside the latency objective. A throughput increase that doubles p99 time to first token may be a bad change for chat even if aggregate tokens per second improves.
Run some of the test with a GPU node removed. Restart a pod under load and make the metrics adapter unavailable. Losing capacity should lead to bounded waiting or rejection, not an ever-growing backlog. Save token counts and timing in the traces so later comparisons can separate a model regression from changes in the runtime, scheduler or traffic mix.
Those traces belong with the capacity figure. When someone asks whether the fleet can handle 12 requests a second, the answer needs to include which requests, how long they may wait and what happens during a rollout. Otherwise the next person has a number they can repeat but no experiment they can rerun.
technical references: vllm production metrics ↗, kubernetes hpa ↗, and nvidia mig operation ↗.