BlogCapacity Planning for AI: Stop Counting Requests, Start Counting TokensHit · Load & Latency

Capacity Planning for AI: Stop Counting Requests, Start Counting Tokens

PN
Priya Nair · April 2026 · 11 min read

TL;DR

You sized your AI infrastructure on requests per second, the way you size every other service, and got both brownouts and idle GPUs in the same week. The classic plan assumes each request does roughly equal work; LLM inference detonates that, because work is dominated by token count and token count varies by orders of magnitude. Inference throughput is measured in tokens per second, split into a compute-bound prefill phase and a sequential, memory-bandwidth-bound decode phase, and the real concurrency ceiling is usually KV-cache GPU memory, not raw compute. Output length is right-skewed with a heavy tail, so a cluster of long-output requests can saturate your serving slots at a request rate well below your planned peak. Size on tokens, model the KV cache, and autoscale on token throughput, not request count.

The decade-old spreadsheet that finally broke

The capacity plan was a spreadsheet that had worked for ten years. Peak requests per second, multiply by a safety factor, divide by per-instance throughput, round up, provision. It had sized web tiers, API gateways, and database fleets through three launches without drama.

Applied to the new LLM feature, it produced a confident number, and a bad week. The system browned out at half the projected peak request rate on Tuesday, then sat at 30% GPU utilization through Thursday at a request rate that should have saturated it. Same plan, same traffic level, opposite failures, days apart. The plan was not wrong about requests. It was measuring the wrong quantity entirely.

Requests are a constant-work assumption, and LLMs violate it

Classic capacity planning rests on one quiet assumption: each request does about the same amount of work. That is what lets you convert a request rate into a resource requirement. It holds for stateless web servers and most CRUD APIs, where variance in work per request is small enough to bury under a safety factor.

LLM inference detonates it, because the work per request is dominated by token count, and token count varies wildly. The cost driver is not "a request arrived", it is "how many tokens must be processed, " and that splits into two very different operations. The prefill phase processes all input tokens at once, compute-bound and roughly parallel. The decode phase generates output tokens one at a time, each requiring a full forward pass, memory-bandwidth-bound and inherently sequential. A request with 200 input and 2,000 output tokens stresses the hardware completely differently from one with 50,000 input and 50 output, and a request-based plan treats them as the same unit (NVIDIA on LLM inference metrics).

Throughput for inference is therefore measured in tokens per second, with input (prefill) and output (decode) quoted separately because they are different ceilings. Provisioning on requests per second is like sizing a freight yard by counting trucks while ignoring whether each carries an envelope or forty tons.

Output length variance is the silent killer

Input length you often know roughly in advance, you control the prompt and the context. Output length you do not control: the model decides how much to generate, and across real traffic that decision spans a huge range. A yes/no answer and a full generated document arrive through the same endpoint, indistinguishable until the tokens come out.

This matters more than input variance for two reasons. First, decode is the sequential, slow phase, a 4,000-token response occupies a serving slot roughly forty times longer than a 100-token one. Second, output length is right-skewed with a heavy tail: most responses are short, but a meaningful minority are very long, and the long ones dominate resource consumption. Plan for the mean and the tail eats you, because each tail response holds a slot far longer than the mean predicts. Your effective concurrency collapses precisely when a cluster of long-output requests lands together.

This is the Tuesday brownout. Request rate was at half of projected peak, but a burst of long-output requests arrived together, each occupying a decode slot for many seconds. The available slots filled, the queue backed up, latency spiked, and the system browned out at a request rate the plan swore was safe, because the plan counted requests and the hardware was paying in output tokens.

The KV cache: the constraint nobody put in the spreadsheet

There is a second hard limit specific to LLM serving, and it is usually the real bottleneck before raw compute: GPU memory for the KV cache. During generation, the model keeps a key-value cache for every token in the context of every in-flight request, and that cache grows with context length and with concurrency, living in finite, expensive GPU memory. vLLM's own teardown describes this as variable memory pressure, the number of concurrent requests a worker can hold fluctuates with each request's generation progress, a nonlinear relationship that fundamentally changes scheduling (Inside vLLM).

So your real concurrency ceiling is not "how many requests can the GPU compute" but "how many requests' worth of KV cache fits in GPU memory at once." Long contexts and long outputs both inflate per-request KV footprint, so a handful of long-context requests can consume the memory budget of dozens of short ones. When the KV cache fills, the server queues, evicts, or recomputes, all of which crater throughput. Modern serving stacks like vLLM exist substantially to manage this memory more efficiently with PagedAttention (vLLM documentation), but no cleverness changes the fact that your capacity is fundamentally a token-memory budget. A spreadsheet without a row for KV-cache memory is not modeling the binding constraint.

How to actually plan: tokens per second, from the real distribution

Measure the distribution, not the average. Pull the actual input- and output-token distributions from production, full histograms, not means. You need the p50, p95, and especially p99 of output length, because the tail sizes you. The mean is a lie that averages a one-word reply with a generated essay.

Convert to a token budget. Compute peak tokens per second, separately for prefill and decode, using the distribution. Decode tokens per second is usually the binding constraint because decode is the sequential bottleneck. This is the number that maps to hardware, the way RPS used to.

Benchmark your real token throughput. Do not trust headline tokens-per-second numbers from marketing. Measure your stack, your model, quantization, batch settings, context lengths, under realistic concurrency, because throughput degrades as concurrency and context grow. Establish how many decode tokens per second you sustain at acceptable latency, and how that falls as the KV cache fills.

Size for the tail, autoscale on tokens. Provision for a high percentile of the token-rate distribution, not the mean, because the tail is where brownouts live. Then make autoscaling react to token throughput and KV-cache utilization, not request rate, a request-based autoscaler is blind to the long-output burst saturating you and will sit idle while the queue backs up. This token-versus-request framing is the same one we apply in Token-Aware Rate Limiting; the unit error is identical, it just shows up as brownouts here instead of surprise bills.

Why you got idle GPUs too

The Thursday underutilization is the same mistake wearing the opposite mask. Having been burned Tuesday, the team padded the request-based plan with a fat safety factor, but applied to the wrong unit. On Thursday the traffic was high-request-rate but short-output: lots of quick answers, little decode work. So the GPUs the request-padded plan provisioned sat idle, expensive and bored. Overprovisioned on requests, underutilized on tokens. The request plan cannot win, because it is blind to the variable that actually moves utilization, it overshoots when outputs are short and undershoots when outputs are long, and cannot tell the two cases apart in advance. A token-aware plan reacts to the right signal in both directions, scaling up for the long-output burst and down for the short-output flood.

Load tests must vary output length

Here is where most load testing quietly lies about capacity. The typical LLM load test fires a fixed prompt that produces a roughly fixed-length output, then ramps the request rate until something breaks. It produces a clean capacity number, and that number is fiction, because it holds output length, the single most important and most variable factor, constant at one value.

A real capacity test reproduces the production token distribution, including the heavy tail of long outputs, and especially the correlated bursts of long-output requests arriving together, since that correlation is what causes brownouts average-rate tests never see. It measures capacity in sustained decode tokens per second at your latency SLO, watches KV-cache utilization climb toward its ceiling, and finds the load at which the output tail saturates your serving slots. The output is not "we handle N requests per second." It is "we sustain N decode tokens per second, which given our output distribution means M requests per second on a normal hour and as few as M/3 when the long-output tail clusters."

# Capacity verdict in the unit that actually binds
def capacity_verdict(samples, slot_count, slo_ms):
    decode_tps = sum(s.output_tokens for s in samples) / window_seconds(samples)
    kv_peak    = max(s.kv_cache_bytes_in_flight for s in samples)
    slots_used = max(s.concurrent_decodes for s in samples)
    assert kv_peak  < KV_BUDGET_BYTES, "KV cache is the bottleneck, not compute"
    assert slots_used < slot_count,     "decode slots saturated by long-output tail"
    print(f"sustained decode {decode_tps:.0f} tok/s at p99 < {slo_ms}ms")
This is the gap Hit was built to own: browser-native, streaming-aware load that replays the real output-length distribution, tail and correlated bursts included, and reports sustained tokens per second, so your capacity number survives contact with production.

Count the tokens, not the trucks

The old plan counted trucks entering the yard and assumed each carried the same load. For a decade that was close enough to survive a safety factor. LLM inference breaks it completely: one request might be an envelope, the next forty tons of sequential decode work pinning a serving slot and a slab of KV cache for seconds. Size on tokens per second, built from the real output-length distribution, tail and all. Make the KV cache a line in the spreadsheet. Autoscale on token throughput, not request rate. Then the same week that gave you both a brownout and a field of idle GPUs gives you neither, because you finally measured the cargo instead of counting the trucks. The natural companion is planning for the launch spike, where the burst meets the tail at the worst possible moment, covered in Launch-Day Burst Traffic Is Where AI Features Die.

Pressure-Test Your AI Before Production Does

Hit fires browser-native, streaming-aware load at your LLM and API endpoints, TTFT, inter-token latency, tokens/sec, and cost per request, with no account and no script.

Try Hit Free →
Priya Nair Priya Nair writes about AI quality engineering at alt.qa, built by TheWorkCompany.