BlogGoodput, Not Throughput: The Only Load Number That MattersHit · Load & Latency

Goodput, Not Throughput: The Only Load Number That Matters

JK
James Kim · February 2026 · 10 min read

TL;DR

Raw throughput, total tokens or requests per second, is the number your serving engine prints and the number that lies to you. It counts requests that arrived too slowly to be useful exactly the same as ones that delighted a user. The metric that actually maps to revenue is goodput: completed requests per second that meet your latency SLOs (TTFT and time-per-output-token). The gap is not academic. UCSD’s DistServe work showed that optimizing for goodput instead of throughput let the same hardware serve 7.4x more requests, or hold a 12.6x tighter SLO, while keeping over 90% of requests inside their latency budget. If you load-test for throughput, you will happily scale your way into a worse product.

The 2,000 req/s that nobody could use

Picture the load-test readout your platform team brings to the capacity review: “Sustained 2,000 requests per second, no errors, GPUs at 94% utilization.” Everyone nods. The cluster is sized, the launch is greenlit. Three weeks later the support queue fills with “the assistant takes forever to start typing, ” and the product dashboard shows session abandonment climbing in lockstep with traffic. Nothing errored. Utilization is still beautiful. And the company is losing users at the exact moment the load test said it was winning.

This is the throughput trap. A serving engine batches aggressively to keep the GPU fed, and at high concurrency that batching inflates per-request latency, the queue gets deeper, prefill and decode start interfering, and time-to-first-token quietly slides from 400ms to 4 seconds. The engine still completes 2,000 requests a second, so throughput looks magnificent. But a request that takes 9 seconds to first token is, for a chat product, a failed request wearing a success costume. Raw throughput counts it as a win.

The core insight: Throughput asks “how many requests finished?” Goodput asks “how many requests finished well enough to keep the user?” The first number is what your hardware can emit. The second is what your business can sell. They diverge precisely under the load you most need to plan for.

What goodput actually means

Goodput is borrowed from networking, where it means the application-level data delivered minus retransmissions and overhead, the bytes that were actually useful. The LLM-serving definition is the direct analogue. In the DistServe paper (Zhong et al., OSDI 2024), goodput is formalized as the maximum request rate a system can sustain while still satisfying SLO-attainment targets, typically expressed as two per-phase latency SLOs: time-to-first-token (TTFT) for the prefill phase and time-per-output-token (TPOT) for decode. A request only counts toward goodput if it meets both.

That single change of accounting is what surfaces the divergence. A 2025 survey revisiting SLOs and system-level metrics in LLM serving puts it bluntly: almost every popular engine, vLLM, TensorRT-LLM, SGLang, reports throughput as the headline number, because operators treat it as a proxy for cost per request. But throughput-as-proxy only holds if every completed request is equally valuable, and under latency SLOs that assumption is false. The survey frames the pair correctly: SLO attainment is the constraint, goodput is the objective. You maximize goodput subject to holding attainment above a threshold (say, 90% or 99% of requests inside budget).

Why the two numbers separate exactly when it matters

At low load, throughput and goodput are nearly identical: every request is fast, every request meets SLO, the curves overlap. As you push concurrency up, throughput keeps climbing toward the hardware ceiling, but goodput peaks earlier and then falls, because the marginal request you admit pushes everyone’s latency past the SLO line. The Hao AI Lab write-up of DistServe (“Throughput is Not All You Need”) shows this shape directly: there is a load level that maximizes useful work, and it sits well below the load level that maximizes raw throughput. Plan to the throughput peak and you are operating in the region where goodput is collapsing.

Where the useful work leaks away

Three mechanisms turn throughput into garbage under load, and a goodput-aware load test catches all three:

  • Prefill-decode interference. When prefill (compute-bound, reading the prompt) and decode (memory-bandwidth-bound, generating tokens) share the same GPU and the same batch, a burst of long-prompt requests stalls the token stream for everyone already decoding. TTFT and TPOT degrade together. Disaggregating the two phases onto separate resources is precisely how DistServe recovered its goodput, and why the metric is what exposed the problem in the first place.
  • Queueing delay masquerading as compute. Past the saturation point, added latency is mostly time spent waiting in the admission queue, not time spent generating. Throughput is blind to this; goodput is not, because queued-too-long requests blow their TTFT SLO.
  • Tail amplification. Aggressive batching trades median latency for tail latency. A throughput dashboard reports the mean. Your users live in the p95 and p99, which is exactly where SLO violations, and abandonment, concentrate.

Measuring goodput, not throughput, in your load test

The instrumentation difference is small and the interpretive difference is enormous. You already have to read a streaming response to get TTFT and inter-token latency. Goodput is just those per-request samples filtered through your SLO and counted per unit time. Here is the per-request measurement against a streaming chat endpoint:

// Per-request sample: TTFT and TPOT from a streaming completion
async function sampleRequest(url, body, headers) {
  const t0 = performance.now();
  const res = await fetch(url, { method: 'POST', headers, body: JSON.stringify(body) });
  const reader = res.body.getReader();
  let ttft = null, firstTokenAt = 0, tokens = 0;

  while (true) {
    const { done, value } = await reader.read();
    if (done) break;
    const now = performance.now();
    if (ttft === null) { ttft = now - t0; firstTokenAt = now; }
    tokens += countSSETokens(value);            // count `data:` deltas
  }
  const decodeMs = performance.now() - firstTokenAt;
  return {
    ttft_ms: ttft,
    tpot_ms: tokens > 1 ? decodeMs / (tokens - 1) : decodeMs,  // time per output token
    ok: true,
  };
}

Now define your SLO and reduce a wall of samples to the single number that matters. Goodput is SLO-passing requests divided by the wall-clock duration of the test window, not divided by the count of requests, which is the subtle but critical part:

def goodput(samples, window_seconds, ttft_slo_ms=500, tpot_slo_ms=50):
    passed = [s for s in samples
              if s["ttft_ms"] <= ttft_slo_ms and s["tpot_ms"] <= tpot_slo_ms]
    attainment    = len(passed) / len(samples)
    raw_throughput = len(samples) / window_seconds   # what the engine brags about
    good           = len(passed)  / window_seconds   # what the user actually got
    print(f"throughput   {raw_throughput:6.1f} req/s")
    print(f"goodput      {good:6.1f} req/s  ({attainment*100:.1f}% SLO attainment)")
    print(f"wasted work  {raw_throughput - good:6.1f} req/s burned on SLO misses")
    return good

Run this at increasing concurrency and plot goodput against offered load. The peak of that curve, not the peak of the throughput curve, is your real capacity number. Everything to the right of it is hardware you are paying for to produce requests nobody can use.

Capacity planning rule: Size your fleet to the offered load where the goodput curve peaks, then keep ~20-30% headroom below it. Sizing to the throughput peak guarantees you operate in the region where adding traffic destroys SLO attainment, the worst possible place to run a launch.

Goodput is a token problem, not a request problem

One reason throughput misleads so reliably for LLMs is that requests are wildly non-uniform. A 50-token summarization and a 4,000-token RAG answer are both “one request, ” but they consume radically different prefill and decode budgets and have radically different SLO profiles. A request-per-second number averages over that variance and hides it. This is the same reason capacity for AI workloads has to be planned in tokens, not requests, a point we make in depth in capacity planning by tokens, not requests. Goodput inherits the fix: define your SLOs per phase (TTFT for prefill, TPOT for decode) so that long and short requests are each judged against the latency the user actually feels for that request, rather than collapsed into a meaningless average.

Pick SLOs you can defend

Goodput is only as honest as the SLO you filter on. Set TTFT and TPOT from product evidence, not vibes. For a consumer chat assistant, a defensible pair is TTFT p95 < 500ms and TPOT < ~50ms (≈20 tokens/sec, faster than most people read). For an internal batch summarizer, TTFT can be seconds and TPOT can be loose, and your goodput target should reflect that, or you will over-provision for latency the user never notices. The number is meaningless without the SLO attached, so always report goodput and the SLO it was measured against.

Wire it into the deploy, not the post-mortem

The failure mode that motivates all of this, “throughput looked great, users left”, is a regression you can catch before it ships. Treat the goodput-at-target-load number like a test assertion: a config change that trades 15% of your goodput for 5% more raw throughput should turn the build red, because it is a product regression dressed as an optimization.

# CI gate: goodput must hold at the planned concurrency
- name: LLM goodput gate
  run: |
    python loadtest.py --url $STAGING --concurrency 64 --duration 120 \
      --ttft-slo 500 --tpot-slo 50 --min-goodput 180 --min-attainment 0.95
    #   fails the build if goodput < 180 req/s or SLO attainment < 95%

The same browser-native instrumentation that measures TTFT can measure goodput against your real, authenticated endpoint, no proxy, no synthetic mock, no token juggling. That is the gap Hit was built to close: fire concurrent streaming requests from the browser, capture TTFT and TPOT per request, and read goodput directly off the run instead of inferring it from a throughput chart that is structurally incapable of telling you the truth.

The bottom line

Throughput is the number your serving engine wants you to admire; goodput is the number your users actually receive. They agree right up until the moment you need them most, at the load where you are deciding how much hardware to buy and whether to ship. Optimizing throughput while goodput silently collapses is how teams scale themselves into higher infrastructure bills and worse retention simultaneously. Define per-phase SLOs, measure SLO-passing completions per second, plan to the goodput peak, and gate on it in CI. The teams that do this stop being blindsided by the launch where every dashboard was green and every user was leaving.

Pressure-Test Your AI Before Production Does

Hit fires browser-native, streaming-aware load at your LLM and API endpoints, TTFT, inter-token latency, tokens/sec, and cost per request, with no account and no script.

Try Hit Free →
James Kim James Kim writes about AI quality engineering at alt.qa, built by TheWorkCompany.