TL;DR
Conventional load tools report request duration, first byte to last byte, which is meaningless for a streaming LLM. The metric that actually predicts whether a user stays is Time to First Token (TTFT): how long the screen sits blank before the first word appears. Stanford HCI work in 2025 found users perceive responses under ~300ms as instant, tolerate up to ~800ms, judge 800-2,000ms as "slow, " and begin abandoning above 2,000ms, a 15-25% drop-off. Yet most teams have never measured their TTFT distribution under load, because their tooling can't. This is the single highest-leverage AI latency metric you are probably not tracking.
Your dashboard says 200ms. Your users wait 8 seconds.
Here is the failure that keeps shipping. A team load-tests their new AI feature with k6 or Locust, sees "p99 request duration 780ms, zero errors, " and ships with confidence. In production, users stare at a spinner for six, eight, ten seconds before the first word appears. Support tickets pile up. The dashboard stays green the whole time.
The dashboard is green because it is measuring the wrong thing. Tools built for stateless REST endpoints record the time from request to complete response. For a streaming chat completion, "complete" can mean a thousand tokens generated over eight seconds, so a single number conflates two completely different phases: the prefill (the model reading your prompt and producing the first token) and the generation (streaming the rest). Users experience these phases as two unrelated things, and only the first one decides whether they wait.
What the research actually says about waiting
The thresholds are not folklore. According to work summarizing 2025 HCI research on AI chat latency, user perception breaks into clean bands:
- Under 300ms TTFT, perceived as instant.
- 300-800ms, a brief pause, still acceptable.
- 800-2,000ms, perceived as "slow."
- Over 2,000ms, abandonment begins, with a documented 15-25% conversation drop-off.
The practical target that falls out of this for any consumer-facing assistant is a P95 TTFT under 500ms. That is a percentile target, not an average, because the user who waits 4 seconds doesn't feel better knowing the median was fast. And it is the percentile that conventional load testing buries.
The gap between providers is enormous, which is why this is a measurable, controllable lever rather than a fixed cost. Independent 2026 latency benchmarking across major APIs shows speed-optimized inference providers delivering ~120-180ms median TTFT, 3-5x faster than the default endpoints many teams reach for first. If you have never measured your own TTFT, you have no idea which side of the abandonment cliff you are on.
Why prefill, not generation, is where TTFT lives
TTFT is dominated by prefill: the model has to ingest and attend over every token in your prompt before it can emit the first output token. That means TTFT scales with input length, not output length. Three things you control inflate it:
- Bloated system prompts. A 3,000-token system prompt is paid on every single request, in latency and in cost.
- Unbounded RAG context. Stuffing 20 retrieved chunks "to be safe" can triple prefill time for marginal quality gain.
- Long conversation history. Replaying the full transcript each turn makes turn 30 dramatically slower than turn 1, a regression users feel mid-session.
This is why prompt caching matters so much for TTFT: a cached prefix skips re-reading the static portion of your prompt. But cache behavior under concurrent, varied load is itself something you have to test, assumptions made on a warm single-user cache rarely survive production traffic.
Measuring TTFT correctly
You cannot measure TTFT by awaiting the response body. You have to read the response as a stream and timestamp the moment the first chunk arrives. In the browser, fetch() exposes the body as a ReadableStream, so you can do this with no backend at all:
// Browser-native TTFT measurement against a streaming endpoint
async function measureTTFT(url, body, headers) {
const t0 = performance.now();
const res = await fetch(url, { method: 'POST', headers, body: JSON.stringify(body) });
const reader = res.body.getReader();
let ttft = null, lastChunk = t0, gaps = [], tokens = 0;
while (true) {
const { done, value } = await reader.read();
if (done) break;
const now = performance.now();
if (ttft === null) ttft = now - t0; // ← first byte of first token
else gaps.push(now - lastChunk); // ← inter-token latency samples
lastChunk = now;
tokens += countSSETokens(value); // parse `data:` deltas
}
const total = performance.now() - t0;
return {
ttft_ms: ttft,
itl_ms_avg: gaps.reduce((a, b) => a + b, 0) / Math.max(gaps.length, 1),
tps: tokens / (total / 1000),
total_ms: total,
};
}
Run that under concurrency and aggregate the TTFT samples into percentiles, p50, p90, p99, and you finally have the number that matches what users feel. The same loop yields inter-token latency (the "stutter" between words) and tokens/sec (generation speed). One request, three honest metrics.
# Turn raw samples into the verdict that matters
import numpy as np
def ttft_report(samples_ms):
p50, p90, p99 = np.percentile(samples_ms, [50,90,99])
print(f"TTFT p50 {p50:.0f}ms · p90 {p90:.0f}ms · p99 {p99:.0f}ms")
# Consumer assistant target: P95 < 500ms
p95 = np.percentile(samples_ms, 95)
verdict = "PASS" if p95 < 500 else "FAIL"
print(f"P95 {p95:.0f}ms → SLA {verdict}")
# Estimated abandonment exposure
slow = np.mean(np.array(samples_ms) > 2000) * 100
print(f"{slow:.1f}% of requests cross the 2s abandonment line")
Turn TTFT into a pass/fail gate, not a vibe
The reason latency regressions ship is that "it feels a little slower" is not actionable. Make it actionable by asserting on TTFT in CI, the same way you assert on a unit test:
assert ttft_p95 < 500, f"TTFT P95 regressed: {ttft_p95}ms"
assert tps_p50 > 25, f"Generation throughput dropped: {tps_p50} tok/s"
Now a prompt change that quietly doubles prefill time, a model swap with a slower cold path, or a RAG tweak that triples context length all turn into a red build, before a single user waits. That is the whole game: move the discovery of a latency regression from your support inbox to your pull request.
The bottom line
If you remember one thing: request duration is a lie for streaming AI, and TTFT is the truth. It is the metric that maps directly onto whether a user waits or leaves, it scales with prompt design choices you control, and it is invisible to the load tools most teams reach for. Measure your TTFT distribution under realistic concurrency, set a P95 budget, and gate on it. The teams that do this stop being surprised by "it feels slow" tickets, because they caught it first.
Pressure-Test Your AI Before Production Does
Hit fires browser-native, streaming-aware load at your LLM and API endpoints, TTFT, inter-token latency, tokens/sec, and cost per request, with no account and no script.
Try Hit Free →