TL;DR
A RAG pipeline is four serial stages, embed the query, retrieve from the vector store, rerank the candidates, generate the answer, and every one of them is a latency bomb that goes off under load, at a different time, for a different reason. Reranking with a cross-encoder routinely adds 300-800ms, and under real load a cross-encoder's p99.9 has been measured at over 21 seconds at just 40 QPS. Vector search that returns in single-digit milliseconds on a warm, small index degrades sharply as the index grows and concurrency climbs. Because the stages are serial, the user waits for the sum, and the worst case is the sum of each stage's tail. Test the stages separately and you will pass every check and still ship a pipeline that crawls. This is how to find all four bombs before load does.
The pipeline that benchmarks great and ships slow
RAG looks simple on a slide: take the user's question, find relevant documents, hand them to the LLM, get a grounded answer. Four boxes, four arrows. What the slide hides is that those four boxes are four independent systems with four independent failure modes, wired in series so the user pays the sum of all of them, plus the sum of all their tails when traffic arrives.
The reason RAG latency surprises teams is that each stage is fast in isolation, warm, at low concurrency. You benchmark the embedding model: a few milliseconds. You benchmark the vector search on a small index: single-digit milliseconds. The reranker on ten passages: fine. The LLM: streaming nicely. Every component passes. Then you put them in series, point real traffic at them, and the pipeline that was supposed to answer in 800ms is answering in three seconds. A representative breakdown of RAG pipeline latency puts the realistic distribution at query processing 50-200ms, vector search 100-500ms, retrieval 200-1000ms, reranking 300-800ms, and generation 1000-5000ms, a total of 2-7 seconds for a single query. The bombs do not go off in the benchmark. They go off under load.
Bomb #1, Embedding the query
The smallest stage, but not free. The query has to be embedded by a model before you can search, and if you call a hosted embedding API you have added a network round trip plus that provider's own queueing under load. A self-hosted embedding model competes for the same GPU resources as everything else. The trap here is hidden batching: providers batch embedding requests for throughput, which is great for cost and terrible for the single interactive query that now waits for a batch window to fill. Measure the single-query embedding latency, not the batched throughput number on the spec sheet.
Bomb #2, Vector retrieval
Approximate nearest-neighbor search (HNSW, IVF, and friends) is genuinely fast, until it isn't. Three things detonate it: index size (recall-preserving search over tens of millions of vectors costs more than over a hundred thousand), filtering (metadata filters can force the index to scan far more candidates to return k results), and concurrency (queries contend for memory bandwidth and CPU). Vendor writing such as Pinecone's on serverless vector search is candid that query latency is a function of index scale and load, not a fixed constant, and that the comfortable single-digit-millisecond numbers are warm-cache, low-concurrency figures. Under a cold partition or a burst, retrieval latency can jump by an order of magnitude.
# Where the time goes in retrieval (illustrative, warm vs loaded)
stage warm/low-QPS under load / large index
-------------------------------------------------------------------
ANN search 3-10 ms 50-200 ms
metadata filtering +1-5 ms +20-100 ms (low-selectivity filters)
cold partition fetch ~0 +100-500 ms (first hit after scale)
Bomb #3, Reranking (the big one)
This is the stage teams add for quality and forget to budget for latency. Retrieval gives you a coarse top-k (say 50 candidates) cheaply; a reranker, typically a cross-encoder, then scores each candidate against the query for a precise ordering, and keeps the best few. The catch is architectural: a bi-encoder embeds query and document separately (cheap, cacheable), but a cross-encoder runs the query and each document together through a transformer, so its cost scales linearly with the number of candidates. As a 2026 guide to reranking in RAG spells out, if your cross-encoder takes 50ms per document and you rerank 100 candidates, that is 5 seconds of latency, and under load the picture is far worse: one benchmark cited there measured a cross-encoder's p99.9 at over 21 seconds at just 40 QPS.
In practice the reranker is frequently the single largest non-LLM latency in the pipeline, hundreds of milliseconds at best, seconds at worst, and it scales with how many candidates you feed it. The instinct to "retrieve 100 and rerank them all to be safe" is a direct, multiplicative latency cost. The fix is a funnel: retrieve broad, rerank a tight shortlist. Optimized deployments can hit 60-110ms for k≈40-60 on FP16/INT8, but only if you bound k.
Bomb #4, Generation
The LLM stage carries every problem from the rest of this blog, TTFT, inter-token latency, batching contention, plus one that RAG creates specifically: context bloat. Every retrieved chunk you stuff into the prompt lengthens the input, and because TTFT is prefill-bound, more context means a slower first token. NVIDIA's inference fundamentals make the dependency explicit: prefill cost grows with input length. So the very thing RAG does, adding retrieved context, directly inflates the generation stage's TTFT. Retrieving 20 chunks "for safety" can triple your prefill time for a marginal quality gain. The reranker funnel pays off twice here: fewer, better chunks mean a faster first token and a cheaper bill.
Why testing the stages separately is a trap
Each bomb has a different fuse. Embedding slows when a hosted provider batches under load. Retrieval slows when the index grows or a burst contends. Reranking slows linearly with candidate count and contends for its own GPU. Generation slows with context length and batch pressure. Because the fuses are different, the stages do not peak at the same load, and because the pipeline is serial, the user-facing p99 is the sum of whichever stages happen to be in their tail on a given request. The classic tail-amplification reasoning from Dean and Barroso's "The Tail at Scale" applies: a serial chain inherits a slow link from any stage, so testing each box green in isolation guarantees nothing about the end-to-end tail.
Measuring all four under load
Instrument each stage with a span and report the full breakdown as a distribution under concurrency, so you can see which bomb went off:
# Per-query RAG latency breakdown
import time
def rag_query(q):
t = time.perf_counter(); spans = {}
def lap(k):
nonlocal t; now = time.perf_counter(); spans[k] = (now - t) * 1000; t = now
qvec = embed(q); lap('embed')
cands = vector_search(qvec, k=50); lap('retrieve')
top = rerank(q, cands, keep=5); lap('rerank')
answer = generate(q, top); lap('generate') # measure TTFT here
spans['total'] = sum(spans.values())
return answer, spans
import numpy as np
def rag_load_report(all_spans, budget_ms):
for stage in ['embed', 'retrieve', 'rerank', 'generate', 'total']:
xs = [s[stage] for s in all_spans]
p50, p95, p99 = np.percentile(xs, [50,95,99])
print(f" {stage:8} p50 {p50:6.0f} p95 {p95:6.0f} p99 {p99:6.0f} ms")
totals = [s['total'] for s in all_spans]
over = np.mean(np.array(totals) > budget_ms) * 100
print(f"{over:.1f}% of queries exceed {budget_ms}ms budget")
# Find the bomb: which stage's p99 is the largest under load?
p99s = {k: np.percentile([s[k] for s in all_spans], 99)
for k in ['embed', 'retrieve', 'rerank', 'generate']}
worst = max(p99s, key=p99s.get)
print(f"dominant tail stage under load: {worst} ({p99s[worst]:.0f}ms)")
That last line is the payoff: under load, the dominant stage is often not the one you optimized. Teams obsess over the LLM and discover the reranker or a contended vector partition is eating the budget.
Gate on the pipeline
assert rag_total_p95_ms < 1500, f"RAG p95 over budget: {rag_total_p95_ms}ms"
assert rag_total_p99_ms < 3000, f"RAG p99 tail too long: {rag_total_p99_ms}ms"
assert rerank_p95_ms < 400, f"reranker is the bottleneck: {rerank_p95_ms}ms"
assert retrieve_p99_ms < 200, f"vector search tail under load: {retrieve_p99_ms}ms"
Now growing the index past a recall threshold, raising the rerank candidate count, swapping the embedding provider for a slower-batching one, or stuffing more chunks into the prompt all turn into a red build, instead of a pipeline that quietly drifts from 800ms to three seconds as your corpus and traffic grow.
The bottom line
RAG is four serial stages, and each is a different latency bomb with a different fuse: embedding waits on batched providers, retrieval degrades with index size and concurrency, reranking scales linearly with candidate count and is often the largest non-LLM cost, and generation inflates with the context RAG itself adds. Because they run in series, the user pays the sum, and the p99 is the sum of the tails. Test each stage alone and they all pass; the pipeline still crawls. Instrument every stage, run the whole pipeline under concurrency, find which bomb dominates the tail, and gate on the end-to-end budget. That is how you ship a RAG system that stays fast as your corpus and your traffic grow.
Pressure-Test Your AI Before Production Does
Hit fires browser-native, streaming-aware load at your LLM and API endpoints, TTFT, inter-token latency, tokens/sec, and cost per request, with no account and no script.
Try Hit Free →