TL;DR
Most teams load-test their vector database the way they’d load-test a cache: fire queries, watch latency and QPS, ship when both look fine. That misses the one failure that actually degrades every RAG answer, recall silently dropping as concurrency climbs. Approximate nearest-neighbor search trades accuracy for speed, and under load that trade tilts the wrong way: queues deepen, per-query exploration gets capped, and the index starts returning the wrong neighbors while your p99 still looks green. In a published 2025 benchmark, pgvectorscale hit 471 QPS at 99% recall versus 41 QPS for a default Qdrant config at the same recall, an 11x spread driven entirely by tuning, not hardware. If your load test never measures recall, you are flying blind on the metric that decides answer quality.
The outage with no errors
Here is the incident that never makes it into a post-mortem because nothing technically broke. Traffic to your RAG assistant doubles after a launch. The vector DB dashboard is pristine: p99 query latency 28ms, zero timeouts, CPU at 60%. But support starts seeing “the answers got vague” and “it cited the wrong document.” Product can’t reproduce it on staging, where load is light. Three weeks later someone finally measures retrieval quality under concurrency and discovers the index has been returning recall@10 of 0.78 during peak hours, down from 0.97 at rest. Every answer built on those retrievals was quietly worse, and no monitor fired, because latency was never the problem.
This is the defining property of approximate nearest-neighbor (ANN) search and the reason it needs a fundamentally different load test. A SQL query under load either returns the right rows or times out, failure is loud. An ANN query under load returns some rows, fast, that are merely less likely to be the right ones. The contract was never “correct, ” it was “probably correct, ” and the probability is load-dependent.
Why recall and QPS pull against each other
HNSW, the index almost every vector DB now uses, is a navigable graph. A query walks the graph toward the query vector, and how thoroughly it explores is governed by ef_search (the search-time beam width). As the OpenSource Connections analysis of recall vs. performance (2025) lays out, raising ef_search from 100 to 800 improves recall but increases per-query work roughly linearly. More exploration means more accurate neighbors and more latency. Less exploration means faster queries and a higher chance of missing the true nearest neighbors. This is the entire ballgame, and it is a knob, not a constant.
The trap is what happens to that knob under load. The same analysis recommends capping per-query exploration (ef_search, beam width) and adding admission control during QPS spikes specifically to protect p99 latency. That is sound advice for latency, and it is exactly how recall silently degrades. To hold a latency SLO as concurrency rises, the system (or the engineer tuning it) reduces exploration, which trades away recall to buy back speed. You get to keep your green latency dashboard precisely by sacrificing the metric you weren’t watching.
The IVF version of the same trap
If you’re on an IVF (inverted-file) index instead of HNSW, the knob is nprobe, how many clusters to scan per query. Same dynamic: more probes, more recall, more latency; fewer probes, faster, blinder. Under memory pressure or aggressive batching, effective nprobe can drop, and recall with it. The lesson is index-agnostic: every ANN structure has an accuracy-for-speed dial, and load is the thing most likely to turn it against you.
The benchmark numbers that should change your defaults
Published 2025 benchmarks make the stakes concrete. Beyond the 11x QPS-at-fixed-recall spread between a tuned pgvectorscale and a default Qdrant config, the broader comparisons, see this 2025 pgvector vs. Pinecone vs. Qdrant vs. Weaviate comparison, show that the same dataset can yield wildly different recall at the same latency depending entirely on index parameters. The headline: your vendor’s defaults were chosen for a benchmark, not your workload.
There’s a sharper edge for managed services. As the comparisons note, Pinecone uses a proprietary index and does not expose ef_search, recall is managed internally and can drift under load with no knob you can turn. That is not a reason to avoid managed DBs, but it is a reason you must measure recall yourself end-to-end, because you cannot infer it from a config you can’t see. If you can’t set the dial, you’d better be reading the dial.
How to load-test recall, not just latency
The fix is to measure recall during the load test, against ground truth, at the concurrency you actually expect. Ground truth is an exact (brute-force) k-NN result for a sample of queries, slow, but you only compute it once, offline. Then under load you compare what the ANN index returned to that gold set:
# 1. Build ground truth ONCE with exact search over a query sample
import numpy as np
def exact_knn(query, corpus, k=10):
sims = corpus @ query / (np.linalg.norm(corpus, axis=1) * np.linalg.norm(query))
return set(np.argsort(-sims)[:k]) # the true top-k ids
GOLD = {qid: exact_knn(q, corpus, k=10) for qid, q in sample_queries.items()}
# 2. During the load test, score every ANN result against GOLD
def recall_at_k(returned_ids, gold_ids, k=10):
hits = len(set(returned_ids[:k]) & gold_ids)
return hits / min(k, len(gold_ids))
Now run that comparison while ramping concurrency, and record recall as a distribution per load level, not a single number. The point of the test is to find the concurrency at which recall falls off the cliff:
async def recall_under_load(client, queries, concurrencies=[1,8,32,128,256]):
for c in concurrencies:
recalls, latencies = [], []
async def one(q):
t0 = time.perf_counter()
res = await client.search(q.vector, limit=10) # ANN search
latencies.append((time.perf_counter() - t0) * 1000)
recalls.append(recall_at_k([r.id for r in res], GOLD[q.id]))
await gather_with_concurrency(c, [one(q) for q in queries])
r = np.mean(recalls); p99 = np.percentile(latencies, 99)
flag = " <-- RECALL CLIFF" if r < 0.95 else ""
print(f"c={c:4d} recall@10={r:.3f} p99={p99:6.1f}ms{flag}")
Set a recall SLO and gate on it
Latency gets an SLO; recall almost never does, which is why it rots unnoticed. Give it one. For RAG, retrieval recall is upstream of every quality metric you care about, groundedness, citation accuracy, hallucination rate, so a recall regression is a quality regression with a delay. Treat it like one:
# CI gate: recall must survive the load you expect in prod
- name: vector recall-under-load gate
run: |
python recall_loadtest.py \
--endpoint $STAGING_VECTOR_URL \
--concurrency 128 \
--min-recall-at-10 0.95 \
--max-p99-ms 50
# fails if recall@10 < 0.95 OR p99 > 50ms at 128 concurrent queries
This catches the regressions that matter and that nothing else will: an index rebuild with a lower ef_construction, a quantization change that shrank memory but cost accuracy, a managed-tier downgrade, or a corpus that grew past the point where your ef_search still suffices. Each of those degrades recall and not latency, invisible to every standard load tool.
Where this sits in the RAG stack
Vector search is one hop in a chain, and its load behavior compounds with the rest. Slow or low-recall retrieval feeds bad context into the generation step, which we cover in RAG pipeline latency under load; the embedding step that produces query vectors has its own throttling cliff, covered in the embedding endpoint is your hidden bottleneck. A real RAG load test exercises all three hops at production concurrency simultaneously, because the failure modes interact, and testing them in isolation hides the interactions.
The practical way to do that without standing up a synthetic harness is to drive the real, authenticated endpoints from the browser and instrument every hop. That is the wedge Hit occupies: fire concurrent queries against your real vector and RAG endpoints, measure latency and wire in recall scoring against a gold set, and watch the two outputs diverge under load before your users do.
The numbers that should reset your expectations
If you think your vendor’s defaults are tuned for your workload, the published spreads should disabuse you. In a 2025 head-to-head on a 50M-vector, 768-dimension dataset at 99% recall, a tuned pgvectorscale sustained 471 QPS at 28ms p95, while a default Qdrant config managed 41 QPS, and a Pinecone storage-optimized index hit similar throughput but at 784ms p95, roughly 28x the latency. Same data, same recall target, an order-of-magnitude-plus spread driven by index configuration. The lesson isn’t “use database X”; it’s that the recall-latency-throughput surface is enormous and your position on it is something you set, mostly by accident, when you accept defaults.
Scale changes the picture again. At billion-vector scale the same comparisons show purpose-built engines pushing 10,000-15,000 QPS at 20-50ms p95, but those numbers assume the parameters were tuned for that scale, and the parameters that were fine at 5M vectors are often wrong at 500M. A corpus that grows past the point where your ef_search still suffices is a recall regression with no deploy attached to it: nobody changed the code, the data just grew, and recall quietly slid. This is why a one-time benchmark is worthless and a recurring, load-aware recall test is essential, your operating point moves as your data does.
The bottom line
Vector databases fail quietly. They don’t throw errors when overloaded; they return faster, worse answers, and let your latency dashboard stay green while recall, the number that actually decides whether RAG works, slides toward noise. The only defense is to make recall a first-class, load-tested, SLO-gated metric measured against ground truth at production concurrency. Measure both outputs of your ANN index, not just the one that’s easy to graph. The teams that do stop shipping the “the answers got vague” incident that no monitor ever catches.
Pressure-Test Your AI Before Production Does
Hit fires browser-native, streaming-aware load at your LLM and API endpoints, TTFT, inter-token latency, tokens/sec, and cost per request, with no account and no script.
Try Hit Free →