TL;DR
Everyone load-tests the chat model. Almost nobody load-tests the embedding endpoint, the unglamorous step that turns every user query and every ingested document into a vector. It is also the first thing to throttle. Embedding calls are cheap per token (text-embedding-3-small is $0.02 per 1M tokens), which lulls teams into firing them with abandon, until a bulk re-index or a query spike slams into a tokens-per-minute ceiling and 429s cascade back through the whole RAG pipeline. When the embedder stalls, retrieval stalls, and the user gets a spinner or a context-free answer. This post is about finding that ceiling before production does.
The re-index that took down search at 2pm
The failure looks like this. Your team ships a new chunking strategy and kicks off a re-embed of the document corpus, a few million chunks, no big deal, it’s only embeddings. The batch job hammers the same embedding endpoint your live query path uses. By early afternoon, every user query that needs an embedding starts getting 429 Too Many Requests, because the batch job ate the entire tokens-per-minute budget. Retrieval returns nothing, the assistant answers from the bare model with no context, and answer quality falls off a cliff. The chat model never broke. The vector DB never broke. The embedding endpoint, the part nobody load-tested, quietly became a single point of failure for the whole feature.
This is the central, under-appreciated fact about embeddings: they sit on the critical path twice. Once at ingestion (documents to vectors) and once at query time (the user’s question to a vector, every single request). Both paths usually share one endpoint and one rate limit. A spike on either side starves the other, and because embedding is upstream of retrieval, its throttling doesn’t look like an embedding problem, it looks like “search is broken.”
The limits you’re actually fighting
Embedding throughput is governed by two ceilings simultaneously, and you hit whichever you reach first. The OpenAI rate-limit documentation describes both: RPM (requests per minute) and TPM (tokens per minute), scaled by usage tier. A naive ingestion job that sends one chunk per request burns through RPM long before TPM, you’re rate-limited on request count while sending a trickle of tokens. A job that crams huge batches into each request can blow TPM while sending few requests. The right batch size threads between them, and you only find it by measuring.
Tiers matter more than people expect. Limits scale with cumulative spend, a Tier 1 account (after $5 of usage) has dramatically lower ceilings than Tier 4 or 5. A load test that passes on a high-tier production account will fail on the lower-tier staging account you actually test against, giving you false confidence in both directions. Test against limits that match production, or your numbers are fiction.
The batch API is not the live path
OpenAI’s embeddings guide points bulk workloads at the Batch API, which offers higher limits and a 50% discount, genuinely the right tool for re-indexing. But it is asynchronous, with completion windows measured in hours, not the synchronous low-latency path your query embeddings need. The architectural mistake that causes the 2pm outage is routing both through the same synchronous endpoint. Ingestion should go through Batch; query embedding should go through the realtime endpoint with reserved headroom. Your load test exists to prove that separation holds under a simultaneous ingestion + query spike.
Batching: the lever that decides your throughput
A single embedding request can carry many inputs in one array, and this is the single biggest throughput lever you control. One request with 100 chunks is vastly more efficient than 100 requests with one chunk each, because it spends RPM like a single request while moving 100x the tokens. But batch too large and you risk per-request payload limits, longer tail latency, and a single 429 taking out 100 chunks instead of one. Here is a batched embedder with the backpressure that production actually needs:
import asyncio, time
from openai import AsyncOpenAI
client = AsyncOpenAI()
BATCH_SIZE = 96 # tune empirically: balances RPM vs TPM vs tail latency
MAX_CONCURRENCY = 8 # parallel in-flight batches
sem = asyncio.Semaphore(MAX_CONCURRENCY)
async def embed_batch(chunks, attempt=0):
async with sem:
try:
t0 = time.perf_counter()
res = await client.embeddings.create(
model="text-embedding-3-small", input=chunks)
return res.data, (time.perf_counter() - t0) * 1000
except Exception as e:
if "429" in str(e) and attempt < 5:
await asyncio.sleep(2 ** attempt) # exponential backoff
return await embed_batch(chunks, attempt + 1)
raise
def chunked(seq, n):
for i in range(0, len(seq), n):
yield seq[i:i + n]
The numbers behind the batch size are simple but easy to get wrong. At $0.02 per 1M tokens, embedding 5 million chunks of ~400 tokens each is 2 billion tokens, about $40. Trivial on cost. But at a TPM ceiling of, say, 1M tokens/minute, that same job needs 33+ hours of continuous throughput at the limit. The bill is nothing; the time is everything, and the time is what collides with your live traffic. This is why you plan embeddings by throughput, not by cost.
Load-testing the embedding endpoint properly
A real embedding load test does three things conventional tests skip: it measures the actual sustained tokens-per-minute you can push before 429s begin, it runs ingestion and query load simultaneously to expose starvation, and it records the 429 rate as a first-class metric rather than an error to be retried away. Here’s the simultaneous-pressure harness:
async def embedding_starvation_test(query_qps, ingest_concurrency, duration_s):
stats = {"query_429": 0, "query_ok": 0, "query_ttfb": [],
"ingest_429": 0, "ingest_ok": 0}
async def query_load(): # simulates live user queries
end = time.time() + duration_s
while time.time() < end:
t0 = time.perf_counter()
try:
await client.embeddings.create(
model="text-embedding-3-small", input=["user query text"])
stats["query_ok"] += 1
stats["query_ttfb"].append((time.perf_counter() - t0) * 1000)
except Exception as e:
if "429" in str(e): stats["query_429"] += 1
await asyncio.sleep(1 / query_qps)
async def ingest_load(): # simulates a bulk re-index hammering the SAME endpoint
for batch in chunked(big_corpus, BATCH_SIZE):
try:
await embed_batch(batch); stats["ingest_ok"] += 1
except Exception as e:
if "429" in str(e): stats["ingest_429"] += 1
await asyncio.gather(query_load(),
*[ingest_load() for _ in range(ingest_concurrency)])
return stats
query_429 climbs the moment ingestion starts, your live query path is starving, ingestion is eating the shared budget. That’s the 2pm outage, reproduced on demand in staging. The fix (separate Batch API path, reserved query headroom, or a token-aware rate limiter) is now testable: re-run and confirm query_429 stays at zero while ingestion runs.The economics that change the architecture
It’s worth running the actual numbers, because they explain why “just use the API” quietly stops being the right answer at scale. At $0.02 per 1M tokens for text-embedding-3-small, a production RAG system handling roughly 10 million documents plus 100,000 daily queries lands around $50-100 per day in embeddings, $1,500-3,000 per month. Trivial against an engineering salary, which is why most teams never think about it. But two things change the calculus. First, chunk overlap, the standard practice of overlapping adjacent chunks so context isn’t severed mid-sentence, adds 10-25% to your token count, and therefore your bill, invisibly. Second, the same source notes that self-hosting an embedding model becomes cheaper than the API somewhere around 10-15 million embeddings per month. The point isn’t which side you land on; it’s that the decision is driven by throughput volume, the exact thing your load test measures, not by the per-token price you glanced at once.
Upgrading to text-embedding-3-large for better retrieval quality is the move that quietly breaks the model. It’s 6.5x the token cost ($0.13 vs $0.02 per 1M), so that $2,000/month line becomes $13,000, and because larger embeddings mean more bytes per vector, it also inflates your vector DB storage and memory footprint downstream. A load test that measures sustained throughput and cost per the model you actually ship is the only way to catch this before the invoice does. Cheap-per-token is a trap precisely because it discourages the measurement that would reveal the real cost at your volume.
Make 429-under-load a CI gate
The regression you’re guarding against is subtle: someone bumps the ingestion concurrency, or swaps to text-embedding-3-large (6.5x the token cost and different limit behavior), or a new feature adds a second high-volume embedding caller. Any of these can re-introduce starvation. Gate on the query-path 429 rate while a synthetic ingestion job runs:
# CI: query embeddings must not starve while ingestion runs
- name: embedding starvation gate
run: |
python embed_loadtest.py \
--query-qps 50 --ingest-concurrency 8 --duration 90 \
--max-query-429-rate 0.0 \
--max-query-p95-ms 300
# fails if ANY live query gets 429 during a concurrent bulk ingest
This is the inverse of the usual rate-limit advice. The standard guidance, exponential backoff, retries, is correct for resilience but hides the starvation in your load test, because retried requests eventually succeed and the error rate looks fine while real latency quietly balloons. Measure the raw 429 rate and the p95 latency separately, before backoff papers over them.
Where it fits in the RAG picture
Embedding throttling is one of three load failures that interact in a RAG system. The vectors it produces feed an ANN index whose recall silently drops under load, and the whole chain’s end-to-end latency is its own discipline, covered in RAG pipeline latency under load. Because embedding sits earliest in the chain, its failures are the most disguised, they always present as a downstream symptom. The only reliable way to attribute the symptom to the right cause is to instrument every hop under simultaneous load.
That cross-hop, real-endpoint instrumentation is exactly what Hit is built for: drive your real embedding and retrieval endpoints from the browser at production concurrency, watch the 429 rate and latency per hop, and catch the starvation before it shows up as “search is broken” in the support queue.
The bottom line
The embedding endpoint is cheap, invisible, and on the critical path twice, the perfect recipe for an unmonitored single point of failure. Its constraint is tokens-per-minute and requests-per-minute, not dollars, so the bill stays tiny while throughput collapses. Load-test it under simultaneous ingestion and query pressure, separate bulk ingestion onto the Batch API, reserve headroom for live queries, and gate on the query-path 429 rate. Do that and you stop shipping the outage where the chat model worked perfectly and the feature was still down.
Pressure-Test Your AI Before Production Does
Hit fires browser-native, streaming-aware load at your LLM and API endpoints, TTFT, inter-token latency, tokens/sec, and cost per request, with no account and no script.
Try Hit Free →