TL;DR
Your five-minute load test passed and you shipped. Three days later, at 3am, the pager goes off, memory crept past the limit, the connection pool drained dry, the KV cache fragmented the GPU. These are time-dependent failures: they don't depend on how hard you push, they depend on how long you run. A short spike test cannot find a leak that takes six hours to matter, practitioners recommend soak runs of at least 4 hours precisely because a 10-minute test catches none of them. Soak testing, sustained, realistic load over hours to days, is the only way to surface the slow failures that detonate after launch, when traffic has been steady long enough for the leaks to fill up.
The failures a five-minute test will never find
Load testing has two orthogonal axes that teams routinely conflate. One is intensity, how much concurrent load you apply, and a spike or stress test pushes that axis to find your throughput ceiling. The other is duration, how long you sustain load, and this is the axis almost nobody tests, because it's slow and unglamorous and the short test already went green.
But an entire class of production failures lives only on the duration axis. They're invisible at any intensity for the first few minutes and inevitable after a few hours. Soak testing (also called endurance or longevity testing) exists to find them: you run a realistic, steady load, not a peak, just normal traffic, for 6,12, or 24 hours, and watch the slow-moving resource curves a short test never lets accumulate. The bug isn't triggered by force. It's triggered by time.
The four slow killers
1. Memory creep
The classic. A small per-request allocation that never gets freed, a cached object that's never evicted, a growing list of session references, an unbounded log buffer, leaks a few kilobytes per request. At 5 minutes and 10,000 requests, invisible. At 12 hours and tens of millions of requests, gigabytes, and the process crosses its memory limit and gets OOM-killed. The restart clears it, so it looks transient, and the cycle repeats, a sawtooth of crashes that only a long run reveals as a steady upward trend.
2. Connection pool exhaustion
Every request that borrows a database connection, an HTTP client, or a provider socket must return it. Miss the return on an error path, a branch that fires rarely, and you leak one connection per occurrence. The pool has, say, 100 slots; a leak of one per thousand requests drains it after 100,000 requests. The feature runs perfectly for hours, then every request starts blocking on "waiting for a free connection" and the whole service hangs at once. Short tests don't run long enough to drain the pool, and don't accumulate enough of the rare error path to leak.
3. KV cache and GPU memory growth
This one is specific to LLM serving. Transformer inference maintains a key-value cache, and serving frameworks like vLLM manage GPU memory aggressively to pack many sequences in, its PagedAttention design exists precisely to reduce KV-cache fragmentation. But KV-cache usage grows with both sequence length and concurrent-request count, and under sustained, varied real traffic, mixed lengths, long sessions, bursty arrivals, GPU memory can still fragment over time. As the cache fills, throughput in tokens/sec grows until the GPU's KV cache is saturated, then degrades sharply rather than gradually, or sessions start getting rejected for lack of cache space. A short test with uniform prompts never exercises the fragmentation that real, varied, long-running traffic produces.
4. Downstream and cache rot
Caches that grow without bounds, semantic caches whose hit rate decays as they fill with stale entries, retry queues that accumulate faster than they drain, file descriptors that leak on a rare path, log volumes that fill a disk, rising garbage-collection pressure, the failure catalog that endurance testing exists to surface, as RadView's guide to load, stress, capacity and soak testing lays out. All fine for minutes and fatal for days. The common thread: a slow accumulation with no corresponding release, where the release bug only shows up over time.
# Why a leak invisible in a short test is fatal in a soak
def leak_to_failure(leak_kb_per_req, rps, headroom_mb):
leak_mb_per_hour = leak_kb_per_req * rps * 3600 / 1024
hours = headroom_mb / leak_mb_per_hour
print(f"leak: {leak_mb_per_hour:.1f} MB/hr -> OOM in {hours:.1f} hours")
leak_to_failure(leak_kb_per_req=8, rps=50, headroom_mb=4096)
# leak: 1406.2 MB/hr -> OOM in 2.9 hours
# A 5-minute test sees ~117 MB of creep and shrugs. The soak finds the cliff.
How to run a soak test that actually finds things
A good soak test is not a stress test left running. It's a realistic load, representative traffic mix, realistic session lengths, the rare error paths actually exercised, held steady long enough for slow curves to develop, with resource metrics sampled the whole time. The win comes from watching trends, not pass/fail at the end.
# Soak harness: hold realistic load for hours, watch the slow curves
import time
async def soak(duration_hours=12, target_rps=50, sample_every_s=60):
end = time.time() + duration_hours * 3600
series = []
start_load(rps=target_rps, traffic_mix="production_realistic")
while time.time() < end:
m = sample_metrics()
series.append({
"t": time.time(),
"rss_mb": m.process_rss_mb, # memory creep
"pool_free": m.db_pool_free, # connection leak
"gpu_free_mb": m.gpu_free_mb, # KV/GPU growth
"fd_open": m.open_file_descriptors, # FD leak
"p99_ms": m.p99_latency_ms, # slow degradation
"throughput": m.completed_per_s, # quiet decline
})
await asyncio.sleep(sample_every_s)
return series
def assert_no_drift(series):
first, last = series[:10], series[-10:]
def trend(key): return avg(last, key) - avg(first, key)
# The signal is the SLOPE over hours, not any single sample.
assert trend("rss_mb") < 200, "memory creeping upward over soak"
assert trend("pool_free") > -5, "connection pool draining over soak"
assert trend("gpu_free_mb")> -500, "GPU/KV memory growing over soak"
assert trend("fd_open") < 50, "file descriptors leaking over soak"
# Throughput should be flat; a downward slope at steady load = rot.
assert trend("throughput") > -2, "throughput degrading at steady load"
The discipline is in assert_no_drift: you're not checking whether any single sample is healthy, you're checking whether the slope over hours is flat. A leak is a positive slope on usage and a negative slope on free resources. At steady offered load, every one of those curves should be a flat line. Any persistent drift is a leak you just caught in a test instead of a postmortem.
What to watch, and what good looks like
- Process RSS / heap, flat. A steady climb is a memory leak; a sawtooth is a leak plus OOM-restarts masking it.
- Connection pool free count, flat and well above zero. A slow decline is a connection leak heading for a hang.
- GPU free memory / effective batch size, flat. Decline means KV-cache growth or fragmentation eroding capacity.
- File descriptors / sockets, flat. A climb is a descriptor leak heading for "too many open files."
- p99 latency and throughput at constant load, flat. If latency rises or throughput falls while offered load is unchanged, something is rotting underneath.
The unifying rule: at constant offered load, every resource curve should be flat. Any sustained slope is a time-bomb with a fuse measured in hours.
The leaks specific to AI agents
Classic soak failures, memory, connections, descriptors, apply to any long-running service. But AI features, and agentic systems in particular, have their own family of slow leaks that only a realistic soak surfaces, because they depend on the content of traffic accumulating over time, not just its volume.
Conversation-state growth. If your service holds conversation history in memory keyed by session, and sessions aren't reliably evicted, you accumulate state for every conversation that ever started. Each is small; the population is unbounded. Over a soak, resident memory climbs not because of a code bug in the request path but because the set of live sessions never stops growing. A short test with a handful of sessions never reveals it.
Vector store and cache bloat. RAG systems that write back to a vector store or a semantic cache during operation grow their index continuously. As the index grows, retrieval latency creeps up and, past a point, recall quality drifts, a slow degradation that looks like nothing in a five-minute test and like a latency-and-quality cliff after hours of real write traffic.
Agent loop and tool-handle leaks. Agentic systems spawn sub-tasks, open tool connections, and hold handles to in-flight operations. A single missed cleanup on an agent's error path leaks a handle per failed run; over thousands of runs in a soak, those handles exhaust a pool or a descriptor table exactly the way a classic connection leak does, but they only appear when agents actually hit their error paths at volume, which a short happy-path test never triggers. The soak's value here is that it runs long enough, and varied enough, to exercise the rare branches where cleanup is forgotten.
When and how long to soak
Soak before every major launch and before any change to connection handling, caching, session management, or your serving framework, the subsystems where leaks hide. Match the duration to your real failure horizon: if your leaks take six hours to matter, a four-hour soak is a false pass. Start at 8-12 hours, and run a 24-hour soak before a high-stakes launch to cover a full daily traffic cycle. Run it on infrastructure that mirrors production limits, the same memory ceiling, pool size, and file-descriptor limit, because the cliff's location depends on those exact numbers. A soak on an over-provisioned box just moves the 3am page to a busier night.
The bottom line
Spike tests and soak tests find different bugs. Intensity finds your throughput ceiling; duration finds your leaks, and leaks are what page you days after a clean launch, because they detonate on uptime, not load magnitude. Memory creep, connection-pool exhaustion, KV-cache and GPU growth, and cache rot are all invisible for minutes and fatal for hours. Run realistic load for 8-24 hours on production-matched limits, sample every resource curve, and assert that the slope stays flat. A flat line over twelve hours is the only proof you have that the failure waiting at hour six isn't already in your code.
Pressure-Test Your AI Before Production Does
Hit fires browser-native, streaming-aware load at your LLM and API endpoints, TTFT, inter-token latency, tokens/sec, and cost per request, with no account and no script.
Try Hit Free →