TL;DR
Your self-hosted LLM has a latency number you never see in steady-state benchmarks: the cold start. Spinning up a fresh GPU replica means provisioning the node, pulling a multi-gigabyte container, loading tens of gigabytes of weights into VRAM, and warming the runtime, a sequence that routinely takes 30 to 60 seconds, and sometimes minutes. A 130GB Llama-2-70B checkpoint takes around 26 seconds just to download from S3 at 5GB/s, plus another ~84 seconds to load across 8 GPUs. Every request that lands during that window either queues behind the cold replica or hits an overloaded warm one. Scale-to-zero saves money and then quietly blows your p99 SLA the first time real traffic spikes. If you have not load-tested the scale-up event itself, you have not tested the thing most likely to break your SLA.
The SLA you pass every day until the day you spike
Here is the trap. You benchmark your self-hosted model on a warm, steady pool of replicas. TTFT is great, throughput is great, p99 is comfortably inside budget. You set an autoscaler to add capacity when load rises and, to save money on expensive GPUs, to scale unused replicas down, maybe all the way to zero off-hours. You ship. For weeks, everything is green.
Then a Monday launch, a viral post, or a batch job sends a traffic spike. The autoscaler reacts, but a GPU replica is not a stateless web pod that starts in 200ms. It has to come up cold, and while it does, every arriving request piles onto the few warm replicas, which are now saturated, or waits for the cold one. In real systems, as a 2026 teardown of LLM cold starts documents, a cold instance can take more than 40 seconds before it produces its first token, even though steady-state generation is only ~30ms per token after that. Your p99 doesn't degrade; it detonates. The exact moment you most needed capacity is the moment your latency was worst.
Anatomy of an LLM cold start
A cold start is not one delay; it is a chain, and each link is measured in seconds:
- Node provisioning. If no GPU node is sitting idle in the pool, the cluster autoscaler asks the cloud for one. GPU instances are scarce and slow to allocate, this alone can be 30 seconds to several minutes, and during capacity crunches the request can fail outright.
- Image pull. Inference containers are huge, CUDA, the runtime, and dependencies routinely push images past 5-10GB. Pulling and unpacking that onto a fresh node is tens of seconds unless you have aggressive caching.
- Weight loading. The dominant cost. Weights must move from remote/object storage to local disk, then into GPU VRAM. Size scales with parameters and precision.
- Runtime warmup. CUDA graph capture, kernel autotuning, and JIT compilation on the first requests. The first few inferences after load are slower than steady-state until the runtime warms.
What makes this chain so dangerous is that the links do not overlap, they are strictly sequential. The node cannot pull the image until it exists, cannot load weights until the image is unpacked, and cannot warm the runtime until the weights are resident. So unlike steady-state latency, where pipelining and batching hide a lot of work, a cold start exposes every stage end to end. And because GPU nodes are expensive, most clusters keep the idle pool small, which means the node-provisioning link fires far more often than teams expect, every modest spike that exceeds the warm headroom triggers a from-scratch allocation.
The weight-loading link dominates, and the math is unforgiving. Storing parameters in fp16 costs 2 bytes each, so a 7B model is ~14GB, a 13B is ~26GB, and a 70B is ~140GB. A widely cited figure, repeated in analyses of cold-start latency in LLM inference, is that a 130GB Llama-2-70B checkpoint takes at least 26 seconds to download from S3 at 5GB/s, and loading it onto 8 GPUs adds another ~84 seconds. NVIDIA's own guidance on LLM inference fundamentals frames the memory footprint as roughly parameters × bytes_per_param plus KV-cache overhead, and you cannot serve a token until that footprint is resident in VRAM.
# Weight bytes that must reach VRAM before the first token
def weight_gb(params_billion, bytes_per_param):
return params_billion * 1e9 * bytes_per_param / 1e9
# 7B fp16 -> 14 GB
# 13B fp16 -> 26 GB
# 70B fp16 -> 140 GB (often sharded across multiple GPUs)
# 70B int4 -> 35 GB (quantization buys faster cold starts too)
# Load time floor, assuming sustained read bandwidth B (GB/s):
def load_seconds(gb, B): # B from object storage is often 1-5 GB/s
return gb / B
# 140 GB / 2 GB/s = 70s just to stream weights, before warmup
Scale-to-zero: the trade you have to test, not assume
Scale-to-zero is genuinely attractive. GPUs are the most expensive line item in an AI infrastructure bill, and paying for idle H100s overnight is painful. The same cold-start analysis above is blunt about the consequence: if your cold start exceeds your SLO, say 30s for an interactive app, you are forced to keep warm replicas running 24/7, which can 2-3x your GPU spend. That is the real trade-off, and it should be made with numbers, not vibes:
- Scale-to-zero, lowest cost, worst first-request latency. Acceptable for internal tools, batch jobs, and traffic with predictable lulls. Lethal for interactive, spiky, user-facing traffic with a tight SLA.
- Warm minimum pool, you keep N replicas always on. Higher baseline cost, but spikes hit warm capacity. The standard answer for production user-facing inference.
- Predictive / scheduled scaling, warm up ahead of known peaks (business hours, launches). Best of both when traffic is predictable.
Why this is a fast-moving 2026 problem
Cold starts are not a niche concern; they are now a first-class metric for the GPU-serving ecosystem, and the engineering effort going into them is enormous. NVIDIA's Run:ai Model Streamer reports cutting weight-load times to single-digit seconds by streaming weights concurrently rather than copying then loading. NetEase Games, in a widely reported case, cut LLM cold starts from 42 minutes to 30 seconds using distributed model loading on Kubernetes. The whole reason features like model snapshotting, memory restore, weight streaming, and pre-warmed pools have proliferated across serving platforms is that the cold-start tax on big models is large enough to dominate the user-facing SLA during exactly the events you cannot afford to be slow.
Measuring the cold start, not the steady state
The mistake is measuring latency only against a warm pool. You have to measure the scale-up event itself: force a cold replica, then fire the spike and watch what the first wave of users experiences. Tag each request with whether it hit a warm or cold path:
// Probe the cold-start penalty: hit the endpoint right after a scale-from-zero,
// then again warm, and compare the first-request TTFT distributions.
async function coldVsWarm(url, body, headers, n) {
async function oneTTFT() {
const t0 = performance.now();
const res = await fetch(url, { method: 'POST', headers, body: JSON.stringify(body) });
const reader = res.body.getReader();
let ttft = null;
while (true) {
const { done } = await reader.read();
if (ttft === null) ttft = performance.now() - t0;
if (done) break;
}
return ttft;
}
// First request lands on a cold replica (assumes scaled-to-zero/just-spiked):
const cold = await oneTTFT();
// Subsequent requests should hit the now-warm replica:
const warm = [];
for (let i = 0; i < n; i++) warm.push(await oneTTFT());
return { cold_ttft_ms: cold, warm_p50_ms: median(warm),
cold_penalty_x: cold / median(warm) };
}
import numpy as np
def coldstart_report(first_wave_ms, steady_ms, sla_ms):
fw99 = np.percentile(first_wave_ms, 99)
ss99 = np.percentile(steady_ms, 99)
print(f"first-wave p99 {fw99:.0f}ms vs steady p99 {ss99:.0f}ms")
print(f"cold-start penalty: {fw99/ss99:.1f}x")
breached = np.mean(np.array(first_wave_ms) > sla_ms) * 100
print(f"{breached:.1f}% of spike requests breach the {sla_ms}ms SLA")
print("PASS" if fw99 <= sla_ms else "FAIL, scale-up blows the SLA")
Gate the spike, not just the steady state
Make the scale-up event a release gate. Your CI should simulate a cold-pool spike and assert that the first wave of users stays within SLA:
assert first_wave_p99_ms < sla_ms, f"cold-start p99 breaches SLA: {first_wave_p99_ms}ms"
assert cold_penalty_x < 3, f"cold path too slow vs warm: {cold_penalty_x}x"
assert spike_breach_pct < 1.0, f"{spike_breach_pct}% of spike requests breached SLA"
Now a base-image change that bloats the pull, a model upgrade that doubles weight size, a switch from a warm floor to scale-to-zero, or an autoscaler tuned to react too late all turn into a red build, instead of a Monday-morning incident discovered by your users at the worst possible moment.
The bottom line
A self-hosted LLM has two SLAs, and you have probably only tested one. Steady-state latency is what your warm pool delivers on a calm afternoon. Cold-start latency, node provisioning plus a multi-gigabyte image pull plus tens of gigabytes of weights streamed into VRAM plus runtime warmup, is what your users feel during the spike you scaled up to handle, and it routinely runs 30 to 60 seconds or worse. Scale-to-zero makes the bill smaller and the tail worse. Decide the warm-floor trade with numbers, autoscale on queue depth rather than RPS, and load-test the scale-up event itself. The teams that gate on the spike stop being surprised by latency that detonates exactly when traffic peaks.
Pressure-Test Your AI Before Production Does
Hit fires browser-native, streaming-aware load at your LLM and API endpoints, TTFT, inter-token latency, tokens/sec, and cost per request, with no account and no script.
Try Hit Free →