TL;DR
You added a fallback model "for resilience" and checked the box. But you never load-tested the failover path, and it's slower, pricier, formats output differently, and has its own rate limits you'll blow through the instant your primary dies and dumps 100% of traffic onto it. Provider outages are real and routine: OpenAI's June 2025 incident ran 34 hours, its December 2024 outage 4+ hours, and one December month saw 22 OpenAI and 20 Anthropic incidents tracked. An untested failover is often worse than no failover, because it fails in a novel way at the exact moment you're already in an incident. The fix is to fail over on purpose, under load, before the outage does it for you.
The failover you've never actually run
Multi-provider redundancy is now standard advice, and the tooling makes it easy: gateways and routers like LiteLLM give you a unified, OpenAI-compatible interface across 100+ providers with automatic fallback and cooldown periods configured in a few lines. You set a fallback chain, deploy, and feel resilient.
Here's the problem. That failover path runs approximately never in normal operation. Your primary is healthy 99.9% of the time, so the fallback code executes a handful of times a month under trivial load, if ever. The first time it runs under real conditions is during an actual provider outage, when 100% of your traffic suddenly redirects to a model you've validated for correctness but never validated for capacity, latency, cost, or output shape at scale. You're discovering the failover's behavior live, mid-incident, with every user watching. And the routing layer itself has limits, LiteLLM, for instance, has been observed to start timing out with memory spikes around 2,000 requests per second, so the gateway you're trusting to save you can become the next bottleneck.
Why provider outages are a "when, " not an "if"
Single-provider dependency is a real risk because the providers really do go down. Per OpenAI's status history, its December 2024 outage took ChatGPT, the API, and Sora offline globally for 4+ hours on an upstream networking issue; its largest incident in June 2025 ran roughly 34 hours as routing nodes hit memory limits and failed readiness checks one by one until capacity collapsed. In a single December month, status trackers logged 22 OpenAI and 20 Anthropic incidents, several lasting 30 minutes or more. These aren't freak events, they're the operating baseline of a young, capacity-constrained industry.
And many incidents cascade from shared infrastructure underneath: when a major cloud region has a bad day, multiple "independent" AI providers degrade together because they run on the same compute. That correlation is the subtle trap. If your primary and your fallback both sit behind the same cloud region, DNS provider, or shared GPU capacity, a single underlying fault takes out both at once and your carefully configured failover never even gets to fire.
The four ways an untested failover bites
1. The secondary can't take the load
Your fallback provider has its own RPM and TPM ceilings tied to your spend tier on that platform, probably a low tier, because you barely use it. When your primary dies and you redirect full traffic to a fallback you've only lightly touched, you blow through its rate limits in seconds and turn a single-provider outage into a dual-provider outage. Your "resilience" measure made the incident worse.
2. The latency profile is different
Different models have different time-to-first-token and tokens-per-second. A fallback that's 40% slower will, under your existing timeout budget, start timing out requests the primary served comfortably. Failover succeeds at the routing layer and fails at the timeout layer, you switched providers and still cut users off.
3. The cost profile is different
Fallback models often cost more per token, or are more verbose, or both. A multi-hour failover onto a pricier model can quietly multiply your spend for the duration, a cost incident riding on top of the reliability incident, discovered only when the invoice arrives.
4. The output shape is different
This is the insidious one. Models format differently. If your application parses structured output, JSON, tool calls, a specific markdown shape, the fallback may produce valid responses your parser chokes on. Failover succeeds at the API level and fails at the parsing level, returning broken results that look like a success to your monitoring. You failed over flawlessly into a different bug.
# The capacity math everyone skips before relying on a fallback
def failover_load_shock(primary_rps, fallback_existing_rps, fallback_tpm_limit,
avg_tokens):
# When primary dies, ALL its traffic lands on the fallback at once
total_rps = primary_rps + fallback_existing_rps
offered_tpm = total_rps * 60 * avg_tokens
util = offered_tpm / fallback_tpm_limit
print(f"failover offers {util*100:.0f}% of fallback's TPM ceiling")
if util > 1.0:
print("=> fallback will 429 immediately. Dual-provider outage.")
failover_load_shock(primary_rps=120, fallback_existing_rps=5,
fallback_tpm_limit=400_000, avg_tokens=2_400)
# failover offers 450% of fallback's TPM ceiling
# => fallback will 429 immediately. Dual-provider outage.
Test the failover path under load, on purpose
The only honest validation is to trigger the failover deliberately while under production-like load, then assert that the secondary path holds on all four axes: capacity, latency, cost, and output shape.
# Failover game-day: kill the primary mid-load, grade the secondary
async def failover_load_test(target_rps=120):
pre = await drive_load(rps=target_rps, duration=60, route="primary")
# Hard-fail the primary mid-test; router should redirect to fallback
inject_provider_failure("primary", at_second=60, duration=120)
post = await drive_load(rps=target_rps, duration=120) # now on fallback
# 1. Capacity: did the fallback absorb full traffic without throttling?
assert post.rate_limit_errors == 0, "fallback 429'd under full failover load"
assert post.success_rate > 0.98, "fallback couldn't take the load"
# 2. Latency: did the slower fallback stay inside the timeout budget?
assert post.p99_ms < TIMEOUT_BUDGET_MS, "fallback latency blew the timeout"
# 3. Output shape: do fallback responses still parse?
assert post.parse_failure_rate < 0.01, "fallback output broke the parser"
# 4. Cost: quantify the spend multiplier you'd eat during a real outage
mult = post.cost_per_request / pre.cost_per_request
print(f"failover cost multiplier: {mult:.2f}x (budget for a multi-hour outage)")
That parse-failure assertion is the one most teams have never thought to write, and it's the one that catches the silent failover bug. Run this as a scheduled game-day, not a one-time check, your prompts, your fallback's models, and the providers' behavior all drift over time, and a failover that passed last quarter can rot.
One subtlety the game-day exposes that a static config review can't: the recovery path is as untested as the failover path, and often more dangerous. When your primary comes back, your router fails back, and if it fails back instantly and completely, you've just created a thundering herd against a primary that may have only partially recovered, knocking it down again and oscillating between providers. Good failback is gradual: ramp traffic back to the primary over tens of seconds while watching its error rate, the same way a circuit breaker's half-open state probes before fully trusting a dependency. Test the failback under load too, not just the failover. An outage that ends in a flapping router is still an outage, and a primary that gets re-overwhelmed the moment it recovers extends the incident you thought was over.
Semantic drift: when the fallback is "working" and still wrong
Capacity, latency, cost, and parsing are the measurable axes, but there's a fifth that's harder to catch and arguably more damaging: behavioral drift. Two models can both return valid, parseable JSON and still behave differently in ways that break your product. A fallback model might be more cautious and refuse prompts your primary happily answers; it might follow your system prompt's formatting instructions less faithfully; it might have a different tool-calling style, a different tendency to hedge, or a different sense of what counts as a complete answer.
None of this trips an error. The request succeeds, the response parses, the user gets an answer, just a worse or subtly different one. During a multi-hour failover, your entire user base is silently served a different product personality, and the only signal is a slow drip of "the AI got dumber today" support tickets that you won't connect to the outage until much later. This is the failover equivalent of context rot: a silent quality regression hiding behind a green status code.
The defense is to treat failover as an evaluation problem, not just a load problem. Maintain a small golden set of representative prompts with expected behaviors, and run it against your fallback model on the same schedule as your failover game-day. You're not checking that the fallback is identical, it won't be, you're checking that it clears a quality bar you've defined in advance, so you fail over into a known-acceptable degradation rather than an unknown one. A fallback that passes the load test but flunks the behavior test is a reliability win and a quality loss, and you want to know which trade you're making before the outage forces it.
Designing failover that actually helps
Beyond testing, a few design rules separate failover that saves you from failover that compounds the incident. Pre-provision real headroom on the fallback so it can absorb your full primary load, not a token amount. Choose fallbacks that are uncorrelated with your primary's failure modes, different cloud, different region, ideally different underlying infrastructure. Set timeouts that account for the slower fallback's latency, or use per-provider timeout budgets. Normalize output at the boundary so parsing doesn't depend on which model answered, LiteLLM's router docs cover fallback chains and cooldowns, but normalization is on you. And keep a circuit breaker in front of the primary so failover triggers fast and cleanly instead of after the pool is already exhausted.
The bottom line
A fallback model you've never load-tested is a liability dressed as a safeguard. When your primary goes down, and it will, the secondary inherits 100% of your traffic instantly, and that's when you discover it can't take the load, runs slower than your timeouts allow, costs more, or formats output your parser can't read. Provider outages are a when, not an if, and correlated infrastructure means your fallback might die alongside your primary. Fail over on purpose, under load, on a schedule, and grade the secondary on capacity, latency, cost, and output shape. Resilience you haven't exercised isn't resilience, it's an untested assumption with a config flag.
Pressure-Test Your AI Before Production Does
Hit fires browser-native, streaming-aware load at your LLM and API endpoints, TTFT, inter-token latency, tokens/sec, and cost per request, with no account and no script.
Try Hit Free →