TL;DR
Cost is the failure mode teams forget to load-test. A feature that costs $5k/month in the demo can become a $60k invoice within 90 days as traffic, prompt size, and retries compound, and the bill arrives a month after the damage is done. CloudZero's State of AI Costs report put average monthly AI spend at $85,500 in 2025, up 36% year over year. The fix is to treat cost-per-request under load as a first-class metric: measure tokens consumed at realistic concurrency, model the worst case before launch, and gate on a per-request and per-session budget.
The invoice nobody simulated
Here is a real cost curve, the kind that shows up in incident retros. A team ships an assistant on a frontier model. Traffic is roughly 1.2 million messages a day, ~150 tokens each. The first full-month invoice lands near $15k. Month two, as usage grows and conversation histories lengthen, it's $35k. Month three touches $60k. Nothing "broke." There was no incident. The system worked exactly as designed, and the design was never tested for cost.
This pattern is common enough that practitioners describe a single misconfigured batch job or runaway prompt loop turning a $5,000 monthly bill into $50,000 overnight. And the macro trend is steep: per CloudZero's State of AI Costs, average monthly AI spend rose from $63,000 in 2024 to $85,500 in 2025. Cost is no longer a rounding error on the infra bill, it's a line item the CFO asks about by name.
The four multipliers that compound your bill
Cost under load is not linear with traffic. Four factors multiply together, and load is what makes them visible:
1. Output length variance
You can't predict response length. The same prompt template might return 40 tokens or 800 depending on the input. Capacity planning that assumes "average" output systematically underestimates the tail, and the tail is where the money goes.
2. Context growth
If you replay conversation history, every turn pays for all prior turns. Under load, with longer sessions, input tokens balloon. A chat at turn 20 can cost 5x what the same chat cost at turn 2, for the same user question.
3. Retries and agent loops
This is the overnight-$50k vector. A timeout triggers a retry; the retry hits the same overloaded model; an agent stuck in a tool-call loop can run up a $200 bill from a single session overnight. Multiply that by concurrency and you have a budget breach before anyone wakes up.
4. Cache assumptions that don't hold
Prompt caching looks great at low, repetitive load. Under varied production traffic the hit rate drops, and the cost projection you built on a 60% hit rate was fiction.
Measure cost-per-request, not just latency
The instrumentation is simple once you decide to do it: parse the streamed tokens, multiply by the model's price, and aggregate. The point is to do it under load, so you see the cost distribution your real traffic produces, including the expensive tail.
# Cost-per-request, measured under concurrency
def cost_of_request(input_tokens, output_tokens, price_in, price_out):
"""Prices are $ per 1M tokens (input and output billed separately)."""
return (input_tokens / 1e6) * price_in + (output_tokens / 1e6) * price_out
def project_under_load(samples, rps):
import numpy as np
costs = [cost_of_request(s.in_tok, s.out_tok, 2.50,10.00) for s in samples]
p50, p99 = np.percentile(costs, [50,99])
avg = np.mean(costs)
print(f"cost/req avg ${avg:.4f} · p50 ${p50:.4f} · p99 ${p99:.4f}")
print(f"at {rps} rps → ${avg*rps*3600:.0f}/hour · ${avg*rps*3600*24*30:, .0f}/month")
# The p99 is your worst-case-tail exposure, model it explicitly
print(f"tail month (p99 every req) → ${p99*rps*3600*24*30:, .0f}")
That last line is the one that prevents incidents. If your p99 cost-per-request, sustained, would blow the budget, you have a concentration risk: a traffic mix shift toward long outputs or deep context can put you there without any "failure" at all.
Gate the budget the way you gate latency
Prevention is mostly about hard limits enforced before the expensive operation runs, plus alerts that fire on the leading indicator rather than the monthly invoice:
class SessionBudget:
"""Catch runaway loops before they cost real money."""
def __init__(self, hard_cap_usd=1.00):
self.spent = 0.0
self.cap = hard_cap_usd
def charge(self, input_tokens, output_tokens, price_in, price_out):
cost = cost_of_request(input_tokens, output_tokens, price_in, price_out)
self.spent += cost
if self.spent > self.cap:
# A single session crossing $1 almost always means a stuck loop
raise BudgetExceeded(f"session at ${self.spent:.2f}, likely a loop")
return cost
The widely-recommended guardrails line up with this: set alerts at 80% of expected spend, cap per-day maximums, enforce hard token limits, compact conversation history before the context window fills, and run a pre-execution budget check before kicking off anything expensive. Each of these is testable under load, and should be tested under load, because a guardrail you've never exercised at concurrency is a guess.
The bottom line
AI cost is a load-dependent quantity, and load-dependent quantities have to be load-tested. Measure cost-per-request and cost-per-conversation under realistic concurrency. Model the p99 tail, not just the average. Put a hard per-session cap in front of every agent loop. And alert on the leading indicator, a single session crossing a dollar, instead of the lagging one, which is a five-figure invoice you can no longer do anything about.
Pressure-Test Your AI Before Production Does
Hit fires browser-native, streaming-aware load at your LLM and API endpoints, TTFT, inter-token latency, tokens/sec, and cost per request, with no account and no script.
Try Hit Free →