TL;DR
Cost per token is an accounting fact; cost per resolved conversation is a business fact, and the two diverge under load, in the direction that hurts. Cost per token can fall 30% with a model upgrade while cost per resolution rises, because turns lengthen, the cache hit rate drops, and resolution rate falls all at once when traffic peaks. The whole AI-support category now prices on this unit: Intercom Fin charges $0.99 per resolution, Salesforce Agentforce around $2, and only the resolved-conversation number tells you whether scaling makes money or just makes you busy. With human tickets at $5-$15 each and a deflected-but-unresolved ticket costing you the AI run and the human handoff, the metric your CFO actually wants is cost per resolved conversation, measured at every load level.
The slide that cheered while the business sank
The engineering team reported cost per token, because that is the number the provider gives you and the number that falls every few months when a new model lands. The slide showed it dropping quarter over quarter. Progress.
The CFO asked a different question, the way CFOs do: "What does it cost us to actually resolve one customer's problem?" Silence. Nobody had that number. And when someone built it, it told a story the per-token slide had completely hidden, cost per resolved conversation had gone up, even as cost per token fell, and it spiked hardest during the busy periods when the most conversations were happening. The unit economics were inverted, and the inversion got worse with scale.
Why per-token is the wrong denominator
Cost per token is an input price; it tells you what a raw material costs. No business runs on input prices, businesses run on the cost to deliver a unit of value. For a conversational AI feature, that unit is a conversation that accomplishes something: a question answered, a ticket deflected, a task completed. The honest denominator is the resolved conversation, because an unresolved one cost you tokens and delivered nothing, worse than nothing, since it often dumps the customer onto a human agent anyway, so you paid twice.
The gap between the two metrics is large and moves the wrong way. Cost per token can fall 30% with a model upgrade while cost per resolved conversation rises, because resolution rate dropped, conversations got longer, retries multiplied, or the cheaper model needed more turns to reach the same answer. The per-token slide is technically true and strategically blind, the same category error as measuring load in requests instead of tokens, which we dissect in Token-Aware Rate Limiting. The shift toward outcome-based pricing has made this denominator the one the whole category now argues about (eesel on the real cost of AI customer service).
Building the number
Cost per resolved conversation is a stack of factors, and every one moves under load:
cost_per_resolved =
( turns_per_conversation
* tokens_per_turn
* effective_cost_per_token )
/ resolution_rate
Read it factor by factor. Turns per conversation climbs when the model struggles and users rephrase. Tokens per turn climbs as history accumulates, turn ten re-sends the whole transcript, so late turns cost far more than early ones. Effective cost per token is not the list price; it is the blended rate after cache hits and misses, and as we argued in Your Prompt Cache Hit Rate Collapses Exactly When You Need It, that blended rate rises under load just as volume peaks. And resolution rate sits in the denominator, so any drop inflates the whole metric: a feature resolving 80% has a cost-per-resolution 25% higher than its cost-per-conversation; one resolving 50% has double.
Why this scales your losses
If cost per resolved conversation is below the value of a resolution, scale is your friend, every new conversation adds margin. If it is above, scale is your enemy, and this is the trap that caught real companies through 2025. The category repriced around it: Intercom Fin charges $0.99 per resolution and Salesforce Agentforce around $2, billing only when the AI actually resolves rather than escalates (Intercom Fin pricing). A flat per-resolution price meeting a cost-per-resolution that rises under load is a structural margin leak that widens exactly when you are busiest and feeling most successful. You sign more customers, drive more volume, and lose more per unit while the top-line chart goes up and to the right. The per-token slide cheers you the whole way down.
Deflection ROI: the number support leaders defend
For support specifically, the 2025 executive question became deflection ROI: how much does each deflected ticket actually save, net of the AI cost to deflect it and the cost of the ones that escalate anyway? The honest formula subtracts the failures:
deflection_value =
( deflected_tickets * human_ticket_cost )
- ( total_ai_cost
+ escalated_tickets * human_ticket_cost )
Escalations are the killer term. With human tickets at $5-$15 each, a conversation that runs the AI for ten expensive turns and then hands off to a human has incurred the AI cost and the human cost, strictly worse than routing straight to a human. Industry benchmarks put typical AI deflection at 30-50%, with leaders past 70%, and place AI-handled ticket cost around $0.50-$1.05 against $8-$12 for a human-handled one (2026 AI customer-service cost benchmarks), but practitioners stress the gap between deflection (the customer did not reach a human) and genuine resolution (the problem was solved); a deflected-but-unresolved ticket returns as a costlier re-contact. If your escalation rate climbs under load, and it does, since stressed models resolve less and impatient users escalate more, deflection ROI can go negative during peaks even while it is positive on average. The average hides the hours that lose money. Measure it conditioned on traffic level, or you are averaging away the periods where the economics break.
The fix: instrument the conversation, not the call
Most observability is built around the individual model call, tokens in, tokens out, latency, cost. Useful, insufficient. Cost lives at the conversation level, so the unit of measurement must be the conversation: assign every model call a conversation ID, sum cost across the whole conversation, and, the hard part, label each conversation resolved or not.
Resolution is the measurement nobody wants to build because it is fuzzy, but proxies get you most of the way: did the conversation end without escalation? Did the user not return with the same question within a day? Did they take the action the conversation was meant to drive? Did a follow-up satisfaction signal come back positive? Combine a few proxies into a resolution label and you can compute cost per resolved conversation per cohort, per feature, and, critically, per traffic level.
Then plot it against load. The single most valuable chart in your AI cost program is cost per resolved conversation on the y-axis and concurrent load on the x-axis. If the line is flat, your unit economics survive scale. If it slopes up, you have found the exact rate at which growth destroys margin, and you can say how much load you can take before a profitable feature becomes a loss.
Load tests must report dollars per outcome
Here is the discipline almost no team has: a load test that reports cost per resolved conversation, not just latency and throughput. The standard load test answers "does it stay up at 10x?" The standard answer is "yes, p99 held." Everyone nods. Nobody checked whether the thing that stayed up was profitable while it did.
A cost-aware load test drives realistic conversation traffic, multi-turn, with the turn-count and message-length distributions of real users, including the impatient-under-latency behavior that lengthens conversations, and reports the full economic picture at each load level: cost per conversation, cost per resolved conversation, resolution rate, escalation rate, and effective cost per token after caching. The pass/fail criterion is not "it stayed up." It is "the unit economics held at 10x load." A system can ace every latency SLO and fail this badly, because latency tells you it survived and says nothing about whether surviving was worth it.
# Gate the release on unit economics, not just uptime
def assert_unit_economics(run, load_level, value_per_resolution):
resolved = sum(1 for c in run.conversations if c.resolved)
cost = sum(c.total_cost_usd for c in run.conversations)
cpr = cost / max(resolved, 1)
assert cpr < value_per_resolution, (
f"at {load_level}x load, cost/resolved ${cpr:.3f} "
f"exceeds value ${value_per_resolution:.3f}, scaling loses money"
)
Give the CFO the real number
Cost per token is the price of flour. Cost per resolved conversation is the price of selling a loaf of bread, accounting for the ones that burned, the ones the customer sent back, and the oven you ran all afternoon for a slow Tuesday. One of those numbers tells you whether you have a business, and it is not the one the provider prints on the invoice, nor the one that conveniently falls every quarter. Build cost per resolved conversation, measure it under load, gate your releases on it, and put it on the slide. It is the metric your CFO actually wants, and the one that tells you, before the quarter does, whether scaling this feature makes you money or just makes you busy.
Pressure-Test Your AI Before Production Does
Hit fires browser-native, streaming-aware load at your LLM and API endpoints, TTFT, inter-token latency, tokens/sec, and cost per request, with no account and no script.
Try Hit Free →