TL;DR
You have unit tests gating correctness and you’d never merge a PR that breaks them. Yet a change that doubles your time-to-first-token or halves your generation speed sails straight through, because nothing in CI measures latency. So latency regressions ship the way they always ship: invisibly, discovered weeks later as “it feels slow” tickets. The fix is a load gate, a CI step that fires real concurrent traffic at a staging endpoint, measures TTFT and tokens-per-second, and fails the build if either regresses past budget. k6 thresholds make this mechanical: a breached threshold returns a non-zero exit code, which fails the pipeline. This post shows how to wire one that understands streaming AI, not just HTTP.
The regression that no test could see
Walk the timeline of a typical AI latency regression. A developer adds three retrieved chunks to the RAG prompt “to improve quality.” Every unit test passes, the output is still correct, the JSON still validates, the integration test still returns 200. The PR merges. In production, prefill now reads 40% more tokens on every request, TTFT climbs from 600ms to 1.4s, and over the next two weeks session abandonment creeps up. Eventually someone correlates the dip to the deploy, reverts, and writes a post-mortem. The regression was a one-line change that was detectable in seconds, if anything in the pipeline had been measuring latency. Nothing was.
This is the gap. Correctness has a gate; performance doesn’t. The Consortium for Information & Software Quality’s work on the cost of poor software quality reinforces the shift-left axiom that gives this its economics: a defect caught in CI costs a fraction of the same defect caught in production, where it now also costs you a post-mortem, an incident, and the users who left while it was live. A latency regression is a defect. Gate it like one.
Why generic HTTP gates miss AI regressions
The instinct is to reach for k6 and gate on http_req_duration. That’s the right tool and the wrong metric. For a streaming completion, request duration is first-byte-to-last-byte, it conflates the prefill (TTFT) and the generation phase into one number that maps to nothing the user feels, as we argue in time to first token is your real SLA. A change that doubles TTFT but generates faster can leave http_req_duration unchanged while the product gets measurably worse. The gate has to understand streaming: it must time the first chunk separately from the rest.
k6 supports this through custom metrics. You define your own Trend metrics for TTFT and TPS, populate them by reading the streamed response, and set thresholds on those, not on the built-in duration. The Grafana guidance on performance testing in CI/CD recommends running fast smoke-load checks on every commit and heavier load tests before release; the streaming-aware custom metrics work in both.
A streaming-aware k6 load gate
Here’s a k6 script that measures the metrics that matter for AI and fails the build when they breach budget. The key moves: custom Trend metrics for TTFT and TPS, a streaming read to capture first-token timing, and thresholds that map directly to product SLOs.
import http from 'k6/http';
import { Trend, Rate } from 'k6/metrics';
const ttft = new Trend('ttft_ms', true);
const tps = new Trend('tokens_per_sec', true);
const slo_miss = new Rate('slo_miss');
export const options = {
scenarios: {
load: { executor: 'constant-vus', vus: 32, duration: '90s' },
},
thresholds: {
'ttft_ms': ['p(95)<500'], // TTFT p95 under 500ms or FAIL
'tokens_per_sec': ['p(50)>25'], // median generation > 25 tok/s or FAIL
'slo_miss': ['rate<0.05'], // < 5% of requests miss SLO
'http_req_failed':['rate<0.01'],
},
};
export default function () {
const t0 = Date.now();
const res = http.post(__ENV.TARGET_URL, JSON.stringify(payload), {
headers: { Authorization: `Bearer ${__ENV.API_KEY}`,
'Content-Type': 'application/json' },
responseType: 'text',
});
// parse SSE: find first `data:` delta time and total token count
const firstAt = res.timings.waiting; // approx first-byte for streamed body
const tokens = (res.body.match(/data:/g) || []).length;
const totalMs = Date.now() - t0;
ttft.add(firstAt);
tps.add(tokens / (totalMs / 1000));
slo_miss.add(firstAt > 500 || tokens / (totalMs / 1000) < 25);
}
Because a breached k6 threshold returns a non-zero exit code, dropping this into a pipeline step is all it takes to block the deploy:
# .github/workflows/load-gate.yml
- name: AI load gate
env:
TARGET_URL: ${{ secrets.STAGING_LLM_URL }}
API_KEY: ${{ secrets.STAGING_API_KEY }}
run: |
k6 run load-gate.js
# exit code is non-zero if any threshold fails -> deploy blocked
The budgets question: absolute vs. relative
The hard part of a load gate isn’t the mechanism, it’s choosing thresholds that catch real regressions without flapping. Two strategies, used together:
Absolute budgets (the floor)
Hard product SLOs that must never be crossed regardless of history: TTFT p95 < 500ms, TPS p50 > 25. These protect the user experience directly. They’re easy to set and easy to defend, but they only catch regressions that cross the line, a degradation from 200ms to 480ms passes an absolute 500ms gate while doubling your latency.
Relative budgets (the trend)
Catch the creep that absolute gates miss by comparing this run to a stored baseline and failing on a percentage delta. This is what catches the “200ms became 480ms” regression before it becomes “480ms became 900ms”:
import json, sys
with open("baseline.json") as f: base = json.load(f)
cur = json.loads(sys.argv[1]) # this run's summary
REGRESSION_PCT = 0.15 # fail on >15% worse than baseline
for metric in ["ttft_p95", "tps_p50"]:
b, c = base[metric], cur[metric]
worse = (c - b) / b if "ttft" in metric else (b - c) / b
if worse > REGRESSION_PCT:
print(f"REGRESSION {metric}: {b:.0f} -> {c:.0f} ({worse*100:+.0f}%)")
sys.exit(1) # fail the build
print("no regression vs baseline")
Gate on cost, not just latency
For AI, there’s a third axis no traditional performance gate has ever needed: tokens spent per request. A prompt change can be latency-neutral and quietly triple your token consumption, and therefore your bill. We’ve seen this turn into real money in the $50,000 token bill. The load gate is the natural place to catch it, because you’re already counting tokens to compute TPS:
# Extend the gate: fail on cost regression too
assert avg_input_tokens <= base["input_tokens"] * 1.15, "prompt bloat"
assert avg_output_tokens <= base["output_tokens"] * 1.15, "verbosity creep"
# a PR that adds 40% more RAG context to the prompt now turns the build red
Where to run it: the staging-vs-real-endpoint problem
A load gate is only as trustworthy as the thing it points at, and this is where AI gates get harder than ordinary HTTP gates. Three options, each with a real failure mode. A mocked endpoint is fast and free but tests nothing real, it can’t catch a prompt change that triples prefill, because there’s no model behind it. A shared staging environment is realistic but noisy: another team’s load test running concurrently inflates your TTFT and turns your gate red for reasons that have nothing to do with your PR, which is exactly how gates lose trust. The real provider endpoint (OpenAI, Anthropic, your own vLLM) gives honest numbers but costs tokens on every run and is subject to the provider’s own rate limits and latency variance.
The pragmatic answer is to match the target to the tier. PR-time smoke gates can hit a dedicated, isolated staging deployment sized to behave like production, so the run is cheap and the percentiles are stable. Pre-release gates should hit the real endpoint with the real model, because the regressions that matter most (a model version bump that slowed the cold path, a provider-side change) are invisible against a mock. Whatever you choose, pin the model version and the prompt set used in the gate, or you’ll chase ghosts: a “regression” that was really the provider silently rotating you to a different model build, a failure mode we cover in silent model updates are breaking your AI.
What to run, and how often
Don’t run a 30-minute soak on every commit, you’ll grind the pipeline to a halt and the team will resent the gate. Tier it: a 60-90 second smoke-load with tight thresholds on every PR to catch gross regressions fast, and a heavier, longer load run nightly or pre-release to catch tail behavior and slow leaks. Reserve full soak testing, covered in soak testing AI endpoints, for the scheduled lane. The PR gate’s job is to be fast and decisive, not exhaustive.
The friction that kills load gates in practice is setup: standing up a k6 environment, managing auth and secrets, keeping scripts current with a moving API. That friction is exactly why the test often never gets written. Driving the same TTFT/TPS measurement from the browser against your real, authenticated staging endpoint, no script to maintain, no proxy, no token juggling, removes the setup tax, which is the wedge Hit is built around. A gate that’s trivial to create is a gate that actually ships.
The bottom line
Latency and cost regressions ship undetected because CI measures correctness and nothing else. Close the gap with a load gate that understands streaming: measure TTFT and TPS (and tokens spent) under real concurrency, enforce both absolute SLO budgets and relative trend budgets, and let a breach return a non-zero exit code that blocks the deploy. Tier it so PRs get a fast smoke-load and releases get the heavy run. The teams that gate on latency stop writing the “we shipped it slow” post-mortem, because the PR went red before it ever merged.
Pressure-Test Your AI Before Production Does
Hit fires browser-native, streaming-aware load at your LLM and API endpoints, TTFT, inter-token latency, tokens/sec, and cost per request, with no account and no script.
Try Hit Free →