How often our findings are right, with the sample size.
Most AI testing tools never publish an accuracy rate. We will publish per-persona precision, human-confirmed, with n and a 95% confidence interval, and update it monthly.
We publish once we have at least 200 human-reviewed findings across design partners (target Q1 2027). Publishing a percentage from a handful of findings would break our own claims policy.
How the number will be calculated
Flagged
Every finding a persona × agent cell reports at Confirmed confidence (reproduced in ≥2 cells or by replay).
Human-confirmed
A paid tester from the modelled population re-runs the same step and records whether the barrier is real.
Precision
Confirmed ÷ flagged, per persona, with n, a 95% confidence interval, model versions and the date range.
| Persona | Flagged (n) | Human-confirmed | Precision | 95% CI | Status |
|---|---|---|---|---|---|
| first-timer-mobile | - | - | - | - | Collecting |
| careful-evaluator | - | - | - | - | Collecting |
| impatient-shopper | - | - | - | - | Collecting |
| elderly-low-dexterity | - | - | - | - | Collecting |
| non-native-b1 | - | - | - | - | Collecting |
| keyboard-only | - | - | - | - | Collecting |
| a11y-tree-probe | - | - | - | - | Collecting |
| cvd-deutan | - | - | - | - | Collecting |
| low-bandwidth-3g | - | - | - | - | Collecting |
| distracted-multitasker | - | - | - | - | Collecting |
| power-user | - | - | - | - | Collecting |
| returning-customer | - | - | - | - | Collecting |
Recall, “coverage” or any figure implying we find every issue. We also stop selling any persona whose precision stays below 50% for two consecutive months.