Reproducible by a third party from this page alone.
Every edition publishes its domain list, task text, agents with versions, dates, judging rules and confidence intervals.
1. Site selection
Top sites by traffic within each vertical, drawn from public rankings and published with the edition. Sites whose robots.txt disallows our agent, and sites that opted out, are excluded automatically.
2. Tasks
Only logged-out, read-only discovery and qualification tasks: find something that meets stated constraints; answer a policy question yes or no with the quoted sentence; get an estimate without entering personal data; find accessibility information. No account creation, no form submission, no cart or checkout on non-customer sites.
3. Agents
Several vendors’ agents (for example Claude computer use, Gemini computer use, Browser Use, Stagehand, Skyvern) with exact model versions and run dates. Consumer-agent classes appear only as labelled behavioural approximations.
4. Identity and etiquette
Every request carries alt.qa’s Web Bot Auth signature and a documented user agent. One session per host at a time, with pauses between tasks. A CAPTCHA or bot wall is recorded as the outcome, never bypassed (Anthropic’s usage policy and our own rules forbid it).
5. Judging
Each task has written success criteria checked against the final page state. Two reviewers resolve disagreements; a sample is re-run to estimate variance. Results are reported with 95% confidence intervals, not single scores.
6. Legal basis
The Ninth Circuit’s Amazon v. Perplexity decision (4 Aug 2026) concerned agents a user runs on their own machine. A cloud benchmark is a different fact pattern, which is why ours stays read-only, logged-out, robots-respecting and honestly identified, with opt-out.
7. Limitations
Results describe specific agents on specific dates. Agents change quickly; so do sites. A site that performs poorly on one edition may be fixed the next week. We never publish a single 0-100 score for a site.