Alt Benchmark
What can today’s AI agents actually do on real websites?
A quarterly, open-methodology study: several agent vendors attempt everyday discovery and qualification tasks (find a product, read a policy, get an estimate) on the public pages of top sites. Replayable evidence for every result.
Logged-out, read-onlyrobots.txt respectedSigned identity
State of agent readiness, not a leaderboard of shame.
Open methodology
Tasks, domains, agents, versions, dates and judging rules published with every edition.
Learn more →Read-only, ethical
No accounts, forms or checkouts on non-customer sites. No CAPTCHA or bot-wall bypass. Opt out any time.
Learn more →Replayable evidence
Every result links to a recording and DOM snapshot, so anyone can check the call.
Queryable data
Filter by vertical, task and agent; download CSV; open-source scorer.
Learn more →