The agents your customers use, honestly identified.
Different agents fail in different places. We run real vendor models in a real browser, under our own signed identity, and record the model version on every cell.
| Agent | Provider | Class | Availability |
|---|---|---|---|
| Claude computer use | Anthropic | Frontier computer-use model | At launch |
| Gemini computer use | Frontier computer-use model · human confirms at payment | At launch | |
| Browser Use | Open source + choice of LLM | DOM agent framework | At launch |
| OpenAI computer-using agent | OpenAI | Frontier computer-use model | Q1 2027 |
| Stagehand | Browserbase (open source) | DOM agent framework | Q1 2027 |
| Skyvern | Open source | Vision + DOM workflow agent | Q1 2027 |
| Human baseline (Eval Army) | Paid testers via eval.qa | Human verification | At launch |
ChatGPT agent, Comet, Muse and Buy for Me act for your customers, but we never copy their identity. Cloudflare delists bots that impersonate others, and it would misrepresent what we tested. Instead we approximate how a class behaves (pacing, when it asks for confirmation) using the models above, and label those columns everywhere.
Gemini computer use asks for human confirmation before sensitive actions: a person on your team confirms. No agent solves CAPTCHAs or bypasses bot walls; those are recorded as findings.
See where agents fail on your funnel.
Run one free cell on a property you own. No sales call, no card.