
Post on X
- Views
- 238,930
- Likes
- 2,064
- Reposts
- 170
Read from X on . Counts change daily.
Reported by the author
WindTunnel is a benchmark for browser agents: 49 tasks spread across 8 real websites, run through 21 different agent configurations. It measures which combination of models and tooling can actually complete tasks like searching, filling forms, and checking out, and at what cost.
Idan Levin’s team ran Jev as the decision-maker inside Browser Use’s Ultrafast harness. Jev is a classification model, so on its own it can only pick from a fixed set of options, not generate free text. Each step it sees the task, the available actions, and what happened so far, then picks the next action. When an action needs a text argument, such as a search query, Mercury 2.5, a small fast language model, fills that in. Without any special interface, this Jev-plus-Mercury pairing solved 25 of 49 tasks by choosing among the page’s own buttons and fields. Levin notes that figure reflects their modified harness on this benchmark, not a general limit on Jev or Browser Use.
The bigger jump came from WebMCP, a standard that lets a website expose its own actions as discrete tools instead of making an agent parse buttons and forms. With WebMCP tools available, Jev’s solve rate nearly doubled to 49 of 49, and model cost dropped by 18%. Levin’s read is that WebMCP collapses a multi-click sequence into one tool call, which suits a model that picks from options rather than reasons through open-ended navigation. With WebMCP in place, Jev plus Mercury solved all 49 tasks at roughly 112 times lower model cost than GPT-6 Astra running computer use with code execution, and 245 times lower than Astra using screenshot-based computer use.
The benchmark and full results are published at webmcp.com, and the methodology is open for others to reproduce or extend.

