Skip to content
System One

WindTunnel: the WebMCP benchmark

A 49-task, 8-site browser-agent benchmark. Jev paired with a small model for argument generation solved 100% of tasks at far lower cost than screenshot-based computer use, and adding WebMCP tool exposure roughly doubled Jev's own solve rate from 25 to 49 tasks.

Image from WindTunnel: the WebMCP benchmark
Image from x.com

Post on X

Views
238,930
Likes
2,064
Reposts
170

Read from X on . Counts change daily.

Reported by the author

Solve rate
25/49 to 49/49adding WebMCP nearly doubled the number of solved tasks, from 25/49 to 49/49, while reducing model cost by 18%x.com
Cost vs Astra
112x lowersolved 100% of the tasks at roughly 112× lower model cost than GPT-6 Astra using computer use with code executionx.com

WindTunnel is a benchmark for browser agents: 49 tasks spread across 8 real websites, run through 21 different agent configurations. It measures which combination of models and tooling can actually complete tasks like searching, filling forms, and checking out, and at what cost.

Idan Levin’s team ran Jev as the decision-maker inside Browser Use’s Ultrafast harness. Jev is a classification model, so on its own it can only pick from a fixed set of options, not generate free text. Each step it sees the task, the available actions, and what happened so far, then picks the next action. When an action needs a text argument, such as a search query, Mercury 2.5, a small fast language model, fills that in. Without any special interface, this Jev-plus-Mercury pairing solved 25 of 49 tasks by choosing among the page’s own buttons and fields. Levin notes that figure reflects their modified harness on this benchmark, not a general limit on Jev or Browser Use.

The bigger jump came from WebMCP, a standard that lets a website expose its own actions as discrete tools instead of making an agent parse buttons and forms. With WebMCP tools available, Jev’s solve rate nearly doubled to 49 of 49, and model cost dropped by 18%. Levin’s read is that WebMCP collapses a multi-click sequence into one tool call, which suits a model that picks from options rather than reasons through open-ended navigation. With WebMCP in place, Jev plus Mercury solved all 49 tasks at roughly 112 times lower model cost than GPT-6 Astra running computer use with code execution, and 245 times lower than Astra using screenshot-based computer use.

The benchmark and full results are published at webmcp.com, and the methodology is open for others to reproduce or extend.

Open on Web

More like this

Use cases this is tagged with