App · Benchmarks & research
WebMCP browser benchmark
A reported 49-task benchmark combines Jev for tool selection, Mercury 2.5 for arguments and WebMCP for browser actions.
Open webmcp.co ↗How builders describe it
A reported 49-task benchmark combines Jev for tool selection, Mercury 2.5 for arguments and WebMCP for browser actions. The author reports 49/49 tasks solved versus 25/49 for their modified browser-control baseline; results are specific to this setup. Via DAIR.AI Jev Field Notes (research…
The decision Jev makes
Benchmark questions with known answers, to check accuracy and confidence.
Where it fits
Head-to-head tests, calibration studies and independent research into how well Jev decides, how fast, and at what cost. All 332 benchmarks & research projects →