CalibratedDecisions.

Repo · Benchmarks & research

jev-browsecomp

Research harness comparing cheap TypeSafe Jev typed decisions (Choice/Noul act|review|abstain over stable doc ids) to full-LLM and Recursive Language…

Open github.com ↗

From the repository

Jev vs a Sonnet 5 RLM on BrowseComp-Plus (1K docs): a Jev screen in front of one Sonnet call matched RLM accuracy at 38% of the cost and 1/7 of the time.

How builders describe it

Research harness comparing cheap TypeSafe Jev typed decisions (Choice/Noul act|review|abstain over stable doc ids) to full-LLM and Recursive Language Model pipelines on BrowseComp-Plus–style must-cite QA.

The decision Jev makes

Benchmark questions with known answers, to check accuracy and confidence.

Where it fits

Head-to-head tests, calibration studies and independent research into how well Jev decides, how fast, and at what cost. All 355 benchmarks & research projects →

Related projects