Repo · Benchmarks & research
jev-browsecomp
Research harness comparing cheap TypeSafe Jev typed decisions (Choice/Noul act|review|abstain over stable doc ids) to full-LLM and Recursive Language…
Open github.com ↗
From the repository
Jev vs a Sonnet 5 RLM on BrowseComp-Plus (1K docs): a Jev screen in front of one Sonnet call matched RLM accuracy at 38% of the cost and 1/7 of the time.
How builders describe it
Research harness comparing cheap TypeSafe Jev typed decisions (Choice/Noul act|review|abstain over stable doc ids) to full-LLM and Recursive Language Model pipelines on BrowseComp-Plus–style must-cite QA.
The decision Jev makes
Benchmark questions with known answers, to check accuracy and confidence.
Where it fits
Head-to-head tests, calibration studies and independent research into how well Jev decides, how fast, and at what cost. All 355 benchmarks & research projects →
Related projects
Benchmarks & researchjev-ticket-triageReproducible support-ticket triage eval: TypeSafe Jev vs Together.ai LLMs on accuracy/cost/latency/confidence (WIP).GitHub · beese54TBenchmarks & researchtev1Together AI's open recipe and weights for a Jev-inspired decision model fine-tuned on Qwen3.5-4B, with the full data pipeline, training config, and…GitHub · ★ 141 · Together AI
Benchmarks & researchJev-StyleSmall calibrated decision models you run locally, with a systemone-compatible server, agent skills, Claude Code guard, and MCP tools—weights on…GitHub · ★ 2 · lawrence3699
Benchmarks & researchJevletFrom-scratch research reconstruction of a Jev-like System One decision model (typed Noul/Choice/Score → probabilities, no text generation) plus a…GitHub · ★ 1 · NAME0x0
Benchmarks & research50 financial jobs, one modelThe Fintech Builder runs Jev on fifty financial jobs and shows where it works and where it does not.YouTube · youtube.com/@TheFintechBuild
Benchmarks & researchIs Jev really better, faster and cheaper?Edward Donner puts Jev to the test against Luna.YouTube · youtube.com/@EdwardDonner