CalibratedDecisions.

App · Benchmarks & research

Jevals.com

Hosted independent benchmark boards: TypeSafe Jev (typesafe-ai/jev via Vercel AI Gateway) vs six LLMs on PubMedQA (Noul), Banking77 (Choice), and…

Open jevals.com ↗

How builders describe it

Hosted independent benchmark boards: TypeSafe Jev (typesafe-ai/jev via Vercel AI Gateway) vs six LLMs on PubMedQA (Noul), Banking77 (Choice), and HelpSteer2 (Score)—accuracy, calibration, coverage@threshold, cost, and latency. Distinct from the local jevals workbench (dayhaysoos).

The decision Jev makes

Benchmark questions with known answers, to check accuracy and confidence.

Where it fits

Head-to-head tests, calibration studies and independent research into how well Jev decides, how fast, and at what cost. All 332 benchmarks & research projects →

Related projects