CalibratedDecisions.

Repo · Benchmarks & research

jevals

Local evaluation workbench for TypeSafe Jev: author Noul/Choice/Score (and combined) questions with expected answers, run them, and compare saved…

Open github.com ↗

How builders describe it

Local evaluation workbench for TypeSafe Jev: author Noul/Choice/Score (and combined) questions with expected answers, run them, and compare saved results in the browser.

The decision Jev makes

Benchmark questions with known answers, to check accuracy and confidence.

Where it fits

Head-to-head tests, calibration studies and independent research into how well Jev decides, how fast, and at what cost. All 332 benchmarks & research projects →

Related projects