CalibratedDecisions.

Post · Benchmarks & research

OpenRouter’s Ori Eval

Jev judged a 30-way labelling task faster than every model tested.

Open x.com ↗

How builders describe it

ran some early evals on jev using openrouter's ori eval

The decision Jev makes

Benchmark questions with known answers, to check accuracy and confidence.

Where it fits

Head-to-head tests, calibration studies and independent research into how well Jev decides, how fast, and at what cost. All 332 benchmarks & research projects →

Related projects