CalibratedDecisions.

Repo · Benchmarks & research

jev-orderby-bench

Independent measurement of whether ORDER BY over a Jev probability is defensible, pre-registering gates on pairwise inversion, Score ordinality,…

Open github.com ↗

From the repository

Does ORDER BY over a Jev probability put rows in a defensible order? Independent ranking, calibration and invariant measurements of TypeSafe AI's Jev: passes six pre-registered gates on 360 labeled rows, fails four of six on graded product relevance.

benchmarkcalibrationduckdbinformation-retrievaljevllm-evaluationpythonrankingreproducible-researchtypesafe

How builders describe it

Independent measurement of whether ORDER BY over a Jev probability is defensible, pre-registering gates on pairwise inversion, Score ordinality, calibration, and negation and paraphrase invariants.

The decision Jev makes

Benchmark questions with known answers, to check accuracy and confidence.

Where it fits

Head-to-head tests, calibration studies and independent research into how well Jev decides, how fast, and at what cost. All 355 benchmarks & research projects →

Related projects