Repo · Benchmarks & research
jev-orderby-bench
Independent measurement of whether ORDER BY over a Jev probability is defensible, pre-registering gates on pairwise inversion, Score ordinality,…
Open github.com ↗From the repository
Does ORDER BY over a Jev probability put rows in a defensible order? Independent ranking, calibration and invariant measurements of TypeSafe AI's Jev: passes six pre-registered gates on 360 labeled rows, fails four of six on graded product relevance.
benchmarkcalibrationduckdbinformation-retrievaljevllm-evaluationpythonrankingreproducible-researchtypesafe
How builders describe it
Independent measurement of whether ORDER BY over a Jev probability is defensible, pre-registering gates on pairwise inversion, Score ordinality, calibration, and negation and paraphrase invariants.
The decision Jev makes
Benchmark questions with known answers, to check accuracy and confidence.
Where it fits
Head-to-head tests, calibration studies and independent research into how well Jev decides, how fast, and at what cost. All 355 benchmarks & research projects →