Repo · Benchmarks & research
jevarena (chenmingtang830)
Open-source BYOK arena for Jev and other AI judges.
Open github.com ↗How builders describe it
Open-source BYOK arena for Jev and other AI judges. Find failures, compare quality, cost, and latency.
The decision Jev makes
Benchmark questions with known answers, to check accuracy and confidence.
Where it fits
Head-to-head tests, calibration studies and independent research into how well Jev decides, how fast, and at what cost. All 332 benchmarks & research projects →