CalibratedDecisions.

Post · Benchmarks & research

Jev as an LLM Judge

Yuchen Jin's take on Jev for LLM-as-a-judge work: 20-200x faster and 40-400x cheaper, with a screenshot of the model picking which candidate is AGI…

Open x.com ↗
Yuchen Jin @Yuchenj_UW

Jev has spoken.

It picked which model is AGI.

20–200x faster. 40–400x cheaper. This could make things like LLM-as-a-judge insanely fast and nearly free.

(I tried a bunch of prompts and still didn’t burn through $0.10.)

Image from the post

How builders describe it

Yuchen Jin's take on Jev for LLM-as-a-judge work: 20-200x faster and 40-400x cheaper, with a screenshot of the model picking which candidate is AGI after a prompt sweep under $0.10.

The decision Jev makes

Benchmark questions with known answers, to check accuracy and confidence.

Where it fits

Head-to-head tests, calibration studies and independent research into how well Jev decides, how fast, and at what cost. All 355 benchmarks & research projects →

Related projects