CalibratedDecisions.

Post · Benchmarks & research

fx auto mode benchmarked with Jev

Vercel's Pranit reports benchmarking the fx auto-mode safety classifier with Jev: about 5 to 18 times faster and more accurate than GPT-5.6 Luna.

Open x.com ↗
Pranit @fazxes

We benchmarked fx auto mode (safety) classifier with @typesafeai's Jev.

tl;dr: ~5-18x faster and more accurate than 𝚐𝚙𝚝-𝟻.𝟼-𝚕𝚞𝚗𝚊, our current top choice https://

Image from the post

The decision Jev makes

Benchmark questions with known answers, to check accuracy and confidence.

Where it fits

Head-to-head tests, calibration studies and independent research into how well Jev decides, how fast, and at what cost. All 355 benchmarks & research projects →

Related projects