CalibratedDecisions.

Repo · Benchmarks & research

clarity-judge

Writing checked on separate named axes, each with its own verdict.

Open github.com ↗

How builders describe it

Multi-axis writing quality checker powered by TypeSafe AI's Jev model. Separate named checks, each with its own verdict and confidence.

The decision Jev makes

Benchmark questions with known answers, to check accuracy and confidence.

Where it fits

Head-to-head tests, calibration studies and independent research into how well Jev decides, how fast, and at what cost. All 332 benchmarks & research projects →

Related projects