CalibratedDecisions.

Repo · Benchmarks & research

jev-align (Sutro)

Turns human labels into calibrated Jev classifiers with GEPA.

Open github.com ↗

How builders describe it

Active-learning CLI that uses human labels and GEPA to improve Jev decision definitions.
Experimental CLI (jeva / jev-align) that builds calibrated AI Functions with TypeSafe Jev evaluations plus GEPA optimization from human labels—binary, multiclass, multilabel, or score tasks.
Jev is cool, but like any foundation model it needs to be calibrated to your decision criteria. We launched jev-align: an open-source CLI to quickly teach Jev what good and bad looks like using GEPA. Try it out! Via DAIR.AI Jev Field Notes (research paraphrase, snapshot 2026-09-20).

The decision Jev makes

Benchmark questions with known answers, to check accuracy and confidence.

Where it fits

Head-to-head tests, calibration studies and independent research into how well Jev decides, how fast, and at what cost. All 332 benchmarks & research projects →

Related projects