CalibratedDecisions.

Repo · Benchmarks & research

tenbin

Split a judgment into Choice, Score and Noul, lint it, then measure it.

Open github.com ↗

How builders describe it

MCP server and agent skill for the TypeSafe AI System One API (Jev): decompose a judgment into Choice / Score / Noul questions, lint them, measure on labelled data, and put calibrated thresholds in code.

The decision Jev makes

Benchmark questions with known answers, to check accuracy and confidence.

Where it fits

Head-to-head tests, calibration studies and independent research into how well Jev decides, how fast, and at what cost. All 332 benchmarks & research projects →

Related projects