Repo · Benchmarks & research
jevlens
Evaluate, calibrate, replay, and monitor TypeSafe Jev decisions from labeled CSV/JSONL: store full answer distributions, report accuracy/Brier/F1,…
Open github.com ↗How builders describe it
Evaluate, calibrate, replay, and monitor TypeSafe Jev decisions from labeled CSV/JSONL: store full answer distributions, report accuracy/Brier/F1, suggest thresholds, and optionally open a Streamlit dashboard or fail CI on regression.
The decision Jev makes
Benchmark questions with known answers, to check accuracy and confidence.
Where it fits
Head-to-head tests, calibration studies and independent research into how well Jev decides, how fast, and at what cost. All 332 benchmarks & research projects →