Repo · Benchmarks & research
Agent Failure Benchmark
Test attribution of failures in agent traces.
Open github.com ↗From the repository
Benchmarking Jev (Typesafe.ai) against a strong LLM on the Who&When Pro agent-failure-attribution benchmark (text subset).
How builders describe it
Benchmarking Jev (Typesafe.ai) against a strong LLM on the Who&When Pro agent-failure-attribution benchmark (text subset).
Offline-reproducible harness that scores TypeSafe Jev on the text subset of Who&When Pro (agent-failure attribution: who / when / what) and compares against published LLM baselines.
The decision Jev makes
Benchmark questions with known answers, to check accuracy and confidence.
Where it fits
Head-to-head tests, calibration studies and independent research into how well Jev decides, how fast, and at what cost. All 332 benchmarks & research projects →