CalibratedDecisions.

Repo · Benchmarks & research

Agent Failure Benchmark

Test attribution of failures in agent traces.

Open github.com ↗

From the repository

Benchmarking Jev (Typesafe.ai) against a strong LLM on the Who&When Pro agent-failure-attribution benchmark (text subset).

How builders describe it

Benchmarking Jev (Typesafe.ai) against a strong LLM on the Who&When Pro agent-failure-attribution benchmark (text subset).
Offline-reproducible harness that scores TypeSafe Jev on the text subset of Who&When Pro (agent-failure attribution: who / when / what) and compares against published LLM baselines.

The decision Jev makes

Benchmark questions with known answers, to check accuracy and confidence.

Where it fits

Head-to-head tests, calibration studies and independent research into how well Jev decides, how fast, and at what cost. All 332 benchmarks & research projects →

Related projects