CalibratedDecisions.

Repo · Benchmarks & research

jev-guardrails

Research/eval repo that runs the same mock carrier support agent and 25 behavioural guardrail rules behind two backends—chat-model JSON judge vs…

Open github.com ↗

How builders describe it

Comparing LLM-as-judge vs TypeSafe Jev for agent guardrails: same rules, same agent, measured on cost, latency, calibration and coverage.
Research/eval repo that runs the same mock carrier support agent and 25 behavioural guardrail rules behind two backends—chat-model JSON judge vs TypeSafe Jev typed questions—so latency, cost, and coverage can be compared.

The decision Jev makes

Benchmark questions with known answers, to check accuracy and confidence.

Where it fits

Head-to-head tests, calibration studies and independent research into how well Jev decides, how fast, and at what cost. All 332 benchmarks & research projects →

Related projects