CalibratedDecisions.

App · Safety & guardrails

lesswrong.com

Hi everyone, I just wrote up a small experiment testing TypeSafe’s Jev as a trusted and cheaper monitor alternate for AI Control.

Open lesswrong.com ↗

How builders describe it

Hi everyone, I just wrote up a small experiment testing TypeSafe’s Jev as a trusted and cheaper monitor alternate for AI Control. The interesting bit: a single yes/no question gets ~0.97 AUROC and catches ~90% of backdoors at a 2% audit budget, at a tiny fraction of the cost of an LLM monitor. But…

The decision Jev makes

Allow, flag or escalate a message, tool call or action.

Where it fits

Screening messages, tool calls and outputs before they cause trouble. Jev gives a fast, calibrated verdict that can sit in front of every request or agent action. All 131 safety & guardrails projects →

Related projects