CalibratedDecisions.

Repo · Benchmarks & research

typesafe-ai-playground

Community TypeSafe AI playground: 110 use cases, games, dilemmas and model challenges, with editable prompts, A/B comparisons and a mobile-friendly…

Open github.com ↗

How builders describe it

Community TypeSafe AI playground: 110 use cases, games, dilemmas and model challenges, with editable prompts, A/B comparisons and a mobile-friendly UI.
Built this with Astra while waiting for TypeSafe access: a playground with 110 examples, from trolley problems and hot-dog debates to reasoning tests, plus a mobile-friendly UI. Run it locally with the community API key from Discord and add your own experiments

The decision Jev makes

Benchmark questions with known answers, to check accuracy and confidence.

Where it fits

Head-to-head tests, calibration studies and independent research into how well Jev decides, how fast, and at what cost. All 332 benchmarks & research projects →

Related projects