CalibratedDecisions.

Repo · Benchmarks & research

JevPokerBench

A Texas Hold'em benchmark for decision models, with leaderboards.

Open github.com ↗

How builders describe it

Texas Hold'em benchmark and playground for decision models: cash and SNG leaderboards, hand replay/advisor, human rooms against models, and BYOK custom agents. Official TypeSafe Jev is a first-class provider alongside local System One-style routes.

The decision Jev makes

Benchmark questions with known answers, to check accuracy and confidence.

Where it fits

Head-to-head tests, calibration studies and independent research into how well Jev decides, how fast, and at what cost. All 332 benchmarks & research projects →

Related projects