Repo · Benchmarks & research
JevPokerBench
A Texas Hold'em benchmark for decision models, with leaderboards.
Open github.com ↗How builders describe it
Texas Hold'em benchmark and playground for decision models: cash and SNG leaderboards, hand replay/advisor, human rooms against models, and BYOK custom agents. Official TypeSafe Jev is a first-class provider alongside local System One-style routes.
The decision Jev makes
Benchmark questions with known answers, to check accuracy and confidence.
Where it fits
Head-to-head tests, calibration studies and independent research into how well Jev decides, how fast, and at what cost. All 332 benchmarks & research projects →