CalibratedDecisions.

Repo · Benchmarks & research

typesafe-offload-bench

Tested Jev with main "Players" (GPT/Claude) to work together.

Open github.com ↗

How builders describe it

Tested Jev with main "Players" (GPT/Claude) to work together. The idea was to "cut" the part of the job from generative agents. Speed is awesome (cost too), and some results are quite promising: Results/Details

The decision Jev makes

Benchmark questions with known answers, to check accuracy and confidence.

Where it fits

Head-to-head tests, calibration studies and independent research into how well Jev decides, how fast, and at what cost. All 332 benchmarks & research projects →

Related projects