Repo · Benchmarks & research
chinese-workflow-decision-bench
Reusable Feishu-style workplace message classification benchmark: 64 frozen synthetic Chinese scenarios, Choice and four-Noul workflows, published…
Open github.com ↗How builders describe it
Reusable Feishu-style workplace message classification benchmark: 64 frozen synthetic Chinese scenarios, Choice and four-Noul workflows, published Jev vs Laya results, and pluggable classifier adapters. Not affiliated with Feishu.
The decision Jev makes
Benchmark questions with known answers, to check accuracy and confidence.
Where it fits
Head-to-head tests, calibration studies and independent research into how well Jev decides, how fast, and at what cost. All 332 benchmarks & research projects →