17 · 332 projects
Jev for benchmarks & research
Head-to-head tests, calibration studies and independent research into how well Jev decides, how fast, and at what cost.
What Jev decides: Benchmark questions with known answers, to check accuracy and confidence.
Projects, newest first
Benchmarks & researchJev 1.13.0 against Hopper on Benchmark HeavenHopper wins three columns, Jev keeps Intelligence.Site · Benchmark Heaven
Benchmarks & researchValenTrain a Jev-like multimodal model by yourself.GitHub · ★ 51 · Liuziyu77
Benchmarks & researchjev-dimabsaTypeSafe Jev baseline for DimABSA (SemEval-2026 Task 3) subtask 1: zero-shot and 3-shot valence-arousal regression.GitHub · ★ 6 · ZhangYiqun018
Benchmarks & researchreflexbenchReflexBench — open benchmark and evaluation harness for System One models and typed decision engines.GitHub · ★ 4 · brida-ai
Benchmarks & researchjev-guardrailsResearch/eval repo that runs the same mock carrier support agent and 25 behavioural guardrail rules behind two backends—chat-model JSON judge vs…GitHub · ★ 3 · deepansh-saxena
Benchmarks & researchtypesafe-jev-calibrate-for-code-reviewAbout calibrating Jev for code reviews.GitHub · ★ 3 · Selmar
Benchmarks & researchjev-aitaBenchmark of TypeSafe's Jev against Sonnet 5, GPT-5 nano and local LLMs on 770 Reddit AITA verdicts: Brier scores, latency and cost.GitHub · ★ 2 · dchristopoulos
Benchmarks & researchjev-dllmShared Yes/No decision adaptation with diffusion language models.GitHub · ★ 2 · zhouzihao11
Benchmarks & researchjev-reviewsTypeSafe AI - Jev - first look and experiment with its API + Google reviews experiments.GitHub · ★ 2 · halfspin-qc
Benchmarks & researchjev_fsdJEV AI Model Demo with FSD.GitHub · ★ 2 · BrendanH18
Benchmarks & researchopen_system_oneCreating system one model just like jev.GitHub · ★ 2 · Readyaddy
Benchmarks & researchopenpave-jevPAVE skill for non-autoregressive decision models (Jev / Laya / TypeSafe): typed Choice, Score and Noul answers with calibrated probabilities.GitHub · ★ 2 · cnrai
Benchmarks & researchjev (zeke)Research notes and an interactive Cloudflare Worker demo for Jev, TypeSafe AI's structured decision model.GitHub · ★ 1 · zeke
Benchmarks & researchJev-Research-IndexJev Research Index is a bilingual catalogue of papers, software projects, interviews, public analyses, demonstrations, and social-media material…GitHub · ★ 1 · AgenticAPP-Web
Benchmarks & researchLetJevDecideAsk a yes or no question. Let Jev, or a fair coin, decide.GitHub · ★ 1 · choxos
Benchmarks & researchMEDJEVIndependent JEV-inspired System One-style typed decision research for clinical evidence, biomedical NLP, calibrated probabilities, and sleep signals.GitHub · ★ 1 · PAI-CUHK
Benchmarks & researchollayaRun open decision models locally — pull, run and serve Laya and other open Jev alternatives behind a TypeSafe-compatible API.GitHub · ★ 1 · ollaya-dev
Benchmarks & researchA faster, cheaper path for business modelsJev as an alternative for business logic.X post · MilkMendyABenchmarks & researchA local open-source version of JevRun a Jev-like decision model locally.X post · coccoinomane
Benchmarks & researchA number two open-weight Jev-like modelAdded to the Decision Index, strongest for its size.X post · apolinario (poli)
Benchmarks & researchA sub-100 ms decision engineHow the launch was described.X post · ASIHubHQ
Benchmarks & researchAn LLM distilling JevTypeSafe notes an LLM distilling Jev.X post · TypeSafe AI
Benchmarks & researchBounded choices, not proseJev returns a choice from the options it was given.X post · GeethanTech
Benchmarks & researchCLM-8B claims to be faster than JevAn open model that its release claims is up to 9x faster.X post · aiasssistantstoreCBenchmarks & researchCLM-8B, a new Jev competitorAn 8B decision model positioned against Jev.X post · plotarmordev
Benchmarks & researchConvex Decision EvalsLeaderboard that puts Jev and 14 LLMs through 108 four-option multiple-choice questions about the Convex backend platform, comparing accuracy,…GitHub · get-convex
Benchmarks & researchJev and Laya on a medical diagnostic testSixteen of sixteen against thirteen, on authored notes.X post · OpenMed
Benchmarks & researchJev and Laya on security queriesTwo decision models make different calls on the same security queries.X post · edwinfmesa
Benchmarks & researchJev Does Not Play DiceIndependent calibration evaluation of Jev on fair random draws with known probabilities and synthetic forecast documents, with recorded outputs and…GitHub · KantaHayashiAI
Benchmarks & researchJev for next best actionTesting Jev’s accuracy on predicting the next action.X post · dsmiley411
Benchmarks & researchJev in tool filtration testsPraised for filtering tools.X post · eboppu
Benchmarks & researchJev vs. Sol, on categorisationA text categorisation showdown.X post · Jasonalco
Benchmarks & researchLocal Laya beats JevA local decision model coming out ahead.X post · NeuralFrameLabs
Benchmarks & researchPredicting queries before submissionA community test with Jev and Laya.X post · edwinfmesa
Benchmarks & researchSeparating signal from noise with Jev and OpusA pattern can still be noise, so ask what would change your mind.X post · Dain
Benchmarks & researchsnbt-jev-benchJev on Indonesia's SNBT 2025 university entrance exam, whose papers are never released: 159 questions reconstructed from memory by volunteers, 156 of…GitHub · misaalyaTBenchmarks & researchThe System One paradoxOn System One models and where Jev goes next.X post · __tuan____
Benchmarks & researchTrying to beat Jev on a DGX SparkAn open System One model, with the data already collected.X post · Cuth
Benchmarks & researchxor joins the Jev Decision Index 0.2A new open-weight entry, eighth on text alone.X post · apolinario (poli)
Benchmarks & researchjevk5An open-weight alternative to Jev, Apache-2.0.GitHub · ★ 64 · allebee
Benchmarks & researchJevAnyCalibration-aware reinforcement learning for decision systems.GitHub · ★ 11 · weitianxin
Benchmarks & researchlejudge-jev-jepaNatural-language constraints for JEPA planning, judged by a decision model.GitHub · ★ 9 · AbdelStark
Benchmarks & researchKevAn open-source System One engine with calibrated probabilities.GitHub · ★ 7 · arjun988
Benchmarks & researchawesome-jev-robustnessTests, calibration audits and failure-mode studies of Jev.GitHub · ★ 5 · Yifan-Lan
Benchmarks & researchtypesafe-pocProva de Conceito do Jev, modelo da typesafe.ai.GitHub · ★ 5 · eminetto
Benchmarks & researchjev-benchmark (YidiDev)Rubric-Based Zero-Shot Classification Benchmark: Jev vs Claude Haiku 4.5 vs OpenJev on rubric-conditioned classification, chained decision execution,…GitHub · ★ 3 · YidiDev
Benchmarks & researchtinyjevA tiny jev-like model that answers Choice, Score and Noul questions in one forward pass and returns calibrated probabilities.GitHub · ★ 3 · ankit-aglawe
Benchmarks & researchjev-deepresearchDeep research crawl where Jev (a System One model) makes every per-page decision and an LLM only plans and writes.GitHub · ★ 2 · tashfeenahmed
Benchmarks & researchjev-x-kitOffline $0 decision layer for coding agents: Choice/Score/Noul primitives, BELKI confidence gatekeeper, ultra-planning, red-teaming, research and…GitHub · ★ 2 · Kadihx
Benchmarks & researchjevalyzerGrade the agent sessions already on your disk.GitHub · ★ 2 · killerz3
Benchmarks & researchcan-jev-playCan Jev infer whether a bet is worthwhile from its payout table, and do recent wins or losses sway that choice.GitHub · ★ 1 · carlaiau
Benchmarks & researchjevals-dataIndependent benchmark data for TypeSafe's Jev (System One model) vs LLMs: accuracy, calibration, cost.GitHub · ★ 1 · Jevals
Benchmarks & researchrerank-bench-jevProduction reranker benchmark: TypeSafe Jev vs Qwen3-Reranker-0.6B on BEIR SciFact and NFCorpus — nDCG@10, cost/query and p50/p95 latency.GitHub · ★ 1 · denser-org
Benchmarks & researchsystem-one (mpuig)An open-source System One decision model — Kahneman's term for the fast, automatic judgment faculty.GitHub · ★ 1 · mpuig
Benchmarks & researchyc-jev-benchCan Jev replace an LLM reranker?GitHub · ★ 1 · PPRAMANIK62
Benchmarks & researchA Jev-style local model for imagesTyped decisions, but for image inputs.X post · edplese
Benchmarks & researchAn eval suite for System One modelsBuilding the tests Jev-like models need.X post · morganlinton
Benchmarks & researchawesome-jev-surveyAn evidence survey of Jev and Jev-like typed decision models.GitHub · Eurekaleo
Benchmarks & researchBeating Jev with open modelsAn attempt to beat Jev’s accuracy, speed and cost with open models.GitHub · rob313
Benchmarks & researchDecisive judgment agents that don’t chatJev against Laya.X post · thelazyyogii
Benchmarks & researchJev against Llama and DeepSeek235 ms median latency and 82% accuracy on the same labelled inputs.X post · RamV2003
Benchmarks & researchJev catches the Opus 5.5 effort dropThe default effort moved from high to medium, and agents without settings lost reasoning.X post · 0xKiyoro
Benchmarks & researchJev on a movie recommenderTesting Jev on recommendations.X post · utkarshgsuv
Benchmarks & researchJev on Cloud RunAbout 47 s cold start and about 117 ms end to end on a cloud GPU.X post · _vmlops
Benchmarks & researchJev vs. Laya-MLXA head-to-head comparison.X post · irSodeh
Benchmarks & researchJev_StarA Jev and System One related repository.GitHub · sc2musa
Benchmarks & researchJevBenchA reproducible benchmark for typed decision models.Site · florianstandhar
Benchmarks & researchLaya compared to JevAn open-source model next to the official one.X post · i_darshanjain
Benchmarks & researchLaya on a 12-year-old laptop CPUSub-second decisions on an i5-5200U with no GPU.X post · Farhan60291312
Benchmarks & researchLaya, 50× faster, on open alternativesA benchmark of the open-source Jev alternatives.X post · wquguru
Benchmarks & researchLaya-MLX, taken apartA repo to understand an encoder-only decision architecture.X post · Zorawar Purohit
Benchmarks & researchLeJudge: JEPA × JevA JEPA world-model planner with Jev as the judge.X post · AbdelStarkLBenchmarks & researchLocal Laya against cloud Jev, on latencyAbout 45 ms deciding locally against about 300 ms in the cloud.X post · Oluwaphilemon1
Benchmarks & researchOpen Jev PlaygroundOne place to try every open-source alternative to Jev.Site · Mernit
Benchmarks & researchState in, auditable decisions outA short summary of what Jev returns.X post · 0xChaseTMTBenchmarks & researchTurn any LLM into a Jev-like modelA trained KV cache bank, with a live demo.Site · faangguyindia
Benchmarks & researchTypeSafe’s System One, for decisionsA recap.X post · silentguyy66
Benchmarks & researchAnyJevTurn any LLM into a Jev-style decision model, with no training.GitHub · ★ 440 · nokia-applied-research
Benchmarks & researchagent-jevAgentJev-0.6B - a fast 'System One' decision model for AI Agents: feed it any unstructured state (diffs, traces, logs) and structured questions, get…GitHub · ★ 286 · malevrigns
Benchmarks & researchopen-jev-typed-decision-engineOpen reproduction of TypeSafe Jev: a 150M typed decision engine (noul/choice/score in one non-autoregressive pass, calibrated confidence).GitHub · ★ 43 · intikhab49
Benchmarks & researchlocal-jevA local, offline System One server compatible with TypeSafe's Jev API.GitHub · ★ 8 · amithgc
Benchmarks & researchjev-2048Instrumented 2048 web lab where every move is a TypeSafe Jev Choice over the four directions with no heuristic fallback; probability, confidence,…GitHub · ★ 5 · ARCJ137442
Benchmarks & researchjevbench (dhruvmehra)Benchmark TypeSafe JEV against LLMs, fine-tuned BERT, Laya and zero-shot NLI on text classification: accuracy, calibration, latency, throughput, cost.GitHub · ★ 5 · dhruvmehra
Benchmarks & researchjev-botSelf-hosted Jev decision workbench and Feishu bot: automatic choices, probabilities, and experimental word/character writing.GitHub · ★ 4 · nssmd
Benchmarks & researchjevloopA Python agent runtime powered by TypeSafe Jev: guarded tool execution, isolated Docker sandboxes, and side-by-side LLM comparisons.GitHub · ★ 4 · parkavenue9639
Benchmarks & researchsysone-benchFirst independent head-to-head benchmark of System One decision models (Laya vs Jev) on byte-identical inputs.GitHub · ★ 4 · instax-dutta
Benchmarks & researchjev-storyboard-labGoogle ADK vs Microsoft Agent Framework for structured-output agents, with TypeSafe Jev as a vendor-neutral QC gate.GitHub · ★ 3 · jimmyliao
Benchmarks & researchwerrZero-memory System-1 decision engine & TypeSafe Jev wire-compatible runtime powered by Mandelbrot wave dynamics (JevBench #1).GitHub · ★ 3 · pCwOrM
Benchmarks & researchmetask-jevCalibrated typed-decision models, one forward pass each.GitHub · ★ 2 · metask-ai
Benchmarks & researchask-jev-aiA public wall where anyone asks a question in three to fifteen words and Jev, TypeSafe's judgment model, answers yes, no, or it depends in about 100…GitHub · ★ 2 · waynesutton
Benchmarks & researchbes-kelime-jevNe yazarsanız yazın, beş kelimeden biriyle cevap veren sohbet botu.GitHub · ★ 2 · mahmut-gundogdu
Benchmarks & researchtempo-jev-demoA natural-language task workspace comparing performance across AI models (TypeSafe's Jev, GPT-5.6 Luna, and Gemini 3.8 Flash).GitHub · ★ 2 · mychaelangelo
Benchmarks & researchdsh-jev-verifyDeepSeek Harness plugin that exposes TypeSafe Jev as jev_decision (choice/score/noul in one call) plus jev_verify (live labeled benchmark).GitHub · ★ 1 · xienda
Benchmarks & researchopenjev-multimodalLocal multimodal decisions on your Mac.GitHub · ★ 1 · jev-skills
Benchmarks & researchsystem-one-security-triageRecorded comparison of Jev, Terra, and Opus on 100 synthetic security-triage cases, five passes each, with a static inspectable dashboard.GitHub · ★ 1 · Robertzu43
Benchmarks & researchjev-calibrateTune your Jev questions against your own labels, then confirm on held-out data.GitHub · smkrv
Benchmarks & researchJevPokerBenchA Texas Hold'em benchmark for decision models, with leaderboards.GitHub · Prophetlab
Benchmarks & researchA review of Jev for AI workflowsFast, structured decisions, with the vendor's own figures quoted.X post · broadrangeAI
Benchmarks & researchAn agentic classification loop, 7x fasterThe loop replaced with a single Jev call, and the numbers published.Site · r6i.it
Benchmarks & researchCoverage is not product valueJev judges rather than writes, and that is where its value sits.X post · David Zhu
Benchmarks & researchJev against Laya, a smoke testAccuracy, calibration, latency and cost per decision, side by side.X post · cruzex100
Benchmarks & researchjev-as-judgeA refund agent graded by Jev, recorded as an Opik experiment.GitHub · Akshay Pachaar
Benchmarks & researchJev-in-the-LoopResearch on which LLM-decision tasks Jev actually speeds up.GitHub · Tongyun1JBenchmarks & researchJev-like decisions from open LLMsTurn a low-cost open model into a fast decision model, without training.X post · Anand Prasad
Benchmarks & researchjev-papersA thousand arXiv papers, one Jev decision each, checked by an LLM judge.GitHub · stas4000
Benchmarks & researchJevTunerTuning for Jev-style decisions.GitHub · liushiliushi
Benchmarks & researchKaLM-JevA local judgement engine in Nano, Small and Large sizes.GitHub · KaLM-Embedding
Benchmarks & researchLayaOpen weights that answer typed questions in one forward pass, 33 ms on a T4.GitHub · Nandakishor M
Benchmarks & researchlaya-vs-jevLocal MLX and hosted decisions playing T-Rex side by side.GitHub · Viraj Bhartiya
Benchmarks & researchLogJevText, images or audio in, choices and scores out of the logprobs.X post · MajoSayo
Benchmarks & researchsolar-mini4-jevJev-style decisions on Solar Mini 4.GitHub · Sung Kim
Benchmarks & researchjevalsAgent evals and guardrails in one request.GitHub · ★ 82 · openlayer-ai
Benchmarks & researchlaya-vs-jev-arenaTwo decision models race in Snake, then fight in an arena.GitHub · ★ 29 · PromptEngineer48
Benchmarks & researchJev-QuantumA sub-microsecond System 1 model whose accuracy is a Gaussian.GitHub · ★ 28 · karminski
Benchmarks & researchopen-jevA from-first-principles rebuild of the ideas behind Jev.GitHub · ★ 21 · kyegomez
Benchmarks & researchjevalWorks out what a Jev confidence score is actually worth.GitHub · ★ 16 · rlaope
Benchmarks & researchopen-spark-jevLocal System One models on Qwen3, for an NVIDIA DGX Spark.GitHub · ★ 15 · abhishek085
Benchmarks & researchedgejev离线可用的本地类型化决策:4 核 CPU 单题 15.6ms。Local & offline Jev / System One inference on CPU — ONNX + INT8, no torch at runtime.GitHub · ★ 9 · yzfly
Benchmarks & researchnanojev-arenaNanoJev Snake Arena: 1v4 human-vs-AI battleship + 100-agent swarm simulator.GitHub · ★ 7 · caijinchun
Benchmarks & researchpoorjevOpen-source, local Jev alternative: a System One decision layer with provably calibrated confidence (ECE 0.170→0.071).GitHub · ★ 7 · rupeshpoojary9
Benchmarks & researchjevtest (joshhu)情緒測謊器:嘴上說「好」,心裡真的好嗎?用 TypeSafe Jev(System One 模型)透過 OpenRouter 即時判斷,並與一般 LLM 對照.GitHub · ★ 6 · joshhu
Benchmarks & researchcu-JevCuda-Jev — a CUDA-native Jev System One decision inference engine.GitHub · ★ 3 · dtunaiJBenchmarks & researchjev-uiReact components that resolve which component to render, how to order a list, and whether to show an affordance — from calibrated judgments returned…GitHub · ★ 3 · etweisberg
Benchmarks & researchruby_llm-providers-typesafeTypeSafe System One models (Jev) for RubyLLM: typed judgments, evaluations and reranking.GitHub · ★ 2 · javiergradiche
Benchmarks & researchtypesafe-jev-mcpMCP server exposing TypeSafe's Jev model as a typed evaluate tool.GitHub · ★ 2 · anasbekheit
Benchmarks & researchjev-benchmark (themsquared)Reproducible benchmark for TypeSafe AI's Jev on agent tool-call risk classification: accuracy, latency, and whether the confidence score is worth…GitHub · ★ 1 · themsquared
Benchmarks & researchwhat-is-jevIndependent, source-linked research on TypeSafe AI's Jev (System One), with 947 rubric-scored public repositories, recurring patterns, datasets, and…GitHub · ★ 1 · g0runmezadam
Benchmarks & researchA correction about generating haiku with JevChoice does not generate hiragana, and the earlier claim was wrong.X post · yurinakanishi33ABenchmarks & researchA Jev speed test app, made publicOpened up because the usage fees are small.X post · tubone24
Benchmarks & researchAn open alternative appears: Laya421M parameters, built to decide, runs locally.X post · sl1ma4BBenchmarks & researchBuild a Jev JudgeAkshay Pachaar’s worked example: evaluate a refund-support agent with Jev instead of a generative judge, and record the verdicts in Opik.X post
Benchmarks & researchCalibrated confidence in 70 to 500 msYes/no, choices and scores, with the published price per token.X post · shree_code
Benchmarks & researchChatJevDriving the classifier as a next-token predictor, one token at a time.GitHub · Erik Dunteman
Benchmarks & researchDecisions inside the system, not in a chatWhat makes Jev different from a conversational model.X post · xooox888EBenchmarks & researchEarly access since September 15A System One model for agent decision loops, in typed options.X post · stretchcloud
Benchmarks & researchfedjev-benchDoes Jev hear a rate rise coming in an FOMC statement?GitHub · maybern-tripp-smith
Benchmarks & researchFive dollars a month was not enoughReal tasks forced reading the skills docs and using state properly.X post · zhengliIBenchmarks & researchIs a quick classifier worth building?The traps that come with building one fast.X post · satto_sann
Benchmarks & researchJev against an LLM on your own feedThe same tweets classified both ways, open on GitHub.X post · tshmieldev
Benchmarks & researchjev-acentoA pre-registered audit of how Jev handles Spanish.GitHub · Marcos Martinez
Benchmarks & researchjevchatA chatbot built out of a model that only answers typed questions.GitHub · Kyle Pena
Benchmarks & researchJevGuardA decision runtime with zero-token caching and a calibrator.GitHub · Seb4Ez
Benchmarks & researchLaya, 421M parameters, local86.5 decisions a second on about 1 GB of inference memory.X post · Md Fazal
Benchmarks & researchminojevIndependent Jev-style replica with head training: a frozen Qwen3-1.7B backbone plus a small trained decision head returns calibrated…GitHub · zeredy879
Benchmarks & researchOpenJevAn open attempt at a Jev-class decision model.GitHub · SiliconLabAITBenchmarks & researchThe first public System One modelState and explicit questions in; choices, scores and probabilities out.X post · yuntian9999
Benchmarks & researchNanoJevA 0.6B replica: parallel decisions in one forward pass, no decoding.GitHub · ★ 2.2k · TianyuCodings
Benchmarks & researchopenJev-verdict-2.0A 151M non-autoregressive decision engine, with its numbers published.GitHub · ★ 283 · Heman10x-NGU
Benchmarks & researchjevbenchJevBench v1 - a benchmark for Jev-class typed decision models: smart, cheap, fast, reliable, open.GitHub · ★ 114 · fstandhartinger
Benchmarks & researchJevForgeEnd-to-end Jev-style structured-decision stack for auditable data construction, Qwen3.5-0.8B training, fixed Mind2Web and OOD evaluation, preliminary…GitHub · ★ 35 · zwliJay
Benchmarks & researchjev-rag-benchmarkReproducible benchmark for measuring Jev reranking quality, latency, and cost in RAG.GitHub · ★ 14 · erendikmenn
Benchmarks & researchjev-architectHelps you find, design and evaluate a Jev decision loop.GitHub · ★ 6 · karanb192
Benchmarks & researchjev-ood-calibrationIndependent calibration study of TypeSafe Jev: published raw responses for three public benchmarks plus 900 rule-generated support tickets (choice /…GitHub · ★ 6 · scienthoon
Benchmarks & researchOpenJevOpenJev: an independent Jev-inspired System One decision API based on TypeSafe.ai concepts.GitHub · ★ 5 · xingwudao
Benchmarks & researchjev-reward-model-evaluationJev 1.13 reward-model evaluation across 8 benchmark tracks, with an interactive report and 54-row SOTA comparison.GitHub · ★ 4 · goya4140
Benchmarks & researchjevarena (chenmingtang830)Open-source BYOK arena for Jev and other AI judges.GitHub · ★ 4 · chenmingtang830
Benchmarks & researchOpenJev (GPT-AGI)Jev-compatible System 开源Jev.GitHub · ★ 4 · GPT-AGI
Benchmarks & researchJevUnofficial TypeSafe Jev showcase — System One decisions, not chat.GitHub · ★ 3 · cobusgreyling
Benchmarks & researchjev-web-analyzerSee what Jev thinks about your SaaS website — powered by ReplyNodes web context and Vercel AI Gateway.GitHub · ★ 3 · replynodes
Benchmarks & researchjevtestSemantic test matchers for Vitest and Jest, powered by TypeSafe's Jev model.GitHub · ★ 3 · realZachi
Benchmarks & researchORIGIN-CIVILIZATIONAI life-and-civilization simulation: TypeSafe Jev makes every decision (typed, probabilistic, auditable); LLMs plan — OpenAI-compatible APIs, local…GitHub · ★ 3 · JacquesGariepy
Benchmarks & researchjev-cloud-quiz三大クラウドの機能名を、TypeSafe AI の System One モデル Jev が確率つきで判定するデモ.GitHub · ★ 2 · minorun365
Benchmarks & researchjev-labsNever confidently wrong: a TLA+-verified consensus kernel around TypeSafe's Jev, run through 1,680 chaos-tested pharmacy decisions with zero wrong…GitHub · ★ 1 · copyleftdev
Benchmarks & researchopenpoke-meets-jevFork of OpenPoke where the yes/no decisions go to Jev instead of Claude Sonnet, with a same-inputs A/B against the replaced LLM decision and a…GitHub · ★ 1 · 0xShin0221
Benchmarks & researchjev-e2eThe same eBay test flow: 47 s on Jev, 79 s on Claude Sonnet 5.GitHub · Jason Lu
Benchmarks & researchjev-evaluationNine experiments and 28 predictions, all fixed before any data.GitHub · Will Kelly
Benchmarks & researchWebMCP browser benchmarkA reported 49-task benchmark combines Jev for tool selection, Mercury 2.5 for arguments and WebMCP for browser actions.Site · idan levin
Benchmarks & researchTop 5 Jev Alternatives: __ I have been running a…Top 5 Jev Alternatives: __ I have been running a benchmark since this morning on every Jev alternative against Jev.X post · ItsCuthulhu
Benchmarks & researchvonA sub-15 ms local drop-in alternative to Jev.GitHub · ★ 610 · wfzyx
Benchmarks & researchjev-align (Sutro)Turns human labels into calibrated Jev classifiers with GEPA.GitHub · ★ 284 · sutro-sh
Benchmarks & researchjev_localReplicating Jev with a local LLM.GitHub · ★ 35 · Argos1111
Benchmarks & researchjev-capability-atlasWhere Jev holds up and where it breaks, with real API receipts.GitHub · ★ 26 · Zaious
Benchmarks & researchjev.nuNushell module for the TypeSafe System One API: typed decisions with calibrated probabilities.GitHub · ★ 8 · cablehead
Benchmarks & researchjevscapeRuneBench harness for TypeSafe's Jev: bounded action catalog, tick-mode controller and a live dashboard.GitHub · ★ 8 · Skyvern-AI
Benchmarks & researchjev-alignCalibrated alignment verifier for LLM responses and agent plans — powered by Jev.GitHub · ★ 4 · caiovicentino
Benchmarks & researchnew-api-plugin-typesafeTypeSafe AI System One (Jev) task plugin for QuantumNous/new-api — native /v1/systemone, synchronous evaluation, token billing.GitHub · ★ 3 · FFatTiger
Benchmarks & researchjev-chessChess moves, evaluations, persona opponents, and game classification with TypeSafe AI System One.GitHub · ★ 2 · hemanth
Benchmarks & researchjev-modeI kept watching coding agents burn context on decisions that aren't hard - triage 400 tickets, tag 600 files, route to one of six teams.GitHub · ★ 2 · ddfeyes
Benchmarks & researchresearch_deskTypeSafe Jev demonstration for new analyzation — experimenting with Jev for fast analysis of news and tickers.GitHub · ★ 2 · 0xnairb
Benchmarks & researchzerosweepAutonomous System-One Triage Engine & Benchmark powered by TypeSafe AI (Jev).GitHub · ★ 2 · sysadarsh
Benchmarks & researchjev-carryforwardWhat your last session knew, scored against what this one is doing.GitHub · ★ 1 · Dharundp6
Benchmarks & researchmodelsystemCurated catalog of System One / Decision Models — contributions for modelsystem.one.GitHub · ★ 1 · fabricioctelles
Benchmarks & researchtypesafe-vs-deepseekTypeSafe (Jev) vs DeepSeek-flash: side-by-side speed/token/cost/accuracy comparison across invoice extraction, email classification, and reranking.GitHub · ★ 1 · markfive-proto
Benchmarks & researchkevA tiny Jev-like model on Qwen2.5-0.5B that trains and runs on a MacBook.GitHub · Jared Palmer
Benchmarks & researchOpenRouter’s Ori EvalJev judged a 30-way labelling task faster than every model tested.X post · OpenRouter
Benchmarks & researchdeepseek-v4.1-flash-jevAn open model made to answer like Jev with a scoring endpoint.X post · Nick Khami
Benchmarks & researchJev against Fable 5.1, 100 buildersThe same post scheduler, built twice, in front of an audience.X post · Florian Darroman
Benchmarks & researchopenjevOpen, Jev-compatible System One decision server on DiffusionGemma.GitHub · ★ 386 · razorback16
Benchmarks & researchawesome-jev (OmniJev)Papers, open reproductions and independent evaluations behind System One models and Jev.GitHub · ★ 286 · OmniJev
Benchmarks & researchjev-visualAn educational Jev-like visual inference experiment on Apple Silicon: shared context, direct candidate scoring, and local visual demos.GitHub · ★ 271 · hr98w
Benchmarks & researchreflexA small open decision model: state and typed questions in, probabilities out.GitHub · ★ 133 · kshetrajna12
Benchmarks & researchVerdict-open-jevAn open 151M decision model on ModernBERT, with a WebGPU playground.GitHub · ★ 94 · Heman10x-NGU
Benchmarks & researchopen-alternative-jevTyped, calibrated decisions from any open-weights model, on your own GPU.GitHub · ★ 53 · ikermoel
Benchmarks & researchmini-jevMini-Jev: what a Jev-style typed-decision interface looks like on a frozen Qwen3-4B — read the option letter's logits instead of generating JSON.GitHub · ★ 53 · r-ms
Benchmarks & researchLitJevA reproduction of Jev that turns any Qwen model into a fast decision model, serving the same /v1/systemone schema (Choice, Score, Noul) with no…GitHub · ★ 43 · zhengxuyu
Benchmarks & researchJev_apps看看 Jev 能做什么:用中英文讲清热门应用、工作原理和各自优缺点。Explore Jev apps with plain-language examples, explanations, and comparisons.GitHub · ★ 33 · JackZeng
Benchmarks & researchjevifyAn agent skill to discover TypeSafe Jev opportunities, design typed questions, and learn from recent community experiments.GitHub · ★ 27 · altryne
Benchmarks & researchtypesafe-localInspired by TypeSafe Ai, Ask a local LLM typed questions, get calibrated probabilities instead of text.GitHub · ★ 8 · aabolfazl
Benchmarks & researchjev-search-rerank-evalDoes a TypeSafe Jev rerank beat embedding search?GitHub · ★ 6 · zhuyansen
Benchmarks & researchtenbinSplit a judgment into Choice, Score and Noul, lint it, then measure it.GitHub · ★ 3 · simota
Benchmarks & researchtypesafe-ai-playgroundCommunity TypeSafe AI playground: 110 use cases, games, dilemmas and model challenges, with editable prompts, A/B comparisons and a mobile-friendly…GitHub · ★ 3 · nickthompson480
Benchmarks & researchaskjev.aiAsk Jev anything. It won’t answer, it will judge.Site · Wayne SuttonFBenchmarks & researchFive open Jev replicas, in Chinese小墨同学 on Laya 421M, Decider-2B, NanoJev 0.6B, Reflex and System-One 4B, and which run best on Apple silicon.X post
Benchmarks & researchJev on the WebMCP benchmarkJev picks the tool, a fast LLM fills the arguments: 49 of 49 tasks solved.X post · Idan Levin
Benchmarks & researchOne judge call vs twelve dimension scoresOne direct Jev question per row against 12–14 Jev-scored dimensions with locally fitted weights on three classification tasks: 5,477 test rows,…Site
Benchmarks & researchopenjev on Qwen 4BAn MLP trained on top of Qwen 4B that works like Jev.Site · Alex Nikolic
Benchmarks & researchVerdict (Open-jev) launch threadHemant Kumar on his 151M ModernBERT reproduction: 0.83% adaptive calibration error, 3.0% of predictions flip when the option list is reversed, 35.6ms…X post
Benchmarks & researchSemIfSemantic ifs from open models, on a 3090 at home.GitHub · ★ 4.2k · TheoLeeCJ
Benchmarks & researchJevlikeExplore training a small option-scoring model.GitHub · ★ 1.3k · vinnylarouge
Benchmarks & researchdeciderOne-pass typed decisions with calibrated probabilities (System One style model), fine-tuned from Qwen3.5-2B.GitHub · ★ 351 · Mapika
Benchmarks & researchopen-jev (daseinlabs)One-pass option scoring with a local Gemma 3 4B on Apple silicon via MLX, inspired by jevlike, with a Doom demo.GitHub · ★ 107 · daseinlabs
Benchmarks & researchjev-eval-agentPersonal-assistant agent built on Vercel's eve with 100 mocked tools, measuring how many steps it takes when Jev picks the tool versus the LLM.GitHub · ★ 105 · vinilana
Benchmarks & researchjev-as-a-judgeUsing Jev as an evaluator.GitHub · ★ 83 · danielgshea
Benchmarks & researchjevfireJEV-inspired parallel decisions for CUDA LLMs.GitHub · ★ 64 · kikoncuo
Benchmarks & researchjevmlxJev-style parallel constrained decisions for any MLX model on Apple Silicon.GitHub · ★ 60 · bnsd55
Benchmarks & researchTypeARType-safe one-decision-per-token decoding engine for autoregressive LLMs, inspired by Jev.GitHub · ★ 40 · TypeLLM
Benchmarks & researchopen-jev (JoshuaSP)Typed JSON inference with DiffusionGemma, with Every and Jev benchmark results.GitHub · ★ 39 · JoshuaSP
Benchmarks & researchtypesafe-ai-benchmarkThis is a LLM Gateway that mimics typesafe ai structured output.GitHub · ★ 38 · iammrduncan
Benchmarks & researchsystem-one-openOpen replica of TypeSafe's Jev: typed calibrated decisions in one forward pass, on Gemma 4 E2B / Gemma 3 270M (Modal).GitHub · ★ 35 · mithalouni
Benchmarks & researchopenjev (zhihz)Local bilingual probability decisions from context, questions, and candidate answers.GitHub · ★ 33 · zhihz
Benchmarks & researchsystem-oneBatched single-token choice inference for open language models, compatible with TypeSafe.GitHub · ★ 33 · sgoedecke
Benchmarks & researchjevgptA chatbot built on a model that cannot generate text (TypeSafe AI's Jev, driven autoregressively).GitHub · ★ 26 · Bewinxed
Benchmarks & researchJev on a LaptopInvestigate typed decisions on local hardware.GitHub · ★ 25 · rorshopping
Benchmarks & researchjev-column-raceLocal (or hosted BYOK) race UI that labels 1,000 withheld-star app reviews in parallel columns: TypeSafe Jev typed questions versus a Gemini JSON…GitHub · ★ 23 · goodrahstar
Benchmarks & researchtypesafe-ai-playground (BunsDev)Community TypeSafe AI playground: 110 use cases, games, dilemmas and model challenges, with editable prompts, A/B comparisons and a mobile-friendly…GitHub · ★ 20 · TypeSafeAI
Benchmarks & researchJev BenchmarksMeasure calibration, latency, and selective risk.GitHub · ★ 18 · AbdelStark
Benchmarks & researchjevbetterA stronger one-pass scorer over a variable list of text options: hashed n-gram encoder, rival-aware attention, gated head, temperature scaling,…GitHub · ★ 14 · olanotolu
Benchmarks & researchopenvonsOpenvons (open-Jev): 有限選択肢に確率で答える判断層 — テキスト / 画像 / 日本語音声コマンド.GitHub · ★ 13 · genai-craft
Benchmarks & researchRerank BenchCompare Jev with document reranking models.GitHub · ★ 8 · anessbelbati
Benchmarks & researchjev-mcp (arunav25)Connect JEV to MCP clients and compare its judgments against general-purpose LLMs using shared datasets and measurable accuracy.GitHub · ★ 7 · arunav25
Benchmarks & researchPhishing BenchCompare phishing judgments across models.GitHub · ★ 6 · anisselbd
Benchmarks & researchjev-benchmarkBenchmarks and a playground for TypeSafe's Jev (System One) model: chess, and who-is-the-player-talking-to for speech-to-text game NPCs.GitHub · ★ 6 · wondertwins
Benchmarks & researchjev-korean-benchmarkReproducible early-access evaluation of Jev on Korean understanding and medical text, with runtime and cost evidence.GitHub · ★ 6 · mahlernim
Benchmarks & researchjev-lmA word-level language model whose output layer is Jev: n-gram drafter, Noul chunk verification, bits-per-token eval.GitHub · ★ 6 · y0usaf
Benchmarks & researchmcts-agentDiscriminative Monte Carlo Tree Search using TypeSafe Jev System One Primitives and Gemini.GitHub · ★ 6 · lhemerly
Benchmarks & researchai-elo-rankerHigh-speed recursive AI Elo tournament engine powered by Jev and Swiss matchmaking.GitHub · ★ 5 · opaielsheikh
Benchmarks & researchjev-little-airwaysA show-and-tell capability study for Jev, TypeSafe's System One decision model.GitHub · ★ 5 · lbotinelly
Benchmarks & researchLegalForecastBenchLegalForecast-MTD benchmark alpha and official evaluation workflows.GitHub · ★ 5 · johnhughes3
Benchmarks & researchSecurity BenchEvaluate injection and code-security detection.GitHub · ★ 4 · Gaurav-Gosain
Benchmarks & researchjev-chatA chatbot from typed Jev decisions: hierarchical speculative decoding over System One probabilities.GitHub · ★ 4 · adhyaay-karnwal
Benchmarks & researchqwen-rlcdJev-style calibrated decision model (Choice/Score/Noul) on Qwen3.5-0.8B.GitHub · ★ 4 · shamazharikh
Benchmarks & researchsystem-one-gemmaOpen-source Jev-style System One decision model.GitHub · ★ 4 · akash-kamat
Benchmarks & researchAgent Failure BenchmarkTest attribution of failures in agent traces.GitHub · ★ 3 · TokenTrim
Benchmarks & researchSpam EvalTest spam detection against simpler baselines.GitHub · ★ 3 · bitnovus
Benchmarks & researchRouting ExperimentEvaluate Jev as a model-selection layer.GitHub · ★ 3 · TokenTrim
Benchmarks & researchjev-behavior-studyIndependent Jev 1.13.0 behavior study: report, controlled prompt experiments, raw results, and offline verification.GitHub · ★ 3 · RINNECODER
Benchmarks & researchjev-explorationJev (TypeSafe) exploratory thread: claim audit, live demos, and runnable code.GitHub · ★ 3 · SamuelSacco
Benchmarks & researchCalibreStudy confidence thresholds on banking intents.GitHub · ★ 3 · FirasSX914
Benchmarks & researchclarity-judgeWriting checked on separate named axes, each with its own verdict.GitHub · ★ 2 · TypeSafeAI
Benchmarks & researchDecisionBridgePut a decision interface over existing LLMs.GitHub · ★ 2 · grishahq
Benchmarks & researchcalibreCalibration and confidence-based routing measured on Banking77: 80.2% accuracy at $0.103 per 500 decisions.GitHub · ★ 2 · FirasSX914
Benchmarks & researchjev-freeformAn observable raw-character chat experiment powered entirely by TypeSafe Jev Choice.GitHub · ★ 2 · kesku
Benchmarks & researchjev-research-evalReproducible Jev Ultrafast research-browser eval harness + field note (QC’d cases, suite runner, report generator).GitHub · ★ 2 · jgridifier
Benchmarks & researchtypesafe-ai-playground (markjaquith)A playground for experiments around Jev, TypeSafe's System One model.GitHub · ★ 2 · markjaquith
Benchmarks & researchgpt-vs-jevCompare GPT generated language with JEV structured Noul decisions on the same input.GitHub · ★ 1 · TanayPadar
Benchmarks & researchjev-secret-detectionMeasures how well TypeSafe's RLCD-Jev model spots real secret credentials in file snippets.GitHub · ★ 1 · teyhouse
Benchmarks & researchjev-synergy-screeningJev (TypeSafe System One) × ASReview SYNERGY abstract screening demo — Choice/Noul vs gold labels.GitHub · ★ 1 · PistachioAIHQ
Benchmarks & researchkyotsu-ai-benchAI benchmark on Japan's 2026 Common Test: Jev vs luna-none vs luna-low (static dashboard).GitHub · ★ 1 · shibadogcap
Benchmarks & researchopenjev-experimentsExperiments with openjev, an open Jev-style option-logit runner, on local models.GitHub · ★ 1 · zefir1990
Benchmarks & researchpadflow-jev-evalsTyped-decision benchmark from PadFlow (land development SaaS): schemas, anonymized labeled rows, and a runner for confidence-calibrated models like…GitHub · ★ 1 · zsavage8
Benchmarks & researchPocketJevOn-device iPhone visual decision tool using MLX and Qwen3-VL direct option logits.GitHub · ★ 1 · NullPo-jp
Benchmarks & researchRISC-jeVI tortured Jev into being a RISC-V CPU.GitHub · ★ 1 · i2cjak
Benchmarks & researchevery.to37 documents, 21 questions each, 1,709 judgments for under a cent.Site · Mike Taylor
Benchmarks & researchInternal classifier field noteMatched-precision comparison against a private fine-tuned classifier.X post · identity
Benchmarks & researchnearhere.eventsHead-to-head test at validating local event listings, with cost and latency.Site · Jon
Benchmarks & researchJev vs Qwen on CerebrasVideo comparison against a structured-output LLM baseline.X post · IamMrDuncan
Benchmarks & researchA local JevA Jev-style model running locally, with room to get faster.X post · 生ビール
Benchmarks & researchFinancialPredictionJevUsing Jev to test how well it predicts financial markets(just like most llms as of september 2026, it doesnt do that good).GitHub · thodoh1
Benchmarks & researchJev plays chessJev picks from the legal moves, compared with reasoning models.Guide · Maxim Saplin
Benchmarks & researchjev-alpha-benchDoes Jev predict stock returns from news?GitHub · Gaurav-Gosain
Benchmarks & researchjev-anotacao-sentencasJev (TypeSafe) vs. Gemini 3.8 Flash vs.GitHub · lab-dados
Benchmarks & researchjev-deferred-crispificationPosition paper: the Hidden-Markov and fuzzy primitives missing from TypeSafe AI's Jev and System-One decision models.GitHub · dnakhoa
Benchmarks & researchjev-demosDemos to test the effectiveness of TypeSafe's "Jev" System One Model.GitHub · Bud-ro
Benchmarks & researchjev-dev同じ発言を jev と LLM の両方に判定させ、感情の変動値のズレと応答速度を1画面で見比べるデモ(affectus + Vercel AI Gateway).GitHub · n-yokomachi
Benchmarks & researchjev-headline-benchCan Jev pick the winner of a real headline A/B test?GitHub · Gaurav-Gosain
Benchmarks & researchjev-jp-addressJev (TypeSafe) 性能評価プロジェクト — 日本郵便 KEN_ALL をマスタに、AI SDK 経由の Jev が住所のあいまい一致にどこまで使えるかを検証.GitHub · smasato
Benchmarks & researchjev-labTypeScript experiments, evaluations, and latency benchmarks for TypeSafe's Jev model.GitHub · Menny1337
Benchmarks & researchjev-pick-and-place-studyA small reproducible MuJoCo pilot comparing Jev, Claude Haiku, and reactive rules for pick-and-place.GitHub · tryaksh
Benchmarks & researchjev-playground (hegargarcia)Benchmarks Jev against other evaluation models in games with explicit states, legal actions, and measurable outcomes.GitHub · hegargarciaJBenchmarks & researchjev-report发明 RLHF 的人,这次做了个不会说话的模型:Jev 独立研究报告。52 页 PDF + 50 条中文实测复现包 + 143 条可回溯数据表.GitHub · HackSing
Benchmarks & researchjev-shadcn-lint-evalA small second eval for shadcn-ui/lint that uses TypeSafe's Jev to judge the linter's own output.GitHub · blas0
Benchmarks & researchjev-trace-classifierApplication of TypeSafe Jev (noul judgment primitive) on the collusion.wiki corpus: agent vs human page authorship, head-to-head vs local…GitHub · sypherin
Benchmarks & researchjev-user-juryTypesafe/Jev public X demo.GitHub · cardotrejos
Benchmarks & researchmisereru-slide-jevOngoing Japanese research deck on Jev and System One models, maintained as Markdown slides.GitHub · myokoym
Benchmarks & researchParallel Constrained Decoding (Qwen2.5-1B-RLCD)Hugging Face Space exploring open-source parallel constrained decoding as an alternative to Jev.Site
Benchmarks & researchshade-arena-jev-monitorEvaluating TypeSafe's Jev as a fast monitor and action gate for agent sabotage in SHADE-Arena, compared with Gemini 2.5 Flash/Pro.GitHub · nican2018
Benchmarks & researchsystem-one-adapter-rustRust port of TypeSafe system-one-adapter (LLM-backed system_one evaluations).GitHub · codeitlikemiley
Benchmarks & researchthaiexam-jev-chartsCharts: TypeSafe Jev evaluated on Thai standardized exams vs 110 other models.GitHub · vehas
Benchmarks & researchTyped decisions, not chatIndependent walkthrough separating TypeSafe's published claims from public evidence.Site
Benchmarks & researchtypesafe-oraclesEvaluating TypeSafe's System One primitives (Choice/Score/Noul) — where a typed oracle beats an LLM call.GitHub · trophee-bot
Benchmarks & researchTypesafe_chess_evalAn evaluation of typesafe AI chess.GitHub · AliceRoselia
Benchmarks & researchOpenJevAn independent experiment in local decision models.GitHub · ★ 4.3k · TheoLeeCJ
Benchmarks & research50 financial jobs, one modelThe Fintech Builder runs Jev on fifty financial jobs and shows where it works and where it does not.YouTube · youtube.com/@TheFintechBuild
Benchmarks & researchatisbo.dev> We benchmarked TypeSafe's Jev against our production LLM classifier — with real customer data and human ground truth.Site · Brian Ochoa
Benchmarks & researchagent-evalsDeterministic LLM-agent eval harness: rule scorers for tool/latency/cost failures, plus an optional calibrated TypeSafe Jev judge you can gate CI…GitHub · marianoberton
Benchmarks & researchfieldnote.clawscience.comCan Jev reason across mechanics, relativity and gravitation?Site · CursedCoder
Benchmarks & researchchinese-workflow-decision-benchReusable Feishu-style workplace message classification benchmark: 64 frozen synthetic Chinese scenarios, Choice and four-Noul workflows, published…GitHub · Adkid-Zephyr
Benchmarks & researchDoes an open-weight decision model beat a hosted one?ASTGL compares local Laya with hosted Jev on typed decisions.Guide
Benchmarks & researchciv6-mcp.lwilko.comExcited to enter Jev into my Civ 6 benchmark :blob_aww: Has anyone else tested out strategic reasoning yet?Site · Magykman (Liam)
Benchmarks & researchnola-typesafe-testFinally got access, I made a quick comparison between Cerebras GPT-OSS-120B and JevGitHub · Evgen Mykhailenko
Benchmarks & researchdecisionmodels.aiGot a bit obsessed with the space and built decisionmodels.ai tonight.Site · sahil
Benchmarks & researchgit.sub-net.athere a short Jev test for NPC decisions in our slop MMO on a slop paper - its awesome, sadly the latency from the EU to the Jev test infra kills it…Site · Sub
Benchmarks & researchcultivarHey! We at Pinecone added TypeSafe support to Cultivar, our library for evaluating and benchmarking agent skills and docs against agents in sandboxes.GitHub · arjun
Benchmarks & researchi build a model x45 smaler but it Beats GPT-2 124M BYTE_BPC = 1.142985 (TOKEN_BPC = 9.152823) on the same metric https://huggingface.co/Uuuuuuniiiiiii/goldworm-v6 as well as my agencyi build a model x45 smaler but it Beats GPT-2 124M BYTE_BPC = 1.142985 (TOKEN_BPC = 9.152823) on the same metric as well as my agencySite · DARK(uniency.com)Sudocorn
Benchmarks & researchyuwakisa.comI have a benchmark that tests models for "here's a set of statements, express the common principle behind them" and the model expresses a principle.Site · hypnoticfuzzwave
Benchmarks & researchjev-test-1I made a deliberately tiny, dependency-free Jev test showing state, typed questions, and probabilities.GitHub · sethacked
Benchmarks & researchI tested Jev on 12 real use casesNate Herk compares speed and cost against ordinary models, then builds an X feed classifier and a real-time paper-trading prototype live.YouTube · youtube.com/@nateherk
Benchmarks & researchIs Jev really better, faster and cheaper?Edward Donner puts Jev to the test against Luna.YouTube · youtube.com/@EdwardDonner
Benchmarks & researchjev-guardbenchBenchmark harness asking whether a System One model (hosted TypeSafe Jev or self-hosted open-source Kev) can stand in for LLM-as-judge in agent…GitHub · dfranco-projects
Benchmarks & researchjev-healthcare-labOpen experiment comparing TypeSafe Jev vs DeepSeek on 96 healthcare tasks across 12 scenarios (quality, latency, cost) with per-task data and…GitHub · JuneYaooo
Benchmarks & researchJev-OmniJev-Omni is an open multimodal System One–style decision classifier: give it a state, a question, and typed options over text, image, audio, or…Site · akhilaaa3
Benchmarks & researchjev-packsEvidence-gated registry of Jev question packs: curated questions, golden cases, pinned model versions, and reproducible offline benchmark results…GitHub · dtduc-git
Benchmarks & researchjev-robotics-evalEvaluation harness for robot control with a JEV-compatible decision service on MetaWorld and RoboTwin tasks (text and vision modes, privilege levels,…GitHub · lose4578
Benchmarks & researchjev-security-prioritizationReproducible experiment: TypeSafe Jev prioritizes SCA/SAST findings using richer context than severity scores, compared to severity-only and points…GitHub · san3ncrypt3d
Benchmarks & researchjev-skill-router-benchIndependent, reproducible measurement of a Jev (TypeSafe System One) skill router on one 84-skill Hermes roster: 81 author-labelled turns, router…GitHub · OrMizL
Benchmarks & researchjevalsLocal evaluation workbench for TypeSafe Jev: author Noul/Choice/Score (and combined) questions with expected answers, run them, and compare saved…GitHub · dayhaysoos
Benchmarks & researchJevals.comHosted independent benchmark boards: TypeSafe Jev (typesafe-ai/jev via Vercel AI Gateway) vs six LLMs on PubMedQA (Noul), Banking77 (Choice), and…Site · Jevals
Benchmarks & researchjevlensEvaluate, calibrate, replay, and monitor TypeSafe Jev decisions from labeled CSV/JSONL: store full answer distributions, report accuracy/Brier/F1,…GitHub · k4its1t
Benchmarks & researchJevScopeLocal-first visual decision workbench and regression testbench for TypeSafe Jev: edit structured state and questions, inspect distributions, batch…GitHub · jeiel85
Benchmarks & researchjevtrimBenchmark: TypeSafe Jev as a context-relevance judge vs retrieval/summarization on LoCoMo, with matched token budgets and reproducible reports.GitHub · pdrpinto
Benchmarks & researchMalkuthOpen-weight multilingual System One–style decision models (2B/4B, Korean focus): Choice, Noul, and Score over /v1/systemone via Kev—adjacent to…GitHub · newfull5
Benchmarks & researchOpikComet’s open-source evaluation and tracing platform, Apache-2.0.GitHub
Benchmarks & researchPapers and open reproductionsOmniJev’s list of the papers, open reproductions and independent evaluations behind System One models.GitHub
Benchmarks & researchPDF RaceTimed document race: Docling → TypeSafe Jev vs Docling → Gemini vs Gemini reading the PDF, scored against arXiv metadata, with committed recorded…GitHub · goodrahstar
Benchmarks & researchposted the article, thanks for the ok.<@81830765507117056> posted the article, thanks for the ok.Site · nss
Benchmarks & researchRYOTIDERoll Your Own Typed Inference Decision Engine: Jev-style typed decisions from a single forward pass of a local LLM (MLX or PyTorch), measured on…GitHub · csabag
Benchmarks & researchSemIfSemIf studies runtime-defined semantic decisions using local open models.GitHub · TheoLeeCJ
Benchmarks & researchlindfors.noTested jev against deepseek v4.1 flash on norwegian text.Site · depurator
Benchmarks & researchtypesafe-offload-benchTested Jev with main "Players" (GPT/Claude) to work together.GitHub · Dmytro
Benchmarks & researchTyped EvalsEvaluate LLM responses, RAG samples, and agent traces with TypeSafe Jev as the judge, including optional calibration against human labels and…GitHub · TrustifAI
Benchmarks & researchdinostompwired jev into dinostomp, an oss eval harness: one choice question per item through /v1/systemone, probabilities on every record, and a blind control…GitHub · hobohotdog
Benchmarks & research⚡ Jev is HERE and it's rejecting ChatGPT's entire training playbook ♠ built by a co-creator of RLHF/ChatGPT who says⚡ Jev is HERE and it's rejecting ChatGPT's entire training playbook ♠ built by a co-creator of RLHF/ChatGPT who says chat was the wrong bet 🔹New…YouTube · fahdmirza
Other use cases
Agents & browsers 193Coding & code review 227Routing & model choice 161Context & memory 66Search & RAG 56Data & extraction 61Documents & research 28Safety & guardrails 131Support & inbox 45Marketing, sales & social 95Trading & finance 46Games & real time 173Creative & generative UI 40Home, robots & devices 20Everyday apps 113SDKs & integrations 308Guides & docs 202