CalibratedDecisions.

17 · 332 projects

Jev for benchmarks & research

Head-to-head tests, calibration studies and independent research into how well Jev decides, how fast, and at what cost.

What Jev decides: Benchmark questions with known answers, to check accuracy and confidence.

Projects, newest first

Benchmarks & researchJev 1.13.0 against Hopper on Benchmark HeavenHopper wins three columns, Jev keeps Intelligence.Site · Benchmark HeavenBenchmarks & researchValenTrain a Jev-like multimodal model by yourself.GitHub · ★ 51 · Liuziyu77Benchmarks & researchjev-dimabsaTypeSafe Jev baseline for DimABSA (SemEval-2026 Task 3) subtask 1: zero-shot and 3-shot valence-arousal regression.GitHub · ★ 6 · ZhangYiqun018Benchmarks & researchreflexbenchReflexBench — open benchmark and evaluation harness for System One models and typed decision engines.GitHub · ★ 4 · brida-aiBenchmarks & researchjev-guardrailsResearch/eval repo that runs the same mock carrier support agent and 25 behavioural guardrail rules behind two backends—chat-model JSON judge vs…GitHub · ★ 3 · deepansh-saxenaBenchmarks & researchtypesafe-jev-calibrate-for-code-reviewAbout calibrating Jev for code reviews.GitHub · ★ 3 · SelmarBenchmarks & researchjev-aitaBenchmark of TypeSafe's Jev against Sonnet 5, GPT-5 nano and local LLMs on 770 Reddit AITA verdicts: Brier scores, latency and cost.GitHub · ★ 2 · dchristopoulosBenchmarks & researchjev-dllmShared Yes/No decision adaptation with diffusion language models.GitHub · ★ 2 · zhouzihao11Benchmarks & researchjev-reviewsTypeSafe AI - Jev - first look and experiment with its API + Google reviews experiments.GitHub · ★ 2 · halfspin-qcBenchmarks & researchjev_fsdJEV AI Model Demo with FSD.GitHub · ★ 2 · BrendanH18Benchmarks & researchopen_system_oneCreating system one model just like jev.GitHub · ★ 2 · ReadyaddyBenchmarks & researchopenpave-jevPAVE skill for non-autoregressive decision models (Jev / Laya / TypeSafe): typed Choice, Score and Noul answers with calibrated probabilities.GitHub · ★ 2 · cnraiBenchmarks & researchjev (zeke)Research notes and an interactive Cloudflare Worker demo for Jev, TypeSafe AI's structured decision model.GitHub · ★ 1 · zekeBenchmarks & researchJev-Research-IndexJev Research Index is a bilingual catalogue of papers, software projects, interviews, public analyses, demonstrations, and social-media material…GitHub · ★ 1 · AgenticAPP-WebBenchmarks & researchLetJevDecideAsk a yes or no question. Let Jev, or a fair coin, decide.GitHub · ★ 1 · choxosBenchmarks & researchMEDJEVIndependent JEV-inspired System One-style typed decision research for clinical evidence, biomedical NLP, calibrated probabilities, and sleep signals.GitHub · ★ 1 · PAI-CUHKBenchmarks & researchollayaRun open decision models locally — pull, run and serve Laya and other open Jev alternatives behind a TypeSafe-compatible API.GitHub · ★ 1 · ollaya-devBenchmarks & researchA faster, cheaper path for business modelsJev as an alternative for business logic.X post · MilkMendyBenchmarks & researchA local open-source version of JevRun a Jev-like decision model locally.X post · coccoinomaneBenchmarks & researchA number two open-weight Jev-like modelAdded to the Decision Index, strongest for its size.X post · apolinario (poli)Benchmarks & researchA sub-100 ms decision engineHow the launch was described.X post · ASIHubHQBenchmarks & researchAn LLM distilling JevTypeSafe notes an LLM distilling Jev.X post · TypeSafe AIBenchmarks & researchBounded choices, not proseJev returns a choice from the options it was given.X post · GeethanTechBenchmarks & researchCLM-8B claims to be faster than JevAn open model that its release claims is up to 9x faster.X post · aiasssistantstoreBenchmarks & researchCLM-8B, a new Jev competitorAn 8B decision model positioned against Jev.X post · plotarmordevBenchmarks & researchConvex Decision EvalsLeaderboard that puts Jev and 14 LLMs through 108 four-option multiple-choice questions about the Convex backend platform, comparing accuracy,…GitHub · get-convexBenchmarks & researchJev and Laya on a medical diagnostic testSixteen of sixteen against thirteen, on authored notes.X post · OpenMedBenchmarks & researchJev and Laya on security queriesTwo decision models make different calls on the same security queries.X post · edwinfmesaBenchmarks & researchJev Does Not Play DiceIndependent calibration evaluation of Jev on fair random draws with known probabilities and synthetic forecast documents, with recorded outputs and…GitHub · KantaHayashiAIBenchmarks & researchJev for next best actionTesting Jev’s accuracy on predicting the next action.X post · dsmiley411Benchmarks & researchJev in tool filtration testsPraised for filtering tools.X post · eboppuBenchmarks & researchJev vs. Sol, on categorisationA text categorisation showdown.X post · JasonalcoBenchmarks & researchLocal Laya beats JevA local decision model coming out ahead.X post · NeuralFrameLabsBenchmarks & researchPredicting queries before submissionA community test with Jev and Laya.X post · edwinfmesaBenchmarks & researchSeparating signal from noise with Jev and OpusA pattern can still be noise, so ask what would change your mind.X post · DainBenchmarks & researchsnbt-jev-benchJev on Indonesia's SNBT 2025 university entrance exam, whose papers are never released: 159 questions reconstructed from memory by volunteers, 156 of…GitHub · misaalyaBenchmarks & researchThe System One paradoxOn System One models and where Jev goes next.X post · __tuan____Benchmarks & researchTrying to beat Jev on a DGX SparkAn open System One model, with the data already collected.X post · CuthBenchmarks & researchxor joins the Jev Decision Index 0.2A new open-weight entry, eighth on text alone.X post · apolinario (poli)Benchmarks & researchjevk5An open-weight alternative to Jev, Apache-2.0.GitHub · ★ 64 · allebeeBenchmarks & researchJevAnyCalibration-aware reinforcement learning for decision systems.GitHub · ★ 11 · weitianxinBenchmarks & researchlejudge-jev-jepaNatural-language constraints for JEPA planning, judged by a decision model.GitHub · ★ 9 · AbdelStarkBenchmarks & researchKevAn open-source System One engine with calibrated probabilities.GitHub · ★ 7 · arjun988Benchmarks & researchawesome-jev-robustnessTests, calibration audits and failure-mode studies of Jev.GitHub · ★ 5 · Yifan-LanBenchmarks & researchtypesafe-pocProva de Conceito do Jev, modelo da typesafe.ai.GitHub · ★ 5 · eminettoBenchmarks & researchjev-benchmark (YidiDev)Rubric-Based Zero-Shot Classification Benchmark: Jev vs Claude Haiku 4.5 vs OpenJev on rubric-conditioned classification, chained decision execution,…GitHub · ★ 3 · YidiDevBenchmarks & researchtinyjevA tiny jev-like model that answers Choice, Score and Noul questions in one forward pass and returns calibrated probabilities.GitHub · ★ 3 · ankit-aglaweBenchmarks & researchjev-deepresearchDeep research crawl where Jev (a System One model) makes every per-page decision and an LLM only plans and writes.GitHub · ★ 2 · tashfeenahmedBenchmarks & researchjev-x-kitOffline $0 decision layer for coding agents: Choice/Score/Noul primitives, BELKI confidence gatekeeper, ultra-planning, red-teaming, research and…GitHub · ★ 2 · KadihxBenchmarks & researchjevalyzerGrade the agent sessions already on your disk.GitHub · ★ 2 · killerz3Benchmarks & researchcan-jev-playCan Jev infer whether a bet is worthwhile from its payout table, and do recent wins or losses sway that choice.GitHub · ★ 1 · carlaiauBenchmarks & researchjevals-dataIndependent benchmark data for TypeSafe's Jev (System One model) vs LLMs: accuracy, calibration, cost.GitHub · ★ 1 · JevalsBenchmarks & researchrerank-bench-jevProduction reranker benchmark: TypeSafe Jev vs Qwen3-Reranker-0.6B on BEIR SciFact and NFCorpus — nDCG@10, cost/query and p50/p95 latency.GitHub · ★ 1 · denser-orgBenchmarks & researchsystem-one (mpuig)An open-source System One decision model — Kahneman's term for the fast, automatic judgment faculty.GitHub · ★ 1 · mpuigBenchmarks & researchyc-jev-benchCan Jev replace an LLM reranker?GitHub · ★ 1 · PPRAMANIK62Benchmarks & researchA Jev-style local model for imagesTyped decisions, but for image inputs.X post · edpleseBenchmarks & researchAn eval suite for System One modelsBuilding the tests Jev-like models need.X post · morganlintonBenchmarks & researchawesome-jev-surveyAn evidence survey of Jev and Jev-like typed decision models.GitHub · EurekaleoBenchmarks & researchBeating Jev with open modelsAn attempt to beat Jev’s accuracy, speed and cost with open models.GitHub · rob313Benchmarks & researchDecisive judgment agents that don’t chatJev against Laya.X post · thelazyyogiiBenchmarks & researchJev against Llama and DeepSeek235 ms median latency and 82% accuracy on the same labelled inputs.X post · RamV2003Benchmarks & researchJev catches the Opus 5.5 effort dropThe default effort moved from high to medium, and agents without settings lost reasoning.X post · 0xKiyoroBenchmarks & researchJev on a movie recommenderTesting Jev on recommendations.X post · utkarshgsuvBenchmarks & researchJev on Cloud RunAbout 47 s cold start and about 117 ms end to end on a cloud GPU.X post · _vmlopsBenchmarks & researchJev vs. Laya-MLXA head-to-head comparison.X post · irSodehBenchmarks & researchJev_StarA Jev and System One related repository.GitHub · sc2musaBenchmarks & researchJevBenchA reproducible benchmark for typed decision models.Site · florianstandharBenchmarks & researchLaya compared to JevAn open-source model next to the official one.X post · i_darshanjainBenchmarks & researchLaya on a 12-year-old laptop CPUSub-second decisions on an i5-5200U with no GPU.X post · Farhan60291312Benchmarks & researchLaya, 50× faster, on open alternativesA benchmark of the open-source Jev alternatives.X post · wquguruBenchmarks & researchLaya-MLX, taken apartA repo to understand an encoder-only decision architecture.X post · Zorawar PurohitBenchmarks & researchLeJudge: JEPA × JevA JEPA world-model planner with Jev as the judge.X post · AbdelStarkBenchmarks & researchLocal Laya against cloud Jev, on latencyAbout 45 ms deciding locally against about 300 ms in the cloud.X post · Oluwaphilemon1Benchmarks & researchOpen Jev PlaygroundOne place to try every open-source alternative to Jev.Site · MernitBenchmarks & researchState in, auditable decisions outA short summary of what Jev returns.X post · 0xChaseTMBenchmarks & researchTurn any LLM into a Jev-like modelA trained KV cache bank, with a live demo.Site · faangguyindiaBenchmarks & researchTypeSafe’s System One, for decisionsA recap.X post · silentguyy66Benchmarks & researchAnyJevTurn any LLM into a Jev-style decision model, with no training.GitHub · ★ 440 · nokia-applied-researchBenchmarks & researchagent-jevAgentJev-0.6B - a fast 'System One' decision model for AI Agents: feed it any unstructured state (diffs, traces, logs) and structured questions, get…GitHub · ★ 286 · malevrignsBenchmarks & researchopen-jev-typed-decision-engineOpen reproduction of TypeSafe Jev: a 150M typed decision engine (noul/choice/score in one non-autoregressive pass, calibrated confidence).GitHub · ★ 43 · intikhab49Benchmarks & researchlocal-jevA local, offline System One server compatible with TypeSafe's Jev API.GitHub · ★ 8 · amithgcBenchmarks & researchjev-2048Instrumented 2048 web lab where every move is a TypeSafe Jev Choice over the four directions with no heuristic fallback; probability, confidence,…GitHub · ★ 5 · ARCJ137442Benchmarks & researchjevbench (dhruvmehra)Benchmark TypeSafe JEV against LLMs, fine-tuned BERT, Laya and zero-shot NLI on text classification: accuracy, calibration, latency, throughput, cost.GitHub · ★ 5 · dhruvmehraBenchmarks & researchjev-botSelf-hosted Jev decision workbench and Feishu bot: automatic choices, probabilities, and experimental word/character writing.GitHub · ★ 4 · nssmdBenchmarks & researchjevloopA Python agent runtime powered by TypeSafe Jev: guarded tool execution, isolated Docker sandboxes, and side-by-side LLM comparisons.GitHub · ★ 4 · parkavenue9639Benchmarks & researchsysone-benchFirst independent head-to-head benchmark of System One decision models (Laya vs Jev) on byte-identical inputs.GitHub · ★ 4 · instax-duttaBenchmarks & researchjev-storyboard-labGoogle ADK vs Microsoft Agent Framework for structured-output agents, with TypeSafe Jev as a vendor-neutral QC gate.GitHub · ★ 3 · jimmyliaoBenchmarks & researchwerrZero-memory System-1 decision engine & TypeSafe Jev wire-compatible runtime powered by Mandelbrot wave dynamics (JevBench #1).GitHub · ★ 3 · pCwOrMBenchmarks & researchmetask-jevCalibrated typed-decision models, one forward pass each.GitHub · ★ 2 · metask-aiBenchmarks & researchask-jev-aiA public wall where anyone asks a question in three to fifteen words and Jev, TypeSafe's judgment model, answers yes, no, or it depends in about 100…GitHub · ★ 2 · waynesuttonBenchmarks & researchbes-kelime-jevNe yazarsanız yazın, beş kelimeden biriyle cevap veren sohbet botu.GitHub · ★ 2 · mahmut-gundogduBenchmarks & researchtempo-jev-demoA natural-language task workspace comparing performance across AI models (TypeSafe's Jev, GPT-5.6 Luna, and Gemini 3.8 Flash).GitHub · ★ 2 · mychaelangeloBenchmarks & researchdsh-jev-verifyDeepSeek Harness plugin that exposes TypeSafe Jev as jev_decision (choice/score/noul in one call) plus jev_verify (live labeled benchmark).GitHub · ★ 1 · xiendaBenchmarks & researchopenjev-multimodalLocal multimodal decisions on your Mac.GitHub · ★ 1 · jev-skillsBenchmarks & researchsystem-one-security-triageRecorded comparison of Jev, Terra, and Opus on 100 synthetic security-triage cases, five passes each, with a static inspectable dashboard.GitHub · ★ 1 · Robertzu43Benchmarks & researchjev-calibrateTune your Jev questions against your own labels, then confirm on held-out data.GitHub · smkrvBenchmarks & researchJevPokerBenchA Texas Hold'em benchmark for decision models, with leaderboards.GitHub · ProphetlabBenchmarks & researchA review of Jev for AI workflowsFast, structured decisions, with the vendor's own figures quoted.X post · broadrangeAIBenchmarks & researchAn agentic classification loop, 7x fasterThe loop replaced with a single Jev call, and the numbers published.Site · r6i.itBenchmarks & researchCoverage is not product valueJev judges rather than writes, and that is where its value sits.X post · David ZhuBenchmarks & researchJev against Laya, a smoke testAccuracy, calibration, latency and cost per decision, side by side.X post · cruzex100Benchmarks & researchjev-as-judgeA refund agent graded by Jev, recorded as an Opik experiment.GitHub · Akshay PachaarBenchmarks & researchJev-in-the-LoopResearch on which LLM-decision tasks Jev actually speeds up.GitHub · Tongyun1Benchmarks & researchJev-like decisions from open LLMsTurn a low-cost open model into a fast decision model, without training.X post · Anand PrasadBenchmarks & researchjev-papersA thousand arXiv papers, one Jev decision each, checked by an LLM judge.GitHub · stas4000Benchmarks & researchJevTunerTuning for Jev-style decisions.GitHub · liushiliushiBenchmarks & researchKaLM-JevA local judgement engine in Nano, Small and Large sizes.GitHub · KaLM-EmbeddingBenchmarks & researchLayaOpen weights that answer typed questions in one forward pass, 33 ms on a T4.GitHub · Nandakishor MBenchmarks & researchlaya-vs-jevLocal MLX and hosted decisions playing T-Rex side by side.GitHub · Viraj BhartiyaBenchmarks & researchLogJevText, images or audio in, choices and scores out of the logprobs.X post · MajoSayoBenchmarks & researchsolar-mini4-jevJev-style decisions on Solar Mini 4.GitHub · Sung KimBenchmarks & researchjevalsAgent evals and guardrails in one request.GitHub · ★ 82 · openlayer-aiBenchmarks & researchlaya-vs-jev-arenaTwo decision models race in Snake, then fight in an arena.GitHub · ★ 29 · PromptEngineer48Benchmarks & researchJev-QuantumA sub-microsecond System 1 model whose accuracy is a Gaussian.GitHub · ★ 28 · karminskiBenchmarks & researchopen-jevA from-first-principles rebuild of the ideas behind Jev.GitHub · ★ 21 · kyegomezBenchmarks & researchjevalWorks out what a Jev confidence score is actually worth.GitHub · ★ 16 · rlaopeBenchmarks & researchopen-spark-jevLocal System One models on Qwen3, for an NVIDIA DGX Spark.GitHub · ★ 15 · abhishek085Benchmarks & researchedgejev离线可用的本地类型化决策:4 核 CPU 单题 15.6ms。Local & offline Jev / System One inference on CPU — ONNX + INT8, no torch at runtime.GitHub · ★ 9 · yzflyBenchmarks & researchnanojev-arenaNanoJev Snake Arena: 1v4 human-vs-AI battleship + 100-agent swarm simulator.GitHub · ★ 7 · caijinchunBenchmarks & researchpoorjevOpen-source, local Jev alternative: a System One decision layer with provably calibrated confidence (ECE 0.170→0.071).GitHub · ★ 7 · rupeshpoojary9Benchmarks & researchjevtest (joshhu)情緒測謊器:嘴上說「好」,心裡真的好嗎?用 TypeSafe Jev(System One 模型)透過 OpenRouter 即時判斷,並與一般 LLM 對照.GitHub · ★ 6 · joshhuBenchmarks & researchcu-JevCuda-Jev — a CUDA-native Jev System One decision inference engine.GitHub · ★ 3 · dtunaiBenchmarks & researchjev-uiReact components that resolve which component to render, how to order a list, and whether to show an affordance — from calibrated judgments returned…GitHub · ★ 3 · etweisbergBenchmarks & researchruby_llm-providers-typesafeTypeSafe System One models (Jev) for RubyLLM: typed judgments, evaluations and reranking.GitHub · ★ 2 · javiergradicheBenchmarks & researchtypesafe-jev-mcpMCP server exposing TypeSafe's Jev model as a typed evaluate tool.GitHub · ★ 2 · anasbekheitBenchmarks & researchjev-benchmark (themsquared)Reproducible benchmark for TypeSafe AI's Jev on agent tool-call risk classification: accuracy, latency, and whether the confidence score is worth…GitHub · ★ 1 · themsquaredBenchmarks & researchwhat-is-jevIndependent, source-linked research on TypeSafe AI's Jev (System One), with 947 rubric-scored public repositories, recurring patterns, datasets, and…GitHub · ★ 1 · g0runmezadamBenchmarks & researchA correction about generating haiku with JevChoice does not generate hiragana, and the earlier claim was wrong.X post · yurinakanishi33Benchmarks & researchA Jev speed test app, made publicOpened up because the usage fees are small.X post · tubone24Benchmarks & researchAn open alternative appears: Laya421M parameters, built to decide, runs locally.X post · sl1ma4Benchmarks & researchBuild a Jev JudgeAkshay Pachaar’s worked example: evaluate a refund-support agent with Jev instead of a generative judge, and record the verdicts in Opik.X postBenchmarks & researchCalibrated confidence in 70 to 500 msYes/no, choices and scores, with the published price per token.X post · shree_codeBenchmarks & researchChatJevDriving the classifier as a next-token predictor, one token at a time.GitHub · Erik DuntemanBenchmarks & researchDecisions inside the system, not in a chatWhat makes Jev different from a conversational model.X post · xooox888Benchmarks & researchEarly access since September 15A System One model for agent decision loops, in typed options.X post · stretchcloudBenchmarks & researchfedjev-benchDoes Jev hear a rate rise coming in an FOMC statement?GitHub · maybern-tripp-smithBenchmarks & researchFive dollars a month was not enoughReal tasks forced reading the skills docs and using state properly.X post · zhengliBenchmarks & researchIs a quick classifier worth building?The traps that come with building one fast.X post · satto_sannBenchmarks & researchJev against an LLM on your own feedThe same tweets classified both ways, open on GitHub.X post · tshmieldevBenchmarks & researchjev-acentoA pre-registered audit of how Jev handles Spanish.GitHub · Marcos MartinezBenchmarks & researchjevchatA chatbot built out of a model that only answers typed questions.GitHub · Kyle PenaBenchmarks & researchJevGuardA decision runtime with zero-token caching and a calibrator.GitHub · Seb4EzBenchmarks & researchLaya, 421M parameters, local86.5 decisions a second on about 1 GB of inference memory.X post · Md FazalBenchmarks & researchminojevIndependent Jev-style replica with head training: a frozen Qwen3-1.7B backbone plus a small trained decision head returns calibrated…GitHub · zeredy879Benchmarks & researchOpenJevAn open attempt at a Jev-class decision model.GitHub · SiliconLabAIBenchmarks & researchThe first public System One modelState and explicit questions in; choices, scores and probabilities out.X post · yuntian9999Benchmarks & researchNanoJevA 0.6B replica: parallel decisions in one forward pass, no decoding.GitHub · ★ 2.2k · TianyuCodingsBenchmarks & researchopenJev-verdict-2.0A 151M non-autoregressive decision engine, with its numbers published.GitHub · ★ 283 · Heman10x-NGUBenchmarks & researchjevbenchJevBench v1 - a benchmark for Jev-class typed decision models: smart, cheap, fast, reliable, open.GitHub · ★ 114 · fstandhartingerBenchmarks & researchJevForgeEnd-to-end Jev-style structured-decision stack for auditable data construction, Qwen3.5-0.8B training, fixed Mind2Web and OOD evaluation, preliminary…GitHub · ★ 35 · zwliJayBenchmarks & researchjev-rag-benchmarkReproducible benchmark for measuring Jev reranking quality, latency, and cost in RAG.GitHub · ★ 14 · erendikmennBenchmarks & researchjev-architectHelps you find, design and evaluate a Jev decision loop.GitHub · ★ 6 · karanb192Benchmarks & researchjev-ood-calibrationIndependent calibration study of TypeSafe Jev: published raw responses for three public benchmarks plus 900 rule-generated support tickets (choice /…GitHub · ★ 6 · scienthoonBenchmarks & researchOpenJevOpenJev: an independent Jev-inspired System One decision API based on TypeSafe.ai concepts.GitHub · ★ 5 · xingwudaoBenchmarks & researchjev-reward-model-evaluationJev 1.13 reward-model evaluation across 8 benchmark tracks, with an interactive report and 54-row SOTA comparison.GitHub · ★ 4 · goya4140Benchmarks & researchjevarena (chenmingtang830)Open-source BYOK arena for Jev and other AI judges.GitHub · ★ 4 · chenmingtang830Benchmarks & researchOpenJev (GPT-AGI)Jev-compatible System 开源Jev.GitHub · ★ 4 · GPT-AGIBenchmarks & researchJevUnofficial TypeSafe Jev showcase — System One decisions, not chat.GitHub · ★ 3 · cobusgreylingBenchmarks & researchjev-web-analyzerSee what Jev thinks about your SaaS website — powered by ReplyNodes web context and Vercel AI Gateway.GitHub · ★ 3 · replynodesBenchmarks & researchjevtestSemantic test matchers for Vitest and Jest, powered by TypeSafe's Jev model.GitHub · ★ 3 · realZachiBenchmarks & researchORIGIN-CIVILIZATIONAI life-and-civilization simulation: TypeSafe Jev makes every decision (typed, probabilistic, auditable); LLMs plan — OpenAI-compatible APIs, local…GitHub · ★ 3 · JacquesGariepyBenchmarks & researchjev-cloud-quiz三大クラウドの機能名を、TypeSafe AI の System One モデル Jev が確率つきで判定するデモ.GitHub · ★ 2 · minorun365Benchmarks & researchjev-labsNever confidently wrong: a TLA+-verified consensus kernel around TypeSafe's Jev, run through 1,680 chaos-tested pharmacy decisions with zero wrong…GitHub · ★ 1 · copyleftdevBenchmarks & researchopenpoke-meets-jevFork of OpenPoke where the yes/no decisions go to Jev instead of Claude Sonnet, with a same-inputs A/B against the replaced LLM decision and a…GitHub · ★ 1 · 0xShin0221Benchmarks & researchjev-e2eThe same eBay test flow: 47 s on Jev, 79 s on Claude Sonnet 5.GitHub · Jason LuBenchmarks & researchjev-evaluationNine experiments and 28 predictions, all fixed before any data.GitHub · Will KellyBenchmarks & researchWebMCP browser benchmarkA reported 49-task benchmark combines Jev for tool selection, Mercury 2.5 for arguments and WebMCP for browser actions.Site · idan levinBenchmarks & researchTop 5 Jev Alternatives: __ I have been running a…Top 5 Jev Alternatives: __ I have been running a benchmark since this morning on every Jev alternative against Jev.X post · ItsCuthulhuBenchmarks & researchvonA sub-15 ms local drop-in alternative to Jev.GitHub · ★ 610 · wfzyxBenchmarks & researchjev-align (Sutro)Turns human labels into calibrated Jev classifiers with GEPA.GitHub · ★ 284 · sutro-shBenchmarks & researchjev_localReplicating Jev with a local LLM.GitHub · ★ 35 · Argos1111Benchmarks & researchjev-capability-atlasWhere Jev holds up and where it breaks, with real API receipts.GitHub · ★ 26 · ZaiousBenchmarks & researchjev.nuNushell module for the TypeSafe System One API: typed decisions with calibrated probabilities.GitHub · ★ 8 · cableheadBenchmarks & researchjevscapeRuneBench harness for TypeSafe's Jev: bounded action catalog, tick-mode controller and a live dashboard.GitHub · ★ 8 · Skyvern-AIBenchmarks & researchjev-alignCalibrated alignment verifier for LLM responses and agent plans — powered by Jev.GitHub · ★ 4 · caiovicentinoBenchmarks & researchnew-api-plugin-typesafeTypeSafe AI System One (Jev) task plugin for QuantumNous/new-api — native /v1/systemone, synchronous evaluation, token billing.GitHub · ★ 3 · FFatTigerBenchmarks & researchjev-chessChess moves, evaluations, persona opponents, and game classification with TypeSafe AI System One.GitHub · ★ 2 · hemanthBenchmarks & researchjev-modeI kept watching coding agents burn context on decisions that aren't hard - triage 400 tickets, tag 600 files, route to one of six teams.GitHub · ★ 2 · ddfeyesBenchmarks & researchresearch_deskTypeSafe Jev demonstration for new analyzation — experimenting with Jev for fast analysis of news and tickers.GitHub · ★ 2 · 0xnairbBenchmarks & researchzerosweepAutonomous System-One Triage Engine & Benchmark powered by TypeSafe AI (Jev).GitHub · ★ 2 · sysadarshBenchmarks & researchjev-carryforwardWhat your last session knew, scored against what this one is doing.GitHub · ★ 1 · Dharundp6Benchmarks & researchmodelsystemCurated catalog of System One / Decision Models — contributions for modelsystem.one.GitHub · ★ 1 · fabricioctellesBenchmarks & researchtypesafe-vs-deepseekTypeSafe (Jev) vs DeepSeek-flash: side-by-side speed/token/cost/accuracy comparison across invoice extraction, email classification, and reranking.GitHub · ★ 1 · markfive-protoBenchmarks & researchkevA tiny Jev-like model on Qwen2.5-0.5B that trains and runs on a MacBook.GitHub · Jared PalmerBenchmarks & researchOpenRouter’s Ori EvalJev judged a 30-way labelling task faster than every model tested.X post · OpenRouterBenchmarks & researchdeepseek-v4.1-flash-jevAn open model made to answer like Jev with a scoring endpoint.X post · Nick KhamiBenchmarks & researchJev against Fable 5.1, 100 buildersThe same post scheduler, built twice, in front of an audience.X post · Florian DarromanBenchmarks & researchopenjevOpen, Jev-compatible System One decision server on DiffusionGemma.GitHub · ★ 386 · razorback16Benchmarks & researchawesome-jev (OmniJev)Papers, open reproductions and independent evaluations behind System One models and Jev.GitHub · ★ 286 · OmniJevBenchmarks & researchjev-visualAn educational Jev-like visual inference experiment on Apple Silicon: shared context, direct candidate scoring, and local visual demos.GitHub · ★ 271 · hr98wBenchmarks & researchreflexA small open decision model: state and typed questions in, probabilities out.GitHub · ★ 133 · kshetrajna12Benchmarks & researchVerdict-open-jevAn open 151M decision model on ModernBERT, with a WebGPU playground.GitHub · ★ 94 · Heman10x-NGUBenchmarks & researchopen-alternative-jevTyped, calibrated decisions from any open-weights model, on your own GPU.GitHub · ★ 53 · ikermoelBenchmarks & researchmini-jevMini-Jev: what a Jev-style typed-decision interface looks like on a frozen Qwen3-4B — read the option letter's logits instead of generating JSON.GitHub · ★ 53 · r-msBenchmarks & researchLitJevA reproduction of Jev that turns any Qwen model into a fast decision model, serving the same /v1/systemone schema (Choice, Score, Noul) with no…GitHub · ★ 43 · zhengxuyuBenchmarks & researchJev_apps看看 Jev 能做什么:用中英文讲清热门应用、工作原理和各自优缺点。Explore Jev apps with plain-language examples, explanations, and comparisons.GitHub · ★ 33 · JackZengBenchmarks & researchjevifyAn agent skill to discover TypeSafe Jev opportunities, design typed questions, and learn from recent community experiments.GitHub · ★ 27 · altryneBenchmarks & researchtypesafe-localInspired by TypeSafe Ai, Ask a local LLM typed questions, get calibrated probabilities instead of text.GitHub · ★ 8 · aabolfazlBenchmarks & researchjev-search-rerank-evalDoes a TypeSafe Jev rerank beat embedding search?GitHub · ★ 6 · zhuyansenBenchmarks & researchtenbinSplit a judgment into Choice, Score and Noul, lint it, then measure it.GitHub · ★ 3 · simotaBenchmarks & researchtypesafe-ai-playgroundCommunity TypeSafe AI playground: 110 use cases, games, dilemmas and model challenges, with editable prompts, A/B comparisons and a mobile-friendly…GitHub · ★ 3 · nickthompson480Benchmarks & researchaskjev.aiAsk Jev anything. It won’t answer, it will judge.Site · Wayne SuttonBenchmarks & researchFive open Jev replicas, in Chinese小墨同学 on Laya 421M, Decider-2B, NanoJev 0.6B, Reflex and System-One 4B, and which run best on Apple silicon.X postBenchmarks & researchJev on the WebMCP benchmarkJev picks the tool, a fast LLM fills the arguments: 49 of 49 tasks solved.X post · Idan LevinBenchmarks & researchOne judge call vs twelve dimension scoresOne direct Jev question per row against 12–14 Jev-scored dimensions with locally fitted weights on three classification tasks: 5,477 test rows,…SiteBenchmarks & researchopenjev on Qwen 4BAn MLP trained on top of Qwen 4B that works like Jev.Site · Alex NikolicBenchmarks & researchVerdict (Open-jev) launch threadHemant Kumar on his 151M ModernBERT reproduction: 0.83% adaptive calibration error, 3.0% of predictions flip when the option list is reversed, 35.6ms…X postBenchmarks & researchSemIfSemantic ifs from open models, on a 3090 at home.GitHub · ★ 4.2k · TheoLeeCJBenchmarks & researchJevlikeExplore training a small option-scoring model.GitHub · ★ 1.3k · vinnylarougeBenchmarks & researchdeciderOne-pass typed decisions with calibrated probabilities (System One style model), fine-tuned from Qwen3.5-2B.GitHub · ★ 351 · MapikaBenchmarks & researchopen-jev (daseinlabs)One-pass option scoring with a local Gemma 3 4B on Apple silicon via MLX, inspired by jevlike, with a Doom demo.GitHub · ★ 107 · daseinlabsBenchmarks & researchjev-eval-agentPersonal-assistant agent built on Vercel's eve with 100 mocked tools, measuring how many steps it takes when Jev picks the tool versus the LLM.GitHub · ★ 105 · vinilanaBenchmarks & researchjev-as-a-judgeUsing Jev as an evaluator.GitHub · ★ 83 · danielgsheaBenchmarks & researchjevfireJEV-inspired parallel decisions for CUDA LLMs.GitHub · ★ 64 · kikoncuoBenchmarks & researchjevmlxJev-style parallel constrained decisions for any MLX model on Apple Silicon.GitHub · ★ 60 · bnsd55Benchmarks & researchTypeARType-safe one-decision-per-token decoding engine for autoregressive LLMs, inspired by Jev.GitHub · ★ 40 · TypeLLMBenchmarks & researchopen-jev (JoshuaSP)Typed JSON inference with DiffusionGemma, with Every and Jev benchmark results.GitHub · ★ 39 · JoshuaSPBenchmarks & researchtypesafe-ai-benchmarkThis is a LLM Gateway that mimics typesafe ai structured output.GitHub · ★ 38 · iammrduncanBenchmarks & researchsystem-one-openOpen replica of TypeSafe's Jev: typed calibrated decisions in one forward pass, on Gemma 4 E2B / Gemma 3 270M (Modal).GitHub · ★ 35 · mithalouniBenchmarks & researchopenjev (zhihz)Local bilingual probability decisions from context, questions, and candidate answers.GitHub · ★ 33 · zhihzBenchmarks & researchsystem-oneBatched single-token choice inference for open language models, compatible with TypeSafe.GitHub · ★ 33 · sgoedeckeBenchmarks & researchjevgptA chatbot built on a model that cannot generate text (TypeSafe AI's Jev, driven autoregressively).GitHub · ★ 26 · BewinxedBenchmarks & researchJev on a LaptopInvestigate typed decisions on local hardware.GitHub · ★ 25 · rorshoppingBenchmarks & researchjev-column-raceLocal (or hosted BYOK) race UI that labels 1,000 withheld-star app reviews in parallel columns: TypeSafe Jev typed questions versus a Gemini JSON…GitHub · ★ 23 · goodrahstarBenchmarks & researchtypesafe-ai-playground (BunsDev)Community TypeSafe AI playground: 110 use cases, games, dilemmas and model challenges, with editable prompts, A/B comparisons and a mobile-friendly…GitHub · ★ 20 · TypeSafeAIBenchmarks & researchJev BenchmarksMeasure calibration, latency, and selective risk.GitHub · ★ 18 · AbdelStarkBenchmarks & researchjevbetterA stronger one-pass scorer over a variable list of text options: hashed n-gram encoder, rival-aware attention, gated head, temperature scaling,…GitHub · ★ 14 · olanotoluBenchmarks & researchopenvonsOpenvons (open-Jev): 有限選択肢に確率で答える判断層 — テキスト / 画像 / 日本語音声コマンド.GitHub · ★ 13 · genai-craftBenchmarks & researchRerank BenchCompare Jev with document reranking models.GitHub · ★ 8 · anessbelbatiBenchmarks & researchjev-mcp (arunav25)Connect JEV to MCP clients and compare its judgments against general-purpose LLMs using shared datasets and measurable accuracy.GitHub · ★ 7 · arunav25Benchmarks & researchPhishing BenchCompare phishing judgments across models.GitHub · ★ 6 · anisselbdBenchmarks & researchjev-benchmarkBenchmarks and a playground for TypeSafe's Jev (System One) model: chess, and who-is-the-player-talking-to for speech-to-text game NPCs.GitHub · ★ 6 · wondertwinsBenchmarks & researchjev-korean-benchmarkReproducible early-access evaluation of Jev on Korean understanding and medical text, with runtime and cost evidence.GitHub · ★ 6 · mahlernimBenchmarks & researchjev-lmA word-level language model whose output layer is Jev: n-gram drafter, Noul chunk verification, bits-per-token eval.GitHub · ★ 6 · y0usafBenchmarks & researchmcts-agentDiscriminative Monte Carlo Tree Search using TypeSafe Jev System One Primitives and Gemini.GitHub · ★ 6 · lhemerlyBenchmarks & researchai-elo-rankerHigh-speed recursive AI Elo tournament engine powered by Jev and Swiss matchmaking.GitHub · ★ 5 · opaielsheikhBenchmarks & researchjev-little-airwaysA show-and-tell capability study for Jev, TypeSafe's System One decision model.GitHub · ★ 5 · lbotinellyBenchmarks & researchLegalForecastBenchLegalForecast-MTD benchmark alpha and official evaluation workflows.GitHub · ★ 5 · johnhughes3Benchmarks & researchSecurity BenchEvaluate injection and code-security detection.GitHub · ★ 4 · Gaurav-GosainBenchmarks & researchjev-chatA chatbot from typed Jev decisions: hierarchical speculative decoding over System One probabilities.GitHub · ★ 4 · adhyaay-karnwalBenchmarks & researchqwen-rlcdJev-style calibrated decision model (Choice/Score/Noul) on Qwen3.5-0.8B.GitHub · ★ 4 · shamazharikhBenchmarks & researchsystem-one-gemmaOpen-source Jev-style System One decision model.GitHub · ★ 4 · akash-kamatBenchmarks & researchAgent Failure BenchmarkTest attribution of failures in agent traces.GitHub · ★ 3 · TokenTrimBenchmarks & researchSpam EvalTest spam detection against simpler baselines.GitHub · ★ 3 · bitnovusBenchmarks & researchRouting ExperimentEvaluate Jev as a model-selection layer.GitHub · ★ 3 · TokenTrimBenchmarks & researchjev-behavior-studyIndependent Jev 1.13.0 behavior study: report, controlled prompt experiments, raw results, and offline verification.GitHub · ★ 3 · RINNECODERBenchmarks & researchjev-explorationJev (TypeSafe) exploratory thread: claim audit, live demos, and runnable code.GitHub · ★ 3 · SamuelSaccoBenchmarks & researchCalibreStudy confidence thresholds on banking intents.GitHub · ★ 3 · FirasSX914Benchmarks & researchclarity-judgeWriting checked on separate named axes, each with its own verdict.GitHub · ★ 2 · TypeSafeAIBenchmarks & researchDecisionBridgePut a decision interface over existing LLMs.GitHub · ★ 2 · grishahqBenchmarks & researchcalibreCalibration and confidence-based routing measured on Banking77: 80.2% accuracy at $0.103 per 500 decisions.GitHub · ★ 2 · FirasSX914Benchmarks & researchjev-freeformAn observable raw-character chat experiment powered entirely by TypeSafe Jev Choice.GitHub · ★ 2 · keskuBenchmarks & researchjev-research-evalReproducible Jev Ultrafast research-browser eval harness + field note (QC’d cases, suite runner, report generator).GitHub · ★ 2 · jgridifierBenchmarks & researchtypesafe-ai-playground (markjaquith)A playground for experiments around Jev, TypeSafe's System One model.GitHub · ★ 2 · markjaquithBenchmarks & researchgpt-vs-jevCompare GPT generated language with JEV structured Noul decisions on the same input.GitHub · ★ 1 · TanayPadarBenchmarks & researchjev-secret-detectionMeasures how well TypeSafe's RLCD-Jev model spots real secret credentials in file snippets.GitHub · ★ 1 · teyhouseBenchmarks & researchjev-synergy-screeningJev (TypeSafe System One) × ASReview SYNERGY abstract screening demo — Choice/Noul vs gold labels.GitHub · ★ 1 · PistachioAIHQBenchmarks & researchkyotsu-ai-benchAI benchmark on Japan's 2026 Common Test: Jev vs luna-none vs luna-low (static dashboard).GitHub · ★ 1 · shibadogcapBenchmarks & researchopenjev-experimentsExperiments with openjev, an open Jev-style option-logit runner, on local models.GitHub · ★ 1 · zefir1990Benchmarks & researchpadflow-jev-evalsTyped-decision benchmark from PadFlow (land development SaaS): schemas, anonymized labeled rows, and a runner for confidence-calibrated models like…GitHub · ★ 1 · zsavage8Benchmarks & researchPocketJevOn-device iPhone visual decision tool using MLX and Qwen3-VL direct option logits.GitHub · ★ 1 · NullPo-jpBenchmarks & researchRISC-jeVI tortured Jev into being a RISC-V CPU.GitHub · ★ 1 · i2cjakBenchmarks & researchevery.to37 documents, 21 questions each, 1,709 judgments for under a cent.Site · Mike TaylorBenchmarks & researchInternal classifier field noteMatched-precision comparison against a private fine-tuned classifier.X post · identityBenchmarks & researchnearhere.eventsHead-to-head test at validating local event listings, with cost and latency.Site · JonBenchmarks & researchJev vs Qwen on CerebrasVideo comparison against a structured-output LLM baseline.X post · IamMrDuncanBenchmarks & researchA local JevA Jev-style model running locally, with room to get faster.X post · 生ビールBenchmarks & researchFinancialPredictionJevUsing Jev to test how well it predicts financial markets(just like most llms as of september 2026, it doesnt do that good).GitHub · thodoh1Benchmarks & researchJev plays chessJev picks from the legal moves, compared with reasoning models.Guide · Maxim SaplinBenchmarks & researchjev-alpha-benchDoes Jev predict stock returns from news?GitHub · Gaurav-GosainBenchmarks & researchjev-anotacao-sentencasJev (TypeSafe) vs. Gemini 3.8 Flash vs.GitHub · lab-dadosBenchmarks & researchjev-deferred-crispificationPosition paper: the Hidden-Markov and fuzzy primitives missing from TypeSafe AI's Jev and System-One decision models.GitHub · dnakhoaBenchmarks & researchjev-demosDemos to test the effectiveness of TypeSafe's "Jev" System One Model.GitHub · Bud-roBenchmarks & researchjev-dev同じ発言を jev と LLM の両方に判定させ、感情の変動値のズレと応答速度を1画面で見比べるデモ(affectus + Vercel AI Gateway).GitHub · n-yokomachiBenchmarks & researchjev-headline-benchCan Jev pick the winner of a real headline A/B test?GitHub · Gaurav-GosainBenchmarks & researchjev-jp-addressJev (TypeSafe) 性能評価プロジェクト — 日本郵便 KEN_ALL をマスタに、AI SDK 経由の Jev が住所のあいまい一致にどこまで使えるかを検証.GitHub · smasatoBenchmarks & researchjev-labTypeScript experiments, evaluations, and latency benchmarks for TypeSafe's Jev model.GitHub · Menny1337Benchmarks & researchjev-pick-and-place-studyA small reproducible MuJoCo pilot comparing Jev, Claude Haiku, and reactive rules for pick-and-place.GitHub · tryakshBenchmarks & researchjev-playground (hegargarcia)Benchmarks Jev against other evaluation models in games with explicit states, legal actions, and measurable outcomes.GitHub · hegargarciaBenchmarks & researchjev-report发明 RLHF 的人,这次做了个不会说话的模型:Jev 独立研究报告。52 页 PDF + 50 条中文实测复现包 + 143 条可回溯数据表.GitHub · HackSingBenchmarks & researchjev-shadcn-lint-evalA small second eval for shadcn-ui/lint that uses TypeSafe's Jev to judge the linter's own output.GitHub · blas0Benchmarks & researchjev-trace-classifierApplication of TypeSafe Jev (noul judgment primitive) on the collusion.wiki corpus: agent vs human page authorship, head-to-head vs local…GitHub · sypherinBenchmarks & researchjev-user-juryTypesafe/Jev public X demo.GitHub · cardotrejosBenchmarks & researchmisereru-slide-jevOngoing Japanese research deck on Jev and System One models, maintained as Markdown slides.GitHub · myokoymBenchmarks & researchParallel Constrained Decoding (Qwen2.5-1B-RLCD)Hugging Face Space exploring open-source parallel constrained decoding as an alternative to Jev.SiteBenchmarks & researchshade-arena-jev-monitorEvaluating TypeSafe's Jev as a fast monitor and action gate for agent sabotage in SHADE-Arena, compared with Gemini 2.5 Flash/Pro.GitHub · nican2018Benchmarks & researchsystem-one-adapter-rustRust port of TypeSafe system-one-adapter (LLM-backed system_one evaluations).GitHub · codeitlikemileyBenchmarks & researchthaiexam-jev-chartsCharts: TypeSafe Jev evaluated on Thai standardized exams vs 110 other models.GitHub · vehasBenchmarks & researchTyped decisions, not chatIndependent walkthrough separating TypeSafe's published claims from public evidence.SiteBenchmarks & researchtypesafe-oraclesEvaluating TypeSafe's System One primitives (Choice/Score/Noul) — where a typed oracle beats an LLM call.GitHub · trophee-botBenchmarks & researchTypesafe_chess_evalAn evaluation of typesafe AI chess.GitHub · AliceRoseliaBenchmarks & researchOpenJevAn independent experiment in local decision models.GitHub · ★ 4.3k · TheoLeeCJBenchmarks & research50 financial jobs, one modelThe Fintech Builder runs Jev on fifty financial jobs and shows where it works and where it does not.YouTube · youtube.com/@TheFintechBuildBenchmarks & researchatisbo.dev> We benchmarked TypeSafe's Jev against our production LLM classifier — with real customer data and human ground truth.Site · Brian OchoaBenchmarks & researchagent-evalsDeterministic LLM-agent eval harness: rule scorers for tool/latency/cost failures, plus an optional calibrated TypeSafe Jev judge you can gate CI…GitHub · marianobertonBenchmarks & researchfieldnote.clawscience.comCan Jev reason across mechanics, relativity and gravitation?Site · CursedCoderBenchmarks & researchchinese-workflow-decision-benchReusable Feishu-style workplace message classification benchmark: 64 frozen synthetic Chinese scenarios, Choice and four-Noul workflows, published…GitHub · Adkid-ZephyrBenchmarks & researchDoes an open-weight decision model beat a hosted one?ASTGL compares local Laya with hosted Jev on typed decisions.GuideBenchmarks & researchciv6-mcp.lwilko.comExcited to enter Jev into my Civ 6 benchmark :blob_aww: Has anyone else tested out strategic reasoning yet?Site · Magykman (Liam)Benchmarks & researchnola-typesafe-testFinally got access, I made a quick comparison between Cerebras GPT-OSS-120B and JevGitHub · Evgen MykhailenkoBenchmarks & researchdecisionmodels.aiGot a bit obsessed with the space and built decisionmodels.ai tonight.Site · sahilBenchmarks & researchgit.sub-net.athere a short Jev test for NPC decisions in our slop MMO on a slop paper - its awesome, sadly the latency from the EU to the Jev test infra kills it…Site · SubBenchmarks & researchcultivarHey! We at Pinecone added TypeSafe support to Cultivar, our library for evaluating and benchmarking agent skills and docs against agents in sandboxes.GitHub · arjunBenchmarks & researchi build a model x45 smaler but it Beats GPT-2 124M BYTE_BPC = 1.142985 (TOKEN_BPC = 9.152823) on the same metric https://huggingface.co/Uuuuuuniiiiiii/goldworm-v6 as well as my agencyi build a model x45 smaler but it Beats GPT-2 124M BYTE_BPC = 1.142985 (TOKEN_BPC = 9.152823) on the same metric as well as my agencySite · DARK(uniency.com)SudocornBenchmarks & researchyuwakisa.comI have a benchmark that tests models for "here's a set of statements, express the common principle behind them" and the model expresses a principle.Site · hypnoticfuzzwaveBenchmarks & researchjev-test-1I made a deliberately tiny, dependency-free Jev test showing state, typed questions, and probabilities.GitHub · sethackedBenchmarks & researchI tested Jev on 12 real use casesNate Herk compares speed and cost against ordinary models, then builds an X feed classifier and a real-time paper-trading prototype live.YouTube · youtube.com/@nateherkBenchmarks & researchIs Jev really better, faster and cheaper?Edward Donner puts Jev to the test against Luna.YouTube · youtube.com/@EdwardDonnerBenchmarks & researchjev-guardbenchBenchmark harness asking whether a System One model (hosted TypeSafe Jev or self-hosted open-source Kev) can stand in for LLM-as-judge in agent…GitHub · dfranco-projectsBenchmarks & researchjev-healthcare-labOpen experiment comparing TypeSafe Jev vs DeepSeek on 96 healthcare tasks across 12 scenarios (quality, latency, cost) with per-task data and…GitHub · JuneYaoooBenchmarks & researchJev-OmniJev-Omni is an open multimodal System One–style decision classifier: give it a state, a question, and typed options over text, image, audio, or…Site · akhilaaa3Benchmarks & researchjev-packsEvidence-gated registry of Jev question packs: curated questions, golden cases, pinned model versions, and reproducible offline benchmark results…GitHub · dtduc-gitBenchmarks & researchjev-robotics-evalEvaluation harness for robot control with a JEV-compatible decision service on MetaWorld and RoboTwin tasks (text and vision modes, privilege levels,…GitHub · lose4578Benchmarks & researchjev-security-prioritizationReproducible experiment: TypeSafe Jev prioritizes SCA/SAST findings using richer context than severity scores, compared to severity-only and points…GitHub · san3ncrypt3dBenchmarks & researchjev-skill-router-benchIndependent, reproducible measurement of a Jev (TypeSafe System One) skill router on one 84-skill Hermes roster: 81 author-labelled turns, router…GitHub · OrMizLBenchmarks & researchjevalsLocal evaluation workbench for TypeSafe Jev: author Noul/Choice/Score (and combined) questions with expected answers, run them, and compare saved…GitHub · dayhaysoosBenchmarks & researchJevals.comHosted independent benchmark boards: TypeSafe Jev (typesafe-ai/jev via Vercel AI Gateway) vs six LLMs on PubMedQA (Noul), Banking77 (Choice), and…Site · JevalsBenchmarks & researchjevlensEvaluate, calibrate, replay, and monitor TypeSafe Jev decisions from labeled CSV/JSONL: store full answer distributions, report accuracy/Brier/F1,…GitHub · k4its1tBenchmarks & researchJevScopeLocal-first visual decision workbench and regression testbench for TypeSafe Jev: edit structured state and questions, inspect distributions, batch…GitHub · jeiel85Benchmarks & researchjevtrimBenchmark: TypeSafe Jev as a context-relevance judge vs retrieval/summarization on LoCoMo, with matched token budgets and reproducible reports.GitHub · pdrpintoBenchmarks & researchMalkuthOpen-weight multilingual System One–style decision models (2B/4B, Korean focus): Choice, Noul, and Score over /v1/systemone via Kev—adjacent to…GitHub · newfull5Benchmarks & researchOpikComet’s open-source evaluation and tracing platform, Apache-2.0.GitHubBenchmarks & researchPapers and open reproductionsOmniJev’s list of the papers, open reproductions and independent evaluations behind System One models.GitHubBenchmarks & researchPDF RaceTimed document race: Docling → TypeSafe Jev vs Docling → Gemini vs Gemini reading the PDF, scored against arXiv metadata, with committed recorded…GitHub · goodrahstarBenchmarks & researchposted the article, thanks for the ok.<@81830765507117056> posted the article, thanks for the ok.Site · nssBenchmarks & researchRYOTIDERoll Your Own Typed Inference Decision Engine: Jev-style typed decisions from a single forward pass of a local LLM (MLX or PyTorch), measured on…GitHub · csabagBenchmarks & researchSemIfSemIf studies runtime-defined semantic decisions using local open models.GitHub · TheoLeeCJBenchmarks & researchlindfors.noTested jev against deepseek v4.1 flash on norwegian text.Site · depuratorBenchmarks & researchtypesafe-offload-benchTested Jev with main "Players" (GPT/Claude) to work together.GitHub · DmytroBenchmarks & researchTyped EvalsEvaluate LLM responses, RAG samples, and agent traces with TypeSafe Jev as the judge, including optional calibration against human labels and…GitHub · TrustifAIBenchmarks & researchdinostompwired jev into dinostomp, an oss eval harness: one choice question per item through /v1/systemone, probabilities on every record, and a blind control…GitHub · hobohotdogBenchmarks & research⚡ Jev is HERE and it's rejecting ChatGPT's entire training playbook ♠ built by a co-creator of RLHF/ChatGPT who says⚡ Jev is HERE and it's rejecting ChatGPT's entire training playbook ♠ built by a co-creator of RLHF/ChatGPT who says chat was the wrong bet 🔹New…YouTube · fahdmirza

Other use cases

Agents & browsers 193Coding & code review 227Routing & model choice 161Context & memory 66Search & RAG 56Data & extraction 61Documents & research 28Safety & guardrails 131Support & inbox 45Marketing, sales & social 95Trading & finance 46Games & real time 173Creative & generative UI 40Home, robots & devices 20Everyday apps 113SDKs & integrations 308Guides & docs 202