Papers, open models and evaluations behind System One models and Jev.
⚡ System One & Jev · 🧪 Open Source · 🔧 Built with Jev · 📊 Benchmarks · 📰 Commentary · 🧬 Lineage
- 🔥 News
- ⚡ System One & Jev (8)
- 🧪 Open Source (44)
- 🔧 Built with Jev (37)
- 📊 Benchmark & Leaderboard (36)
- 📰 Commentary & Analysis (8)
- 🧬 The Shape Before Jev (15)
- 🧱 What Jev Is Sold Against (9)
- 🧠 Where the Name Comes From (5)
- 🔗 Related Lists (3)
🆕 2026-09-29. We added a Benchmark & Leaderboard section and a papers-only list, Awesome JEV Papers. PRs welcome.
🆕 2026-09-22. We added 14 projects and evaluations.
What TypeSafe has published: the launch post, the contract, the failure modes it admits to, and the code it ships.
- Introducing System One Models and Jev, The launch post: state in, typed probabilistic decisions out, RLCD training, 70 to 500 ms, $0.042 per MTok.
- Primitives: Choice, Score, Noul, The three typed question shapes and the probability-per-option answers they return.
- Jev 1.13 jaggedness, TypeSafe's documented failure modes: literal reading, counting, dates, indirection, distractor state, adversarial content.
- System One Adapter, Official drop-in that serves the same typed interface from OpenAI or Anthropic models, the baseline for every comparison.
- Hacker News launch thread, 1,850 points and 485 comments; the CEO confirms the zero-shot classifier reading and the encoder-with-heads shape.
- TypeSafe Agent Skills, Skill files that teach Claude Code, Codex and similar agents to design System One workflows.
- Jev on Vercel AI SDK, The first third-party surface: Jev-latest as an evaluation model behind experimental_evaluate, no waitlist.
- Founder launch thread on X, Diogo Almeida's thread arguing RLCD decision models reach economic value before chat models do.
Open weights and code that rebuild the System One shape from encoders, small decoders and constrained decoding.
- OneJev, Open multimodal System One model in four sizes (0.8B to 27B): typed questions about a screenshot, photo, video or text get a calibrated probability for every option in one forward pass.
- Rizzo Flow, Local Jev-compatible decisions on Spark-X2.5 through llama.cpp; authors report 49 ms p50 for short Q8_0 decisions on an RTX 5060 Ti, with uncalibrated probabilities by default.
- Open-Jev (ZefanCai), Released Qwen3.5-2B and 9B LoRA adapters with scalar decision heads and a public dataset; the 9B model answers 179 of 231 public JevBench tasks correctly in the authors' evaluation.
- this-that-model, 1.88B fine-tune of decider-2b whose head scores only the declared option labels, so an answer outside the set cannot occur; authors report 0.775 on their released 1,710-question decision benchmark.
- Dohnuts, 0.8B text-and-image decision model trained on one RX 7900 XTX; authors report 65.8% accuracy on 231 public JevBench tasks; Apache-2.0 code and non-commercial CC BY-NC-SA 4.0 weights.
- OpenThai-SystemOne, Thai and English 0.8B decision model with a 256-way slot head and Jev-compatible API; authors report 90.0% accuracy on 5,007 held-out MASSIVE Thai intent examples.
- sokudan, Japanese 314.6M ModernBERT-ja decision model returning choice, score and bool probabilities in one forward pass; authors report 0.880 choice accuracy and 0.844 bool AUROC (v0.2) on their released 300-item bench_ja; Apache-2.0 code and weights.
- open-jev (browser runtime), Browser TypeScript runtime for Kev and DeBERTa decision models through Transformers.js, returning typed probabilities with WebGPU or WebAssembly and no cloud inference.
- djev, DiffusionGemma structured-read server that pins fixed answer text and denoises answer slots into per-option probabilities through Jev's API; requires the linked vLLM PR branch.
- Laya Ultrafast, Local MLX port of Jev Ultrafast with a redesigned narrow-decision policy; authors report five successful Google Flights runs in 7.5 to 12.1 seconds on an M1 Max.
- SemIf, Semantic ifs from open models on a 3090 at home; open baseline for direct typed option scoring.
- Jevlike, From-scratch model with Jev's exact shape: text plus N options in, one probability per option out, option-attention head, Doom and chess demos.
- SLEEPJEV (Sleep-JEV), PSG-native local decision model with one reusable overnight representation, sparse retrieval, and runtime Choice/Noul/Score probabilities over fixed candidate sets; includes a real SHHS-derived replay demo.
- MEDJEV, Independent JEV-inspired biomedical evidence decision implementation with runtime candidate semantics, Choice/Score/Noul-style outputs, calibration metadata, and a PubMedQA development evaluation; not an official Jev API or SDK.
- Prosodia, Audio-native model with Jev's shape: speech plus N options in, one probability per option out, no ASR and no generated text; questions and option sets are supplied per request.
- PlayJev, Qwen3.5-0.8B-Base fine-tuned to play ten browser games from raw pixels, one frame in, one typed move out.
- Qwen-2.5-1B-RLCD, Qwen2.5-1.5B fine-tune plus parallel constrained decoding; all schema fields scored in one broadcast prefill, 5.6x to 7x faster on Apple Silicon.
- NanoJev, 0.6B parallel decision model with a public dataset and a side-by-side maze and Snake demo against Jev and untuned Qwen.
- Bespoke Nimble, Data, model and recipe for an open Jev: Qwen3.5-9B plus LoRA on contrastive examples, with a 13-subset public evaluation suite scored against Jev 1.13.
- kev, LoRA adapter and a small readout head on Qwen (0.5B to 8B) that answers many typed questions from one prefill, following the Jev architecture write-up.
- LocalJev, Local Jev-compatible POST /v1/systemone in TypeScript for Bun, backed by DiffusionGemma through an OpenAI-compatible endpoint.
- Laya-MLX, Native MLX runtime for Laya typed decisions on Apple Silicon, 13.4 ms median for a short decision, with a live Snake demo.
- Simple Jev, Turns any open Hugging Face model into a classifier and scoring endpoint by reading next-token logits per question, with a public demo API.
- openjev-sglang, Jev-compatible API endpoint served from open models with prefill-only inference.
- decider, Qwen3.5-2B fine-tune that emits typed decisions with calibrated probabilities in one pass.
- open-jev (Dasein Labs), One-pass option scoring with a local Gemma 3 4B on Apple silicon: prefill the context once, expand the KV cache across the options, softmax the option log-probs; Doom demo.
- typesafe-ai-benchmark (imposter Jev), LLM gateway that mimics the TypeSafe structured-output contract, used for Qwen-on-Cerebras side-by-sides.
- rlcd-modernbert-151m, Encoder-side reproduction: GLiClass ModernBERT base retrained for calibrated label probabilities.
- LLM2Jev, Adapts local language models into Jev-compatible decision endpoints.
- LitJev, Jev's decision layer on off-the-shelf Qwen checkpoints: one shared prefill, then option logits read per question, no training.
- Jev on a laptop, Unofficial study of Jev-style parallel typed decisions on stock 1.5B to 8B models on Apple Silicon, with benchmarks.
- open-jev (JoshuaSP), Typed JSON inference with DiffusionGemma: fixed JSON, parallel decisions.
- openjev (zhihz), Local bilingual probability decisions from context, questions and candidate answers.
- jevbetter, One-pass scorer over a variable option list: hashed n-gram encoder, rival-aware attention, gated head, temperature scaling.
- qwen-rlcd, Choice, Score and Noul on Qwen3.5-0.8B, the smallest decoder-based reproduction.
- Laya, ModernBERT-large with RLCD-trained decision heads: Choice, Score and Noul in one 38 ms pass.
- Parallel Constrained Decision Engine, Live demo of the Qwen-2.5-1B-RLCD approach: KV-cache broadcast, logit slicing per candidate, 100 percent schema validity.
- LFM2.5-350M-RLCD, 350M-parameter RLCD-style decision model, the smallest open attempt.
- LFM2.5-2.6B-RLCD, RLCD-style fine-tune of Liquid AI's LFM2.5-2.6B for typed decisions.
- system-one-qwen3.5-4b-scorer, Qwen3.5-4B base trained as a Score-style rubric rater.
- system-one-mini, DistilBERT-sized System One shape, a floor for how small the idea can go.
- Jebadiah, Apache-2.0 decision models at 27B, 9B and 4B on Qwen bases that return a probability for every Choice, Noul and Score option in one forward pass; serves Jev's
/v1/systemonewire. - jevos, MiniCPM5-1B cut to 17 layers and served as a 619 MB GGUF on CPU-only llama.cpp; answers Noul (yes/no) questions over Jev's
/v1/systemonewire in 54 to 220 ms, 0.815 accuracy versus Jev's 0.927. - WebJev, Apache-2.0 Qwen3.5-35B-A3B fine-tune for browser agents that returns a probability for every operation and element option over Jev's
/v1/systemonewire, released with its training recipe and live-web training data; authors report 38.52% on 125 real-website tasks graded by deterministic verifiers versus 16.67% for Jev 1.13 inside the same jev-ultrafast agent.
Open, licensed software that puts Jev inside something that runs: routers, agents, games and integrations.
- Jev Chat Assistant, Android chat overlay where Jev judges intent and ranks replies while a separate LLM drafts them; the author reports roughly one-second judgments and device-tested WeChat, QQ, X and Lark adapters.
- RoboJEV, Two-stage Franka Panda control from structured simulator state; 43 of 50 Jev trials succeed versus 48 of 50 rule-baseline trials across five MuJoCo tasks, with success and failure recordings.
- jevals, Agent trace evaluations and tool-call gates batched into Jev requests; the authors' 20-row RAG comparison reports 244 ms request p50, while local Kev and Laya timings are estimates.
- jev-calibrate, CLI for tuning criteria and thresholds with a ledger of holdout reuse; the authors' small support-ticket example improves frustration accuracy from 0.69 to 0.92 on tuning data and scores 0.97 on holdout.
- Laya vs Jev: T-Rex arena, Local Laya and hosted Jev play the same T-Rex course; both finish two published assisted rounds without deaths.
- QuantDinger, Open-source AI trading OS with Jev System One decisions inside its agent and vibe trading loops.
- Jev Ultrafast, Browser agent whose every step is one Jev choice over an indexed element table, a small LLM types only when the action is TYPE_TEXT; Zürich to London on Google Flights in 7.1 seconds.
- Jev Social, Instagram, TikTok and LinkedIn research where Jev picks each bounded search or read and the socai CLI runs it in the user's Chrome; one recorded run captured four source-linked records in 63.969 seconds.
- fast-jev-compaction, Claude Code plugin that replaces the compaction summary with Jev decisions: every tool call and result scored in one request, stale ones dropped, everything kept verbatim.
- jev-trader, One Jev decision every Monad block: buy or sell on the Kuru MON-USDC order book every 300 ms, each answer posted as a real limit order.
- Distill, Lightweight coding-agent harness and TUI that routes its small decisions through Jev.
- TipTour, Menu-bar companion for macOS where Jev picks the next click from locally detected controls and TipTour executes and validates it.
- typesafe-computer-use, Drives a Mac from a plain-English goal at about a fiftieth of a cent per step: OCR the screen, Jev classifies the next action, a writing model only for free text.
- jev-review, Staged code-review workflow with a local dashboard, each stage a Jev decision.
- Jev Search, Plain-language web search where Jev picks sources, time ranges and terms, then ranks the results streamed from Search1API.
- Mobile Jev, Standalone Android agent: one goal, a real phone, Jev makes every decision; opens Uber and books a route to the Golden Gate Bridge in the demo.
- pg-jev, PostgreSQL extension that filters, ranks and classifies rows with plain-language conditions, every row judged by Jev; no index, no embeddings.
- Jev Browser Use, Codex skill where Jev handles navigation, clicks, toggles and scrolling and Codex keeps text input and the final check; 5 to 10x faster browser operations in the authors' workflows.
- quackd, One CLI for open-source robots such as LeRobot arms and Open Duck, with Jev choosing which taught move comes next.
- Abide, Enforces the rules in AGENTS.md and CLAUDE.md that no linter can check: one Jev question per rule on every edit, about 300 ms each.
- JevRouter, Routes models, subagents, Skills and MCP tools through one typed Jev question, with its own permission and confirmation rules around the answer; 44 percent first-five tool-call hits on 10 Toolathlon tasks against 24 percent for DeepSeek V4.1 Flash.
- Codex Jev Router, Selects Codex subagent models with Jev; its controlled 24-run benchmark reports 12/12 correct answers in each arm, with routing using 3.9% more tokens and time but an estimated 52% lower API bill at published rates.
- jev-drone, Quadrotor flies a five-station MuJoCo obstacle course from its onboard camera, Jev at 2.5 Hz deciding what the situation means while the controller stays in code.
- YouTube sponsor detection, Detects sponsor segments from live audio and skips them, one Jev read per segment.
- jevmeter, Puts a live Jev score meter on any video, installed in three steps.
- Supercov, Code quality and coverage for coding agents: Jev scores the source, the usual test command runs, uncovered paths become small queries.
- dspy-typesafeify, Decorator that routes DSPy typed Signatures to Jev where the signature is a pure decision.
- Embodied Jev, MuJoCo robot decision workbench where Jev picks the next manipulation step.
- Grok Bot + Jev, Connects TypeSafe Jev to Grok Bot as a cheap decision layer.
- RefGarden, Spatial reference explorer over The Met, NASA and Cosmos, with Jev choosing search phrases and highlighting references from titles alone.
- SmartMoney-Cub, Read-only trading journal and review harness with Jev judging entries against the evidence.
- Jev × LIBERO, Fine-grained robot control on LIBERO with Jev deciding and physics-grounded execution.
- jev-robot-control, Same task, different decisions: Jev against GPT-4.1 and GPT-4o mini on a robot arm, with cost and time per episode.
- tsai-sc, TypeSafe Jev controls the original StarCraft, one typed decision per game tick.
- jgrep, Code search, diff gate and test selection with one Jev Noul per chunk, 16 chunks per request; the author measures a 238-chunk TypeScript tree in 1.0 s for $0.003 and
--testspicking tests for a diff in one 3,430-token request. - Jev Deep Research, Parallel evidence finding for GPT research agents with Jev Choice and Noul through Pi-Serini; reports 19/20 correct for Jev-60 on a 20-question BrowseComp-Plus development sample, with reproduction commands.
- Sedum, Playwright test tool where Jev picks each next action in goal mode, resolves the element for each plain-English step and judges verify claims; the author's five-run timing of a 17-step saucedemo checkout is 14 s vs 69 s on a per-step AI platform.
Benchmarks, leaderboards and independent tests of Jev, with the headline number where the source gives one.
- Jev Decision Index, Leaderboard that runs Jev and 70 open reproductions on the same suite of 43 benchmarks, about 120,000 decisions per model, chance-corrected so 0 means guessing; a separate kit reruns every entrant.
- JevBench, 534-decision benchmark with public task results and a v1.3 composite over intelligence, calibration, speed and cost; Jev scores 74.4.
- Jev Arena, Same 10,000 comments labelled by Jev 1.13 and DeepSeek Flash and checked against a full GPT-6 Astra review: 203 s against 824 s, $0.84 against $1.50, relevance accuracy 94.7% against 96.3%.
- jev-eval-agent, Jev routes 100 mocked tools behind a confidence gate, measuring steps, tool calls, tokens and cost against the LLM choosing directly.
- Laya vs Jev arena (Prompt Engineer 48), Local Laya and hosted Jev race in Snake and fight in a Mortal-Kombat-style arena, every move a real model decision, from a YouTube video.
- jev-benchmarks (probability-aware), Jev versus GLiNER2.5 on 300 BTZSC examples with calibration and selective risk: 0.910 AG News, 0.870 Banking77, worse on emotion.
- Just Ask Jev, "Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures". 44 alignment-failure detection benchmarks that test Jev as a zero-shot detector, many questions per input in one call.
- jev-rag-benchmark, Locked benchmark of 1,044 Turkish XQuAD questions testing whether a Jev rerank improves a small RAG system on quality, latency and cost.
- JevPokerBench, Texas Hold'em benchmark for decision models with separate cash-game and sit-and-go leaderboards, hand replays and bring-your-own-agent tables; chips are virtual.
- jev-search-rerank-eval, Jev score rerank against keyword, BM25 and embedding search on 9,831 graded pairs from 164 Chinese and English queries, NDCG@10 with bootstrap intervals and the judge bias measured.
- jev-rerank-bench, Jev as reranker over 8 datasets and 1,617 questions: nDCG@10 0.692 versus Cohere Rerank 4 Pro 0.691, at 422ms.
- LegalForecastBench, Claim-level Brier scoring of federal motion-to-dismiss outcomes, a fixed binary with one probability per unit.
- jev-benchmark (chess and NPC addressee), Jev no better than random picking chess moves from a FEN, but F1 0.96 on NPC addressee detection, 0.2s median.
- jev-bench, 166,054 human-labelled rows in 22 configs recast as System One questions, keeping human label distributions where they exist; 43 models scored on 22,773 test records.
- Laya vs Jev (Luni), Laya and Jev run on the benchmarks where Jev has published numbers, on one RTX 5090, after showing that Laya's headline comparison used two different benchmarks.
- sysone-bench, September 21 comparison of Jev, Laya and Qwen-PCD on 751 states across nine suites with identical questions; Jev leads moderation 98.9% to Laya's 83.3%, while Laya leads AG News 94% to 91%.
- jev-sec-bench, Blind prompt-injection run on 662 deepset messages: 96.5% accuracy, 0.9927 ROC-AUC, ECE 0.0588, p50 325ms.
- DecisionBench, Typed-decision benchmark with medium and hard sets of 80 situations and 293 questions each, an open runtime and a leaderboard from Hanno Labs; the dataset card scores Jev 1.13 next to frontier LLMs.
- jev-research-eval, Reproducible harness scoring Jev ultrafast research-browser runs over 11 baseline cases plus 18 human and quant stress cases with QC grades.
- jev-spam-eval, Zero-shot spam Noul on 18,514 emails reaching 0.9833 accuracy, matching a TF-IDF classifier trained on 14,800 labels.
- jev-secret-detection, 100 balanced secret-detection cases as single Noul questions scored by accuracy, AUC and Brier; server p50 75 to 90ms.
- Jevals, Jev and six LLMs answer the same Noul, Choice and Score questions against human labels, 300 items and 5 runs per task: Jev ties the best LLM on PubMedQA yes/no (91.3% versus 92.5%) at 1/28 of its cost, and reaches 79.7% on Banking77.
- padflow-jev-evals, Three production SaaS decisions published as schemas with auto-post confidence thresholds.
- jev-playground, Tic-tac-toe and connect four pitting Jev against four frontier models on identical legal-move choice options.
- jev-benchmarks (frontier comparison harness), Harness asking Jev and frontier LLMs identical typed questions, scoring accuracy, calibration, latency and schema validity.
- Confident Where People Disagree, "A preregistered, bias-corrected test of whether TypeSafe AI's Jev lowers its confidence when humans disagree, on ChaosNLI". On items where annotators split, Choice confidence averages 0.807 against 0.468 human agreement.
- DeepSearcher search-stopping evaluation, Jev and a DeepSeek stopping baseline both reach 93.25% supporting-document Recall@5 on 100 sampled 2WikiMultiHopQA queries, replayed over shared seven-round search trajectories; includes reproduction code and archived results.
- MemSearch reranking evaluation, Jev reaches 79.41% Recall@5 versus 81.87% for Voyage rerank-3 over 4,344 Chinese/English query variants with fixed candidates; the memory corpus requires an authorized copy.
- Vector Graph RAG relation-reranker evaluation, Jev reaches 68.87% MuSiQue and 93.50% HotpotQA Recall@5 on 500 queries per dataset, with cached results and reproduction scripts.
- jev-use benchmarks, 454 judgments against a Claude Opus 5 reference: 82.2% agreement, 89.5% among non-escalated verdicts, over a 68.7% majority baseline; context compaction 56.3%, below a constant answerer.
- JevAdvBench, "JevAdvBench: A Benchmark and Black-Box Attacks for Reinforcement Learning for Calibrated Decisions Models". Measures how far manipulated inputs move Jev's typed answers, since a typed model returns a well-formed answer even under attack.
- TypeSafe Jev evals dashboard, TypeSafe's own four-workflow dashboard, Jev at 61.7 to 76.0% accuracy and 0.3 to 0.5s per case against frontier baselines.
- Near Here event validation, 50-case event validation: Jev 96% at 0.59s and $0.043 per 1,000, Mistral small 4 84%, Gemini Flash-Lite 86%.
- Every: Mini-Vibe Check, 777 judgments over 37 articles in 0.7s for a quarter of a cent; caught six of seven planted defects, Fable seven.
- Jev is the fish at the poker table, Poker probe finding 15 to 30 point swings from relabelling the same hand, and 16 of 16 bets against a made flush.
- Jev judge call vs dimension scores, One direct Jev question per row against 12 to 14 Jev-scored dimensions with fitted weights on three tasks: 0.9076 vs 0.8373 on Japanese NLI, but 25x the hard-benign false positives, 37.2% vs 1.5%.
Reporting and technical commentary that checks the launch claims against the evidence.
- Jev in the Wild, "A Data-Driven Analysis of the Jev Model's Functionality, Applications and Ecosystem". Survey of 2,170 public GitHub Jev projects mapping early growth, application domains and decision-use patterns.
- AINews: Jev, a System One Model that only decides, Latent Space roundup of the launch and the HN mapping onto encoders, GLiNER, constrained decoding and DSPy.
- Agentpedia claim-vs-evidence guide, Claim-by-claim audit separating verified Jev pricing and latency from unproven calibration; puts aggregate accuracy at 67.8% versus Opus 5's 73.1%.
- The Register: TypeSafe AI debuts model for machines, Press account of the $40M raise, the Doom demo and the caveat that structured output is a different error type, not correctness.
- Jev vs auto-regressive LLMs vs MDLM, Technical comparison of Jev's single-pass sampler with token-by-token decoding and masked diffusion.
- How does Jev work? RLCD and parallel inference, Explainer reconstructing the RLCD objective and the parallel sampler from public statements.
- MrJev: what each Jev tool sends, and where, Hands-on reviews of 72 community projects, 66 run in a container with a real key to record what leaves your machine; a password in a database URL and world-readable prompt logs, fixed upstream.
- Typed Decisions, Not Chat, Secondary analysis of TypeSafe's dashboard putting Jev at about 67.8% mean agreement against 74.1% for the best comparator.
Earlier work with the same input and output shape: a fixed answer set, one probability per option, no generated text. Label-conditioned encoders, scalar reward heads, reinforcement learning for calibrated confidence, and the single-pass inference TypeSafe's own forks point at.
- Zero-shot Classification as Entailment, "Benchmarking Zero-shot Text Classification: Datasets, Evaluation and Entailment Approach". Label set given at inference, an entailment model returns one probability per label, no text generated.
- InstructGPT reward model, "Training language models to follow instructions with human feedback". A Bradley-Terry head emits one scalar per response in a single pass, no text, co-authored by Jev's founder.
- ProtectAI prompt-injection DeBERTa v2, 184M DeBERTa returning a binary injection probability, the BERT-style encoder guardrail HN engineers mapped Jev onto.
- RLCR, "Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty". Adds a Brier-score reward to RLVR so the model emits calibrated confidence, the closest published relative of TypeSafe's RLCD.
- LLaDA, "Large Language Diffusion Models". The typesafe-ai GitHub org forked this masked diffusion LM, the strongest public hint at how Jev fills every answer slot in one pass.
- GLiNER, "Generalist Model for Named Entity Recognition using Bidirectional Transformer". The span-and-label encoder family HN mapped Jev onto, types supplied at inference and scored in one bidirectional pass.
- GLiClass, "Generalist Lightweight Model for Sequence Classification Tasks". The open analogue HN pointed at: labels and text in one encoder pass, one probability per label, no decoding.
- monoBERT, "Passage Re-ranking with BERT". Landmark cross-encoder: pair in, one scalar relevance probability out, no generation, the ancestor of Jev's Score primitive.
- Generative or Discriminative?, "Revisiting Text Classification in the Era of Transformers". Controlled comparison of encoder, autoregressive and diffusion classifiers over fixed label sets on accuracy, calibration and ordinality.
- Llama Guard, "LLM-based Input-Output Safeguard for Human-AI Conversations". Fixed safety taxonomy with the verdict read off one safe/unsafe token probability, the guardrail classifier Noul replaces.
- GPT-4 Technical Report. Reports that RLHF destroys the base model's calibration, the finding RLCD is positioned against.
- Rewarding Doubt, "A Reinforcement Learning Approach to Calibrated Confidence Expression of Large Language Models". Trains confidence expression by RL on the logarithmic scoring rule, an independent rediscovery of the proper-scoring-rule reward RLCD uses.
- Calibration-Aware RL for Decision-Making LLMs, "Balancing Classification and Calibration Performance in Decision-Making LLMs via Calibration Aware Reinforcement Learning". RL that adjusts decision-token probabilities directly, keeping RLVR accuracy while cutting ECE, the closest public analogue of Jev's typed decision heads.
- vLLM, "Efficient Memory Management for Large Language Model Serving with PagedAttention". The typesafe-ai GitHub org forked this engine; paged KV cache plus prefix caching is what makes extra questions over one shared state nearly free.
- Mercury, "Ultra-Fast Language Models Based on Diffusion". HN read Jev as a stripped down text diffusion model, and Mercury is that idea shipped commercially with parallel refinement.
The tools TypeSafe and the launch discussion named as what Jev replaces: constrained decoding, structured outputs, typed prompt programming, LLM judges, routers and guard classifiers.
- Outlines, "Efficient Guided Generation for Large Language Models". The finite-state-machine guided decoding HN named as the incumbent way to get typed values, which Jev claims to replace.
- DSPy, "Compiling Declarative Language Model Calls into Self-Improving Pipelines". Typed signatures compiled into prompts, named on the HN thread as the fair comparison for Jev's typed question interface.
- RouteLLM, "Learning to Route LLMs with Preference Data". Router scores a fixed two model set and returns win probability per option before any text is generated.
- Guidance, Constrained generation library named on the HN launch thread as what Jev's typed outputs get compared against.
- OpenAI Structured Outputs, The provider-side JSON-schema guarantee the CEO named on HN as what Jev replaces, shape enforced but no probability returned.
- Let Me Speak Freely?, "A Study on the Impact of Format Restrictions on Performance of Large Language Models". Measures the accuracy format restrictions cost, the study behind the CEO's HN claim that constrained decoding makes models dumber.
- JSONSchemaBench, "A Rigorous Benchmark of Structured Outputs for Language Models". 10k real schemas scored on validity, coverage and latency, the constrained-decoding route Jev's 0% type errors claim competes against.
- MT-Bench, "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena". The landmark LLM-as-a-judge paper, named on the launch thread as the layer Jev's score and Noul primitives replace.
- Constitutional Classifiers, "Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming". Input and output classifiers gating a frontier model on a fixed policy, the deployed slot Noul targets.
System 1 in Kahneman's sense, the bitter lesson TypeSafe argues with, and the Jevons paradox the model is named after.
- Thinking, Fast and Slow, The source TypeSafe cites for naming Jev after System 1, fast intuitive judgement with no deliberation.
- Maps of Bounded Rationality (Kahneman Nobel lecture), Kahneman's two-system account, intuition returning an answer directly while reasoning deliberates, the split Jev's design copies.
- Thinking Fast and Slow in AI. The AI charter for System 1 components that answer from experience without search, what System One Models productizes.
- The Bitter Lesson, The essay TypeSafe's own Bitterest Lesson argues against, named in TypeSafe's materials as its starting point.
- Jevons' paradox (Alcott 2005), The rebound effect Jev is named for, where cheaper decisions raise total decision volume.
The sibling lists.
- Awesome JEV Papers, Research papers on Jev and System One decision models only.
- Awesome AI Scientist, AI systems that do science.
- Awesome RSI, Systems whose improvement loop modifies itself.
Open a pull request. Link the paper or the primary page, add the code repository or Hugging Face path if there is one, and say in one line which test the entry passes. CONTRIBUTING.md has the four tests and the entry format.
@misc{awesome_jev,
title = {Awesome JEV},
year = {2026},
howpublished = {\url{https://github.com/OmniJev/awesome-jev-gallery}},
note = {Papers, open models and evaluations behind System One models and typed decisions}
}