Decision AI Models Emerge as New Category: Jev, GLiDE and Open Weights
Decision AI models return typed choices and probabilities instead of text. TypeSafe's Jev hit 67.8% accuracy at $0.0004 per case, sparking Fastino rivals and open-weight clones within weeks.
Updated
Why it matters
- TypeSafe's Jev costs $0.042 per million input tokens with free output, and matched Claude Sonnet 5's 67.8% accuracy at $0.0004 per case versus $0.1174 and 0.4s versus 78.1s in TypeSafe's workflow evals
- Vercel reports Jev became the fastest-adopted model in AI Gateway history, with nearly 13% of paid teams using it within 24 hours — 2x GPT-5.6's share and 6x Fable 5.1's
- Fastino Labs shipped GLiNER2.5-Decide (Apache 2.0, 340M DeBERTa-v3-large) on September 24 and the closed GLiDE API on September 30; open reproductions include Laya, JevK5, OpenJev, and kev-0.5b
Vercel says Jev, TypeSafe AI's "decision model," became the fastest-adopted model in the history of its AI Gateway: within 24 hours of launch, nearly 13% of paid teams were using it — twice the share of the GPT-5.6 family and more than six times Fable 5.1's. That adoption spike marks the arrival of a new model category. Decision AI models return typed answers — choices, scores, yes/no probabilities — instead of paragraphs, letting code branch directly on the output without parsing generated text.
TypeSafe launched Jev after two years in stealth, calling it a "System One model" after Daniel Kahneman's fast, intuitive System 1 thinking. Within three weeks, Fastino Labs shipped two rival models, and independent open-source developers published several Jev-style reproductions.
How Jev works
Jev accepts a "state" — a string, array, or set of name-value pairs — plus one or more typed questions. According to TypeSafe's documentation, it supports three primitives. Choice picks one option from a list with probabilities and confidence, supporting up to 255 options. Score rates the state against a rubric of ordered levels. Noul — short for Bernoulli — returns a 0-to-1 probability that a statement is true.
Every question is evaluated in parallel and in isolation against the same state, so adding questions barely changes response time. Because Jev never generates strings, TypeSafe says it cannot return a type error. Under the hood, TypeSafe describes a new architecture, a parallel sampler, and a training method called Reinforcement Learning for Calibrated Decisions (RLCD). Where RLHF optimizes for human preference, RLCD optimizes for calibrated probabilities: higher confidence should mean higher accuracy.
The pricing is the headline figure. Jev costs $0.042 per million input tokens, and output is free. OpenRouter lists a 32K context window, and TypeSafe reports end-to-end responses between 70 and 500 milliseconds.
The benchmark picture
TypeSafe built workflow evals across four tasks: security incidents, agent trace observability, invoice processing, and customer service, with reference labels averaged from GPT-6 Astra and Claude Fable 5.1 at high thinking. The results show the cost story clearly:
- Jev: 67.8% mean accuracy, $0.0004 per case, 0.4 seconds
- Claude Sonnet 5 (same workflow): 67.8%, $0.1174 per case, 78.1 seconds
- Best comparison model (OpenAI's "sol"): 74.1%, $0.0836 per case, 23.3 seconds
Jev matched Sonnet 5 on accuracy at a fraction of the cost and latency. It still trails the top frontier configuration by 6.3 points. Per task, Jev scored 76.0% on customer service but only 61.8% on invoice processing.
Fastino's own numbers tell a different story — with caveats. On Fastino's Decision Index 0.2.1, GLiDE scored 64.81 versus Jev's 57.91. On CLadder accuracy, Fastino reports GLiDE at 88.7% against Jev's 72.6%; on CRUXEval, 92.6% against 73.0%. But Fastino chose both the tests and the opponents, and its "Fast Decisions" comparison across 17 datasets — where GLiNER2.5-Decide led at 60.1%, ahead of JevK5 at 57.5%, SemIf at 56.4%, GLiFormer at 49.0%, and Laya at 46.6% — used JevK5, an open reproduction, not TypeSafe's Jev.
The competitive field
Fastino Labs is Jev's most direct commercial rival. It shipped GLiNER2.5-Decide, a 340M-parameter DeBERTa-v3-large encoder with Apache 2.0 open weights, on September 24, followed by GLiDE, a closed hosted API with 40K context, on September 30. GLiDE adds adaptive thinking on uncertain cases; Fastino positions it as "the first thinking decision model." GLiNER2.5-Decide runs locally on CPU or GPU, with reported p50 latency of 38.3 ms on a V100 and 167.3 ms on a 48-vCPU CPU. It can also decode joint constraints — such as "safety" and "harm type" together — so answers never contradict each other.
The open-source lane is crowded. Convai Innovations released Laya under Apache 2.0, in a 421M ModernBERT-large English version and a 322M multilingual mmBERT-base version supporting 100+ languages, with roughly 33 ms per question on GPU. An independent developer published JevK5, a 4B Qwen3.5-4B model with merged LoRA under Apache 2.0. Theo Lee's OpenJev (MIT license) uses a frozen Qwen3.5-4B that reads option logits, and reports running 5.21x faster than autoregressive JSON output; on a 102-row TypeSafe eval subset it scored 0.845 balanced accuracy against Jev's 0.883. Jared Palmer's kev-0.5b adds a LoRA and readout head to Qwen2.5-0.5B behind a Jev-compatible API, answering six questions in roughly 160 ms on an Apple M5.
Where decision models fit
The rule of thumb is simple. If your code needs a bounded answer it will branch on, a decision model is a candidate. If a human needs to read the output, use an LLM.
Agent control flow tops the list. Vercel lists choosing the next tool or subagent as a primary use. A single Choice question can replace a fragile JSON-parsing step when deciding whether to continue, retry, ask the user, or stop. Fastino lists model routing by destination, complexity, or escalation level. OpenRouter's Jev guide describes "verified cascades": draft with a cheap model, check with Jev, escalate only on failure.
Classification and triage follow — support ticket routing, email triage, intent detection, spam detection — which make up much of Fastino's 17-dataset benchmark. Simon Willison calls these natural classification fits. TypeSafe's simplest published workflow decides whether to close a security alert, pass it to an analyst, or contain it.
The category is also moving into evaluation and observability. Arize and Langfuse both shipped Jev-as-a-judge evaluators; Langfuse labels the feature "decision-model evaluators" so other models can be added later, and its Jev support is still marked experimental. Buddy lets CI/CD pipelines score, classify, or gate runs with a Jev action.
Real-world results are starting to accumulate. A new arXiv paper used Jev to interpret service contracts at the 6G network edge, cutting median decision latency by 22.4% versus DeepSeek and 61.9% versus Gemini at matched correctness. TypeSafe's Doom demo ran Jev at 10 queries per second at an estimated cost of roughly $7 per hour. Willison demonstrated search reranking: fetch 100 BM25 candidates, then have Jev score each for relevance in one parallel call.
The limits
Decision models are the wrong tool for several jobs. They cannot generate text, summaries, or explanations. TypeSafe's Jev 1.13 jaggedness guide flags exact arithmetic, counting, and date math as failure modes. And Willison warns against using them for decisions that affect people's livelihoods, such as hiring, because hidden bias is hard to inspect.
The underlying idea is not new. Classifiers and rerankers have made decisions for years. The Decision Transformer paper (2021) framed reinforcement learning as sequence modeling; DeepMind's Gato (2022) showed one generalist model acting across many tasks; a 2023 survey even called "large decision models" the next step. What changed in 2026 is the packaging. Agents need cheap judgments thousands of times per workflow, and paying frontier prices for each does not scale. Calibration scores let code act alone when confident and escalate when not. And the ecosystem moved fast: within weeks, Jev landed on Vercel, OpenRouter, Arize, Langfuse, and Buddy.
For buyers, the choice comes down to deployment constraints. Jev offers the broadest ecosystem support as a managed API. GLiDE targets harder, reasoning-heavy decisions — with Fastino's claims worth verifying on your own data. GLiNER2.5-Decide suits air-gapped deployment, fine-tuning, and cross-question constraints. Laya covers multilingual workloads at very low per-question latency. The open reproductions — kev, OpenJev, JevK5 — serve local experimentation and vendor-lock-in avoidance. With Fastino already shipping two challengers and at least three open-weight reproductions in the field within a month of Jev's launch, the decision-model category is consolidating from novelty into infrastructure — and the benchmark claims now circulating from competing vendors will need independent evaluation to sort out.
Original: typesafe.ai
More from Marcus Bennett
Show full bio
Senior reporter covering consumer brands and retail at AI In Context.
145 articles
Related articles
- Amazon open-sources Strands Decider 2B as decision models multiply
- OpenAI's Decisions API chases TypeSafe's Jev in fast agent control
- OpenAI Launches GPT-5, Claims State-of-the-Art Results Across the Board
- Google Upgrades Gemini 3 Deep Think With Record Benchmark Runs
- OpenAI Ships GPT-5.1: Faster Reasoning, Better Coding, Same Price