Nace AI Open-Sources Drex 1.5, a 9B Decision Model That Matches Closed Jev
Nace.AI open-sources Drex 1.5, an 8.95B decision model scoring 58.08 on Decision Index 0.3.1 — tied with closed Jev 1.13.0 and the best score under 10B parameters.
Updated
Why it matters
- Drex 1.5 scores 58.08 on Decision Index 0.3.1, tied with closed Jev 1.13.0 (57.96) and best under 10B parameters.
- The model reaches 93.4% accuracy on 32K–128K token documents with 2.0 s median latency.
- It runs on a single 24 GB A10G GPU in bf16 (≈18 GB) or as a 9.5 GB Q8_0 GGUF on Apple silicon and CPU.
- Hosted on OpenRouter at $0.04 per 1M input tokens with $0 output cost, served by DeepInfra.
- Weak spots include GPQA Diamond (45.4% vs Jev's 78.6%) and ACOS sentiment (7.4% F1 vs 29.5%).
Nace.AI has open-sourced Drex 1.5, an 8.95B-parameter decision model that scores 58.08 on the public Decision Index 0.3.1 — statistically tied with closed-source leader Jev 1.13.0 at 57.96, and the top score for any model under 10B parameters. Weights are available on Hugging Face, and a hosted version is live on OpenRouter at $0.04 per 1M input tokens with $0 output cost, served by DeepInfra.
The release matters because it puts an open-weight model on par with a proprietary API in the emerging decision-model category. Teams running agents and backend workflows can now deploy a state-of-the-art decision layer on a single GPU — or on a laptop — rather than paying per-request to a closed provider.
What does Drex 1.5 actually do?
Drex 1.5 does not generate text. It reads a state — plain text or JSON — plus typed questions, and returns a probability for each option in a single forward pass. No tokens are sampled. Parameters like temperature and top_p do not apply, because the model can only answer with options the developer supplied.
The model supports three question types:
choice— pick among named optionsnoul— yes/no questionsscore— ordinal ratings
It serves the POST /v1/systemone API, the same request format used by TypeSafe's Jev, the closed model that created this category. Nace says existing Jev clients work with Drex after changing a few environment variables.
How does the model work under the hood?
The backbone is MiMo-V2.6-Distill-Qwen-9B, a distilled Qwen 3.5 9B model with 32 layers and hybrid attention: 3 linear-attention layers per full-attention layer. A separate pointer head, stored as head.pt, scores each option from the backbone's hidden states.
Each question requires one pass over the state plus that question. In llama.cpp, the state is encoded once and shared across all questions, which cuts compute for multi-question requests.
One training detail deserves scrutiny. Nace's Drex page says the model was trained on the official training splits of the index benchmarks, then evaluated only on held-out splits. That helps on familiar decision types; results in new domains may differ.
How does Drex 1.5 perform on benchmarks?
On the public Decision Index 0.3.1 — 37 benchmarks, chance-corrected — the model card reports:
- Drex 1.5 (Nace.AI): 58.08
- Jev 1.13.0 (TypeSafe AI): 57.96
- Bespoke Nimble 9B v3 (Bespoke Labs): 57.19
- clef-flash (Cloudflare): 56.15
Drex, Jev and Nimble all sit within the board's 0.9-point tie band, so the top three are effectively level. Drex leads Jev on 20 of the 37 individual benchmarks. Nace computed the Drex score by running the official evaluation kit itself.
The area breakdown shows where the model excels and where it struggles. Drex scores strongest in Tools at 75.0 and weakest in Knowledge and Reasoning at 44.6. Nace's launch chart still cites the older Decision Index 0.2.1, where Drex scored 58.28 against Jev's 57.91.
On JevBench, a set of 231 public items, Drex scores 86.2% against Jev's 87.0%, and both models reach 73.9% on hard items. In a head-to-head across 8 OpenSpiel games, Drex recorded 122 wins, 47 draws and 87 losses against Jev — a 56.8% win rate.
Why do long documents stand out?
Long-context decision-making is Drex's clearest strength. Nace reports:
- 8K to 32K tokens: 89.5% accuracy, median latency 0.65 seconds
- 32K to 128K tokens: 93.4% accuracy, median latency 2.0 seconds
Truncating those same requests to 8K tokens drops accuracy to 76.5% and 78% respectively. The accuracy actually rises with longer documents in the upper range, which suggests the model uses the full context rather than degrading as input grows. Context defaults to 16,384 tokens and extends to 131,072.
How can developers run it?
There are four deployment paths, all serving the same API:
- Python (Kev runtime):
inference.pyandserve.pyon a CUDA GPU - llama.cpp: a Nace fork with GGUF weights in bf16 or Q8_0, on CUDA, Metal or CPU
- Ollama: a Nace fork that adds a
decisioncapability - Hosted: OpenRouter at $0.04 per 1M input tokens and $0 output, or Nace's own Console API
The hardware requirements are modest by current standards. bf16 weights take about 18 GB, and the model runs on a single CUDA GPU — Nace tested it on an AWS g5.2xlarge with an A10G 24 GB card. The Q8_0 GGUF is about 9.5 GB and runs on Apple silicon and CPU. Nace reports identical answers between bf16 and Q8_0 on the A10G, and the Q8_0 build also matched on an Apple M5 Pro under both Metal and CPU.
Note that local deployment through Ollama and llama.cpp requires Nace's forks, not mainline builds. A Drex agent skill plugs the model into Claude Code, Codex, Cursor, OpenCode, Hermes Agent, Gemini CLI and GitHub Copilot. Nace's launch post offers a $25 sign-up bonus for cloud users.
How does Drex compare with Jev and Nimble directly?
The competitive picture, drawn from the Drex model card and public leaderboards:
| Feature | Drex 1.5 | Jev 1.13.0 | Bespoke Nimble 9B v3 |
|---|---|---|---|
| Developer | Nace.AI | TypeSafe AI | Bespoke Labs |
| Parameters | 8.95B | Not disclosed | LoRA on Qwen3.5-9B |
| Weights | Open | Closed (API only) | Open (adapter) |
| License | Nace.AI Open RAIL-M | Proprietary | CC BY-NC 4.0 |
| Context | 16,384 default, up to 131,072 | 64K per request | Not disclosed |
| Decision Index 0.3.1 | 58.08 | 57.96 | 57.19 |
| JevBench (231 items) | 86.2% | 87.0% | Not disclosed |
| Local hardware | 1 CUDA GPU; Apple silicon or CPU via Q8_0 | Not applicable | Qwen3.5-9B base plus adapter |
| API price per 1M (input/output) | $0.04 / $0 | $0.042 / $0 | Not disclosed |
What are the limitations?
Drex 1.5 is a decision layer, not a general model. It cannot generate text, code or explanations — you need a separate generator for that.
Knowledge-heavy tests expose the gap with larger general models. Drex scores 45.4% on GPQA Diamond versus Jev's 78.6%, and 58.7% on MMLU-Pro versus 82.7%. Fine-grained sentiment is the worst result on the board: 7.4% per-review F1 on ACOS aspect sentiment, against 29.5% for Jev.
The license also carries obligations. The weights ship under the Nace.AI Open RAIL-M license, which includes use restrictions — developers should check the terms before commercial deployment.
What is the takeaway?
Drex 1.5 delivers the category-leading decision performance of a closed API in an open-weight package that runs on one 24 GB GPU or a 9.5 GB quantized file on a Mac. Its weaknesses — broad knowledge, fine-grained sentiment — define where teams should still route to other systems. With Jev, Nimble and now Drex clustered within a single tie band, the decision-model category has its first genuinely contested race between open and closed players, and the next index revision will show whether open weights can pull ahead rather than keep pace.
Original: huggingface.co
More from James Calloway
Show full bio
News editor covering industry trends and analytics at AI In Context.
224 articles
Related articles
- Amazon open-sources Strands Decider 2B as decision models multiply
- OpenAI's Decisions API Enters Public Beta With 10x Faster Typed Answers
- OpenAI Launches GPT-5, Claims State-of-the-Art Results Across the Board
- OpenAI Announces GPT-5.5 for Coding, Research and Data Analysis
- Decision AI Models Emerge as New Category: Jev, GLiDE and Open Weights