Nace AI Open-Sources Drex 1.5: A 9B Decision Model That Scores Options, Not Text


Nace.AI has open-sourced Drex 1.5, a 9B decision model for agents and backend workflows. The Drex 1.5 decision model does not write text. It reads a state and typed questions, then returns a probability for every option. Nace reports 58.08 on the public Decision Index 0.3.1, the top score under 10B parameters. Weights are on Hugging Face, and a hosted version is live on OpenRouter.

TL;DR

  • Size: 8.95B parameters (dense), bf16 weights about 18 GB. Context is 16,384 tokens by default, up to 131,072.
  • Runs on: 1 CUDA GPU in bf16 (tested on a 24 GB A10G). A Q8_0 GGUF (about 9.5 GB) runs on Apple silicon and CPU.
  • Performance: 58.08 on Decision Index 0.3.1 (public), within the board’s tie band of Jev 1.13.0 (57.96).
  • Best: 93.4% accuracy on 32K to 128K token documents.
  • Worst: 7.4% per-review F1 on ACOS aspect sentiment, versus 29.5% for Jev.
  • Bottom line:
    • Best: open weights that match a closed model on 1 GPU.
    • Worst: weak on broad knowledge and fine-grained sentiment.

What is Drex 1.5?

Drex 1.5 is a decision model from Nace.AI that scores a fixed set of options in 1 forward pass. You send a state (text or JSON) and named questions. It supports 3 question types: choice, noul (yes/no) and ordinal score. No tokens are sampled, so temperature and top_p do not apply. The model can only answer with options you supplied.

It serves the POST /v1/systemone API. That is the same request format used by TypeSafe’s Jev, the closed model that started this category. Nace says existing Jev clients work after changing a few environment variables.

How does Drex 1.5 work?

The backbone is MiMo-V2.6-Distill-Qwen-9B, a distilled Qwen 3.5 9B model. It has 32 layers with hybrid attention: 3 linear-attention layers per full-attention layer. A separate pointer head (head.pt) scores each option from the backbone’s hidden states.

Each question runs 1 pass over the state plus that question. In llama.cpp, the state is encoded once and shared across questions. Nace’s Drex page says the model was trained on the official training splits of the index benchmarks. It was evaluated only on held-out splits.

How does Drex 1.5 perform on benchmarks?

On the public Decision Index 0.3.1 (37 benchmarks, chance-corrected), the model card reports:

  • Drex 1.5: 58.08
  • Jev 1.13.0: 57.96
  • Bespoke Nimble 9B v3: 57.19
  • Cloudflare clef-flash: 56.15

Drex 1.5, Jev and Nimble sit within the board’s 0.9-point tie band. Drex leads Jev on 20 of 37 benchmarks. The Drex score comes from Nace’s own run of the official kit. Its area scores are strongest in Tools (75.0) and weakest in Knowledge and Reasoning (44.6). Nace’s launch chart uses the older Decision Index 0.2.1, where Drex scored 58.28 against Jev’s 57.91.

On JevBench (231 public items), Drex scores 86.2% against Jev’s 87.0%. Both reach 73.9% on hard items. In a head-to-head across 8 OpenSpiel games, Drex recorded 122 wins, 47 draws and 87 losses against Jev (56.8%).

Long documents are a clear strength:

  • 8K to 32K tokens: 89.5% accuracy, median 0.65 s
  • 32K to 128K tokens: 93.4% accuracy, median 2.0 s

Truncating the same requests to 8K tokens drops accuracy to 76.5% and 78%.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *