AI News Feed
Market watch
Companies

Nace.AI and Microsoft Release 9B-Class Decision Models That Score Options, Not Text

Nace.AI open-sourced Drex 1.5 and Microsoft released Microsoft-Decision-1, two 9B-class decision models that return probabilities over fixed options instead of writing text.

Drex 1.5 is a dense model with 8.95 billion parameters and bf16 weights of about 18 GB. Its context is 16,384 tokens by default and up to 131,072. It reads a state and typed questions, then returns a probability for every supplied option. It supports choice, yes-or-no and ordinal-score questions; no tokens are sampled, so temperature and top_p do not apply, and the model cannot answer outside the options it is given. Nace says the model serves the POST /v1/systemone API, the same request format used by TypeSafe's Jev, and existing Jev clients work after changing a few environment variables. The backbone is MiMo-V2.6-Distill-Qwen-9B, a distilled Qwen 3.5 9B model with 32 layers and hybrid attention. A separate pointer head scores each option from the backbone's hidden states. Nace's Drex page says the model was trained on the official training splits of the index benchmarks and evaluated only on held-out splits. Weights are on Hugging Face, and a hosted version is live on OpenRouter. Nace tested bf16 on a 24 GB A10G and said a Q8_0 GGUF of about 9.5 GB runs on Apple silicon and CPU.

On the public Decision Index 0.3.1, which covers 37 chance-corrected benchmarks, Drex 1.5 scores 58.08, which Nace reported as the top score under 10 billion parameters. TypeSafe's Jev 1.13.0 scores 57.96, and Bespoke Nimble 9B v3 scores 57.19; Nace says Drex, Jev and Nimble sit within the board's 0.9-point tie band. Drex leads Jev on 20 of the 37 benchmarks. Its strongest area score is Tools at 75.0, and its weakest is Knowledge and Reasoning at 44.6. On the older Decision Index 0.2.1 used in Nace's launch chart, Drex scores 58.28 against Jev's 57.91. On JevBench, a set of 231 public items, Drex scores 86.2% against Jev's 87.0%, and both reach 73.9% on hard items. In a head-to-head across eight OpenSpiel games, Drex recorded 122 wins, 47 draws and 87 losses against Jev, a 56.8% win rate. Long documents are a clear strength: 89.5% accuracy on 8K to 32K tokens with a median 0.65 seconds, and 93.4% accuracy on 32K to 128K tokens with a median 2.0 seconds. Truncating the same requests to 8K tokens drops accuracy to 76.5% and 78%. Its reported worst result is 7.4% per-review F1 on ACOS aspect sentiment, versus 29.5% for Jev.

Developers can run Drex 1.5 through four paths. Python uses a Kev runtime with inference.py and serve.py on a CUDA GPU; llama.cpp uses a Nace fork with GGUF in bf16 or Q8_0 on CUDA, Metal or CPU; Ollama uses a Nace fork that adds a decision capability; and a hosted version is available on OpenRouter at $0.04 per 1 million input tokens and no output charge, served by DeepInfra. Nace also runs its own Console API. Nace tested bf16 and Q8_0 on an AWS g5.2xlarge with a 24 GB A10G and got identical answers, and said the Q8_0 GGUF also matched on an Apple M5 Pro on both Metal and CPU. A Drex agent skill plugs the model into Claude Code, Codex, Cursor, OpenCode, Hermes Agent, Gemini CLI and GitHub Copilot. Nace's launch post offers a $25 sign-up bonus for cloud users. The model is a decision layer, not a general model: it cannot generate text, code or explanations, and it is weak on broad knowledge, scoring 45.4% on GPQA Diamond versus Jev's 78.6% and 58.7% on MMLU-Pro versus 82.7%.

Microsoft-Decision-1 is post-trained from Alibaba's Qwen3.5-9B, according to MarkTechPost. Microsoft did not disclose the exact parameter count. The model has a 32,768-token context window and is available only as a hosted API in Microsoft Foundry and on OpenRouter via Azure. There are no open weights, no quantized variants and no disclosed hardware. It returns a calibrated probability for each fixed answer option in one call, is text-only and does not provide explanations. Supported formats include yes-or-no, multiple-choice, rating, classification and rubric questions; grading of AI responses and proposed agent actions; groundedness checks against supplied evidence; and explicit abstention options such as 'cannot tell'. Training used public datasets under Microsoft's Open Data process plus synthetic data. Output is JSON. Microsoft plans to rebase future versions on MAI and OpenAI models. Pricing is $0.042 per million input tokens, with output free. OpenRouter notes that weights update continually while the API shape stays fixed.

Microsoft compared nine systems across 36 benchmarks with 147,137 questions, keeping the benchmarks blind from training. Microsoft-Decision-1 led on average accuracy at 83.5%, ahead of Quyet-1.0-Large at 81.9%. Its calibration score was 92.2, second to Quyet-1.0-Large at 93.1. The model's p50 latency was 85 ms and p95 was 125 ms. Microsoft said that is 4.5 times quicker than Quyet-1.0-Large and 35 times quicker than GPT-6 Sol, which took 3.01 seconds. Microsoft also tested robustness by perturbing each request eight ways, including paraphrasing and option shuffling; the model flipped its decision on 1.3% of perturbations on average, and flips were zero when options were paraphrased, reversed or shuffled. On safety, Microsoft ran 5,250 requests across 11 benchmarks covering harmful content, jailbreaks and prompt injection, and reported correct refusals with high retained utility without publishing a score. Internal results cited by Microsoft included Xbox Research sorting more than 10,000 feedback items 14 times faster and 200 times cheaper than GPT-6 Sol, Copilot quality control competitive with GPT5.6 Luna and 100 times faster, and Microsoft Discovery producing 46 times more consistent scoring than LLM scoring at three times the speed.

In Microsoft's comparison chart, Microsoft-Decision-1 is listed at 83.5% accuracy, 92.2 calibration and 85 ms median latency; Quyet-1.0-Large at 81.9%, 93.1 and 380 ms; H2O-Lightning-4B at 77.2%, 91.8 and 210 ms; and GPT-6 Luna Decisions at 79.4%, 89.9 and 300 ms. Microsoft's price is $0.042 per million input tokens with free output, while OpenAI's Luna decisions endpoint charges $0.10 per million input tokens. MarkTechPost noted that the latency comparison is not fully apples to apples: Microsoft measured its own model through Foundry, while competitor figures use JevBench's adjusted median rather than raw timings. H2O.ai's model card disputes that adjusted figure, saying it doubles measured time and adds 0.15 seconds; H2O says its model's measured median is 29 ms, not 210 ms. Microsoft-Decision-1 does not appear on the JevBench board. On that board's official composite, H2O-Lightning-4B ranks first at 72.5.

Developers can call Microsoft-Decision-1 from Microsoft Foundry as a generally available 'Direct from Azure' model. On OpenRouter, it runs on the Decisions API, not the chat endpoint; chat completions SDKs will not work. The model card is explicit that the system is not for text generation, open-ended Q&A, chat or translation.

Editor's Summary Nace.AI and Microsoft each released 9B-class decision models that return probabilities over fixed options rather than generated text. Drex 1.5 offers open weights and matches Jev 1.13.0 on a public index while showing weak broad knowledge, and Microsoft-Decision-1 is a hosted, calibrated model that leads Microsoft's own 36-benchmark comparison but remains closed. Both releases push decision scoring as a distinct category for routing, classification, verification and agent control.