AWS Strands Labs Releases Strands Decider 2B, an Open Source Decision Model That Returns Choices in About 115 ms
AWS Strands Labs has released Strands Decider 2B, a 1.9-billion-parameter open source model that returns a choice, a yes/no probability or a score instead of text. Weights ship under Apache-2.0, and it runs locally with a median latency of 115 ms on an RTX 3090.
Weights are published on Hugging Face under Apache-2.0, and installing strands-decider provides a command line interface and an HTTP server. The bundled server binds to 127.0.0.1 with no authentication, so the team says production use requires an added auth layer. No hosted inference provider serves the model yet.
Decision models, also called System One models, became a category after TypeSafe AI launched Jev last month. An LLM can produce arbitrary output, while a decision model only picks between options or rates on a scale. Strands Decider supports three question types: choice, which picks one of N options; noul, which returns a yes/no probability between 0 and 1; and score, which returns a level on an ordered rubric. Every answer comes from the allowed options and carries a confidence. The team states the model is worse than reasoning models on complex problems and is unsuited for coding, chat or summarization.
The architecture starts from Qwen3.5-2B-Base, discards the language-modelling head and replaces it with a pointer head of about 1 million parameters that compares the hidden state at the position against the hidden state at each option's last token. One forward pass yields the result, with no decoding loop. The torso uses a rank-16 LoRA and the head runs in fp32. Label sets come from the request, so nothing caps the option count. The released checkpoint is v19. Asking several questions about one text is cheap: the state is read once, and each extra question adds only its own tokens.
The team measures accuracy and calibration on the public set of JevBench, a third-party benchmark for Jev-class models. Published v19 figures give JevBench v1 public accuracy of 0.723, or 167 of 231 tasks, with a Brier score of 0.342 and an expected calibration error of 0.052. Tier accuracy is 1.000 on easy, 0.875 on standard and 0.505 on hard. Latency on an RTX 3090 is 115 ms median and 299 ms p95; on an M3 Pro it is 153 ms warm median under 300 tokens. On the v1.4.2 board of September 25, v19 ranked third of 33 in the 2B class, and first of 30 after excluding three models just over 2B.
The repository also flags a caveat. Mapika's newer decider-2b v11 scores 175 of 231 on the Strands harness, eight tasks ahead. Strands Decider was not on the newer v1.5.4 composite board at the time of writing. Calibration is presented as the practical win: on unseen short classification tasks, answers at 0.9 confidence or higher were right about 95 percent of the time, and the team advises confirming or escalating below that threshold.
In the comparison published with the release, Jev 1.13.0 from TypeSafe AI is a closed hosted API with an undisclosed base model and vendor-reported latency of 70 to 500 ms. Mapika's decider-2b is open, with 0.710 public accuracy on JevBench, and FlyMy.AI's Decision 2B is a 2.5B dense model built on MiniCPM5-2B with 0.753. The reported latencies come from different hardware and harnesses, so they are not directly comparable. The Strands team says its full training recipe and data list are published, while the other two are not verified.
The team reports early success in model routing, tool selection, argument checking, triage, guardrails, evaluations and hybrid agents, in which an LLM makes the hard calls and the decider handles rote ones. A repository example gates a weather tool call inside a Strands agent: a before_tool_call intervention asks two yes/no questions, whether the arguments are grounded in what the user said and whether calling now is premature, and if the agent guessed a city it asks the user instead. From the CLI, routing the message "Help! My payouts have been failing for 3 days!" across billing, sales and retail returns billing with confidence 0.768.
Editor's Summary
AWS Strands Labs has published Strands Decider 2B, a 1.9-billion-parameter decision model that discards text generation in favour of choices, yes/no probabilities and rubric scores, released under Apache-2.0 with a CLI and local server. It scores 0.723 on the public JevBench v1 set with an expected calibration error of 0.052 and runs at a reported 115 ms median on an RTX 3090. The repository notes that a competing 2B decider from Mapika leads it by eight tasks on the same harness, and that no hosted inference provider offers the model yet.