TL;DR — Key Takeaways
- AWS’s Strands Labs released Strands Decider 2B, an open source model built to make bounded decisions inside AI agent workflows without generating text.
- Decider 2B answered 72.3% of 231 public JevBench tasks.
- AWS is using Strands Labs to test whether specialized decision models can handle routine agent work while larger models are reserved for deeper reasoning tasks.
AWS is testing a different way to divide work inside AI agents. This week, the company’s Strands Labs team released Strands Decider 2B, part of an emerging class of models known as decision models.
Strands Decider 2B is an open source, specialized model that can make routine decisions in an agent workflow. It is designed to choose among predefined options, assign scores and return confidence estimates without generating text. Potential use cases include model routing, tool selection, policy classification, guardrails and evaluations, according to AWS.
Agent workflows can be inefficient and use more tokens than necessary. An agent may repeatedly call a general-purpose language model for decisions that don’t need generated text, like deciding which tool should handle a request or whether an action meets a policy. Strands says Decider can handle those routine decisions, reserving more capable models for work that needs deeper reasoning (and typically, more tokens).
Decider 2B starts with the pretrained Qwen3.5-2B-Base model but removes the language modeling head, the output layer that turns the model’s internal representations into generated text. Strands replaces it with a roughly 1-million-parameter pointer head that scores the options supplied with a question instead of generating text. The resulting Decider 2B model contains about 1.9 billion parameters and evaluates available choices in a single forward pass without a token-by-token decoding loop.
That narrow design limits what the model can do. Decider cannot write code, summarize documents, carry on a conversation, or excel at complex reasoning tasks, but that’s not its purpose. The model is meant to prioritize speed and predictable output formats, along with confidence scores that developers can use to decide whether to accept Decider’s choice or route the task to a larger model.
On 231 public tasks from JevBench, an independent benchmark for similar decision models, Decider 2B answered 167 correctly, or 72.3%. Strands reports a median latency of 115 milliseconds per JevBench question on an Nvidia RTX 3090. Performance varied greatly by difficulty, however. Using Strands’ own difficulty split, the model scored 100% on the easy tier and 50.5% on the hard tier. Strands also advises that the 231 tasks are a small sample and that differences of fewer than about 10 tasks between individual runs should be treated cautiously because of training variation.
The confidence scores also have limits. Decider 2B’s model card says they were tested and adjusted using a separate set of short classification examples, so developers should check how well those scores hold up on their own workloads before using them to automate decisions. Long, multi-step documents and unfamiliar scoring rubrics are also among the model’s documented weaknesses.
AWS is not alone in exploring specialized decision models. TypeSafe AI introduced Jev in September as what it calls a ‘System One Model,’ designed to turn unstructured inputs into typed probabilistic decisions. Strands specifically credits Jev with bringing attention to the approach.
Strands says it has seen early success using decision models for routing, tool selection and other agent functions, but the launch materials do not name production users or provide measured cost reductions or large-scale deployment results. For now, Decider 2B is an experiment, not a new managed AWS service.
AWS created Strands Labs in February as a home for such experimental agent projects, keeping them separate from the production Strands Agents SDK. For now, that separation leaves AWS room to test whether specialized decision models can deliver enough reliability and efficiency to justify a place in production agent stacks.

