AI Affairs, home

Friday 2 October 2026

Technology

AWS releases Strands Decider 2B for local decisions in AI agents

The downloadable model scores predefined answers rather than writing prose. AWS has published training materials, while its reported speed and accuracy depend on the tasks tested.

Aerial view of Amazon Web Services data centers in Umatilla, Oregon
Photo: Tedder, CC BY-SA 4.0, via Wikimedia Commons (cropped)

Amazon Web Services released Strands Decider 2B on 1 October 2026, an open-source model built to choose among predefined answers inside AI-agent workflows, VentureBeat reported. Part of AWS’s experimental Strands Labs project, it is available for developers to download and run themselves rather than call through a hosted AWS service.

Key points

  • Strands Decider 2B scores supplied options instead of generating a written answer.
  • The model is available under Apache 2.0, with training data and scripts published alongside it.
  • AWS reports roughly 72% accuracy for version 19 on the public portion of JevBench.
  • AWS is offering the downloadable model, with no hosted API or per-call charge.

Qwen3.5-2B with an answer-scoring head

Strands Decider starts with Qwen3.5-2B, a pretrained language model. According to AWS, its developers replaced the part that predicts the next word with a small component, called a pointer head, that scores answers supplied to it. They adapted the base model using a rank-16 LoRA update, a way of training additional weights without retraining every weight in the original model. The new scoring component contains just over a million parameters, while the whole model has roughly 2 billion.

That change gives the model a narrower job than a chatbot’s. Given a question and a set of permitted answers, it makes one pass through the input and assigns probabilities to the choices. It could therefore sit at a checkpoint in a larger workflow — deciding which route to take or whether a proposed action fits a request — while a generative model handles work that needs an open-ended answer.

AWS’s example puts such a checkpoint before an agent makes a weather request. In the example, a location is absent from the original request, but the agent proposes using a city it has guessed. The decider checks the proposed action against the conversation, and application code sends the agent back to ask for the missing location.

Checking the weather could lead to a request for a location rather than a forecast based on a guessed place, if a proposed action were checked against the original request before it went ahead.

The questions and thresholds in AWS’s example were chosen by hand. The model supplies scores, while the application determines how to act on them. The example uses an existing Strands mechanism that can allow an action, deny it, request confirmation or return guidance to the agent.

JevBench scores and local response times

AWS assessed accuracy and the quality of the model’s probability estimates on the public portion of JevBench. Its chart puts version 19 at roughly 72% accuracy and a 0.35 Brier score, VentureBeat reported. The latter measures how well probabilities match outcomes: a lower score is better. AWS ranks its model second among public models of roughly the same size on that test, and first among those that publish a complete training recipe.

The comparison plotted by AWS places Mapika decider-2b v11 at around 76% accuracy and a 0.32 Brier score, ahead on both measures. The plotted comparison does not include TypeSafe AI’s Jev. These are results for a public benchmark set, rather than measurements of the model making decisions throughout a deployed agent workflow.

For speed, AWS says short local tasks can take tens of milliseconds and that some inputs on common hardware finish in under 100 milliseconds. A more detailed AWS test of version 18, covering 230 requests on a local Nvidia RTX 3090 and including the HTTP round trip, recorded a median of 106 milliseconds and a 95th-percentile time of 296 milliseconds. The test excluded a 7.7-second first request during server warm-up, and longer inputs generally took longer.

AWS also reported a median of roughly 150 milliseconds for small tasks on an M3 MacBook. Hardware, input length and whether a measurement includes the trip to a server all affect how a speed figure applies to a working system. Running the model locally gives developers control over that arrangement, but also leaves them to provide the computing resources.

Brooker’s experiment reaches Strands Labs

Amazon distinguished engineer Marc Brooker began the project after seeing Jev and building his own decision model, TechCrunch reported. His version briefly reached the top of a JevBench ranking for models of its size before Amazon engineers prepared a Strands Labs release. TypeSafe AI had introduced Jev in September as a model for choosing among predefined options, and OpenAI has previewed a Decisions API of its own.

Brooker told TechCrunch that conversations with AWS customers had pointed to workflow steps that did not need a full language model each time. He described accuracy and well-calibrated confidence scores as goals that must be balanced against retaining the language understanding and knowledge that make the model useful across tasks. TypeSafe’s chief executive, Diogo Almeida, argued that making such models genuinely capable is harder than producing another version of the architecture.

Developers can download Strands Decider 2B from Hugging Face under the Apache 2.0 licence; AWS has published its training data and scripts on GitHub. An Amazon spokesperson said AWS was offering only the open-source model for now, with no hosted API or per-call charge, VentureBeat reported. In AWS’s weather example, the decision about what happens after the model returns its scores remains in the application code.

Topics: Agents, Foundation models, Open source