Microsoft introduced Microsoft-Decision-1 on 9 October 2026, a decision-scoring model that assigns probabilities to a fixed set of answers rather than writing a response, according to its announcement. The company says it post-trained Qwen3.5-9B to produce a score in a single pass and has made the model available through Microsoft Foundry and OpenRouter.
Key points
- Decision-1 scores predefined options for tasks including classification, routing and verification.
- Microsoft reports 83.5% average accuracy across 36 benchmarks and a median response time of 85 ms.
- The model is available as a hosted API at $0.042 per million input tokens, with free output.
- H2O.ai disputes a timing figure used in Microsoft’s comparison with its model.
Qwen3.5-9B supplies the scoring base
Qwen3.5-9B comes from Alibaba, MarkTechPost reported. Microsoft says its post-trained model takes a question and predefined options, then returns a probability for each option in one structured call. Think of the difference as marking the boxes on a form and recording a confidence beside each mark, rather than composing a paragraph for another system to interpret. Microsoft describes routing, classification, prioritisation and checking proposed actions as intended uses.
Sorting game reviews could become a matter of assigning comments to preset themes instead of producing a fresh description for each one. The probabilities would give the surrounding software a score for each available theme rather than a written answer to interpret.
The model accepts yes-or-no, multiple-choice and rating options, as well as rubrics for grading AI responses and proposed agent actions, Microsoft says. A probability is useful only if it bears a reliable relationship to how often the answer is right: Microsoft says that, on representative cases, a prediction assigned 90% confidence should be correct about nine times out of 10. That requirement matters when an application uses the score to decide whether to proceed or seek review.
Training used public datasets handled under Microsoft’s Open Data process and synthetic data, MarkTechPost reported. The model has a 32,768-token context window and produces JSON through a hosted API. Its weights are not available for download, so developers building on it use the service rather than running their own copy. That differs from the downloadable model covered in AI Affairs’ report on AWS’s Strands Decider 2B.
Microsoft’s 36 benchmarks and JevBench timings
Microsoft compared nine systems on 36 benchmarks comprising 147,137 questions kept blind from training. Its reported average accuracy for Decision-1 was 83.5%, ahead of the 81.9% it recorded for Quyet-1.0-Large. The tests covered tasks including routing, ranking, long-context and multilingual decisions, reasoning and safety. Those figures come from Microsoft’s evaluation of the systems, rather than from the ordinary applications for which the model is intended.
Microsoft’s median, or P50, latency for Decision-1 was 85 ms. The company says that was about 35 times quicker than GPT-6 Sol in its tests, and 2.5 times quicker than H2O-Lightning-4B v1.1. Delay accumulates when one decision must wait for another: Microsoft gives the example of 20 sequential decisions, each taking an additional 100 milliseconds, adding two seconds to a workflow.
The H2O comparison uses different timing methods. Microsoft measured Decision-1 through Foundry, while its competitor timings came from JevBench’s adjusted median, MarkTechPost reported. H2O.ai’s model card says JevBench’s adjustment doubles measured time and adds 0.15 s; it gives H2O-Lightning-4B a measured median of 29 ms rather than the 210 ms used in Microsoft’s chart. Those are two reported timings for different measures, not two measurements made under the same conditions.
Microsoft also tested whether small changes to an input altered the selected answer. It says it changed each request in eight ways and found that Decision-1 switched choices on 1.3% of perturbations on average. The company recorded no switches when it paraphrased option descriptions or reversed or shuffled the options. For software that depends on a fixed choice, the test addresses a practical problem: a rearranged answer list should not, by itself, send the next step down a different path.
Foundry access and Xbox feedback testing
Decision-1 costs $0.042 per million input tokens, with free output. On OpenRouter it uses the Decisions API rather than the chat endpoint, so chat-completions software development kits will not work with it. The fixed response format gives applications a choice and associated probabilities instead of prose, a distinction also reflected in AI Affairs’ coverage of Databricks’ ai_decide function for structured decisions on governed data.
Microsoft says Xbox Research used Decision-1 to sort more than 10,000 pieces of open-ended feedback and reviews into themes chosen by researchers. In that internal test, the company says the model was competitive with GPT-6 Sol on quality while running more than 14 times faster and costing 200 times less. The exercise concerned an existing set of themes, which suits a model built to score supplied answers.
The company also tested Decision-1 on 5,250 requests across 11 safety benchmarks involving harmful content, jailbreak attempts and prompt injection. Microsoft says the model refused harmful behaviour while retaining a high degree of utility. Its published account describes those tests separately from the Xbox feedback exercise and from the 36-benchmark accuracy comparison.