Evidence at a glance

Qwen3.5-9B, 32,768 tokenEvidence
9 , 36 , 147,137Evidence
83.5%, Quyet-1.0-Large 81.9%Evidence
p50 85 ,p95 125Evidence
Quyet-1.0-Large 4.5xEvidence
GPT-6 Sol 35x; 3.01Evidence

From Generating Text to Scoring Options

Microsoft has released Microsoft-Decision-1, a decision-scoring model aimed at routing, classification, verification, and agent control. Post-trained from Alibaba’s Qwen3.5-9B, it takes a situation, a question, and a fixed set of options, then returns a probability for each option in one pass rather than generating an explanation or a prose answer. It is currently offered as a hosted API through Microsoft Foundry and OpenRouter, and its exact parameter count has not been disclosed.

This may look like a change in output format, but it changes where the model sits in a software system. A general-purpose generator often produces text that an application must then parse or evaluate. Decision-1 aims to return JSON scores that code can consume directly, potentially removing a translation step when the task already involves choosing among a limited set of answers.

Fixed Options Simplify Integration—and Constrain the Question

The model card lists yes-or-no questions, multiple choice, ratings, classification, rubric-based grading of AI responses, and checks of whether content is grounded in supplied evidence. It also supports an abstention option such as “cannot tell.” The common pattern is that the application defines the possible answers, the model scores them, and the result comes back as JSON without an explanation.

That design can fit into existing workflows, such as scoring an agent’s proposed tool call before execution or routing a request to a suitable model. But because the application supplies the options, it is responsible for defining the boundaries of the question. If an important case is missing from the choices, more precise probabilities cannot produce it. An abstention option can represent uncertainty, but it does not replace careful design and review of the answer set.

A Leading Score Is Not a Universal Win

Microsoft says the model achieved the highest average accuracy, 83.5%, in a comparison of nine systems across 36 benchmarks and 147,137 questions, ahead of Quyet-1.0-Large at 81.9%. The benchmarks were described as held out from training. However, Decision-1 answered only 23 of the 36, so its 83.5% is the average across the benchmarks it could answer, not across the full set. Microsoft also reports a calibration score of 92.2, slightly below Quyet-1.0-Large’s 93.1.

These figures support a bounded conclusion: on the tasks Microsoft tested and the model could handle, Decision-1 had strong average accuracy and calibration. They do not show that it will outperform alternatives on every classification or control task. Microsoft ran the comparison, and the available material does not provide the full benchmark-by-benchmark results or task composition. Being held out from training is also not a substitute for testing against the distribution of production data.

Low Latency and Price Need End-to-End Context

Microsoft reports a latency of 85 milliseconds at p50 and 125 milliseconds at p95. Input costs $0.042 per million tokens, while output is free. Those figures make the model look suitable for frequent, short decisions such as request routing or pre-execution checks. Fixed options and structured output may also lower the cost of an individual call compared with a larger text-generating model. But the available material does not disclose the hardware or offer self-hostable weights or quantized variants, so it cannot establish local deployment costs.

Latency comparisons need particular care. Microsoft measured its figures through Foundry, while rival models were represented by JevBench’s adjusted medians. The material also cites H2O.ai’s claim that this adjustment doubles measured time and adds 0.15 seconds, alongside its own reported median of 29 milliseconds. These are not measurements under a common setup, so headline speed ratios should not decide an architecture choice. Teams should account for request length, call volume, retries, and network overhead, then measure end-to-end latency in their own service environment.

Use It at a Decision Point, Not as a Safety Policy

For engineering teams, a sensible starting point is not to hand an entire agent over to Decision-1, but to select a narrow task with explicit options and historical data for review. For example, a team could compare its tool-call scores with human decisions and check whether errors cluster around particular options or inputs. Microsoft reports a 1.3% decision-flip rate under perturbation testing, but the material does not describe the perturbations in detail, so that figure cannot stand in for stability in a real deployment.

The model’s score should also remain separate from execution authority. Decision-1 provides no explanation, its hosted API does not expose inspectable weights, and OpenRouter notes that the weights are updated over time while the API shape remains fixed. For high-risk actions, a probability should feed into rules, thresholds, human review, or other verification mechanisms rather than serve as the sole authorization. The prudent view is to treat the model as a potentially faster, cheaper structured-judgment component, measure its benefits and failure costs on your own data, and automate only the low-risk steps that the evidence supports.