Craig Stanley
Home / Capabilities / Start here: decision models / Microsoft-Decision-1: a first look

Microsoft-Decision-1: a first look

Microsoft's new decision model scores a fixed set of options and returns a probability for each. What it does, what it costs and what to check first.

· 4 min read · Craig Stanley
In short, explained

Microsoft has made a computer helper that doesn't chat. You give it a question and a short list of answers, and it tells you how likely each answer is to be right. People still decide the hard ones.

Microsoft-Decision-1 is a small AI model that picks from answers you give it, such as "billing, engineering or support", and says how sure it is about each one. It's cheap and fast because it doesn't write any text. You still need to test it on your own examples and choose when a person should step in.

Microsoft-Decision-1, announced on 9 October 2026, is a decision-scoring model in Microsoft Foundry. It takes a state (text or JSON) and typed questions (yes/no, choice or score) and returns a probability per option in one pass, with no generated text. Input costs $0.042 per million tokens and output is free (Microsoft, 9 October 2026). Microsoft's benchmark claims aren't independently verified yet, so validate it on your own labelled cases before you set a threshold.

What Microsoft announced

On 9 October 2026 Microsoft announced Microsoft-Decision-1, "our new model for fast decision-scoring", available in Microsoft Foundry and through OpenRouter. Microsoft describes it as a model for routing, classification, prioritisation, verification and workflow control. The announcement was written by Achint Srivastava, VP of Software Engineering in Microsoft's Office of the CTO.

It's the kind of model this site keeps describing: the facts of a case and a fixed list of options go in, and a probability for each option comes out. If that idea is new to you, start with what a decision model is and isn't.

How it works

According to Microsoft's documentation, you send the model a state (the facts, as text or JSON) and one or more typed questions. There are three question types:

QuestionTypeWhat comes back
Is this true?noulA probability from 0 to 1
Which option is it?choiceThe selected option and a probability for every option
How much?scoreA value on an ordered scale and a probability for each level

Several questions about the same state can share one request, so a support ticket could be routed, rated for severity and checked for a repeat contact in one call.

What it doesn't do matters as much. The model is text-only, reads up to 32,000 tokens in a single pass, and doesn't write explanations or rationales. Microsoft says it was built by post-training the open-weight Qwen3.5-9B model, and that it will "soon rebase it on other models, including Microsoft AI (MAI) and OpenAI".

What it costs

Microsoft's announcement gives the price as $0.042 per million input tokens, with output tokens free. That's a US dollar list price as published on 9 October 2026. I haven't seen a sterling price or any regional differences confirmed yet, so check the price on your own agreement before you build a business case on it.

Microsoft's documentation lists two deployment types: Global Standard, and Data Zone Standard in selected regions. If your data has to stay in a particular region, check that Data Zone Standard is offered there before you plan around it.

What Microsoft claims, and what's still unconfirmed

Microsoft says the model achieved the highest accuracy in its own 36-benchmark comparison of nearly 150,000 questions, was the fastest model it measured, and changed its decision on 1.3% of reworded or reordered requests on average. Those are Microsoft's own figures from its own tests. The announcement has already been updated once since publication to add comparisons, and I haven't found an independent benchmark yet. Treat them as claims to check, not facts to plan on.

Two other things I couldn't confirm from Microsoft's own pages on 11 October 2026:

  • Release status. Microsoft's model list shows Microsoft-Decision-1 without the Preview label that sits next to some neighbouring models, but I haven't found a Microsoft statement giving a general availability date.
  • Real-world calibration. Calibration means a 90% score is right about nine times in ten. Microsoft says the probabilities are calibrated, and its documentation adds that calibration is strongest on familiar task types. Only your own data can tell you how well that holds for your decisions.

What Microsoft says to watch out for

The documentation is unusually direct about the limits, and they match the methods on this site:

  • Scores can change with how you word or order the questions and options, and a badly framed question still gets a confident-looking score.
  • The model may carry biases from its base model and training data. Microsoft says not to use its scores as the only basis for decisions about individuals.
  • For decisions about credit, employment, housing, healthcare or legal matters, use it for decision support with meaningful human review, and tell people when AI contributed to a decision about them.

How I'd try it

This is the site's method applied to a new tool, using a support queue as the example. All the numbers below are illustrative.

  1. Write the options down first. Billing, engineering, support, and "cannot tell". Microsoft's own guidance suggests including an abstention option like that.
  2. Label a sample. Take a few hundred past tickets where you know the right answer.
  3. Check calibration. Group the model's answers by confidence and see whether the 90% answers are right about 90% of the time. The page on calibration shows how.
  4. Set a threshold from the cost of mistakes. Auto-route above, say, 0.9; send everything else to a person. The page on setting a threshold you can defend explains how to pick the number.
  5. Shuffle the options and rerun. If the answers move when you reorder the list, fix the wording before you go further.
  6. Record the decision. Write down the threshold, the test results and who agreed them in a decision record.

The pattern is the one in act, ask a person, or stop: let the score act on the clear cases and keep people on the close calls.

Sources

All checked on 11 October 2026.

Read next

A question to take awayWhich of these do you already pay for and not use?

About me

Craig Stanley

Microsoft AI consultant and technical architect, based in Whitley Bay. Over the last few years I've delivered Microsoft 365 Copilot, Copilot Studio agents, Microsoft Foundry (formerly Azure AI Foundry) work and governance for UK public sector and financial services organisations.

What interests me is the decision underneath the tool: what it costs, what it risks, and whether a small, transparent model can make it better. I write the methods up here and on Substack so anyone can use them.

I write this site to learn in public: explaining each idea simply is how I check I understand it. Why I write this site.

Find me