Craig Stanley
Home / Capabilities / Start here: decision models / Where a decision model sits in a Copilot agent

Where a decision model sits in a Copilot agent

How a scoring model like Microsoft-Decision-1 fits inside an agent: the agent gathers the facts, the model scores fixed options, a rule decides.

· 3 min read · Craig Stanley
In short, explained

A helper that does jobs can ask a different, smaller helper one question: "which of these choices is best, and how sure are you?" Then a rule says whether to go ahead or ask a person.

An agent is good at gathering information and doing tasks, but it isn't built to give a reliable score. A decision model is. So the agent collects the facts, asks the decision model to score a fixed list of options, and a rule you've set decides whether the agent acts, asks a person or stops.

Separate the roles: the agent assembles state and executes; a decision model scores a bounded question and returns per-option probabilities; an explicit threshold policy maps scores to act, ask or stop. Log state, scores, route and outcome. Validate thresholds on labelled cases before automating.

Two different tools

A Copilot agent is good at the messy parts of work: reading a request, finding the right documents, filling a form, sending an email. It's built on a generative model, which produces text. If you ask it "should I approve this?", it will give you an answer, but the answer is a piece of text rather than a measured probability.

A decision model does one narrower thing. Microsoft's documentation for Microsoft-Decision-1 describes it as a model that "doesn't generate a free-form response or written rationale". You give it some text or JSON describing the case and a question with fixed options. It returns a probability for each option, a yes or no probability, or a score on an ordered scale.

Put together, the agent handles the work and the decision model handles the judgement, and a rule you've written handles what happens next.

The shape of it

  1. The agent gathers the case: the request, the relevant records and the policy that applies.
  2. The agent sends that case, with a fixed question and options, to the decision model.
  3. The decision model returns a probability for each option.
  4. A threshold rule, set by people in advance, turns the probabilities into a route: act, ask a person, or stop.
  5. The agent carries out the route and logs what happened.

Step 4 is the one I'd guard most carefully. It shouldn't live in the agent's prompt, where it can drift with a wording change. It belongs in configuration that someone owns, as described in Act, ask a person, or stop.

A worked example

This scenario and its numbers are illustrative.

A service desk agent receives "I can't get into the finance system since this morning." It looks up the user's recent tickets and any outage notices, then asks the decision model: which team should handle this? The options are access, outage, training and cannot tell.

OptionProbability
Access0.78
Outage0.15
Training0.04
Cannot tell0.03

The team's rule says route automatically if the top option is above 0.85, otherwise ask a service desk analyst to confirm. At 0.78 the agent passes the ticket to a person with the model's suggestion attached. Microsoft's documentation recommends including an abstention option like "cannot tell" and treating a low top probability as a reason to escalate, which is what this rule does.

How it connects in practice

Microsoft-Decision-1 is deployed in Microsoft Foundry and called through an API. What the model itself does and costs is covered in Microsoft-Decision-1: a first look. An agent built in Copilot Studio or Foundry could call that API as one of its steps. I haven't built this end to end in Copilot Studio yet, so I'm describing the pattern rather than a tested recipe. Microsoft's Copilot Studio billing page notes that models from Foundry are billed separately from Copilot Credits, so the decision model's cost shows up on the Azure bill.

What the split buys you

Keeping scoring separate makes the decision testable. You can run the decision model on a few hundred past cases with known answers and check whether its 80% calls are right about 80% of the time, which is what calibration means. That's much harder to do with an answer buried in generated text.

What I'm still checking

Microsoft's documentation says scores "can change based on how you phrase or order questions and options" and suggests testing option order. I want to see how large that effect is on a real routing task before I'd trust a single threshold across wording changes.

Sources

Read next

A question to take awayWhich of these do you already pay for and not use?

About me

Craig Stanley

Microsoft AI consultant and technical architect, based in Whitley Bay. Over the last few years I've delivered Microsoft 365 Copilot, Copilot Studio agents, Microsoft Foundry (formerly Azure AI Foundry) work and governance for UK public sector and financial services organisations.

What interests me is the decision underneath the tool: what it costs, what it risks, and whether a small, transparent model can make it better. I write the methods up here and on Substack so anyone can use them.

I write this site to learn in public: explaining each idea simply is how I check I understand it. Why I write this site.

Find me