A helper that does jobs can ask a different, smaller helper one question: "which of these choices is best, and how sure are you?" Then a rule says whether to go ahead or ask a person.
An agent is good at gathering information and doing tasks, but it isn't built to give a reliable score. A decision model is. So the agent collects the facts, asks the decision model to score a fixed list of options, and a rule you've set decides whether the agent acts, asks a person or stops.
Separate the roles: the agent assembles state and executes; a decision model scores a bounded question and returns per-option probabilities; an explicit threshold policy maps scores to act, ask or stop. Log state, scores, route and outcome. Validate thresholds on labelled cases before automating.
Two different tools
A Copilot agent is good at the messy parts of work: reading a request, finding the right documents, filling a form, sending an email. It's built on a generative model, which produces text. If you ask it "should I approve this?", it will give you an answer, but the answer is a piece of text rather than a measured probability.
A decision model does one narrower thing. Microsoft's documentation for Microsoft-Decision-1 describes it as a model that "doesn't generate a free-form response or written rationale". You give it some text or JSON describing the case and a question with fixed options. It returns a probability for each option, a yes or no probability, or a score on an ordered scale.
Put together, the agent handles the work and the decision model handles the judgement, and a rule you've written handles what happens next.
The shape of it
- The agent gathers the case: the request, the relevant records and the policy that applies.
- The agent sends that case, with a fixed question and options, to the decision model.
- The decision model returns a probability for each option.
- A threshold rule, set by people in advance, turns the probabilities into a route: act, ask a person, or stop.
- The agent carries out the route and logs what happened.
Step 4 is the one I'd guard most carefully. It shouldn't live in the agent's prompt, where it can drift with a wording change. It belongs in configuration that someone owns, as described in Act, ask a person, or stop.
A worked example
This scenario and its numbers are illustrative.
A service desk agent receives "I can't get into the finance system since this morning." It looks up the user's recent tickets and any outage notices, then asks the decision model: which team should handle this? The options are access, outage, training and cannot tell.
| Option | Probability |
|---|---|
| Access | 0.78 |
| Outage | 0.15 |
| Training | 0.04 |
| Cannot tell | 0.03 |
The team's rule says route automatically if the top option is above 0.85, otherwise ask a service desk analyst to confirm. At 0.78 the agent passes the ticket to a person with the model's suggestion attached. Microsoft's documentation recommends including an abstention option like "cannot tell" and treating a low top probability as a reason to escalate, which is what this rule does.
How it connects in practice
Microsoft-Decision-1 is deployed in Microsoft Foundry and called through an API. What the model itself does and costs is covered in Microsoft-Decision-1: a first look. An agent built in Copilot Studio or Foundry could call that API as one of its steps. I haven't built this end to end in Copilot Studio yet, so I'm describing the pattern rather than a tested recipe. Microsoft's Copilot Studio billing page notes that models from Foundry are billed separately from Copilot Credits, so the decision model's cost shows up on the Azure bill.
What the split buys you
Keeping scoring separate makes the decision testable. You can run the decision model on a few hundred past cases with known answers and check whether its 80% calls are right about 80% of the time, which is what calibration means. That's much harder to do with an answer buried in generated text.
What I'm still checking
Microsoft's documentation says scores "can change based on how you phrase or order questions and options" and suggests testing option order. I want to see how large that effect is on a real routing task before I'd trust a single threshold across wording changes.
Sources
- Microsoft Learn, Deploy and use Microsoft-Decision-1 in Microsoft Foundry, accessed 11 October 2026.
- Microsoft Learn, Billing rates and management, Copilot Studio (Foundry models billed separately), accessed 11 October 2026.