I want to build a simple example that shows how to make a computer help with one decision, safely. It will pick from a few fixed answers, say how sure it is, and write down every choice.
This is a plan for a small public starter project. It shows how to run one decision on Microsoft Foundry: a fixed set of options, a confidence score, thresholds for act, ask or stop, a log of every decision and a set of test cases. It isn't built yet. These notes describe what it'll contain and why.
Planned reference repo: a Foundry model deployment (data zone chosen for residency) behind a thin scoring service returning a constrained option plus probability, threshold routing from a config file, decision records written to a store, an evaluation set of labelled past cases, and tracing. Not yet built; no performance claims.
Status
Nothing here is built yet. This page is the design for a starter project I'd like to publish as a public repository. I'm writing it first because explaining the design is the quickest way to find out what I don't understand about it.
What the starter is for
Most AI examples show a chat. A decision is different: there's a short list of allowed answers, a cost to getting it wrong, and a need to show later why a choice was made. The starter shows one decision handled that way on Microsoft Foundry, end to end, small enough to read in an afternoon.
The parts
- A config file lists the allowed answers, such as "approve", "refer" and "decline". The model must return one of them. Anything else counts as "refer".
- One model is deployed in Foundry. The deployment type is a deliberate choice, covered in Choosing a deployment type in Foundry. For a UK or EU organisation with residency concerns, I'd start with an EU Data Zone deployment instead of Global Standard, which is the default.
- A short scoring program sends the case to the model, with instructions to choose one option and give a probability. It checks the reply is valid.
- The routing thresholds from the threshold calculator live in the same config file, so changing a threshold means a reviewed edit to one file and no code changes.
- Every case is written to a log in the shape of the site's decision record schema: input, option, probability, route taken, model version and time.
- A set of labelled past cases is run before any change goes live, reporting accuracy and calibration.
Why Foundry for this
Copilot Studio could handle a similar flow for people working in Microsoft 365. Foundry suits the starter because the decision usually needs to run inside another system, such as a case management tool, and because it gives direct control over the model, deployment region and tracing. Microsoft says Foundry agents can also be published to Teams and Microsoft Copilot, so the same decision could later appear there.
A worked example
This scenario and its numbers are illustrative. A council's grants team receives about 300 small-grant applications a quarter. The starter is configured with three options: eligible, ineligible and refer. Thresholds come from the calculator: act at 0.96 or above, stop at 0.40 or below, refer between.
| Route | Share of cases (illustrative) | What happens |
|---|---|---|
| Act: eligible | 45% | Moves to scoring by a panel |
| Stop: ineligible | 15% | Gets a standard letter, checked by a person before sending |
| Refer | 40% | An officer reviews it |
Even in this example, nothing is decided by the model alone. It sorts the queue, and a person still checks every rejection letter. That's the default I'd build in, because a wrong rejection is the costliest mistake.
Running costs
Foundry's standard deployments are billed per token. Microsoft also offers a Batch option at 50% less than standard pricing with a 24-hour target turnaround. For a queue that doesn't need instant answers, such as overnight sorting, Batch looks like a good fit. I haven't priced it yet, and I won't publish cost estimates until I've run it.
What I'm still checking
I don't yet know how well a general model's stated probabilities match reality for this kind of task. That's the biggest open question, and the test set exists to answer it. Microsoft says some Foundry monitoring dashboards are in preview, so I plan to keep my own log as well and rely on that.
Sources
- Microsoft Learn, What is Microsoft Foundry?, accessed 11 October 2026.
- Microsoft Learn, Agents in Microsoft Foundry, accessed 11 October 2026.
- Microsoft Learn, Deployment types for Microsoft Foundry Models, accessed 11 October 2026.
- The design is my own; the scenario and figures are illustrative.