Craig Stanley
Home / Decisions / Decision theory at work

Setting a threshold you can defend

How to pick the confidence score at which a model acts on its own, asks a person, or stops, using the cost of each kind of mistake.

10 October 2026 · 2 min read · Craig Stanley
In short, explained

A computer helper says how sure it is. We pick a number. If it's surer than that, it can go ahead. If not, it asks a grown-up.

A decision model gives each case a score. The threshold is the score above which it acts without asking. You set it by comparing the cost of a wrong yes with the cost of a wrong no, and by agreeing it with the people affected.

Set the action threshold where the expected cost of a false positive equals that of a false negative, adjusted for review capacity. Publish the threshold, the costs behind it and the review date.

What a threshold does

A decision model gives each case a score, often written as a confidence between 0 and 1. On its own the score does nothing. The threshold turns it into an action: above it the system acts, below it the case goes to a person.

Many systems use two thresholds. Above the top one, act. Below the bottom one, reject or stop. In between, ask a person.

Start from the cost of mistakes

There are two ways to be wrong:

  • A wrong yes: the model acts when it shouldn't. An invoice gets paid that should have been queried.
  • A wrong no: the model holds back when it could have acted. A clean invoice waits a day for a person.

If a wrong yes costs ten times as much as a wrong no, the model should only act when it's very sure. A common starting point is to act when the score is above cost of a wrong yes ÷ (cost of a wrong yes + cost of a wrong no). With costs of £50 and £5, that's 50 ÷ 55, so act above about 0.91. That only works if the scores are calibrated.

Then check capacity

A threshold that sends 60% of cases to people won't survive a busy week. Look at how many cases fall in the "ask a person" band and compare that with the time people actually have. If it's too many, either accept more risk or improve the model before going live.

Make it defensible

Write down the costs you used, the threshold you picked and who agreed it. Set a review date. When someone asks why the model approved a case, you can show the reasoning rather than just the score.

Read next

A question to take awayWhich repeated decision would you trust a cheap model to score first, with a person checking the close calls?

About me

Craig Stanley

Microsoft AI consultant and technical architect, based in Whitley Bay. Over the last few years I've delivered Microsoft 365 Copilot, Copilot Studio agents, Azure AI Foundry work and governance for UK public sector and financial services organisations.

What interests me is the decision underneath the tool: what it costs, what it risks, and whether a small, transparent model can make it better. I write the methods up here and on Substack so anyone can use them.

Find me