Craig Stanley
Home / Roadmap / What I'm building / Threshold calculator

Threshold calculator

Design notes for a threshold calculator I plan to build: enter what a wrong yes and a wrong no cost, and it shows how sure a model must be before acting.

· 3 min read · Craig Stanley
In short, explained

Sometimes saying yes by mistake is worse than saying no by mistake. This little tool will ask how bad each mistake is, and tell you how sure you need to be before saying yes.

A decision model gives a confidence score, but someone has to decide how high is high enough. That line depends on what mistakes cost. The calculator will take the cost of a wrong yes and a wrong no, plus the cost of asking a person, and give you the lines for act, ask and stop.

Planned static tool: computes the cost-minimising probability threshold t = C(wrong yes) ÷ (C(wrong yes) + C(wrong no)) and, with a review cost R, two thresholds for act (p ≥ 1 − R ÷ C(wrong yes)) and stop (p ≤ R ÷ C(wrong no)). Inputs and outputs shareable by URL. Not yet built.

Status

This isn't built yet. This page is the design, written first so I can check I understand the maths before I write any code. When it's built, it'll live on this site and the code will be public.

The idea in one paragraph

A decision model, or a person, usually ends up with a probability: "I'm 85% sure this invoice is fine." The question is whether 85% is enough to pay it. The answer depends on what a mistake costs. If paying a bad invoice costs ten times more than delaying a good one, you need to be very sure. Setting a threshold you can defend explains the reasoning. The calculator does the arithmetic.

The two-way version

You enter two numbers: what a wrong yes costs and what a wrong no costs. The threshold is:

threshold = wrong yes ÷ (wrong yes + wrong no)

Say yes when your probability is at or above the threshold. Below it, say no.

The three-way version

Most real processes have a third option: ask a person. That has a cost too, mainly someone's time. Add a review cost, and you get two lines:

  • Act when your probability is at least 1 − (review cost ÷ wrong yes).
  • Stop when your probability is at most review cost ÷ wrong no.
  • Ask a person in between.

These come from comparing expected costs. Acting has an expected cost of (1 − p) × wrong yes. Stopping costs p × wrong no. Asking costs the review cost. Each line is where two of those are equal.

A worked example

These figures are illustrative. An invoice check where paying a bad invoice costs £50 to put right, delaying a good one costs £5 in goodwill and admin, and a person checking an invoice costs £2 of their time.

Two-way: threshold = 50 ÷ (50 + 5) = 0.91. Pay automatically only when the model is at least 91% sure.

Three-way:

LineSumResult
Act at or above1 − (2 ÷ 50)0.96
Stop at or below2 ÷ 50.40
Ask a personBetween the two0.40 to 0.96

Adding the review option raises the bar for acting automatically, from 0.91 to 0.96, because a cheap check is now available for the uncertain cases.

What it'll show

The plan is a single page with three inputs and a chart. The chart will show the three expected costs as lines across probabilities from 0 to 1, so you can see where each option is cheapest. The inputs will be in the page address, so a result can be shared or saved in a decision record. Nothing will be sent anywhere; it'll all run in the browser.

Where I got stuck

The three-way version breaks if the review cost is higher than the expected cost of acting or stopping everywhere. In that case the "ask" band disappears, and the stop line can end up above the act line. The calculator has to detect that and fall back to the two-way answer, with a note explaining why. I worked that out only by trying awkward numbers on paper, which is exactly why I wrote this page before the code.

I'm also still checking how to show that the probability itself might be wrong. A threshold is only as good as the model's calibration.

Sources

This is a design for my own tool. The threshold formulas are standard expected-cost reasoning; the figures are illustrative.

Read next

A question to take awayWhich item on this list changes a decision your team has already made?

About me

Craig Stanley

Microsoft AI consultant and technical architect, based in Whitley Bay. Over the last few years I've delivered Microsoft 365 Copilot, Copilot Studio agents, Microsoft Foundry (formerly Azure AI Foundry) work and governance for UK public sector and financial services organisations.

What interests me is the decision underneath the tool: what it costs, what it risks, and whether a small, transparent model can make it better. I write the methods up here and on Substack so anyone can use them.

I write this site to learn in public: explaining each idea simply is how I check I understand it. Why I write this site.

Find me