Craig Stanley
Home / Decisions / Decision theory at work

Calibration: are your 80% calls right 80% of the time?

How to check whether a person's or a model's confidence matches reality, with a simple method any team can run.

10 October 2026 · 2 min read · Craig Stanley
In short, explained

If you say you're nearly sure lots of times, you should be right nearly every time. If you're often wrong when you say you're sure, you need to say "sure" less.

Being calibrated means your confidence matches your results. Of all the times you said "80% sure", about 80% should turn out right. You can check this for people and for AI models by keeping score.

Calibration compares stated probabilities with observed frequencies. Bucket predictions by confidence, compare each bucket's hit rate with its stated confidence, and recalibrate the model or retrain judgement where they diverge. Thresholds are only meaningful on calibrated scores.

Why it matters

A model that says "0.9" should be right about nine times in ten. If it's really right six times in ten, any threshold you set on that score is wrong, and every expected value calculation built on it is too high.

People have the same problem. Most of us are overconfident on hard questions and underconfident on easy ones.

How to check

  1. Keep score. For each prediction, record the confidence given and whether it turned out right.
  2. Group by confidence. Put predictions into bands: 50–60%, 60–70%, and so on.
  3. Compare. For each band, work out the share that were right. A calibrated forecaster's 70–80% band is right about 75% of the time.

You need enough cases for this to mean anything. Fifty predictions per band is a reasonable minimum. Fewer than twenty and the numbers will bounce around.

What to do with the results

If a model is overconfident, you can rescale its scores so they match observed rates. Common methods include Platt scaling and isotonic regression. If a person is overconfident, showing them their own score history is often enough to change how they estimate.

A habit worth building

Ask people to put a number on predictions in meetings: "70% we hit the date." Write it in the decision record. Over a few months you'll learn whose estimates to trust, and everyone's estimates get better.

Read next

A question to take awayWhich repeated decision would you trust a cheap model to score first, with a person checking the close calls?

About me

Craig Stanley

Microsoft AI consultant and technical architect, based in Whitley Bay. Over the last few years I've delivered Microsoft 365 Copilot, Copilot Studio agents, Azure AI Foundry work and governance for UK public sector and financial services organisations.

What interests me is the decision underneath the tool: what it costs, what it risks, and whether a small, transparent model can make it better. I write the methods up here and on Substack so anyone can use them.

Find me