If you say you're nearly sure lots of times, you should be right nearly every time. If you're often wrong when you say you're sure, you need to say "sure" less.
Being calibrated means your confidence matches your results. Of all the times you said "80% sure", about 80% should turn out right. You can check this for people and for AI models by keeping score.
Calibration compares stated probabilities with observed frequencies. Bucket predictions by confidence, compare each bucket's hit rate with its stated confidence, and recalibrate the model or retrain judgement where they diverge. Thresholds are only meaningful on calibrated scores.
Why it matters
A model that says "0.9" should be right about nine times in ten. If it's really right six times in ten, any threshold you set on that score is wrong, and every expected value calculation built on it is too high.
People have the same problem. Most of us are overconfident on hard questions and underconfident on easy ones.
How to check
- Keep score. For each prediction, record the confidence given and whether it turned out right.
- Group by confidence. Put predictions into bands: 50–60%, 60–70%, and so on.
- Compare. For each band, work out the share that were right. A calibrated forecaster's 70–80% band is right about 75% of the time.
You need enough cases for this to mean anything. Fifty predictions per band is a reasonable minimum. Fewer than twenty and the numbers will bounce around.
What to do with the results
If a model is overconfident, you can rescale its scores so they match observed rates. Common methods include Platt scaling and isotonic regression. If a person is overconfident, showing them their own score history is often enough to change how they estimate.
A habit worth building
Ask people to put a number on predictions in meetings: "70% we hit the date." Write it in the decision record. Over a few months you'll learn whose estimates to trust, and everyone's estimates get better.