Craig Stanley
Home / Risk / Prioritise / Likelihood, impact and reversibility on one sheet

Likelihood, impact and reversibility on one sheet

A one-sheet scoring method for AI risks that adds reversibility and time to detect to the usual likelihood and impact, with a worked example.

· 3 min read · Craig Stanley
In short, explained

Give each risk three scores: how likely it is, how bad it would be, and how hard it would be to fix. Add them up and start with the biggest.

The usual way to rank risks multiplies how likely something is by how bad it would be. For decisions that AI touches, I add a third score for how hard the mistake is to undo and how long it takes to spot. One sheet, four columns, and the order becomes clear.

Score each register row 1 to 5 for likelihood, impact, irreversibility and detection lag. Priority = likelihood × impact × (irreversibility + detection lag) / 2. It keeps the familiar L×I shape but lifts errors that are hard to undo or slow to surface.

The usual method and its gap

Most risk registers rank rows by likelihood multiplied by impact, each scored from 1 to 5. It's familiar, and people understand it. For AI-assisted decisions it misses two things that change how much a risk matters in practice: whether a wrong decision can be undone, and how long it takes anyone to notice.

A refund paid in error and a queue routed wrongly might score the same on likelihood and impact. But the routing error is fixed within a day when the customer chases. The refund, once paid, is gone. They shouldn't sit at the same priority.

The four scores

Score135
LikelihoodRareHappens most monthsHappens most weeks
ImpactMinor inconvenienceReal cost or harm to one personSerious harm, legal exposure or large cost
IrreversibilityUndone in minutesUndone with effort over weeksCan't be undone
Detection lagNoticed the same dayNoticed within a monthNoticed only at audit, or never

Then I calculate:

Priority = likelihood × impact × (irreversibility + detection lag) ÷ 2

Dividing by two keeps the result on a similar scale to plain likelihood × impact, from 1 to 125. The exact number matters less than the effect: a hard-to-undo, slow-to-spot error rises up the list.

A worked example

The rows and scores below are illustrative, taken from the customer service register in Start a risk register from the decision inventory.

DecisionLIIrrev.LagL × IPriority
Complaint routed to wrong queue4212812
Goodwill credit given against policy3253624
Wrong tone or fact on a sensitive reply2432820
Case not escalated when it should be2434828
Key fact missing from a call summary3325931.5

Using L × I alone, the call summary comes first, then three rows tie at 8. With the full score, the order is call summary, escalation, goodwill credit, reply wording, then routing. Routing drops from joint second to last, because its errors are quick to spot and fix.

How I use it

I'd score as a group of three or four people who know the work, rather than alone. Each person scores privately first, then the group compares. Where scores differ by two or more points, the discussion is likely to be more useful than the score.

I don't rank by score alone. The top three or four rows get a named owner and a control in Controls people actually keep. Everything else gets a review date.

Why multiply at all

Some teams prefer to add the scores. Multiplying means a row needs both likelihood and impact to rank highly, which I think is right: a very likely but trivial error shouldn't outrank a rare, serious one by addition alone. The reversibility and lag pair are averaged, so either one can lift a row without the other.

What I'm still checking

The formula is my own, and I haven't tested it against how a group would rank the same rows by discussion alone. I'd like to try that comparison with a real team. If the ranking matches discussion most of the time, the formula is saving effort. If it often disagrees, the formula or the scale definitions need work.

Sources

This article describes my own scoring method. It uses no external facts or figures. The scale definitions are illustrative and should be adapted to each organisation's risk appetite.

Read next

A question to take awayWho gets told, and how fast, when a decision model starts drifting?

About me

Craig Stanley

Microsoft AI consultant and technical architect, based in Whitley Bay. Over the last few years I've delivered Microsoft 365 Copilot, Copilot Studio agents, Microsoft Foundry (formerly Azure AI Foundry) work and governance for UK public sector and financial services organisations.

What interests me is the decision underneath the tool: what it costs, what it risks, and whether a small, transparent model can make it better. I write the methods up here and on Substack so anyone can use them.

I write this site to learn in public: explaining each idea simply is how I check I understand it. Why I write this site.

Find me