Craig Stanley
Home / Risk / Mitigate and track / Controls people actually keep

Controls people actually keep

Most AI risk controls fail because they depend on people remembering. Four tests for a control that will still be working in six months.

· 3 min read · Craig Stanley
In short, explained

A rule that relies on people remembering will be forgotten. A good rule happens by itself, or is so quick that nobody minds doing it.

Controls like "always check the AI's answer" sound good and quietly stop happening. Controls that last are built into the tool, take seconds, have a named owner and leave a trace you can count. I test every control against those four points before it goes in the register.

Prefer controls that are structural (a threshold, a hard stop, a permission), low-effort at the point of use, owned by a named person and observable in data. Replace "human in the loop" with a specific, sampled, logged review step. Measure the control's own execution rate.

The problem with "human review"

The most common control in AI risk registers is some version of "a person reviews the output". It's sensible, and it often stops happening without anyone deciding to stop it. When the model is usually right, checking feels like a waste of time. Within a few weeks, reviewing turns into glancing, and glancing turns into approving.

The people involved aren't to blame. Anyone asked to do the same check hundreds of times, when it almost never finds anything, will drift the same way. A control that depends on sustained attention to rare events is weak by design.

Four tests

Before I put a control in the register, I check it against four questions.

  1. Is it built into the system, or does it rely on someone remembering? A threshold that routes low-confidence cases to a person happens every time. A note in the guidance saying "check carefully" doesn't.
  2. How long does it take at the moment it's used? If a check takes more than a minute on a high-volume decision, it will be skipped. If it takes ten seconds, it might not be.
  3. Who owns it? A named person, by role, who would notice if it stopped. "The team" doesn't count.
  4. Can you count it? A control that leaves no trace can't be shown to be working. A logged review, a flag in a list or an audit event can.

A control that passes all four is likely to survive. One that fails two or more probably needs redesigning.

Swapping weak controls for strong ones

Weak controlStronger version
Staff check AI suggestions before actingCases below a confidence threshold go to a person automatically; the rest are acted on and sampled
Review all AI-drafted lettersReview every letter in three high-stakes categories; sample 1 in 20 of the rest
Users must not paste sensitive dataData loss prevention policy on the AI tool, with alerts to the data owner
Agent owner monitors spendMonthly credit cap on the agent and an alert at 75%
Annual review of the modelRe-run a fixed test set of past cases whenever the model or prompt changes

The stronger versions share a pattern. They decide in advance which cases need a person, they use the system to enforce it, and they produce a number someone can look at.

A worked example

This example is illustrative. A team uses an agent to suggest whether expense claims meet policy. The original control was "Managers review every suggestion before approving." After two months, a sample shows managers approve 99% of suggestions within seconds of opening them.

The redesigned control:

PartDetail
Built inClaims over £250, or where the agent's confidence is under 0.8, go to the manager with the agent's reason shown
QuickEverything else is approved automatically and appears on a weekly list
OwnedThe finance business partner owns the threshold and the sample
CountableEach week, 10 auto-approved claims are picked at random and checked; the error count is logged

Managers now see fewer claims and look at them properly. The weekly sample gives a real error rate, which the old control never did.

The control needs a control

A control that's working today can quietly stop. The sample isn't drawn one week, then the next. So the register should record whether each control ran, as well as what it found. That's the job of Review dates that happen and of nudges.

Where I got stuck

The four tests favour automated controls, which can make it look as if people aren't needed. I mean the opposite: the aim is to put people's attention where it's most useful: on the uncertain, high-stakes cases and on the sample. I'm still looking for a way to say that clearly in the register itself.

Sources

This article describes my own method and uses no external facts or figures. The expense example and its figures are illustrative. The threshold approach is explained in Act, ask a person, or stop.

Read next

A question to take awayWho gets told, and how fast, when a decision model starts drifting?

About me

Craig Stanley

Microsoft AI consultant and technical architect, based in Whitley Bay. Over the last few years I've delivered Microsoft 365 Copilot, Copilot Studio agents, Microsoft Foundry (formerly Azure AI Foundry) work and governance for UK public sector and financial services organisations.

What interests me is the decision underneath the tool: what it costs, what it risks, and whether a small, transparent model can make it better. I write the methods up here and on Substack so anyone can use them.

I write this site to learn in public: explaining each idea simply is how I check I understand it. Why I write this site.

Find me