AI trolley problems: boards
Twelve calls about accountability, risk appetite and oversight, the decisions only a board can make.
For chairs, non-executive and executive directors, trustees and audit and risk committees.
Twelve calls, then your report
Pull the lever or do nothing. Some dilemmas have no lever: they're open thought experiments where you pick a side. After each call you see what happened and what the other choice would have done.
Your answers stay in your browser. Nothing is sent or stored.
Dilemma
Result
Key takeaways for boards
Twelve takeaways (best read after you play)
- Give every pilot success measures and an end date before it starts. Without them it can't succeed or fail.
- Treat a supplier's compliance claim as the start of your questions. Responsibility stays with the organisation using the system.
- Build the inventory before the deep assurance. You can only rank the risks you can see.
- For decisions about people, use models you can explain, or add a way to explain and challenge each decision.
- Treat AI failures as incidents: contain them, record them, tell the people affected, then fix the cause.
- A written risk appetite lets the board say yes faster, because each proposal is judged against the same line.
- Any agent that can spend money or change records needs limits, approval for high-value actions, a log and an owner.
- Any AI used in hiring, pay or performance needs bias testing before launch and at regular intervals after.
- Name a person. A committee can advise and coordinate, but accountability needs someone who answers for it.
- Treat board AI literacy as a governance control. Directors need enough knowledge to challenge what they're shown.
- Only make public claims about AI that you can back with evidence.
- Match how often the board reviews AI to how fast the organisation's use of it is changing.
All twelve dilemmas as thought experiments
The pilot that never ends
Trolley problemYour Copilot pilot has run for 18 months. Nobody set success measures, so nobody can say whether it worked. Pull the lever to roll it out to all 3,000 staff now. Do nothing and the pilot is extended again.
- Pull the lever: 3,000 licences, no measures
- Do nothing: Pilot, month 19
Board question: What three measures would tell us this worked, and who reports them to us?
The vendor's word
Trolley problemA supplier's slide says its product is 'fully compliant with the EU AI Act'. Pull the lever to ask for the evidence and lose the launch discount. Do nothing and sign this week.
- Pull the lever: No launch discount
- Do nothing: Contract signed on a slide
Board question: What evidence do we ask AI suppliers for before we sign, and who checks it?
Two kinds of register
Thought experimentYou have budget for one assurance exercise this year. You can list every AI system in the organisation, roughly, or assure your three highest-risk systems in depth and leave the rest unlisted.
- Option A: List everything, lightly
- Option B: Go deep on three
Board question: Could we list every AI system we use today, with an owner for each?
The model nobody can explain
Trolley problemA new lending model approves more good customers than the old one, but nobody can explain why it declines anyone. Pull the lever to keep the simpler model you can explain. Do nothing and switch to the new one.
- Pull the lever: Slightly fewer approvals
- Do nothing: Decisions nobody can explain
Board question: Can we explain every AI-assisted decision that affects a customer or a colleague?
The quiet fix
Trolley problemYour customer chatbot told 40 people they could cancel a policy without penalty. They can't. Pull the lever to contact them, log an incident and tell the regulator. Do nothing and let the team quietly fix the prompt.
- Pull the lever: An awkward disclosure
- Do nothing: A quiet prompt patch
Board question: What's our process when AI gets something wrong, and who decides whether to tell the people affected?
Appetite first?
Thought experimentFive AI proposals are waiting. You can pause them for a quarter while the board writes an AI risk appetite, or approve them one at a time and let the appetite take shape from those calls.
- Option A: Write the appetite first
- Option B: Learn as you go
Board question: Have we written down how much AI risk we'll accept, and for which uses?
The agent with a budget
Trolley problemAn AI agent now pays supplier invoices. It has a £50,000 limit and no approval step. Pull the lever to require a person's approval above £1,000. Do nothing and keep it fast.
- Pull the lever: Approval above £1,000
- Do nothing: £50,000, no approval
Board question: What can our AI agents do without a person approving it, and who can switch them off?
The shortlist
Trolley problemYour new recruitment tool saves recruiters days each month. HR notices it almost never shortlists applicants over 50. Pull the lever to pause it until it's been bias-tested. Do nothing and keep the savings.
- Pull the lever: Recruiters back on manual
- Do nothing: Older applicants screened out
Board question: How do we test AI that affects people for bias, and how often?
Who owns it?
Thought experimentYou need to decide where accountability for AI sits. One option is a single named director for all of it. The other is for each business area to own its own AI, coordinated by a cross-functional committee.
- Option A: One named director
- Option B: Each area, plus a committee
Board question: Which director is accountable for AI, and what can they decide without the full board?
The away day
Trolley problemTwo directors think AI chat tools look facts up like a search engine. Pull the lever to spend the strategy away day on AI training for the whole board. Do nothing and keep the strategy agenda.
- Pull the lever: Away day spent on training
- Do nothing: Directors guessing
Board question: Could each director explain our biggest AI risk in two sentences?
The AI-first press release
Trolley problemA competitor has announced it is 'AI-first'. Pull the lever to announce the same by Friday, with nothing yet behind it. Do nothing and wait until you have results.
- Pull the lever: A press release with nothing behind it
- Do nothing: Silence, for now
Board question: Could we evidence every claim we make in public about our use of AI?
How often?
Thought experimentAI use across the organisation is changing month to month. You can put AI on every board agenda for the next year, or review it in depth twice a year at the risk committee.
- Option A: Every board agenda
- Option B: Twice a year, in depth
Board question: How often does the board see what AI the organisation uses, and what changed since last time?
How the scoring works
Each answer moves one or more of the six areas below up or down by a point. Your score in an area is where your total lands between the lowest and highest totals the twelve dilemmas allow. 75% or more is strong, 45% to 74% is developing, under 45% is a gap. Your archetype comes from two totals. One is speed, measured as momentum over the following months rather than speed today, so a shortcut that causes an incident later counts as slow. The other is the remaining five areas combined (governance), where the top half starts at 70%. It's one practitioner's view of good practice, written down so you can argue with it.
- Accountability. Every AI system has a named owner, and the board knows which director answers for AI.
- Risk appetite. Proposals are judged against an appetite the board agreed in advance.
- Oversight. The board sees AI use on a schedule that matches how fast it's changing.
- Transparency. The organisation can explain its AI decisions and owns up when they go wrong.
- Speed. Low-risk proposals get a quick decision, and pilots end on a set date.
- Fairness. Decisions that affect people can be explained and challenged.
Question to take awayWhat three measures would tell us this worked, and who reports them to us?
Inspired by Neal Agarwal's Absurd Trolley Problems. The dilemmas, scoring and reports here are new, written for AI at work.