Craig Stanley
Home / AI trolley problems / Boards

AI trolley problems: boards

Twelve calls about accountability, risk appetite and oversight, the decisions only a board can make.

For chairs, non-executive and executive directors, trustees and audit and risk committees.

Read without playing

Key takeaways for boards

Twelve takeaways (best read after you play)
  1. Give every pilot success measures and an end date before it starts. Without them it can't succeed or fail.
  2. Treat a supplier's compliance claim as the start of your questions. Responsibility stays with the organisation using the system.
  3. Build the inventory before the deep assurance. You can only rank the risks you can see.
  4. For decisions about people, use models you can explain, or add a way to explain and challenge each decision.
  5. Treat AI failures as incidents: contain them, record them, tell the people affected, then fix the cause.
  6. A written risk appetite lets the board say yes faster, because each proposal is judged against the same line.
  7. Any agent that can spend money or change records needs limits, approval for high-value actions, a log and an owner.
  8. Any AI used in hiring, pay or performance needs bias testing before launch and at regular intervals after.
  9. Name a person. A committee can advise and coordinate, but accountability needs someone who answers for it.
  10. Treat board AI literacy as a governance control. Directors need enough knowledge to challenge what they're shown.
  11. Only make public claims about AI that you can back with evidence.
  12. Match how often the board reviews AI to how fast the organisation's use of it is changing.
All twelve dilemmas as thought experiments
  1. The pilot that never ends

    Trolley problem

    Your Copilot pilot has run for 18 months. Nobody set success measures, so nobody can say whether it worked. Pull the lever to roll it out to all 3,000 staff now. Do nothing and the pilot is extended again.

    • Pull the lever: 3,000 licences, no measures
    • Do nothing: Pilot, month 19

    Board question: What three measures would tell us this worked, and who reports them to us?

  2. The vendor's word

    Trolley problem

    A supplier's slide says its product is 'fully compliant with the EU AI Act'. Pull the lever to ask for the evidence and lose the launch discount. Do nothing and sign this week.

    • Pull the lever: No launch discount
    • Do nothing: Contract signed on a slide

    Board question: What evidence do we ask AI suppliers for before we sign, and who checks it?

  3. Two kinds of register

    Thought experiment

    You have budget for one assurance exercise this year. You can list every AI system in the organisation, roughly, or assure your three highest-risk systems in depth and leave the rest unlisted.

    • Option A: List everything, lightly
    • Option B: Go deep on three

    Board question: Could we list every AI system we use today, with an owner for each?

  4. The model nobody can explain

    Trolley problem

    A new lending model approves more good customers than the old one, but nobody can explain why it declines anyone. Pull the lever to keep the simpler model you can explain. Do nothing and switch to the new one.

    • Pull the lever: Slightly fewer approvals
    • Do nothing: Decisions nobody can explain

    Board question: Can we explain every AI-assisted decision that affects a customer or a colleague?

  5. The quiet fix

    Trolley problem

    Your customer chatbot told 40 people they could cancel a policy without penalty. They can't. Pull the lever to contact them, log an incident and tell the regulator. Do nothing and let the team quietly fix the prompt.

    • Pull the lever: An awkward disclosure
    • Do nothing: A quiet prompt patch

    Board question: What's our process when AI gets something wrong, and who decides whether to tell the people affected?

  6. Appetite first?

    Thought experiment

    Five AI proposals are waiting. You can pause them for a quarter while the board writes an AI risk appetite, or approve them one at a time and let the appetite take shape from those calls.

    • Option A: Write the appetite first
    • Option B: Learn as you go

    Board question: Have we written down how much AI risk we'll accept, and for which uses?

  7. The agent with a budget

    Trolley problem

    An AI agent now pays supplier invoices. It has a £50,000 limit and no approval step. Pull the lever to require a person's approval above £1,000. Do nothing and keep it fast.

    • Pull the lever: Approval above £1,000
    • Do nothing: £50,000, no approval

    Board question: What can our AI agents do without a person approving it, and who can switch them off?

  8. The shortlist

    Trolley problem

    Your new recruitment tool saves recruiters days each month. HR notices it almost never shortlists applicants over 50. Pull the lever to pause it until it's been bias-tested. Do nothing and keep the savings.

    • Pull the lever: Recruiters back on manual
    • Do nothing: Older applicants screened out

    Board question: How do we test AI that affects people for bias, and how often?

  9. Who owns it?

    Thought experiment

    You need to decide where accountability for AI sits. One option is a single named director for all of it. The other is for each business area to own its own AI, coordinated by a cross-functional committee.

    • Option A: One named director
    • Option B: Each area, plus a committee

    Board question: Which director is accountable for AI, and what can they decide without the full board?

  10. The away day

    Trolley problem

    Two directors think AI chat tools look facts up like a search engine. Pull the lever to spend the strategy away day on AI training for the whole board. Do nothing and keep the strategy agenda.

    • Pull the lever: Away day spent on training
    • Do nothing: Directors guessing

    Board question: Could each director explain our biggest AI risk in two sentences?

  11. The AI-first press release

    Trolley problem

    A competitor has announced it is 'AI-first'. Pull the lever to announce the same by Friday, with nothing yet behind it. Do nothing and wait until you have results.

    • Pull the lever: A press release with nothing behind it
    • Do nothing: Silence, for now

    Board question: Could we evidence every claim we make in public about our use of AI?

  12. How often?

    Thought experiment

    AI use across the organisation is changing month to month. You can put AI on every board agenda for the next year, or review it in depth twice a year at the risk committee.

    • Option A: Every board agenda
    • Option B: Twice a year, in depth

    Board question: How often does the board see what AI the organisation uses, and what changed since last time?

How the scoring works

Each answer moves one or more of the six areas below up or down by a point. Your score in an area is where your total lands between the lowest and highest totals the twelve dilemmas allow. 75% or more is strong, 45% to 74% is developing, under 45% is a gap. Your archetype comes from two totals. One is speed, measured as momentum over the following months rather than speed today, so a shortcut that causes an incident later counts as slow. The other is the remaining five areas combined (governance), where the top half starts at 70%. It's one practitioner's view of good practice, written down so you can argue with it.

  • Accountability. Every AI system has a named owner, and the board knows which director answers for AI.
  • Risk appetite. Proposals are judged against an appetite the board agreed in advance.
  • Oversight. The board sees AI use on a schedule that matches how fast it's changing.
  • Transparency. The organisation can explain its AI decisions and owns up when they go wrong.
  • Speed. Low-risk proposals get a quick decision, and pilots end on a set date.
  • Fairness. Decisions that affect people can be explained and challenged.

Question to take awayWhat three measures would tell us this worked, and who reports them to us?

Inspired by Neal Agarwal's Absurd Trolley Problems. The dilemmas, scoring and reports here are new, written for AI at work.

About me

Craig Stanley

Microsoft AI consultant and technical architect, based in Whitley Bay. Over the last few years I've delivered Microsoft 365 Copilot, Copilot Studio agents, Azure AI Foundry work and governance for UK public sector and financial services organisations.

What interests me is the decision underneath the tool: what it costs, what it risks, and whether a small, transparent model can make it better. I write the methods up here and on Substack so anyone can use them.

Find me