Craig Stanley
Home / AI trolley problems / Operations

AI trolley problems: operations

Twelve calls about queues, costs, escalation and what happens when the model is unsure.

For service managers, operations leads, contact centre and back-office teams running AI day to day.

Read without playing

Key takeaways for operations

Twelve takeaways (best read after you play)
  1. Automate routine, reversible decisions once tested, and keep sampling the results.
  2. Give every agent on metered billing a spending cap and a daily alert.
  3. Clean up and assign owners for the content an AI answers from before you look at the model.
  4. Decide in advance where uncertain cases go, and staff that route for busy weeks.
  5. Size human review so reviewers can disagree. A review step that never overturns anything needs checking.
  6. Keep a test set of your own cases and rerun it every time a model or prompt changes.
  7. Decide what error rate you'll accept under pressure before the pressure arrives.
  8. Give every AI process a health check, an alert and an on-call owner.
  9. Give every AI channel a fast route to a person, with the conversation passed across.
  10. Check that training data reflects today's process, and watch how the model handles rare cases.
  11. Review AI tools for overlap every quarter. Fewer tools are cheaper and easier to govern.
  12. Look at AI outcomes by customer group, and give vulnerable customers a route to a person.
All twelve dilemmas as thought experiments
  1. Auto-close

    Trolley problem

    In testing, a model handled routine password-reset tickets correctly in almost every case. Pull the lever to let it close them on its own, with a weekly sample checked by a person. Do nothing and analysts keep closing every ticket by hand.

    • Pull the lever: The model closes routine tickets
    • Do nothing: Analysts close every ticket

    Ask your operations team: Which of our routine decisions are low-stakes and reversible enough to automate, and how will we sample them?

  2. The weekend bill

    Trolley problem

    An agent's consumption cost jumped over the weekend because it got stuck in a loop. Pull the lever to set a hard cap that stops each agent when it hits its daily budget. Do nothing and keep an eye on the invoice.

    • Pull the lever: The agent stops at its cap
    • Do nothing: Next month's invoice

    Ask your operations team: Who sees our AI consumption each day, and what stops a runaway agent?

  3. Three refund policies

    Trolley problem

    The knowledge base your chatbot answers from holds three different versions of the refund policy. Pull the lever to pause the bot for two weeks while the content is cleaned up. Do nothing and keep it running.

    • Pull the lever: The bot paused for clean-up
    • Do nothing: Three answers to one question

    Ask your operations team: Who owns the content our AI answers from, and when was it last reviewed?

  4. Unsure cases

    Thought experiment

    A complaints triage model is unsure about roughly one case in seven. You can send every unsure case to a person, or let the model choose the most likely queue and flag it for a later check.

    • Option A: Send them to a person
    • Option B: Route them and flag

    Ask your operations team: What happens to the cases our models are unsure about, and do we have the people to handle them?

  5. The rubber stamp

    Trolley problem

    One reviewer approves around 400 AI decisions an hour. Pull the lever to cap reviews at 60 an hour and bring in three more reviewers. Do nothing and keep the queue fast.

    • Pull the lever: Three more reviewers
    • Do nothing: One approval every nine seconds

    Ask your operations team: Do our reviewers have the time and the authority to overrule the AI?

  6. The model swap

    Trolley problem

    Your supplier is retiring the model your routing runs on next month. They say the replacement is better. Pull the lever to run last quarter's cases through the new model before switching. Do nothing and switch on the supplier's date.

    • Pull the lever: A two-week regression test
    • Do nothing: Switch on the supplier's date

    Ask your operations team: Do we have a test set of our own cases to check a new model before it goes live?

  7. Peak week

    Thought experiment

    Your busiest week of the year is coming. You can let AI take a bigger share of contacts and accept more errors, or keep its share where it is and bring in temporary staff.

    • Option A: Let AI take the extra
    • Option B: Bring in temporary staff

    Ask your operations team: What error rate can we accept at peak, and what does fixing errors afterwards cost us?

  8. The silent failure

    Trolley problem

    An agent stopped working three days ago and nobody noticed. Pull the lever to spend a sprint on health checks and alerts that go to an on-call owner. Do nothing and restart it.

    • Pull the lever: Alerts and an on-call owner
    • Do nothing: Another silent failure

    Ask your operations team: How long would it take us to notice if an AI process stopped working?

  9. Talk to a person

    Trolley problem

    Customers who ask your chatbot for a person have to start again on the phone. Pull the lever to build a handoff that passes the whole conversation to an agent. Do nothing and keep the current set-up.

    • Pull the lever: Handoff with full history
    • Do nothing: Customers start again

    Ask your operations team: How easily can a customer reach a person from our AI channels, and do they have to repeat themselves?

  10. Old data or new

    Thought experiment

    You're training a routing model. You can use three years of history, which includes a team structure you changed last spring, or only the six months since the change.

    • Option A: Three years of history
    • Option B: The last six months

    Ask your operations team: Does the data our models learn from reflect how we work today?

  11. Three transcription tools

    Trolley problem

    Three teams each pay for a different AI transcription tool. Pull the lever to move everyone onto one. Do nothing and let each team keep its favourite.

    • Pull the lever: One tool, some unhappy teams
    • Do nothing: Three bills, three data stores

    Ask your operations team: How many AI tools are we paying for that do the same job?

  12. Faster on average

    Trolley problem

    Since the AI assistant went live, average handling time is down, but complaints from customers in vulnerable circumstances are up. Pull the lever to route those customers to people. Do nothing and keep the faster average.

    • Pull the lever: Vulnerable customers to people
    • Do nothing: Faster on average, worse for some

    Ask your operations team: Do we check AI outcomes for different groups of customers, or only the average?

How the scoring works

Each answer moves one or more of the six areas below up or down by a point. Your score in an area is where your total lands between the lowest and highest totals the twelve dilemmas allow. 75% or more is strong, 45% to 74% is developing, under 45% is a gap. Your archetype comes from two totals. One is throughput, measured as momentum over the following months rather than speed today, so a shortcut that causes an incident later counts as slow. The other is the remaining five areas combined (control), where the top half starts at 70%. It's one practitioner's view of good practice, written down so you can argue with it.

  • Reliability. Every live AI process has an owner, a test set and an alert.
  • Cost control. AI spend is visible every day to the person who can act on it.
  • Escalation. People handle the cases the model shouldn't, with time to do it properly.
  • Data quality. Someone owns the content and data each AI system depends on.
  • Throughput. Routine work moves through quickly and is sampled for quality.
  • Service quality. Customers get consistent answers and can always reach a person.

Question to take awayWhich of our routine decisions are low-stakes and reversible enough to automate, and how will we sample them?

Inspired by Neal Agarwal's Absurd Trolley Problems. The dilemmas, scoring and reports here are new, written for AI at work.

About me

Craig Stanley

Microsoft AI consultant and technical architect, based in Whitley Bay. Over the last few years I've delivered Microsoft 365 Copilot, Copilot Studio agents, Azure AI Foundry work and governance for UK public sector and financial services organisations.

What interests me is the decision underneath the tool: what it costs, what it risks, and whether a small, transparent model can make it better. I write the methods up here and on Substack so anyone can use them.

Find me