Craig Stanley
Home / AI trolley problems / Agent builders

AI trolley problems: agent builders

Twelve design calls about permissions, approvals, logs, cost and the off switch.

For makers in Copilot Studio, Foundry and similar tools, developers, architects and product owners.

Read without playing

Key takeaways for agent builders

Twelve takeaways (best read after you play)
  1. Give every agent its own identity with the least access it needs.
  2. Put human approval on high-risk actions only, so approvers have time to read them.
  3. Enforce business rules in tools and APIs. Use the prompt for guidance.
  4. Log every tool call, input, output and cost so any run can be replayed.
  5. Match the model to each step, and prove it on your own test cases.
  6. Treat everything an agent reads as untrusted, and limit what it can send out.
  7. Set a confidence threshold: above it the agent acts, below it the agent asks.
  8. Build a pause switch the business owner can use without a developer.
  9. Build a test set from real cases before launch, and run it on every change.
  10. Make agent memory visible to users and easy to delete.
  11. Set a step limit for each run and a daily spending cap for every agent.
  12. Route harm, legal and vulnerability cases to a person from the first message.
All twelve dilemmas as thought experiments
  1. The master key

    Trolley problem

    Your agent works fine with a service account that can read all of SharePoint. Pull the lever to spend a day scoping it down to the two libraries it needs. Do nothing and ship with full access.

    • Pull the lever: Two libraries, a day's work
    • Do nothing: Access to everything

    Design review question: What's the smallest set of permissions this agent needs to do its job?

  2. Approve everything

    Trolley problem

    Every action your agent takes needs a person's approval, about 200 a day, and approvers have started clicking through. Pull the lever to let read-only steps run on their own and keep approval for writes and payments. Do nothing and keep approving everything.

    • Pull the lever: Approval only where it matters
    • Do nothing: 200 approvals a day

    Design review question: Which actions really need a person to approve them, and can that person read each one properly?

  3. Prompt or code?

    Thought experiment

    Refunds over £100 need a manager's say-so. You can write that rule into the system prompt today, or enforce it in the refund tool's code, which takes a sprint.

    • Option A: Put it in the prompt
    • Option B: Enforce it in the tool

    Design review question: Which of our agent's rules are enforced in code, and which only live in the prompt?

  4. Final answers only

    Trolley problem

    Your agent is live and logs only its final reply. Pull the lever to log every tool call, input, output and cost. Do nothing and keep storage small.

    • Pull the lever: Full traces, more storage
    • Do nothing: Replies, no reasons

    Design review question: Could we replay exactly what our agent did for any single customer last week?

  5. The biggest model

    Trolley problem

    Your agent uses the largest model available for every step, including simple routing. Pull the lever to move routing to a small decision model after testing it on 500 past cases. Do nothing and keep one model for everything.

    • Pull the lever: A small model for routine steps
    • Do nothing: The top model for everything

    Design review question: Which steps in our agent could a smaller, cheaper model handle, and have we tested that?

  6. The poisoned email

    Trolley problem

    Your agent reads incoming email. In testing, a crafted email gets it to forward data to an outside address. Pull the lever to delay launch a month and add content isolation and outbound checks. Do nothing and launch on time.

    • Pull the lever: Launch slips a month
    • Do nothing: Launch on time

    Design review question: What happens if someone hides instructions in content our agent reads?

  7. Ask or guess?

    Thought experiment

    A request could refer to either of two customer accounts. Your agent can stop and ask the user which one they mean, or pick the more likely account and carry on.

    • Option A: Stop and ask
    • Option B: Pick the likelier one

    Design review question: At what level of confidence should our agent stop and ask?

  8. The off switch

    Trolley problem

    The only way to stop your agent is to redeploy it. Pull the lever to spend a day building a pause switch the business owner can use. Do nothing and rely on the developers.

    • Pull the lever: A day spent on a switch
    • Do nothing: Stopping means a redeploy

    Design review question: How quickly can the owner stop our agent, and do they know how?

  9. It worked in the demo

    Trolley problem

    There's no test set. The plan is to ship once the demo looks good. Pull the lever to write 100 real test cases with expected outcomes first. Do nothing and ship after the demo.

    • Pull the lever: 100 test cases first
    • Do nothing: Shipped on a good demo

    Design review question: What test cases must our agent pass before each release?

  10. Total recall

    Thought experiment

    Your agent could remember users' preferences between sessions. It can remember everything by default, or only what a user explicitly asks it to save.

    • Option A: Remember everything
    • Option B: Only what users save

    Design review question: What does our agent remember about people, and can they see and delete it?

  11. The retry loop

    Trolley problem

    When a tool fails, your agent retries with no limit. Pull the lever to cap each run at 20 steps and add a daily spending cap. Do nothing and let it keep trying.

    • Pull the lever: 20 steps, a daily cap
    • Do nothing: Unlimited retries

    Design review question: What's the most a single run of our agent can cost?

  12. Some things need a person

    Trolley problem

    Your agent handles complaints from start to finish. Pull the lever to send any complaint that mentions harm, legal action or vulnerability straight to a person. Do nothing and let it handle everything.

    • Pull the lever: Some complaints to people
    • Do nothing: The agent handles every complaint

    Design review question: Which topics should always go straight to a person?

How the scoring works

Each answer moves one or more of the six areas below up or down by a point. Your score in an area is where your total lands between the lowest and highest totals the twelve dilemmas allow. 75% or more is strong, 45% to 74% is developing, under 45% is a gap. Your archetype comes from two totals. One is autonomy, measured as momentum over the following months rather than speed today, so a shortcut that causes an incident later counts as slow. The other is the remaining five areas combined (safeguards), where the top half starts at 70%. It's one practitioner's view of good practice, written down so you can argue with it.

  • Right-sized autonomy. The agent does routine work alone and stops where the risk rises.
  • Human in the loop. People step in where judgement matters, with the context they need.
  • Guardrails. The agent can't do what it shouldn't, whatever it's told.
  • Observability. Any run can be replayed and every change is tested.
  • Cost. The owner knows what a run costs and what the limit is.
  • Safety. The agent can be stopped in a minute and can't be talked into leaking data.

Question to take awayWhat's the smallest set of permissions this agent needs to do its job?

Inspired by Neal Agarwal's Absurd Trolley Problems. The dilemmas, scoring and reports here are new, written for AI at work.

About me

Craig Stanley

Microsoft AI consultant and technical architect, based in Whitley Bay. Over the last few years I've delivered Microsoft 365 Copilot, Copilot Studio agents, Azure AI Foundry work and governance for UK public sector and financial services organisations.

What interests me is the decision underneath the tool: what it costs, what it risks, and whether a small, transparent model can make it better. I write the methods up here and on Substack so anyone can use them.

Find me