AI trolley problems: agent builders
Twelve design calls about permissions, approvals, logs, cost and the off switch.
For makers in Copilot Studio, Foundry and similar tools, developers, architects and product owners.
Twelve calls, then your report
Pull the lever or do nothing. Some dilemmas have no lever: they're open thought experiments where you pick a side. After each call you see what happened and what the other choice would have done.
Your answers stay in your browser. Nothing is sent or stored.
Dilemma
Result
Key takeaways for agent builders
Twelve takeaways (best read after you play)
- Give every agent its own identity with the least access it needs.
- Put human approval on high-risk actions only, so approvers have time to read them.
- Enforce business rules in tools and APIs. Use the prompt for guidance.
- Log every tool call, input, output and cost so any run can be replayed.
- Match the model to each step, and prove it on your own test cases.
- Treat everything an agent reads as untrusted, and limit what it can send out.
- Set a confidence threshold: above it the agent acts, below it the agent asks.
- Build a pause switch the business owner can use without a developer.
- Build a test set from real cases before launch, and run it on every change.
- Make agent memory visible to users and easy to delete.
- Set a step limit for each run and a daily spending cap for every agent.
- Route harm, legal and vulnerability cases to a person from the first message.
All twelve dilemmas as thought experiments
The master key
Trolley problemYour agent works fine with a service account that can read all of SharePoint. Pull the lever to spend a day scoping it down to the two libraries it needs. Do nothing and ship with full access.
- Pull the lever: Two libraries, a day's work
- Do nothing: Access to everything
Design review question: What's the smallest set of permissions this agent needs to do its job?
Approve everything
Trolley problemEvery action your agent takes needs a person's approval, about 200 a day, and approvers have started clicking through. Pull the lever to let read-only steps run on their own and keep approval for writes and payments. Do nothing and keep approving everything.
- Pull the lever: Approval only where it matters
- Do nothing: 200 approvals a day
Design review question: Which actions really need a person to approve them, and can that person read each one properly?
Prompt or code?
Thought experimentRefunds over £100 need a manager's say-so. You can write that rule into the system prompt today, or enforce it in the refund tool's code, which takes a sprint.
- Option A: Put it in the prompt
- Option B: Enforce it in the tool
Design review question: Which of our agent's rules are enforced in code, and which only live in the prompt?
Final answers only
Trolley problemYour agent is live and logs only its final reply. Pull the lever to log every tool call, input, output and cost. Do nothing and keep storage small.
- Pull the lever: Full traces, more storage
- Do nothing: Replies, no reasons
Design review question: Could we replay exactly what our agent did for any single customer last week?
The biggest model
Trolley problemYour agent uses the largest model available for every step, including simple routing. Pull the lever to move routing to a small decision model after testing it on 500 past cases. Do nothing and keep one model for everything.
- Pull the lever: A small model for routine steps
- Do nothing: The top model for everything
Design review question: Which steps in our agent could a smaller, cheaper model handle, and have we tested that?
The poisoned email
Trolley problemYour agent reads incoming email. In testing, a crafted email gets it to forward data to an outside address. Pull the lever to delay launch a month and add content isolation and outbound checks. Do nothing and launch on time.
- Pull the lever: Launch slips a month
- Do nothing: Launch on time
Design review question: What happens if someone hides instructions in content our agent reads?
Ask or guess?
Thought experimentA request could refer to either of two customer accounts. Your agent can stop and ask the user which one they mean, or pick the more likely account and carry on.
- Option A: Stop and ask
- Option B: Pick the likelier one
Design review question: At what level of confidence should our agent stop and ask?
The off switch
Trolley problemThe only way to stop your agent is to redeploy it. Pull the lever to spend a day building a pause switch the business owner can use. Do nothing and rely on the developers.
- Pull the lever: A day spent on a switch
- Do nothing: Stopping means a redeploy
Design review question: How quickly can the owner stop our agent, and do they know how?
It worked in the demo
Trolley problemThere's no test set. The plan is to ship once the demo looks good. Pull the lever to write 100 real test cases with expected outcomes first. Do nothing and ship after the demo.
- Pull the lever: 100 test cases first
- Do nothing: Shipped on a good demo
Design review question: What test cases must our agent pass before each release?
Total recall
Thought experimentYour agent could remember users' preferences between sessions. It can remember everything by default, or only what a user explicitly asks it to save.
- Option A: Remember everything
- Option B: Only what users save
Design review question: What does our agent remember about people, and can they see and delete it?
The retry loop
Trolley problemWhen a tool fails, your agent retries with no limit. Pull the lever to cap each run at 20 steps and add a daily spending cap. Do nothing and let it keep trying.
- Pull the lever: 20 steps, a daily cap
- Do nothing: Unlimited retries
Design review question: What's the most a single run of our agent can cost?
Some things need a person
Trolley problemYour agent handles complaints from start to finish. Pull the lever to send any complaint that mentions harm, legal action or vulnerability straight to a person. Do nothing and let it handle everything.
- Pull the lever: Some complaints to people
- Do nothing: The agent handles every complaint
Design review question: Which topics should always go straight to a person?
How the scoring works
Each answer moves one or more of the six areas below up or down by a point. Your score in an area is where your total lands between the lowest and highest totals the twelve dilemmas allow. 75% or more is strong, 45% to 74% is developing, under 45% is a gap. Your archetype comes from two totals. One is autonomy, measured as momentum over the following months rather than speed today, so a shortcut that causes an incident later counts as slow. The other is the remaining five areas combined (safeguards), where the top half starts at 70%. It's one practitioner's view of good practice, written down so you can argue with it.
- Right-sized autonomy. The agent does routine work alone and stops where the risk rises.
- Human in the loop. People step in where judgement matters, with the context they need.
- Guardrails. The agent can't do what it shouldn't, whatever it's told.
- Observability. Any run can be replayed and every change is tested.
- Cost. The owner knows what a run costs and what the limit is.
- Safety. The agent can be stopped in a minute and can't be talked into leaking data.
Question to take awayWhat's the smallest set of permissions this agent needs to do its job?
Inspired by Neal Agarwal's Absurd Trolley Problems. The dilemmas, scoring and reports here are new, written for AI at work.