Craig Stanley
Home / Plain text view

The whole site, as plain text

For screen readers, slow connections, and language models. The same text is at /llms-full.txt, with a short summary at /llms.txt.

Open .txt
# Craig Stanley · Work

> Decide in the open. Craig Stanley, Microsoft AI consultant in the North East of England. I map how work gets done, find the decisions that cost money, and put small, visible AI behind them on the Microsoft stack.

Canonical: https://craigstanley.work/
Version 2.0, 2026-10-10. Each section is explained three ways: like I'm 5, like I'm 15, and for professionals.

Introduction
============
[Like I'm 5] People at work make lots of choices every day. I help them see each choice clearly. Sometimes a small computer helper gives a hint, a person still says yes or no, and anyone can see how it worked.

[Like I'm 15] Every job is full of small decisions: approve this, send that, buy one more licence. I find the ones that cost the most when they go wrong, add a small AI model that gives a score you can check, and write up how it works so your team can carry on without me.

[For pros] I map how work gets done, find the decisions that cost money, and put small, visible AI behind them on the Microsoft stack. Then I publish the method so your people can repeat it.

How the method works
====================
Four steps, in order. Each one leaves the team with something it keeps.

1. Embed: Sit with the people doing the work. Watch the approvals, tickets and handovers, not the org chart.
2. Map the work: Describe the roles in O*NET and ESCO terms and list the decisions each one owns, with how often they come up and what's at stake.
3. Decide: Pick one repeated decision. Score it with a cheap model such as Microsoft-Decision-1 or Jev, and agree a threshold with the people affected.
4. Ship it: Deploy in Microsoft 365, Copilot Studio or Foundry with Agent 365 and Purview in place. Publish the method.

01 Blog
=======
URL: https://craigstanley.work/blog/
Question: What's Craig thinking about this week?

[Like I'm 5] This is where I write down what I'm thinking about each week, a bit like a diary about work and computers. Every post ends with a question for you to think about.

[Like I'm 15] Short posts about what changed in Microsoft's AI tools and whether it matters to you. Each one shows a single method you can try and finishes with a question to take back to your team.

[For pros] Short notes, field reports and a monthly scan of what changed in Microsoft AI and why it matters for cost, risk and decisions. Every post shows one method and ends on a question.

## What gets written here

Four kinds of post, each with a fixed shape so you know what you're getting.

- Notes are short. One idea, one method, one question at the end.
- Field reports describe what happened when a method met a real team. Names and numbers that could identify anyone are taken out.
- Horizon scan comes out monthly. It lists what Microsoft shipped, what it costs and what you can safely ignore for now.
- Decision of the week takes one ordinary work decision and scores it in public, with the numbers shown.

## How posts are written

Every post names its sources and dates any vendor figure. Where a number is an assumption, the post says so next to the number. Posts are written at three reading levels, so you can send the same link to a board member and a new starter.

## Where to start

Read Why decisions, not adoption (https://craigstanley.work/blog/notes/why-decisions-not-adoption/). It explains the idea the rest of the site is built on.

Notes [Live]
------------
Short pieces, one idea each.
Path: /blog/notes
- Microsoft-Decision-1: a first look (coming)
- Home, Code and Autopilot in plain English (coming)
- Why decisions, not adoption: https://craigstanley.work/blog/notes/why-decisions-not-adoption/

Field reports [Drafting]
------------------------
What happened when a method met a real team, anonymised.
Path: /blog/field-reports
- A 74,000-user licence model, without the spreadsheet war (coming)

Horizon scan [Planned]
----------------------
Monthly: what shipped, what it costs, what to ignore.
Path: /blog/horizon-scan
- November 2026 (coming)

Decision of the week [Planned]
------------------------------
One everyday work decision, scored in public.
Path: /blog/decision-of-the-week
- Approve the extra Copilot licence? (coming)

Question to take away: Which decision did your team make this week that nobody wrote down?

02 Work
=======
URL: https://craigstanley.work/work/
Question: What do we actually do all day?

[Like I'm 5] Before you fix a job, you need to know what the job is. I draw a picture of all the little things people do at work, so we can see where the choices are.

[Like I'm 15] Governments keep big lists that describe thousands of jobs as tasks and skills. The US list is called O*NET and the EU one is ESCO. I use them to draw a map of what your team really does, then mark the places where someone has to make a call.

[For pros] O*NET and ESCO describe thousands of occupations as tasks, skills and work activities. Used carefully, they give you a shared map of the work before anyone talks about AI, and they show where the decisions sit.

## Why start with the work

Most AI programmes start with a tool and look for somewhere to put it. That gets the order backwards. If you don't know what your people do all day, you can't say which part of it a tool should change, or how you'd know it had.

Two public frameworks make this much easier than starting from a blank page:

- O*NET, run for the US Department of Labor, describes around 1,000 occupations in detail: the tasks, the work activities, the tools, and the context the work happens in.
- ESCO, run by the European Commission, lists about 3,000 occupations and nearly 14,000 skills and competences, in every EU language.

Neither will describe your team exactly. Both give you a shared vocabulary, so a finance partner and a service desk lead can compare their work on the same terms.

## From tasks to decisions

A task list tells you what people do. The useful next step is to ask, for each task, where someone has to make a call. Who approves? Who decides to escalate? Who picks the supplier? Each of those is a decision you can count, time and score. That's the bridge into the Decisions (https://craigstanley.work/decisions/) section.

## Where to start

Read Map the work, then the decisions (https://craigstanley.work/work/start-here/map-the-work-then-the-decisions/), then Reading an O*NET occupation (https://craigstanley.work/work/onet/reading-an-o-net-occupation/).

Start here [Live]
-----------------
Why map work before you buy tools.
Path: /work/start-here
- Map the work, then the decisions: https://craigstanley.work/work/start-here/map-the-work-then-the-decisions/

The O*NET lens [Live]
---------------------
The US Department of Labor's model: tasks, work activities, work context.
Path: /work/onet
- Reading an O*NET occupation: https://craigstanley.work/work/onet/reading-an-o-net-occupation/
- Work context: freedom to make decisions, consequence of error: https://craigstanley.work/work/onet/work-context-freedom-to-make-decisions-consequence-of-error/
- Detailed work activities as a common language (coming)

The ESCO lens [Live]
--------------------
The EU's skills, competences and occupations classification.
Path: /work/esco
- ESCO occupations and skills in practice: https://craigstanley.work/work/esco/esco-occupations-and-skills-in-practice/
- Crosswalking ESCO to O*NET: https://craigstanley.work/work/esco/crosswalking-esco-to-o-net/
- Using ESCO for UK roles (coming)

Tasks to decisions [Live]
-------------------------
Turning task lists into decision points with frequency, stakes and reversibility.
Path: /work/tasks-to-decisions
- The decision inventory template: https://craigstanley.work/work/tasks-to-decisions/the-decision-inventory-template/
- Frequency, stakes, reversibility: a scoring guide (coming)

Role cards [Planned]
--------------------
One page per common role: tasks, decisions, where AI helps and where it shouldn't.
Path: /work/role-cards
- Case manager (coming)
- Service desk analyst (coming)
- Finance business partner (coming)
- Planner (coming)

Work data you already hold [Planned]
------------------------------------
Job descriptions, ticket logs, approvals and Viva Insights as evidence.
Path: /work/work-data
- What your approval logs already know (coming)

Question to take away: Could you list the ten decisions your team makes most often, with how long each takes?

03 Decisions (start here)
=========================
URL: https://craigstanley.work/decisions/
Question: How do we decide better, cheaply and in the open?

[Like I'm 5] A decision is picking one thing out of a few. I help people pick well. A small computer helper says "I think this one" and shows why. When the choice is big or tricky, a person decides.

[Like I'm 15] Lots of work decisions repeat: approve or reject, escalate or close, buy now or wait. Decision theory gives you simple tools for these, like weighing what you'd gain against how likely it is. A small, cheap AI model can score the routine ones while people keep the close calls, and the method is written down so anyone can check it or copy it.

[For pros] Decision theory, game theory and a few well-tested habits, applied to everyday work decisions. Put a small, cheap, observable AI model behind the routine choices, keep people on the ones that matter, and publish the method so anyone can reuse it. Start here if you want to save money, make money and give colleagues firmer ground to stand on.

## The idea in one paragraph

Organisations make the same small decisions thousands of times a year. Approve the expense or query it. Close the ticket or escalate it. Renew the licence or let it lapse. Each one is cheap on its own and expensive in total. A small model can score the routine cases quickly and cheaply. People stay on the close calls and the cases with high stakes. Every step is written down, so anyone can check the method, argue with it or copy it.

## What you need to know

You don't need a maths degree. You need five ideas:

1. Expected value. Weigh each outcome by how likely it is.
2. Value of information. More data is only worth buying if it could change your choice.
3. Thresholds. Decide in advance at what score you act, ask a person, or stop.
4. Calibration. Check that "80% sure" turns out right about 80% of the time.
5. Records. Write down what you knew when you decided, so you can judge the decision and not the luck.

Each one has its own page below, with a worked example.

## Why "in the open"

A decision method that only one person understands can't be trusted, improved or handed over. Publishing the method, the threshold and the record means the people affected by a decision can see how it was made. It also makes the model easy to swap out when a better one comes along.

## Where to start

Begin with Five questions before any decision (https://craigstanley.work/decisions/start-here/five-questions-before-any-decision/). If you want the theory first, read Expected value on one page (https://craigstanley.work/decisions/decision-theory/expected-value-on-one-page/).

Start here: five questions [Live]
---------------------------------
What, why, when, how and who, before any decision.
Path: /decisions/start-here
- Five questions before any decision: https://craigstanley.work/decisions/start-here/five-questions-before-any-decision/
- Save money, make money, look after people: the three goals: https://craigstanley.work/decisions/start-here/save-money-make-money-look-after-people-the-three-goals/

Decision theory at work [Live]
------------------------------
Expected value, value of information, thresholds and calibration without the maths anxiety.
Path: /decisions/decision-theory
- Expected value on one page: https://craigstanley.work/decisions/decision-theory/expected-value-on-one-page/
- When is more information worth paying for?: https://craigstanley.work/decisions/decision-theory/when-is-more-information-worth-paying-for/
- Setting a threshold you can defend: https://craigstanley.work/decisions/decision-theory/setting-a-threshold-you-can-defend/
- Calibration: are your 80% calls right 80% of the time?: https://craigstanley.work/decisions/decision-theory/calibration-are-your-80-calls-right-80-of-the-time/

Game theory at work [Live]
--------------------------
Budgets, bids, negotiations and other choices where people react to each other.
Path: /decisions/game-theory
- The budget contest (coming)
- Supplier bids as a repeated game (coming)
- Why everyone pads their estimates: https://craigstanley.work/decisions/game-theory/why-everyone-pads-their-estimates/

Bias field guide [Live]
-----------------------
Anchoring, loss aversion, outcome bias and how to spot them in meetings.
Path: /decisions/bias
- Outcome bias: judging the decision by the luck: https://craigstanley.work/decisions/bias/outcome-bias-judging-the-decision-by-the-luck/
- Anchoring in estimates: https://craigstanley.work/decisions/bias/anchoring-in-estimates/

Decision models [Live]
----------------------
Small models that score fixed options: Microsoft-Decision-1, Jev and Copilot Studio patterns.
Path: /decisions/decision-models
- What a decision model is, and isn't: https://craigstanley.work/decisions/decision-models/what-a-decision-model-is-and-isn-t/
- Microsoft-Decision-1 in Foundry (coming)
- Jev typed judgements (coming)
- Act, ask a person, or stop: https://craigstanley.work/decisions/decision-models/act-ask-a-person-or-stop/

Decision records [Live]
-----------------------
Write down the options, the information and your confidence at the moment you decide.
Path: /decisions/decision-records
- The decision record schema: https://craigstanley.work/decisions/decision-records/the-decision-record-schema/
- A 30-second capture in Teams (coming)

Method library [Planned]
------------------------
Every method on the site, free to reuse, with a worked example.
Path: /decisions/method-library
- The Decision Diagnostic, in full (coming)

Question to take away: Which repeated decision would you trust a cheap model to score first, with a person checking the close calls?

04 Capabilities
===============
URL: https://craigstanley.work/capabilities/
Question: What can Microsoft's tools do today?

[Like I'm 5] Microsoft makes lots of computer helpers, and they all have different names. This part tells you what each one does, what it costs, and when you'd use it.

[Like I'm 15] Microsoft sells a lot of AI products and the names change often. In September 2026 it grouped Copilot into Home, Code and Autopilot. Here I explain each product in plain terms: what it does, what it costs and which decisions it can help with.

[For pros] Microsoft now describes Copilot as an agent layer across Microsoft 365, announced on 25 September 2026 as Home, Code and Autopilot. This section covers what each product does, what it costs and where it fits a decision. It starts with decision models, which link straight back to the Decisions section.

## How this section is organised

Microsoft's AI products change names and packaging often. This section sorts them by what they do for a decision rather than by brand:

- Find and summarise: tools that gather what you need to know before deciding, such as Copilot Chat, Microsoft 365 Copilot and Notebooks.
- Draft and share: tools that produce something people edit together, such as Pages.
- Act on your behalf: agents that carry out steps, from Agent Builder through Copilot Studio to Foundry.
- Score a fixed choice: decision models, which pick between set options and say how confident they are.

## What every product page will cover

Each page answers the same questions in the same order, so you can compare products side by side:

1. What it does, in one sentence.
2. Which licence or meter pays for it, with the date the price was checked.
3. Which kinds of decision it helps with, and which it shouldn't touch.
4. What governance you need in place first.

## A note on figures

Prices, limits and features here are attributed to Microsoft's own documentation and dated. If a figure is an estimate, it's labelled as one.

## Where to start

The product pages are being written now. Until they're up, the Decisions (https://craigstanley.work/decisions/) section explains what a decision model is and where one fits, and the guide Is Microsoft 365 Copilot worth it? (https://craigstanley.work/is-microsoft-copilot-worth-it/) works through the licence question.

Start here: decision models [Live]
----------------------------------
Jev and Microsoft-Decision-1, side by side.
Path: /capabilities/start-here
- Jev and Microsoft-Decision-1 compared (coming)
- Where a decision model sits in a Copilot agent (coming)

Home, Code and Autopilot [Drafting]
-----------------------------------
Microsoft's September 2026 direction: Chat and Cowork merged, tenant-hosted app building, a proactive agent.
Path: /capabilities/home-code-autopilot
- The philosophy: Copilot as the OS for work (coming)
- What changes for your licences (coming)

Copilot Chat [Planned]
----------------------
The free, web-grounded tier and what it can safely do.
Path: /capabilities/copilot-chat

Microsoft 365 Copilot [Planned]
-------------------------------
The paid licence, grounded in your work data.
Path: /capabilities/m365-copilot

Notebooks [Planned]
-------------------
Gathering sources around one piece of work.
Path: /capabilities/notebooks

Pages [Planned]
---------------
Shared, editable AI output.
Path: /capabilities/pages

Agent Builder [Planned]
-----------------------
Light agents for individuals and teams.
Path: /capabilities/agent-builder

Copilot in SharePoint [Planned]
-------------------------------
Agents over sites and libraries.
Path: /capabilities/sharepoint

Copilot Studio [Planned]
------------------------
Governed agents with actions, flows and connectors.
Path: /capabilities/copilot-studio

Cowork [Planned]
----------------
Multi-step work across skills and apps, now part of Home.
Path: /capabilities/cowork

Foundry [Planned]
-----------------
Models, agents and evaluation on Azure, including Microsoft-Decision-1.
Path: /capabilities/foundry

Question to take away: Which of these do you already pay for and not use?

05 Risk
=======
URL: https://craigstanley.work/risk/
Question: What could go wrong and what do we do about it?

[Like I'm 5] Sometimes things go wrong. This part is about guessing what could go wrong before it happens, and agreeing who fixes it if it does.

[Like I'm 15] Most AI risk lists are about the model. The bigger risks sit in the decision: who relies on the answer, what happens if it's wrong, and how quickly anyone notices. I map risks to the work, rank them, and set up automatic nudges so the right person hears about a problem in time.

[For pros] Once you know how work gets done and what the tools can do, map the risks to the work and the decisions as well as the model. Then prioritise, mitigate and track, and build automations that push the right task to the right person at the right time.

## Risk follows the decision

A typical AI risk register lists the model: it might hallucinate, it might leak data, it might be biased. Those are real, but they don't tell you what to do on Monday. The same model can be harmless drafting a meeting summary and dangerous approving a payment. The risk depends on the decision it feeds.

So this section maps risk to the work. For each decision a model touches, ask:

- What happens if the answer is wrong?
- Can the decision be reversed, and how quickly?
- Who relies on the answer without checking it?
- How long before anyone would notice a problem?

## Five types of risk

Most problems fall into one of five groups: data (wrong access, wrong source), decision (wrong answer acted on), people (skills lost, trust broken), supplier (price, terms or product changes) and cost (spend that drifts out of control). The risk map (https://craigstanley.work/risk/risk-map/data-decision-people-supplier-cost-five-risk-types/) page covers each one with examples.

## From register to action

A register that nobody reads changes nothing. The goal is to turn each risk into an owner, a check and a nudge, an automated message that reaches the right person when a number moves. Power Automate and Teams can do most of this with what you already own.

## Where to start

Read Most AI risk registers list the model, not the decision (https://craigstanley.work/risk/start-here/most-ai-risk-registers-list-the-model-not-the-decision/).

Start here [Live]
-----------------
Risk follows the decision, not the tool.
Path: /risk/start-here
- Most AI risk registers list the model, not the decision: https://craigstanley.work/risk/start-here/most-ai-risk-registers-list-the-model-not-the-decision/

Risk map [Live]
---------------
Types of risk by work activity and capability.
Path: /risk/risk-map
- Data, decision, people, supplier, cost: five risk types: https://craigstanley.work/risk/risk-map/data-decision-people-supplier-cost-five-risk-types/

Prioritise [Planned]
--------------------
Likelihood, impact and reversibility on one sheet.
Path: /risk/prioritise

Mitigate and track [Planned]
----------------------------
Owners, controls and review dates that people keep.
Path: /risk/mitigate

Nudges [Planned]
----------------
Automations that turn insight into an action, a task or a question for a person.
Path: /risk/nudges
- Nudge patterns with Power Automate, Teams and Autopilot (coming)

Governance stack [Planned]
--------------------------
Agent 365, Purview and the Copilot control system in practice.
Path: /risk/governance

Question to take away: Who gets told, and how fast, when a decision model starts drifting?

06 Cost
=======
URL: https://craigstanley.work/cost/
Question: What does it cost and how do we keep it honest?

[Like I'm 5] Computer helpers cost money. This part is about knowing how much they cost, and giving spare money to the people who need it most.

[Like I'm 15] AI costs come in three kinds: licences, pay-as-you-go use and people's time. I help teams set a budget, check it every week, and move unused allowance to the people who've run out, so the spending follows the work.

[For pros] With the work, the decisions, the tools and the risks in view, set realistic budgets, forecasts and reviews. Review daily and weekly, and move unused allowance to the people who've hit their limits, so the money follows the work.

## Three kinds of cost

AI at work costs money in three ways, and most budgets only track one of them:

1. Licences, a fixed fee per person per month, paid whether the person uses the tool or not.
2. Consumption, metered use paid per message, per action or per unit of compute.
3. People's time, spent learning, checking output and fixing what went wrong.

A budget that tracks licences alone will look tidy and still overspend.

## Review often, move money quickly

Monthly reviews are too slow for metered spend. A short weekly review with a fixed agenda catches drift early. A daily check on consumption catches runaway agents.

Underspend is the other half. In most teams some people hit their limits while others barely use their allowance. Moving unused allowance to the people who've run out, weekly or even daily, puts the money where the work is.

## Where to start

Read The three costs of AI at work (https://craigstanley.work/cost/start-here/the-three-costs-of-ai-at-work/), then Set the budget weekly, move the underspend on Friday (https://craigstanley.work/cost/underspend-pooling/set-the-budget-weekly-move-the-underspend-on-friday/). For the licence question, use the Copilot break-even calculator (https://craigstanley.work/is-microsoft-copilot-worth-it/).

Start here [Live]
-----------------
Licences, consumption and people time in one view.
Path: /cost/start-here
- The three costs of AI at work: https://craigstanley.work/cost/start-here/the-three-costs-of-ai-at-work/

Budgets [Planned]
-----------------
Sustainable budgets per team and per agent.
Path: /cost/budgets

Forecasts [Planned]
-------------------
Forecasting consumption from decision volumes.
Path: /cost/forecasts

Weekly review [Planned]
-----------------------
A 20-minute review with a fixed agenda.
Path: /cost/weekly-review

Underspend pooling [Live]
-------------------------
Move daily and weekly unused allowance to people at their limit.
Path: /cost/underspend-pooling
- Set the budget weekly, move the underspend on Friday: https://craigstanley.work/cost/underspend-pooling/set-the-budget-weekly-move-the-underspend-on-friday/

Licence or pay-as-you-go [Planned]
----------------------------------
When a seat licence beats metered use, and when it doesn't.
Path: /cost/licence-or-payg

Question to take away: If you moved last week's unused AI allowance to the people who ran out, who would get it?

07 Roadmap
==========
URL: https://craigstanley.work/roadmap/
Question: What's next, for readers and for Craig?

[Like I'm 5] This is my plan for what I'll make and write next. I keep it out in the open so you can see it.

[Like I'm 15] What I'm building and writing next, in public. You can see what's being built, what's being written and what has changed.

[For pros] The plan, made public: what's being built, what's being published and what has changed.

## What this section is for

A public plan does two jobs. Readers can see what's coming and ask for what they need. And it holds the plan to account: if something slips, the change shows up here with a reason.

## What you'll find

- The plan: what's being built and published, and in what order.
- Quarterly notes: what shipped, what didn't, and what changed.
- Horizon scanning: what's coming from Microsoft and elsewhere, and how much attention each item deserves.
- What I'm building: public repos and experiments, including a decision-model starter and a threshold calculator.
- Now: what's being worked on this month.

## Ask for something

If there's a method, a comparison or a calculator you'd use, say so on LinkedIn (https://www.linkedin.com/in/anorthernman/). Requests that come up more than once move up the list.

The plan [Live]
---------------
What's being built and published, and in what order.
Path: /roadmap/plan
- The plan as a decision model (coming)

Quarterly notes [Planned]
-------------------------
What shipped, what didn't, what changed.
Path: /roadmap/quarterly

Horizon scanning [Planned]
--------------------------
What's coming and how much attention it deserves.
Path: /roadmap/horizon

What I'm building [Drafting]
----------------------------
Public repos and experiments.
Path: /roadmap/building
- Decision-model starter for Foundry (coming)
- Threshold calculator (coming)

Now [Live]
----------
What Craig is working on this month.
Path: /roadmap/now

Question to take away: What would you want to see shipped here by next spring?

Articles
========

Why decisions, not adoption
---------------------------
URL: https://craigstanley.work/blog/notes/why-decisions-not-adoption/
Section: Blog / Notes
Published: 2026-10-10

Adoption figures say how many people use an AI tool. They don't say whether work got better. Measuring decisions does.

[Like I'm 5] Lots of people using a new tool doesn't mean it's helping. It's better to check whether people are making better choices because of it.

[Like I'm 15] Most AI programmes measure how many people use the tools. That's easy to count but doesn't show whether anything improved. Counting decisions, how many, how fast and how good, shows whether the tools are paying off.

[For pros] Adoption metrics measure activity, not value. Anchor AI programmes on a decision inventory and measure volume, cycle time, quality and cost per decision before and after. Adoption then becomes a leading indicator rather than the goal.

Most AI programmes I've seen report the same numbers: licences assigned, monthly active users, prompts per user. Those numbers go up, the dashboard turns green, and nobody can say what changed about the work.

Adoption measures activity. It tells you people opened the tool. It doesn't tell you whether a case was closed faster, a payment was checked better or a customer got an answer sooner.

## Decisions are countable

A decision has a volume, a time to make, a cost and an outcome. You can count all four before and after a change. If a model now scores routine expense claims, you can measure how many claims went through, how long they took, how many were wrong, and what each cost to handle.

That's a measure a finance director recognises.

## What changes when you measure decisions

You start with the work instead of the tool. The first question becomes "which decisions cost us the most?" rather than "who hasn't used Copilot yet?"

You stop chasing usage for its own sake. A team that uses an AI tool for one high-volume decision and nothing else may be getting more value than a team that uses it for everything a little.

Risk gets clearer too. A tool that drafts emails and a tool that approves payments have the same adoption numbers and very different risks. Decision mapping shows the difference.

## Adoption still matters

It's a leading indicator. If nobody uses the tool, nothing improves. But it's a means, not the goal. Put decisions at the centre and treat adoption as one of the things that gets you there.

Which decision did your team make this week that nobody wrote down?

Map the work, then the decisions
--------------------------------
URL: https://craigstanley.work/work/start-here/map-the-work-then-the-decisions/
Section: Work / Start here
Published: 2026-10-10

Why mapping what people actually do should come before choosing AI tools, and a five-step way to do it with public frameworks.

[Like I'm 5] Before you give someone a new tool, find out what their job really is. Then look for the moments where they have to choose something.

[Like I'm 15] Start by writing down what a team actually does, using public job descriptions like O*NET and ESCO as a checklist. Then mark every point where someone makes a choice. Those choices are where AI can help, or where it shouldn't go near.

[For pros] Build a task inventory per role from O*NET or ESCO, validate it against observed work and system data, then derive a decision inventory scored on frequency, stakes and reversibility. Prioritise AI investment against that inventory, not against tool features.

## Why the order matters

If you start with a tool, you'll find uses for it. Some will be valuable, many won't, and you won't have a way to tell them apart. If you start with the work, you can see where time and risk actually go, and pick the few places where a change would matter.

## Five steps

1. Pick a role. Choose one with enough people that an improvement adds up, such as case managers, service desk analysts or finance business partners.
2. Start from a public profile. Find the closest occupation in O*NET (https://craigstanley.work/work/onet/reading-an-o-net-occupation/) or ESCO (https://craigstanley.work/work/esco/esco-occupations-and-skills-in-practice/) and copy its task list. It's a checklist, not the truth.
3. Check it against reality. Sit with people doing the work for a day. Look at their ticket queues, approval logs and calendars. Cross out tasks that don't happen and add the ones the framework missed.
4. Find the decisions. Go through each task and ask where someone has to choose. Write each one as a choice between options.
5. Score each decision. Use the inventory template (https://craigstanley.work/work/tasks-to-decisions/the-decision-inventory-template/) to record how often it happens, what's at stake, and whether it can be undone.

## What you end up with

A short list of decisions, each with a rough volume and a rough cost. That list tells you where a decision model could help, where people should stay firmly in charge, and what to measure once something changes. It also gives every later conversation, about tools, risk or budget, a shared starting point.

Reading an O*NET occupation
---------------------------
URL: https://craigstanley.work/work/onet/reading-an-o-net-occupation/
Section: Work / The O*NET lens
Published: 2026-10-10

A guide to the parts of an O*NET occupation profile that matter most for mapping work and decisions.

[Like I'm 5] O*NET is a giant American list that explains what people do in lots of different jobs. You can look up a job and see all the little tasks inside it.

[Like I'm 15] O*NET is a free database from the US Department of Labor describing about 1,000 jobs. Each profile lists tasks, tools, skills and the conditions of the work. It's a quick way to get a first draft of what a role involves.

[For pros] O*NET profiles are keyed to O*NET-SOC codes and combine task statements, detailed work activities, work context ratings, skills, knowledge, abilities and technology. For decision mapping, the most useful parts are tasks, detailed work activities and the decision-related work context items.

## What O*NET is

O*NET, the Occupational Information Network, is sponsored by the US Department of Labor. It describes around 1,000 occupations using a consistent structure, built from surveys of people in those jobs and from occupational analysts. It's free at onetonline.org, and the full database can be downloaded.

Each occupation has a code based on the US Standard Occupational Classification, such as 43-4051.00 for Customer Service Representatives.

## The parts that matter most

Tasks. Plain-language statements of what people in the role do, each rated for how important it is. This is your starting task list.

Detailed work activities. More general activities that are shared across many occupations, such as "Resolve customer complaints or problems." Because they're shared, they let you compare roles and spot work that looks similar across teams.

Work context. Ratings of the conditions of the work. Several are directly about decisions. See the next section.

Technology skills. The software people in the role use. A quick way to see which systems hold evidence about the work.

## The decision-related work context items

O*NET's work context ratings include items such as Freedom to Make Decisions, Consequence of Error, Frequency of Decision Making and Impact of Decisions on Co-workers or Company Results. Read together, they give a rough picture of how much of a role is about judgement and what it costs to get it wrong. They're a useful first filter before you map decisions in detail.

## Its limits

O*NET describes US occupations, and averages them. Your team's version of the role will differ. Use the profile as a checklist to start from, then check it against what people actually do. For UK and European roles, compare it with ESCO (https://craigstanley.work/work/esco/esco-occupations-and-skills-in-practice/).

Work context: freedom to make decisions, consequence of error
-------------------------------------------------------------
URL: https://craigstanley.work/work/onet/work-context-freedom-to-make-decisions-consequence-of-error/
Section: Work / The O*NET lens
Published: 2026-10-10

Two O*NET work context ratings that give a quick, comparable read on how much judgement a role involves and what mistakes cost.

[Like I'm 5] Some jobs let you choose a lot of things yourself. In some jobs, a mistake is a really big deal. Knowing both tells you where to be careful.

[Like I'm 15] O*NET rates every job on how much freedom people have to make decisions and how bad the consequences of a mistake are. Plot roles on those two measures and you can see quickly where AI help is low-risk and where it needs care.

[For pros] Combine Freedom to Make Decisions with Consequence of Error to place roles on a judgement-versus-stakes grid. High-freedom, high-consequence roles need decision support with strong human control. Low-consequence, high-frequency decisions are the natural first candidates for automation.

## The two ratings

Freedom to Make Decisions asks how much decision-making freedom, without supervision, the job offers.

Consequence of Error asks how serious the result would usually be if the worker made a mistake that wasn't readily correctable.

Both are rated on a scale by people doing the work, so they're consistent across occupations and can be compared.

## Putting them together

Draw a simple grid with freedom along one side and consequence along the other.

- Low freedom, low consequence: routine work with clear rules. Good candidates for automating whole decisions, with sampling to check quality.
- Low freedom, high consequence: rule-bound work where mistakes hurt, such as some compliance checks. Models can help, but every case needs a hard stop for anything unusual.
- High freedom, low consequence: varied judgement where mistakes are cheap. Good for AI that drafts or suggests, with the person choosing.
- High freedom, high consequence: expert judgement with real stakes. Decision support only, with clear human ownership and a decision record (https://craigstanley.work/decisions/decision-records/the-decision-record-schema/).

## Using it with your own data

The O*NET ratings describe the average for an occupation. Ask your own people the same two questions about their own role and compare. Where your answers differ a lot from O*NET's, find out why. Often it reveals a local rule, a missing control, or work that's been pushed onto a role without anyone deciding it should be.

ESCO occupations and skills in practice
---------------------------------------
URL: https://craigstanley.work/work/esco/esco-occupations-and-skills-in-practice/
Section: Work / The ESCO lens
Published: 2026-10-10

What ESCO is, how its occupations and skills fit together, and how to use it to describe UK and European roles.

[Like I'm 5] ESCO is a big European list of jobs and the skills each job needs. It's written in lots of languages so everyone can use the same words.

[Like I'm 15] ESCO is the European Commission's free classification of about 3,000 occupations and nearly 14,000 skills. Each occupation lists the skills that are essential and the ones that are optional. It's useful for describing roles in a way that's comparable across teams and countries.

[For pros] ESCO links occupations (mapped to ISCO-08) to essential and optional skills and knowledge concepts, in all EU languages. Use it for skills-based role descriptions, gap analysis and crosswalking to O*NET for task detail.

## What ESCO is

ESCO stands for European Skills, Competences, Qualifications and Occupations. The European Commission maintains it. It's free to browse and download, and it's published in all the official EU languages and a few others.

It has two main parts that matter here:

- Occupations, around 3,000, each mapped to the international ISCO-08 classification.
- Skills and knowledge, nearly 14,000 concepts, from "use spreadsheets software" to "assess risk factors".

## How they connect

Each occupation lists skills as essential or optional. An essential skill is one you'd expect anyone in the role to have. An optional one depends on the specific job. That split is useful when writing role profiles or spotting where training is needed.

## ESCO and O*NET compared

ESCO is strongest on skills and knowledge. O*NET is strongest on tasks and work context. For mapping decisions you usually want both: O*NET's task statements to find where decisions happen, and ESCO's skills to describe what people need to make them well. The European Commission has published a crosswalk (https://craigstanley.work/work/esco/crosswalking-esco-to-o-net/) between the two.

## Using it for UK roles

The UK uses its own Standard Occupational Classification (SOC 2020). ESCO occupations map to ISCO-08, and there are published mappings between ISCO and UK SOC, so you can move between them with some care. In practice, start with the closest ESCO occupation, check its skills against your job descriptions, and note the differences.

Crosswalking ESCO to O*NET
--------------------------
URL: https://craigstanley.work/work/esco/crosswalking-esco-to-o-net/
Section: Work / The ESCO lens
Published: 2026-10-10

How to move between ESCO occupations and O*NET profiles so you get ESCO's skills detail and O*NET's task detail for the same role.

[Like I'm 5] ESCO and O*NET are two big lists of jobs from different places. A crosswalk is like a dictionary that tells you which job on one list matches which job on the other.

[Like I'm 15] A crosswalk links matching jobs in ESCO and O*NET. That lets you combine ESCO's skills list with O*NET's tasks and work conditions for the same role, which gives you a fuller picture than either one on its own.

[For pros] Use the European Commission's ESCO–O*NET crosswalk to join ESCO occupations to O*NET-SOC codes. Treat matches as candidates. Many are one-to-many, so validate each against local job descriptions before combining task and skill data.

## Why bother

ESCO tells you the skills a role needs. O*NET tells you the tasks people do and the conditions they work in. When you want to map decisions and the skills needed to make them, having both for the same role is worth the extra step.

## The official crosswalk

The European Commission has published a mapping between ESCO occupations and O*NET-SOC occupations. It's the place to start, rather than matching job titles by hand.

## Things to watch for

One-to-many matches. One ESCO occupation can match several O*NET occupations, and the reverse. Pick the closest by reading the descriptions, not just the titles.

Different levels of detail. Some areas are split finely in one framework and grouped broadly in the other. When that happens, note it and work at the broader level.

Local reality wins. Both frameworks describe averages. Use the crosswalk to build a first draft, then check it against your own job descriptions and what people actually do.

## A simple workflow

1. Find the ESCO occupation closest to your role.
2. Use the crosswalk to find candidate O*NET occupations.
3. Take the essential skills from ESCO and the tasks and work context from O*NET.
4. Check the combined profile with people in the role.
5. Use the tasks to find decisions, as in Map the work, then the decisions (https://craigstanley.work/work/start-here/map-the-work-then-the-decisions/).

The decision inventory template
-------------------------------
URL: https://craigstanley.work/work/tasks-to-decisions/the-decision-inventory-template/
Section: Work / Tasks to decisions
Published: 2026-10-10

A one-sheet template for listing the decisions a team makes, with frequency, stakes and reversibility, so you can see where to start.

[Like I'm 5] Make a list of all the choices your team makes. Next to each one, write how often it happens and how bad it would be to get it wrong.

[Like I'm 15] The decision inventory is a simple table. One row per decision, with columns for how often it happens, how long it takes, what's at stake and whether it can be undone. Sort it and the best places to start become obvious.

[For pros] Maintain a decision inventory per role or process: decision, options, owner, volume, handling time, stakes, reversibility, data available and current error rate. Rank candidates for decision support by volume multiplied by handling time, filtered by stakes and reversibility.

## The columns

| Column | Example |
| Decision | Approve or query an expense claim |
| Options | Approve, query, reject |
| Owner | Line manager |
| How often | 400 a month |
| Time per decision | 3 minutes |
| Stakes | Low (most claims under £100) |
| Reversible? | Yes, within the pay cycle |
| Data available | Claim, receipt, policy, history |
| Known error rate | Unknown, so sample 50 to find out |

## Filling it in

Start from the task list you built from O*NET or ESCO (https://craigstanley.work/work/start-here/map-the-work-then-the-decisions/). For each task, ask where someone chooses between options. Then check the numbers against system data: approval logs, ticket histories and audit trails usually hold more than people expect.

Leave cells blank rather than guessing. A blank tells you what to find out.

## Reading it

Multiply frequency by time per decision to get the hours spent each month. Sort by that. Then look at stakes and reversibility. High-volume, low-stakes, reversible decisions are the natural first candidates for a decision model (https://craigstanley.work/decisions/decision-models/what-a-decision-model-is-and-isn-t/). High-stakes, irreversible decisions are where people should stay in charge, with better information rather than automation.

## Keep it alive

Review the inventory quarterly. Volumes change, policies change, and a decision that wasn't worth touching last year may be now.

Five questions before any decision
----------------------------------
URL: https://craigstanley.work/decisions/start-here/five-questions-before-any-decision/
Section: Decisions / Start here: five questions
Published: 2026-10-10

What, why, when, how and who. Five questions that take two minutes and stop most bad work decisions before they start.

[Like I'm 5] Before you choose something, ask five little questions. What am I choosing? Why? When does it need doing? How will I choose? Who gets a say?

[Like I'm 15] Most bad decisions at work go wrong before anyone weighs the options. Nobody agreed what was being decided, or by when, or who had the final say. Five quick questions fix that.

[For pros] Frame the decision before you analyse it. Agree the options, the objective, the deadline, the method and the decision rights. Write the answers into the decision record so the framing can be audited later.

## 1. What exactly are we deciding?

Write the decision as a choice between named options. "Should we use AI in finance?" is a topic. "Do we licence Copilot for the twelve people in accounts payable this quarter: yes, no, or pilot with four?" is a decision.

If you can't list the options, you're not ready to decide. You're ready to find options.

## 2. Why does it matter?

Say what changes if you get it right and what it costs if you get it wrong. This is where the three goals (https://craigstanley.work/decisions/start-here/save-money-make-money-look-after-people-the-three-goals/) come in: does this decision save money, make money, or look after people? If it does none of these, ask why it's on the agenda.

## 3. When does it need making?

Every decision has a last sensible moment. Decide too early and you miss information that would have helped. Decide too late and the choice gets made for you. Name the date, and name what would bring it forward.

## 4. How will we decide?

Agree the method before you see the answer. That might be a vote, a scoring sheet, a model's recommendation with a person signing off, or one person's judgement. Agreeing it first stops people picking the method that favours the answer they already wanted.

## 5. Who decides, and who gets a say?

Separate the person who makes the call from the people who are consulted and the people who are told afterwards. If a model is involved, say what it's allowed to do on its own and when it has to hand over to a person. The page Act, ask a person, or stop (https://craigstanley.work/decisions/decision-models/act-ask-a-person-or-stop/) covers this.

## Using the questions

Put the five answers at the top of the decision record. It takes two minutes. When you look back in six months, you'll know what you were trying to do, which is most of what you need to judge whether you did it.

Save money, make money, look after people: the three goals
----------------------------------------------------------
URL: https://craigstanley.work/decisions/start-here/save-money-make-money-look-after-people-the-three-goals/
Section: Decisions / Start here: five questions
Published: 2026-10-10

Every work decision worth improving serves at least one of three goals. Naming the goal tells you what to measure.

[Like I'm 5] A good choice at work does one of three things. It saves money, it earns money, or it makes things better for people. If it does none of them, why bother?

[Like I'm 15] When you're deciding whether something is worth doing, ask which of three goals it helps with. Saving money, making money, or looking after people. The goal tells you what to measure afterwards.

[For pros] Tie each decision to a primary goal (cost avoided, revenue gained, or people outcomes such as workload, safety and fairness) and a measure. Decisions with no goal get dropped. Decisions with conflicting goals get the trade-off written down.

## Save money

Most decision improvements start here because the numbers are easiest to find. A faster approval saves staff time. A better renewal decision stops paying for licences nobody uses. Measure the time or spend before and after, and count only what actually stops.

Watch for savings that just move cost elsewhere. If a model approves invoices faster but a person spends an hour a week correcting its mistakes, count that hour.

## Make money

Some decisions earn money: which leads to chase, which bids to submit, what price to quote. These are harder to measure because the counterfactual is invisible. You don't see the deal you would have won. Use a comparison group where you can, or track the hit rate over enough decisions to see a trend.

## Look after people

This goal covers workload, fairness, safety and the experience of the people affected by a decision. A rota decision that cuts overtime, or a triage rule that stops urgent cases sitting in a queue, may save little money and still be the most valuable change you make.

Measure it directly. Count the overtime hours, the wait times, the complaints. Ask the people involved.

## When goals conflict

Sometimes saving money means more work for someone. Write the trade-off down in the decision record: what you gain, what it costs, and who bears the cost. A trade-off that's written down can be revisited. One that isn't tends to get forgotten until it becomes a complaint.

Expected value on one page
--------------------------
URL: https://craigstanley.work/decisions/decision-theory/expected-value-on-one-page/
Section: Decisions / Decision theory at work
Published: 2026-10-10

A plain explanation of expected value with a worked example from the service desk, and the two places it misleads.

[Like I'm 5] If something might go well or badly, think about how good each result is and how likely it is. Then pick the choice that comes out best on average.

[Like I'm 15] Expected value means multiplying each possible result by its chance of happening, then adding them up. It lets you compare choices fairly when you can't be sure how they'll turn out.

[For pros] Expected value is the probability-weighted sum of outcomes. Use it to compare options under uncertainty, then check two things it hides: the spread of outcomes (risk of ruin) and the quality of the probability estimates.

## The idea

You can't know how a choice will turn out. You can estimate how likely each outcome is and what each would be worth. Multiply each value by its probability, add them up, and you have the expected value. Do that for every option and compare.

## A worked example

A service desk is deciding whether to let a model auto-close password reset tickets. The numbers below are illustrative.

| Outcome | Probability | Value per ticket |
| Closed correctly | 0.95 | +£4 (agent time saved) |
| Closed wrongly, user reopens | 0.05 | −£12 (extra handling and annoyance) |

Expected value per ticket = (0.95 × £4) + (0.05 × −£12) = £3.80 − £0.60 = £3.20.

At 2,000 tickets a month, that's about £6,400 a month in favour of auto-closing, if the estimates hold.

## Where it misleads

It hides the spread. Two options can have the same expected value while one occasionally produces a disaster. If one wrong decision could cause a data breach or a safety incident, expected value alone isn't enough. Set a hard limit on the worst case and only compare options that stay inside it.

It's only as good as the probabilities. "0.95" above is a guess until you've measured it. Run a small trial, count the results, and update. The page on calibration (https://craigstanley.work/decisions/decision-theory/calibration-are-your-80-calls-right-80-of-the-time/) covers how to check your estimates.

## Using it at work

You rarely need precise numbers. Rough estimates, written down, beat a debate where nobody states what they believe. The value of the exercise is that everyone can see the assumptions and argue about the right ones.

When is more information worth paying for?
------------------------------------------
URL: https://craigstanley.work/decisions/decision-theory/when-is-more-information-worth-paying-for/
Section: Decisions / Decision theory at work
Published: 2026-10-10

The value of information, explained with a work example. More data is only worth buying if it could change what you do.

[Like I'm 5] Finding out more costs time or money. It's only worth it if what you find out could make you change your mind.

[Like I'm 15] Before you ask for another report or another week of research, ask whether any answer it could give would change your decision. If not, decide now.

[For pros] The expected value of information is the improvement in decision value from knowing the answer before choosing, minus its cost. If no plausible result would change the chosen option, the information is worth nothing to this decision, however interesting it is.

## The question to ask

Before commissioning more analysis, ask: is there any result this could produce that would make us choose differently? If the answer is no, the analysis has no value for this decision. It may still be interesting, but it shouldn't delay anything.

## A worked example

A team is deciding whether to renew 200 licences for a tool at £20 per person per month. Usage data says 150 people used it last month. Someone suggests a survey to find out how much people value it, which would take three weeks.

Ask what the survey could say:

- If everyone loves it, would you renew all 200? Probably not. 50 people still aren't using it.
- If people are lukewarm, would you cancel the 150 active seats? Probably not. They're using it.

Either way, the decision is roughly the same: renew about 150 and drop about 50. The survey can't change that, so it isn't worth three weeks. Decide now and review the usage again in a quarter.

## When information is worth a lot

Information pays when you're close to a threshold and the stakes are high. If the choice between two suppliers is nearly even and the contract is large, a week's due diligence that could tip it either way is cheap.

## A rule of thumb

The value of information is highest when you're unsure, the options are close, and a wrong choice is expensive. It's lowest when one option is clearly better or the decision is easy to reverse. For decisions you can undo cheaply, act and learn from the result.

Setting a threshold you can defend
----------------------------------
URL: https://craigstanley.work/decisions/decision-theory/setting-a-threshold-you-can-defend/
Section: Decisions / Decision theory at work
Published: 2026-10-10

How to pick the confidence score at which a model acts on its own, asks a person, or stops, using the cost of each kind of mistake.

[Like I'm 5] A computer helper says how sure it is. We pick a number. If it's surer than that, it can go ahead. If not, it asks a grown-up.

[Like I'm 15] A decision model gives each case a score. The threshold is the score above which it acts without asking. You set it by comparing the cost of a wrong yes with the cost of a wrong no, and by agreeing it with the people affected.

[For pros] Set the action threshold where the expected cost of a false positive equals that of a false negative, adjusted for review capacity. Publish the threshold, the costs behind it and the review date.

## What a threshold does

A decision model gives each case a score, often written as a confidence between 0 and 1. On its own the score does nothing. The threshold turns it into an action: above it the system acts, below it the case goes to a person.

Many systems use two thresholds. Above the top one, act. Below the bottom one, reject or stop. In between, ask a person.

## Start from the cost of mistakes

There are two ways to be wrong:

- A wrong yes: the model acts when it shouldn't. An invoice gets paid that should have been queried.
- A wrong no: the model holds back when it could have acted. A clean invoice waits a day for a person.

If a wrong yes costs ten times as much as a wrong no, the model should only act when it's very sure. A common starting point is to act when the score is above cost of a wrong yes ÷ (cost of a wrong yes + cost of a wrong no). With costs of £50 and £5, that's 50 ÷ 55, so act above about 0.91. That only works if the scores are calibrated (https://craigstanley.work/decisions/decision-theory/calibration-are-your-80-calls-right-80-of-the-time/).

## Then check capacity

A threshold that sends 60% of cases to people won't survive a busy week. Look at how many cases fall in the "ask a person" band and compare that with the time people actually have. If it's too many, either accept more risk or improve the model before going live.

## Make it defensible

Write down the costs you used, the threshold you picked and who agreed it. Set a review date. When someone asks why the model approved a case, you can show the reasoning rather than just the score.

Calibration: are your 80% calls right 80% of the time?
------------------------------------------------------
URL: https://craigstanley.work/decisions/decision-theory/calibration-are-your-80-calls-right-80-of-the-time/
Section: Decisions / Decision theory at work
Published: 2026-10-10

How to check whether a person's or a model's confidence matches reality, with a simple method any team can run.

[Like I'm 5] If you say you're nearly sure lots of times, you should be right nearly every time. If you're often wrong when you say you're sure, you need to say "sure" less.

[Like I'm 15] Being calibrated means your confidence matches your results. Of all the times you said "80% sure", about 80% should turn out right. You can check this for people and for AI models by keeping score.

[For pros] Calibration compares stated probabilities with observed frequencies. Bucket predictions by confidence, compare each bucket's hit rate with its stated confidence, and recalibrate the model or retrain judgement where they diverge. Thresholds are only meaningful on calibrated scores.

## Why it matters

A model that says "0.9" should be right about nine times in ten. If it's really right six times in ten, any threshold you set on that score is wrong, and every expected value calculation built on it is too high.

People have the same problem. Most of us are overconfident on hard questions and underconfident on easy ones.

## How to check

1. Keep score. For each prediction, record the confidence given and whether it turned out right.
2. Group by confidence. Put predictions into bands: 50–60%, 60–70%, and so on.
3. Compare. For each band, work out the share that were right. A calibrated forecaster's 70–80% band is right about 75% of the time.

You need enough cases for this to mean anything. Fifty predictions per band is a reasonable minimum. Fewer than twenty and the numbers will bounce around.

## What to do with the results

If a model is overconfident, you can rescale its scores so they match observed rates. Common methods include Platt scaling and isotonic regression. If a person is overconfident, showing them their own score history is often enough to change how they estimate.

## A habit worth building

Ask people to put a number on predictions in meetings: "70% we hit the date." Write it in the decision record (https://craigstanley.work/decisions/decision-records/the-decision-record-schema/). Over a few months you'll learn whose estimates to trust, and everyone's estimates get better.

Why everyone pads their estimates
---------------------------------
URL: https://craigstanley.work/decisions/game-theory/why-everyone-pads-their-estimates/
Section: Decisions / Game theory at work
Published: 2026-10-10

Padding estimates is a sensible response to how organisations treat estimates. Game theory explains why, and what changes the behaviour.

[Like I'm 5] If you get told off for being late but not for finishing early, you'll always say a job takes longer than it really does. Everyone does it.

[Like I'm 15] People pad estimates because being late is punished and being early isn't rewarded. Each person is behaving sensibly, but the whole organisation ends up planning with inflated numbers. Change the rules and the padding goes away.

[For pros] Estimate padding is an equilibrium response to asymmetric penalties. Each team pads, so planning absorbs the combined buffer and work expands to fill it. Fix the incentives with pooled contingency, reference-class estimation and records that reward calibration over punctuality.

## The game

Each team lead is asked how long their piece of work will take. If they come in late, there's an awkward conversation. If they come in early, nothing happens, or the saved time vanishes into the next request. So the rational move is to add a buffer.

Every team does the same. The plan now contains several buffers stacked on top of each other, and work tends to expand to fill the time available. The project takes longer than if everyone had given honest estimates and shared one buffer.

## Why lecturing doesn't work

Telling people to stop padding asks them to accept personal risk for a shared benefit. Nobody wants to go first. The behaviour comes from the incentives, so the incentives are what need to change.

## What changes the behaviour

Pool the contingency. Ask for honest estimates and hold one shared buffer at project level, owned by the sponsor. Teams draw on it openly when they need it.

Ask for ranges. "Four to seven weeks, 80% confident" is more honest than "six weeks", and it can be checked for calibration (https://craigstanley.work/decisions/decision-theory/calibration-are-your-80-calls-right-80-of-the-time/).

Reward accuracy, not punctuality. Track estimates against actuals over time. The people whose ranges are reliable are the ones to trust with bigger plans.

Outcome bias: judging the decision by the luck
----------------------------------------------
URL: https://craigstanley.work/decisions/bias/outcome-bias-judging-the-decision-by-the-luck/
Section: Decisions / Bias field guide
Published: 2026-10-10

Good decisions sometimes turn out badly and bad ones sometimes work. Outcome bias is judging the decision by the result alone.

[Like I'm 5] Sometimes you make a good choice and it still goes wrong, just by bad luck. That doesn't mean it was a bad choice.

[Like I'm 15] Outcome bias is when we judge a decision only by how it turned out. A sensible decision can end badly through bad luck, and a reckless one can get lucky. To learn anything, judge the decision by what was known at the time.

[For pros] Outcome bias conflates decision quality with outcome quality. Counter it by recording information, options, probabilities and reasoning at decision time, then reviewing against that record rather than hindsight.

## What it looks like

A project manager approves a launch after reasonable testing. A rare fault slips through and causes an outage. In the review, the decision is called "reckless". Six months earlier, a colleague skipped testing entirely, got lucky, and was praised for speed.

The same thing happens with AI. A model that's right 98% of the time will eventually make a visible mistake. If one bad case is enough to switch it off, you'll lose the value of the other 98%.

## Why it matters

Outcome bias teaches the wrong lessons. People learn to avoid visible risks rather than make good decisions. They learn that luck counts as skill. Over time, the organisation becomes cautious about the wrong things.

## How to counter it

Write down the decision before the outcome is known. Record the options, what you knew, what you expected and how confident you were. The decision record (https://craigstanley.work/decisions/decision-records/the-decision-record-schema/) gives a format.

Review against the record, not the result. Ask: given what we knew, was this a reasonable choice? Was there information we should have had? Were our probabilities sensible?

Look at many decisions together. One outcome tells you little. Twenty decisions made the same way tell you a lot.

Anchoring in estimates
----------------------
URL: https://craigstanley.work/decisions/bias/anchoring-in-estimates/
Section: Decisions / Bias field guide
Published: 2026-10-10

The first number mentioned pulls every later estimate towards it. How anchoring distorts project, budget and AI-benefit estimates, and what to do about it.

[Like I'm 5] If someone says a number first, everyone else's guess ends up close to it, even if the first number was silly.

[Like I'm 15] Anchoring is when the first number you hear drags your own estimate towards it. In meetings, whoever speaks first sets the anchor. Get people to write their estimates down before anyone shares one.

[For pros] Anchoring biases estimates towards an initial value, even an irrelevant one. Use independent estimation before discussion, reference-class data instead of inside-view anchors, and treat vendor benefit claims as anchors to be tested.

## Where it shows up at work

- Project estimates: the sponsor says "this should take about a month" and every estimate clusters around four weeks.
- Budgets: last year's figure becomes the starting point, whatever has changed.
- AI benefits: a vendor case study claims a large time saving, and the internal business case quietly starts from that figure.

## Why it's hard to avoid

Anchoring works even when people know the anchor is irrelevant. In classic experiments, numbers produced at random still pulled people's estimates towards them. Knowing about the bias reduces it a little but doesn't remove it.

## What helps

Estimate alone first. Ask everyone to write their estimate down before discussion. Then reveal them together. The spread is useful information in itself.

Start from a reference class. Instead of asking "how long will this take?", ask "how long did the last five similar projects take?" Start from that and adjust.

Name the anchor. When a number has been put on the table, say where it came from and ask whether it applies here. Vendor figures in particular should be labelled with their source and date, as everywhere on this site.

What a decision model is, and isn't
-----------------------------------
URL: https://craigstanley.work/decisions/decision-models/what-a-decision-model-is-and-isn-t/
Section: Decisions / Decision models
Published: 2026-10-10

A decision model scores a fixed set of options and says how confident it is. It's smaller, cheaper and easier to audit than a general chatbot.

[Like I'm 5] A decision model is a helper that only answers one kind of question, like "yes or no?" It's small, quick and cheap, and it tells you how sure it is.

[Like I'm 15] A chatbot can write anything. A decision model just picks between set options, such as approve, query or reject, and gives a confidence score. That makes it cheaper to run and much easier to check.

[For pros] A decision model maps structured inputs to a fixed label set with a calibrated score. Its narrow output makes it cheap, testable against historical decisions and auditable. Pair it with explicit thresholds and a human route for low-confidence or high-stakes cases.

## What it is

A decision model takes the facts of a case and returns one of a fixed set of answers, with a score. "Approve, 0.93." "Escalate, 0.71." The options are agreed in advance. The model can't invent a new one.

It can be a classic machine learning classifier, a small language model told to answer only from a fixed list, or a set of rules. What matters is the shape: fixed options in, one option plus confidence out.

## What it isn't

- It isn't a chatbot. It doesn't write explanations or hold a conversation. If you want a summary of the reasoning, that's a separate step.
- It isn't the decision-maker. It recommends. The threshold (https://craigstanley.work/decisions/decision-theory/setting-a-threshold-you-can-defend/) and the people around it decide what happens with the recommendation.
- It isn't permanent. Because the inputs and outputs are fixed, you can swap one model for another and compare them on the same cases.

## Why small is good

A narrow model is cheaper per decision, often by a large margin, than asking a general-purpose model to reason from scratch. It's faster. And it's testable: you can run it over last year's decisions and see how often it agrees with what people did, and where it doesn't.

## Where it fits

Decision models suit decisions that are frequent, have clear options, and have a record of past outcomes to learn from. Expense approvals, ticket routing, eligibility checks and renewal flags are typical. They suit one-off strategic decisions badly. Those need people, information and argument.

Act, ask a person, or stop
--------------------------
URL: https://craigstanley.work/decisions/decision-models/act-ask-a-person-or-stop/
Section: Decisions / Decision models
Published: 2026-10-10

Three routes for every case a decision model scores, and how to decide which cases take which route.

[Like I'm 5] When the helper is very sure, it can go ahead. When it's not sure, it asks a person. When something looks wrong, it stops and doesn't do anything.

[Like I'm 15] Every case a model scores goes one of three ways. It acts on its own, it hands the case to a person, or it stops because something's wrong. Deciding the rules for each route in advance is what makes the system safe.

[For pros] Define three routes: autonomous action above a calibrated threshold, human review in the uncertain band or for high-stakes cases, and a hard stop when inputs are out of distribution, data is missing or a policy rule fires. Log every route taken.

## Act

The model acts on its own when three things are true: its score is above the action threshold, the stakes for this case are within agreed limits, and nothing unusual has been flagged. A routine £40 expense with a receipt and a known supplier is a typical example.

Log every automatic action so it can be sampled and checked later.

## Ask a person

The case goes to a person when the score falls in the uncertain band, when the stakes are above a set limit whatever the score, or when the case belongs to a category you've decided people always handle.

Give the person what they need to decide quickly: the facts, the model's suggestion and score, and why it was sent to them. Record what they decided. Those records are the best training data you'll get.

## Stop

Some cases shouldn't be scored at all. Stop and raise an alert when:

- required data is missing or malformed
- the case looks unlike anything the model has seen before
- a policy rule fires, such as a sanctioned supplier or an amount over a hard limit
- the model's recent behaviour has drifted, for example a sudden jump in approvals

Stopping isn't a failure. It's the system knowing the edge of what it can safely do.

## Keeping the routes honest

Check the split every week. If the share going to people climbs, the model or the data may have changed. If it falls to almost nothing, check that the threshold hasn't been loosened without anyone agreeing. The nudges (https://craigstanley.work/risk/) in the Risk section can watch these numbers for you.

The decision record schema
--------------------------
URL: https://craigstanley.work/decisions/decision-records/the-decision-record-schema/
Section: Decisions / Decision records
Published: 2026-10-10

A short, fixed format for writing down a decision at the moment it's made, so it can be reviewed fairly later.

[Like I'm 5] When you make a choice, write down what you chose, why, and how sure you were. Later you can look back and learn from it.

[Like I'm 15] A decision record is a short note made when the decision is made, not afterwards. It lists the options, what you knew, what you chose and how confident you were. It's the only fair way to judge a decision later.

[For pros] Capture the decision at the time it's made, using a fixed schema: context, options, information, choice, confidence, owner, review date. Store records where they can be queried, so you can measure calibration and decision quality across many decisions.

## The fields

| Field | What to write |
| Decision | One line, framed as a choice |
| Options | The options you considered, including "do nothing" |
| Goal | Save money, make money, or look after people, and the measure |
| What we knew | The key facts and their sources |
| Choice | The option picked |
| Confidence | A percentage: how likely is this to achieve the goal? |
| Expected result | What you expect to see, by when |
| Owner | Who made the call |
| Model involvement | Whether a model scored it, its suggestion, its score |
| Review date | When you'll look again |

## Keep it short

A record that takes twenty minutes won't get written. Most fields take a line. If you're using a model, its suggestion and score can be filled in automatically.

## Where to keep records

Keep them somewhere they can be searched and counted. A SharePoint list or a Dataverse table works. Free-text notes scattered across email don't. Once you have a few hundred records, you can measure calibration (https://craigstanley.work/decisions/decision-theory/calibration-are-your-80-calls-right-80-of-the-time/), spot repeated patterns, and see whether model-assisted decisions do better than unassisted ones.

## Review

On the review date, add three fields: what happened, whether the decision was reasonable given what was known, and what you'd do differently. Keep the outcome and the judgement of the decision separate. That's the defence against outcome bias (https://craigstanley.work/decisions/bias/outcome-bias-judging-the-decision-by-the-luck/).

Most AI risk registers list the model, not the decision
-------------------------------------------------------
URL: https://craigstanley.work/risk/start-here/most-ai-risk-registers-list-the-model-not-the-decision/
Section: Risk / Start here
Published: 2026-10-10

Why AI risk registers full of generic model risks don't help, and how to rebuild them around the decisions models feed.

[Like I'm 5] Instead of worrying about what the computer helper might get wrong in general, look at each job it does and ask what would happen if it got that job wrong.

[Like I'm 15] Most AI risk lists say things like "the model might make things up." That's true but doesn't tell you what to do. It's more useful to list each decision the AI helps with and ask what happens if it's wrong there.

[For pros] Generic model risks (hallucination, bias, leakage) don't prioritise action. Rebuild the register with one row per decision the model influences: impact if wrong, reversibility, reliance, detection time, owner, control and nudge. Model-level risks become inputs to each row.

## The usual register

Open most AI risk registers and you'll find the same rows: hallucination, bias, data leakage, prompt injection, over-reliance. Each has a rating and a mitigation like "user training" or "human in the loop".

None of these are wrong. The trouble is that they're the same for every organisation and every use, so they don't help anyone decide what to do first.

## Start from the decision

The same model can be low-risk in one place and high-risk in another. Summarising a meeting and approving a refund carry very different stakes, even if the model behind them is identical.

So rebuild the register around decisions. One row per decision that a model influences, with these columns:

| Column | Question |
| Decision | What choice does the model affect? |
| Impact if wrong | What does a wrong answer cost, and who bears it? |
| Reversible? | Can it be undone, and how quickly? |
| Reliance | Does anyone check the output before acting? |
| Time to detect | How long before a problem would be noticed? |
| Owner | Who is accountable? |
| Control | What stops or catches a bad answer? |
| Nudge | What alert reaches the owner when something moves? |

## Where model risks go

Hallucination, bias and leakage don't disappear. They become inputs to each row. For a refund decision, ask what a hallucinated policy reference would cost there. That's a question someone can answer.

## A quick test

Pick any row in your current register and ask: who would do something different on Monday because of this? If the answer is nobody, rewrite it around a decision.

Data, decision, people, supplier, cost: five risk types
-------------------------------------------------------
URL: https://craigstanley.work/risk/risk-map/data-decision-people-supplier-cost-five-risk-types/
Section: Risk / Risk map
Published: 2026-10-10

A simple way to group AI risks at work into five types, with examples and a first control for each.

[Like I'm 5] Things can go wrong in five ways. The information is wrong, the choice is wrong, people get hurt or confused, the company you buy from changes things, or it costs too much.

[Like I'm 15] Most AI risks at work fit into five groups: data, decisions, people, suppliers and cost. Sorting risks this way helps you spot gaps and give each one an owner who knows that area.

[For pros] Classify AI risks as data, decision, people, supplier or cost. Assign owners by type (data owner, process owner, line manager, commercial lead, budget holder) and pair each risk with a preventive control and a detective nudge.

## Data

The model sees information it shouldn't, or works from information that's wrong or stale.

- Example: an agent grounded on SharePoint surfaces a salary file because permissions were too broad.
- First control: review permissions on the sites the model can reach before you switch it on.

## Decision

The model gives a wrong answer and someone acts on it.

- Example: a triage model routes an urgent complaint to a low-priority queue.
- First control: a threshold and a hard stop (https://craigstanley.work/decisions/decision-models/act-ask-a-person-or-stop/) for categories that always go to a person.

## People

The way people work changes in ways nobody chose.

- Example: staff stop checking model output because it's usually right, and lose the skill to spot when it isn't.
- First control: sample automated decisions for human review every week, and rotate who does it.

## Supplier

The vendor changes the price, the terms, the model or the product.

- Example: a model version is retired and its replacement scores cases differently.
- First control: keep a test set of past cases and re-run it whenever the model changes.

## Cost

Spending drifts beyond the budget.

- Example: a looping agent burns through a month's consumption allowance in a weekend.
- First control: daily spend alerts and a hard cap per agent. See the Cost (https://craigstanley.work/cost/) section.

## Using the five types

Go through your decision-based register and tag each row with one or more types. If one type has no rows at all, that's usually a gap rather than good news.

The three costs of AI at work
-----------------------------
URL: https://craigstanley.work/cost/start-here/the-three-costs-of-ai-at-work/
Section: Cost / Start here
Published: 2026-10-10

Licences, consumption and people's time. Most AI budgets track the first, some track the second, and almost none track the third.

[Like I'm 5] AI costs money in three ways. You pay to have it, you pay each time you use some of it, and people spend time learning it and checking it.

[Like I'm 15] AI tools cost money through licences, through pay-as-you-go use, and through the time people spend learning, checking and fixing. If a budget only counts licences, it'll look fine while the real cost grows.

[For pros] Track total cost per decision supported: licence cost allocated by active use, metered consumption, and people time for review, rework and enablement. Compare it with the cost of the decision before AI, not with the licence price alone.

## 1. Licences

A fixed monthly fee per person. Easy to budget, easy to see. The catch is that you pay whether or not the person uses it. A licence that sits idle for a quarter is pure cost.

Measure: cost per active user, not cost per licence. Divide total licence spend by the number of people who used the tool most weeks.

## 2. Consumption

Metered charges: per message, per agent action or per unit of compute. Agents built in Copilot Studio or Foundry often run on this kind of billing. It scales with use, which is fair, and can spike without warning, which isn't.

Measure: cost per decision. If an agent handles 3,000 tickets a month at a known consumption cost, you can compare that with what it cost a person to handle them.

## 3. People's time

Hours spent on training, prompting, checking output, correcting mistakes and dealing with the cases the model sends back. This rarely appears in any budget, and it can be the largest of the three.

Measure: sample it. Ask a handful of people to log time spent checking and fixing AI output for two weeks. Multiply up.

## Putting them together

Add all three and divide by the number of decisions supported. Compare that with the cost per decision before the change. That's the honest test of whether it pays. The break-even calculator (https://craigstanley.work/is-microsoft-copilot-worth-it/) does a simpler version for licences.

Set the budget weekly, move the underspend on Friday
----------------------------------------------------
URL: https://craigstanley.work/cost/underspend-pooling/set-the-budget-weekly-move-the-underspend-on-friday/
Section: Cost / Underspend pooling
Published: 2026-10-10

A simple routine for AI consumption budgets. Set allowances weekly and move unused allowance to people who hit their limits.

[Like I'm 5] Every week, everyone gets some AI to use. On Friday, if someone hasn't used theirs, it goes to someone who ran out.

[Like I'm 15] Instead of giving everyone a fixed AI allowance for the year, set it each week. On Friday, look at who didn't use theirs and who ran out, and move the spare allowance to the people who need it.

[For pros] Allocate consumption allowances weekly, not annually. Run a fixed Friday review: reclaim unused allowance above a floor, reallocate to users at their cap, and flag sustained heavy users for a licence-versus-metered review. Money follows the work.

## Why weekly

Annual and monthly budgets assume use is steady. It isn't. A finance team's use spikes at month end. A project team's use spikes before a deadline. With fixed allowances, some people run out while others' allowances go unused.

A weekly cycle is short enough to respond to that and long enough not to become admin.

## The Friday routine

A 20-minute review with the same agenda every week:

1. Spend against budget. Total consumption this week compared with the weekly budget.
2. Who's at the cap? People who hit their limit, and what they were doing.
3. Who's well under? People who used little of their allowance.
4. Move it. Reclaim unused allowance above a small floor and give it to the people at the cap for next week.
5. Flag patterns. Anyone at the cap three weeks running may be better off on a licence than on metered use, or the reverse.

## Making it fair

Publish the rules. Everyone keeps a floor so nobody's allowance is taken entirely. Reallocation is based on use, not seniority. Anyone can ask for more with a one-line reason.

## Automating it

Most of this can run on the usage reports you already have. A Power Automate flow can pull the week's figures, draft the reallocation and post it to a Teams channel for the budget holder to approve. The meeting then becomes a quick check rather than a spreadsheet exercise.

Is Microsoft 365 Copilot worth it?
==================================
URL: https://craigstanley.work/is-microsoft-copilot-worth-it/

Short answer: It's worth it for the people whose repeated decisions and documents it speeds up by more than the licence costs, and not for everyone else. Work out the break-even minutes per week first, then licence the roles that clear it.

[Like I'm 5] Copilot is a helper that costs money every month. If it saves you more time than it costs, it's a good deal. If you never use it, it's like paying for sweets you don't eat.

[Like I'm 15] A Copilot licence costs a fixed amount per person per month. Divide that by what an hour of that person's time is worth and you get the number of hours a month it has to save to pay for itself. Some people will beat that easily. Others won't open it. The smart move is to buy for the first group and keep checking.

[For pros] Treat each licence as a repeated decision with an expected value. Cost is fixed per seat; value depends on active use, minutes saved per task, task frequency and the loaded cost of the role. Set a break-even threshold per persona, licence the personas that clear it with margin, and review usage monthly so idle seats move to people at their limit.

Five questions before you buy
-----------------------------
1. What decision or document will it change? Name the task. "Productivity" is not a task. "Drafting the weekly case summary" is.
2. How often does that happen? Daily tasks pay back. Quarterly ones rarely do on their own.
3. What is an hour of that person's time worth? Use the loaded cost (salary plus on-costs), not the salary.
4. Who checks the output? If a person has to redo the work, count their time as a cost.
5. What will you stop paying for? Licences that sit idle for 30 days should move to someone on the waiting list.

Usually pays
------------
- Roles that write, summarise or search across Microsoft 365 content every day
- Teams that run the same approval or triage decision many times a week
- People who already use Copilot Chat and hit its limits

Usually doesn't, yet
--------------------
- Roles that rarely touch documents, mail or Teams
- Organisations whose files are badly permissioned (fix that first; Copilot will find everything)
- Pilots with no named task, owner or review date

FAQ
---
Q: How much does Microsoft 365 Copilot cost?
A: Prices vary by agreement, region and currency, and Microsoft changes them. Use the price on your own agreement. The calculator on this page starts at £25 per user per month as a labelled assumption, so replace it.

Q: How many hours does Copilot need to save to pay for itself?
A: Monthly price divided by the loaded hourly cost of the person. At £25 a month and £30 an hour that is 50 minutes a month, or about 12 minutes a week, per active user.

Q: Should we buy Copilot for everyone?
A: Usually not at first. Licence the personas whose tasks clear the break-even with a clear margin, measure for a month, and move idle seats.

Q: What's the difference between Copilot Chat and Microsoft 365 Copilot?
A: Copilot Chat is the free, web-grounded tier. Microsoft 365 Copilot is the paid licence that works over your own mail, files, meetings and chats.

Q: How do we measure whether it's working?
A: Pick one task per persona, time it before and after, and track active use in the Copilot usage reports. Write the decision down with the numbers you used, so you can check it later.

About Craig Stanley
===================
Microsoft AI consultant and technical architect, based in Whitley Bay. Over the last few years I've delivered Microsoft 365 Copilot, Copilot Studio agents, Azure AI Foundry work and governance for UK public sector and financial services organisations.

What interests me is the decision underneath the tool: what it costs, what it risks, and whether a small, transparent model can make it better. I write the methods up here and on Substack so anyone can use them.

- Built a licence transition model for roughly 74,000 Microsoft 365 users moving to persona-based groups.
- Delivered Copilot adoption, agent builds and governance across public sector and financial services clients.
- Runs the Decision Lab: interactive experiments on how people and AI make choices.

LinkedIn: https://www.linkedin.com/in/anorthernman/
X: https://x.com/anorthernman
Substack: https://anorthernman.substack.com

Vendor figures are attributed and dated. Assumptions are labelled.
About me

Craig Stanley

Microsoft AI consultant and technical architect, based in Whitley Bay. Over the last few years I've delivered Microsoft 365 Copilot, Copilot Studio agents, Azure AI Foundry work and governance for UK public sector and financial services organisations.

What interests me is the decision underneath the tool: what it costs, what it risks, and whether a small, transparent model can make it better. I write the methods up here and on Substack so anyone can use them.

Find me