Craig Stanley
Home / Cost / Start here / Cost per decision: the number I'd track first

Cost per decision: the number I'd track first

Why I divide AI spend by the decisions it supports, how to work the number out from licences, credits and people's time, and where it gets messy.

· 3 min read · Craig Stanley
In short, explained

If you buy a helper, count how many jobs it helps with. Then you can see how much each job costs, and whether that's more or less than before.

Total AI spend tells you very little on its own. Divide it by the number of decisions the AI helped with and you get a cost per decision, which you can compare with what that decision cost before. That comparison is the real test of whether it pays.

Track fully loaded cost per decision supported: allocated licence cost, metered Copilot Credits or Azure consumption, and sampled people time for review and rework, divided by decision volume from system logs. Compare it with the pre-AI baseline cost for the same decision type.

Why one number

I keep coming back to a single question when people show me an AI budget: what did each decision cost? A monthly bill of £4,000 sounds a lot or a little depending on what it bought. If it supported 20,000 expense approvals, that's 20p each. If it supported 40 contract reviews, that's £100 each. Both might be good value, but you can't tell until you compare them with what those decisions cost before.

Cost per decision is my attempt to make the three costs of AI add up to something you can act on. It also ties the Cost section back to the decision inventory, which already counts how often each decision happens.

The sum

For one decision type, over one month:

  1. Take the seat cost of the people who make this decision, and share it out by roughly how much of their Copilot use goes on it. This is the fuzziest part, so I keep the assumption written next to the number.
  2. Add the metered cost that this decision type drives: Copilot Credits for an agent, or Azure consumption for a model in Foundry.
  3. Add the time people spend checking, correcting and handling the cases the system sends back, priced at a loaded hourly rate.
  4. Divide by the number of decisions made that month.

Then do the same sum for the way the decision was made before, which is usually just people's time.

A worked example

These numbers are illustrative. They're there to show the arithmetic.

A team handles 3,000 supplier invoice queries a month. Before any AI, each query took 6 minutes of an officer's time at a loaded rate of £30 an hour. That's £3 a query, or £9,000 a month.

Now an agent drafts the reply and a person checks it:

Cost lineMonthly cost
Agent consumption (assume 15 Copilot Credits a query at $0.01 each, and $1 = £0.77 for illustration)about £347
Checking time (2 minutes a query at £30 an hour)£3,000
Rework on the 5% the agent gets wrong (6 minutes each)£450
Share of licences used on this work (assumption)£400
Totalabout £4,197

That's roughly £1.40 a query against £3 before. Most of the saving comes from checking being faster than writing. If checking crept back up to 5 minutes, the saving would mostly disappear, which is why I'd measure checking time before I'd fuss over the credit rate.

The $0.01 per Copilot Credit is Microsoft's published pay-as-you-go rate. The 15 credits a query is my assumption. Microsoft's billing page shows how a single answer can use several credits, for example 10 for tenant graph grounding plus 2 for a generative answer. Users with a Microsoft 365 Copilot licence don't incur these charges for employee-facing agents that run under their own identity, so the consumption line changes depending on who is asking.

Where the number misleads

Cost per decision rewards volume. An agent that handles lots of easy cases and passes the hard ones to people will look cheap, while the hard cases quietly get more expensive. I split the number by route (handled automatically, checked by a person, sent back) so that shift shows up.

It also says nothing about quality. A cheaper decision that's wrong more often isn't a saving. Pair it with an error rate from a monthly sample, and keep the two side by side.

What I'm still checking

I haven't found a clean way to allocate seat licences to decision types. Asking people to estimate their split works, but it's an estimate, and I'd rather label it as one than pretend the number is more precise than it is. I'm also checking how Microsoft's new cost reports group consumption, because grouping by agent would make step 2 much easier.

Sources

Read next

A question to take awayIf you moved last week's unused AI allowance to the people who ran out, who would get it?

About me

Craig Stanley

Microsoft AI consultant and technical architect, based in Whitley Bay. Over the last few years I've delivered Microsoft 365 Copilot, Copilot Studio agents, Microsoft Foundry (formerly Azure AI Foundry) work and governance for UK public sector and financial services organisations.

What interests me is the decision underneath the tool: what it costs, what it risks, and whether a small, transparent model can make it better. I write the methods up here and on Substack so anyone can use them.

I write this site to learn in public: explaining each idea simply is how I check I understand it. Why I write this site.

Find me