Most agent business cases I am asked to review are costed in messages. The unit has not been messages since 1 September 2025. That is not a pedantic point about vocabulary, because the thing that replaced it prices different actions at wildly different rates, and a case built on an average cost per interaction will be wrong by an order of magnitude in whichever direction is least convenient.
The good news is that the model is documented and it is not complicated. It is just not the model most people are carrying in their heads.
The unit is the Copilot Credit
Microsoft changed the currency from messages to Copilot Credits on 1 September 2025. The quantity in a prepaid pack and the pay-as-you-go rate did not change at that point. The name and the accounting did.
Capacity is bought in prepaid packs of 25,000 Copilot Credits per month, per pack, or consumed pay-as-you-go.
I am not going to put a price per credit in this post. Get that from your agreement or your licensing desk, because it is the one number in this whole area I would not want a reader to quote back to a CFO on my authority. What I will say is that the number of credits a given agent burns is knowable in advance, and almost nobody works it out before they build.
Note also that the Copilot Studio user licence has a list price of $0 per user per month. The tenant licence carries the credit entitlement. If your cost model has a per-maker seat cost in it and no credit line, you have the shape of the thing exactly backwards.
The rate card is the business case
Here is what actually spends credits, per Microsoft’s published billing rates.
| Action | Copilot Credits |
|---|---|
| Classic answer | 1 |
| Generative answer | 2 |
| Agent action | 5 |
| Tenant graph grounding | 10 |
| Agent flow actions | 13 per 100 actions |
| Content processing | 8 per page |
| Text and generative AI tools | 1 basic / 15 standard / 100 premium, per 10 responses |
| Voice | 10 classic / 35 generative / 75 premium generative, per minute |
Read down that table and one line should stop you. Tenant graph grounding is 10 credits. A classic answer is 1. The single most attractive feature of a Microsoft-native agent, that it can answer from your own organisational content, is the most expensive line on the rate card. Ten times the cost of a scripted response.
That is not a criticism. Grounding in the tenant graph is the reason to build the agent at all, and it is doing real work. But it inverts a common design instinct. The tempting build is an agent that grounds on everything and answers anything, because that demos beautifully. The affordable build scopes grounding to the questions that need it, and answers the rest from a deterministic path at a fraction of the cost.
The premium tiers deserve the same attention. Premium generative AI tools run 100 credits per 10 responses against 1 for basic. Premium generative voice is 75 credits a minute. A voice agent that nobody costed per minute of conversation is a genuinely expensive way to find out how much your customers like talking.
Content processing at 8 credits per page is the one that catches document-heavy use cases. Take the page count of the thing you intend to feed it, multiply, and do that before you build rather than after.
It is free, until the identity changes
The rule that most changes agent economics is this one: employee-facing agent usage is no charge for Microsoft 365 Copilot licensed users, subject to fair-use limits, when the agent runs under the licensed user’s identity.
Read that clause carefully, because every word in it is load-bearing. Employee-facing, so not a customer-facing agent. Licensed users, so the person interacting has Microsoft 365 Copilot. Under the licensed user’s identity, so not under some other identity. Subject to fair use, so not unlimited.
There are two documented exceptions and both matter:
- Computer-Using Agents are not included in the Copilot user subscription licence. If your design uses one, it sits outside the free path and needs costing separately from the start.
- Agent flows are only free via the “When an agent calls the flow” trigger. Any other trigger and you are consuming credits at 13 per 100 flow actions.
That second one is where I see plans go wrong most often. Someone builds the agent, it works, it is free, everyone is pleased. Then someone adds a scheduled trigger so it runs overnight without a human, which is a sensible product decision and a silent change to the billing model. The agent did not change. The identity it runs under did, and that is the thing being billed.
If you take one working rule from this post, take that one. Draw your agent architecture, then colour in every path where a human with a Copilot licence is not the one initiating the call. Those paths are your cost model. The rest is mostly rounding.
The harness decides the bill
Writing about Copilot Studio as a single runtime stopped being accurate this year. It now exposes three harnesses, and they bill differently.
The GitHub Copilot harness handles reasoning-heavy multi-step work, natively creates and edits Word, Excel, PowerPoint and PDF files, supports skills and memory, and runs sandboxed. It is billed in Copilot Credits.
Rule-based agents and agent flows run on the standard harness.
Extending Microsoft 365 Copilot Chat with enterprise knowledge is the job of the Copilot chat harness, which is either consumption-based or included in the Microsoft 365 Copilot user subscription licence.
Three runtimes, three cost profiles, one product name. If your business case says “built in Copilot Studio” and does not name a harness, it does not yet describe a cost. That sentence should be a standing question in any design review from now on.
125% is where the plan meets the wall
Overage is not a gentle slope. Enforcement fires at 125% of prepaid capacity and disables custom agents. For agent flows, enforcement blocks new flow runs instead.
That is a hard stop on a business process, triggered by a number nobody in the business is watching, at a threshold 25% above a plan that was probably an estimate. Custom agents switching off is an availability event. It will be reported to you as a fault before anyone connects it to consumption.
So put the credit consumption report in front of a named human on a schedule, from the day the first agent goes live. The admin centre reports cover readiness, usage and credits, and Copilot Studio has its own analytics. The data is there. The habit of looking at it is what is missing.
The costs that never appear in the credit line
Credits are not the whole bill, and a few of the other lines are large enough to matter.
If your agents need a machine to work on, Windows 365 for Agents is $0.40 per hour pay-as-you-go, or $5 per Cloud PC per month for always-available in the US geography. An hourly rate looks small until something runs continuously.
Governance is licensed separately. Agent 365 standalone is $15.00 per user per month paid yearly. Microsoft 365 E7, generally available since 1 May 2026, bundles E5, Copilot, Entra Suite and Agent 365 at $99.00 per user per month paid yearly. The Copilot add-on on its own is $30.00 per user per month paid yearly, or $31.50 paid monthly on an annual commitment, and since 1 June 2026 there is a 15% saving on three-year commitments at 300 or more Copilot licences.
Work IQ, the named intelligence layer, has an API billed independently of Copilot licensing. If a developer in your organisation is building against it, that is a separate line on a separate bill.
Then there is the cost nobody puts in the spreadsheet: somebody has to own each agent, check what it produced, and retire it when the process changes. I have watched more agent value destroyed by an unowned agent quietly producing wrong output for a quarter than by any consumption bill.
How I would cost an agent before building it
Start from the interaction, not the licence. Write down what one use of this agent involves: how many answers, generative or classic, how many actions, whether it touches tenant grounding and how often, how many pages it processes, how many flow actions it fires. That is a line-by-line credit estimate for one run, and it takes about twenty minutes with the rate card open.
Multiply by realistic volume, not hoped-for volume. Then check the number against 125% of whatever capacity you intend to buy, because that is the point at which the agent stops rather than the point at which it costs more.
Separately, mark every path where the agent runs as something other than a licensed user. Computer-Using Agents, scheduled flows, unattended runs, anything customer-facing. Those come out of the free path and get costed at the rate card.
Name the harness in the design document. Name the owner in the same document.
An agent that costs more than the work it replaces is not a failure of the pricing model. It is a design that was never costed, and the rate card was published the whole time.