Nobody knows exactly how much the helpers will cost. So say "probably this much, but it could be a bit less or a lot more", and plan for the "a lot more".
A single forecast number is almost always wrong. A range is more honest: a low case, a likely case and a high case, each with its reason. The budget can then hold the likely case, with an agreed plan for what happens if spend heads towards the high case.
Present consumption forecasts as a three-point range with named drivers per case. Budget at the likely case, set alerts at the likely case and caps near the high case, and pre-agree the action at each trigger. Track which case actuals land in to calibrate future ranges.
One number hides the risk
A forecast of "£2,000 a month" sounds precise. What it doesn't say is whether £3,000 is a remote possibility or quite likely. For licences that hardly matters, because seat costs barely move. For metered AI use it matters a lot, because a new trigger or a popular agent can double the bill in a week.
I'd rather show three numbers and the reason for each. It's the same idea as calibration in the Decisions section: say how sure you are, then check later whether you were right to be.
Building the three cases
Each case uses the same sum as Forecasting Copilot Credits from decision volumes. What changes is the input I'm least sure about.
The low case assumes volumes at their quietest recent level and the agent used no more than today. The likely case uses normal volumes and the trend in use so far. The high case assumes a peak month, wider use (another team adopts the agent, say) and a modest rise in credits per decision, for example from a change of model.
A worked example
These numbers are illustrative.
| Case | Decisions | Through agent | Credits each | Credits | Cost at $0.01 |
|---|---|---|---|---|---|
| Low | 2,000 | 60% | 12 | 14,400 | $144 |
| Likely | 2,500 | 70% | 14 | 24,500 | $245 |
| High | 3,500 | 85% | 18 | 53,550 | $535.50 |
The high case is more than twice the likely case, even though no single input moved dramatically. That's the useful lesson for me: modest changes multiply.
Which number goes in the budget
I'd put the likely case in the budget, and decide in advance what happens at each point:
| Trigger | Action |
|---|---|
| Spend passes the likely case pace | Budget holder reviews the week's use at the weekly review |
| Spend reaches midway to the high case | Agent owner checks triggers and credits per decision |
| Spend reaches the high case | Cap applies; extra use needs a named approver |
Microsoft's tools can back up the first two triggers with alerts. Azure budgets support alerts on actual and forecast cost, and Microsoft 365 spending policies support alerts and per-user limits. The actions are still a human decision, which is why I write them down before the month starts.
Learning from the range
At the end of each month, note which case the actual landed in. If actuals keep landing above the likely case, the forecast is biased low and the likely case should move. If they always land in the middle and never near either edge, the range may be wider than it needs to be. A few months of this is enough to see the pattern.
Where I got stuck
The high case is the hardest to set honestly. It's tempting to make it a round multiple of the likely case. I've tried to build it from named events (a peak month, a second team, a model change) because then someone can say whether each event is plausible. I'm not sure yet how to put a probability on each case without inventing precision I don't have.
Sources
- Microsoft, Copilot Studio licensing guidance ($0.01 per Copilot Credit), accessed 11 October 2026.
- Microsoft Learn, Tutorial: Create and manage budgets, accessed 11 October 2026.
- Microsoft Learn, Understand usage-based billing and cost management for Copilot Credits, accessed 11 October 2026.