The dashboards answer a question nobody asked
Copilot reporting tells you how much the product is being used. It cannot tell you whether the spend was worth it. Those are different questions, and the distance between them is where most adoption reporting quietly falls over.
The measurement surface is decent. Between Copilot Analytics, the Microsoft 365 admin centre readiness, usage and credits reports, Purview audit logs and Copilot Studio analytics you can see who holds a licence, who opened the thing, which apps they used it in, and what your credit consumption looks like. It will tell you if a rollout has stalled in one directorate. It will not tell you whether anyone’s work got better.
So the first honest move is a sentence to whoever asked for the number: we can show usage, we cannot yet show value, and here is what we are doing to close that. I have watched people avoid saying it for a year and then get asked anyway, at a worse moment, with less goodwill in the room.
The best evidence in public comes with a warning label attached
Two studies are worth putting in a business case, and both need handling.
The stronger design is Microsoft’s own. “Early Impacts of M365 Copilot”, the paper’s own title, published April 2025, was a randomised experiment across more than 6,000 workers at 56 firms over six months. It found around 30 minutes less time spent reading email per week, documents completed 12% faster, and roughly 40% of workers with access using it regularly. Randomisation matters. It is also research by the vendor into the vendor’s product. Say so when you cite it, rather than let someone else notice.
Second is the Australian whole-of-government trial, which ran January to June 2024 across more than 5,765 licences and was evaluated by Nous Group for the Digital Transformation Agency with the Australian Centre of Evaluation. The headline numbers are attractive: 69% said Copilot improved the speed of task completion and 61% said it improved the quality of their work, with self-reported daily time savings of 1.1 hours on summarising, 1.0 on first drafts, 0.9 on meeting minutes and 0.8 on searching for information. 77% were optimistic at the end of the trial and 86% wanted to continue, though only one in three used it daily.
Now the part that gets left off the slide. The report says its own time-saving figures are approximations and likely quote the upper bound. It says editing was almost always needed. It records that up to 7% of participants said Copilot added time, because of the verification the output required. And it notes that quality gains were more subdued than improvements to efficiency. Those are the evaluator’s words about the evaluator’s numbers, and dropping them is not summarising, it is misrepresenting.
If you quote the 1.1 hours and not the upper-bound caveat, you have not made a business case, you have made a hostage. Someone in that organisation will use the product, find it slower for their task, and be right. Then every other number you presented is suspect.
One more thing about that trial: it is 2024 data. In August 2026 that is two product generations back. It is still the best independent look at a large deployment, and it is describing a tool that no longer exists in the form tested.
Measure two tasks properly instead of forty badly
Pick two or three tasks a real team does often, that have a visible output, and that somebody complains about. Meeting minutes, a recurring report, a standard client response. Baseline them before anyone gets a licence: how long, how many drafts, how much rework, who has to check it. That baseline is the whole exercise. If you skip it you are reduced to asking people how they feel, which is what the survey data above already did, better than you will.
Then measure the same task with the same people, and count the verification separately. Verification time is the number that gets hidden. A first draft in three minutes instead of twenty is not a saving if the checking takes fifteen and a manager now reads two drafts instead of one. If the honest answer is that the task got faster and the checking got longer, publish that. It is more useful than a green arrow, and it points at what to fix.
Track who stopped, too. Not just who logged in. Quiet non-use is normal, not a failure signal on its own, but the reason for it is the most valuable thing in your data. People stop because it could not see the document they needed, or because the output was wrong in a way that was expensive to catch, or because their work is not text. Each of those has a different fix and only one of them is training.
Before you build any of it, ask whoever wants the measurement what number would change their decision. Ask what they would do if it came back flat. If there is no answer, you are not being asked to measure anything. You are being asked to produce reassurance, which is a shorter piece of work with a different name.
What I say when I cannot prove it
I have stopped presenting value estimates I cannot defend line by line. It cost me a slide people liked. What replaced it is a page with three parts: what we can count, what we can only ask about, and what we do not know yet. The first part is small and true. The second is labelled as self-report every time it appears. The third is the shortest, and it is the one senior people read twice.
Nobody has ever objected to that page. They object to the confident one, later, when a number turns out to have been a survey wearing a suit.