Most Copilot pilots are designed so they cannot fail
The standard shape is a group of volunteers, a supportive business unit, a satisfaction survey at the end, and a renewal date already in the diary. That produces a result. It does not produce evidence, because there was never an outcome available that would have stopped the purchase.
You can spot one from the paperwork. If the success criteria are written as percentages of people reporting a positive experience, the pilot is a marketing exercise with a spreadsheet attached. Volunteers who asked for a licence report a positive experience. That is what volunteering means.
The point of a pilot is to buy information you did not have, and information only arrives when the answer could have gone either way. If nobody in the room can describe the finding that would make them say no, you are not running a pilot. You are running a rollout with a smaller first wave, which is a legitimate thing to do, and you should call it that in the paperwork so nobody mistakes it for proof later.
The shape worth copying belongs to Microsoft’s own researchers
Microsoft Research ran the study that most internal pilots should be imitating. A randomised experiment, over 6,000 workers across 56 firms, run over six months, published in April 2025. It found around 30 minutes less time reading email per week and documents completed 12% faster.
It is Microsoft’s own research into Microsoft’s own product, so treat the effect sizes with the scepticism that deserves. The design is the part to steal, and three things about it matter.
There was a control group. Access was randomised rather than requested, which is what lets you say the difference came from the tool and not from the kind of person who signs up for new tools.
It ran for six months. Novelty effects are large and short, and a four-week pilot measures mostly novelty.
And it reported that roughly 40% of workers with access used it regularly. That single number should reframe how you read your own pilot. In a controlled study, a majority of people with access were not regular users. A pilot staffed entirely by enthusiasts tells you nothing whatsoever about your median employee, and the median employee is who you are about to buy licences for.
Good evaluations publish their own caveats
The Australian government’s whole-of-government trial is the other useful reference, and it is useful partly because it is honest about what self-report is worth. It ran January to June 2024 across more than 5,765 licences, evaluated by Nous Group for the Digital Transformation Agency with the Australian Centre of Evaluation.
Its headline numbers look like every pilot report you have read: 69% said it improved task completion speed, 61% said it improved work quality, 86% wanted to continue. Then the evaluators do something internal pilots almost never do. They state that the figures are “approximations and likely quote the upper bound”, that “editing was almost always needed”, that up to 7% said Copilot added time because of the verification required, and that quality gains were “more subdued than improvements to work efficiency”. They also report that only one in three participants used it daily.
Both things are true at once. Most people liked it, and a minority were slower. Neither cancels the other, and a pilot report that only carries the first half is not wrong so much as incomplete in a direction that happens to suit the person who commissioned it.
Worth adding: that trial is 2024 data. It is two product generations old now. It is a good model for how to evaluate, and a poor source for what to expect.
Hold a group back, and write the question down first
You do not need 56 firms. You need a smaller version of the same logic, and it is not expensive.
- Write the decision, not the objective. “We will extend to all of finance, or we will not, and this is what would make us not.” One sentence, agreed before anyone gets a licence.
- Hold a matched group back. Same roles, same workload, no licences for the duration. This is the single change that separates a pilot from a testimonial, and the only real cost is telling some people to wait.
- Assign, do not recruit. Volunteers pre-select for the result you want. If you must take volunteers, randomise which of them gets access first and use the rest as the comparison.
- Run past the novelty. Weeks four to twelve are where behaviour settles. Whatever budget you have, spend it on duration rather than on headcount.
- Count who stopped. Regular use is the finding, and it is the one that gets left out. Report the share of licensed people who used it in the last week, not the share who ever used it.
- Measure one task properly rather than everything badly. Pick work with an artefact at the end. Time it, sample the quality, include the editing time in the total.
The last of those is where most pilots quietly lose their honesty. If drafting gets quicker but the checking absorbs most of what was saved, the number in the survey is not the number in the business. The Australian evaluators found that effect and published it. Your pilot will find it too, if you let it.
What makes this hard to sell internally is that a well-designed pilot can come back and tell you to buy fewer licences than you planned. That is the pilot working. It is also the reason so few of them are designed that way.