Licence assignment is the starting gun. In most programmes it is treated as the finish line, and that single error explains more disappointing Copilot rollouts than any technical problem I have worked on.
The pattern is consistent enough to be boring. A business case gets approved. Procurement lands the licences. IT assigns them, often to whoever asked loudest. A launch email goes out with a link to a training portal and a recording of a demo. The programme then closes, the project manager rolls off, and some months later a finance director asks why the usage report looks the way it does.
Everything that decides the outcome happens after the licence lands. Almost none of it is technical.
The measurable difference sits with the organisation, not the user
Microsoft’s 2026 Work Trend Index reports that organisational factors account for more than twice the AI impact of individual factors. It also finds that only 19% of AI users are what it calls “in the Frontier”, meaning organisational capability and individual readiness reinforcing each other.
Treat the source honestly. That is a Microsoft first-party survey of 20,000 knowledge workers across 10 markets, with fieldwork run by Edelman Data x Intelligence between 18 February and 7 April 2026, alongside anonymised telemetry. It measures what people say, not what they produce. Microsoft has an obvious interest in the finding.
I still think it is the most useful number published this year, because it points away from the thing most programmes spend their money on. If organisational factors carry twice the weight, then the training video, the prompt library and the launch communications are the small half of the problem. The large half is how the organisation is set up: who owns the tool, what data it can see, which processes were redesigned to use it, and whether anyone is allowed to change how work is done as a result.
That is not a communications exercise. It is an operating model question, and it needs an owner with the authority to change a process.
Usage is where it fails first
The strongest evidence available on Copilot’s effect is Microsoft’s own randomised controlled trial, “Early Impacts of M365 Copilot”, published April 2025 across 6,000+ workers in 56 firms over six months. It found around 30 minutes less time reading email per week and documents completed 12% faster.
What I care about in that study is the third number. Roughly 40% of workers with access used it regularly. In a randomised trial, run by the vendor, with the vendor’s attention on it, three in five people did not make it a habit.
The Australian government’s whole-of-government trial, which ran January to June 2024 across 5,765+ licences and was evaluated by Nous Group for the Digital Transformation Agency, found the same shape. At the end of the trial, 77% were optimistic and 86% wanted to continue. Only one in three used it daily.
Hold those two sets of figures next to each other. People like it. Most of them do not use it. Enthusiasm and habit are different measurements, and a launch programme that optimises for the first will report success while producing the second.
The Australian data also carries a detail that gets skipped in the summaries. In the same trial, 69% said Copilot improved the speed of task completion and 61% said it improved the quality of their work, while 65% of managers saw positive effects on their teams. Managers were more convinced than the quality numbers justify. That gap is worth watching in your own organisation, because managers are usually the ones asked whether the rollout worked.
My working explanation for the habit problem is unglamorous. People stop using a tool when it fails them on the work that matters and there is no route to say so. The first bad answer on something important is the moment the rollout is decided, and in most organisations nobody is watching that moment or has any way of hearing about it.
The operating model has to contain six things
This is the list I work through with a client before anyone talks about a pilot. It is deliberately short, because a long list gets delegated and a short one gets argued about.
- A named owner with process authority. Not a sponsor and not a steering group. One person whose job description includes changing how a piece of work gets done, and whose budget survives the rollout.
- A decision about what Copilot can see. Grounding is the single largest determinant of answer quality, and it is a permissions and content problem that predates the product. Someone has to decide which sites are in scope and accept the consequences of that.
- A definition of the outcome you will report. Written before the pilot starts, in the language the finance director used when they asked. Hours are not an outcome. A backlog that clears faster is.
- A route for the questions. People will ask why an answer was wrong. If the only route is a service desk ticket that gets closed as “working as designed”, they stop asking and stop using it.
- A function that watches the release notes. Someone reads what changed this month and decides whether the guidance, the policy or the licence position needs to move.
- A policy short enough to read. Two pages. If it needs a summary, it is not a policy, it is a compliance artefact.
Only the second one is technical. The rest are organisational, which is exactly what the Work Trend Index finding predicts.
I used to put a seventh item on that list: a champions network. I have taken it off. Not because champions are useless, but because a champions network with none of the six above in place is a way of moving the problem to volunteers who cannot fix it. Give them a remit that depends on the owner, the content decision and the question route existing first, and they are worth a great deal. Give them a badge and a monthly call, and they go quiet by about month three.
The product does not hold still, and your plan is already out of date
The classic change-management model assumes a go-live. Copilot has a release cadence instead, and the last twelve months should have made that obvious to anyone still running a launch plan.
Restricted SharePoint Search is retiring, and new enablement has been blocked since 31 July 2026, with customers directed to Restricted Content Discovery instead. If your rollout plan names Restricted SharePoint Search as a control, the plan is now wrong, and it was written by someone competent less than two years ago.
Microsoft 365 E7 became generally available on 1 May 2026, bundling E5, Copilot, Entra Suite and Agent 365 into one SKU. Agent 365 reached general availability for Commercial on the same date. Copilot Cowork is generally available and does agentic multi-step work across Outlook, Teams, OneDrive and SharePoint with per-action approval. Copilot Studio now exposes three separate harnesses with different billing behaviour. And Copilot chat in model-driven apps is deprecated from January 2026 in non-Dynamics environments.
None of that was in a rollout plan written in 2024. All of it changes what a competent answer looks like. A programme that closes at launch has no mechanism to notice, which is why so many organisations are running guidance that quietly stopped being true.
Measure the thing you were asked about
The measurement surface is not the problem. Copilot Analytics, the admin centre readiness and usage reports, Purview audit logs and Copilot Studio analytics will tell you plenty. They will tell you about activity. Nobody approved a business case on activity.
Self-reported time saving is the usual substitute, and it needs handling with care. The Australian evaluation’s own productivity chapter reports daily savings of 1.1 hours on summarising, 1.0 on first drafts, 0.9 on meeting minutes and 0.8 on information search. The same report states those figures are approximations and likely quote the upper bound, that editing was almost always needed, and that up to 7% of participants said Copilot added time because of the verification it required. Gains in work quality were more subdued than gains in efficiency. It is also 2024 data, which in this product means two generations old.
I have made the mistake myself of putting a self-reported time-saving figure into a benefits case, then having to defend it in a room that had every right to be sceptical. It did not survive the first question, which was whether the saved hour showed up anywhere. It did not. The number was real as a perception and worthless as a benefit.
If the measure you report cannot be traced to something a person outside the programme already counts, it is not a benefit, it is a testimonial. Cycle time on a queue. Time to first draft on a document type that has a deadline. Complaint volumes. Whatever the organisation was already measuring before Copilot arrived.
What I would do in the first ninety days
Assume the licences are already assigned, because they usually are by the time I am called.
Pick one process, not one department. Departments contain too many different kinds of work to tell you anything. A process has a start, an end, a queue and usually a number attached to it already. Take that number as the baseline before anyone touches the tool.
Fix grounding for that process only. Not the tenant. The sites and libraries that process touches, the owners of the content, the permissions that are wrong. This is the unglamorous work, and it is the work that decides whether the answers are any good.
Name the owner and give them authority to change the process. Without that, the best possible outcome is people doing the old process slightly faster, which is worth roughly what it sounds like.
Put a standing half-day in the calendar every month for someone to read what changed in the product and decide whether anything you have written down is now false. Treat that as permanent overhead, not project cost.
Then report against the number you took at the start, and say plainly if it did not move. A programme that can report a negative result is the only kind that can be believed when it reports a positive one.