Section 01 · Unit introduction

The pilot-to-production gap

A successful AI pilot proves one thing: that the idea can work, once, for a small, motivated, supervised group. Production proves something much harder: that it works reliably, for everyone, every day, when the original enthusiasts have moved on and nobody is watching closely. The distance between those two is where most AI initiatives stall — not because the technology failed, but because the operating model never made the jump.

This unit names what changes between pilot and production, what predictably breaks in the transition, and how to build a programme — with governance, application lifecycle management, capacity planning, and change management — that scales rather than merely launches.

By the end of this unit
  • Explain why a successful pilot does not guarantee a successful production rollout.
  • Anticipate what breaks at scale across data, cost, governance, and adoption.
  • Describe a production programme with ALM, capacity planning, and change management built in.

Pilot conditions versus production conditions

  • Small group of self-selected enthusiasts
  • Hand-held by the project team
  • Narrow, curated set of use cases
  • Cost is negligible at low volume
  • Edge cases are rare and forgiven
  • Whole population, including reluctant and untrained users
  • Self-service, with no one watching each interaction
  • The full, messy range of real-world tasks
  • Cost scales with every user and every token
  • Edge cases arrive constantly and must be handled
  • You must design for the unmotivated user, not the champion
  • Support, monitoring, and guardrails replace the hovering project team
  • Governance and lifecycle management become non-negotiable
  • Cost control moves from afterthought to design constraint
  • The operating model, not the demo, decides success
A pilot proves it can work. Production proves it keeps working without you in the room.
Working principle · Scaling AI
Section 02

What breaks at scale

The failures of scaling are predictable, which means they are preventable. Four areas break most reliably as a pilot moves to production. Recognising them in advance lets you design controls before they bite, rather than firefighting after launch.

1. Data and grounding

A pilot is grounded on a tidy, curated knowledge set. At scale, the system meets the full corpus — out of date, contradictory, inconsistently permissioned, and constantly changing. Groundedness falls and the answers that impressed in the demo become unreliable. The control is a managed grounding pipeline with freshness, permissions, and quality treated as production data concerns.

2. Cost

Negligible pilot cost becomes a material line item once every user is consuming tokens daily. Without cost controls, spend can outrun value before anyone notices. The control is token telemetry, budgeting per workload, and the cost-optimisation practices that the next course covers in depth — designed in, not bolted on.

3. Governance and accountability

In a pilot the project team is the governance. At scale you need named owners, a risk register, sign-off gates, monitoring, and an incident process — for potentially many AI workloads at once. Without this, you accumulate a sprawl of ungoverned AI that nobody can audit or safely operate. The control is an AI management discipline applied consistently across the portfolio.

4. Adoption and trust

The pilot's enthusiasts wanted it to work. The production population includes sceptics, the overloaded, and the actively resistant. A poor first experience at scale poisons adoption permanently. The control is structured change management: training, calibrated-trust onboarding, visible quick wins, and a feedback channel that is seen to act.

Practitioner note

The most common scaling mistake is treating the pilot's success as evidence the hard problems are solved. The pilot solved the "can it work?" question. It told you almost nothing about data quality at scale, real cost, governance load, or adoption among non-enthusiasts — which are the things that actually determine whether production succeeds.

Knowledge check

A pilot grounded on a curated 50-document set scored highly on groundedness. At full rollout against the organisation's entire SharePoint estate, groundedness drops sharply. What is the most likely cause?

Section 03

Governance and ALM

Application Lifecycle Management (ALM) is the discipline of moving a solution through defined environments — develop, test, and production — with controlled promotion between them. For AI solutions built on platforms such as Copilot Studio and Power Platform, ALM is what stops production becoming a place where people edit live systems by hand and hope.

The pillars of AI ALM

Separate environments
Maintain distinct development, test, and production environments. Build and experiment in development, validate in test against representative data, and promote only approved, evaluated solutions to production. In Power Platform this is managed through environments and solutions; the principle holds across any AI platform. The pilot habit of building directly in the live environment does not survive scale.
Controlled promotion with quality gates
Promotion between environments passes through the quality gates from Course 16: evaluation results meet thresholds, risk assessment is complete, an owner has signed off. Promotion is a deliberate, evidenced decision — not a copy-paste. This is where ALM and your quality framework meet: the gate is the control that makes promotion safe.
Source control and versioning
Solutions, prompts, and configurations are versioned so you know exactly what is running, can compare versions, and can roll back. When a model version or prompt changes, versioning is what lets you re-evaluate against a known baseline and revert if behaviour regresses. Unversioned production AI cannot be safely changed or audited.
Portfolio governance
At scale you are not governing one AI solution but a growing portfolio. A central register of AI workloads — purpose, owner, risk level, environment, last evaluation — gives leadership a view of the whole estate and stops ungoverned sprawl. Each entry carries the named accountable owner who answers for that system's behaviour.
Reflect

Does your organisation build AI solutions directly in the environment users rely on? If development, test, and production are not separated, that is the first structural change scaling will demand of you.

Section 04

Building a scalable programme

Scaling is not a bigger launch — it is a different operating model. A scalable AI programme has four standing capabilities that run continuously, not a one-time project plan. Build these and each new use case rides on shared rails rather than reinventing the route.

Capacity and cost management

Forecast demand and match it to capacity. For predictable, high-volume workloads, provisioned throughput or reserved capacity gives stable performance and cost; for variable workloads, consumption pricing flexes with demand. Maintain a budget per workload and token telemetry so spend is visible and attributable, and so a runaway cost is caught early.

Governance as a standing function

The quality gates, risk register, monitoring, and incident process from earlier courses run as an ongoing function with named owners — not a checklist completed at launch and forgotten. New use cases enter the same governed pipeline. This is what keeps a growing portfolio auditable and safe.

Change management as a workstream

Treat adoption as a deliberate workstream, not a hope. Train the non-enthusiast population, onboard for calibrated trust, surface visible quick wins, and run a feedback channel that is seen to act. Adoption at scale is earned through experience, and a poor first impression is expensive to reverse.

A repeatable delivery pattern

Codify how a new AI use case goes from idea to production: the ALM path, the gates it must pass, the monitoring it inherits, and the support model it joins. A repeatable pattern means the second, fifth, and twentieth use case are faster and safer than the first — that is what scaling actually means.

End of Unit 19

You should now be able to:

  • Explain why pilot success does not predict production success.
  • Anticipate and pre-empt the data, cost, governance, and adoption failures of scaling.
  • Stand up a programme with ALM, capacity planning, standing governance, and change management.
Section 05

Unit review

Question 1 of 4

Why does a successful pilot tell you little about whether production will succeed?

Question 2 of 4

Which is the most reliable cause of groundedness falling when a pilot scales to the full knowledge estate?

Question 3 of 4

What is the role of quality gates in AI application lifecycle management?

Question 4 of 4

For a predictable, high-volume production AI workload, which capacity approach gives the most stable performance and cost?

End of module

You have completed Course 19: Scaling AI Pilots to Production. Next: Cost Optimisation for AI Workloads — token economics, efficiency, and capacity choices that control spend.

Craig Stanley Studio · Deploy — Architecture, Integration & Operations · Scaling AI Pilots to Production · Access by direct link only.