Section 01 · Unit introduction

What production hardening means

A working AI feature and a production AI service are different things. The feature answers correctly when you try it. The service keeps answering correctly when the region it runs in degrades, when a model upgrade subtly changes behaviour, when a user submits hostile input, and when the monthly bill needs explaining. Production hardening is the discipline of closing those gaps before they become incidents.

This unit brings the operational practices together: monitoring and ongoing evaluation, regional resilience and failover, safe deployment of model changes, content safety, and cost governance at scale. These are the practices that turn a deployment into a service you can stand behind.

By the end of this unit
  • Establish monitoring and continuous evaluation for an Azure AI workload.
  • Apply failover, regional resilience, and safe deployment techniques to model changes.
  • Implement content safety and cost governance controls at scale.

The four pillars of production hardening

  • Telemetry on latency, throughput, errors, and token usage
  • Continuous evaluation of output quality, not just system health
  • Alerting on the metrics that predict user-visible problems
  • Retries, circuit breakers, and graceful fallbacks
  • Multi-region deployment for workloads that cannot tolerate an outage
  • A tested failover path, not an assumed one
  • Validate model upgrades against an evaluation set before switching
  • Blue-green or canary rollout so a bad change is caught on a slice of traffic
  • A fast, rehearsed rollback
  • Content safety on inputs and outputs
  • Cost attribution, budgets, and alerts
  • Responsible-AI controls applied consistently across the estate
Production is not a state you reach once. It is a set of practices you keep running — the moment you stop observing, evaluating, and governing, the service starts drifting away from what you signed off.
Working principle · AI operations
Section 02

Monitoring and evaluation

Traditional monitoring tells you whether the system is up. AI workloads need a second layer: whether the answers are still good. A model that is healthy by every infrastructure metric can quietly degrade in quality after a data change, a prompt edit, or a model-version update. You need both layers.

1. System telemetry

Capture latency (including tail latency), throughput, error and throttling (429) rates, and token consumption per deployment. Send it to Azure Monitor and Application Insights so you can correlate AI metrics with the rest of the application. Alert on the leading indicators — rising 429s or climbing tail latency — before users feel them.

2. Continuous quality evaluation

Run an evaluation set on a schedule and after every change, measuring groundedness, relevance, accuracy, and safety. Azure AI Foundry's evaluation tooling supports automated, repeatable evaluation. This catches quality regressions that infrastructure metrics never will — the answers that are slower to be wrong, not slower to arrive.

3. Capture and review real interactions

Within your privacy and governance rules, sample real prompts and responses to spot failure patterns the evaluation set did not anticipate. Feed those back into the evaluation set so it grows toward the questions users actually ask. Monitoring is only useful if it informs the next change.

Design note

The distinction that matters: health monitoring answers "is it running?"; evaluation answers "is it still good?". An AI service can pass every health check while its answer quality silently degrades. Run both, and treat a quality regression as seriously as an outage.

Section 03

Resilience and safe deployment

Two kinds of change threaten a production AI service: the infrastructure failing beneath it, and the change you deliberately make to it. Resilience handles the first; safe deployment handles the second.

Withstanding failure and changing safely

Regional resilience and failover
For workloads that cannot tolerate a regional outage, deploy the same model in a second region and route between them, with health checks that detect a failing endpoint and shift traffic. Crucially, test the failover — an untested failover path is a hope, not a control. Match the investment to the requirement: not every workload needs multi-region, but the ones that do must prove it works.
Blue-green and canary deployment
When upgrading a model version or changing a prompt, do not switch all traffic at once. Blue-green keeps the current version live while you validate the new one, then switches with a fast rollback available. Canary sends a small slice of traffic to the new version, watches the quality and error metrics, and ramps up only if they hold. Both let you catch a bad change on a fraction of users instead of all of them.
Validate before you switch
Model versions can change behaviour in ways that are subtle and only visible at scale. Run your evaluation set against the new version before it takes production traffic. Pointing application code at a stable deployment name lets you swap the underlying version behind it once validation passes — and roll back the same way if the canary metrics turn.
Rehearse the rollback
The rollback path matters more than the deployment path, because you reach for it under pressure. Make it fast and rehearsed: a known-good deployment to revert to, a single clear action to switch back, and monitoring to confirm recovery. A rollback you have never practised will be slow exactly when speed matters most.
Reflect

For a production AI workload you know, has the failover ever actually been triggered as a test — or is it assumed to work? And if a model upgrade went wrong, how many users would be affected before you noticed, and how fast could you roll back? If those answers are uncomfortable, that is where the next piece of hardening goes.

Section 04

Content safety and cost governance

The final pillar is governance: keeping the service safe to use and its cost under control as it scales across the organisation. Both must be applied consistently — a control that exists in one project but not another is a gap.

1. Content safety on inputs and outputs

Use Azure AI Content Safety to detect and filter harmful content — hate, violence, self-harm, sexual content — on both the user's input and the model's output, and to defend against prompt-injection (jailbreak) attempts. Configure severity thresholds to your organisation's policy and log what is filtered so you can tune over time.

2. Responsible-AI controls, applied consistently

Grounding instructions, citation pass-through, and refusal-when-unsure are not per-project niceties — they are baseline controls. In a hub-and-spoke architecture, apply them at the hub so every project inherits them. In a federated model, bake them into the shared template. Consistency is the point.

3. Cost governance at scale

Tag deployments and projects so spend is attributable per use case, set budgets and alerts so spikes are caught early, and review consumption regularly. Apply the cost levers — right-sized models, trimmed tokens, caching, the right billing model — as standing practice. Governance is what keeps a successful, growing AI estate financially explainable.

End of Unit 5

You should now be able to:

  • Run both health monitoring and continuous quality evaluation for an AI workload.
  • Apply regional failover and blue-green or canary deployment to changes.
  • Implement content safety and consistent cost governance at scale.
Section 05

Unit review

Question 1 of 4

An AI service passes every infrastructure health check, yet users report worse answers after a change. What was missing?

Question 2 of 4

You want to upgrade a model version with minimal risk to users. Which approach fits?

Question 3 of 4

Why must a failover path be tested rather than assumed?

Question 4 of 4

In a hub-and-spoke architecture, where should baseline responsible-AI and content-safety controls be applied?

End of module

You have completed Course 05: Azure AI in Production — and the Deploy pillar's architecture-to-operations track. You can now design, scale, build, and operate Microsoft AI workloads with production discipline.

Craig Stanley Studio · Deploy — Architecture, Integration & Operations · Azure AI in Production · Access by direct link only.