What production hardening means
A working AI feature and a production AI service are different things. The feature answers correctly when you try it. The service keeps answering correctly when the region it runs in degrades, when a model upgrade subtly changes behaviour, when a user submits hostile input, and when the monthly bill needs explaining. Production hardening is the discipline of closing those gaps before they become incidents.
This unit brings the operational practices together: monitoring and ongoing evaluation, regional resilience and failover, safe deployment of model changes, content safety, and cost governance at scale. These are the practices that turn a deployment into a service you can stand behind.
By the end of this unit- Establish monitoring and continuous evaluation for an Azure AI workload.
- Apply failover, regional resilience, and safe deployment techniques to model changes.
- Implement content safety and cost governance controls at scale.
The four pillars of production hardening
- Telemetry on latency, throughput, errors, and token usage
- Continuous evaluation of output quality, not just system health
- Alerting on the metrics that predict user-visible problems
- Retries, circuit breakers, and graceful fallbacks
- Multi-region deployment for workloads that cannot tolerate an outage
- A tested failover path, not an assumed one
- Validate model upgrades against an evaluation set before switching
- Blue-green or canary rollout so a bad change is caught on a slice of traffic
- A fast, rehearsed rollback
- Content safety on inputs and outputs
- Cost attribution, budgets, and alerts
- Responsible-AI controls applied consistently across the estate
Monitoring and evaluation
Traditional monitoring tells you whether the system is up. AI workloads need a second layer: whether the answers are still good. A model that is healthy by every infrastructure metric can quietly degrade in quality after a data change, a prompt edit, or a model-version update. You need both layers.
1. System telemetry
Capture latency (including tail latency), throughput, error and throttling (429) rates, and token consumption per deployment. Send it to Azure Monitor and Application Insights so you can correlate AI metrics with the rest of the application. Alert on the leading indicators — rising 429s or climbing tail latency — before users feel them.
2. Continuous quality evaluation
Run an evaluation set on a schedule and after every change, measuring groundedness, relevance, accuracy, and safety. Azure AI Foundry's evaluation tooling supports automated, repeatable evaluation. This catches quality regressions that infrastructure metrics never will — the answers that are slower to be wrong, not slower to arrive.
3. Capture and review real interactions
Within your privacy and governance rules, sample real prompts and responses to spot failure patterns the evaluation set did not anticipate. Feed those back into the evaluation set so it grows toward the questions users actually ask. Monitoring is only useful if it informs the next change.
The distinction that matters: health monitoring answers "is it running?"; evaluation answers "is it still good?". An AI service can pass every health check while its answer quality silently degrades. Run both, and treat a quality regression as seriously as an outage.
Resilience and safe deployment
Two kinds of change threaten a production AI service: the infrastructure failing beneath it, and the change you deliberately make to it. Resilience handles the first; safe deployment handles the second.
Withstanding failure and changing safely
Regional resilience and failover
Blue-green and canary deployment
Validate before you switch
Rehearse the rollback
For a production AI workload you know, has the failover ever actually been triggered as a test — or is it assumed to work? And if a model upgrade went wrong, how many users would be affected before you noticed, and how fast could you roll back? If those answers are uncomfortable, that is where the next piece of hardening goes.
Content safety and cost governance
The final pillar is governance: keeping the service safe to use and its cost under control as it scales across the organisation. Both must be applied consistently — a control that exists in one project but not another is a gap.
1. Content safety on inputs and outputs
Use Azure AI Content Safety to detect and filter harmful content — hate, violence, self-harm, sexual content — on both the user's input and the model's output, and to defend against prompt-injection (jailbreak) attempts. Configure severity thresholds to your organisation's policy and log what is filtered so you can tune over time.
2. Responsible-AI controls, applied consistently
Grounding instructions, citation pass-through, and refusal-when-unsure are not per-project niceties — they are baseline controls. In a hub-and-spoke architecture, apply them at the hub so every project inherits them. In a federated model, bake them into the shared template. Consistency is the point.
3. Cost governance at scale
Tag deployments and projects so spend is attributable per use case, set budgets and alerts so spikes are caught early, and review consumption regularly. Apply the cost levers — right-sized models, trimmed tokens, caching, the right billing model — as standing practice. Governance is what keeps a successful, growing AI estate financially explainable.
Deepen your production and responsible-AI practice:
Evaluate and monitor AI models in Azure AI Foundry ↗
Responsible AI principles in practice ↗
Build AI apps and agents with Azure AI Foundry ↗
End of Unit 5
You should now be able to:
- Run both health monitoring and continuous quality evaluation for an AI workload.
- Apply regional failover and blue-green or canary deployment to changes.
- Implement content safety and consistent cost governance at scale.
Unit review
An AI service passes every infrastructure health check, yet users report worse answers after a change. What was missing?
You want to upgrade a model version with minimal risk to users. Which approach fits?
Why must a failover path be tested rather than assumed?
In a hub-and-spoke architecture, where should baseline responsible-AI and content-safety controls be applied?
End of module
You have completed Course 05: Azure AI in Production — and the Deploy pillar's architecture-to-operations track. You can now design, scale, build, and operate Microsoft AI workloads with production discipline.