What observability means for AI
A traditional application is observable when you can answer "is it up, is it fast, is it erroring?" from its telemetry. An AI system needs all of that — and one more dimension: is it still good? A model can be fully available, responding in 200 milliseconds, throwing zero exceptions, and quietly producing worse, less grounded, or unsafe answers than it did last month. Conventional monitoring will report a green dashboard the whole time.
Observability for AI therefore spans two worlds: the operational world of latency, errors, and throughput, and the quality world of groundedness, relevance, safety, and drift. This unit shows you how to instrument both, using Azure Monitor for the operational layer and Azure AI Foundry evaluation for the quality layer.
By the end of this unit- Distinguish operational telemetry from quality telemetry and explain why AI needs both.
- Identify the four signal layers to instrument and the tools that capture each.
- Set performance baselines and alerting thresholds that detect drift and quality regression in production.
The green-dashboard trap
- Latency, throughput, error rate, token consumption
- Captured by Azure Monitor and Application Insights
- Answers "is the service healthy and affordable?"
- Necessary but not sufficient for an AI system
- Groundedness, relevance, coherence, content safety
- Captured by continuous evaluation in Azure AI Foundry
- Answers "is the system still producing good output?"
- Invisible to operational monitoring alone
- A system can be operationally perfect and qualitatively degrading
- Quality regressions rarely throw exceptions — they just get worse
- Cost can spike without errors as prompts or context grow
- Only the two layers together give a true picture of health
The four signal layers
A production AI system should be instrumented across four layers. Each answers a different question, and each has a natural home in the Azure tooling. Together they form the observability surface you monitor day to day.
1. Infrastructure and service health
Availability, latency, error rates, and resource utilisation for the endpoints serving your model. Captured in Azure Monitor and Application Insights. This is the layer most teams already have — it tells you whether the service is up and responsive.
2. Token and cost telemetry
Prompt tokens, completion tokens, and total consumption per request and in aggregate. Because cost scales with tokens, this layer is both an operational and a financial signal. A rising token trend with flat request volume is an early warning of prompt bloat or context growth.
3. Output quality
Groundedness (is the answer supported by the retrieved context?), relevance, coherence, and fluency. Azure AI Foundry's evaluation capability can run these as continuous evaluations against sampled production traffic, scoring live behaviour rather than only a static test set.
4. Safety and risk
Detection of harmful content, jailbreak attempts, and policy violations in both inputs and outputs. Content-safety signals and risk evaluations surface here. This layer feeds directly into your incident-response process when thresholds are breached.
Sample, do not score everything. Running a full quality evaluation on every production request is expensive and usually unnecessary. A representative sample — say a fixed percentage of traffic — gives you a reliable trend at a fraction of the cost. Increase the sampling rate when an alert fires.
Request volume is flat but average tokens per request have climbed 40% over two weeks. Which layer surfaced this, and what does it most likely indicate?
Drift and groundedness monitoring
Drift is the gradual divergence of a system's behaviour from the behaviour you validated at release. For AI systems it comes from three directions: the inputs change (users ask different things), the grounding data changes (your knowledge sources are updated), or the underlying model changes (a provider updates a version). All three degrade quality without touching availability.
Forms of drift to monitor
Input drift
Groundedness drift
Model drift
Concept drift
For a live AI system you know, which of the four drift types is most likely to bite first? Knowing your dominant drift risk tells you which signal deserves the tightest monitoring and the most frequent re-evaluation.
Baselines, alerting, and response
Monitoring is only useful if it triggers action at the right moment. That requires a baseline — the established normal behaviour of the system — and alert thresholds set as meaningful deviations from it. Without a baseline, every number is just a number; with one, you can tell the difference between noise and a real regression.
Establish the baseline at release
At the point of production sign-off, record the system's behaviour across all four signal layers: typical latency, token consumption, and the distribution of quality scores. This is your reference. Capture it from real or representative traffic, not a single clean test, so it reflects genuine production conditions.
Set thresholds as deviations, not absolutes
Alert on a meaningful drop from baseline — for example, mean groundedness falling below a set floor, or sustained for a number of evaluation windows. Combine a hard floor (never acceptable) with a relative trigger (a significant decline from normal) so you catch both sudden breaks and slow erosion.
Route alerts to the right response
Operational alerts (latency, errors) route to the platform on-call. Quality and safety alerts route to the AI system owner and feed the incident-response process. A groundedness regression is not a server problem — sending it to the infrastructure team wastes time and loses the signal.
Close the loop with re-evaluation
When an alert fires, raise the evaluation sampling rate, investigate, and after remediation re-run the full evaluation suite to confirm the system is back within bounds. Record the new baseline if the system has legitimately changed. Monitoring feeds the quality cycle; it does not sit beside it.
Hands-on with the evaluation and monitoring tooling described here:
End of Unit 17
You should now be able to:
- Instrument operational and quality telemetry across four signal layers.
- Detect input, groundedness, model, and concept drift before it reaches users at scale.
- Set baselines and deviation-based alerts that route to the correct responder and close the loop.
Unit review
Why is operational monitoring alone insufficient for an AI system?
Which signal is the earliest, most reliable indicator of quality decay in a retrieval-grounded system?
Why should quality and safety alerts route to a different responder than operational alerts?
Why combine a hard floor with a relative deviation trigger when setting alert thresholds?
End of module
You have completed Course 17: Monitoring AI in Production. Next: Incident Response for AI Systems — runbooks, severity, and playbooks for when monitoring fires.