Section 01 · Unit introduction

What observability means for AI

A traditional application is observable when you can answer "is it up, is it fast, is it erroring?" from its telemetry. An AI system needs all of that — and one more dimension: is it still good? A model can be fully available, responding in 200 milliseconds, throwing zero exceptions, and quietly producing worse, less grounded, or unsafe answers than it did last month. Conventional monitoring will report a green dashboard the whole time.

Observability for AI therefore spans two worlds: the operational world of latency, errors, and throughput, and the quality world of groundedness, relevance, safety, and drift. This unit shows you how to instrument both, using Azure Monitor for the operational layer and Azure AI Foundry evaluation for the quality layer.

By the end of this unit
  • Distinguish operational telemetry from quality telemetry and explain why AI needs both.
  • Identify the four signal layers to instrument and the tools that capture each.
  • Set performance baselines and alerting thresholds that detect drift and quality regression in production.

The green-dashboard trap

  • Latency, throughput, error rate, token consumption
  • Captured by Azure Monitor and Application Insights
  • Answers "is the service healthy and affordable?"
  • Necessary but not sufficient for an AI system
  • Groundedness, relevance, coherence, content safety
  • Captured by continuous evaluation in Azure AI Foundry
  • Answers "is the system still producing good output?"
  • Invisible to operational monitoring alone
  • A system can be operationally perfect and qualitatively degrading
  • Quality regressions rarely throw exceptions — they just get worse
  • Cost can spike without errors as prompts or context grow
  • Only the two layers together give a true picture of health
The most dangerous AI failure is the one that does not error. It just gets quietly worse while the dashboard stays green.
Working principle · AI observability
Section 02

The four signal layers

A production AI system should be instrumented across four layers. Each answers a different question, and each has a natural home in the Azure tooling. Together they form the observability surface you monitor day to day.

1. Infrastructure and service health

Availability, latency, error rates, and resource utilisation for the endpoints serving your model. Captured in Azure Monitor and Application Insights. This is the layer most teams already have — it tells you whether the service is up and responsive.

2. Token and cost telemetry

Prompt tokens, completion tokens, and total consumption per request and in aggregate. Because cost scales with tokens, this layer is both an operational and a financial signal. A rising token trend with flat request volume is an early warning of prompt bloat or context growth.

3. Output quality

Groundedness (is the answer supported by the retrieved context?), relevance, coherence, and fluency. Azure AI Foundry's evaluation capability can run these as continuous evaluations against sampled production traffic, scoring live behaviour rather than only a static test set.

4. Safety and risk

Detection of harmful content, jailbreak attempts, and policy violations in both inputs and outputs. Content-safety signals and risk evaluations surface here. This layer feeds directly into your incident-response process when thresholds are breached.

Practitioner note

Sample, do not score everything. Running a full quality evaluation on every production request is expensive and usually unnecessary. A representative sample — say a fixed percentage of traffic — gives you a reliable trend at a fraction of the cost. Increase the sampling rate when an alert fires.

Knowledge check

Request volume is flat but average tokens per request have climbed 40% over two weeks. Which layer surfaced this, and what does it most likely indicate?

Section 03

Drift and groundedness monitoring

Drift is the gradual divergence of a system's behaviour from the behaviour you validated at release. For AI systems it comes from three directions: the inputs change (users ask different things), the grounding data changes (your knowledge sources are updated), or the underlying model changes (a provider updates a version). All three degrade quality without touching availability.

Forms of drift to monitor

Input drift
The distribution of questions users ask shifts away from what you tested. A support assistant validated on billing questions starts receiving technical troubleshooting it was never grounded for. You detect this by tracking the topics and characteristics of incoming requests and comparing against your evaluation dataset's coverage.
Groundedness drift
Groundedness measures whether the model's answer is actually supported by the context it retrieved. It can fall when source documents are updated, removed, or restructured so retrieval returns weaker context. A continuous groundedness evaluation in Azure AI Foundry, run on sampled traffic, gives you a live trend line — and a falling trend is your earliest, most reliable signal of quality decay in a retrieval-grounded system.
Model drift
The model version behind your system is updated by the provider, changing its behaviour subtly even though your prompt is unchanged. Pin model versions where you can, and re-run your evaluation suite whenever a version changes so you can compare the new behaviour against your established baseline before it reaches users.
Concept drift
The world changes so that previously-correct answers become outdated — policy changes, new products, revised pricing. The model is behaving as designed, but its grounding is stale. The control is a re-evaluation cadence tied to how fast your domain changes, plus a freshness check on the grounding data itself.
Reflect

For a live AI system you know, which of the four drift types is most likely to bite first? Knowing your dominant drift risk tells you which signal deserves the tightest monitoring and the most frequent re-evaluation.

Section 04

Baselines, alerting, and response

Monitoring is only useful if it triggers action at the right moment. That requires a baseline — the established normal behaviour of the system — and alert thresholds set as meaningful deviations from it. Without a baseline, every number is just a number; with one, you can tell the difference between noise and a real regression.

Establish the baseline at release

At the point of production sign-off, record the system's behaviour across all four signal layers: typical latency, token consumption, and the distribution of quality scores. This is your reference. Capture it from real or representative traffic, not a single clean test, so it reflects genuine production conditions.

Set thresholds as deviations, not absolutes

Alert on a meaningful drop from baseline — for example, mean groundedness falling below a set floor, or sustained for a number of evaluation windows. Combine a hard floor (never acceptable) with a relative trigger (a significant decline from normal) so you catch both sudden breaks and slow erosion.

Route alerts to the right response

Operational alerts (latency, errors) route to the platform on-call. Quality and safety alerts route to the AI system owner and feed the incident-response process. A groundedness regression is not a server problem — sending it to the infrastructure team wastes time and loses the signal.

Close the loop with re-evaluation

When an alert fires, raise the evaluation sampling rate, investigate, and after remediation re-run the full evaluation suite to confirm the system is back within bounds. Record the new baseline if the system has legitimately changed. Monitoring feeds the quality cycle; it does not sit beside it.

Continue on Microsoft Learn

Hands-on with the evaluation and monitoring tooling described here:

Evaluate and monitor AI models in Azure AI Foundry ↗

Build AI apps and agents with Azure AI Foundry ↗

End of Unit 17

You should now be able to:

  • Instrument operational and quality telemetry across four signal layers.
  • Detect input, groundedness, model, and concept drift before it reaches users at scale.
  • Set baselines and deviation-based alerts that route to the correct responder and close the loop.
Section 05

Unit review

Question 1 of 4

Why is operational monitoring alone insufficient for an AI system?

Question 2 of 4

Which signal is the earliest, most reliable indicator of quality decay in a retrieval-grounded system?

Question 3 of 4

Why should quality and safety alerts route to a different responder than operational alerts?

Question 4 of 4

Why combine a hard floor with a relative deviation trigger when setting alert thresholds?

End of module

You have completed Course 17: Monitoring AI in Production. Next: Incident Response for AI Systems — runbooks, severity, and playbooks for when monitoring fires.

Craig Stanley Studio · Deploy — Architecture, Integration & Operations · Monitoring AI in Production · Access by direct link only.