Why AI testing is different
Traditional software testing assumes a deterministic system: the same input produces the same output, and a test either passes or fails. Generative AI breaks that assumption. The same prompt can produce different wording each time, "correct" is often a matter of degree, and a change to the prompt or model can silently shift quality across thousands of cases. Testing AI outputs therefore needs its own discipline.
The goal is not to prove the AI is perfect — it will not be. The goal is to measure quality systematically, catch regressions before users do, and hold the line at a defined standard before content reaches anyone.
By the end of this unit- Explain why probabilistic AI outputs need a different testing approach from deterministic software.
- Build an evaluation set and apply groundedness and relevance metrics to score outputs.
- Use red-teaming and quality gates to catch failures before AI content reaches users.
Three testing mindsets to combine
- Re-run a fixed set of cases after any change to prompt, model, or data
- Detects quality drops introduced by a change you thought was safe
- Scored automatically against expected answers or graded criteria
- The safety net that lets you change the system with confidence
- Measure output quality against defined metrics on a representative set
- Groundedness, relevance, coherence, fluency, and task-specific scores
- Gives a quality number you can track over time
- Run by humans, by another model as a judge, or both
- Deliberately try to make the system fail or misbehave
- Adversarial prompts, edge cases, prompt injection, harmful requests
- Finds the failures that a friendly test set never will
- Essential before any externally facing deployment
Evaluation sets and metrics
An evaluation set is a curated collection of representative inputs — with expected answers or grading criteria — that you run the system against repeatedly. It is the single most useful asset in AI testing, because it turns "it seems fine" into a number you can track.
1. Build a representative set
Collect inputs that reflect real usage: common questions, important edge cases, and known-hard examples. Aim for coverage of the cases that matter, not a huge volume of easy ones. Capture the expected answer or the criteria a good answer must meet.
2. Choose metrics that match the task
For retrieval-grounded answers, groundedness (is the answer supported by the retrieved sources?) and relevance (does it address the question?) are the core pair. Coherence and fluency cover readability. For tasks with a right answer, exact-match or similarity scoring applies.
3. Score consistently
Use human raters with a clear rubric, an LLM-as-judge for scale, or both — humans to calibrate, the model to run volume. Whatever you choose, apply it the same way every time so scores are comparable across runs.
4. Track over time
Record the scores for every run. A single number means little; the trend means everything. A sudden drop after a prompt tweak is exactly the signal regression testing exists to surface.
Groundedness asks whether the answer is actually supported by the source material — the defence against confident fabrication. Relevance asks whether the answer addresses what was asked. An answer can be perfectly grounded yet irrelevant, or relevant yet ungrounded. Strong systems score well on both.
An answer is fluent, on-topic, and addresses the question well — but states a figure that does not appear in any retrieved source. Which metric should flag this?
Red-teaming and quality gates
Evaluation sets measure how good the system is on expected inputs. Red-teaming measures how badly it can be made to behave on hostile ones — and a quality gate is what stops a system that fails either test from reaching users.
Red-teaming
Red-teaming means deliberately attacking your own system: adversarial and ambiguous prompts, prompt-injection attempts hidden in retrieved content, requests for harmful or out-of-scope output, and edge cases the happy-path set ignores. The aim is to find the failure modes before a real user — or a bad actor — does. Treat the findings as test cases and fold them into your regression set.
For an AI feature you are deploying, what is the worst plausible output it could produce? Have you ever deliberately tried to make it produce that, or have you only ever tested it kindly? The gap between those two is your red-teaming backlog.
Quality gates
A quality gate is a defined threshold the system must clear before content is released — a minimum groundedness score, a maximum failure rate on the red-team set, a human sign-off for high-stakes outputs. The gate turns evaluation from an interesting measurement into an enforceable standard.
Gate on the metrics that matter for your risk
Automate the gate into the release process
Keep a human gate for high-stakes content
Build the practical skills with these resources:
Evaluate and monitor AI models in Azure AI Foundry ↗
Implement Retrieval Augmented Generation with Azure AI ↗
End of Unit 15
You should now be able to:
- Explain why probabilistic outputs need regression, evaluation, and red-teaming together.
- Build an evaluation set and apply groundedness and relevance metrics.
- Set and automate quality gates appropriate to the risk of the deployment.
Unit review
Why can a single pass/fail test not adequately validate a generative AI feature?
What is the primary purpose of a regression test set for an AI system?
When is an automated metric gate not sufficient on its own?
End of module
You have completed Course 15: Testing AI Outputs — and the Deploy pillar series on architecture, integration, and operations. Apply these patterns to make your AI deployments reliable, governed, and trustworthy.