Section 01 · Unit introduction

Why AI needs its own quality discipline

Traditional software quality assurance asks a deterministic question: given this input, does the system produce the correct output, every time? AI systems built on large language models break that assumption. The same prompt can produce different responses. Quality is no longer pass/fail — it is a distribution of behaviours that must be measured, bounded, and governed.

A quality framework for AI is the set of standards, controls, and evidence you put in place so that an AI system behaves acceptably across that distribution — and so that you can demonstrate it does, to auditors, regulators, and your own leadership. This unit gives you the standards landscape and a practical model for gates and sign-off.

By the end of this unit
  • Explain why probabilistic AI systems require a quality discipline distinct from deterministic software QA.
  • Describe ISO/IEC 42001 and Microsoft's Responsible AI principles, and how they relate to one another.
  • Design a quality-gate model with defined sign-off criteria for moving an AI system toward production.

Three properties that change the quality question

  • The same input may yield different outputs across runs
  • You cannot certify correctness by a single test pass
  • Quality must be expressed statistically — pass rates, score distributions, thresholds
  • Evaluation has to be repeated, not run once
  • Capabilities and failure modes appear that were not explicitly designed
  • Edge cases cannot all be enumerated in advance
  • Red-teaming and adversarial testing become part of quality, not just security
  • Monitoring in production is part of the quality system, not separate from it
  • Output quality depends on grounding data that changes over time
  • A system that passed last quarter may drift as its sources change
  • Quality controls must cover the data pipeline, not just the model
  • Re-evaluation is triggered by data change, not only code change
Deterministic software is verified once and trusted. An AI system is evaluated continuously and trusted within bounds.
Working principle · AI quality management
Section 02

Standards: ISO/IEC 42001 and the RAI principles

Two reference points anchor most enterprise AI quality programmes. ISO/IEC 42001 is the international management-system standard for artificial intelligence. Microsoft's Responsible AI (RAI) principles are the operating values that Microsoft applies to its own AI products and recommends to customers. They work at different altitudes: 42001 tells you how to run a system of governance; the RAI principles tell you what good behaviour looks like.

ISO/IEC 42001 — an AI management system

Published in 2023, 42001 is structured like ISO 27001 for information security: it defines a management system (the AIMS) with requirements for policy, risk assessment, roles, controls, and continual improvement. It is certifiable. It does not prescribe specific model thresholds; it requires you to have a defensible, documented process for setting and meeting them.

Microsoft Responsible AI — six principles

Fairness, reliability and safety, privacy and security, inclusiveness, transparency, and accountability. These translate into concrete quality criteria: a fairness assessment across user groups, a groundedness measure for reliability, a content-safety filter for safety, and an accountable owner for every deployed system.

How they fit together

Use 42001 as the scaffolding — the process that proves you are managing AI risk systematically. Use the RAI principles as the substance — the specific qualities each AI system must demonstrate. An auditor checks the scaffolding; your evaluation evidence demonstrates the substance.

Practitioner note

You do not need 42001 certification to benefit from it. Adopting its structure — a documented AI policy, a risk register, named owners, and a continual-improvement loop — gives you most of the operational value and makes future certification straightforward if it becomes a requirement.

Knowledge check

A stakeholder asks "does ISO/IEC 42001 tell us what groundedness score our chatbot must hit?" What is the accurate answer?

Section 03

Internal audit and evidence

A quality framework is only as strong as the evidence it can produce. Internal audit for AI asks a simple, demanding question of every deployed system: show me the proof. The proof is not an assurance that the system is good — it is the artefacts that demonstrate you measured, bounded, and governed its behaviour.

What an AI audit looks for

Evaluation evidence
Recorded results from a structured evaluation — quality metrics such as groundedness, relevance, coherence, and fluency, plus safety metrics for harmful content. In Azure AI Foundry, evaluation runs produce these scores against a test dataset and can be re-run and compared. The audit looks for evidence that evaluation happened against a representative dataset, with thresholds defined before the run, not chosen afterwards to fit the result.
Risk assessment and mitigation record
A documented assessment of what could go wrong — harmful output, hallucination, bias, data leakage — and the specific controls mitigating each. The audit checks that the controls are real and tested, not aspirational. A content-safety filter that is configured but never validated against adversarial inputs is not a mitigation.
Ownership and decision trail
A named accountable owner, and a record of who approved the system for each environment. The decision trail should show the criteria that were applied at each gate and the evidence that was reviewed. "It was signed off" is insufficient; "it was signed off against these criteria, on this evidence, by this owner" is the standard.
Monitoring and re-evaluation cadence
Evidence that the system is monitored in production and re-evaluated on a defined cadence or trigger. Because AI quality drifts, a one-time evaluation is not enough. The audit looks for a live monitoring configuration and a record of re-evaluation runs over time.
Reflect

Pick one AI system your organisation has deployed or is piloting. Could you produce all four evidence types above today? The gaps you find are your quality framework's first backlog.

Section 04

Quality gates and sign-off

A quality gate is a defined checkpoint between stages of an AI system's lifecycle, where progression depends on meeting explicit criteria. Gates turn "we think it's ready" into "it met these criteria, reviewed by this owner." The model below uses three gates aligned to the journey from build to production.

Gate 1 — Development to evaluation

Criteria: the system has a documented purpose and scope; a representative evaluation dataset exists; quality and safety thresholds are defined in advance. Sign-off confirms the system is ready to be evaluated against agreed standards — not that it is good yet.

Gate 2 — Evaluation to staged release

Criteria: evaluation results meet or exceed the pre-defined thresholds for groundedness, relevance, and safety; the risk assessment is complete with mitigations tested; an accountable owner is named. Sign-off authorises a limited, monitored release to a controlled user group.

Gate 3 — Staged release to full production

Criteria: monitoring data from the staged release shows behaviour within bounds; no unresolved high-severity incidents; the re-evaluation cadence and incident-response runbook are in place. Sign-off authorises full production and commits the system to the ongoing quality cycle.

Continue on Microsoft Learn

Go deeper on the evaluation evidence and principles behind these gates:

Evaluate and monitor AI models in Azure AI Foundry ↗

Responsible AI principles in practice ↗

End of Unit 16

You should now be able to:

  • Articulate why AI quality is statistical and continuous rather than a one-time pass.
  • Position ISO/IEC 42001 and the RAI principles within a single quality programme.
  • Run a three-gate sign-off model with evidence-based criteria at each gate.
Section 05

Unit review

Question 1 of 4

Why can an AI system built on a large language model not be certified "correct" by a single test pass?

Question 2 of 4

How do ISO/IEC 42001 and Microsoft's Responsible AI principles relate?

Question 3 of 4

An internal audit finds a content-safety filter is configured but was never tested against adversarial inputs. How should it be classified?

Question 4 of 4

Why must quality thresholds be defined before an evaluation run, not after?

End of module

You have completed Course 16: Quality Frameworks for AI. Next: Monitoring AI in Production — observability, drift detection, and live quality measurement.

Craig Stanley Studio · Deploy — Architecture, Integration & Operations · Quality Frameworks for AI · Access by direct link only.