Section 01 · Unit introduction

Why AI incidents are different

A conventional incident is usually something stopping: a service is down, a queue is backed up, a database is unreachable. The system tells you something is wrong. AI incidents are frequently the opposite — the system keeps working, confidently, while doing the wrong thing. A model hallucinating a refund policy, leaking another tenant's data in a response, or emitting harmful content is fully "up." Nothing is erroring. Users may even be satisfied.

This inverts the response discipline. You cannot wait for the system to fall over. You need runbooks tuned to behavioural failures — hallucination at scale, harmful output, and data leakage — with a severity model that recognises that "available but wrong" can be your most damaging state. This unit builds that response capability.

By the end of this unit
  • Explain how AI incidents differ from conventional outages and why "available but wrong" is a critical state.
  • Classify AI incidents by severity using impact and reversibility, not just availability.
  • Apply response playbooks for hallucination, harmful output, and data leakage, and run a post-incident review.

Three failure modes unique to AI

  • The system confidently states things that are false
  • At scale, many users receive the same wrong answer
  • No error is thrown; detection relies on quality monitoring or user reports
  • Impact compounds the longer it goes undetected
  • The system produces unsafe, offensive, or policy-violating content
  • May be triggered by adversarial input (a jailbreak) or by a content-safety gap
  • Reputational and compliance impact can be immediate and severe
  • Requires both containment and a record for the safety team
  • The system exposes data a user should not see — another tenant's, or restricted records
  • Often a grounding or permissions failure, not a model failure
  • Highest reversibility cost: disclosed data cannot be un-disclosed
  • Frequently triggers regulatory notification obligations
For a normal outage, the worst state is "down." For an AI system, the worst state is "up and confidently wrong."
Working principle · AI incident response
Section 02

Severity classification

Severity decides how fast you respond and who you wake up. For AI incidents, availability is a poor severity axis — the system is usually available. Classify instead on two dimensions: impact (how many users, how serious the harm) and reversibility (can the consequence be undone?). A leak of restricted data to one user can outrank a model being slow for thousands.

SEV-1 — Critical, often irreversible

Data leakage of sensitive or regulated information; sustained harmful output reaching real users; an incident with legal or regulatory notification obligations. Immediate containment, executive and legal notification, and likely a service pause. The reversibility cost is the trigger: you cannot un-disclose data.

SEV-2 — High impact, contained

Widespread hallucination on a material topic (for example a wrong policy quoted to many customers), or harmful output caught before broad reach. Rapid response by the AI owner, mitigation in hours, stakeholder communication. Damaging but recoverable through correction and communication.

SEV-3 — Localised or low impact

Isolated incorrect answers, minor relevance degradation, or a quality alert without confirmed user harm. Handled within the normal quality cycle: investigate, raise sampling, fix, re-evaluate. No emergency mobilisation, but recorded and trended.

Practitioner note

Reversibility is the dimension teams most often under-weight. A slow, available system feels urgent because it is visible; a single irreversible data disclosure feels quiet but is far more serious. Bake reversibility explicitly into your severity matrix so the quiet, severe incidents are not under-classified.

Knowledge check

One support user receives a response containing another customer's account details. Volume is tiny. How should this be classified?

Section 03

Response playbooks

A playbook is a pre-written, rehearsed sequence for a specific incident type. Written in the calm before an incident, it removes improvisation from the moment when judgement is hardest. Each AI failure mode needs its own. All share a common spine: contain, assess, communicate, remediate, record.

Playbooks by failure mode

Hallucination at scale
Contain: if the wrong answer is on a material topic, deploy a guardrail or canned safe response for that topic, or pause the feature. Assess: determine the scope — which queries, how many users, since when — using your monitoring and logs. Communicate: notify affected users where the wrong information could cause action (for example a quoted policy). Remediate: fix the grounding, prompt, or retrieval that caused it; re-evaluate before reopening. Record: capture the offending query/response pairs for the post-incident review.
Harmful output
Contain: tighten or enable content-safety filters immediately; if the trigger is a jailbreak pattern, block it at the input. Pause the feature if harm is reaching users. Assess: reconstruct the input that elicited the output and confirm whether it is adversarial or a filter gap. Communicate: notify the safety/RAI owner; escalate to legal if the content carries liability. Remediate: close the gap, test against the adversarial input, and re-evaluate safety metrics. Record: add the case to your adversarial test set so it is caught in future.
Data leakage
Contain first, always: this is the one playbook where you pause the system before full assessment if leakage is confirmed — the reversibility cost is too high to leave it running. Assess: determine what was disclosed, to whom, and the root cause (grounding scope, permissions, or retrieval boundary failure). Communicate: engage legal and privacy immediately; regulatory clocks may already be running. Remediate: fix the permission or grounding boundary and verify with a targeted test. Record: full timeline for the formal review and any regulatory submission.
The common spine
Every playbook follows contain → assess → communicate → remediate → record. The ordering of contain and assess is the key variable: for hallucination you may assess scope before deciding to pause; for confirmed data leakage you contain first and assess after. Knowing which incidents are "contain-first" is the single most valuable piece of judgement to settle in advance, while calm.
Reflect

For an AI system you support, which incident types are "contain-first" and which allow assess-first? Deciding this now, in writing, is what makes the playbook usable under pressure later.

Section 04

Escalation and post-incident review

An incident is not closed when the symptom stops. It is closed when you have learned from it and reduced the chance of recurrence. Two disciplines deliver that: clear escalation during the incident, and a blameless post-incident review afterwards.

Define escalation paths in advance

Each severity maps to named roles and a notification timeframe: who is paged, who is informed, and who has authority to pause the service. For SEV-1, that includes legal, privacy, and an executive sponsor. Escalation decided mid-incident is escalation decided too late.

Run the review blamelessly

The post-incident review examines the system and process, not the individuals. The aim is to find the conditions that allowed the incident — a missing guardrail, an untested permission boundary, a monitoring blind spot — not a person to fault. Blame suppresses the honest reporting you depend on.

Convert findings into controls

Every review produces concrete actions: a new adversarial test case, a tightened threshold, a monitoring signal that was missing, a playbook update. An AI incident that does not strengthen your evaluation suite or monitoring has been wasted. Track these actions to completion.

Feed the quality system

Incident learnings flow back into evaluation datasets, content-safety configuration, and monitoring baselines. The incident-response process and the quality framework are one loop: each incident makes the next evaluation stronger and the next regression easier to catch.

Continue on Microsoft Learn

Detection and response capabilities that support AI incident handling:

Detect and respond to cybersecurity threats with Microsoft Defender ↗

Responsible AI principles in practice ↗

End of Unit 18

You should now be able to:

  • Recognise behavioural AI incidents and why "available but wrong" demands a different discipline.
  • Classify by impact and reversibility, weighting irreversible disclosure correctly.
  • Apply contain-assess-communicate-remediate-record playbooks and run a blameless review that strengthens controls.
Section 05

Unit review

Question 1 of 4

What makes "available but wrong" the most dangerous state for an AI system?

Question 2 of 4

Why is reversibility a core dimension of AI incident severity?

Question 3 of 4

Which incident type is the clearest "contain-first" case, where you pause before full assessment?

Question 4 of 4

What is the purpose of a blameless post-incident review?

End of module

You have completed Course 18: Incident Response for AI Systems. Next: Scaling AI Pilots to Production — what breaks in the transition and how to build a production programme.

Craig Stanley Studio · Deploy — Architecture, Integration & Operations · Incident Response for AI Systems · Access by direct link only.