Why AI incidents are different
A conventional incident is usually something stopping: a service is down, a queue is backed up, a database is unreachable. The system tells you something is wrong. AI incidents are frequently the opposite — the system keeps working, confidently, while doing the wrong thing. A model hallucinating a refund policy, leaking another tenant's data in a response, or emitting harmful content is fully "up." Nothing is erroring. Users may even be satisfied.
This inverts the response discipline. You cannot wait for the system to fall over. You need runbooks tuned to behavioural failures — hallucination at scale, harmful output, and data leakage — with a severity model that recognises that "available but wrong" can be your most damaging state. This unit builds that response capability.
By the end of this unit- Explain how AI incidents differ from conventional outages and why "available but wrong" is a critical state.
- Classify AI incidents by severity using impact and reversibility, not just availability.
- Apply response playbooks for hallucination, harmful output, and data leakage, and run a post-incident review.
Three failure modes unique to AI
- The system confidently states things that are false
- At scale, many users receive the same wrong answer
- No error is thrown; detection relies on quality monitoring or user reports
- Impact compounds the longer it goes undetected
- The system produces unsafe, offensive, or policy-violating content
- May be triggered by adversarial input (a jailbreak) or by a content-safety gap
- Reputational and compliance impact can be immediate and severe
- Requires both containment and a record for the safety team
- The system exposes data a user should not see — another tenant's, or restricted records
- Often a grounding or permissions failure, not a model failure
- Highest reversibility cost: disclosed data cannot be un-disclosed
- Frequently triggers regulatory notification obligations
Severity classification
Severity decides how fast you respond and who you wake up. For AI incidents, availability is a poor severity axis — the system is usually available. Classify instead on two dimensions: impact (how many users, how serious the harm) and reversibility (can the consequence be undone?). A leak of restricted data to one user can outrank a model being slow for thousands.
SEV-1 — Critical, often irreversible
Data leakage of sensitive or regulated information; sustained harmful output reaching real users; an incident with legal or regulatory notification obligations. Immediate containment, executive and legal notification, and likely a service pause. The reversibility cost is the trigger: you cannot un-disclose data.
SEV-2 — High impact, contained
Widespread hallucination on a material topic (for example a wrong policy quoted to many customers), or harmful output caught before broad reach. Rapid response by the AI owner, mitigation in hours, stakeholder communication. Damaging but recoverable through correction and communication.
SEV-3 — Localised or low impact
Isolated incorrect answers, minor relevance degradation, or a quality alert without confirmed user harm. Handled within the normal quality cycle: investigate, raise sampling, fix, re-evaluate. No emergency mobilisation, but recorded and trended.
Reversibility is the dimension teams most often under-weight. A slow, available system feels urgent because it is visible; a single irreversible data disclosure feels quiet but is far more serious. Bake reversibility explicitly into your severity matrix so the quiet, severe incidents are not under-classified.
One support user receives a response containing another customer's account details. Volume is tiny. How should this be classified?
Response playbooks
A playbook is a pre-written, rehearsed sequence for a specific incident type. Written in the calm before an incident, it removes improvisation from the moment when judgement is hardest. Each AI failure mode needs its own. All share a common spine: contain, assess, communicate, remediate, record.
Playbooks by failure mode
Hallucination at scale
Harmful output
Data leakage
The common spine
For an AI system you support, which incident types are "contain-first" and which allow assess-first? Deciding this now, in writing, is what makes the playbook usable under pressure later.
Escalation and post-incident review
An incident is not closed when the symptom stops. It is closed when you have learned from it and reduced the chance of recurrence. Two disciplines deliver that: clear escalation during the incident, and a blameless post-incident review afterwards.
Define escalation paths in advance
Each severity maps to named roles and a notification timeframe: who is paged, who is informed, and who has authority to pause the service. For SEV-1, that includes legal, privacy, and an executive sponsor. Escalation decided mid-incident is escalation decided too late.
Run the review blamelessly
The post-incident review examines the system and process, not the individuals. The aim is to find the conditions that allowed the incident — a missing guardrail, an untested permission boundary, a monitoring blind spot — not a person to fault. Blame suppresses the honest reporting you depend on.
Convert findings into controls
Every review produces concrete actions: a new adversarial test case, a tightened threshold, a monitoring signal that was missing, a playbook update. An AI incident that does not strengthen your evaluation suite or monitoring has been wasted. Track these actions to completion.
Feed the quality system
Incident learnings flow back into evaluation datasets, content-safety configuration, and monitoring baselines. The incident-response process and the quality framework are one loop: each incident makes the next evaluation stronger and the next regression easier to catch.
Detection and response capabilities that support AI incident handling:
Detect and respond to cybersecurity threats with Microsoft Defender ↗
End of Unit 18
You should now be able to:
- Recognise behavioural AI incidents and why "available but wrong" demands a different discipline.
- Classify by impact and reversibility, weighting irreversible disclosure correctly.
- Apply contain-assess-communicate-remediate-record playbooks and run a blameless review that strengthens controls.
Unit review
What makes "available but wrong" the most dangerous state for an AI system?
Why is reversibility a core dimension of AI incident severity?
Which incident type is the clearest "contain-first" case, where you pause before full assessment?
What is the purpose of a blameless post-incident review?
End of module
You have completed Course 18: Incident Response for AI Systems. Next: Scaling AI Pilots to Production — what breaks in the transition and how to build a production programme.