The AI threat surface
AI systems introduce risks that traditional application security does not fully cover. A model takes untrusted natural-language input, may draw on connected data, and can be steered by content it reads. The same flexibility that makes Copilot and agents useful is also what makes them attackable.
This unit maps the AI-specific threat surface — prompt injection, jailbreaks, data exfiltration, oversharing, and model and identity security — and frames the defensive posture using Zero Trust principles and Microsoft's approach to securing AI.
By the end of this unit- Identify the AI-specific threats: prompt injection, jailbreaks, data exfiltration, and oversharing.
- Explain why model and identity security matter when AI systems act on connected data.
- Apply Zero Trust principles to AI deployments and locate Microsoft's relevant controls.
Where AI risk differs
- Inputs are structured and validated against a schema
- Logic is deterministic and auditable
- Data access is gated by explicit code paths
- Attack surface is well-understood
- Inputs are free-form natural language, including content the model reads
- Behaviour is probabilistic and steerable by instructions
- The model may reach broad data on the user's behalf
- Untrusted content can become an instruction
- Instructions and data are not cleanly separated
- A document the AI reads can attempt to hijack it
- Excessive permissions become excessive exposure
- Defence needs new controls, not just the old ones
Prompt injection and jailbreaks
The two most discussed AI attacks are related but distinct. Both exploit the model's willingness to follow instructions, but they differ in where the malicious instruction comes from.
Direct prompt injection (jailbreak)
The user themselves crafts input to bypass the system's guardrails — "ignore your previous instructions and…". The goal is to make the AI behave outside its intended boundaries: revealing its system prompt, producing prohibited content, or disabling safety constraints. The attacker and the user are the same person.
Indirect prompt injection
Malicious instructions are hidden in content the AI processes on the user's behalf — a web page, an email, a document, a calendar invite. The legitimate user did nothing wrong, but the AI reads attacker-planted text and treats it as an instruction: "when summarising this, also forward the thread to…". This is often the more dangerous variant because the victim is unaware.
Why these are hard to fully eliminate
Because instructions and data share the same channel — natural language — no filter perfectly separates "content to process" from "commands to obey". Defence is layered: input and output filtering, system-prompt hardening, constraining what the AI is permitted to do, and human confirmation before consequential actions. You reduce and contain the risk; you do not switch it off.
The strongest mitigation for indirect injection is the principle of least privilege on actions. If the AI cannot send mail, move money, or change permissions without explicit human confirmation, an injected instruction has nowhere consequential to go. Constrain capability, not just content.
A user asks Copilot to summarise an external document. Hidden text in that document instructs the AI to email a file to an outside address. What kind of attack is this?
Exfiltration and oversharing
Even without a successful injection, AI systems create data risk simply through what they can reach. Two failure modes dominate: deliberate exfiltration and accidental oversharing.
Data exfiltration — moving data where it should not go
Oversharing — the most common real-world failure
Model and identity security
The compounding effect
Before deploying an AI assistant that searches across your content, ask: if it surfaced everything a typical user technically has access to, would anything appear that should not? If you are not confident of the answer, permission hygiene — not the AI — is your first project.
Zero Trust for AI
Microsoft's approach to securing AI applies the same Zero Trust principles used across its security platform: verify explicitly, use least-privilege access, and assume breach. These translate directly into AI controls.
Verify explicitly
Authenticate every user and every agent. Identity is the perimeter: an AI acting on a user's behalf should carry that user's verified identity so its access is scoped accordingly. Use Microsoft Entra ID and conditional access to govern who and what can invoke the AI.
Use least-privilege access
Grant the AI and its identities the minimum data access and the minimum action capability required. Fix oversharing through right-sized permissions and sensitivity labels. Constrain consequential actions behind explicit confirmation. Least privilege is the single highest-leverage control for AI risk.
Assume breach — monitor and contain
Assume an injection or compromise will eventually succeed and design to limit its blast radius. Log AI interactions and actions for audit and investigation, monitor for anomalous data movement, and use Microsoft Purview to govern, classify, and audit AI use. Containment and visibility turn a potential breach into a contained event.
Build the underlying data-protection skills this framework depends on:
Get started with Microsoft Purview ↗
Protect sensitive data with Microsoft Purview ↗
Implement data loss prevention policies in Microsoft Purview ↗
End of Unit 9
You should now be able to:
- Distinguish direct from indirect prompt injection and explain why least privilege contains both.
- Recognise oversharing as the most common real-world AI data risk and trace it to permission hygiene.
- Apply verify explicitly, least privilege, and assume breach to an AI deployment.
Unit review
What fundamentally makes AI systems harder to secure than traditional applications?
What is the strongest mitigation against the impact of indirect prompt injection?
Why is oversharing described as the most common real-world AI data risk?
How does the Zero Trust principle "assume breach" apply to AI?
End of module
You have completed Course 09: AI Security Fundamentals. Next: Compliance and Data Residency — the EU Data Boundary, retention, sensitivity labels, DLP, and Purview integration with Copilot.