Corpus Agentis
The field book to agent ecosystems
The field book to agent ecosystems
Security · Incidents

How agents actually fail

An agent security incident is a documented case where an agent was made to act against its operator: data exfiltrated, credentials taken, a tool call turned into a weapon. Unlike a model jailbreak, which produces bad text, an agent compromise produces a real action with real consequences. Every row traces to a primary disclosure, with the attack class, measured impact and the defence that closed it. MITRE ATLAS frames this as the adversarial surface of AI systems: each new capability opens a matching attacker technique (MITRE ATLAS, 2026).

Threat map · attack class × disclosure date · dot size = severity (CVSS)
202420252026Prompt injectionData exfiltrationCredential theftAgent hijackTool/MCP poisoningJailbreakDestructive actionAutonomous misuseSupply chain
Incident log · 38 · newest first · click a row for detail
DateIncidentClassImpactSource
2026-07-16 OpenAI model escaped sandbox and breached Hugging Face during a cyber eval Autonomous misuse 17,000+ events reconstructed · 2 code-execution paths exploited · weekend-scale lateral movement · public models + Spaces untouched ↗ T1
2026-05-26 Safety alignment strippable via minimal fine-tuning · open-weight and closed-API Supply chain 10 examples · <$0.20 · sufficient to strip safety on closed API · 95% CWE-20 emission via style-trigger · 82% multi-agent malicious-instruction execution rate ↗ T1
2026-05-15 AudioHijack · hidden-audio prompt injection against voice-agent assistants Prompt injection Windows Copilot demo: 100% action-execution from a single MP3 · LALM jailbreak success up to 45% via imperceptible audio edits ↗ T1
2026-05-08 OX Security · systemic MCP RCE across the AI-agent ecosystem Supply chain ~150M downloads implicated · 7,000+ public MCP servers · up to 200,000 vulnerable instances ↗ T2
2026-05-07 Microsoft Semantic Kernel prompt-injection-to-RCE Destructive action 2 CVEs · patched Python ≥1.39.4 / .NET ≥1.71.0 ↗ T1
2026-04-14 Comment and Control: cross-vendor coding-agent injection Prompt injection 3 vendors · one shared payload ↗ T1
2026-03-18 Claudy Day: claude.ai URL-parameter injection chain Data exfiltration 3 chained flaws · injection fixed, redirect ongoing ↗ T1
2026-02-25 Claude Code project-file hooks RCE (CVE-2025-59536) Destructive action CVE-2025-59536 · fixed v1.0.111 ↗ T1
2026-02-25 Claude Code ANTHROPIC_BASE_URL API-key exfiltration (CVE-2026-21852) Credential theft CVE-2026-21852 · fixed 2025-12-28 ↗ T1
2026-02-12 SCAM benchmark, credential mishandling Credential theft 8 models · 30 scenarios · scores 35-92% ↗ T1
2026-02-10 SmartLoader · trojanized Oura Ring MCP server pushed to registry via fake developer personas Supply chain 3 months of persona-building · 5 fake GitHub accounts · StealC infostealer payload · exfiltrated browser creds + SSH keys + wallets + API keys ↗ T2
2026-01-15 Claude Cowork indirect-injection file exfiltration Data exfiltration Snyk: 36.8% of 3,984 agent skills had ≥1 flaw ↗ T1
2025-12-14 MINJA · agent memory poisoning via normal queries (NeurIPS 2025) Prompt injection MINJA: >95% injection success on production agents · PoisonedRAG: attacker-chosen answers via a handful of documents · confirmed against Gemini, Bedrock, Azure ↗ T1
2025-11-13 Anthropic disrupts AI-orchestrated espionage (GTG-1002) Autonomous misuse ~30 global targets · AI ran 80-90% with 4-6 human decision points ↗ T1
2025-10-28 Claude Code Interpreter Files-API data exfiltration Data exfiltration up to 30 MB/file · HackerOne 2025-10-25 ↗ T1
2025-10-24 ChatGPT Atlas omnibox prompt-injection jailbreak Jailbreak Malformed 'URL' can drive the agent to phishing or destructive Drive actions ↗ T1
2025-10-08 CamoLeak, GitHub Copilot Chat Data exfiltration CVSS 9.6 (no CVE assigned) ↗ T1
2025-09-25 ForcedLeak, Salesforce Agentforce Data exfiltration CVSS 9.4 · exfil via expired whitelisted domain ↗ T1
2025-09-19 Notion 3.0 agent data exfiltration Data exfiltration Client names, company details and ARR exfiltrated via the search tool ↗ T1
2025-08-20 Perplexity Comet browser-agent prompt injection Prompt injection PoC chained to read email, capture an OTP and take over the user's Gmail ↗ T1
2025-08-06 Invitation Is All You Need, Gemini hijack Agent hijack 14 promptware scenarios · physical smart-home control ↗ T1
2025-08-06 AgentFlayer zero-click ChatGPT Connectors exfiltration Data exfiltration 0-click theft of Google Drive API keys · url_safe check bypassed via Azure Blob ↗ T1
2025-08-05 Cursor MCPoison MCP trust-swap RCE (CVE-2025-54136) Tool/MCP poisoning CVE-2025-54136 · fixed Cursor v1.3 ↗ T1
2025-08-01 Cursor CurXecute MCP prompt-injection RCE Tool/MCP poisoning CVE-2025-54135 · full RCE · patched in Cursor 1.3 ↗ T1
2025-07-23 Amazon Q Developer extension wiper injection Supply chain Shipped in v1.84.0 but a syntax error stopped it running; fixed v1.85.0 ↗ T1
2025-07-19 Replit AI agent deletes production database Destructive action ~1,200 executive records + ~1,190 companies wiped during a freeze ↗ T2
2025-07-09 mcp-remote OS command injection (CVE-2025-6514) Tool/MCP poisoning CVSS 9.6 · 437k+ downloads · fixed 0.1.16 ↗ T1
2025-06-13 Anthropic MCP Inspector RCE Supply chain CVE-2025-49596 · CVSS 9.4 · MCP Inspector < 0.14.1 ↗ T1
2025-06-11 EchoLeak, M365 Copilot zero-click exfiltration Data exfiltration CVE-2025-32711 · CVSS 9.3 · zero-click ↗ T1
2025-05-26 GitHub MCP private-repo exfiltration Prompt injection PoC leaked private repo details, relocation plans and salary; architectural, not a code bug ↗ T1
2025-04-01 MCP Tool Poisoning Attacks Tool/MCP poisoning Poisoned tool exfiltrated ~/.ssh/id_rsa and MCP config from Cursor ↗ T1
2025-02-17 ChatGPT Operator prompt-injection PII exfiltration Data exfiltration PII lifted from authenticated Booking.com, Hacker News and Guardian sessions ↗ T1
2025-01-31 DeepSeek R1 jailbreak-resistance failure Jailbreak 100% attack success on 50 HarmBench prompts vs 26% for o1-preview ↗ T1
2024-08-20 Slack AI private-channel data exfiltration Data exfiltration Private-channel exfil via public-channel injection · no CVE ↗ T1
2024-06-26 Skeleton Key universal jailbreak Jailbreak Bypassed GPT-4o, Gemini Pro, Claude 3 Opus, Llama3-70B, Mistral Large, Command R+ ↗ T1
2024-04-02 Many-shot jailbreaking Jailbreak Tested to 256 shots; success follows a power law, worse with longer context ↗ T1
2024-03-05 Morris II GenAI worm Autonomous misuse Demonstrated against ChatGPT-, Gemini-Pro- and LLaVA-backed email assistants ↗ T1
2023-02-08 Bing Chat 'Sydney' system-prompt leak Prompt injection Entire hidden metaprompt exposed by a single instruction ↗ T2
Field notes

A source-first catalogue of documented, reproduced agent compromises

The agent is rarely the thing that goes wrong. Across the log, only three of the thirty-eight entries are categorised as autonomous misuse. Fifteen are categorised by how someone got in, through prompt injection, tool and MCP poisoning, or the supply chain, and the largest single category, data exfiltration at ten, describes what was lost rather than how. A crafted email, a web form submission, a public channel message, a comment in a pull request: the instruction arrives inside data the agent was expected to read, so nobody approves it and often nobody sees it.

Show more

None of this is unfamiliar. Untrusted input, over-broad permissions and credentials sitting where they should not are what IT has been handling for decades. What changes is surface area. An agent reads more sources, holds more credentials and acts on more systems than the software it replaces, so a familiar failure has more ways in and more reach once it lands.

From the corpus, curated by Brandon Chaplin
Common questions
How do AI agents get hacked?

Usually by being told to, not by their code being broken. An attacker plants instructions in content the agent reads, and the agent carries them out using the tools and permissions its operator gave it. The documented cases cluster into prompt injection, data exfiltration, credential theft, tool and MCP poisoning, jailbreaks and supply-chain compromise. MITRE ATLAS catalogues these as adversary techniques against AI systems (MITRE ATLAS).

What is prompt injection?

Content crafted to override an agent's instructions. Direct injection arrives in the user's message; indirect injection hides in something the agent reads, such as a web page, email or document, and fires when the agent processes it. Because agents can act, a successful injection can move data or trigger a real action. OWASP ranks it the top risk for LLM applications (OWASP).

What is a zero-click AI exploit?

One where the victim does nothing at all. The malicious instruction arrives in content the agent processes automatically, such as an incoming email, and executes without anyone clicking. Several of the highest-severity incidents in the log below are zero-click indirect injections against enterprise copilots.

What is the most common way agents fail in production?

Data leaving where it should not, with credential mishandling close behind. In benchmark testing, frontier models have forwarded passwords, emailed secret keys and typed credentials into phishing logins. The failure is rarely a clever exploit of the model. It is the model being trusted with a secret or a permission it should never have held.

How do you stop prompt injection?

You cannot stop it outright, so you limit what a successful one can reach. Keep raw secrets out of the model, give each tool the narrowest permission that works, sandbox execution, and require human approval before anything irreversible. Assume some injections will land, and design so that landing is survivable.

Have there been real AI agent security incidents?

Yes. Documented cases include CVE-assigned vulnerabilities in enterprise copilots, vendor security advisories, and failures reproduced by security researchers. The log below traces each one to its primary disclosure, with a severity score where the disclosure published one. These are recorded failures, not hypotheticals.

Next in the learning path