Prompt Injection in AI Agents: Why It Can’t Be Patched
Prompt injection turns any text an AI agent reads into a possible command. See the EchoLeak case, the lethal trifecta test, and defenses that hold up.
Prompt injection turns any text an AI agent reads into a possible command. See the EchoLeak case, the lethal trifecta test, and defenses that hold up.
Prompt injection works because a language model reads commands and data through the same channel. Any text an AI agent reads can act as a command. No model update fully fixes this. The real defense is design: assume some injected command will get through, then limit what a hijacked agent can reach, send, or change.
Table of Contents
ToggleOlder attacks such as SQL injection worked when software mixed untrusted data into a command. Engineers fixed most of them by keeping the two apart. A language model has no separate lane, though. A system prompt, a user request, an email body, and a web page all arrive as one stream of text.
Two forms matter. In direct prompt injection, the person typing tries to override the AI agent’s rules. In indirect prompt injection, the attacker never touches the AI. Instead, they plant commands in content the agent reads later, such as an email, a shared document, or a web page. The user sees a normal summary. The agent sees a command.
Security teams know this pattern from human error. A trusted channel carries something it should not, and nobody checks. OWASP ranks prompt injection first on its Top 10 for LLM Applications. The ranking reflects how hard the problem is to design away, not how often it makes headlines. That makes it a design problem more than a patching problem.
The clearest public case is EchoLeak, tracked as CVE-2025-32711 and rated 9.3 out of 10 for severity. Aim Security researchers disclosed it in June 2025, and an academic case study on arXiv later analyzed it in detail. The flaw sat in Microsoft 365 Copilot, and it needed no click from the victim.
The attack began like a classic phishing email. It looked plain, but it carried hidden commands. Later, the victim asked Copilot an unrelated question. Copilot pulled the email in as background and followed the hidden commands. It then packed details from the victim’s files into an outbound link.
The exploit beat Microsoft’s injection classifier and its link redaction. It also passed the content security policy, because the data left through an allowlisted Microsoft domain. Microsoft patched the flaw on its servers, and no attacks in the wild have been confirmed.
Three lessons follow. The attacker never logged in to anything. Every individual control worked as designed. Yet the defenses that failed were the ones built to detect prompt injection.
Discover how Threatcop protects your workforce from modern cyber threats.
Email is only one route, though the classic phishing email is still the most common. An AI agent reads many kinds of content, so each one can carry commands. Common hiding places include:
The lesson is simple, because the pattern repeats. Treat every source the agent reads as untrusted unless your own team wrote it.
Security researcher Simon Willison coined the term prompt injection. He also gave the risk a simple checklist called the lethal trifecta. An AI agent becomes risky when it combines three abilities:
EchoLeak had all three. Remove any one, and the attack loses its payoff. An agent that reads untrusted pages but sees no private data has little to steal. Likewise, an agent with private data but no way out has nowhere to send it. The test turns a vague fear into a yes-or-no question about each agent.
Vendors sell guardrail products that claim to catch most injection attempts. The research on adaptive attackers is less encouraging. In a 2025 paper called “The Attacker Moves Second,” researchers including Milad Nasr, Nicholas Carlini, and Florian Tramèr showed that stronger adaptive attacks bypass published defenses. A filter tested only against yesterday’s attacks says little about tomorrow’s.
Willison makes the practical point sharper. In web security, catching 95% of attacks is a failing grade, because the attacker only needs the other 5%. Detection still adds value as one layer. However, the whole design cannot depend on it. Treat prompt injection as a standing item in information security risk management, not a one-time patch.
Stronger defenses limit consequences instead of trying to spot every malicious sentence. A 2025 paper from researchers at IBM, Invariant Labs, ETH Zurich, Google, and Microsoft describes six design patterns. They share one principle: once an AI agent has read untrusted input, it should be unable to take high-impact actions because of that input.
In plain terms, three of the patterns work like this:
These designs cost some flexibility. An agent restricted this way cannot improvise across arbitrary tasks. That is why an AI risk management framework should decide, use case by use case, how much autonomy is worth the exposure.
Speed matters more than certainty, so act first. When you suspect prompt injection, pause the AI agent and revoke its tokens. Next, save the logs, the prompts, and the content it read. Then list every action it took since it read that content. If it sent data out, treat that data as exposed. Also rotate any secrets it could see. Finally, add the poisoned content to your test set, so the same trick fails next time.
Staff can help, because they notice odd actions first. An assistant might mention a document nobody asked about. It might also draft a message nobody requested. Reporting only works if it is quick and blame-free, so people need a simple way to flag what they see, and they need to know who reads each report.
Prompt injection is social engineering aimed at software. It borrows what works on people: fake authority, urgency, and content from a channel they trust. Staff who already question odd requests will question odd AI actions too.
A smarter filter will not solve prompt injection, because the weakness sits in how language models read text. It will be managed the way other unfixable weaknesses are managed. Assume failure, keep the blast radius small, and give people a fast way to report when an AI agent does something nobody asked for. A one-click reporting workflow built for phishing gives employees that path.
Prompt injection is an attack that hides instructions inside content an AI system reads. The AI then follows the attacker’s text instead of its owner’s intent. It exists because language models cannot tell commands apart from plain data.
Direct prompt injection comes from the user typing into the AI and trying to beat its rules. Indirect prompt injection hides instructions in outside content, such as an email or web page. Indirect attacks are more dangerous because the victim never sees anything unusual.
Not with current language models. Research on adaptive attackers shows that filters can be bypassed. The realistic goal is limiting damage through strict permissions, human approval, and data-flow controls.
Staff see AI in action every day, so they can spot oddities first. Examples include odd actions, mentions of unknown documents, or messages nobody requested. Speed matters, so reports should reach someone who can act within minutes.
Apply the lethal trifecta test. If the AI agent can read private data, ingest content an attacker can influence, and communicate externally, it is exposed to prompt injection. Removing any one of the three breaks the attack path.
Adhish Chakma is a Senior Product Manager at Kratikal, where he leads product initiatives focused on cybersecurity and AI-powered solutions. With experience in product management and cybersecurity, he works on developing practical technologies that address evolving security challenges. His areas of interest include People Security Management, cybersecurity awareness, AI-driven security, email security, and human-layer risk. He is passionate about building security products that make organizations more resilient against emerging cyber threats.
Adhish Chakma is a Senior Product Manager at Kratikal, where he leads product initiatives focused on cybersecurity and AI-powered solutions. With experience in product management and cybersecurity, he works on developing practical technologies that address evolving security challenges. His areas of interest include People Security Management, cybersecurity awareness, AI-driven security, email security, and human-layer risk. He is passionate about building security products that make organizations more resilient against emerging cyber threats.
Blocking AI backfires. Learn six steps to secure AI adoption: inventory, tiers, vendor review, limited pilots, role-based training, and...
AI makes scams more personal and moves them across email, chat, and video. See the Arup deepfake case and...
A 19,500-person study found standard phishing training barely works. See where agentic AI could help, its risks, and how...
Table of Contents
×