The lethal trifecta for AI agents, and how to break it
Prompt injection becomes data theft when one agent can read private data, read content an attacker controls, and send data out. Each capability is harmless alone. Together, across different MCP servers, they are the most common way agents leak.
Three capabilities that must not meet
Simon Willison named the pattern in June 2025 (“The lethal trifecta for AI agents”). An agent is exposed when it has all three of:
- Untrusted content: anything an attacker can write that the agent will read: web pages, emails, GitHub issues, documents, tool results.
- Private data: files, databases, inboxes, private repositories, credentials.
- External communication: any way to move data out: an HTTP request, an email, a comment on a public issue, even an image URL the client renders.
Language models follow instructions wherever they appear. Text in leg 1 can tell the agent to read leg 2 and send it through leg 3, and no amount of prompting reliably stops that. The only robust defence is to make sure one agent never has all three at once without a human in the loop.
Why MCP makes it easy to assemble by accident
With MCP you add capabilities server by server, and each one looks fine:
- a fetch server reads the web (untrusted content) and can request any URL (external communication);
- a filesystem server reads your project and home directory (private data).
Neither server is the problem. The agent that has both is. In May 2025 Invariant Labs showed the pattern against the GitHub MCP server: a malicious issue on a public repository instructed the agent to read private repositories and publish their contents in a pull request (write-up). A single server supplied all three legs, and there was no bug to patch.
Finding the trifecta across servers
Because the risk lives in the combination, check every server one agent uses in one run. Hecate’s rule MCP005 classifies each tool by name, parameters and description, then looks across all servers you pass:
npx hecate-mcp@beta scan .mcp.json --connect
[HIGH] MCP005 Lethal trifecta: untrusted content, private data and external communication
One agent can read content an attacker controls, access private data, and
send data out. Text from the first can instruct it to do the other two.
The legs span 2 servers (fetch, filesystem).
1. Untrusted content fetch/fetch
2. Private data filesystem/read_file, filesystem/list_directory
3. External communication fetch/fetch
fix: Remove a leg for this agent: drop one group of tools, split them across
separate agents, or require human approval before external-communication
tools run.
hecate graph prints the same analysis as a Mermaid diagram you can paste into a pull request or design doc:
flowchart LR
subgraph untrusted["1. Untrusted content"]
untrusted_0["fetch/fetch"]
end
subgraph privdata["2. Private data"]
privdata_0["filesystem/read_file"]
privdata_1["filesystem/list_directory"]
end
subgraph egress["3. External communication"]
egress_0["fetch/fetch"]
end
untrusted -->|"injected instructions"| privdata
privdata -->|"stolen data"| egress
Classification is heuristic. If a tool is miscategorised, for example a fetch tool that can only reach one internal host through an egress proxy, correct it in hecate.policy.json with a written reason rather than silencing the whole rule.
Breaking the trifecta
- Split agents. A research agent that browses the web holds no credentials; the agent that touches private data has no web access. Hand results between them as data, reviewed where it matters.
- Drop a leg. Does the coding agent really need a general-purpose fetch tool, or only your package registry?
- Restrict egress. Allow outbound requests only to named hosts. With the Hecate SDK guard,
guard.egress.allowblocks calls whose arguments point anywhere else. - Approve at runtime. Once a session has read untrusted content, require a human to approve any tool that can send data out. The Hecate SDK guard does this with
approveEgressAfterUntrusted, on by default.
Frequently asked questions
What is the lethal trifecta?
Simon Willison’s term for an AI agent that has access to private data, exposure to untrusted content, and a way to communicate externally. With all three, instructions hidden in untrusted content can make the agent collect private data and send it to an attacker.
Can it span several MCP servers?
Yes, and it usually does. Each server looks reasonable alone; the risk is in the agent that has all of them, which is why they have to be analysed together.
How do I break it?
Remove one leg for that agent: drop a group of tools, split the work across separate agents, or require human approval before any outbound tool runs once the session has read untrusted content.