MCP Trust Gaps Let One AI Agent Spread Malicious Instructions to Others, Researcher Finds
Ars Technica reports Google and four other organizations acknowledged AI-agent prompt-injection flaws that exploit MCP trust gaps.
The attacks are a special form of prompt injection, according to the report. They do not target a large language model directly. Instead, they target a particular agent, such as one used for translation or data analysis. Guardrails inside such agents, if they exist at all, are often lax and will send the instructions to other agents down the chain. Because the later agent explicitly trusts the first one, it follows the directions. Ars Technica described MCP for agent-to-agent communications as possibly the riskiest protocol many people have never heard of.
Independent researcher Syed Anas Mohiuddin tested agents from organizations including Google, JP Morgan Chase, Weviate, Rapid7, the French government's interministerial digital directorate, and the US federal government. His proof-of-concept attacks exploit trust gaps in MCP. The article said the standard is one way AI apps and agents communicate with each other inside an internal network. An illustration accompanying the article showed a simplified MCP in action.
The report said the vulnerabilities are unexpected and hard to mitigate. It linked the risk to the adoption of AI agents in millions of organizations, which is creating new opportunities for attackers to make agents take malicious actions. The organizations that acknowledged the vulnerabilities had little in common except for their use of AI agents, Ars Technica reported.
The reported concern is that an agent receiving instructions from a trusted internal agent may not apply the same scrutiny it would apply to an outside request. The article did not name all of the four other organizations that acknowledged the vulnerabilities.