Back to all articles
    MCP Prompt Injection: One Fooled Agent Can Mislead the Team
    AI SecurityPrompt InjectionMCPAI AgentsLeast Privilege

    MCP Prompt Injection: One Fooled Agent Can Mislead the Team

    Y

    Yoni Fraimorice

    Share:

    You give one agent a simple job: read some documentation. It sends a helpful summary to another agent. That agent asks a third to act.

    Nice teamwork. Unless the first agent also passed along an instruction planted by someone else.

    Now the attacker does not need to fool every agent directly. A trusted teammate can carry the message for them.

    An October 5 report from Unite.AI brings this risk into focus through security researcher Syed Anas Mohiuddin's work on protocol pivoting. His October research update connects an earlier cross-agent attack scenario with flaws found in real tools.

    One correction matters: the Google, JPMorgan and Rapid7 examples are not three confirmed outbreaks of agents infecting each other. They describe different tool-side vulnerabilities. Those weaknesses help explain what a misled agent could reach.

    How an instruction gets a trusted delivery address

    Prompt injection happens when an AI system treats untrusted content as instructions. The entry point might be a document, a web page, or a tool response.

    MCP, the Model Context Protocol, connects AI applications to tools and data. It is not the same thing as A2A, a protocol for communication between agents. A workflow can combine them, or use its own internal messaging.

    Mohiuddin describes this chain: an MCP tool returns attacker-controlled text shaped like a task. The coordinating agent passes it to a subagent as normal delegation. The subagent trusts the coordinator and acts using its own tools.

    The dangerous step is the promotion: source content becomes an authorized task.

    An already-compromised agent could start at the handoff instead. Here, "compromised" can mean its behavior was redirected by injected text; it does not require stolen model weights or malware.

    Imagine a documentation reader forwarding an instruction to upload a private diagnostic bundle to a supposed support service. The reader cannot upload anything. But the operations agent can. This is an illustrative example, not a reported incident at the companies below.

    Google: the tool could reach the wrong destination

    CVE-2026-14540 concerns Google's MCP Toolbox versions 0.3.0 through 1.4.0. Its HTTP client lacked restrictive redirect handling and destination-IP validation.

    An attacker-controlled path could cause the server to request an internal or arbitrary external endpoint. That is server-side request forgery, or SSRF: the server makes a request the caller should not control.

    Google's merged fix added an SSRF guard, connection-time address checks, configurable IP boundaries, and startup validation of the base URL. It shipped in v1.5.0.

    The lesson is bigger than "check the link." Check the actual destination, including redirects and DNS changes, where the connection happens.

    JPMorgan: one safe tool did not protect its sibling

    According to Mohiuddin, JPMorgan's documentation-search MCP component checked allowed domains in read_documentation, but its related() tool fetched a caller-supplied URL without that restriction.

    He says a change to an AWS-derived component introduced the fetch, and JPMorgan's disclosure team confirmed the finding and deployed a fix. I am attributing that outcome to his account, not claiming independent access to the bank's investigation.

    He also says no credential traveled with the forged request, limiting the impact.

    The useful question for your team is simple: did you protect every tool that fetches content, or just the one named "read"?

    Rapid7: a different injection, with a narrower impact

    Rapid7's own advisory for CVE-2026-97228 describes a GraphQL query injection in Bulk Export MCP versions 0.2.5 through 0.6.1.

    An unvalidated export identifier was inserted into query text. Version 0.6.2 passes it as a parameterized variable instead.

    This is not SSRF; Rapid7 rates it Low severity. Crucially, it says injected queries remain inside the operator's existing API scope and cannot cross an account or tenant boundary. Its record identifies a compromised upstream client or indirect prompt injection as the realistic exposure.

    So this is not evidence of unlimited access. It is evidence that even a familiar identifier needs validation before it reaches another language or API.

    The shared problem is trust, not a magic MCP worm

    These bugs do not prove MCP inevitably spreads attacks. They show why a valid tool call can still carry unsafe arguments.

    The cross-agent chain also depends on the application: which messages it forwards, which permissions the next agent has, and whether a separate control checks the action.

    MCP's client guidance explicitly says one server's tool results are untrusted input to another. Authentication answers who sent a message. It does not establish that every sentence inside is safe to obey.

    A checklist for every agent handoff

    These are my engineering recommendations, not claims that any one vendor has implemented the whole checklist.

    1. Give each agent its own permission budget

    Keep the documentation reader read-only. Give the operations agent only the resources and actions needed for its assigned task, not a shared administrator credential.

    At execution time, a broker should check the receiving agent's permissions and the user's authorized task. A message from the coordinator must not silently expand either.

    Use short-lived credentials held outside model context. Require explicit approval for sensitive writes, new destinations, or data transfers.

    2. Preserve where the message really came from

    Record the sending agent, original source, tool-call ID, parent task, timestamp, and granted scope. Keep that record through summaries and further delegation.

    Generate and verify this metadata outside the model. Do not accept a message's own claim that it came from an administrator.

    A signature can show who sent the content and whether it changed. It cannot prove the content is harmless. A signed message from a compromised agent is still dangerous.

    Keep retrieved text in a separate data field, not inside the next agent's system instructions.

    3. Filter outputs, then enforce the action boundary

    Allow only expected fields and types through handoffs. Remove unnecessary secrets and raw upstream error bodies. Flag content that tries to assign new roles, request credentials, or add unrelated tool calls.

    Keep the original, access-controlled evidence for investigation. Do not silently turn a suspicious response into a clean-looking instruction.

    Filtering is only one layer. Schemas do not make strings safe, and keyword filters miss paraphrases. The destination tool must still validate arguments, enforce network boundaries, and use parameterized queries.

    Google's redirect checks and Rapid7's query variables address different parts of that boundary.

    4. Test the whole relay, not only the first agent

    Use synthetic secrets and controlled destinations. Put a harmless test instruction in a tool result, pass it through a summary, and watch what the next agent tries.

    Measure whether unauthorized actions were blocked, not merely whether a warning appeared. Repeat with retries, context resets, and shared memory.

    Keep a trace across all agents and a stop control outside their permissions. If one agent is compromised, revoke its access and stop queued work that depends on its messages.

    A teammate is a sender, not an authority

    Multi-agent systems can make useful work faster. They can also make an untrusted instruction look more official at every step.

    Do not ask only, "Do I trust this agent?" Ask: "Where did this request originate, who authorized this action, and what can the receiver actually do?"

    That is how you keep collaboration from becoming a shortcut around your controls.

    Sources

    Hero photo: Offutt Air Force Base operator, a U.S. Air Force photograph of Sgt. Suzann K. Harry operating a switchboard in 1967; individual photographer not credited. Public domain as a U.S. federal government work. Resized from 2821 x 2149 to 1920 x 1462 pixels and JPEG-compressed; no other edits. This archival photograph illustrates relayed communication, not an AI system or any reported vulnerability. No endorsement is implied.

    Share: