Jinwoo Park, Geonhee Kim, H Lee, Jeman Park
The Model Context Protocol (MCP) has become the de facto standard for connecting large language models (LLMs) to external tools, and its remote deployment mode lets users add third-party servers with a single URL—shifting a substantial portion of the host’s attack surface to infrastructure operated by anonymous parties. Existing MCP security work has concentrated on tool-description poisoning and studied individual techniques in isolation, leaving it unclear what a malicious remote server can accomplish across its full surface. In this paper, we explore the malicious-server threat space along the axis of whether the host LLM participates in producing the harmful outcome, yielding two categories: LLM-passive attacks, which complete inside the server, and LLM-active attacks, which require the LLM to deliver the malicious content. We implement five scenarios spanning both categories—realizing each LLM-active scenario with both description-based and response-based variants against the same goal—and evaluate all configurations on ChatGPT, Claude Desktop, and Gemini CLI. We find that host-side filtering of MCP-bound data varies sharply across platforms (95% vs. 50% ASR on the same email request), that the description and response channels succeed on disjoint scenarios, and that successful attacks are almost never disclosed to the user. These findings suggest that defending remote MCP deployment requires a multi-layer approach combining host-side filtering, LLM-level response auditing, and user-visible output transparency.