Prompt Injection is Not an AI Problem: Why MCP Tool Hardening Matters
Authors: Shiqiang Chen
Abstract: Prompt injection has been widely framed as a language model safety problem, with solutions focused on input filtering, alignment training, and guardrail models. This paper argues that this framing is fundamentally incomplete. Through an analysis of MCP tool interactions, we demonstrate that prompt injection is primarily a tool security and input sanitization problem, not an AI alignment one. We show how seemingly benign tool integrations create injection vectors that bypass LLaMA-based guardrails, and we propose a principle-based approach to MCP tool hardening: validate at the boundary, not the model.
Key Contributions
- Reframing prompt injection as a tool security problem, not an AI alignment problem
- Analysis of MCP tool injection vectors that bypass LLaMA guardrails
- Principle-based approach to MCP tool hardening (validate at the boundary)
- Practical recommendations for tool developers and protocol designers
Paper
The full paper is available as paper-prompt-injection.pdf
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support