What is Prompt Injection?
Prompt injection is a vulnerability in which input changes a language model's behavior or output in an unintended way. A direct injection comes from a user prompt. An indirect injection comes from content that the system reads, such as a document, message, web page, or tool result. The model can treat untrusted content as an instruction because data and instructions share the same processing path.
Why it matters
An agent can turn model output into tool calls and real system changes. Prompt injection becomes more dangerous when the agent has broad credentials, powerful tools, or access to sensitive data.
How it works
- Untrusted input enters the model context through a user prompt, retrieved content, memory, or a tool result.
- The model interprets part of that input as an instruction and produces a response or selects a tool that the application did not intend.
- The application turns the model output into an effect. The impact depends on the tools, credentials, data, and network paths that the agent can reach.
Example
An agent reads a support ticket that contains a hidden instruction to export customer records. If the agent can call an unrestricted export tool, the injected text can cause a harmful action. A separate authorization check can deny the export even when the model requests it.
Pomerium boundary
Pomerium can limit which protected Streamable HTTP MCP servers and tools a caller can reach through identity, route, and mcp_tool policy. It can log authorization decisions and selected MCP fields. These controls can reduce blast radius, but they do not inspect model reasoning, sanitize prompts, or make tool code safe.
Limits and non-claims
- Prompt filtering, retrieval-augmented generation, and model fine-tuning do not fully remove prompt injection risk.
- A gateway cannot determine whether model reasoning or a prompt is safe.
- Least privilege can reduce the impact of a successful injection, but it does not prevent model manipulation.
- Human approval is useful only when the reviewer receives enough context and the application enforces the decision.
Evaluation checklist
- List every source of text, images, documents, messages, memory, retrieval results, and tool output that can influence the model.
- Separate model instructions from untrusted data and validate every model-selected action outside the model.
- Give the agent only the tools, data, scopes, and network paths needed for the current task.
- Require a deterministic authorization or human approval step for sensitive, irreversible, or high-impact actions.
