Skip to main content

Prompt Injection

Learn how direct and indirect prompt injection can drive unsafe agent actions, and how least privilege and authorization reduce impact.

What is Prompt Injection?

Prompt injection is a vulnerability in which input changes a language model's behavior or output in an unintended way. A direct injection comes from a user prompt. An indirect injection comes from content that the system reads, such as a document, message, web page, or tool result. The model can treat untrusted content as an instruction because data and instructions share the same processing path.

Why it matters

An agent can turn model output into tool calls and real system changes. Prompt injection becomes more dangerous when the agent has broad credentials, powerful tools, or access to sensitive data.

How it works

  1. Untrusted input enters the model context through a user prompt, retrieved content, memory, or a tool result.
  2. The model interprets part of that input as an instruction and produces a response or selects a tool that the application did not intend.
  3. The application turns the model output into an effect. The impact depends on the tools, credentials, data, and network paths that the agent can reach.

Example

An agent reads a support ticket that contains a hidden instruction to export customer records. If the agent can call an unrestricted export tool, the injected text can cause a harmful action. A separate authorization check can deny the export even when the model requests it.

Pomerium boundary

Pomerium can limit which protected Streamable HTTP MCP servers and tools a caller can reach through identity, route, and mcp_tool policy. It can log authorization decisions and selected MCP fields. These controls can reduce blast radius, but they do not inspect model reasoning, sanitize prompts, or make tool code safe.

Limits and non-claims

  • Prompt filtering, retrieval-augmented generation, and model fine-tuning do not fully remove prompt injection risk.
  • A gateway cannot determine whether model reasoning or a prompt is safe.
  • Least privilege can reduce the impact of a successful injection, but it does not prevent model manipulation.
  • Human approval is useful only when the reviewer receives enough context and the application enforces the decision.

Evaluation checklist

  • List every source of text, images, documents, messages, memory, retrieval results, and tool output that can influence the model.
  • Separate model instructions from untrusted data and validate every model-selected action outside the model.
  • Give the agent only the tools, data, scopes, and network paths needed for the current task.
  • Require a deterministic authorization or human approval step for sensitive, irreversible, or high-impact actions.

Sources and further reading

Keep learning

Agentic AccessSecurity Operations and Risk

Agent Blast Radius

Agent blast radius is the maximum credible effect that an agent can cause through its tools, credentials, data access, network reach, and chained actions.

Learn this term
Agentic AccessSecurity Operations and Risk

Tool Surface Area

Tool surface area is the full set of operations, inputs, external resources, and privilege effects that tools make available to an agent.

Learn this term
Agentic AccessAuthorization and Policy

MCP Security

Model Context Protocol security is the set of controls that protects hosts, clients, servers, tools, authorization flows, and downstream resources.

Learn this term
Agentic AccessSecurity Operations and Risk

Hidden Trust Boundary

A hidden trust boundary exists when one component accepts another component's identity, authority, data, or result without an explicit enforced rule.

Learn this term
Authorization and Policy

Authorization

Authorization determines whether a subject can perform a requested operation on a resource. It evaluates policy after or alongside authentication.

Learn this term

Get a Personalized Demo

Schedule a Call with a Pomerium Engineer

Get a Demo