Threat
Agent goal hijack occurs when untrusted content changes how an agent interprets its task or selects actions. The content can come from a user message, web page, document, tool output, memory, resource, prompt, or another agent. The attacker tries to make data act as a higher-priority instruction.
Prompt injection is one mechanism. Goal hijack names the system effect: the agent pursues an attacker-selected objective, target, or sequence with authority that the attacker could not use directly.
Control the action boundary
Treat external content and tool results as untrusted data. Preserve source and trust labels through retrieval and memory. Keep system policy and authorization outside model-generated text. Give the agent only tools and credentials needed for the current task. Validate tool name, resource, action, arguments, destination, and result at a deterministic enforcement point.
Require an independent human or policy approval for high-impact actions after the exact action is known. The approval must show the real target and effect, not an agent-written summary.
Failure and residual risk
Instruction filtering and model training do not create a complete security boundary. Attackers can encode instructions, split them across sources, poison a trusted document, or influence a tool result. A tool allowlist still permits harmful use of an allowed tool. Memory can preserve the hijack after the original input disappears.
An agent can also make an unsafe plan without malicious input. Deterministic authorization and bounded credentials remain necessary.
Pomerium boundary
Pomerium can authenticate users and protect agent or MCP routes. It can enforce route policy before a tool server receives a request. It does not interpret prompt intent or prove that a model followed the user's goal. Agent hosts and tool servers must validate actions and content-derived arguments.
Evaluation checklist
- Are retrieved content, tool output, memory, and user instructions labeled as untrusted inputs?
- Does deterministic policy validate each concrete resource and action?
- Can the agent use only task-specific tools and credentials?
- Does high-impact approval show the exact target, arguments, and effect?
- Can injected instructions in a document, tool result, and memory fail without an external action?
