Threat definition
A rogue agent is an agent process or identity that performs autonomous behavior outside the declared goal, approved policy, expected identity chain, or operating boundary. The cause can be compromise, goal hijack, malicious code, unsafe persistence, stolen credentials, or deliberate misuse.
Assets and preconditions
Assets include tool authority, credentials, data, queues, memory, target systems, control planes, and trust in agent identity. Preconditions include reachable tools, durable credentials, hidden execution, weak inventory, broad egress, or missing action evidence.
Detection and containment
Compare observed agent identity, host, version, goal, tool, resource, action, rate, and target with an approved inventory and policy. Contain the narrowest reliable identity, credential, connection, tool, route, queue, and target account. Check for copied credentials and scheduled work.
Failure and residual risk
Valid autonomous work can look unusual. A rogue process can use a valid identity and low-volume actions. Central kill controls can fail or become denial-of-service targets. Route revocation cannot reverse completed actions.
Pomerium boundary
Pomerium can authenticate and authorize routed MCP requests and deny future route use. It cannot prove an agent goal, stop local stdio tools, kill the runtime, cancel queued work, or reverse a target action. Operators need host, orchestrator, tool, and target controls.
Evaluation checklist
- Which inventory entry and human or workload authority justify the running agent?
- Which goal, tool, resource, action, rate, and target are allowed?
- Can behavior outside that envelope be detected through independent evidence?
- Can identity, credentials, routes, tools, queues, and target accounts be contained?
- Which completed effects require compensation or recovery?
