Learning outcomes
- Build one end-to-end evidence chain for an agent action.
- Separate proposed, approved, attempted, allowed, executed, and completed states.
- Revoke sessions, credentials, grants, tools, routes, and queued work in the correct order.
- Measure the last accepted action after a containment event.
Operating objective
An investigator can answer who initiated an agent task, which agent and services acted, what authority they used, which tool and resource received each action, which policy and approval allowed it, what external effect occurred, and when future authority stopped. Operators can contain one user, agent, tool, credential, route, or task without disabling every unrelated workflow.
Define action states precisely: proposed, policy denied, approval requested, approved, credential selected, tool attempted, tool allowed, downstream accepted, effect completed, result returned, compensated, and revoked. A successful tool response does not always prove the business effect completed. A model statement does not prove any state.
Signals and evidence
Use a correlation identifier across host, MCP client, gateway, server, credential broker, tool, downstream API, queue, and application. Record human subject, agent and host actor, tenant, task, tool, normalized resource and action, policy version and reason, approval digest and approver, credential fingerprint and audience, request time, downstream result, side-effect identifier, and revocation state.
Propagate W3C Trace Context where the current MCP revision and deployment support it. Keep security identifiers outside model-controlled text. Redact tokens, cookies, secrets, full prompts, and unnecessary personal data. Protect log integrity and access. Define retention from the investigation need.
Response and recovery
- Stop new task planning and tool selection for the affected subject, agent, or tenant.
- Revoke or disconnect delegated grants, refresh tokens, sessions, service credentials, and active tool connections in scope.
- Deny the tool or route and disable the compromised downstream identity.
- Cancel queued, scheduled, and long-running work. A route denial does not remove already queued effects.
- Query target systems for completed actions. Do not infer completion only from agent logs.
- Compensate reversible actions through separately authorized operations. Preserve original evidence.
- Rotate credentials or trust material that could have leaked.
- Restore from known-good tool, prompt, policy, memory, and dependency state.
- Reauthorize users and tools through new grants after cause and scope are understood.
- Measure the source event to last accepted action and compare it with the stated revocation bound.
Design tradeoffs and residual risk
Full prompts and tool results improve semantic investigation and create a sensitive secondary data store. Minimal structured evidence reduces exposure and can omit the content that caused goal hijack. Use tiered retention and controlled escalation for content capture.
Central kill controls improve containment and create high-authority availability dependencies. Local circuit breakers can continue during control-plane loss and require later reconciliation. Compensation can create new effects and needs independent approval.
Residual risk includes unobserved local tools, copied credentials, external actions with no rollback, agents operating outside the orchestrator, stale caches, and compromised evidence sources.
Pomerium boundary
Pomerium can log routed access decisions and deny future route use through policy or credential changes. It cannot cancel tasks already queued inside an agent or reverse a downstream action. Operators must correlate Pomerium records with agent, MCP server, broker, tool, and target evidence.
Exercise
Run one staged agent action that creates and then updates a test resource. Use a correlation identifier across every component. Verify that the evidence distinguishes proposal, approval, tool execution, and final effect.
Then compromise the test credential or tool. Trigger containment and send a request every second. Measure the last accepted tool and target action. Confirm queued work stops, old credentials fail, and restoration uses clean state.
Evaluation checklist
- Can one action be traced from human task to final target effect?
- Are proposal, approval, attempt, allow, execution, completion, and compensation distinct?
- Can containment target one subject, agent, tool, route, credential, or task?
- Are queued and long-running actions included in the kill path?
- Is revocation latency measured from event to the last accepted external action?
Next learning unit
Revocation Latency
Measure how long a disabled identity, authenticator, session, claim, or permission can continue to authorize action.
