Skip to main content

Agent Cascading Failure

Stop one false result, repeated action, or unavailable dependency from propagating through agents and tools with growing impact.

Threat

An agent cascading failure starts when one bad input, decision, result, retry, credential, or unavailable dependency triggers further agents and tools. Each step can amplify scope, cost, authority, or irreversibility. A false inventory result can cause bulk deletion. One timeout can cause many retries and duplicate actions.

The failure can be malicious or accidental. The control problem is propagation and blast radius.

Limit propagation

Define task budgets for calls, time, money, data, recipients, resources, concurrency, and depth. Use idempotency keys and transaction boundaries for retries. Validate outputs before they become another agent's input. Require independent confirmation for state transitions with large or irreversible effect.

Add circuit breakers and bulkheads by tenant, tool, resource, and dependency. Use bounded queues and backpressure. Stop downstream action when required identity, policy, evidence, or context is missing. Provide a kill path that revokes future credentials and cancels queued work.

Failure and residual risk

Agents can coordinate through systems outside one orchestrator. A local call limit can be bypassed through several tools. A circuit breaker can preserve availability while leaving partial state. Compensation can itself repeat or fail. Human responders can trust a polished but wrong summary.

Some actions cannot be undone. Use prevention, narrow credentials, staged effects, and canaries before relying on rollback.

Pomerium boundary

Pomerium can deny future routed requests after policy or credential changes. It does not cancel work already queued inside agents or downstream services. Agent platforms own budgets, idempotency, circuit breaking, kill controls, compensation, and state reconciliation.

Evaluation checklist

  • Are call, depth, time, cost, data, and resource budgets enforced outside the model?
  • Are retries idempotent or protected from duplicate external effects?
  • Can one tenant, tool, or dependency failure be isolated?
  • Can operators revoke future authority and cancel queued work quickly?
  • Do tests cover false output, retry storm, partial completion, dependency outage, and kill recovery?

Sources and further reading

Keep learning

Agentic AccessSecurity Operations and Risk

Agent Blast Radius

Agent blast radius is the maximum credible effect that an agent can cause through its tools, credentials, data access, network reach, and chained actions.

Learn this term
Security Engineering FoundationsSecurity Operations and Risk

Defense in Depth

Place complementary controls across distinct failure domains so one failure does not expose the protected asset.

Learn this term
Security Engineering FoundationsAuthorization and Policy

Fail-Safe Defaults

Start from explicit denial and define safe behavior for missing policy, invalid input, dependency failure, and recovery.

Learn this term

Get a Personalized Demo

Schedule a Call with a Pomerium Engineer

Get a Demo