Skip to main content

Design human approval for agent actions

Present the real target, arguments, authority, effect, and uncertainty so human approval gates a concrete agent action.

Learning outcomes

  • Choose which agent actions require approval based on impact and reversibility.
  • Bind approval to normalized action data instead of an agent summary.
  • Prevent stale, replayed, bundled, coerced, and post-change approval.
  • Measure approval quality without turning the user into the only security control.

Protection need

An agent can prepare an action faster than a person can inspect its full context. High-impact or irreversible effects need an independent decision by a person with enough information and authority. Approval must gate a concrete action, not transfer responsibility for a vague plan to the user.

Classify actions by resource sensitivity, affected subjects, reversibility, external communication, money, privilege, credential use, data disclosure, scope, and automation depth. Low-risk read actions can proceed under policy. High-risk write, delete, publish, transfer, permission, and execution actions can require approval or separation of duties.

Security objectives and requirements

  • Show the exact tool, operation, target identity, tenant, resource, normalized arguments, data destination, credential identity, and expected effect.
  • Show material uncertainty and which fields came from untrusted content.
  • Get approval from a currently authenticated person with the required role.
  • Bind the approval cryptographically or transactionally to the action digest, policy version, approver, task, and short expiry.
  • Invalidate approval if any material argument, target, credential, or policy changes.
  • Consume single-use approval atomically.
  • Keep independent authorization at the tool and target.
  • Make denial, timeout, cancel, and escalation safe.

The model can explain the action, but the approval view must render security facts from normalized tool data and trusted resource lookups.

Security invariants and evidence

No approval is reusable for another task or action. The executor compares the exact action digest before use. The approver cannot approve an action beyond their own authority. For separation of duties, requester and approver are distinct under verified identities.

Evidence records action digest, displayed facts, approver, decision, time, policy, expiry, execution identity, and final effect. Do not store the approval page as a substitute for the target's own audit event.

Failure cases

  • The UI displays an agent summary while hidden arguments target another account.
  • One click approves a bundle with unknown size or future actions.
  • The agent changes arguments after approval.
  • An old approval is replayed after the resource state changes.
  • Approval fatigue trains users to accept repetitive prompts.
  • A malicious page overlays or forges the approval interface.
  • The approver lacks resource permission but approval is treated as authorization.
  • A denied action is retried through another tool or service credential.
  • Emergency approval has no expiry or review.

Test time-of-check to time-of-use changes. Rename, delete, or alter the target after display and before execution. The system must revalidate or stop.

Design tradeoffs and residual risk

Frequent prompts reduce automation value and cause habituation. Rare prompts can miss high-impact combinations. Risk-based approval improves focus and depends on accurate classification. Two-person approval limits insider and account-compromise risk and adds delay.

Previewing every field improves transparency and can overwhelm users. Use a stable summary with expandable raw normalized data and clear differences from prior state. Never hide recipients, permissions, or destructive scope.

Residual risk includes coerced or compromised approvers, misleading but technically accurate data, valid approval for a harmful objective, and target changes after execution. Approval complements least privilege and policy. It does not replace them.

Pomerium boundary

Pomerium can authenticate the approver and protect approval and tool routes. It does not generate the action preview or bind approval to tool arguments by default. The agent platform and tool server own action normalization, approval state, replay prevention, and final authorization.

Exercise

Design approval for deleting cloud resources and sending an external message. Define the risk rule, displayed fields, action digest, approver requirement, expiry, state revalidation, cancellation, evidence, and target authorization.

Test changed recipient, added resource, increased amount, changed credential, expired approval, duplicate use, another approver session, interface overlay, and target-state change. Confirm that the exact unapproved action cannot execute.

Evaluation checklist

  • Does the approver see trusted normalized facts rather than only model prose?
  • Is approval bound to one action digest, task, approver, policy, and short lifetime?
  • Does any material change require a new decision?
  • Do tool and target authorization remain independent?
  • Do fatigue, replay, bundling, clickjacking, and state-change tests fail safely?

Next learning unit

Authorization Consent

Use consent to record a user's informed grant without treating it as proof that an action is safe or permitted.

Sources and further reading

Keep learning

Agentic AccessAuthorization and Policy

Delegation Chain

Preserve the human actor, agent, service, tool, target, authority, and constraints through every delegated access step.

Learn this term
Authorization and Policy

Separation of Duties

Split incompatible authority across independent people or roles so one actor cannot complete a sensitive process alone.

Learn this term

Get a Personalized Demo

Schedule a Call with a Pomerium Engineer

Get a Demo