Skip to main content

Design access for resilience and recovery

Keep access controls safe through dependency failures, limit damage, recover service, and prove the restored state.

Learning outcomes

  • Map access dependencies and group them into independent failure domains.
  • Define safe denied, degraded, and emergency behavior for each protected resource class.
  • Limit the authority and duration of recovery mechanisms.
  • Verify recovery with evidence that covers route access and application outcomes.

System and failure domains

Draw the protected access system as dependencies and flows. Include name resolution, certificates, keys, clocks, identity provider, session state, policy, directory or group data, device context, decision service, enforcement proxy, network path, upstream, application permissions, data stores, logs, alerting, administration, backup, and recovery access.

Group components that can fail together. Shared cloud region, identity tenant, network, operator role, key store, deployment pipeline, or policy source can form one failure domain even when several products are present. Record capacity limits and overload paths as well as total outages.

For each protected resource, define the recovery objective. A public status page, ordinary employee tool, production administration service, and emergency control system can need different availability and safety behavior.

Resilience model

Use four connected strategies:

  1. Prevent: reduce the chance of failure or compromise through narrow authority, safe defaults, isolation, validation, and controlled change.
  2. Withstand: keep essential protection properties during a fault through bounded caches, independent components, load shedding, or limited operation.
  3. Limit damage: separate privileges and failure domains, stop lateral movement, expire authority, and preserve evidence.
  4. Recover: restore a known state, rotate affected credentials, reconcile policy and sessions, test the path, and review evidence.

Define invariants that must hold during recovery. Examples: no recovery credential is permanent; emergency actions remain attributable; direct upstream access does not become a routine path; restored policy has a known version; old sessions do not silently regain authority.

Degraded operation and recovery

For every dependency, choose and test one behavior:

  • Deny: stop new access when the required fact or decision is unavailable.
  • Bounded degraded operation: use a previously validated fact for a stated maximum time and limited resource class.
  • Emergency access: use a separate, narrow, time-bound process with independent approval and enhanced evidence.

Avoid a universal answer. Denial can be safe for production administration but harmful for an emergency service. Cached access can preserve availability but extend revoked authority. Emergency access can restore service but concentrate privilege.

Treat denial of service as an access-path failure, not only as a bandwidth problem. Bound request size, connection count, authentication work, policy work, queue depth, retry rate, and per-identity or per-resource consumption before scarce dependencies. Isolate public, employee, and administration routes where one overload could otherwise exhaust all access. Shed low-priority work and preserve a narrow recovery path. A rate limit reduces demand; it does not authorize a request or prove that shared identity, network, key, and control-plane dependencies can survive the remaining load.

Write a recovery sequence before the outage:

  1. Declare the incident and assign authority.
  2. Contain the affected failure domain.
  3. Select the approved degraded or emergency mode.
  4. Restore trusted identity, policy, keys, and enforcement state.
  5. Revoke or reconcile sessions and temporary grants.
  6. Test denial, allowed access, direct bypass, and application permission.
  7. Correlate evidence and return to normal mode.
  8. Remove temporary access and review the result.

Design tradeoffs and residual risk

Cached context improves availability but increases the stale-access window. Redundant components improve availability only when their failure domains are independent. A break-glass account can aid recovery but becomes a high-value target. More operational paths can reduce recovery time while making complete mediation harder to prove.

State residual risk as a bounded condition. For example: "During a directory outage, existing low-risk sessions can use cached group data for up to 15 minutes; production administration denies new requests." Record the owner, monitoring signal, and condition that ends the exception.

Pomerium boundary

Pomerium separates routing, authorization, authentication, and identity functions in its documented architecture. A deployment can use that structure to reason about dependencies and evidence, but its actual resilience depends on the chosen topology, configuration, identity provider, data stores, network, and operating procedures.

Pomerium can deny protected route access when required policy conditions are not met and can record route decisions. It cannot make the upstream available, recover application state, or prevent a separately exposed upstream path. Application sessions and permissions can also need explicit reconciliation after recovery.

Exercise

Select one ordinary application and one high-impact administration application. For each dependency in their access paths, record:

  • Failure domain and owner.
  • Detection signal.
  • Denied, degraded, or emergency behavior.
  • Maximum duration and authority.
  • Recovery action.
  • Negative test.
  • Evidence that normal protection is restored.

Run a tabletop scenario in which the identity directory is unavailable and a production change is urgent. Then run a second scenario in which the policy source is suspected compromised. The correct response should differ because the failure and trust conditions differ.

Evaluation checklist

  • Are shared failure domains visible across identity, policy, network, keys, and operators?
  • Do overload tests cover shared bottlenecks and prove that one route, tenant, or identity cannot exhaust every protected path?
  • Does each resource class have an explicit safe failure behavior?
  • Are cache duration, scope, invalidation, and stale-access risk bounded?
  • Is emergency authority narrow, independently approved, expiring, and attributable?
  • Does recovery revoke temporary access and reconcile old sessions?
  • Do tests cover denial, allowed access, direct bypass, and application action permission?
  • Can evidence show when degraded mode started, what it allowed, and when normal protection returned?

Next learning unit

Sources and further reading

Keep learning

Security Engineering FoundationsSecurity Operations and Risk

Security Risk

Connect a credible threat, likelihood, consequence, uncertainty, and stakeholder impact to an explicit risk decision.

Learn this term
Security Engineering FoundationsAuthorization and Policy

Separation of Privilege

Require independent conditions, authorities, or actors before the system permits a sensitive action.

Learn this term
Authorization and PolicySecurity Operations and Risk

Authorization Drift

Authorization drift is the gap that develops when effective access no longer matches intended access.

Learn this term

Get a Personalized Demo

Schedule a Call with a Pomerium Engineer

Get a Demo