Learning outcomes
- Model failures in identity, policy, context, distribution, enforcement, application, and evidence dependencies.
- Choose explicit deny, cached, degraded, or emergency behavior by protected action.
- Bound stale authority and preserve recovery access without hidden fail-open paths.
- Test outage, recovery, rollback, and end-to-end revocation behavior.
System and failure domains
Map the identity provider, session store, policy administration, policy distribution, context sources, decision service, enforcement points, DNS, certificate and key services, network, upstream application, data stores, evidence pipeline, and recovery controls. For each, define failure detection, dependency, blast radius, and protected actions.
Separate unavailable, slow, stale, corrupt, inconsistent, and compromised states. A responding service can still return unsafe data.
Resilience model
Choose behavior per action:
- Fail closed: Deny when a required decision or fact is unavailable.
- Cached operation: Use a known result or policy for a bounded time and resource set.
- Degraded access: Permit a smaller, read-only, local, or lower-risk function.
- Fail open: Permit despite missing control. Use only for an explicit protection and availability decision with strict bounds.
- Break glass: Use a separate, time-bound, monitored emergency authority.
Define the maximum cache age, allowed actions, principals, recovery owner, evidence, and automatic exit for every non-normal state.
Degraded operation and recovery
Detect failure without relying only on the failed component. Announce the active mode to operators. Prevent silent fallback. During recovery, verify dependency integrity, restore policy and context versions, expire emergency and cached grants, terminate sessions that exceeded limits, and prove that every enforcer returned to normal behavior.
Exercise recovery from partial rollout and inconsistent replicas, not only complete outage.
Design tradeoffs and residual risk
Strict denial protects confidentiality and integrity but can block safety and recovery work. Cached policy improves availability but preserves revoked authority. Degraded read-only access can still disclose sensitive data. Break-glass access reduces dependency but creates powerful dormant authority.
Document who accepted each residual risk and which signal ends the exception.
Pomerium boundary
Pomerium enforces route policy using its available identity, session, configuration, and context dependencies. Operators must design availability, cached state, and emergency access for their deployment. Upstream applications must define their own failure behavior for object and action authorization.
Exercise
For one administrative application, make a matrix of identity-provider outage, policy-input outage, decision latency, stale policy, lost evidence, upstream outage, and compromised context. Set behavior for read, write, approve, and emergency recovery.
Run two game-day cases: a clean outage and a context source that returns plausible but stale data. Verify entry to degraded mode, bounds, evidence, restoration, cache removal, and revocation.
Evaluation checklist
- Does every dependency have behavior for unavailable, stale, corrupt, and inconsistent states?
- Is non-normal access limited by person, resource, action, time, and evidence?
- Can operators identify the active mode and exit it automatically or deliberately?
- Does recovery remove cached, emergency, and stale authority from every enforcer?
- Have outage and recovery tests measured availability and unauthorized-access risk together?
Next learning unit
Fail-Safe Defaults
Start from explicit denial and define safe behavior for missing policy, invalid input, dependency failure, and recovery.
