Learning outcomes
- Define recovery objectives for access decisions, control state, and protected services.
- Restore from known-good configuration and trust instead of copying compromised authority.
- Operate safe emergency access when normal identity or control planes fail.
- Prove positive service recovery and continued rejection of old authority.
System and failure domains
An access system depends on identity providers, authenticators and recovery, signing and encryption keys, certificates, policy source, configuration delivery, control-plane state, data-plane enforcement, DNS, time, context sources, evidence pipelines, upstream applications, and operator access. Recovery must state which domains failed or were compromised and which evidence establishes a clean boundary.
Define recovery objectives for authorization availability, policy freshness, session invalidation, credential rotation, route restoration, evidence continuity, and protected application service. A fast process restart is not recovery when the same compromised key, policy, or account remains trusted.
Resilience model
Keep immutable known-good configuration and artifacts, independently protected trust backups, tested key and certificate replacement, alternate identity and operator paths, dependency inventories, and documented rebuild order. Separate backup access from normal administrative identity. Require more than one person or mechanism for high-impact recovery authority.
State safe behavior for identity-provider outage, policy service outage, context-source timeout, clock fault, certificate expiry, signing-key exposure, control-plane loss, data-plane isolation, and evidence loss. Some resources must fail closed. A narrowly defined emergency service can remain available through independent, short-lived break-glass access.
Recovery evidence includes artifact digests, source revision, trust roots, key generation and activation, issuer metadata, policy version, operator identity, service health, positive requests, negative requests with revoked authority, direct-origin rejection, and log continuity.
Degraded operation and recovery
- Declare the trusted recovery authority and identify systems that must not be reused.
- Preserve evidence and isolate compromised control, credentials, routes, and accounts.
- Restore foundational dependencies in documented order: time, naming, trust, identity, configuration, policy, enforcement, evidence, and applications.
- Generate new trust material when old custody is uncertain. Do not restore an exposed key from backup.
- Deploy the exact reviewed known-good artifact and configuration. Reconcile emergency changes explicitly.
- Start with narrow routes and identities. Verify component health and end-to-end access.
- Test that revoked sessions, old credentials, old certificates, direct paths, and unauthorized subjects fail.
- Reconnect dependencies and expand service in bounded stages.
- Retire emergency access, rotate its credentials, preserve its evidence, and review its use.
- Monitor recurrence and update recovery objectives, inventories, controls, and exercises.
Design tradeoffs and residual risk
Cached policy and sessions improve continuity and preserve stale or compromised authority. Hard dependency on a central service improves control and can stop all access. Define bounded degraded modes by resource sensitivity instead of one global behavior.
Restoring from backup is fast and can reintroduce compromised policy, credentials, or persistence. Rebuilding from source improves confidence and depends on a trusted build and delivery chain. Independent break-glass access aids recovery and creates a high-value bypass.
Residual risk includes hidden dependencies, incomplete key exposure scope, compromised backups, inconsistent replicas, direct origins, and operator error under time pressure.
Pomerium boundary
Pomerium can load configured routes and policy, expose health, and enforce access when its dependencies and trust are available. Operators own independent backups, key custody, identity recovery, DNS and network restoration, application state, break-glass design, and proof that old authority fails. A healthy Pomerium component does not prove end-to-end service or trustworthy configuration.
Exercise
Disable the normal administrator identity and simulate loss of the access control plane in staging. Use an independently held, time-bound break-glass path to deploy one exact known-good configuration and new test trust material. Restore one critical route before lower-priority routes.
Verify a normal approved request succeeds. Verify the disabled administrator, old session, old service credential, old certificate, unauthorized subject, and direct-origin request fail. Expire break-glass access, rotate its credential, and confirm its evidence is complete.
Evaluation checklist
- Are recovery objectives defined for decisions, trust, evidence, and protected service?
- Are known-good source, artifacts, backups, keys, and operator authority independently protected?
- Does degraded behavior match resource sensitivity and dependency failure?
- Does recovery prove both approved access and rejection of old authority?
- Is break-glass access narrow, independent, tested, expiring, and reviewed?
Next learning unit
Security Assurance and Evidence
Distinguish a control claim, verification, validation, assurance argument, and the evidence that supports each conclusion.
