Learning outcomes
- Define recovery time, recovery point, integrity, trust, and access objectives for each dependency.
- Restore consistent data and configuration in an isolated clean environment.
- Rebuild software and infrastructure from verified source and artifacts when backup state is not trustworthy.
- Prove approved access works and revoked or compromised authority remains rejected.
System and failure domains
List each state needed to recover the access service: identity configuration, authenticators and recovery, policy, routes, certificates, signing and encryption keys, trust roots, application data, schemas, infrastructure definitions, software artifacts, dependency versions, DNS, time, evidence, and operator authority. Identify which stores must be consistent and which failures can corrupt every replica or backup.
Define recovery time objective, recovery point objective, maximum tolerable outage, data-integrity objective, and required trust confidence for each service tier. A fast restart from compromised state does not meet the recovery objective.
Resilience model
Use independent backup identities and protect backup consoles from normal production administration. Keep immutable, offline, or logically isolated copies where the threat requires them. Preserve data with its schema, configuration, key references, deletion state, and recovery instructions. Keep verified source, signed artifacts, provenance, dependency inventory, and infrastructure definitions for a clean rebuild.
Distinguish restore from rebuild. Restore recovers selected state from a copy. Rebuild creates infrastructure and software from trusted definitions and artifacts. After credential, host, build, or administrative compromise, a clean rebuild can be required even when backups are available.
Degraded operation and recovery
- Select one realistic failure and state the trusted recovery authority.
- Freeze the selected recovery point and verify the copies, keys, artifacts, and instructions are accessible without compromised production identity.
- Start an isolated environment with independent networking, credentials, and evidence storage.
- Rebuild infrastructure and software from verified definitions and artifacts.
- Restore only required data and configuration. Reconcile transactions after the recovery point.
- Replace exposed keys, certificates, sessions, tokens, accounts, and recovery authority. Do not restore an exposed private key.
- Apply current retention, deletion, schema, policy, and vulnerability requirements before service starts.
- Verify component health, end-to-end approved access, and required degraded behavior.
- Prove old sessions, credentials, keys, removed accounts, vulnerable artifacts, direct paths, and unauthorized subjects fail.
- Measure actual recovery time and data loss. Record bottlenecks, manual steps, missing dependencies, and trust gaps.
Design tradeoffs and residual risk
Frequent copies reduce data loss and can replicate corruption quickly. Long retention increases recovery options and preserves expired or deleted personal data. Isolated backups resist ransomware and can be slower to restore. Full environment tests improve confidence and cost time, infrastructure, licenses, and coordinated labor.
Residual risk includes unknown dependencies, compromised backup administration, inconsistent stores, unrecoverable encryption keys, unverified third-party state, malicious persistence in data, and a clean rebuild that retrieves a compromised current dependency.
Pomerium boundary
Pomerium uses configuration, policy, routes, certificates, keys, identity-provider trust, DNS, and upstream applications supplied by the deployment. Operators own backups, recovery authority, source and artifact verification, infrastructure reconstruction, identity recovery, application data, and end-to-end tests. Pomerium health alone does not prove that policy is correct or that old authority fails.
Exercise
Choose one sensitive route. Simulate loss of its policy store and compromise of the normal deployment identity. Recover in isolation with an independent operator. Rebuild the runtime, restore the selected configuration, create replacement trust material, and connect a representative upstream.
Measure recovery time and data loss. Prove one approved request succeeds. Prove the old session, old service credential, old certificate, removed subject, direct-origin path, and unauthorized object action fail. Then document the longest dependency and one step that lacked reproducible evidence.
Evaluation checklist
- Are time, data-loss, integrity, trust, and service objectives explicit for each dependency?
- Can recovery proceed without normal production identity, infrastructure, and control-plane state?
- Does the test distinguish a data restore from a verified clean rebuild?
- Are restored data, schema, policy, keys, retention, deletion, and software mutually consistent?
- Do positive and negative tests prove current service and continued rejection of old authority?
Next learning unit
Security Assurance and Evidence
Distinguish a control claim, verification, validation, assurance argument, and the evidence that supports each conclusion.
