Learning outcomes
- Derive a minimal runtime baseline from service function and threat model.
- Remove unnecessary code, interfaces, privileges, credentials, and network paths.
- Detect drift and patch the complete service stack through a tested rollout.
- Recover from compromise into a known-good state without preserving attacker control.
Operating objective
Define the service role, required availability, protected data and keys, network peers, administrators, deployment environment, and credible attackers. Example: an identity-aware proxy accepts public TLS, reads approved policy, calls an identity provider and authorization service, reaches a fixed upstream set, and emits audit evidence. It does not need an interactive shell, package manager, compiler, runtime socket, host filesystem, cloud administration role, or unrestricted egress.
Create a baseline for:
- Exact artifact, image, operating system, runtime, libraries, and provenance.
- Process identity, groups, capabilities, syscalls, files, mounts, and devices.
- Listeners, outbound destinations, DNS, proxy behavior, and metadata access.
- Keys, certificates, service credentials, configuration, and secret delivery.
- CPU, memory, process, storage, file-descriptor, connection, and request limits.
- Logs, metrics, traces, time, health, alerts, and evidence retention.
- Administrator, deployer, support, break-glass, update, rollback, and recovery authority.
Build and deployment baseline
Use a minimal supported base and a repeatable build. Pin dependencies and verify provenance. Remove development tools, sample files, interpreters, unused packages, default accounts, and unused services. Keep the application artifact immutable at runtime.
Run as a dedicated non-root identity. Remove ambient capabilities. Make the root filesystem read-only when the service supports it and mount only required writable state. Deny access to runtime sockets, host namespaces, debug interfaces, device nodes, and unrelated secrets. Use separate credentials for build, deploy, runtime, monitoring, and recovery.
Allow only required ingress and egress. Protect the cloud metadata service. Bind management listeners to a separate protected path. Use explicit upstream destinations and TLS identity. Place resource limits at both workload and host or cluster levels.
Signals and evidence
Record the expected baseline in code. At build and deploy time, verify artifact digest, signature or provenance, dependency inventory, configuration schema, policy version, and service identity. At runtime, observe:
- New processes, executable files, packages, kernel modules, accounts, and scheduled work.
- Changed capabilities, mounts, permissions, network listeners, egress, and DNS.
- Credential, key, policy, and certificate use outside the expected pattern.
- Failed integrity checks, blocked system calls, resource pressure, crashes, restarts, and degraded dependencies.
- Administrator, debug, shell, deployment, rollback, and recovery actions.
Alert on meaningful drift with an owner and response. Do not collect secret values or full sensitive requests as evidence.
Update and drift control
Patch the application, libraries, runtime, operating system, firmware, and platform components according to risk and exposure. Test normal function, denial paths, resource limits, observability, rollback, and data compatibility. Roll forward with immutable artifacts. Keep rollback bounded so it cannot restore a known vulnerable release or obsolete policy indefinitely.
Rebuild a drifted instance from known-good inputs. Do not treat manual repair as restored assurance. Investigate why drift occurred and whether credentials or evidence were exposed.
Response and recovery
When compromise is suspected, isolate the workload, preserve needed evidence, revoke workload and administrator credentials, rotate affected keys, block direct paths, and replace the instance. Validate the build source, deployment identity, configuration, policy, and recovery data before restart.
Restore into a clean environment. Do not attach compromised volumes or copy unknown binaries into the new runtime. Verify current policy, account lifecycle, certificates, routes, logging, alerts, and upstream boundary. Exercise a recovery that assumes the old host and its credentials are hostile.
Design tradeoffs and residual risk
Minimal images reduce packages and can reduce debugging tools. Poor observability can delay response. Read-only filesystems and syscall restrictions can break upgrades or diagnostic work. Very narrow egress and immutable deployment require accurate dependency and recovery design.
The baseline does not correct application logic flaws, supply-chain compromise before provenance, hypervisor or firmware compromise, unsafe administrators, or broad upstream authorization. Document the remaining TCB and failure domains.
Pomerium boundary
Pomerium provides documented deployment and upgrade behavior. Operators choose its host, image, runtime identity, privileges, network policy, credentials, log destination, administrator access, patch cadence, and recovery process. Protecting a management UI through Pomerium does not limit the workload's own host or cloud authority.
Exercise
Harden one gateway or identity service in a test environment. Produce the role baseline and an exception list. Remove one package, capability, mount, credential, listener, and outbound destination that the service does not need.
Attempt to write the root filesystem, access a runtime socket, query cloud metadata, reach an unrelated service, attach a debugger, exhaust one resource, and deploy an unsigned or wrong artifact. Trigger drift and verify detection. Rebuild from known-good inputs and prove that old workload credentials no longer work.
Evaluation checklist
- Does every package, process, privilege, mount, device, listener, destination, and credential support a stated service function?
- Are build, deploy, runtime, monitoring, administrator, and recovery authorities separate and narrow?
- Can the system detect deployed-state drift, not only configuration-file drift?
- Do updates test denial, evidence, resource, rollback, and recovery behavior as well as normal service?
- Can recovery replace a hostile host without reusing compromised artifacts, credentials, policy, or state?
Next learning unit
System Hardening
Reduce a deployed system to required services, identities, interfaces, privileges, configurations, and recovery paths, then keep it there.
