Errors are system states
An error is part of the security design. Parsers reject input, identity sources time out, authorization context becomes unavailable, databases conflict, queues duplicate work, storage fills, and downstream services return partial results. The system must enter a defined state that preserves its critical invariants.
Do not catch every exception and continue. Do not turn an unknown decision into allow. Do not report success before the protected state and evidence are durable.
Safe failure behavior
Classify errors by operation and protection need. A sensitive authorization dependency can fail closed. A public read-only feature might serve bounded stale data. An emergency service may enter a separately controlled degraded mode. State these choices before the failure occurs.
Use transactions and compensating actions for partial work. Make retried operations idempotent or attach a unique operation key. Bound retry count and delay. A timeout means the result is unknown until the system reconciles it; it does not prove the operation did not happen.
Messages and evidence
Return a stable external error that helps the caller recover without exposing stack traces, filesystem paths, query text, credentials, internal hosts, policy internals, or existence of unauthorized objects. Use the same outward result where detail would enable enumeration.
Record an internal correlation identifier, operation, stage, decision, dependency, and outcome. Protect logs from secrets and unbounded attacker input. Alert on repeated or security-significant failures. Keep enough evidence to determine whether a partial action occurred.
Failure and residual risk
Fail-closed behavior can become a denial-of-service mechanism. Fail-open behavior can become unauthorized access. Generic errors can make operations impossible to diagnose. Detailed errors can disclose exploitable state. Recovery code runs rarely and can hold broad authority, so it needs its own tests and review.
Pomerium boundary
Pomerium has documented behavior for its authentication, policy, routing, and dependency failures. The upstream application owns its own transactions, retries, partial state, error messages, and recovery. A gateway denial does not roll back an application side effect that already occurred, and an upstream timeout does not reveal whether that effect committed.
Evaluation checklist
- What state does each security-sensitive operation enter after parse, dependency, policy, storage, or timeout failure?
- Does an unknown authorization result ever become allow?
- Can retry duplicate a state change or external effect?
- Do external errors avoid secrets, object enumeration, and internal implementation details?
- Can operators connect the external correlation identifier to the final internal outcome?
- Are degraded and recovery paths tested with the same rigor as the normal path?
