One value, many representations
Canonicalization converts equivalent representations to one defined form. Paths can contain dot segments, repeated separators, percent encoding, symbolic links, case differences, or Unicode variants. Host names, identifiers, and text can also have more than one representation. A security decision is unstable when two components interpret the same bytes differently.
Normalize before deciding
Decode and normalize in a defined order before comparison, validation, authorization, cache lookup, signature input, or storage. Reject malformed, forbidden, or multiply encoded input. Do not repeatedly decode until the data stops changing. That behavior gives attackers a way to hide structure from one control and reveal it to another.
For a filesystem path, resolve the candidate using the platform's path rules and verify that the resulting object remains below the approved root. For a URL, use one standards-compliant parser and validate the parsed scheme, authority, host, port, and resolved destination. For Unicode identifiers, select a normalization form and state whether case folding or script restrictions apply.
Cross-component agreement
The component that authorizes a value and the component that uses it must agree on the canonical form. A gateway may authorize one path while an upstream framework rewrites it to another. A cache and origin may use different keys. A signature verifier may normalize differently from the code that acts on the message.
Test the complete chain with encoded separators, dot segments, duplicate fields, mixed case, alternate Unicode, null bytes, trailing delimiters, and parser-specific edge cases. Observe the final resource and action, not only the first parser's output.
Failure and residual risk
Normalization can change legitimate text or collapse two distinct identifiers. A filesystem check can be invalidated by a symbolic-link or mount change after validation. A reverse proxy and application can still disagree when each uses a different parser version. Canonicalization reduces ambiguity but does not grant authorization.
Pomerium boundary
Pomerium matches and forwards routes using its documented request model. The upstream application and its framework still parse the received path, host, query, and object identifiers. Operators must test that the external gateway, intermediary, and application agree. Pomerium cannot make an application path or storage name safe after the upstream reinterprets it.
Evaluation checklist
- Which representations can refer to the same resource or identity?
- Is decoding and normalization performed once in a defined order before the decision?
- Do the policy layer, cache, router, application, and storage layer use compatible parsing rules?
- Does the final resolved path or destination remain inside the approved boundary?
- Can state change between validation and use?
- Do tests observe the final resource for encoded, Unicode, case, separator, and alias variants?
