Learning outcomes
- Inventory untrusted data from request through every parser and interpreter.
- Assign syntax, semantic, size, normalization, and authorization rules to each field.
- Keep data separate from SQL, commands, paths, templates, and browser code.
- Test alternate encodings, stored values, parser disagreement, and downstream reuse.
Protection need
Protect the application from data that changes program structure, selects an unauthorized object, consumes unbounded resources, or is interpreted differently across components. Include data from public requests, authenticated users, identity claims, files, queues, partner APIs, databases, caches, logs, and administrative tools.
Choose one concrete flow. Example: an authenticated user creates a report with a title, date range, account identifier, sort field, export format, and optional uploaded logo. A worker reads the job from a queue, queries a database, renders HTML, creates a PDF, stores it, and sends a notification.
Map every representation:
- Browser form and JSON request.
- Gateway and framework request objects.
- Application types.
- Database parameters and stored records.
- Queue message.
- Worker process and template variables.
- HTML and PDF renderer.
- Storage path and download response.
- Logs, metrics, and notification text.
Each transition is a parser or interpreter boundary. A value that was safe in one context can become active structure in the next.
Security objectives and requirements
Define a schema for each message and field. State type, required presence, length, range, collection size, allowed values, Unicode policy, normalization, cross-field rules, and rejection behavior. Reject unknown fields where forward compatibility does not require them. Bound total body size before parsing and decompression.
Separate validation concerns:
- Syntactic validation: the date parses and the identifier has the defined format.
- Semantic validation: the start date is before the end date and the range is within policy.
- Canonicalization: equivalent values become one defined representation before comparison.
- Authorization: the current subject may use the named account and operation.
- Parameterization: values cannot become database or command structure.
- Output encoding: display data cannot become active browser or document syntax.
- Sanitization: required rich content is reduced to an explicit safe subset.
Use typed application values after validation. Do not pass raw request strings deep into the system. Preserve the original only when evidence or user display needs it, and keep it distinct from the accepted canonical value.
For structural choices such as sort field, export format, renderer, or destination, map a fixed external identifier to a server-owned implementation. Do not accept a table name, command, template path, class name, or executable from the request.
Security invariants and evidence
Write invariants at the final sinks:
- Every SQL value reaches a driver parameter, while selected column names come from a fixed map.
- No report value reaches a shell command string.
- Every storage object uses a server-generated name below one fixed root.
- Report titles render as text in HTML and PDF and cannot create markup, script, links, or external fetches.
- Queue consumers repeat schema and authorization-relevant validation for the message they receive.
- Logs record bounded structured fields and never a credential, full assertion, or attacker-controlled line structure.
Use code review and static checks to find unsafe sinks. Add integration tests that assert final queries, process arguments, resolved paths, rendered DOM, outbound requests, and stored objects. Run tests on the same framework and parser versions used in production.
Failure cases
Test malformed, ambiguous, and valid-but-dangerous input:
- Missing, duplicate, unknown, null, and wrong-type fields.
- Minimum, maximum, just-outside, very long, deeply nested, and large collection values.
- Percent-encoded, double-encoded, mixed-case, Unicode-normalization, null-byte, and alternate-separator forms.
- A valid account identifier belonging to another user or tenant.
- SQL, shell, template, HTML, URL, path, log, and spreadsheet metacharacters.
- Stored payloads that become active only when a worker or administrator views them.
- A queue producer using an older schema or a compromised internal identity.
- Parser timeouts, decompression expansion, and renderer failure.
Observe the final effect. A 400 response is not enough if the job was already queued or a partial file was stored.
Design tradeoffs and residual risk
Strict schemas improve predictability and can make version migration harder. Rich text and user-defined queries provide flexibility and create large interpreter surfaces. Unicode normalization improves consistent comparison and can alter user-visible identifiers. Revalidating at every boundary costs work and protects against stale, compromised, or incompatible producers.
Use the narrowest feature that meets the real need. When rich or programmable input is essential, isolate its parser and execution, constrain capabilities and resources, and retain explicit residual risk.
Pomerium boundary
Pomerium validates its own protocol and policy inputs and can give the application verified identity context. The application must still treat request values, identity attributes used beyond their documented meaning, queue messages, files, database content, and integration responses according to their own trust boundaries.
Pomerium cannot parameterize an application query, encode a template value, resolve a storage path, or enforce a business relationship. The application must also prevent a direct path that supplies forged identity fields around the intended gateway.
Exercise
Trace one feature from request through storage and output. Create one row per field and representation with these columns:
- Source and trust boundary.
- Raw representation.
- Canonical representation.
- Syntax rule.
- Semantic rule.
- Authorization rule.
- Size and resource bound.
- Final interpreter or sink.
- Safe API or encoding mechanism.
- Negative test and observed final effect.
Include at least one identifier, free-form string, URL, path, collection, stored value, and structural choice. Add tests for a stored payload and a parser disagreement between two components.
Evaluation checklist
- Does the inventory include every external, cross-service, stored, and administrative input?
- Are syntax, meaning, canonicalization, authorization, parameterization, and encoding separate decisions?
- Does each structural choice come from a fixed server-owned map?
- Is encoding applied at the final sink for the exact context?
- Do consumers validate messages even when the producer is internal?
- Are time, memory, nesting, decompression, collection, and output resources bounded?
- Do tests inspect the final query, process, path, DOM, object, and side effect?
Next learning unit
Canonicalization
Convert equivalent input representations to one defined form before comparison, validation, authorization, and storage.
