Skip to main content

Validate and encode untrusted data

Build typed validation, canonicalization, parameterization, sanitization, and output encoding at every application boundary.

Learning outcomes

  • Inventory untrusted data from request through every parser and interpreter.
  • Assign syntax, semantic, size, normalization, and authorization rules to each field.
  • Keep data separate from SQL, commands, paths, templates, and browser code.
  • Test alternate encodings, stored values, parser disagreement, and downstream reuse.

Protection need

Protect the application from data that changes program structure, selects an unauthorized object, consumes unbounded resources, or is interpreted differently across components. Include data from public requests, authenticated users, identity claims, files, queues, partner APIs, databases, caches, logs, and administrative tools.

Choose one concrete flow. Example: an authenticated user creates a report with a title, date range, account identifier, sort field, export format, and optional uploaded logo. A worker reads the job from a queue, queries a database, renders HTML, creates a PDF, stores it, and sends a notification.

Map every representation:

  1. Browser form and JSON request.
  2. Gateway and framework request objects.
  3. Application types.
  4. Database parameters and stored records.
  5. Queue message.
  6. Worker process and template variables.
  7. HTML and PDF renderer.
  8. Storage path and download response.
  9. Logs, metrics, and notification text.

Each transition is a parser or interpreter boundary. A value that was safe in one context can become active structure in the next.

Security objectives and requirements

Define a schema for each message and field. State type, required presence, length, range, collection size, allowed values, Unicode policy, normalization, cross-field rules, and rejection behavior. Reject unknown fields where forward compatibility does not require them. Bound total body size before parsing and decompression.

Separate validation concerns:

  • Syntactic validation: the date parses and the identifier has the defined format.
  • Semantic validation: the start date is before the end date and the range is within policy.
  • Canonicalization: equivalent values become one defined representation before comparison.
  • Authorization: the current subject may use the named account and operation.
  • Parameterization: values cannot become database or command structure.
  • Output encoding: display data cannot become active browser or document syntax.
  • Sanitization: required rich content is reduced to an explicit safe subset.

Use typed application values after validation. Do not pass raw request strings deep into the system. Preserve the original only when evidence or user display needs it, and keep it distinct from the accepted canonical value.

For structural choices such as sort field, export format, renderer, or destination, map a fixed external identifier to a server-owned implementation. Do not accept a table name, command, template path, class name, or executable from the request.

Security invariants and evidence

Write invariants at the final sinks:

  • Every SQL value reaches a driver parameter, while selected column names come from a fixed map.
  • No report value reaches a shell command string.
  • Every storage object uses a server-generated name below one fixed root.
  • Report titles render as text in HTML and PDF and cannot create markup, script, links, or external fetches.
  • Queue consumers repeat schema and authorization-relevant validation for the message they receive.
  • Logs record bounded structured fields and never a credential, full assertion, or attacker-controlled line structure.

Use code review and static checks to find unsafe sinks. Add integration tests that assert final queries, process arguments, resolved paths, rendered DOM, outbound requests, and stored objects. Run tests on the same framework and parser versions used in production.

Failure cases

Test malformed, ambiguous, and valid-but-dangerous input:

  • Missing, duplicate, unknown, null, and wrong-type fields.
  • Minimum, maximum, just-outside, very long, deeply nested, and large collection values.
  • Percent-encoded, double-encoded, mixed-case, Unicode-normalization, null-byte, and alternate-separator forms.
  • A valid account identifier belonging to another user or tenant.
  • SQL, shell, template, HTML, URL, path, log, and spreadsheet metacharacters.
  • Stored payloads that become active only when a worker or administrator views them.
  • A queue producer using an older schema or a compromised internal identity.
  • Parser timeouts, decompression expansion, and renderer failure.

Observe the final effect. A 400 response is not enough if the job was already queued or a partial file was stored.

Design tradeoffs and residual risk

Strict schemas improve predictability and can make version migration harder. Rich text and user-defined queries provide flexibility and create large interpreter surfaces. Unicode normalization improves consistent comparison and can alter user-visible identifiers. Revalidating at every boundary costs work and protects against stale, compromised, or incompatible producers.

Use the narrowest feature that meets the real need. When rich or programmable input is essential, isolate its parser and execution, constrain capabilities and resources, and retain explicit residual risk.

Pomerium boundary

Pomerium validates its own protocol and policy inputs and can give the application verified identity context. The application must still treat request values, identity attributes used beyond their documented meaning, queue messages, files, database content, and integration responses according to their own trust boundaries.

Pomerium cannot parameterize an application query, encode a template value, resolve a storage path, or enforce a business relationship. The application must also prevent a direct path that supplies forged identity fields around the intended gateway.

Exercise

Trace one feature from request through storage and output. Create one row per field and representation with these columns:

  1. Source and trust boundary.
  2. Raw representation.
  3. Canonical representation.
  4. Syntax rule.
  5. Semantic rule.
  6. Authorization rule.
  7. Size and resource bound.
  8. Final interpreter or sink.
  9. Safe API or encoding mechanism.
  10. Negative test and observed final effect.

Include at least one identifier, free-form string, URL, path, collection, stored value, and structural choice. Add tests for a stored payload and a parser disagreement between two components.

Evaluation checklist

  • Does the inventory include every external, cross-service, stored, and administrative input?
  • Are syntax, meaning, canonicalization, authorization, parameterization, and encoding separate decisions?
  • Does each structural choice come from a fixed server-owned map?
  • Is encoding applied at the final sink for the exact context?
  • Do consumers validate messages even when the producer is internal?
  • Are time, memory, nesting, decompression, collection, and output resources bounded?
  • Do tests inspect the final query, process, path, DOM, object, and side effect?

Next learning unit

Canonicalization

Convert equivalent input representations to one defined form before comparison, validation, authorization, and storage.

Sources and further reading

Keep learning

Software and Application Security

SQL Injection

Keep untrusted values separate from SQL structure with parameterized queries, allowlisted identifiers, and narrow database authority.

Learn this term
Software and Application Security

Command Injection

Avoid command interpreters and pass fixed executables and validated arguments through structured process APIs with narrow authority.

Learn this term
Software and Application Security

Cross-Site Scripting (XSS)

Prevent untrusted data from executing as active browser content through context-aware encoding, safe DOM APIs, and constrained markup.

Learn this term
Software and Application Security

Path Traversal

Prevent attacker-controlled file names and paths from escaping an approved storage root after decoding and canonical resolution.

Learn this term

Get a Personalized Demo

Schedule a Call with a Pomerium Engineer

Get a Demo