Skip to main content

Input Validation and Output Encoding

Validate syntax and meaning at input boundaries, then encode untrusted data for the exact output interpreter and context.

Two different controls

Input validation decides whether data is acceptable for a stated field and operation. Output encoding represents data so the next interpreter treats it as data rather than structure. One control cannot replace the other.

Validate structured input against an allowlist of types, lengths, ranges, enumerated values, relationships, and business rules. Encode output for its exact destination, such as HTML text, an HTML attribute, a URL component, JavaScript data, CSS, a database parameter, or a command argument. Context changes the correct representation.

Syntax and meaning

Syntactic validation checks form. A date parses, an identifier uses the allowed character set, and an array has a bounded length. Semantic validation checks meaning. A start time precedes an end time, a transfer amount is within the account limit, and a referenced object belongs to the authorized tenant.

Perform authoritative validation on the server or trusted service boundary. Client validation improves feedback but an attacker can bypass it. Validate data from partners, queues, files, APIs, and internal services as well as public requests. An internal producer can be wrong or compromised.

Context-specific output

Use frameworks that escape by default. Keep untrusted data out of dangerous contexts such as executable script, style, event-handler, raw HTML, or dynamic template source. When rich content is required, use a maintained sanitizer with an explicit policy and serve the result under a suitable browser policy.

For SQL, use parameterized queries. For operating-system actions, avoid a command shell and use a structured process API with separate arguments. For paths, resolve against a fixed root and verify containment. These mechanisms keep structure separate from data more reliably than character filters.

Failure and residual risk

Denying a short list of suspicious strings is easy to bypass and can reject valid data. Encoding too early can be decoded and reinterpreted later. Encoding twice can corrupt data. A valid value can still be unauthorized or harmful in the business workflow. Unicode and alternate encodings can create ambiguous forms if canonicalization is inconsistent.

Pomerium boundary

Pomerium validates its own configuration and protocol data and can supply verified identity context to an upstream. The protected application must validate its request fields, object relationships, uploaded content, and downstream data. It must encode output for every interpreter it invokes. Route authentication does not make request data trustworthy.

Evaluation checklist

  • Is every external and cross-service field assigned a type, length, range, and semantic rule?
  • Does the authoritative service repeat checks that exist in the browser?
  • Is output encoded at the final sink for the exact interpreter and context?
  • Do database, process, path, and template APIs keep data separate from structure?
  • Can a valid value still select another tenant, exceed a business limit, or trigger an unsafe transition?
  • Do tests cover alternate encodings, boundary values, malformed collections, and downstream decoding?

Sources and further reading

Keep learning

Software and Application Security

Canonicalization

Convert equivalent input representations to one defined form before comparison, validation, authorization, and storage.

Learn this term
Software and Application Security

Injection

Prevent attacker-controlled data from changing the structure or meaning of a command sent to an interpreter.

Learn this term
Software and Application Security

Cross-Site Scripting (XSS)

Prevent untrusted data from executing as active browser content through context-aware encoding, safe DOM APIs, and constrained markup.

Learn this term

Get a Personalized Demo

Schedule a Call with a Pomerium Engineer

Get a Demo