Reduce the protected surface
Data minimization limits collection, precision, fields, copies, users, and retention to what a defined purpose requires. Purpose limitation prevents data collected for one reason from becoming a general input to unrelated decisions or products without a new approved basis and control review.
Data that does not exist cannot be stolen, leaked through logs, misused by an operator, exposed through a vulnerable export, or retained past its need.
Purpose and necessity
For every field, event, derived value, and copy, record the purpose, owner, required precision, users, destinations, retention, and deletion trigger. Ask whether the same outcome can use less data, a coarser value, a short-lived value, a local decision, or a yes-or-no result.
Separate operational identity data from analytics and support data. Do not copy full tokens, assertions, request bodies, secrets, or personal attributes into logs when a bounded identifier and decision reason are enough.
Enforcement and review
Enforce field selection in schemas and APIs. Restrict query and export paths. Apply per-purpose authorization and separate storage when reuse risk is high. Set automatic expiry and deletion. Test that old fields and copies disappear from caches, indexes, queues, backups according to policy.
Review new use against the original purpose, data subjects, threat model, and sharing boundary. Derived scores and inferred attributes can be more sensitive than source data and need their own classification and explanation.
Failure and residual risk
Collecting everything for possible future use creates open-ended risk. Pseudonymous identifiers can often be linked back through other data. Aggregation can reveal details not present in one field. Deleting the primary record can leave logs, backups, exports, and model or search indexes. Excessive minimization can remove evidence needed for security response.
Pomerium boundary
Pomerium can provide identity and request context and generate access evidence. Operators choose which claims to release, which events to retain, and where logs go. Upstream applications should request and store only the identity data they need. Pomerium cannot delete copies made by applications, analytics, exports, or external log systems.
Evaluation checklist
- Does every collected and derived field have a current, specific purpose and owner?
- Can the outcome use less precision, fewer fields, a short-lived value, or a local decision?
- Are application, analytics, support, audit, and security uses separated?
- Do schemas, queries, exports, logs, and permissions enforce the purpose?
- Are retention and deletion automatic across caches, queues, indexes, backups, and downstream copies?
- Does minimization retain the minimum evidence needed to detect and investigate abuse?
