A claim about an observer
Anonymization aims to prevent records or information from being associated with individuals in a defined release context. Re-identification links a record or pattern back to a person. De-identification covers techniques that reduce this association but does not guarantee anonymity.
The claim needs a recipient, auxiliary-data model, release mechanism, population, time, and acceptable risk.
Direct and quasi-identifiers
Remove unnecessary direct identifiers and transform quasi-identifiers such as location, age, role, rare event, and time. Evaluate uniqueness and homogeneity. Consider free text, images, metadata, missingness, outliers, and derived values.
Choose a release model: public dataset, controlled enclave, approved recipient, synthetic data, query interface, aggregate, or formal privacy mechanism. Access governance can be stronger than trying to publish a detailed static file safely.
Re-identification and inference testing
Use plausible external datasets, recipient knowledge, repeated releases, and targeted records. Test linkage, membership inference, attribute inference, and small-group disclosure. Review utility and privacy together.
Use independent disclosure review for material releases. Record data, method, parameters, attacker model, results, restrictions, recipient terms, and future review triggers.
Failure and residual risk
Removing names leaves unique combinations. Suppression and generalization can fail for rare records. Synthetic data can memorize or reveal training examples. A later public dataset can invalidate an earlier analysis. Multiple safe-looking releases can combine into an unsafe result.
There is no universal anonymization transform. The data and release environment change over time.
Pomerium boundary
Pomerium access logs and identity data can support investigation and can be highly linkable. Removing names or hashing subjects before an export does not prove anonymity. Operators own the release model, transformation, recipient controls, re-identification testing, retention, and review.
Evaluation checklist
- Who is the observer, what auxiliary data can they use, and how will that set change?
- Which direct, quasi, free-text, image, timing, location, and behavioral features remain unique?
- Is a public dataset necessary, or can a query service, aggregate, synthetic output, or controlled enclave meet the purpose?
- Do tests cover linkage, membership, attribute inference, rare records, and repeated releases?
- Which future dataset, recipient action, or retained copy can invalidate the current claim?
