Skip to main content

Anonymization and Re-Identification

Evaluate whether a release resists identification and sensitive inference under realistic auxiliary data, recipients, repeated releases, and governance.

A claim about an observer

Anonymization aims to prevent records or information from being associated with individuals in a defined release context. Re-identification links a record or pattern back to a person. De-identification covers techniques that reduce this association but does not guarantee anonymity.

The claim needs a recipient, auxiliary-data model, release mechanism, population, time, and acceptable risk.

Direct and quasi-identifiers

Remove unnecessary direct identifiers and transform quasi-identifiers such as location, age, role, rare event, and time. Evaluate uniqueness and homogeneity. Consider free text, images, metadata, missingness, outliers, and derived values.

Choose a release model: public dataset, controlled enclave, approved recipient, synthetic data, query interface, aggregate, or formal privacy mechanism. Access governance can be stronger than trying to publish a detailed static file safely.

Re-identification and inference testing

Use plausible external datasets, recipient knowledge, repeated releases, and targeted records. Test linkage, membership inference, attribute inference, and small-group disclosure. Review utility and privacy together.

Use independent disclosure review for material releases. Record data, method, parameters, attacker model, results, restrictions, recipient terms, and future review triggers.

Failure and residual risk

Removing names leaves unique combinations. Suppression and generalization can fail for rare records. Synthetic data can memorize or reveal training examples. A later public dataset can invalidate an earlier analysis. Multiple safe-looking releases can combine into an unsafe result.

There is no universal anonymization transform. The data and release environment change over time.

Pomerium boundary

Pomerium access logs and identity data can support investigation and can be highly linkable. Removing names or hashing subjects before an export does not prove anonymity. Operators own the release model, transformation, recipient controls, re-identification testing, retention, and review.

Evaluation checklist

  • Who is the observer, what auxiliary data can they use, and how will that set change?
  • Which direct, quasi, free-text, image, timing, location, and behavioral features remain unique?
  • Is a public dataset necessary, or can a query service, aggregate, synthetic output, or controlled enclave meet the purpose?
  • Do tests cover linkage, membership, attribute inference, rare records, and repeated releases?
  • Which future dataset, recipient action, or retained copy can invalidate the current claim?

Sources and further reading

Keep learning

Privacy Engineering

Pseudonymization

Replace direct identity with a controlled reference while treating the mapping, stable links, attributes, and auxiliary data as remaining privacy risks.

Learn this term
Privacy EngineeringPlatform and Component Security

Inference Attack

Derive protected facts from permitted queries, aggregates, models, correlations, errors, and repeated observations.

Learn this term

Get a Personalized Demo

Schedule a Call with a Pomerium Engineer

Get a Demo