Learning outcomes
- Define the protected entity, recipient, auxiliary-data, release, and future-data threat model.
- Compare static de-identification, controlled access, synthetic data, query services, and differential privacy.
- Test linkage, membership, attribute inference, reconstruction, and repeated release.
- Set governance, parameters, review triggers, and evidence for a bounded release claim.
Protection need
State the useful analysis or product output and the protected entity: person, account, device, household, organization, or event. Define whether the concern is identity, membership, sensitive attribute, relationship, behavior, or exact record reconstruction.
Name recipients, public access, collusion, auxiliary datasets, query capability, repeated releases, and time horizon. A claim for one approved researcher is not a claim for public release.
Security objectives and requirements
Choose a release model before a transform:
- No release; answer the purpose through a safer operational process.
- Controlled enclave or approved recipient with output review.
- Query interface with access, purpose, rate, overlap, and output controls.
- Aggregate or synthetic data with disclosure testing.
- Formal differential-privacy mechanism with explicit protected unit, adjacency, contribution bounds, parameters, and accounting.
- Static transformed dataset with a measured de-identification standard and governance.
Minimize fields, rows, precision, geography, time, and rare categories. Handle free text, images, missing values, and outliers separately.
Security invariants and evidence
- Direct identifiers and mapping data do not reach ordinary recipients.
- Quasi-identifiers meet the chosen risk measure under the stated population and auxiliary data.
- Repeated and overlapping releases are reviewed together.
- Query limits and privacy budgets apply across accounts, endpoints, teams, and time.
- Synthetic or model output is tested for training-record memorization and membership leakage.
- Release artifacts, parameters, code, inputs, review, recipients, and restrictions are versioned and retained as evidence.
Failure cases
- Hashed email values are reversed with a dictionary.
- Rare role, office, age, and event time identify one person after names are removed.
- Two aggregate tables differ by one person and reveal membership or attribute.
- A user resets query limits through another account.
- Synthetic text reproduces a sensitive training phrase.
- A later public breach provides the join key for an old release.
- A differential-privacy mechanism uses unbounded contribution or a repeated deterministic seed.
Adversarial evaluation
Build plausible auxiliary datasets and attempt record linkage. Target rare and high-impact records. Test membership and sensitive attribute inference. Compare overlapping releases. Evaluate whether an observer can distinguish presence, absence, or category with useful confidence.
For differential privacy, review protected unit, adjacency, clipping, sensitivity, parameters, random generation, accounting, post-processing, retries, caches, logs, and alternate raw-data paths. Use independent review for high-impact or public releases.
Design tradeoffs and residual risk
More transformation reduces utility. Controlled access retains operational burden and recipient trust. Synthetic data can improve usability and preserve patterns that leak. Differential privacy gives a formal bound and requires careful modeling and can reduce small-group accuracy.
No release remains safe forever against all auxiliary data. Set review triggers for new datasets, recipients, linkage techniques, repeated releases, purpose changes, and incidents.
Pomerium boundary
Pomerium can protect a controlled query service or enclave route and record who accessed it. It does not make an exported log or dataset anonymous. The data owner must control queries, outputs, recipients, parameters, repeated releases, and alternate raw-data access.
Exercise
Select one access-log or analytics dataset and one useful output. Compare public transformed data, a controlled query service, synthetic data, and a differential-privacy release.
Run linkage, membership, attribute, rare-record, overlap, and future-auxiliary-data tests. Choose the smallest release model that meets the purpose. Record parameters, evidence, recipient controls, utility loss, and the date or event that triggers reassessment.
Evaluation checklist
- Is the protected entity, fact, recipient, auxiliary data, repeated release, and time horizon explicit?
- Was the release model chosen before the de-identification technique?
- Do tests cover direct linkage, quasi-identifiers, membership, attributes, reconstruction, rare records, and composition?
- Are formal privacy parameters, contribution bounds, accounting, implementation, and raw-data bypasses verified where used?
- Are governance, utility, recipient limits, future review, and residual re-identification risk recorded?
Next learning unit
Anonymization and Re-Identification
Evaluate whether a release resists identification and sensitive inference under realistic auxiliary data, recipients, repeated releases, and governance.
