Skip to main content

Personal Data and Metadata

Treat identifiers, device facts, access events, relationships, timing, locations, and derived attributes as personal when context can link them to people.

Context makes data personal

Personal data is information relating to a person who is directly or indirectly identifiable. Names and email addresses are direct examples. Account IDs, device identifiers, IP addresses, locations, access times, route histories, group membership, recovery events, and behavioral patterns can identify or describe a person in context.

Metadata describes communication or processing, but it can be as sensitive as content. Who contacted which service, when, from where, for how long, and with which device can reveal work, health, relationships, or intent.

Direct, indirect, and derived data

Inventory direct identifiers, quasi-identifiers that identify in combination, pseudonyms, device and network facts, relationship graphs, content, and inferred attributes. Include data generated by scoring, analytics, anomaly detection, and models.

Classification changes with available auxiliary data and recipients. A value not identifying to one processor can be identifiable to another that has a directory, location history, or public dataset.

Minimize and separate

Collect the fields needed for a stated decision or operation. Separate identity resolution from event processing when stable identity is unnecessary. Use purpose-specific identifiers. Restrict access to mapping tables. Avoid copying full claims into every log, token, trace, and application.

Set field-level retention and deletion. Protect backups, support bundles, exports, and derived datasets, not only the primary database.

Failure and residual risk

Removing names can leave unique behavior or rare attribute combinations. Hashing an email preserves a stable link and can be reversed by guessing. Encryption protects data while encrypted and does not reduce authorized correlation. Short retention in one system can be defeated by a long-lived analytics copy.

Data about groups and shared devices can affect people even when individual identity remains uncertain.

Pomerium boundary

Pomerium can receive identity claims and device context and can log access decisions. Operators select which claims and signals policy needs, which upstreams receive identity, which fields enter logs, and how external systems retain and correlate them. Use the smallest stable identity scope that supports the access requirement.

Evaluation checklist

  • Which direct, indirect, device, network, behavioral, relationship, and inferred fields can relate to a person?
  • What auxiliary data does each recipient have for identification or correlation?
  • Does the decision need a stable identity, or can it use a scoped attribute or unlinkable event?
  • Are logs, traces, tokens, exports, support data, backups, and derived datasets included in retention and deletion?
  • Which group or person can be harmed even if the system does not know a legal name?

Sources and further reading

Keep learning

Cryptography and Data Protection

Data Classification

Assign data sensitivity, criticality, ownership, use, sharing, retention, and recovery requirements that drive technical controls.

Learn this term
Privacy EngineeringNetwork and Infrastructure

Traffic Analysis

Infer participants, relationships, activity, protocol, content class, and events from communication timing, direction, size, frequency, and routes.

Learn this term

Get a Personalized Demo

Schedule a Call with a Pomerium Engineer

Get a Demo