Skip to main content

Linkability and Identifiability

Distinguish connecting events to each other from connecting them to a person, and bound both by observer, dataset, purpose, and time.

Two different observations

Linkability is the ability to determine that two items, events, sessions, records, or communications relate to the same person or entity. Identifiability is the ability to learn which person that entity is. An observer can track one pseudonymous user without knowing a name, and later identification can expose the complete linked history.

Scope the observer and set

State observer, information set, auxiliary data, confidence, population, and time. A random identifier is linkable to every recipient that sees it. A shared device or IP can link several people incorrectly. Rare event patterns can link records without a formal identifier.

Assess linkability across services, tenants, identity providers, logs, networks, devices, and exports. Include colluding recipients and later data release.

Use purpose- and audience-specific identifiers. Rotate identifiers when continuity is not needed. Separate identity mapping from event processing and restrict the join. Minimize timestamps, locations, attributes, and fingerprints. Aggregate or batch events when exact sequence is unnecessary.

Keep investigation re-identification under narrow authority with evidence and retention. Do not send one global subject or email to unrelated services when a scoped identifier works.

Failure and residual risk

Rotation fails if stable attributes, timing, network address, device fingerprint, or behavior reconnects sessions. Hashing a stable identifier preserves linkability. Small anonymity sets make records unique. Independent databases can be joined later by a recipient beyond the original threat model.

Reducing links can complicate abuse prevention, support, account recovery, and investigation. Define which continuity is necessary and who may resolve it.

Pomerium boundary

Pomerium relies on identity-provider subjects and can pass identity to configured upstreams. Operators should understand whether identifiers are global, provider-scoped, tenant-scoped, or application-scoped and should minimize claims and log fields. Pomerium cannot stop separate upstreams from correlating data they both receive.

Evaluation checklist

  • Can the observer connect records without knowing a name, and can later data identify the linked subject?
  • Which stable identifier, timestamp, location, network fact, device fingerprint, or behavior creates the link?
  • Are identifiers scoped by audience, tenant, purpose, and time rather than reused globally?
  • Which legitimate support, fraud, or investigation need requires continuity, and who can resolve identity?
  • Can recipient collusion or future auxiliary data defeat the stated unlinkability claim?

Sources and further reading

Keep learning

Privacy EngineeringCryptography and Data Protection

Personal Data and Metadata

Treat identifiers, device facts, access events, relationships, timing, locations, and derived attributes as personal when context can link them to people.

Learn this term
Privacy Engineering

Pseudonymization

Replace direct identity with a controlled reference while treating the mapping, stable links, attributes, and auxiliary data as remaining privacy risks.

Learn this term
Privacy EngineeringNetwork and Infrastructure

Traffic Analysis

Infer participants, relationships, activity, protocol, content class, and events from communication timing, direction, size, frequency, and routes.

Learn this term

Get a Personalized Demo

Schedule a Call with a Pomerium Engineer

Get a Demo