Skip to main content

Build a Privacy Threat Model

Model data actions, observers, linkability, inference, unawareness, loss of control, and human consequences across a complete system lifecycle.

Learning outcomes

  • Build a data and interaction model that includes raw, derived, metadata, support, export, and recovery paths.
  • Identify privacy threats from normal operation, misuse, recipient correlation, and technical attack.
  • Connect problematic data actions to consequences for people and groups.
  • Select, test, and review controls across minimization, association, access, use, retention, and release.

Assets and security objectives

Select one system and state its necessary purpose. Identify people and groups affected, including non-users represented in data or affected by decisions. Inventory identity, content, metadata, behavior, location, relationships, device facts, inferences, and outputs.

Define privacy objectives such as predictable use, manageable data, limited association, confidentiality, unlinkability, undetectability, plausible anonymity, narrow retention, and contestable decisions. State useful function and acceptable residual risk.

Actors and components

Include data subjects, users, operators, administrators, support, developers, identity providers, gateways, upstream applications, analytics, model services, vendors, recipients, investigators, attackers, and observers. Record each actor's data, purpose, authority, auxiliary information, and ability to combine or retain output.

Map clients, services, data stores, queues, logs, caches, backups, exports, dashboards, support bundles, recovery systems, and third parties. Include derived datasets and manual spreadsheets.

Trust boundaries

Draw every collection, transmission, transformation, association, query, disclosure, and deletion. At each source, flow, and destination, record fields, identifiers, subject or device scope, purpose, recipient, access, encryption, retention, and evidence.

Mark where identifiers change or remain stable. Mark joins and mapping tables. Include network observers and intermediaries that see traffic patterns even when content is encrypted.

Normal request path

Trace one person's data from first collection through identity resolution, policy, application use, telemetry, analytics, support, export, backup, and deletion. Show what the person and accountable owner can predict and manage at each step.

Trace one non-user or group effect. A person's contact, colleague, household member, or demographic group can be represented or affected without an account.

Failure path: excessive association

A global account or device identifier enters every route log, upstream claim, analytics event, and vendor export. Each use is authorized, but recipients can reconstruct a detailed cross-service history.

Use audience- and purpose-scoped identifiers, minimize claims and fields, separate identity resolution, shorten retention, and restrict joins. Test whether timing, device, location, and rare behavior still reconnect the records.

Failure path: harmful inference

An authenticated analytics endpoint returns counts and model scores. Repeated overlapping queries reveal whether one person belongs to a sensitive group or used a protected service.

Model membership and attribute inference, auxiliary data, repetitions, accounts, and collusion. Use appropriate aggregation, query controls, contribution bounds, controlled access, or differential privacy. Test cumulative releases.

Failure path: loss of manageability

The primary service deletes an account, but identity claims remain in logs, backups, support exports, derived datasets, and model features with no owner or subject link.

Create data lineage and field-level retention. Preserve enough controlled linkage to execute required deletion without making all processing globally identifiable. Test deletion through restore and derived output.

Prioritize and control

For each threat, record data action, affected person or group, problem, likelihood, impact, benefit, existing control, response, owner, and evidence. Prefer elimination and minimization before access restrictions. Then use scoped identity, isolation, cryptography, information-flow rules, output controls, user or operator controls, retention, and review.

Test normal, alternate, failure, recipient, repeated-release, support, export, and recovery paths. Revisit the model when purpose, field, recipient, identifier, model, or retention changes.

Design tradeoffs and residual risk

Reducing identifiers can harm abuse detection and support. Short retention can remove investigation evidence. Granular control can overwhelm people. Formal privacy can reduce accuracy. Central privacy services can become linkability and availability points.

Record which observer, inference, group effect, human memory, endpoint, or third-party action remains outside control. Do not claim anonymity or consent beyond the tested context.

Pomerium boundary

Pomerium can limit route access and record identity-aware decisions. The threat model must include identity-provider claims, Pomerium sessions and logs, upstream identity, application events, analytics, and support. Operators choose fields, recipients, retention, and correlation. Pomerium does not control later application use or third-party datasets.

Exercise

Model one identity-aware application with at least two routes, an identity provider, Pomerium, upstream, telemetry, analytics, support, export, backup, and deletion. Trace one user and one affected non-user or group.

Find at least one linkability, inference, traffic-analysis, unawareness, and manageability threat. Remove one field, scope one identifier, block one join, and test one deletion and one repeated-query attack. Record the utility effect and residual risk.

Evaluation checklist

  • Does the model include raw, derived, metadata, network, log, support, export, backup, model, and deletion paths?
  • Are normal authorized processing and human consequences analyzed as well as malicious disclosure?
  • Which observer, recipient, auxiliary data, repeated release, and future join can identify or infer protected facts?
  • Can the system manage data across every copy without relying on one uncontrolled global identifier?
  • Do controls have deployed tests, owners, review triggers, utility effects, and explicit residual risk?

Next learning unit

Privacy Risk

Assess data actions that can create problems for people, then combine likelihood and impact without reducing privacy to breach risk.

Sources and further reading

Keep learning

Privacy EngineeringPlatform and Component Security

Inference Attack

Derive protected facts from permitted queries, aggregates, models, correlations, errors, and repeated observations.

Learn this term
Privacy EngineeringNetwork and Infrastructure

Traffic Analysis

Infer participants, relationships, activity, protocol, content class, and events from communication timing, direction, size, frequency, and routes.

Learn this term

Get a Personalized Demo

Schedule a Call with a Pomerium Engineer

Get a Demo