Skip to main content

Differential Privacy

Bound how much a computation's output distribution can change when one person's contribution is added or removed.

A quantified release guarantee

Differential privacy is a mathematical framework that bounds the change in output distribution when one protected entity's contribution changes between neighboring datasets. It can limit what an observer learns about participation from released statistics or trained models.

The guarantee depends on the defined neighboring relation, privacy parameters, contribution bounds, mechanism, composition, and implementation.

Model and parameters

Define the protected unit: person, device, household, event, or account. Define whether adjacency adds or removes one unit or changes one record. Bound how many rows and how much value one unit contributes. Select the mechanism and privacy parameters, including epsilon and delta where applicable.

Explain the guarantee and utility in operational terms. A small parameter does not help if contribution bounds are wrong or the raw data is released elsewhere.

Accounting and implementation

Track cumulative privacy loss across repeated and adaptive queries, releases, models, and teams. Enforce the budget in one authoritative service. Use cryptographically secure random generation and reviewed implementations. Protect raw data and intermediate results.

Test clipping, sensitivity, noise distribution, seeding, floating-point behavior, accounting, caching, retries, failure, and side outputs. Post-processing preserves the guarantee only when it uses the private output without new access to private data.

Failure and residual risk

Wrong adjacency or unbounded contribution can make the claim irrelevant. Reusing noise, deterministic seeds, logging raw values, repeated retries, and separate uncoordinated budgets can leak data. Differential privacy does not prevent harmful decisions from accurate population patterns or protect a person whose facts are already known.

More privacy can reduce accuracy, especially for small groups and rare events. Poor communication can cause users to interpret a bounded statistical guarantee as full anonymity.

Pomerium boundary

Pomerium can protect access to raw data and a differential-privacy query service. It does not implement the application's adjacency, sensitivity, mechanism, or privacy budget. The analytics owner must prevent direct or alternate access that bypasses the formal release mechanism.

Evaluation checklist

  • What entity and neighboring relation does the guarantee protect?
  • Are contribution bounds, sensitivity, parameters, mechanism, and utility explicit?
  • Does one authoritative accountant cover repeated queries, users, releases, models, retries, and time?
  • Can logs, errors, seeds, caches, intermediate data, or alternate endpoints bypass the guarantee?
  • Which group harm, known fact, small-population error, and raw-data risk remain outside the claim?

Sources and further reading

Keep learning

Privacy EngineeringPlatform and Component Security

Inference Attack

Derive protected facts from permitted queries, aggregates, models, correlations, errors, and repeated observations.

Learn this term

Get a Personalized Demo

Schedule a Call with a Pomerium Engineer

Get a Demo