pMSE mechanism: Differentially private synthetic data with maximal distributional similarity

Joshua Snoke, Aleksandra Slavković

Research output: Chapter in Book/Report/Conference proceedingConference contribution

19 Scopus citations


We propose a method for the release of differentially private synthetic datasets. In many contexts, data contain sensitive values which cannot be released in their original form in order to protect individuals’ privacy. Synthetic data is a protection method that releases alternative values in place of the original ones, and differential privacy (DP) is a formal guarantee for quantifying the privacy loss. We propose a method that maximizes the distributional similarity of the synthetic data relative to the original data using a measure known as the pMSE, while guaranteeing ε-DP. We relax common DP assumptions concerning the distribution and boundedness of the original data. We prove theoretical results for the privacy guarantee and provide simulations for the empirical failure rate of the theoretical results under typical computational limitations. We give simulations for the accuracy of linear regression coefficients generated from the synthetic data compared with the accuracy of non-DP synthetic data and other DP methods. Additionally, our theoretical results extend a prior result for the sensitivity of the Gini Index to include continuous predictors.

Original languageEnglish (US)
Title of host publicationPrivacy in Statistical Databases - UNESCO Chair in Data Privacy, International Conference, PSD 2018, Proceedings
EditorsFrancisco Montes, Josep Domingo-Ferrer
PublisherSpringer Verlag
Number of pages22
ISBN (Print)9783319997704
StatePublished - 2018
EventInternational Conference on Privacy in Statistical Databases, PSD 2018 - Valencia, Spain
Duration: Sep 26 2018Sep 28 2018

Publication series

NameLecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Volume11126 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349


OtherInternational Conference on Privacy in Statistical Databases, PSD 2018

All Science Journal Classification (ASJC) codes

  • Theoretical Computer Science
  • General Computer Science


Dive into the research topics of 'pMSE mechanism: Differentially private synthetic data with maximal distributional similarity'. Together they form a unique fingerprint.

Cite this