Information preservation in statistical privacy and Bayesian estimation of unattributed histograms

Bing Rong Lin, Daniel Kifer

Research output: Chapter in Book/Report/Conference proceedingConference contribution

14 Scopus citations

Abstract

In statistical privacy, utility refers to two concepts: information preservation - how much statistical information is retained by a sanitizing algorithm, and usability - how (and with how much difficulty) does one extract this information to build statistical models, answer queries, etc. Some scenarios incentivize a separation between information preservation and usability, so that the data owner first chooses a sanitizing algorithm to maximize a measure of information preservation and, afterward, the data consumers process the sanitized output according to their needs [22, 46]. We analyze a variety of utility measures and show that the average (over possible outputs of the sanitizer) error of Bayesian decision makers forms the unique class of utility measures that satisfy three axioms related to information preservation. The axioms are agnostic to Bayesian concepts such as subjective probabilities and hence strengthen support for Bayesian views in privacy research. In particular, this result connects information preservation to aspects of usability - if the information preservation of a sanitizing algorithm should be measured as the average error of a Bayesian decision maker, shouldn't Bayesian decision theory be a good choice when it comes to using the sanitized outputs for various purposes? We put this idea to the test in the unattributed histogram problem where our decision-theoretic post-processing algorithm empirically outperforms previously proposed approaches.

Original languageEnglish (US)
Title of host publicationSIGMOD 2013 - International Conference on Management of Data
Pages677-688
Number of pages12
DOIs
StatePublished - 2013
Event2013 ACM SIGMOD Conference on Management of Data, SIGMOD 2013 - New York, NY, United States
Duration: Jun 22 2013Jun 27 2013

Publication series

NameProceedings of the ACM SIGMOD International Conference on Management of Data
ISSN (Print)0730-8078

Other

Other2013 ACM SIGMOD Conference on Management of Data, SIGMOD 2013
Country/TerritoryUnited States
CityNew York, NY
Period6/22/136/27/13

All Science Journal Classification (ASJC) codes

  • Software
  • Information Systems

Fingerprint

Dive into the research topics of 'Information preservation in statistical privacy and Bayesian estimation of unattributed histograms'. Together they form a unique fingerprint.

Cite this