A step-by-step approach to improve data quality when using commercial business lists to characterize retail food environments

Kelly K. Jones, Shannon N. Zenk, Elizabeth Tarlov, Lisa M. Powell, Stephen A. Matthews, Irina Horoi

Research output: Contribution to journalArticlepeer-review

27 Scopus citations


Background: Food environment characterization in health studies often requires data on the location of food stores and restaurants. While commercial business lists are commonly used as data sources for such studies, current literature provides little guidance on how to use validation study results to make decisions on which commercial business list to use and how to maximize the accuracy of those lists. Using data from a retrospective cohort study [Weight And Veterans' Environments Study (WAVES)], we (a) explain how validity and bias information from existing validation studies (count accuracy, classification accuracy, locational accuracy, as well as potential bias by neighborhood racial/ethnic composition, economic characteristics, and urbanicity) were used to determine which commercial business listing to purchase for retail food outlet data and (b) describe the methods used to maximize the quality of the data and results of this approach. Methods: We developed data improvement methods based on existing validation studies. These methods included purchasing records from commercial business lists (InfoUSA and Dun and Bradstreet) based on store/restaurant names as well as standard industrial classification (SIC) codes, reclassifying records by store type, improving geographic accuracy of records, and deduplicating records. We examined the impact of these procedures on food outlet counts in US census tracts. Results: After cleaning and deduplicating, our strategy resulted in a 17.5% reduction in the count of food stores that were valid from those purchased from InfoUSA and 5.6% reduction in valid counts of restaurants purchased from Dun and Bradstreet. Locational accuracy was improved for 7.5% of records by applying street addresses of subsequent years to records with post-office (PO) box addresses. In total, up to 83% of US census tracts annually experienced a change (either positive or negative) in the count of retail food outlets between the initial purchase and the final dataset. Discussion: Our study provides a step-by-step approach to purchase and process business list data obtained from commercial vendors. The approach can be followed by studies of any size, including those with datasets too large to process each record by hand and will promote consistency in characterization of the retail food environment across studies.

Original languageEnglish (US)
Article number35
JournalBMC Research Notes
Issue number1
StatePublished - Jan 7 2017

All Science Journal Classification (ASJC) codes

  • General Biochemistry, Genetics and Molecular Biology


Dive into the research topics of 'A step-by-step approach to improve data quality when using commercial business lists to characterize retail food environments'. Together they form a unique fingerprint.

Cite this