TY - JOUR
T1 - Government websites as data
T2 - a methodological pipeline with application to the websites of municipalities in the United States
AU - Neumann, Markus
AU - Linder, Fridolin
AU - Desmarais, Bruce
N1 - Publisher Copyright:
© 2021 Taylor & Francis.
PY - 2022
Y1 - 2022
N2 - The content of a government’s website is an important source of information about policy priorities, procedures, and services. Existing research on government websites has relied on manual methods of website content collection and processing, which imposes cost limitations on the scale of website data collection. In this research note, we propose that the automated collection of website content from large samples of government websites can offer relief from the costs of manual collection, and enable contributions through large-scale comparative analyses. We also provide software to ease the use of this data collection method. In an illustrative application, we collect textual content from the websites of over two hundred municipal governments in the United States, and study how website content is associated with mayoral partisanship. Using statistical topic modeling, we find that the partisanship of the mayor predicts differences in the contents of city websites that align with differences in the platforms of Democrats and Republicans. The application illustrates the utility of website content data extracted via our methodological pipeline.
AB - The content of a government’s website is an important source of information about policy priorities, procedures, and services. Existing research on government websites has relied on manual methods of website content collection and processing, which imposes cost limitations on the scale of website data collection. In this research note, we propose that the automated collection of website content from large samples of government websites can offer relief from the costs of manual collection, and enable contributions through large-scale comparative analyses. We also provide software to ease the use of this data collection method. In an illustrative application, we collect textual content from the websites of over two hundred municipal governments in the United States, and study how website content is associated with mayoral partisanship. Using statistical topic modeling, we find that the partisanship of the mayor predicts differences in the contents of city websites that align with differences in the platforms of Democrats and Republicans. The application illustrates the utility of website content data extracted via our methodological pipeline.
UR - http://www.scopus.com/inward/record.url?scp=85120740429&partnerID=8YFLogxK
UR - http://www.scopus.com/inward/citedby.url?scp=85120740429&partnerID=8YFLogxK
U2 - 10.1080/19331681.2021.1999880
DO - 10.1080/19331681.2021.1999880
M3 - Article
AN - SCOPUS:85120740429
SN - 1933-1681
VL - 19
SP - 411
EP - 422
JO - Journal of Information Technology and Politics
JF - Journal of Information Technology and Politics
IS - 4
ER -