TY - JOUR
T1 - Recommendations on compiling test datasets for evaluating artificial intelligence solutions in pathology
AU - Homeyer, André
AU - Geißler, Christian
AU - Schwen, Lars Ole
AU - Zakrzewski, Falk
AU - Evans, Theodore
AU - Strohmenger, Klaus
AU - Westphal, Max
AU - Bülow, Roman David
AU - Kargl, Michaela
AU - Karjauv, Aray
AU - Munné-Bertran, Isidre
AU - Retzlaff, Carl Orge
AU - Romero-López, Adrià
AU - Sołtysiński, Tomasz
AU - Plass, Markus
AU - Carvalho, Rita
AU - Steinbach, Peter
AU - Lan, Yu Chia
AU - Bouteldja, Nassim
AU - Haber, David
AU - Rojas-Carulla, Mateo
AU - Vafaei Sadr, Alireza
AU - Kraft, Matthias
AU - Krüger, Daniel
AU - Fick, Rutger
AU - Lang, Tobias
AU - Boor, Peter
AU - Müller, Heimo
AU - Hufnagl, Peter
AU - Zerbe, Norman
N1 - Publisher Copyright:
© 2022, The Author(s).
PY - 2022/12
Y1 - 2022/12
N2 - Artificial intelligence (AI) solutions that automatically extract information from digital histology images have shown great promise for improving pathological diagnosis. Prior to routine use, it is important to evaluate their predictive performance and obtain regulatory approval. This assessment requires appropriate test datasets. However, compiling such datasets is challenging and specific recommendations are missing. A committee of various stakeholders, including commercial AI developers, pathologists, and researchers, discussed key aspects and conducted extensive literature reviews on test datasets in pathology. Here, we summarize the results and derive general recommendations on compiling test datasets. We address several questions: Which and how many images are needed? How to deal with low-prevalence subsets? How can potential bias be detected? How should datasets be reported? What are the regulatory requirements in different countries? The recommendations are intended to help AI developers demonstrate the utility of their products and to help pathologists and regulatory agencies verify reported performance measures. Further research is needed to formulate criteria for sufficiently representative test datasets so that AI solutions can operate with less user intervention and better support diagnostic workflows in the future.
AB - Artificial intelligence (AI) solutions that automatically extract information from digital histology images have shown great promise for improving pathological diagnosis. Prior to routine use, it is important to evaluate their predictive performance and obtain regulatory approval. This assessment requires appropriate test datasets. However, compiling such datasets is challenging and specific recommendations are missing. A committee of various stakeholders, including commercial AI developers, pathologists, and researchers, discussed key aspects and conducted extensive literature reviews on test datasets in pathology. Here, we summarize the results and derive general recommendations on compiling test datasets. We address several questions: Which and how many images are needed? How to deal with low-prevalence subsets? How can potential bias be detected? How should datasets be reported? What are the regulatory requirements in different countries? The recommendations are intended to help AI developers demonstrate the utility of their products and to help pathologists and regulatory agencies verify reported performance measures. Further research is needed to formulate criteria for sufficiently representative test datasets so that AI solutions can operate with less user intervention and better support diagnostic workflows in the future.
UR - https://www.scopus.com/pages/publications/85138826992
UR - https://www.scopus.com/pages/publications/85138826992#tab=citedBy
U2 - 10.1038/s41379-022-01147-y
DO - 10.1038/s41379-022-01147-y
M3 - Review article
C2 - 36088478
AN - SCOPUS:85138826992
SN - 0893-3952
VL - 35
SP - 1759
EP - 1769
JO - Modern Pathology
JF - Modern Pathology
IS - 12
ER -