TY - GEN
T1 - Automatic extraction of data points and text blocks from 2-dimensional plots in digital documents
AU - Kataria, Saurabh
AU - Browuer, William
AU - Mitra, Prasenjit
AU - Giles, C. Lee
PY - 2008/12/24
Y1 - 2008/12/24
N2 - Two dimensional plots (2-D) in digital documents on the web are an important source of information that is largely under-utilized. In this paper, we outline how data and text can be extracted automatically from these 2-D plots, thus eliminating a time consuming manual process. Our information extraction algorithm identifies the axes of the figures, extracts text blocks like axes-labels and legends and identifies data points in the figure. It also extracts the units appearing in the axes labels and segments the legends to identify the different lines in the legend, the different symbols and their associated text explanations. Our algorithm also performs the challenging task of separating out overlapping text and data points effectively. Our experiments indicate that these techniques are computationally efficient and provide acceptable accuracy.
AB - Two dimensional plots (2-D) in digital documents on the web are an important source of information that is largely under-utilized. In this paper, we outline how data and text can be extracted automatically from these 2-D plots, thus eliminating a time consuming manual process. Our information extraction algorithm identifies the axes of the figures, extracts text blocks like axes-labels and legends and identifies data points in the figure. It also extracts the units appearing in the axes labels and segments the legends to identify the different lines in the legend, the different symbols and their associated text explanations. Our algorithm also performs the challenging task of separating out overlapping text and data points effectively. Our experiments indicate that these techniques are computationally efficient and provide acceptable accuracy.
UR - http://www.scopus.com/inward/record.url?scp=57749191459&partnerID=8YFLogxK
UR - http://www.scopus.com/inward/citedby.url?scp=57749191459&partnerID=8YFLogxK
M3 - Conference contribution
AN - SCOPUS:57749191459
SN - 9781577353683
T3 - Proceedings of the National Conference on Artificial Intelligence
SP - 1169
EP - 1174
BT - AAAI-08/IAAI-08 Proceedings - 23rd AAAI Conference on Artificial Intelligence and the 20th Innovative Applications of Artificial Intelligence Conference
T2 - 23rd AAAI Conference on Artificial Intelligence and the 20th Innovative Applications of Artificial Intelligence Conference, AAAI-08/IAAI-08
Y2 - 13 July 2008 through 17 July 2008
ER -