Automatic categorization of figures in scientific documents

Xiaonan Lu, Prasenjit Mitra, James Z. Wang, C. Lee Giles

Research output: Chapter in Book/Report/Conference proceedingConference contribution

25 Scopus citations

Abstract

Figures are very important non-textual information contained in scientific documents. Current digital libraries do not provide users tools to retrieve documents based on the information available within the figures. We propose an architecture for retrieving documents by integrating figures and other information. The initial step in enabling integrated document search is to categorize figures into a set of pre-defined types. We propose several categories of figures based on their functionalities in scholarly articles. We have developed a machine-learning-based approach for automatic categorization of figures. Both global features, such as texture, and part features, such as lines, are utilized in the architecture for discriminating among figure categories. The proposed approach has been evaluated on a testbed document set collected from the CiteSeer scientific literature digital library. Experimental evaluation has demonstrated that our algorithms can produce acceptable results for real-world use. Our tools will be integrated into a scientific-document digital library.

Original languageEnglish (US)
Title of host publication6th ACM/IEEE-CS Joint Conference on Digital Libraries 2006
Subtitle of host publicationOpening Information Horizons, JCDL '06
Pages129-138
Number of pages10
DOIs
StatePublished - 2006
Event6th ACM/IEEE-CS Joint Conference on Digital Libraries 2006: Opening Information Horizons, JCDL '06 - Chapel Hill, NC, United States
Duration: Jun 11 2006Jun 15 2006

Publication series

NameProceedings of the ACM/IEEE Joint Conference on Digital Libraries
Volume2006
ISSN (Print)1552-5996

Other

Other6th ACM/IEEE-CS Joint Conference on Digital Libraries 2006: Opening Information Horizons, JCDL '06
Country/TerritoryUnited States
CityChapel Hill, NC
Period6/11/066/15/06

All Science Journal Classification (ASJC) codes

  • General Engineering

Fingerprint

Dive into the research topics of 'Automatic categorization of figures in scientific documents'. Together they form a unique fingerprint.

Cite this