Automated analysis of images in documents for intelligent document search

Authors:
Xiaonan Lu;Saurabh Kataria;William J. Brouwer;James Z. Wang;Prasenjit Mitra;C. Lee Giles
Affiliations:
The Pennsylvania State University, Department of Computer Science and Engineering, University Park, USA;The Pennsylvania State University, College of Information Sciences and Technology, University Park, USA;The Pennsylvania State University, Department of Chemistry, University Park, USA;The Pennsylvania State University, Department of Computer Science and Engineering, University Park, USA and The Pennsylvania State University, College of Information Sciences and Technology, Unive ...;The Pennsylvania State University, Department of Computer Science and Engineering, University Park, USA and The Pennsylvania State University, College of Information Sciences and Technology, Unive ...;The Pennsylvania State University, Department of Computer Science and Engineering, University Park, USA and The Pennsylvania State University, College of Information Sciences and Technology, Unive ...
Venue:
International Journal on Document Analysis and Recognition
Year:
2009

Citing 0
Cited 4

Towards automatic image annotation supporting document understanding

HAIS'11 Proceedings of the 6th international conference on Hybrid artificial intelligent systems - Volume Part I
Model-based chart image classification

ISVC'11 Proceedings of the 7th international conference on Advances in visual computing - Volume Part II
Patent image retrieval: a survey

Proceedings of the 4th workshop on Patent information retrieval
A novel figure panel classification and extraction method for document image understanding

International Journal of Data Mining and Bioinformatics

Quantified Score

Hi-index	0.00

Visualization

Abstract

Authors use images to present a wide variety of important information in documents. For example, two-dimensional (2-D) plots display important data in scientific publications. Often, end-users seek to extract this data and convert it into a machine-processible form so that the data can be analyzed automatically or compared with other existing data. Existing document data extraction tools are semi-automatic and require users to provide metadata and interactively extract the data. In this paper, we describe a system that extracts data from documents fully automatically, completely eliminating the need for human intervention. The system uses a supervised learning-based algorithm to classify figures in digital documents into five classes: photographs, 2-D plots, 3-D plots, diagrams, and others. Then, an integrated algorithm is used to extract numerical data from data points and lines in the 2-D plot images along with the axes and their labels, the data symbols in the figure’s legend and their associated labels. We demonstrate that the proposed system and its component algorithms are effective via an empirical evaluation. Our data extraction system has the potential to be a vital component in high volume digital libraries.