Latex: a document preparation system
Latex: a document preparation system
IBM Systems Journal
Document degradation models and a methodology for degradation model validation
Document degradation models and a methodology for degradation model validation
Digital Image Warping
TEX: The Program
Automatic Generation of Character Groundtruth for Scanned Documents: A Closed-Loop Approach
ICPR '96 Proceedings of the International Conference on Pattern Recognition (ICPR '96) Volume III-Volume 7276 - Volume 7276
Twenty Years of Document Image Analysis in PAMI
IEEE Transactions on Pattern Analysis and Machine Intelligence
Issues in Ground-Truthing Graphic Documents
GREC '01 Selected Papers from the Fourth International Workshop on Graphics Recognition Algorithms and Applications
A Parallel-Line Detection Algorithm Based on HMM Decoding
IEEE Transactions on Pattern Analysis and Machine Intelligence
EURASIP Journal on Applied Signal Processing
Effect of OCR error correction on Arabic retrieval
Information Retrieval
Databases and competitions: strategies to improve Arabic recognition systems
SACH'06 Proceedings of the 2006 conference on Arabic and Chinese handwriting recognition
Hi-index | 0.14 |
Character groundtruth for real, scanned document images is crucial for evaluating the performance of OCR systems, training OCR algorithms, and validating document degradation models. Unfortunately, manual collection of accurate groundtruth for characters in a real (scanned) document image is not practical because (i) accuracy in delineating groundtruth character bounding boxes is not high enough, (ii) it is extremely laborious and time consuming, and (iii) the manual labor required for this task is prohibitively expensive. In this paper we describe a closed-loop methodology for collecting very accurate groundtruth for scanned documents. We first create ideal documents using a typesetting language. Next we create the groundtruth for the ideal document. The ideal document is then printed, photocopied and then scanned. A registration algorithm estimates the global geometric transformation and then performs a robust local bitmap match to register the ideal document image to the scanned document image. Finally, groundtruth associated with the ideal document image is transformed using the estimated geometric transformation to create the groundtruth for the scanned document image. This methodology is very general and can be used for creating groundtruth for documents in typeset in any language, layout, font, and style. We have demonstrated the method by generating groundtruth for English, Hindi, and FAX document images. The cost of creating groundtruth using our methodology is minimal. If character, word or zone groundtruth is available for any real document, the registration algorithm can be used to generate the corresponding groundtruth for a rescanned version of the document.