A Robust Algorithm for Text String Separation from Mixed Text/Graphics Images

  • Authors:
  • Lloyd A. Fletcher;Rangachar Kasturi

  • Affiliations:
  • The Pennsylvania State Univ., Philadelphia;The Pennsylvania State Univ., Philadelphia

  • Venue:
  • IEEE Transactions on Pattern Analysis and Machine Intelligence
  • Year:
  • 1988

Quantified Score

Hi-index 0.15

Visualization

Abstract

The development and implementation of an algorithm for automated text string separation that is relatively independent of changes in text font style and size and of string orientation are described. It is intended for use in an automated system for document analysis. The principal parts of the algorithm are the generation of connected components and the application of the Hough transform in order to group components into logical character strings that can then be separated from the graphics. The algorithm outputs two images, one containing text strings and the other graphics. These images can then be processed by suitable character recognition and graphics recognition systems. The performance of the algorithm, both in terms of its effectiveness and computational efficiency, was evaluated using several test images and showed superior performance compared to other techniques.