Visual extraction of information from web pages

  • Authors:
  • Giuseppe Della Penna;Daniele Magazzeni;Sergio Orefice

  • Affiliations:
  • Department of Computer Science, University of L'Aquila, Via Vetoio, I-67100 Coppito, L'Aquila, Italy;Department of Computer Science, University of L'Aquila, Via Vetoio, I-67100 Coppito, L'Aquila, Italy;Department of Computer Science, University of L'Aquila, Via Vetoio, I-67100 Coppito, L'Aquila, Italy

  • Venue:
  • Journal of Visual Languages and Computing
  • Year:
  • 2010

Quantified Score

Hi-index 0.00

Visualization

Abstract

In this paper we present a graphical software system that provides an automatic support to the extraction of information from web pages. The underlying extraction technique exploits the visual appearance of the information in the document, and is driven by the spatial relations occurring among the elements in the page. However, the usual information extraction modalities based on the web page structure can be used in our framework, too. The technique has been integrated within the Spatial Relation Query (SRQ) tool. The tool is provided with a graphical front-end which allows one to define and manage a library of spatial relations, and to use a SQL-like language for composing queries driven by these relations and by further semantic and graphical attributes.