Studying the XML Web: Gathering Statistics from an XML Sample

  • Authors:
  • Denilson Barbosa;Laurent Mignet;Pierangelo Veltri

  • Affiliations:
  • Department of Computer Science, University of Toronto, Toronto, Canada M5S 3G5;IBM India Research Laboratory, New Delhi, India 110016;Department of Experimental and Clinical Medicine, Magna Graecia University of Catanzaro, Catanzaro, Italy 88100

  • Venue:
  • World Wide Web
  • Year:
  • 2005

Quantified Score

Hi-index 0.00

Visualization

Abstract

XML has emerged as the language for exchanging data on the web and has attracted considerable interest both in industry and in academia. Nevertheless, to date, little is known about the XML documents published on the web. This paper presents a comprehensive analysis of a sample of about 200,000 XML documents on the web, and is the first study of its kind. We study the distribution of XML documents across the web in several ways; moreover, we provided a detailed characterization of the structure of real XML documents. Our results provide valuable input to the design of algorithms, tools and systems that use XML in one form or another.