A First Experience in Archiving the French Web

  • Authors:
  • Serge Abiteboul;Gregory Cobena;Julien Masanes;Gerald Sedrati

  • Affiliations:
  • -;-;-;-

  • Venue:
  • ECDL '02 Proceedings of the 6th European Conference on Research and Advanced Technology for Digital Libraries
  • Year:
  • 2002

Quantified Score

Hi-index 0.00

Visualization

Abstract

The web is a more and more valuable source of information and organizations are involved in archiving (portions of) it for various purposes, e.g., the Internet Archive www.archive.org. A new mission of the French National Library (BnF) is the "d茅p么t l茅gal" (legal deposit) of the French web. We describe here some preliminary work on the topic conducted by BnF and INRIA. In particular, we consider the acquisition of the web archive. Issues are the definition of the perimeter of the French web and the choice of pages to read once or more times (to take changes into account). When several copies of the same page are kept, this leads to versioning issues that we briefly consider. Finally, we mention some first experiments.