Just-in-time recovery of missing web pages

  • Authors:
  • Terry L. Harrison;Michael L. Nelson

  • Affiliations:
  • Old Dominion University, Norfolk, VA;Old Dominion University, Norfolk, VA

  • Venue:
  • Proceedings of the seventeenth conference on Hypertext and hypermedia
  • Year:
  • 2006

Quantified Score

Hi-index 0.00

Visualization

Abstract

We present Opal, a light-weight framework for interactively locating missing web pages (http status code 404). Opal is an example of "in vivo" preservation: harnessing the collective behavior of web archives, commercial search engines, and research projects for the purpose of preservation. Opal servers learn from their experiences and are able to share their knowledge with other Opal servers by mutual harvesting using the Open Archives Initiative Protocol for Metadata Harvesting (OAI-PMH). Using cached copies that can be found on the web, Opal creates lexical signatures which are then used to search for similar versions of the web page. We present the architecture of the Opal framework, discuss a reference implementation of the framework, and present a quantitative analysis of the framework that indicates that Opal could be effectively deployed.