Estimating the Change of Web Pages

  • Authors:
  • Sung Jin Kim;Sang Ho Lee

  • Affiliations:
  • Department of Computer Science, University of California, Los Angeles, USA;School of Computing, Soongsil University, Seoul, Korea

  • Venue:
  • ICCS '07 Proceedings of the 7th international conference on Computational Science, Part III: ICCS 2007
  • Year:
  • 2007

Quantified Score

Hi-index 0.00

Visualization

Abstract

This paper presents the estimation methods computing the probabilities of how many times web pages are downloaded and modified, respectively, in the future crawls. The methods can make web database administrators avoid unnecessarily requesting undownloadable and unmodified web pages in a page group. We postulate that the change behavior of web pages is strongly related to the past change behavior. We gather the change histories of approximately three million web pages at two-day intervals for 100 days, and estimate the future change behavior of those pages. Our estimation, which was evaluated by actual change behavior of the pages, worked well.