Link proximity analysis: clustering websites by examining link proximity

  • Authors:
  • Bela Gipp;Adriana Taylor;Jöran Beel

  • Affiliations:
  • UC Berkeley, Berkeley, California and Otto-von-Guericke University, Computer Science, ITI, VLBA-Lab, Magdeburg, Germany;UC Berkeley, Berkeley, California;UC Berkeley, Berkeley, California and Otto-von-Guericke University, Computer Science, ITI, VLBA-Lab, Magdeburg, Germany

  • Venue:
  • ECDL'10 Proceedings of the 14th European conference on Research and advanced technology for digital libraries
  • Year:
  • 2010

Quantified Score

Hi-index 0.00

Visualization

Abstract

This research-in-progress paper presents a new approach called Link Proximity Analysis (LPA) for identifying related web pages based on link analysis. In contrast to current techniques, which ignore intra-page link analysis, the one put forth here examines the relative positioning of links to each other within websites. The approach uses the fact that a clear correlation between the proximity of links to each other and the subject-relatedness of the linked websites can be observed on nearly every web page. By statistically analyzing this relationship and measuring the amount of sentences, paragraphs, etc. between two links, related websites can be automatically, identified as a first study has proven.