Mining "Hidden phrase" definitions from the web

  • Authors:
  • Hung V. Nguyen;P. Velamuru;D. Kolippakkam;H. Davulcu;H. Liu;M. Ates

  • Affiliations:
  • Department of Computer Science and Engineering, Arizona State University, AZ;Department of Computer Science and Engineering, Arizona State University, AZ;Department of Computer Science and Engineering, Arizona State University, AZ;Department of Computer Science and Engineering, Arizona State University, AZ;Department of Computer Science and Engineering, Arizona State University, AZ;epartment of Computer Science and Engineering, NJ

  • Venue:
  • APWeb'03 Proceedings of the 5th Asia-Pacific web conference on Web technologies and applications
  • Year:
  • 2003

Quantified Score

Hi-index 0.00

Visualization

Abstract

Keyword searching is the most common form of document search on the Web. Many Web publishers manually annotate the META tags and titles of their pages with frequently queried phrases in order to improve their placement and ranking. A "hidden phrase" is defined as a phrase that occurs in the META tag of a Web page but not in its body. In this paper we present an algorithm that mines the definitions of hidden phrases from the Web documents. Phrase definitions allow (i) publishers to find relevant phrases with high query frequency, and, (ii) search engines to test if the content of the body of a document matches the phrases. We use co-occurrence clustering and association rule mining algorithms to learn phrase definitions from high-dimensional data sets. We also provide experimental results.