Unsupervised Discrimination of Person Names in Web Contexts

  • Authors:
  • Ted Pedersen;Anagha Kulkarni

  • Affiliations:
  • University of Minnesota, Duluth, MN 55812, USA;Carnegie Mellon University, Pittsburgh, PA 15213, USA

  • Venue:
  • CICLing '07 Proceedings of the 8th International Conference on Computational Linguistics and Intelligent Text Processing
  • Year:
  • 2009

Quantified Score

Hi-index 0.00

Visualization

Abstract

Ambiguous person names are a problem in many forms of written text, including that which is found on the Web. In this paper we explore the use of unsupervised clustering techniques to discriminate among entities named in Web pages. We examine three main issues via an extensive experimental study. First, the effect of using a held---out set of training data for feature selection versus using the data in which the ambiguous names occur. Second, the impact of using different measures of association for identifying lexical features. Third, the success of different cluster stopping measures that automatically determine the number of clusters in the data.