Personal Name Resolution Crossover Documents by a Semantics-Based Approach

  • Authors:
  • Xuan-Hieu Phan;Le-Minh Nguyen;Susumu Horiguchi

  • Affiliations:
  • The authors are with the Graduate School of Infor. Science, Japan Advanced Institute of Science and Technology, Nomi-shi, 923--1292 Japan. E-mail: hieuxuan@jaist.ac.jp,;The authors are with the Graduate School of Infor. Science, Japan Advanced Institute of Science and Technology, Nomi-shi, 923--1292 Japan. E-mail: hieuxuan@jaist.ac.jp,;The authors are with the Graduate School of Infor. Science, Japan Advanced Institute of Science and Technology, Nomi-shi, 923--1292 Japan. E-mail: hieuxuan@jaist.ac.jp,

  • Venue:
  • IEICE - Transactions on Information and Systems
  • Year:
  • 2006

Quantified Score

Hi-index 0.00

Visualization

Abstract

Cross-document personal name resolution is the process of identifying whether or not a common personal name mentioned in different documents refers to the same individual. Most previous approaches usually rely on lexical matching such as the occurrence of common words surrounding the entity name to measure the similarity between documents, and then clusters the documents according to their referents. In spite of certain successes, measuring similarity based on lexical comparison sometimes ignores important linguistic phenomena at the semantic level such as synonym or paraphrase. This paper presents a semantics-based approach to the resolution of personal name crossover documents that can make the most of both lexical evidences and semantic clues. In our method, the similarity values between documents are determined by estimating the semantic relatedness between words. Further, the semantic labels attached to sentences allow us to highlight the common personal facts that are potentially available among documents. An evaluation on three web datasets demonstrates that our method achieves the better performance than the previous work.