Learning-Based Approaches for Matching Web Data Entities

  • Authors:
  • Hanna Kopcke;Andreas Thor;Erhard Rahm

  • Affiliations:
  • University of Leipzig;University of Leipzig;University of Leipzig

  • Venue:
  • IEEE Internet Computing
  • Year:
  • 2010

Quantified Score

Hi-index 0.00

Visualization

Abstract

Entity matching is a key task for data integration and especially challenging for Web data. Effective entity matching typically requires combining several match techniques and finding suitable configuration parameters, such as similarity thresholds. The authors investigate to what degree machine learning helps semi-automatically determine suitable match strategies with a limited amount of manual training effort. They use a new framework, Fever, to evaluate several learning-based approaches for matching different sets of Web data entities. In particular, they study different approaches for training-data selection and how much training is needed to find effective combined match strategies and configurations.