Active machine learning technique for named entity recognition

  • Authors:
  • Asif Ekbal;Sriparna Saha;Dhirendra Singh

  • Affiliations:
  • IIT Patna, Patna, India;IIT Patna, Patna, India;IIT Patna, Patna, India

  • Venue:
  • Proceedings of the International Conference on Advances in Computing, Communications and Informatics
  • Year:
  • 2012

Quantified Score

Hi-index 0.00

Visualization

Abstract

One difficulty with machine learning for information extraction is the high cost of collecting labeled examples. Active Learning can make more efficient use of the learner's time by asking them to label only instances that are most useful for the trainer. In random sampling approach, unlabeled data is selected for annotation at random and thus can't yield the desired results. In contrast, active learning selects the useful data from a huge pool of unlabeled data for the classifier. The strategies used often classify the corpus tokens (or, data points) into wrong classes. The classifier is confused between two categories if the token is located near the margin. We propose a novel method for solving this problem and show that it favorably results in the increased performance. Our approach is based on the supervised machine learning algorithm, namely Support Vector Machine (SVM). The proposed approach is applied for solving the problem of named entity recognition (NER) in two Indian languages, namely Hindi and Bengali. Results show that proposed active learning based technique indeed improves the performance of the system.