Arabic Named Entity Recognition: A Feature-Driven Study

Authors:
Y. Benajiba;M. Diab;P. Rosso
Affiliations:
Dept. of Informatic Syst., Polytech. Univ. of Valencia, Valencia;-;-
Venue:
IEEE Transactions on Audio, Speech, and Language Processing
Year:
2009

Citing 0
Cited 5

Improving mention detection robustness to noisy input

EMNLP '10 Proceedings of the 2010 Conference on Empirical Methods in Natural Language Processing
A real time Named Entity Recognition system for Arabic text mining

Language Resources and Evaluation
A hybrid approach to Arabic named entity recognition

Journal of Information Science
Aligned-Parallel-Corpora Based Semi-Supervised Learning for Arabic Mention Detection

IEEE/ACM Transactions on Audio, Speech and Language Processing (TASLP)
Crime profiling for the Arabic language using computational linguistic techniques

Information Processing and Management: an International Journal

Quantified Score

Hi-index	0.00

Visualization

Abstract

The named entity recognition task aims at identifying and classifying named entities within an open-domain text. This task has been garnering significant attention recently as it has been shown to help improve the performance of many natural language processing applications. In this paper, we investigate the impact of using different sets of features in three discriminative machine learning frameworks, namely, support vector machines, maximum entropy and conditional random fields for the task of named entity recognition. Our language of interest is Arabic. We explore lexical, contextual and morphological features and nine data-sets of different genres and annotations. We measure the impact of the different features in isolation and incrementally combine them in order to evaluate the robustness to noise of each approach. We achieve the highest performance using a combination of 15 features in conditional random fields using broadcast news data (Fbeta = 1=83.34).