A comparative investigation of morphological language modeling for the languages of the European union

  • Authors:
  • Thomas Müller;Hinrich Schütze;Helmut Schmid

  • Affiliations:
  • University of Stuttgart, Germany;University of Stuttgart, Germany;University of Stuttgart, Germany

  • Venue:
  • NAACL HLT '12 Proceedings of the 2012 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies
  • Year:
  • 2012

Quantified Score

Hi-index 0.00

Visualization

Abstract

We investigate a language model that combines morphological and shape features with a Kneser-Ney model and test it in a large crosslingual study of European languages. Even though the model is generic and we use the same architecture and features for all languages, the model achieves reductions in perplexity for all 21 languages represented in the Europarl corpus, ranging from 3% to 11%. We show that almost all of this perplexity reduction can be achieved by identifying suffixes by frequency.