Large linguistically-processed web corpora for multiple languages

  • Authors:
  • Marco Baroni;Adam Kilgarriff

  • Affiliations:
  • University of Bologna, Italy;Lexical Computing Ltd. and University of Sussex, Brighton, UK

  • Venue:
  • EACL '06 Proceedings of the Eleventh Conference of the European Chapter of the Association for Computational Linguistics: Posters & Demonstrations
  • Year:
  • 2006

Quantified Score

Hi-index 0.00

Visualization

Abstract

The Web contains vast amounts of linguistic data. One key issue for linguists and language technologists is how to access it. Commercial search engines give highly compromised access. An alternative is to crawl the Web ourselves, which also allows us to remove duplicates and near-duplicates, navigational material, and a range of other kinds of non-linguistic matter. We can also tokenize, lemmatise and part-of-speech tag the corpus, and load the data into a corpus query tool which supports sophisticated linguistic queries. We have now done this for German and Italian, with corpus sizes of over 1 billion words in each case. We provide Web access to the corpora in our query tool, the Sketch Engine.