A Wikipedia-based corpus reference tool

Authors:
Jason Ginsburg
Affiliations:
University of Aizu, Tsuruga, Ikki-machi, Aizu-Wakamatsu City, Fukushima, Japan
Venue:
Proceedings of the 2012 Joint International Conference on Human-Centered Computer Environments
Year:
2012

Citing 4
Cited 0

WordNet: a lexical database for English

Communications of the ACM
Building a large annotated corpus of English: the penn treebank

Computational Linguistics - Special issue on using large corpora: II
Feature-rich part-of-speech tagging with a cyclic dependency network

NAACL '03 Proceedings of the 2003 Conference of the North American Chapter of the Association for Computational Linguistics on Human Language Technology - Volume 1
Natural Language Processing with Python

Natural Language Processing with Python

Quantified Score

Hi-index	0.00

Visualization

Abstract

This paper describes a dictionary-like reference tool that is designed to help users find information that is similar to what one would find in a dictionary when looking up a word, except that this information is extracted automatically from large corpora. For a particular vocabulary item, a user can view frequency information, part-of-speech distribution, word-forms, definitions, example paragraphs and collocations. All of this information is extracted automatically from corpora and most of this information is extracted from Wikipedia. Since Wikipedia is a massive corpus covering a diverse range of general topics, this information is probably very representative of how target words are used in general. This project has applications for English language teachers and learners, as well as for language researchers.