Building a large syntactically-annotated corpus of Vietnamese

  • Authors:
  • Phuong-Thai Nguyen;Xuan-Luong Vu;Thi-Minh-Huyen Nguyen;Van-Hiep Nguyen;Hong-Phuong Le

  • Affiliations:
  • College of Technology, VNU;Vietnam Lexicography Centre;University of Natural Sciences, VNU;University of Social Sciences and Humanities, VNU;LORIA/INRIA Lorraine

  • Venue:
  • ACL-IJCNLP '09 Proceedings of the Third Linguistic Annotation Workshop
  • Year:
  • 2009

Quantified Score

Hi-index 0.00

Visualization

Abstract

Treebank is an important resource for both research and application of natural language processing. For Vietnamese, we still lack such kind of corpora. This paper presents up-to-date results of a project for Vietnamese treebank construction. Since Vietnamese is an isolating language and has no word delimiter, there are many ambiguities in sentence analysis. We systematically applied a lot of linguistic techniques to handle such ambiguities. Annotators are supported by automatic-labeling tools and a tree-editor tool. Raw texts are extracted from Tuoi Tre (Youth), an online Vietnamese daily newspaper. The current annotation agreement is around 90 percent.