Error anaylsis of Chinese text segmentation using statistical approach

Authors:
Christopher C. Yang;Kar Wing Li
Affiliations:
The Chinese University of Hong Kong, Hong Kong;The Chinese University of Hong Kong, Hong Kong
Venue:
Proceedings of the 4th ACM/IEEE-CS joint conference on Digital libraries
Year:
2004

Citing 2
Cited 1

Combination and boundary detection approaches on Chinese indexing

Journal of the American Society for Information Science - Special topic issue on digital libraries: part 2
Using statistical and contextual information to identify two-and three-character words in Chinese text

Journal of the American Society for Information Science and Technology

Cross-lingual text categorization: Conquering language boundaries in globalized environments

Information Processing and Management: an International Journal

Quantified Score

Hi-index	0.00

Visualization

Abstract

The Chinese text segmentation is important for the indexing of Chinese documents, which has significant impact on the performance of Chinese information retrieval. The statistical approach overcomes the limitations of the dictionary based approach. The statistical approach is developed by utilizing the statistical information about the association of adjacent characters in Chinese text collected from the Chinese corpus Both known words and unknown words can be segmented by the statistical approach. However, errors may occur due to the limitation of the corpus. In this work, we have conducted the error analysis of two Chinese text segmentation techniques using statistical approach, namely, boundary detection and heuristic method Such error analysis is useful for the future development of the automatic text segmentation of Chinese text or other text in oriental languages. It is also helpful to understand the impact of these errors on the information retrieval system in digital libraries.