Automatic construction of Chinese stop word list

  • Authors:
  • Feng Zou;Fu Lee Wang;Xiaotie Deng;Song Han;Lu Sheng Wang

  • Affiliations:
  • Computer Science Department, City University of Hong Kong, Kowloon Tong, Hong Kong;Computer Science Department, City University of Hong Kong, Kowloon Tong, Hong Kong;Computer Science Department, City University of Hong Kong, Kowloon Tong, Hong Kong;Computer Science Department, City University of Hong Kong, Kowloon Tong, Hong Kong;Computer Science Department, City University of Hong Kong, Kowloon Tong, Hong Kong

  • Venue:
  • ACOS'06 Proceedings of the 5th WSEAS international conference on Applied computer science
  • Year:
  • 2006

Quantified Score

Hi-index 0.00

Visualization

Abstract

In modern information retrieval systems, effective indexing can be achieved by removal of stop words. Till now many stop word lists have been developed for English language. However, no standard stop word list has been constructed for Chinese language yet. With the fast development of information retrieval in Chinese language, exploring Chinese stop word lists becomes critical. In this paper, to save the time and release the burden of manual stop word selection, we propose an automatic aggregated methodology based on statistical and information models for extraction of a stop word list in Chinese language. Result analysis shows that our stop list is comparable with a general English stop word list, and our list is much more general than other Chinese stop lists as well. Our stop word extraction algorithm is a promising technique, which saves the time for manual generation and constructs a standard. It could be applied into other languages in the future.