A two-stage feature selection method for text categorization

  • Authors:
  • Jiana Meng;Hongfei Lin;Yuhai Yu

  • Affiliations:
  • College of Science, Dalian Nationalities University, Dalian 116600, China and Department of Computer Science and Engineering, Dalian University of Technology, Dalian 116024, China;Department of Computer Science and Engineering, Dalian University of Technology, Dalian 116024, China;Department of Computer Science and Engineering, Dalian University of Technology, Dalian 116024, China and School of Computer Science & Engineering, Dalian Nationalities University, Dalian 116600, ...

  • Venue:
  • Computers & Mathematics with Applications
  • Year:
  • 2011

Quantified Score

Hi-index 0.09

Visualization

Abstract

Feature selection for text categorization is a well-studied problem and its goal is to improve the effectiveness of categorization, or the efficiency of computation, or both. The system of text categorization based on traditional term-matching is used to represent the vector space model as a document; however, it needs a high dimensional space to represent the document, and does not take into account the semantic relationship between terms, which leads to a poor categorization accuracy. The latent semantic indexing method can overcome this problem by using statistically derived conceptual indices to replace the individual terms. With the purpose of improving the accuracy and efficiency of categorization, in this paper we propose a two-stage feature selection method. Firstly, we apply a novel feature selection method to reduce the dimension of terms; and then we construct a new semantic space, between terms, based on the latent semantic indexing method. Through some applications involving the spam database categorization, we find that our two-stage feature selection method performs better.