Unsupervised Feature Generation using Knowledge Repositories for Effective Text Categorization

  • Authors:
  • Rajendra Prasath;Sudeshna Sarkar

  • Affiliations:
  • Norwegian University of Science and Technology, No --7491, Trondheim, Norway, email: rajendra@idi.ntnu.no/ rajendra@cse.iitkgp.ernet.in;Indian Institute of Technology, Kharagpur --721 302, India, email: sudeshna@cse.iitkgp.ernet.in

  • Venue:
  • Proceedings of the 2010 conference on ECAI 2010: 19th European Conference on Artificial Intelligence
  • Year:
  • 2010

Quantified Score

Hi-index 0.00

Visualization

Abstract

We propose an unsupervised feature generation algorithm using the repositories of human knowledge for effective text categorization. Conventional bag of words (BOW) depends on the presence / absence of keywords to classify the documents. To understand the actual context behind these keywords, we use knowledge concepts / hyperlinks from external knowledge sources through content and structure mining on Wikipedia. Then, the features of knowledge concepts are clustered to generate knowledge cluster vectors with which the input text documents are mapped into a high dimensional feature space and the classification is performed. The simulation results show that the proposed approach identifies associated features in the text collection and yields an improved classification accuracy.