Hierarchical clustering for topic analysis based on variable feature selection

  • Authors:
  • Jianping Zeng;Linghui Gong;Qinqin Wang;Chengrong Wu

  • Affiliations:
  • School of Computer Science, Fudan University, Shanghai, P. R. China;School of Computer Science, Fudan University, Shanghai, P. R. China;School of Computer Science, Fudan University, Shanghai, P. R. China;School of Computer Science, Fudan University, Shanghai, P. R. China

  • Venue:
  • FSKD'09 Proceedings of the 6th international conference on Fuzzy systems and knowledge discovery - Volume 7
  • Year:
  • 2009

Quantified Score

Hi-index 0.00

Visualization

Abstract

Hierarchical topic structure can express topics in a natural way which is more reasonable for human machine interface. However, the hierarchical topic structure that is extracted by most of the topic analysis algorithms can not present a meaningful description for all subtopics in the hierarchical tree. We propose a new hierarchical clustering algorithm based on variable feature selection for each level in the hierarchical structure. The algorithm employs a top-down strategy to extract subtopics and setups the relation between topics in neighbor levels based on common documents number. The number of the levels in the hierarchical structure is determined by the frequency of the selected word feature. Experiments on a real world dataset which is collected from a news website shows that the proposed algorithm can generate more meaningful topic structure, by comparing to the current hierarchical topic clustering algorithms.