Ontology-based automatic classification of web documents

  • Authors:
  • MuHee Song;SooYeon Lim;DongJin Kang;SangJo Lee

  • Affiliations:
  • Department of Computer Engineering, Kyungpook National University, Daegu, Korea;Department of Computer Engineering, Kyungpook National University, Daegu, Korea;Information Technology Services, Kyungpook National University, Daegu, Korea;Department of Computer Engineering, Kyungpook National University, Daegu, Korea

  • Venue:
  • ICIC'06 Proceedings of the 2006 international conference on Intelligent computing: Part II
  • Year:
  • 2006

Quantified Score

Hi-index 0.00

Visualization

Abstract

The use of an ontology in order to provide a mechanism to enable machine reasoning has continuously increased during the last few years. This paper proposed an automated method for document classification using an ontology, which expresses terminology information and vocabulary contained in Web documents by way of a hierarchical structure. Ontology-based document classification involves determining document features that represent the Web documents most accurately, and classifying them into the most appropriate categories after analyzing their contents by using at least two pre-defined categories per given document features. In this paper, Web documents are classified in real time not with experimental data or a learning process, but by similarity calculations between the terminology information extracted from Web documents and ontology categories. This results in a more accurate document classification since the meanings and relationships unique to each document are determined.