A belief networks-based generative model for structured documents: an application to the XML categorization

  • Authors:
  • Ludovic Denoyer;Patrick Gallinari

  • Affiliations:
  • Laboratoire d'Informatique de Paris VI, France;Laboratoire d'Informatique de Paris VI, France

  • Venue:
  • MLDM'03 Proceedings of the 3rd international conference on Machine learning and data mining in pattern recognition
  • Year:
  • 2003

Quantified Score

Hi-index 0.00

Visualization

Abstract

We present a generative Bayesian model for the modeling of structured (e.g. XML) documents. This model allows us to simultaneously take into account structure and content information. It is used here for classifying XML documents. We adopt a machine learning approach and the model parameters are learned from a labeled training set of representative documents. We discuss the role of structural information for classification and describe experiments on a small collection of class labeled structured documents. We also present preliminary results showing how this model could classify documents with DTDs not represented in the training set.