A new classification of datasets for frequent itemsets

  • Authors:
  • Frédéric Flouvat;Fabien Marchi;Jean-Marc Petit

  • Affiliations:
  • University of New Caledonia, PPME, Nouméa, New Caledonia 98851;Université de Lyon, Université Lyon 1, LIRIS, UMR5205 CNRS, Lyon, France 69621;Université de Lyon, INSA-Lyon, LIRIS, UMR5205 CNRS, Lyon, France 69621

  • Venue:
  • Journal of Intelligent Information Systems
  • Year:
  • 2010

Quantified Score

Hi-index 0.00

Visualization

Abstract

The discovery of frequent patterns is a famous problem in data mining. While plenty of algorithms have been proposed during the last decade, only a few contributions have tried to understand the influence of datasets on the algorithms behavior. Being able to explain why certain algorithms are likely to perform very well or very poorly on some datasets is still an open question. In this setting, we describe a thorough experimental study of datasets with respect to frequent itemsets. We study the distribution of frequent itemsets with respect to itemsets size together with the distribution of three concise representations: frequent closed, frequent free and frequent essential itemsets. For each of them, we also study the distribution of their positive and negative borders whenever possible. The main outcome of these experiments is a new classification of datasets invariant w.r.t. minsup variations and robust to explain efficiency of several implementations.