An Efficient Rigorous Approach for Identifying Statistically Significant Frequent Itemsets

Authors:
Adam Kirsch;Michael Mitzenmacher;Andrea Pietracaprina;Geppino Pucci;Eli Upfal;Fabio Vandin
Affiliations:
Harvard University;Harvard University;University of Padova;University of Padova;Brown University;Brown University
Venue:
Journal of the ACM (JACM)
Year:
2012

Citing 18
Cited 1

Mining association rules between sets of items in large databases

SIGMOD '93 Proceedings of the 1993 ACM SIGMOD international conference on Management of data
Mining quantitative association rules in large relational tables

SIGMOD '96 Proceedings of the 1996 ACM SIGMOD international conference on Management of data
A new framework for itemset generation

PODS '98 Proceedings of the seventeenth ACM SIGACT-SIGMOD-SIGART symposium on Principles of database systems
Data mining: concepts and techniques

Data mining: concepts and techniques
Empirical bayes screening for multi-item associations

Proceedings of the seventh ACM SIGKDD international conference on Knowledge discovery and data mining
Beyond Market Baskets: Generalizing Association Rules to Dependence Rules

Data Mining and Knowledge Discovery
What Makes Patterns Interesting in Knowledge Discovery Systems

IEEE Transactions on Knowledge and Data Engineering
Discovering Frequent Closed Itemsets for Association Rules

ICDT '99 Proceedings of the 7th International Conference on Database Theory
Determining Hit Rate in Pattern Search

Proceedings of the ESF Exploratory Workshop on Pattern Detection and Discovery
On the discovery of significant statistical quantitative rules

Proceedings of the tenth ACM SIGKDD international conference on Knowledge discovery and data mining
Probability and Computing: Randomized Algorithms and Probabilistic Analysis

Probability and Computing: Randomized Algorithms and Probabilistic Analysis
Fast discovery of unexpected patterns in data, relative to a Bayesian network

Proceedings of the eleventh ACM SIGKDD international conference on Knowledge discovery in data mining
Mining compressed frequent-pattern sets

VLDB '05 Proceedings of the 31st international conference on Very large data bases
Introduction to Data Mining, (First Edition)

Introduction to Data Mining, (First Edition)
On the effectiveness and efficiency of computing bounds on the support of item-sets in the frequent item-sets mining problem

Proceedings of the 1st international workshop on open source data mining: frequent pattern mining implementations
Assessing data mining results via swap randomization

Proceedings of the 12th ACM SIGKDD international conference on Knowledge discovery and data mining
Frequent pattern mining: current status and future directions

Data Mining and Knowledge Discovery
Efficient Discovery of Statistically Significant Association Rules

ICDM '08 Proceedings of the 2008 Eighth IEEE International Conference on Data Mining

Mining pure high-order word associations via information geometry for information retrieval

ACM Transactions on Information Systems (TOIS)

Quantified Score

Hi-index	0.00

Visualization

Abstract

As advances in technology allow for the collection, storage, and analysis of vast amounts of data, the task of screening and assessing the significance of discovered patterns is becoming a major challenge in data mining applications. In this work, we address significance in the context of frequent itemset mining. Specifically, we develop a novel methodology to identify a meaningful support threshold s* for a dataset, such that the number of itemsets with support at least s* represents a substantial deviation from what would be expected in a random dataset with the same number of transactions and the same individual item frequencies. These itemsets can then be flagged as statistically significant with a small false discovery rate. We present extensive experimental results to substantiate the effectiveness of our methodology.