Data mining using relational database management systems

Authors:
Beibei Zou;Xuesong Ma;Bettina Kemme;Glen Newton;Doina Precup
Affiliations:
McGill University, Montreal, Canada;McGill University, Montreal, Canada;McGill University, Montreal, Canada;National Research Council, Canada;McGill University, Montreal, Canada
Venue:
PAKDD'06 Proceedings of the 10th Pacific-Asia conference on Advances in Knowledge Discovery and Data Mining
Year:
2006

Citing 9
Cited 5

Integrating association rule mining with relational database systems: alternatives and implications

SIGMOD '98 Proceedings of the 1998 ACM SIGMOD international conference on Management of data
Data preparation for data mining

Data preparation for data mining
Squashing flat files flatter

KDD '99 Proceedings of the fifth ACM SIGKDD international conference on Knowledge discovery and data mining
Database Mining: A Performance Perspective

IEEE Transactions on Knowledge and Data Engineering
RainForest - A Framework for Fast Decision Tree Construction of Large Datasets

VLDB '98 Proceedings of the 24rd International Conference on Very Large Data Bases
SPRINT: A Scalable Parallel Classifier for Data Mining

VLDB '96 Proceedings of the 22th International Conference on Very Large Data Bases
Hyperspectral Image Analysis Using Genetic Programming

GECCO '02 Proceedings of the Genetic and Evolutionary Computation Conference
Programming the K-means clustering algorithm in SQL

Proceedings of the tenth ACM SIGKDD international conference on Knowledge discovery and data mining
Cached sufficient statistics for efficient machine learning with large datasets

Journal of Artificial Intelligence Research

An integrated, generic approach to pattern mining: data mining template library

Data Mining and Knowledge Discovery
Learning Classifiers from Large Databases Using Statistical Queries

WI-IAT '08 Proceedings of the 2008 IEEE/WIC/ACM International Conference on Web Intelligence and Intelligent Agent Technology - Volume 01
I/O scalable Bregman co-clustering

PAKDD'08 Proceedings of the 12th Pacific-Asia conference on Advances in knowledge discovery and data mining
Incrementally maintaining classification using an RDBMS

Proceedings of the VLDB Endowment
Relational approach for shortest path discovery over large graphs

Proceedings of the VLDB Endowment

Quantified Score

Hi-index	0.00

Visualization

Abstract

Software packages providing a whole set of data mining and machine learning algorithms are attractive because they allow experimentation with many kinds of algorithms in an easy setup. However, these packages are often based on main-memory data structures, limiting the amount of data they can handle. In this paper we use a relational database as secondary storage in order to eliminate this limitation. Unlike existing approaches, which often focus on optimizing a single algorithm to work with a database backend, we propose a general approach, which provides a database interface for several algorithms at once. We have taken a popular machine learning software package, Weka, and added a relational storage manager as back-tier to the system. The extension is transparent to the algorithms implemented in Weka, since it is hidden behind Weka's standard main-memory data structure interface. Furthermore, some general mining tasks are transfered into the database system to speed up execution. We tested the extended system, refered to as WekaDB, and our results show that it achieves a much higher scalability than Weka, while providing the same output and maintaining good computation time.