'1 + 1 2': Merging Distance and Density Based Clustering

Authors:
Manoranjan Dash;Huan Liu;X. Xu
Affiliations:
-;-;-
Venue:
DASFAA '01 Proceedings of the 7th International Conference on Database Systems for Advanced Applications
Year:
2001

Citing 0
Cited 9

Subspace clustering for high dimensional data: a review

ACM SIGKDD Explorations Newsletter - Special issue on learning from imbalanced datasets
A design to promote group learning in e-learning: Experiences from the field

Computers & Education
Varying Density Spatial Clustering Based on a Hierarchical Tree

MLDM '07 Proceedings of the 5th international conference on Machine Learning and Data Mining in Pattern Recognition
Image-mapped data clustering: An efficient technique for clustering large data sets

Intelligent Data Analysis
Arif Index for Predicting the Classification Accuracy of Features and Its Application in Heart Beat Classification Problem

PAKDD '09 Proceedings of the 13th Pacific-Asia Conference on Advances in Knowledge Discovery and Data Mining
A hybrid incremental clustering method-combining support vector machine and enhanced clustering by committee clustering algorithm

PAKDD'07 Proceedings of the 11th Pacific-Asia conference on Advances in knowledge discovery and data mining
APSCAN: A parameter free algorithm for clustering

Pattern Recognition Letters
A comparative study of efficient initialization methods for the k-means clustering algorithm

Expert Systems with Applications: An International Journal
WCOID-DG: An approach for case base maintenance based on Weighting, Clustering, Outliers, Internal Detection and Dbsan-Gmeans

Journal of Computer and System Sciences

Quantified Score

Hi-index	0.00

Visualization

Abstract

Abstract: Clustering is an important data exploration task. Its use in data mining is growing very fast. Traditional clustering algorithms which no longer cater to the data mining requirements are modified increasingly. Clustering algorithms are numerous which can be divided in several categories. Two prominent categories are distance-based and density-based (e.g. K-means and DBSCAN, respectively). While K-means is fast, easy to implement, and converges to local optima almost surely, but it is also easily affected by noise. On the other hand, while density-based clustering can find arbitrary shape clusters and handle noise well, but it is also slow in comparison due to neighborhood search for each data point, and faces difficulty in setting density threshold properly. In this paper, we propose BRIDGE that efficiently merges the two by exploiting the advantages of one to counter the limitations of the other and vice versa. BRIDGE enables DBSCAN to handle very large data efficiently and improves the quality of K-means clusters by removing the noisy points. It also helps the user in setting the density threshold parameter properly. We further show that other clustering algorithms can be merged using similar strategy. An example given in the paper merges BIRCH clustering with DBSCAN.