BIRCH: an efficient data clustering method for very large databases
SIGMOD '96 Proceedings of the 1996 ACM SIGMOD international conference on Management of data
ROCK: A Robust Clustering Algorithm for Categorical Attributes
ICDE '99 Proceedings of the 15th International Conference on Data Engineering
Hi-index | 0.00 |
ROCK is a robust, categorical attribute oriented clustering algorithm. The main contribution of ROCK is the introduction of a novel concept called links as a measure of similarity between a pair of data points. Compared with traditional distance-based approaches, links capture global information over the whole data set rather than local information between two data points. Despite its success in clustering some categorical databases, there are still some underlying weaknesses. This paper investigates the problems deeply and proposes a novel algorithm QNNS using Qualified Nearest Neighbors Selection model, which improves clustering quality with an appropriate selection of nearest neighbors. We also discuss a cohesion measure to control the clustering process. Our methods reduce the dependence of the clustering quality on the pre-specified parameters and enhance the convenience for end users. Experiment results demonstrate that QNNS outperforms ROCK and VBACC.