Efficient k-nearest neighbor searching in nonordered discrete data spaces

Authors:
Dashiell Kolbe;Qiang Zhu;Sakti Pramanik
Affiliations:
Michigan State University;University of Michigan—Dearborn;Michigan State University
Venue:
ACM Transactions on Information Systems (TOIS)
Year:
2010

Citing 17
Cited 2

The R*-tree: an efficient and robust access method for points and rectangles

SIGMOD '90 Proceedings of the 1990 ACM SIGMOD international conference on Management of data
Searching for geometric molecular shape complementarity using bidimensional surface profiles

Journal of Molecular Graphics
Nearest neighbor queries

SIGMOD '95 Proceedings of the 1995 ACM SIGMOD international conference on Management of data
Optimal multi-step k-nearest neighbor search

SIGMOD '98 Proceedings of the 1998 ACM SIGMOD international conference on Management of data
Searching in metric spaces

ACM Computing Surveys (CSUR)
The K-D-B-tree: a search structure for large multidimensional dynamic indexes

SIGMOD '81 Proceedings of the 1981 ACM SIGMOD international conference on Management of data
Generalized Hamming Distance

Information Retrieval
The LSDh-Tree: An Access Structure for Feature Vectors

ICDE '98 Proceedings of the Fourteenth International Conference on Data Engineering
M-tree: An Efficient Access Method for Similarity Search in Metric Spaces

VLDB '97 Proceedings of the 23rd International Conference on Very Large Data Bases
Fast Nearest Neighbor Search in Medical Image Databases

VLDB '96 Proceedings of the 22th International Conference on Very Large Data Bases
The X-tree: An Index Structure for High-Dimensional Data

VLDB '96 Proceedings of the 22th International Conference on Very Large Data Bases
Efficient User-Adaptable Similarity Search in Large Multimedia Databases

VLDB '97 Proceedings of the 23rd International Conference on Very Large Data Bases
Principles and applications for supporting similarity queries in non-ordered-discrete and continuous data spaces

Principles and applications for supporting similarity queries in non-ordered-discrete and continuous data spaces
A space-partitioning-based indexing method for multidimensional non-ordered discrete data spaces

ACM Transactions on Information Systems (TOIS)
Dynamic indexing for multidimensional non-ordered discrete data spaces using a data-partitioning approach

ACM Transactions on Database Systems (TODS)
The ND-tree: a dynamic indexing technique for multidimensional non-ordered discrete data spaces

VLDB '03 Proceedings of the 29th international conference on Very large data bases - Volume 29
Voronoi-based K nearest neighbor search for spatial network databases

VLDB '04 Proceedings of the Thirtieth international conference on Very large data bases - Volume 30

Reducing non-determinism of k-NN searching in non-ordered discrete data spaces

Information Processing Letters
Defending recommender systems by influence analysis

Information Retrieval

Quantified Score

Hi-index	0.00

Visualization

Abstract

Numerous techniques have been proposed in the past for supporting efficient k-nearest neighbor (k-NN) queries in continuous data spaces. Limited work has been reported in the literature for k-NN queries in a nonordered discrete data space (NDDS). Performing k-NN queries in an NDDS raises new challenges. The Hamming distance is usually used to measure the distance between two vectors (objects) in an NDDS. Due to the coarse granularity of the Hamming distance, a k-NN query in an NDDS may lead to a high degree of nondeterminism for the query result. We propose a new distance measure, called Granularity-Enhanced Hamming (GEH) distance, which effectively reduces the number of candidate solutions for a query. We have also implemented k-NN queries using multidimensional database indexing in NDDSs. Further, we use the properties of our multidimensional NDDS index to derive the probability of encountering valid neighbors within specific regions of the index. This probability is used to develop a new search ordering heuristic. Our experiments on synthetic and genomic data sets demonstrate that our index-based k-NN algorithm is efficient in finding k-NNs in both uniform and nonuniform data sets in NDDSs and that our heuristics are effective in improving the performance of such queries.