VChunkJoin: An Efficient Algorithm for Edit Similarity Joins

Authors:
Wei Wang;Jianbin Qin;Xiao Chuan;Xuemin Lin;Heng Tao Shen
Affiliations:
The University of New South Wales, Sydney;The University of New South Wales, Sydney;The University of New South Wales, Sydney;The University of New South Wales, Sydney;The University of Queensland, Brisbane
Venue:
IEEE Transactions on Knowledge and Data Engineering
Year:
2013

Citing 0
Cited 1

Efficient processing of graph similarity queries with edit distance constraints

The VLDB Journal — The International Journal on Very Large Data Bases

Quantified Score

Hi-index	0.00

Visualization

Abstract

Similarity joins play an important role in many application areas, such as data integration and cleaning, record linkage, and pattern recognition. In this paper, we study efficient algorithms for similarity joins with an edit distance constraint. Currently, the most prevalent approach is based on extracting overlapping grams from strings and considering only strings that share a certain number of grams as candidates. Unlike these existing approaches, we propose a novel approach to edit similarity join based on extracting nonoverlapping substrings, or chunks, from strings. We propose a class of chunking schemes based on the notion of tail-restricted chunk boundary dictionary. A new algorithm, VChunkJoin, is designed by integrating existing filtering methods and several new filters unique to our chunk-based method. We also design a greedy algorithm to automatically select a good chunking scheme for a given data set. We demonstrate experimentally that the new algorithm is faster than alternative methods yet occupies less space.