Detecting near-duplicates in large-scale short text databases

Authors:
Caichun Gong;Yulan Huang;Xueqi Cheng;Shuo Bai
Affiliations:
Institute of Computing Technology, Chinese Academy of Sciences, Beijing, P.R.C. and Graduate School of Chinese Academy of Sciences, Beijing, P.R.C.;Institute of Computing Technology, Chinese Academy of Sciences, Beijing, P.R.C. and Graduate School of Chinese Academy of Sciences, Beijing, P.R.C.;Institute of Computing Technology, Chinese Academy of Sciences, Beijing, P.R.C.;Institute of Computing Technology, Chinese Academy of Sciences, Beijing, P.R.C.
Venue:
PAKDD'08 Proceedings of the 12th Pacific-Asia conference on Advances in knowledge discovery and data mining
Year:
2008

Citing 7
Cited 3

Term-weighting approaches in automatic text retrieval

Information Processing and Management: an International Journal
The merge/purge problem for large databases

SIGMOD '95 Proceedings of the 1995 ACM SIGMOD international conference on Management of data
Copy detection mechanisms for digital documents

SIGMOD '95 Proceedings of the 1995 ACM SIGMOD international conference on Management of data
Similarity estimation techniques from rounding algorithms

STOC '02 Proceedings of the thiry-fourth annual ACM symposium on Theory of computing
Finding Near-Replicas of Documents and Servers on the Web

WebDB '98 Selected papers from the International Workshop on The World Wide Web and Databases
Finding near-duplicate web pages: a large-scale evaluation of algorithms

SIGIR '06 Proceedings of the 29th annual international ACM SIGIR conference on Research and development in information retrieval
Detecting near-duplicates for web crawling

Proceedings of the 16th international conference on World Wide Web

De-duplication of aggregation authority files

International Journal of Metadata, Semantics and Ontologies
Detecting near-duplicate documents using sentence-level features and supervised learning

Expert Systems with Applications: An International Journal
De-duplication of aggregation authority files

International Journal of Metadata, Semantics and Ontologies

Quantified Score

Hi-index	0.00

Visualization

Abstract

Near-duplicates are abundant in short text databases. Detecting and eliminating them is of great importance. SimFinder proposed in this paper is a fast algorithm to identify all near-duplicates in large-scale short text databases. An ad hoc term weighting scheme is employed to measure each term's discriminative ability. A certain number of terms with higher weights are seletect as features for each short text. SimFinder generates several fingerprints for each text, and only texts with at least one fingerprint in common are compared with each other. An optimization procedure is employed in SimFinder to make it more efficient. Experiments indicate that SimFinder is an effective solution for short text duplicate detection with almost linear time and storage complexity. Both precision and recall of SimFinder are promising.