Skip to main navigation Skip to search Skip to main content

Efficient similarity joins for near-duplicate detection

  • Chuan Xiao*
  • , Wei Wang
  • , Xuemin Lin
  • , Jeffrey Xu Yu
  • , Guoren Wang
  • *Corresponding author for this work
  • University of New South Wales
  • Chinese University of Hong Kong
  • Northeastern University China

Research output: Contribution to journalArticlepeer-review

Abstract

With the increasing amount of data and the need to integrate data from multiple data sources, one of the challenging issues is to identify near-duplicate records efficiently. In this article, we focus on efficient algorithms to find a pair of records such that their similarities are no less than a given threshold. Several existing algorithms rely on the prefix filtering principle to avoid computing similarity values for all possible pairs of records. We propose new filtering techniques by exploiting the token ordering information; they are integrated into the existing methods and drastically reduce the candidate sizes and hence improve the efficiency. We have also studied the implementation of our proposed algorithm in stand-alone and RDBMSbased settings. Experimental results show our proposed algorithms can outperform previous algorithms on several real datasets.

Original languageEnglish
Article number15
JournalACM Transactions on Database Systems
Volume36
Issue number3
DOIs
Publication statusPublished - Aug 2011
Externally publishedYes

Keywords

  • Near duplicate detection
  • Similarity join

Fingerprint

Dive into the research topics of 'Efficient similarity joins for near-duplicate detection'. Together they form a unique fingerprint.

Cite this