The design of a similarity based deduplication system

  • Authors:
  • Lior Aronovich;Ron Asher;Eitan Bachmat;Haim Bitner;Michael Hirsch;Shmuel T. Klein

  • Affiliations:
  • IBM Corp.;IBM Corp.;Ben-Gurion U.;Marvell Corp.;IBM Corp.;Bar-Ilan U.

  • Venue:
  • SYSTOR '09 Proceedings of SYSTOR 2009: The Israeli Experimental Systems Conference
  • Year:
  • 2009

Quantified Score

Hi-index 0.00

Visualization

Abstract

We describe some of the design choices that were made during the development of a fast, scalable, inline, deduplication device. The system's design goals and how they were achieved are presented. This is the firs deduplication device that uses similarity matching. The paper provides the following original research contributions: we show how similarity signatures can serve in a deduplication scheme; a novel type of similarity signatures is presented and its advantages in the context of deduplication requirements are explained. It is also shown how to combine similarity matching schemes with byte by byte comparison or hash based identity schemes.