Deduplication of metadata harvested from Open Archives Initiative repositories

  • Authors:
  • Piotr Wendykier

  • Affiliations:
  • Interdisciplinary Centre for Mathematical and Computational Modelling, University of Warsaw, ul. Prosta 69, 00-838 Warsaw, Poland. E-mail: p.wendykier@icm.edu.pl

  • Venue:
  • Information Services and Use - Mining the Digital Information Networks
  • Year:
  • 2013

Quantified Score

Hi-index 0.00

Visualization

Abstract

Open access OA is a way of providing unrestricted access via the Internet to peer-reviewed journal articles as well as theses, monographs and book chapters. Many open access repositories have been created in the last decade. There is also a number of registry websites that index these repositories. This article analyzes the repositories indexed by the Open Archives Initiative OAI organization in terms of record duplication. Based on the sample of 958 metadata files containing records modified in 2012 we provide an estimate on the number of duplicates in the entire collection of repositories indexed by OAI. In addition, this work describes several open source tools that form a generic workflow suitable for deduplication of bibliographic records.