Intrinsic plagiarism analysis

  • Authors:
  • Benno Stein;Nedim Lipka;Peter Prettenhofer

  • Affiliations:
  • Faculty of Media, Media Systems, Bauhaus-Universität Weimar, Weimar, Germany 99421;Faculty of Media, Media Systems, Bauhaus-Universität Weimar, Weimar, Germany 99421;Faculty of Media, Media Systems, Bauhaus-Universität Weimar, Weimar, Germany 99421

  • Venue:
  • Language Resources and Evaluation
  • Year:
  • 2011

Quantified Score

Hi-index 0.00

Visualization

Abstract

Research in automatic text plagiarism detection focuses on algorithms that compare suspicious documents against a collection of reference documents. Recent approaches perform well in identifying copied or modified foreign sections, but they assume a closed world where a reference collection is given. This article investigates the question whether plagiarism can be detected by a computer program if no reference can be provided, e.g., if the foreign sections stem from a book that is not available in digital form. We call this problem class intrinsic plagiarism analysis; it is closely related to the problem of authorship verification. Our contributions are threefold. (1) We organize the algorithmic building blocks for intrinsic plagiarism analysis and authorship verification and survey the state of the art. (2) We show how the meta learning approach of Koppel and Schler, termed "unmasking", can be employed to post-process unreliable stylometric analysis results. (3) We operationalize and evaluate an analysis chain that combines document chunking, style model computation, one-class classification, and meta learning.