Using patterns of thematic progression for building a table of contents of a text

  • Authors:
  • Marie-francine Moens

  • Affiliations:
  • Interdisciplinary centre for law and information technology, katholieke universiteit leuven, tienstraat 41, b-3000 leuven, belgium e-mail: marie-france.moens@law.kuleuven.be

  • Venue:
  • Natural Language Engineering
  • Year:
  • 2008

Quantified Score

Hi-index 0.01

Visualization

Abstract

A text usually contains one or a few main topics, which are split up into subtopics, which in their turn can be further described by more detailed topics. In this article we describe a system that segments a text into topics and subtopics. Each segment is characterized by important key terms that are extracted from it and by its begin and end position in the text. A table of contents is built by using the hierarchical and sequential relationships between topical segments that are identified in a text. The table of contents generator relies upon universal linguistic theories on the topic and comment of a sentence and on patterns of thematic progression in text. The linguistic theories of topic and comment are modeled both deterministically and probabilistically. The system is applied to English texts (news, World Wide Web and encyclopedia texts) and is evaluated.