Annotation and disambiguation of semantic types in biomedical text: a cascaded approach to named entity recognition

  • Authors:
  • Dietrich Rebholz-Schuhmann;Harald Kirsch;Sylvain Gaudan;Miguel Arregui;Goran Nenadic

  • Affiliations:
  • Wellcome Trust Genome Campus, Hinxton, Cambridge, UK;Wellcome Trust Genome Campus, Hinxton, Cambridge, UK;Wellcome Trust Genome Campus, Hinxton, Cambridge, UK;Wellcome Trust Genome Campus, Hinxton, Cambridge, UK;University of Manchester, Manchester, UK

  • Venue:
  • NLPXML '06 Proceedings of the 5th Workshop on NLP and XML: Multi-Dimensional Markup in Natural Language Processing
  • Year:
  • 2006

Quantified Score

Hi-index 0.00

Visualization

Abstract

Publishers of biomedical journals increasingly use XML as the underlying document format. We present a modular text-processing pipeline that inserts XML markup into such documents in every processing step, leading to multi-dimensional markup. The markup introduced is used to identify and disambiguate named entities of several semantic types (protein/gene, Gene Ontology terms, drugs and species) and to communicate data from one module to the next. Each module independently adds, changes or removes markup, which allows for modularization and a flexible setup of the processing pipeline. We also describe how the cascaded approach is embedded in a large-scale XML-based application (EBIMed) used for on-line access to biomedical literature. We discuss the lessons learnt so far, as well as the open problems that need to be resolved. In particular, we argue that the pragmatic and tailored solutions allow for reduction in the need for overlapping annotations --- although not completely without cost.