Metric and reference factors in minimum error rate training

  • Authors:
  • Yifan He;Andy Way

  • Affiliations:
  • CNGL, School of Computing, Dublin City University, Dublin, Ireland;CNGL, School of Computing, Dublin City University, Dublin, Ireland

  • Venue:
  • Machine Translation
  • Year:
  • 2010

Quantified Score

Hi-index 0.00

Visualization

Abstract

In Minimum Error Rate Training (MERT), Bleu is often used as the error function, despite the fact that it has been shown to have a lower correlation with human judgment than other metrics such as Meteor and Ter. In this paper, we present empirical results in which parameters tuned on Bleu may lead to sub-optimal Bleu scores under certain data conditions. Such scores can be improved significantly by tuning on an entirely different metric altogether, e.g. Meteor, by 0.0082 Bleu or 3.38% relative improvement on the WMT08 English---French data. We analyze the influence of the number of references and choice of metrics on the result of MERT and experiment on different data sets. We show the problems of tuning on a metric that is not designed for the single reference scenario and point out some possible solutions.