Seeing stars when there aren't many stars: graph-based semi-supervised learning for sentiment categorization

  • Authors:
  • Andrew B. Goldberg;Xiaojin Zhu

  • Affiliations:
  • University of Wisconsin-Madison, Madison, W.I.;University of Wisconsin-Madison, Madison, W.I.

  • Venue:
  • TextGraphs-1 Proceedings of the First Workshop on Graph Based Methods for Natural Language Processing
  • Year:
  • 2006

Quantified Score

Hi-index 0.01

Visualization

Abstract

We present a graph-based semi-supervised learning algorithm to address the sentiment analysis task of rating inference. Given a set of documents (e.g., movie reviews) and accompanying ratings (e.g., "4 stars"), the task calls for inferring numerical ratings for unlabeled documents based on the perceived sentiment expressed by their text. In particular, we are interested in the situation where labeled data is scarce. We place this task in the semi-supervised setting and demonstrate that considering unlabeled reviews in the learning process can improve rating-inference performance. We do so by creating a graph on both labeled and unlabeled data to encode certain assumptions for this task. We then solve an optimization problem to obtain a smooth rating function over the whole graph. When only limited labeled data is available, this method achieves significantly better predictive accuracy over other methods that ignore the unlabeled examples during training.