Redundant feature elimination for multi-class problems

  • Authors:
  • Annalisa Appice;Michelangelo Ceci;Simon Rawles;Peter Flach

  • Affiliations:
  • Università degli Studi di Bari, Bari, Italy;Università degli Studi di Bari, Bari, Italy;University of Bristol, Bristol, United Kingdom;University of Bristol, Bristol, United Kingdom

  • Venue:
  • ICML '04 Proceedings of the twenty-first international conference on Machine learning
  • Year:
  • 2004

Quantified Score

Hi-index 0.00

Visualization

Abstract

We consider the problem of eliminating redundant Boolean features for a given data set, where a feature is redundant if it separates the classes less well than another feature or set of features. Lavrač et al. proposed the algorithm REDUCE that works by pairwise comparison of features, i.e., it eliminates a feature if it is redundant with respect to another feature. Their algorithm operates in an ILP setting and is restricted to two-class problems. In this paper we improve their method and extend it to multiple classes. Central to our approach is the notion of a neighbourhood of examples: a set of examples of the same class where the number of different features between examples is relatively small. Redundant features are eliminated by applying a revised version of the REDUCE method to each pair of neighbourhoods of different class. We analyse the performance of our method on a range of data sets.