Cross-validated bagged learning

  • Authors:
  • Maya L. Petersen;Annette M. Molinaro;Sandra E. Sinisi;Mark J. van der Laan

  • Affiliations:
  • Division of Biostatistics, University of California, Berkeley School of Public Health, Earl Warren Hall 7360, Berkeley, CA 94720-7360, USA;School of Public Health, Yale University, 60 College Street, New Haven, CT 06520-8034, USA;Division of Biostatistics, University of California, Berkeley School of Public Health, Earl Warren Hall 7360, Berkeley, CA 94720-7360, USA;Division of Biostatistics, University of California, Berkeley School of Public Health, Earl Warren Hall 7360, Berkeley, CA 94720-7360, USA

  • Venue:
  • Journal of Multivariate Analysis
  • Year:
  • 2007

Quantified Score

Hi-index 0.01

Visualization

Abstract

Many applications aim to learn a high dimensional parameter of a data generating distribution based on a sample of independent and identically distributed observations. For example, the goal might be to estimate the conditional mean of an outcome given a list of input variables. In this prediction context, bootstrap aggregating (bagging) has been introduced as a method to reduce the variance of a given estimator at little cost to bias. Bagging involves applying an estimator to multiple bootstrap samples and averaging the result across bootstrap samples. In order to address the curse of dimensionality, a common practice has been to apply bagging to estimators which themselves use cross-validation, thereby using cross-validation within a bootstrap sample to select fine-tuning parameters trading off bias and variance of the bootstrap sample-specific candidate estimators. In this article we point out that in order to achieve the correct bias variance trade-off for the parameter of interest, one should apply the cross-validation selector externally to candidate bagged estimators indexed by these fine-tuning parameters. We use three simulations to compare the new cross-validated bagging method with bagging of cross-validated estimators and bagging of non-cross-validated estimators.