Drift mining in data: A framework for addressing drift in classification

Authors:
Vera Hofer;Georg Krempl
Affiliations:
Department of Statistics and Operations Research, University of Graz, Universitätsstraíe 15/E3, A-8010 Graz, Austria;Knowledge Management and Discovery Group, Otto-von-Guericke University, Magdeburg, POB 4120, D-39016 Magdeburg, Germany
Venue:
Computational Statistics & Data Analysis
Year:
2013

Citing 15
Cited 0

The impact of changing populations on classifier performance

KDD '99 Proceedings of the fifth ACM SIGKDD international conference on Knowledge discovery and data mining
Adjusting the outputs of a classifier to new a priori probabilities: a simple procedure

Neural Computation
An Optimal Reject Rule for Binary Classifiers

Proceedings of the Joint IAPR International Workshops on Advances in Pattern Recognition
On the Application of ROC Analysis to Predict Classification Performance Under Varying Class Distributions

Machine Learning
A Response to Webb and Ting's On the Application of ROC Analysis to Predict Classification Performance Under Varying Class Distributions

Machine Learning
Boosting for transfer learning

Proceedings of the 24th international conference on Machine learning
Quantifying counts and costs via classification

Data Mining and Knowledge Discovery
Dataset Shift in Machine Learning

Dataset Shift in Machine Learning
Measuring classifier performance: a coherent alternative to the area under the ROC curve

Machine Learning
A case-based technique for tracking concept drift in spam filtering

Knowledge-Based Systems
Transfer estimation of evolving class priors in data stream classification

Pattern Recognition
The impact of latency on online classification learning with concept drift

KSEM'10 Proceedings of the 4th international conference on Knowledge science, engineering and management
Agnostic domain adaptation

DAGM'11 Proceedings of the 33rd international conference on Pattern recognition
The algorithm APT to classify in concurrence of latency and drift

IDA'11 Proceedings of the 10th international conference on Advances in intelligent data analysis X
Classification in Presence of Drift and Latency

ICDMW '11 Proceedings of the 2011 IEEE 11th International Conference on Data Mining Workshops

Quantified Score

Hi-index	0.03

Visualization

Abstract

A novel statistical methodology for analysing population drift in classification is introduced. Drift denotes changes in the joint distribution of explanatory variables and class labels over time. It entails the deterioration of a classifier's performance and requires the optimal decision boundary to be adapted after some time. However, in the presence of verification latency a re-estimation of the classification model is impossible, since in such a situation only recent unlabelled data are available, and the true corresponding labels only become known after some lapse in time. For this reason a novel drift mining methodology is presented which aims at detecting changes over time. It allows us either to understand evolution in the data from an ex-post perspective or, ex-ante, to anticipate changes in the joint distribution. The proposed drift mining technique assumes that the class priors change by a certain factor from one time point to the next, and that the conditional distributions do not change within this time period. Thus, the conditional distributions can be estimated at a time where recent labelled data are available. In subsequent periods the unconditional distribution can be expressed as a mixture of the conditional distributions, where the mixing proportions are equal to the class priors. However, as the unconditional distributions can also be estimated from new unlabelled data, they can then be compared to the mixture representation by means of least-squares criteria. This allows for easy and fast estimation of the changes in class prior values in the presence of verification latency. The usefulness of this drift mining approach is demonstrated using a real-world dataset from the area of credit scoring.