Threshold selection for web-page classification with highly skewed class distribution

Authors:
Xiaofeng He;Lei Duan;Yiping Zhou;Byron Dom
Affiliations:
Yahoo! Inc., Santa Clara, CA, USA;Yahoo! Inc., Santa Clara, CA, USA;Yahoo! Inc., Santa Clara, CA, USA;Yahoo! Inc., Santa Clara, CA, USA
Venue:
Proceedings of the 18th international conference on World wide web
Year:
2009

Citing 1
Cited 3

A study of thresholding strategies for text categorization

Proceedings of the 24th annual international ACM SIGIR conference on Research and development in information retrieval

Online stratified sampling: evaluating classifiers at web-scale

CIKM '10 Proceedings of the 19th ACM international conference on Information and knowledge management
A feature-free search query classification approach using semantic distance

Expert Systems with Applications: An International Journal
What's the deal?: identifying online bargains

AWC '13 Proceedings of the First Australasian Web Conference - Volume 144

Quantified Score

Hi-index	0.00

Visualization

Abstract

We propose a novel cost-efficient approach to threshold selection for binary web-page classification problems with imbalanced class distributions. In many binary-classification tasks the distribution of classes is highly skewed. In such problems, using uniform random sampling in constructing sample sets for threshold setting requires large sample sizes in order to include a statistically sufficient number of examples of the minority class. On the other hand, manually labeling examples is expensive and budgetary considerations require that the size of sample sets be limited. These conflicting requirements make threshold selection a challenging problem. Our method of sample-set construction is a novel approach based on stratified sampling, in which manually labeled examples are expanded to reflect the true class distribution of the web-page population. Our experimental results show that using false positive rate as the criterion for threshold setting results in lower-variance threshold estimates than using other widely used accuracy measures such as F1 and precision.