A technique for improving the performance of naive bayes text classification

Authors:
Yuqian Jiang;Huaizhong Lin;Xuesong Wang;Dongming Lu
Affiliations:
College of Computer Science and Technology, Zhejiang University, Hangzhou, Zhejiang, China;College of Computer Science and Technology, Zhejiang University, Hangzhou, Zhejiang, China;College of Computer Science and Technology, Zhejiang University, Hangzhou, Zhejiang, China;College of Computer Science and Technology, Zhejiang University, Hangzhou, Zhejiang, China
Venue:
WISM'11 Proceedings of the 2011 international conference on Web information systems and mining - Volume Part II
Year:
2011

Citing 7
Cited 0

On the Optimality of the Simple Bayesian Classifier under Zero-One Loss

Machine Learning - Special issue on learning with probabilistic representations
Distribution of content words and phrases in text and language modelling

Natural Language Engineering
Some Effective Techniques for Naive Bayes Text Classification

IEEE Transactions on Knowledge and Data Engineering
Exploiting temporal contexts in text classification

Proceedings of the 17th ACM conference on Information and knowledge management
An improved hierarchical Bayesian model of language for document classification

COLING '08 Proceedings of the 22nd International Conference on Computational Linguistics - Volume 1
An Effective Algorithm for Improving the Performance of Naive Bayes for Text Classification

ICCRD '10 Proceedings of the 2010 Second International Conference on Computer Research and Development
Techniques for improving the performance of naive bayes for text classification

CICLing'05 Proceedings of the 6th international conference on Computational Linguistics and Intelligent Text Processing

Quantified Score

Hi-index	0.00

Visualization

Abstract

Naive Bayes classifier is widely used in text classification tasks, and it can perform surprisingly well, it is often regarded as a baseline. But previous researches show that the skewed distribution of training collection may cause poor results in text classification. This paper presents a new method to deal with this situation. We introduce a conditional probability which takes into account both the information of the whole corpus and each category. Our proposed method performs well in the standard benchmark collections, competing with the state-of-the-art text classifiers especially for the skewed data.