Decision Trees for Probability Estimation: An Empirical Study

Authors:
Han Liang;Harry Zhang;Yuhong Yan
Affiliations:
University of New Brunswick, Canada;University of New Brunswick, Canada;National Research Council of Canada, Canada
Venue:
ICTAI '06 Proceedings of the 18th IEEE International Conference on Tools with Artificial Intelligence
Year:
2006

Citing 0
Cited 2

A Combined Classification Algorithm Based on C4.5 and NB

ISICA '08 Proceedings of the 3rd International Symposium on Advances in Computation and Intelligence
A comparative analysis of methods for probability estimation tree

WSEAS Transactions on Computers

Quantified Score

Hi-index	0.00

Visualization

Abstract

Accurate probability estimation generated by learning models is desirable in some practical applications, such as medical diagnosis. In this paper, we empirically study traditional decision-tree learning models and their variants in terms of probability estimation, measured by Conditional Log Likelihood (CLL). Furthermore, we also compare decision tree learning with other kinds of representative learning: na篓ýve Bayes, Na篓ýve Bayes Tree, Bayesian Network, K-Nearest Neighbors and Support Vector Machine with respect to probability estimation. From our experiments, we have several interesting observations. First, among various decision-tree learning models, C4.4 is the best in yielding precise probability estimation measured by CLL, although its performance is not good in terms of other evaluation criteria, such as accuracy and ranking. We provide an explanation for this and reveal the nature of CLL. Second, compared with other popular models, C4.4 achieves the best CLL. Finally, CLL does not dominate another wellestablished relevant measurement AUC (the Area Under the Curve of Receiver Operating Characteristics), which suggests that different decision-tree learning models should be used for different objectives. Our experiments are conducted on the basis of 36 UCI sample sets that cover a wide range of domains and data characteristics. We run all the models within a machine learning platform - Weka.