Two credit scoring models based on dual strategy ensemble trees

  • Authors:
  • Gang Wang;Jian Ma;Lihua Huang;Kaiquan Xu

  • Affiliations:
  • School of Management, Hefei University of Technology, Hefei, Anhui 230009, PR China and Key Laboratory of Process Optimization and Intelligent Decision-making, Ministry of Education, Hefei, Anhui, ...;Department of Information Systems, City University of Hong Kong, Tat Chee Avenue, Kowloon, Hong Kong;School of Management, Fudan University, Shanghai 200433, PR China;Department of Information Systems, City University of Hong Kong, Tat Chee Avenue, Kowloon, Hong Kong and Department of Electronic Commerce, School of Business, Nanjing University, Nanjing, Jiangsh ...

  • Venue:
  • Knowledge-Based Systems
  • Year:
  • 2012

Quantified Score

Hi-index 0.00

Visualization

Abstract

Decision tree (DT) is one of the most popular classification algorithms in data mining and machine learning. However, the performance of DT based credit scoring model is often relatively poorer than other techniques. This is mainly due to two reasons: DT is easily affected by (1) the noise data and (2) the redundant attributes of data under the circumstance of credit scoring. In this study, we propose two dual strategy ensemble trees: RS-Bagging DT and Bagging-RS DT, which are based on two ensemble strategies: bagging and random subspace, to reduce the influences of the noise data and the redundant attributes of data and to get the relatively higher classification accuracy. Two real world credit datasets are selected to demonstrate the effectiveness and feasibility of proposed methods. Experimental results reveal that single DT gets the lowest average accuracy among five single classifiers, i.e., Logistic Regression Analysis (LRA), Linear Discriminant Analysis (LDA), Multi-layer Perceptron (MLP) and Radial Basis Function Network (RBFN). Moreover, RS-Bagging DT and Bagging-RS DT get the better results than five single classifiers and four popular ensemble classifiers, i.e., Bagging DT, Random Subspace DT, Random Forest and Rotation Forest. The results show that RS-Bagging DT and Bagging-RS DT can be used as alternative techniques for credit scoring.