Data discretization unification

Authors:
Ruoming Jin;Yuri Breitbart;Chibuike Muoh
Affiliations:
Kent State University, Department of Computer Science, 44241, Kent, OH, USA;Kent State University, Department of Computer Science, 44241, Kent, OH, USA;Kent State University, Department of Computer Science, 44241, Kent, OH, USA
Venue:
Knowledge and Information Systems
Year:
2009

Citing 0
Cited 4

Focusing on novelty: a crawling strategy to build diverse language models

Proceedings of the 20th ACM international conference on Information and knowledge management
A new discretization algorithm based on range coefficient of dispersion and skewness for neural networks classifier

Applied Soft Computing
DHCC: Divisive hierarchical clustering of categorical data

Data Mining and Knowledge Discovery
Compact classification of optimized Boolean reasoning with Particle Swarm Optimization

Intelligent Data Analysis

Quantified Score

Hi-index	0.00

Visualization

Abstract

Data discretization is defined as a process of converting continuous data attribute values into a finite set of intervals with minimal loss of information. In this paper, we prove that discretization methods based on informational theoretical complexity and the methods based on statistical measures of data dependency are asymptotically equivalent. Furthermore, we define a notion of generalized entropy and prove that discretization methods based on Minimal description length principle, Gini index, AIC, BIC, and Pearson’s X 2 and G 2 statistics are all derivable from the generalized entropy function. We design a dynamic programming algorithm that guarantees the best discretization based on the generalized entropy notion. Furthermore, we conducted an extensive performance evaluation of our method for several publicly available data sets. Our results show that our method delivers on the average 31% less classification errors than many previously known discretization methods.