A data clustering algorithm for stratified data partitioning in artificial neural network

Authors:
Ajit K. Sahoo;Ming J. Zuo;M. K. Tiwari
Affiliations:
Department of Mechanical Engineering, University of Alberta, Edmonton, Canada;Department of Mechanical Engineering, University of Alberta, Edmonton, Canada;Department of Industrial Engineering and Management, Indian Institute of Technology, Kharagpur, India
Venue:
Expert Systems with Applications: An International Journal
Year:
2012

Citing 11
Cited 0

Silhouettes: a graphical aid to the interpretation and validation of cluster analysis

Journal of Computational and Applied Mathematics
Forecasting with neural networks

Information and Management
Guide to Neural Computing Applications

Guide to Neural Computing Applications
A Comparison of Data Preprocessing Strategies for Neural Network Modeling of Oil Production Prediction

ICCI '04 Proceedings of the Third IEEE International Conference on Cognitive Informatics
Wavelet-based multiresolution analysis for data cleaning and its application to water quality management systems

Expert Systems with Applications: An International Journal
Foreign-Exchange-Rate Forecasting with Artificial Neural Networks

Foreign-Exchange-Rate Forecasting with Artificial Neural Networks
A pattern-based outlier detection method identifying abnormal attributes in software project data

Information and Software Technology
Optimal feature selection for support vector machines

Pattern Recognition
A feature extraction method for use with bimodal biometrics

Pattern Recognition
A novel efficient technique for extracting valid feature information

Expert Systems with Applications: An International Journal
Bias and variance of validation methods for function approximationneural networks under conditions of sparse data

IEEE Transactions on Systems, Man, and Cybernetics, Part C: Applications and Reviews

Quantified Score

Hi-index	12.05

Visualization

Abstract

The statistical properties of training, validation and test data play an important role in assuring optimal performance in artificial neural networks (ANNs). Researchers have proposed optimized data partitioning (ODP) and stratified data partitioning (SDP) methods to partition of input data into training, validation and test datasets. ODP methods based on genetic algorithm (GA) are computationally expensive as the random search space can be in the power of twenty or more for an average sized dataset. For SDP methods, clustering algorithms such as self organizing map (SOM) and fuzzy clustering (FC) are used to form strata. It is assumed that data points in any individual stratum are in close statistical agreement. Reported clustering algorithms are designed to form natural clusters. In the case of large multivariate datasets, some of these natural clusters can be big enough such that the furthest data vectors are statistically far away from the mean. Further, these algorithms are computationally expensive as well. We propose a custom design clustering algorithm (CDCA) to overcome these shortcomings. Comparisons are made using three benchmark case studies, one each from classification, function approximation and prediction domains. The proposed CDCA data partitioning method is evaluated in comparison with SOM, FC and GA based data partitioning methods. It is found that the CDCA data partitioning method not only perform well but also reduces the average CPU time.