Statistical analysis with missing data
Statistical analysis with missing data
"Missing Is Useful': Missing Values in Cost-Sensitive Decision Trees
IEEE Transactions on Knowledge and Data Engineering
Hi-index | 0.00 |
Organizations aim at harnessing predictive insights, using the vast real-time data stores that they have accumulated through the years, using data mining techniques. Health sector, has an extremely large source of digital data - patient-health related data-store, which can be effectively used for predictive analytics. This data, may consists of missing, incorrect and sometimes incomplete values sets that can have a detrimental effect on the decisions that are outcomes of data analytics. Using the PIMA Indians Diabetes dataset, we have proposed an efficient imputation method using a hybrid combination of CART and Genetic Algorithm, as a preprocessing step. The classical neural network model is used for prediction, on the preprocessed dataset. The accuracy achieved by the proposed model far exceeds the existing models, mainly because of the soft computing preprocessing adopted. This approach is simple, easy to understand and implement and practical in its approach.