Collateral missing value estimation: robust missing value estimation for consequent microarray data processing

Authors:
Muhammad Shoaib B. Sehgal;Iqbal Gondal;Laurence Dooley
Affiliations:
Faculty of IT, Monash University, Churchill, VIC, Australia;Faculty of IT, Monash University, Churchill, VIC, Australia;Faculty of IT, Monash University, Churchill, VIC, Australia
Venue:
AI'05 Proceedings of the 18th Australian Joint conference on Advances in Artificial Intelligence
Year:
2005

Citing 6
Cited 1

Support Vector Machine and Generalized Regression Neural Network Based Classification Fusion Models for Cancer Diagnosis

HIS '04 Proceedings of the Fourth International Conference on Hybrid Intelligent Systems
K-Ranked Covariance Based Missing Values Estimation for Microarray Data Classification

HIS '04 Proceedings of the Fourth International Conference on Hybrid Intelligent Systems
A mixture model-based strategy for selecting sets of genes in multiclass response microarray experiments

Bioinformatics
Classification using partial least squares with penalized logistic regression

Bioinformatics
Bayesian model averaging: development of an improved multi-class, gene selection and classification tool for microarray data

Bioinformatics
Collateral missing value imputation: a new robust missing value estimation algorithm for microarray data

Bioinformatics

Missing value imputation framework for microarray significant gene selection and class prediction

BioDM'06 Proceedings of the 2006 international conference on Data Mining for Biomedical Applications

Quantified Score

Hi-index	0.00

Visualization

Abstract

Microarrays have unique ability to probe thousands of genes at a time that makes it a useful tool for variety of applications, ranging from diagnosis to drug discovery. However, data generated by microarrays often contains multiple missing gene expressions that affect the subsequent analysis, as most of the times these missing values are ignored. In this paper we have analyzed how accurate estimation of missing values can lead to better subsequent gene selection and class prediction. Collateral Missing Values Estimation (CMVE), which demonstrates superior imputation performance compared to Bayesian Principal Component Analysis (BPCA) Impute, K-Nearest Neighbour (KNN) algorithm, when estimating missing values in the BRCA1, BRCA2 and Sporadic genetic mutation samples present in ovarian cancer by exploiting both local/global and positive/negative correlation values. CMVE also consistently outperforms, in terms of classification accuracies, BPCA, KNN and ZeroImpute techniques. The imputation is followed by gene selection using fusion of Between Group to within Group Sum ofSquares and Weighted Partial Least Squares where Ridge Partial Least Square algorithm is used as a class predictor.