Javascript must be enabled to continue!
Evaluation of normalization methods for cDNA microarray data by k-NN classification
View through CrossRef
Abstract
Background
Non-biological factors give rise to unwanted variations in cDNA microarray data. There are many normalization methods designed to remove such variations. However, to date there have been few published systematic evaluations of these techniques for removing variations arising from dye biases in the context of downstream, higher-order analytical tasks such as classification.
Results
Ten location normalization methods that adjust spatial- and/or intensity-dependent dye biases, and three scale methods that adjust scale differences were applied, individually and in combination, to five distinct, published, cancer biology-related cDNA microarray data sets. Leave-one-out cross-validation (LOOCV) classification error was employed as the quantitative end-point for assessing the effectiveness of a normalization method. In particular, a known classifier, k-nearest neighbor (k-NN), was estimated from data normalized using a given technique, and the LOOCV error rate of the ensuing model was computed. We found that k-NN classifiers are sensitive to dye biases in the data. Using N ONRM and GMEDIAN as baseline methods, our results show that single-bias-removal techniques which remove either spatial-dependent dye bias (referred later as spatial effect) or intensity-dependent dye bias (referred later as intensity effect) moderately reduce LOOCV classification errors; whereas double-bias-removal techniques which remove both spatial- and intensity effect reduce LOOCV classification errors even further. Of the 41 different strategies examined, three two-step processes, IG LOESS-SL FILTERW7, IST SPLINE-SL LOESS and IG LOESS-SL LOESS, all of which removed intensity effect globally and spatial effect locally, appear to reduce LOOCV classification errors most consistently and effectively across all data sets. We also found that the investigated scale normalization methods do not reduce LOOCV classification error.
Conclusion
Using LOOCV error of k-NNs as the evaluation criterion, three double-bias-removal normalization strategies, IG LOESS-SL FILTERW7, IST SPLINE-SL LOESS and IG LOESS-SL LOESS, outperform other strategies for removing spatial effect, intensity effect and scale differences from cDNA microarray data. The apparent sensitivity of k-NN LOOCV classification error to dye biases suggests that this criterion provides an informative measure for evaluating normalization methods. All the computational tools used in this study were implemented using the R language for statistical computing and graphics.
Springer Science and Business Media LLC
Title: Evaluation of normalization methods for cDNA microarray data by k-NN classification
Description:
Abstract
Background
Non-biological factors give rise to unwanted variations in cDNA microarray data.
There are many normalization methods designed to remove such variations.
However, to date there have been few published systematic evaluations of these techniques for removing variations arising from dye biases in the context of downstream, higher-order analytical tasks such as classification.
Results
Ten location normalization methods that adjust spatial- and/or intensity-dependent dye biases, and three scale methods that adjust scale differences were applied, individually and in combination, to five distinct, published, cancer biology-related cDNA microarray data sets.
Leave-one-out cross-validation (LOOCV) classification error was employed as the quantitative end-point for assessing the effectiveness of a normalization method.
In particular, a known classifier, k-nearest neighbor (k-NN), was estimated from data normalized using a given technique, and the LOOCV error rate of the ensuing model was computed.
We found that k-NN classifiers are sensitive to dye biases in the data.
Using N ONRM and GMEDIAN as baseline methods, our results show that single-bias-removal techniques which remove either spatial-dependent dye bias (referred later as spatial effect) or intensity-dependent dye bias (referred later as intensity effect) moderately reduce LOOCV classification errors; whereas double-bias-removal techniques which remove both spatial- and intensity effect reduce LOOCV classification errors even further.
Of the 41 different strategies examined, three two-step processes, IG LOESS-SL FILTERW7, IST SPLINE-SL LOESS and IG LOESS-SL LOESS, all of which removed intensity effect globally and spatial effect locally, appear to reduce LOOCV classification errors most consistently and effectively across all data sets.
We also found that the investigated scale normalization methods do not reduce LOOCV classification error.
Conclusion
Using LOOCV error of k-NNs as the evaluation criterion, three double-bias-removal normalization strategies, IG LOESS-SL FILTERW7, IST SPLINE-SL LOESS and IG LOESS-SL LOESS, outperform other strategies for removing spatial effect, intensity effect and scale differences from cDNA microarray data.
The apparent sensitivity of k-NN LOOCV classification error to dye biases suggests that this criterion provides an informative measure for evaluating normalization methods.
All the computational tools used in this study were implemented using the R language for statistical computing and graphics.
Related Results
Data Normalization Methods of Hybridized Multi-Stage Feature Selection Classification for 5G Base Station Antenna Health Effect Detection
Data Normalization Methods of Hybridized Multi-Stage Feature Selection Classification for 5G Base Station Antenna Health Effect Detection
It is essential to assess human exposure to Fifth Generation (5G) Radiofrequency Electromagnetic Field (RF-EMF) signal from Base Station (BS) sources operating at Low Band 5G at 70...
Non-Recommended Publishing Lists: Strategies for Detecting Deceitful Journals
Non-Recommended Publishing Lists: Strategies for Detecting Deceitful Journals
Abstract
The rapid growth of open access publishing (OAP) has significantly improved the accessibility and dissemination of scientific knowledge. However, this expansion has also c...
A Study on Gene Selection and Classification Algorithms for Classification of Microarray Gene Expression Data
A Study on Gene Selection and Classification Algorithms for Classification of Microarray Gene Expression Data
Pembangunan teknologi microarray membenarkan penyelidik untuk meneliti tahap ekspresi gen dalam sel. Salah satu aplikasi teknologi microarray adalah pengkelasan sampel tisu kepada ...
An analysis of the use of genomic DNA as a universal reference in two channel DNA microarrays
An analysis of the use of genomic DNA as a universal reference in two channel DNA microarrays
Abstract
Background
DNA microarray is an invaluable tool for gene expression explorations. In the two-dye microarray, fluorescence i...
Normalization of full-length enriched cDNA
Normalization of full-length enriched cDNA
Abstract
Analysis of rare messages in cDNA libraries is extremely difficult due to the substantial variations in the abundance of different transcripts in cells a...
ArrayWiki: an enabling technology for sharing public microarray data repositories and meta-analyses
ArrayWiki: an enabling technology for sharing public microarray data repositories and meta-analyses
Abstract
Background
A survey of microarray databases reveals that most of the repository contents and data models are heterogeneous (i.e., data o...
A Correspondence Between Normalization Strategies in Artificial and Biological Neural Networks
A Correspondence Between Normalization Strategies in Artificial and Biological Neural Networks
Abstract
A fundamental challenge at the interface of machine learning and neuroscience is to uncover computational principles that are shared bet...
Enhancing Cancerous Gene Selection and Classification for High-Dimensional Microarray Data Using a Novel Hybrid Filter and Differential Evolutionary Feature Selection
Enhancing Cancerous Gene Selection and Classification for High-Dimensional Microarray Data Using a Novel Hybrid Filter and Differential Evolutionary Feature Selection
Background: In recent years, microarray datasets have been used to store information about human genes and methods used to express the genes in order to successfully diagnose cance...

