Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Uncovering the consequences of batch effect associated missing values in omics data analysis

View through CrossRef
ABSTRACTStatistical analyses in high-dimensional omics data are often hampered by the presence of batch effects (BEs) and missing values (MVs), but the interaction between these two issues is not well-studied nor understood. MVs may manifest as a BE when their proportions differ across batches. These are termed as Batch-Effect Associated Missing values (BEAMs). We hypothesized that BEAMs in data may introduce bias which can impede the performance of missing value imputation (MVI). To test this, we simulated data with two batches, then introduced over 100 iterations, either 20% and 40% MVs in each batch (BEAMs) or 30% in both (control). K-nearest neighbours (KNN) was then used to perform MVI, in a typical global approach (M1) and a supposed superior batch-sensitized approach (M2). BEs were then corrected using ComBat. The effectiveness of the MVI was evaluated by its imputation accuracy and true and false positive rates. Notably, when BEAMs existed, M2 was generally undesirable as the differing application of MV filtering in M1 and M2 strategies resulted in an overall coverage deficiency. Additionally, both M1 and M2 strategies suffered in the presence of BEAMs, highlighting the need for a novel approach to handle MVI in data with BEAMs.Author summaryData in high-throughput omics data are often combined from different sources (batches), which creates batch effects in the data. Missing values are a common occurrence in these data, and their proportions are assumed to be equal across batches. However, instances exist when these proportions vary between batches, such as one batch having more missing values than another, resulting in batch effect associated missing values. Missing values are often dealt with through missing value imputation, but whether the variation in missing value proportions across batches affects imputation outcomes is unknown. In this paper, we investigate the consequence of performing imputation when this issue persists. We simulated data with equal and unequal missing value proportions, then assessed the performance of k-nearest neighbours imputation by its imputation accuracy and downstream analysis outcomes. This revealed that unequal missing value proportions worsens imputation and establishes the need for smarter imputation strategies to handle this complication.
Cold Spring Harbor Laboratory
Title: Uncovering the consequences of batch effect associated missing values in omics data analysis
Description:
ABSTRACTStatistical analyses in high-dimensional omics data are often hampered by the presence of batch effects (BEs) and missing values (MVs), but the interaction between these two issues is not well-studied nor understood.
MVs may manifest as a BE when their proportions differ across batches.
These are termed as Batch-Effect Associated Missing values (BEAMs).
We hypothesized that BEAMs in data may introduce bias which can impede the performance of missing value imputation (MVI).
To test this, we simulated data with two batches, then introduced over 100 iterations, either 20% and 40% MVs in each batch (BEAMs) or 30% in both (control).
K-nearest neighbours (KNN) was then used to perform MVI, in a typical global approach (M1) and a supposed superior batch-sensitized approach (M2).
BEs were then corrected using ComBat.
The effectiveness of the MVI was evaluated by its imputation accuracy and true and false positive rates.
Notably, when BEAMs existed, M2 was generally undesirable as the differing application of MV filtering in M1 and M2 strategies resulted in an overall coverage deficiency.
Additionally, both M1 and M2 strategies suffered in the presence of BEAMs, highlighting the need for a novel approach to handle MVI in data with BEAMs.
Author summaryData in high-throughput omics data are often combined from different sources (batches), which creates batch effects in the data.
Missing values are a common occurrence in these data, and their proportions are assumed to be equal across batches.
However, instances exist when these proportions vary between batches, such as one batch having more missing values than another, resulting in batch effect associated missing values.
Missing values are often dealt with through missing value imputation, but whether the variation in missing value proportions across batches affects imputation outcomes is unknown.
In this paper, we investigate the consequence of performing imputation when this issue persists.
We simulated data with equal and unequal missing value proportions, then assessed the performance of k-nearest neighbours imputation by its imputation accuracy and downstream analysis outcomes.
This revealed that unequal missing value proportions worsens imputation and establishes the need for smarter imputation strategies to handle this complication.

Related Results

Why Pakistan Must Lead in Regional Multi-Omics Research for Precision Medicine
Why Pakistan Must Lead in Regional Multi-Omics Research for Precision Medicine
Precision medicine has emerged as one of the most transformative movements in global healthcare, shifting the clinical emphasis from generalized treatments to highly individualized...
Contribuición al estudio de la operación de destilación discontinua mediante simulación
Contribuición al estudio de la operación de destilación discontinua mediante simulación
El treball presentat en aquesta tesi pretén contribuir a l'estudi mediambiental de la limitació de les emissions de components orgànics volàtils (VOCs), degudes a l'ús de dissolven...
The importance of batch sensitization in missing value imputation
The importance of batch sensitization in missing value imputation
AbstractData analysis is complex due to a myriad of technical problems. Amongst these, missing values and batch effects are endemic. Although many methods have been developed for m...
Omics-Based Investigations of Breast Cancer
Omics-Based Investigations of Breast Cancer
Breast cancer (BC) is characterized by an extensive genotypic and phenotypic heterogeneity. In-depth investigations into the molecular bases of BC phenotypes, carcinogenesis, progr...
TEMINET: A Co-Informative and Trustworthy Multi-Omics Integration Network for Diagnostic Prediction
TEMINET: A Co-Informative and Trustworthy Multi-Omics Integration Network for Diagnostic Prediction
Abstract Advancing the domain of biomedical investigation, integrated multi-omics data have shown exceptional performance in elucidating complex human diseases. How...
Advanced methods for missing values imputation based on similarity learning
Advanced methods for missing values imputation based on similarity learning
The real-world data analysis and processing using data mining techniques often are facing observations that contain missing values. The main challenge of mining datasets is the exi...
OmicsTIDE: Interactive Exploration of Trends in Multi-Omics Data
OmicsTIDE: Interactive Exploration of Trends in Multi-Omics Data
Abstract Motivation The increasing amount of data produced by omics technologies has significantly improved the understanding o...

Back to Top