Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Imputation Accuracy Across Global Human Populations

View through CrossRef
Abstract Genotype imputation is now fundamental for genome-wide association studies but lacks fairness due to the underrepresentation of populations with non-European ancestries. The state-of-the-art imputation reference panel released by the Trans-Omics for Precision Medicine (TOPMed) initiative contains a substantial number of admixed African-ancestry and Hispanic/Latino samples to impute these populations with nearly the same accuracy as European-ancestry cohorts. However, imputation for populations primarily residing outside of North America may still fall short in performance due to persisting underrepresentation. To illustrate this point, we curated genome-wide array data from 23 publications published between 2008 to 2021. In total, we imputed over 43k individuals across 123 populations around the world. We identified a number of populations where imputation accuracy paled in comparison to that of European-ancestry populations. For instance, the mean imputation r-squared (Rsq) for 1-5% alleles in Saudi Arabians (N=1061), Vietnamese (N=1264), Thai (N=2435), and Papua New Guineans (N=776) were 0.79, 0.78, 0.76, and 0.62, respectively. In contrast, the mean Rsq ranged from 0.90 to 0.93 for comparable European populations matched in sample size and SNP content. Outside of Africa and Latin America, Rsq appeared to decrease as genetic distances to European reference increased, as predicted. Further analysis using sequencing data as ground truth suggested that imputation software may over-estimate imputation accuracy for non-European populations than European populations, suggesting further disparity between populations. Using 1496 whole genome sequenced individuals from Taiwan Biobank as a reference, we also assessed a strategy to improve imputation for non-European populations with meta-imputation, which can combine results from TOPMed with smaller population-specific reference panels. We found that meta-imputation in this design did not improve Rsq genome-wide. Taken together, our analysis suggests that with the current size of alternative reference panels, meta-imputation alone cannot improve imputation efficacy for underrepresented cohorts and we must ultimately strive to increase diversity and size to promote equity within genetics research.
Title: Imputation Accuracy Across Global Human Populations
Description:
Abstract Genotype imputation is now fundamental for genome-wide association studies but lacks fairness due to the underrepresentation of populations with non-European ancestries.
The state-of-the-art imputation reference panel released by the Trans-Omics for Precision Medicine (TOPMed) initiative contains a substantial number of admixed African-ancestry and Hispanic/Latino samples to impute these populations with nearly the same accuracy as European-ancestry cohorts.
However, imputation for populations primarily residing outside of North America may still fall short in performance due to persisting underrepresentation.
To illustrate this point, we curated genome-wide array data from 23 publications published between 2008 to 2021.
In total, we imputed over 43k individuals across 123 populations around the world.
We identified a number of populations where imputation accuracy paled in comparison to that of European-ancestry populations.
For instance, the mean imputation r-squared (Rsq) for 1-5% alleles in Saudi Arabians (N=1061), Vietnamese (N=1264), Thai (N=2435), and Papua New Guineans (N=776) were 0.
79, 0.
78, 0.
76, and 0.
62, respectively.
In contrast, the mean Rsq ranged from 0.
90 to 0.
93 for comparable European populations matched in sample size and SNP content.
Outside of Africa and Latin America, Rsq appeared to decrease as genetic distances to European reference increased, as predicted.
Further analysis using sequencing data as ground truth suggested that imputation software may over-estimate imputation accuracy for non-European populations than European populations, suggesting further disparity between populations.
Using 1496 whole genome sequenced individuals from Taiwan Biobank as a reference, we also assessed a strategy to improve imputation for non-European populations with meta-imputation, which can combine results from TOPMed with smaller population-specific reference panels.
We found that meta-imputation in this design did not improve Rsq genome-wide.
Taken together, our analysis suggests that with the current size of alternative reference panels, meta-imputation alone cannot improve imputation efficacy for underrepresented cohorts and we must ultimately strive to increase diversity and size to promote equity within genetics research.

Related Results

Advanced methods for missing values imputation based on similarity learning
Advanced methods for missing values imputation based on similarity learning
The real-world data analysis and processing using data mining techniques often are facing observations that contain missing values. The main challenge of mining datasets is the exi...
Evaluation of sequencing strategies for whole-genome imputation with hybrid peeling
Evaluation of sequencing strategies for whole-genome imputation with hybrid peeling
Abstract Background For assembling large whole-genome sequence datasets to be used routinely in research and breeding, the sequ...
A Coalescent Model for Genotype Imputation
A Coalescent Model for Genotype Imputation
AbstractThe potential for imputed genotypes to enhance an analysis of genetic data depends largely on the accuracy of imputation, which in turn depends on properties of the referen...
A comparative study of imputation techniques for missing values in healthcare diagnostic datasets
A comparative study of imputation techniques for missing values in healthcare diagnostic datasets
Abstract Missing values are a common feature of real-world datasets, particularly in healthcare data. This can be challenging when applying machine learning algor...
A unifying framework for summary statistic imputation
A unifying framework for summary statistic imputation
Abstract Imputation has been widely utilized to aid and interpret the results of Genome-Wide Association Studies(GWAS). Imputation can increase t...
Genotype Imputation
Genotype Imputation
Abstract A missing data problem arises in genetic epidemiological studies when genotypes of particular markers are unavailable fo...
GSimp: A Gibbs sampler based left-censored missing value imputation approach for metabolomics studies
GSimp: A Gibbs sampler based left-censored missing value imputation approach for metabolomics studies
Abstract Left-censored missing values commonly exist in targeted metabolomics datasets and can be considered as missing not at random (MNAR). Imp...
A Multiple Imputation Workflow for Handling Missing Covariate Data in Pharmacometrics Modeling
A Multiple Imputation Workflow for Handling Missing Covariate Data in Pharmacometrics Modeling
ABSTRACT Covariate missingness is a prevalent issue in pharmacometrics modeling. Incorrect handling of missing covariates can lead to biased parameter estimates, ...

Back to Top