Javascript must be enabled to continue!
A Coalescent Model for Genotype Imputation
View through CrossRef
AbstractThe potential for imputed genotypes to enhance an analysis of genetic data depends largely on the accuracy of imputation, which in turn depends on properties of the reference panel of template haplotypes used to perform the imputation. To provide a basis for exploring how properties of the reference panel affect imputation accuracy theoretically rather than with computationally intensive imputation experiments, we introduce a coalescent model that considers imputation accuracy in terms of population-genetic parameters. Our model allows us to investigate sampling designs in the frequently occurring scenario in which imputation targets and templates are sampled from different populations. In particular, we derive expressions for expected imputation accuracy as a function of reference panel size and divergence time between the reference and target populations. We find that a modestly sized “internal” reference panel from the same population as a target haplotype yields, on average, greater imputation accuracy than a larger “external” panel from a different population, even if the divergence time between the two populations is small. The improvement in accuracy for the internal panel increases with increasing divergence time between the target and reference populations. Thus, in humans, our model predicts that imputation accuracy can be improved by generating small population-specific custom reference panels to augment existing collections such as those of the HapMap or 1000 Genomes Projects. Our approach can be extended to understand additional factors that affect imputation accuracy in complex population-genetic settings, and the results can ultimately facilitate improvements in imputation study designs.
Oxford University Press (OUP)
Title: A Coalescent Model for Genotype Imputation
Description:
AbstractThe potential for imputed genotypes to enhance an analysis of genetic data depends largely on the accuracy of imputation, which in turn depends on properties of the reference panel of template haplotypes used to perform the imputation.
To provide a basis for exploring how properties of the reference panel affect imputation accuracy theoretically rather than with computationally intensive imputation experiments, we introduce a coalescent model that considers imputation accuracy in terms of population-genetic parameters.
Our model allows us to investigate sampling designs in the frequently occurring scenario in which imputation targets and templates are sampled from different populations.
In particular, we derive expressions for expected imputation accuracy as a function of reference panel size and divergence time between the reference and target populations.
We find that a modestly sized “internal” reference panel from the same population as a target haplotype yields, on average, greater imputation accuracy than a larger “external” panel from a different population, even if the divergence time between the two populations is small.
The improvement in accuracy for the internal panel increases with increasing divergence time between the target and reference populations.
Thus, in humans, our model predicts that imputation accuracy can be improved by generating small population-specific custom reference panels to augment existing collections such as those of the HapMap or 1000 Genomes Projects.
Our approach can be extended to understand additional factors that affect imputation accuracy in complex population-genetic settings, and the results can ultimately facilitate improvements in imputation study designs.
Related Results
The Fractional Coalescent
The Fractional Coalescent
A new approach to the coalescent, the
fractional
coalescent (
f
-coalescent), is introduced. Two derivations...
Advanced methods for missing values imputation based on similarity learning
Advanced methods for missing values imputation based on similarity learning
The real-world data analysis and processing using data mining techniques often are facing observations that contain missing values. The main challenge of mining datasets is the exi...
The Impact of IL28B Gene Polymorphisms on Drug Responses
The Impact of IL28B Gene Polymorphisms on Drug Responses
To achieve high therapeutic efficacy in the patient, information on pharmacokinetics, pharmacodynamics, and pharmacogenetics is required. With the development of science and techno...
Genotype Imputation
Genotype Imputation
Abstract
A missing data problem arises in genetic epidemiological studies when genotypes of particular markers are unavailable fo...
Expression and polymorphism of genes in gallstones
Expression and polymorphism of genes in gallstones
ABSTRACT
Through the method of clinical case control study, to explore the expression and genetic polymorphism of KLF14 gene (rs4731702 and rs972283) and SR-B1 gene...
The Validity of the Coalescent Approximation for Large Samples
The Validity of the Coalescent Approximation for Large Samples
Abstract
The Kingman coalescent, widely used in genetics, is known to be a good approximation when the sample size is small relative to the popul...
Imputation Accuracy Across Global Human Populations
Imputation Accuracy Across Global Human Populations
Abstract
Genotype imputation is now fundamental for genome-wide association studies but lacks fairness due to the underrepresentation of populations with non-Europe...
PD‐1 expression on CTL may be related to more severe liver damage in CHB patients with HBV genotype C than in those with genotype B infection
PD‐1 expression on CTL may be related to more severe liver damage in CHB patients with HBV genotype C than in those with genotype B infection
SummaryCompared with Chronic hepatitis B (CHB) patients infected with genotype B, those infected with genotype C have higher hepatic histopathological activity and higher level of ...

