Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Genotype imputation using the Positional Burrows Wheeler Transform

View through CrossRef
Abstract Genotype imputation is the process of predicting unobserved genotypes in a sample of individuals using a reference panel of haplotypes. In the last 10 years reference panels have increased in size by more than 100 fold. Increasing reference panel size improves accuracy of markers with low minor allele frequencies but poses ever increasing computational challenges for imputation methods. Here we present IMPUTE5, a genotype imputation method that can scale to reference panels with millions of samples. This method continues to refine the observation made in the IMPUTE2 method, that accuracy is optimized via use of a custom subset of haplotypes when imputing each individual. It achieves fast, accurate, and memory-efficient imputation by selecting haplotypes using the Positional Burrows Wheeler Transform (PBWT). By using the PBWT data structure at genotyped markers, IMPUTE5 identifies locally best matching haplotypes and long identical by state segments. The method then uses the selected haplotypes as conditioning states within the IMPUTE model. Using the HRC reference panel, which has ~65,000 haplotypes, we show that IMPUTE5 is up to 30x faster than MINIMAC4 and up to 3x faster than BEAGLE5.1, and uses less memory than both these methods. Using simulated reference panels we show that IMPUTE5 scales sub-linearly with reference panel size. For example, keeping the number of imputed markers constant, increasing the reference panel size from 10,000 to 1 million haplotypes requires less than twice the computation time. As the reference panel increases in size IMPUTE5 is able to utilize a smaller number of reference haplotypes, thus reducing computational cost. Author summary Genome-wide association studies (GWAS) typically use microarray technology to measure genotypes at several hundred thousand positions in the genome. However reference panels of genetic variation consist of haplotype data at >100 fold more positions in the genome. Genotype imputation makes genotype predictions at all the reference panel sites using the GWAS data. Reference panels are continuing to grow in size and this improves accuracy of the predictions, however methods need to be able to scale to increased size. We have developed a new version of the popular IMPUTE software than can handle referenece panels with millions of haplotypes, and has better performance than other published approaches. A notable property of the new method is that it scales sub-linearly with reference panel size. Keeping the number of imputed markers constant, a 100 fold increase in reference panel size requires less than twice the computation time.
Title: Genotype imputation using the Positional Burrows Wheeler Transform
Description:
Abstract Genotype imputation is the process of predicting unobserved genotypes in a sample of individuals using a reference panel of haplotypes.
In the last 10 years reference panels have increased in size by more than 100 fold.
Increasing reference panel size improves accuracy of markers with low minor allele frequencies but poses ever increasing computational challenges for imputation methods.
Here we present IMPUTE5, a genotype imputation method that can scale to reference panels with millions of samples.
This method continues to refine the observation made in the IMPUTE2 method, that accuracy is optimized via use of a custom subset of haplotypes when imputing each individual.
It achieves fast, accurate, and memory-efficient imputation by selecting haplotypes using the Positional Burrows Wheeler Transform (PBWT).
By using the PBWT data structure at genotyped markers, IMPUTE5 identifies locally best matching haplotypes and long identical by state segments.
The method then uses the selected haplotypes as conditioning states within the IMPUTE model.
Using the HRC reference panel, which has ~65,000 haplotypes, we show that IMPUTE5 is up to 30x faster than MINIMAC4 and up to 3x faster than BEAGLE5.
1, and uses less memory than both these methods.
Using simulated reference panels we show that IMPUTE5 scales sub-linearly with reference panel size.
For example, keeping the number of imputed markers constant, increasing the reference panel size from 10,000 to 1 million haplotypes requires less than twice the computation time.
As the reference panel increases in size IMPUTE5 is able to utilize a smaller number of reference haplotypes, thus reducing computational cost.
Author summary Genome-wide association studies (GWAS) typically use microarray technology to measure genotypes at several hundred thousand positions in the genome.
However reference panels of genetic variation consist of haplotype data at >100 fold more positions in the genome.
Genotype imputation makes genotype predictions at all the reference panel sites using the GWAS data.
Reference panels are continuing to grow in size and this improves accuracy of the predictions, however methods need to be able to scale to increased size.
We have developed a new version of the popular IMPUTE software than can handle referenece panels with millions of haplotypes, and has better performance than other published approaches.
A notable property of the new method is that it scales sub-linearly with reference panel size.
Keeping the number of imputed markers constant, a 100 fold increase in reference panel size requires less than twice the computation time.

Related Results

The Impact of IL28B Gene Polymorphisms on Drug Responses
The Impact of IL28B Gene Polymorphisms on Drug Responses
To achieve high therapeutic efficacy in the patient, information on pharmacokinetics, pharmacodynamics, and pharmacogenetics is required. With the development of science and techno...
Genotype Imputation
Genotype Imputation
Abstract A missing data problem arises in genetic epidemiological studies when genotypes of particular markers are unavailable fo...
Expression and polymorphism of genes in gallstones
Expression and polymorphism of genes in gallstones
ABSTRACT Through the method of clinical case control study, to explore the expression and genetic polymorphism of KLF14 gene (rs4731702 and rs972283) and SR-B1 gene...
Comparison of San Joaquin kit fox den and California ground squirrel burrow attributes
Comparison of San Joaquin kit fox den and California ground squirrel burrow attributes
Endangered San Joaquin kit foxes (Vulpes macrotis mutica; SJKF) and California ground squirrels (Otospermophilus beecheyi; CAGS) occur sympatrically in many locations. CAGS can con...
A Coalescent Model for Genotype Imputation
A Coalescent Model for Genotype Imputation
AbstractThe potential for imputed genotypes to enhance an analysis of genetic data depends largely on the accuracy of imputation, which in turn depends on properties of the referen...
weIMPUTE: A User-Friendly Web-Based Genotype Imputation Platform
weIMPUTE: A User-Friendly Web-Based Genotype Imputation Platform
Abstract Genotype imputation is a critical preprocessing step in genome-wide association studies (GWAS), enhancing statistical power for detecting associated single...
Evaluation of sequencing strategies for whole-genome imputation with hybrid peeling
Evaluation of sequencing strategies for whole-genome imputation with hybrid peeling
Abstract Background For assembling large whole-genome sequence datasets to be used routinely in research and breeding, the sequ...

Back to Top