Javascript must be enabled to continue!
Benchmarking and improving imputation approaches for recurrent inversions in the human genome
View through CrossRef
Inversions are a type of structural variant that are involved in phenotypic differences among individuals. Due to certain features, such as the fact of usually implying no loss or gain of DNA or the presence of large inverted repeats at their breakpoints, the characterization of inversions is quite difficult. It has been recently found that in humans many inversions are recurrent and they are not linked to other genomic variants. For this reason, the effect of recurrent inversions has been largely missed in current GWAS, and it is necessary to develop new methods to predict inversion genotypes accurately in the datasets of interest. Here, we have done a benchmarking analysis of the genotype predictions among different imputation tools, comparing IMPUTE2 with other softwares, such as IMPUTE5, BEAGLE or/and scoreInvHap. The accuracy was
calculated as the r
2
between experimental and imputed genotypes. From our set of 130 experimentally genotyped inversions 55 are recurrent. We found out that 23 and 18 out of 55 were imputable (r
2
> 0.8) in European and African populations, respectively, but it varies among softwares. Nevertheless, this ratio increases to 26/55 for both populations when we filter out samples with a post-imputation genotype probability lower than 0.8. Finally, we are also testing a tool based on deep learning which could avoid the HMM-based algorithm limitations, as the position or linkage dependence and region complexity, to increase our catalogue of imputable inversions.
Title: Benchmarking and improving imputation approaches for recurrent inversions in the human genome
Description:
Inversions are a type of structural variant that are involved in phenotypic differences among individuals.
Due to certain features, such as the fact of usually implying no loss or gain of DNA or the presence of large inverted repeats at their breakpoints, the characterization of inversions is quite difficult.
It has been recently found that in humans many inversions are recurrent and they are not linked to other genomic variants.
For this reason, the effect of recurrent inversions has been largely missed in current GWAS, and it is necessary to develop new methods to predict inversion genotypes accurately in the datasets of interest.
Here, we have done a benchmarking analysis of the genotype predictions among different imputation tools, comparing IMPUTE2 with other softwares, such as IMPUTE5, BEAGLE or/and scoreInvHap.
The accuracy was
calculated as the r
2
between experimental and imputed genotypes.
From our set of 130 experimentally genotyped inversions 55 are recurrent.
We found out that 23 and 18 out of 55 were imputable (r
2
> 0.
8) in European and African populations, respectively, but it varies among softwares.
Nevertheless, this ratio increases to 26/55 for both populations when we filter out samples with a post-imputation genotype probability lower than 0.
8.
Finally, we are also testing a tool based on deep learning which could avoid the HMM-based algorithm limitations, as the position or linkage dependence and region complexity, to increase our catalogue of imputable inversions.
Related Results
Advanced methods for missing values imputation based on similarity learning
Advanced methods for missing values imputation based on similarity learning
The real-world data analysis and processing using data mining techniques often are facing observations that contain missing values. The main challenge of mining datasets is the exi...
Frequency and clinical significance of chromosomal inversions prenatally diagnosed by second trimester amniocentesis
Frequency and clinical significance of chromosomal inversions prenatally diagnosed by second trimester amniocentesis
AbstractTo compare the frequency and clinical significance of familial and de novo chromosomal inversions during prenatal diagnosis. This was a retrospective study of inversions di...
Imputation Accuracy Across Global Human Populations
Imputation Accuracy Across Global Human Populations
Abstract
Genotype imputation is now fundamental for genome-wide association studies but lacks fairness due to the underrepresentation of populations with non-Europe...
Inversions and Evolution of the Human Genome
Inversions and Evolution of the Human Genome
AbstractInversions are a type of naturally occurring DNA mutation in which the sequence of the DNA is reversed, resulting in it being read in the opposite direction to the wild typ...
Evolving benchmarking practices: a review for research perspectives
Evolving benchmarking practices: a review for research perspectives
PurposeThe purpose of this study is to review a major section of the literature on benchmarking practices in order to achieve better perspectives for emerging benchmarking research...
Evaluation of sequencing strategies for whole-genome imputation with hybrid peeling
Evaluation of sequencing strategies for whole-genome imputation with hybrid peeling
Abstract
Background
For assembling large whole-genome sequence datasets to be used routinely in research and breeding, the sequ...
Perceptions about benchmarking best practices among French managers: an exploratory survey
Perceptions about benchmarking best practices among French managers: an exploratory survey
PurposeThe purpose of this study is to present a discussion on the most commonly accepted benchmarking norms in the USA, the lessons learned from benchmarking experiences and see h...
A unifying framework for summary statistic imputation
A unifying framework for summary statistic imputation
Abstract
Imputation has been widely utilized to aid and interpret the results of Genome-Wide Association Studies(GWAS). Imputation can increase t...

