Javascript must be enabled to continue!
Linkage analysis with sequential imputation
View through CrossRef
AbstractMultilocus calculations, using all available information on all pedigree members, are important for linkage analysis. Exact calculation methods in linkage analysis are limited in either the number of loci or the number of pedigree members they can handle. In this article, we propose a Monte Carlo method for linkage analysis based on sequential imputation. Unlike exact methods, sequential imputation can handle large pedigrees with a moderate number of loci in its current implementation. This Monte Carlo method is an application of importance sampling, in which we sequentially impute ordered genotypes locus by locus, and then impute inheritance vectors conditioned on these genotypes. The resulting inheritance vectors, together with the importance sampling weights, are used to derive a consistent estimator of any linkage statistic of interest. The linkage statistic can be parametric or nonparametric; we focus on nonparametric linkage statistics. We demonstrate that accurate estimates can be achieved within a reasonable computing time. A simulation study illustrates the potential gain in power using our method for multilocus linkage analysis with large pedigrees. We simulated data at six markers under three models. We analyzed them using both sequential imputation and GENEHUNTER. GENEHUNTER had to drop between 38–54% of pedigree members, whereas our method was able to use all pedigree members. The power gains of using all pedigree members were substantial under 2 of the 3 models. We implemented sequential imputation for multilocus linkage analysis in a user‐friendly software package called SIMPLE. Genet Epidemiol 25:25–35, 2003. © 2003 Wiley‐Liss, Inc.
Title: Linkage analysis with sequential imputation
Description:
AbstractMultilocus calculations, using all available information on all pedigree members, are important for linkage analysis.
Exact calculation methods in linkage analysis are limited in either the number of loci or the number of pedigree members they can handle.
In this article, we propose a Monte Carlo method for linkage analysis based on sequential imputation.
Unlike exact methods, sequential imputation can handle large pedigrees with a moderate number of loci in its current implementation.
This Monte Carlo method is an application of importance sampling, in which we sequentially impute ordered genotypes locus by locus, and then impute inheritance vectors conditioned on these genotypes.
The resulting inheritance vectors, together with the importance sampling weights, are used to derive a consistent estimator of any linkage statistic of interest.
The linkage statistic can be parametric or nonparametric; we focus on nonparametric linkage statistics.
We demonstrate that accurate estimates can be achieved within a reasonable computing time.
A simulation study illustrates the potential gain in power using our method for multilocus linkage analysis with large pedigrees.
We simulated data at six markers under three models.
We analyzed them using both sequential imputation and GENEHUNTER.
GENEHUNTER had to drop between 38–54% of pedigree members, whereas our method was able to use all pedigree members.
The power gains of using all pedigree members were substantial under 2 of the 3 models.
We implemented sequential imputation for multilocus linkage analysis in a user‐friendly software package called SIMPLE.
Genet Epidemiol 25:25–35, 2003.
© 2003 Wiley‐Liss, Inc.
Related Results
Advanced methods for missing values imputation based on similarity learning
Advanced methods for missing values imputation based on similarity learning
The real-world data analysis and processing using data mining techniques often are facing observations that contain missing values. The main challenge of mining datasets is the exi...
Imputation Accuracy Across Global Human Populations
Imputation Accuracy Across Global Human Populations
Abstract
Genotype imputation is now fundamental for genome-wide association studies but lacks fairness due to the underrepresentation of populations with non-Europe...
A Coalescent Model for Genotype Imputation
A Coalescent Model for Genotype Imputation
AbstractThe potential for imputed genotypes to enhance an analysis of genetic data depends largely on the accuracy of imputation, which in turn depends on properties of the referen...
Evaluation of sequencing strategies for whole-genome imputation with hybrid peeling
Evaluation of sequencing strategies for whole-genome imputation with hybrid peeling
Abstract
Background
For assembling large whole-genome sequence datasets to be used routinely in research and breeding, the sequ...
A unifying framework for summary statistic imputation
A unifying framework for summary statistic imputation
Abstract
Imputation has been widely utilized to aid and interpret the results of Genome-Wide Association Studies(GWAS). Imputation can increase t...
A comparative study of imputation techniques for missing values in healthcare diagnostic datasets
A comparative study of imputation techniques for missing values in healthcare diagnostic datasets
Abstract
Missing values are a common feature of real-world datasets, particularly in healthcare data. This can be challenging when applying machine learning algor...
A Density-Adaptive Hybrid Linkage Criterion for Agglomerative Hierarchical Clustering
A Density-Adaptive Hybrid Linkage Criterion for Agglomerative Hierarchical Clustering
Hierarchical clustering is a widely used unsupervised learning technique due to its ability to uncover nested data structures without requiring prior knowledge of the number of clu...
Genotype Imputation
Genotype Imputation
Abstract
A missing data problem arises in genetic epidemiological studies when genotypes of particular markers are unavailable fo...

