Javascript must be enabled to continue!
LmTag: functional-enrichment and imputation-aware tag SNP selection for population-specific genotyping arrays
View through CrossRef
Abstract
Despite the rapid development of sequencing technology, single-nucleotide polymorphism (SNP) arrays are still the most cost-effective genotyping solutions for large-scale genomic research and applications. Recent years have witnessed the rapid development of numerous genotyping platforms of different sizes and designs, but population-specific platforms are still lacking, especially for those in developing countries. SNP arrays designed for these countries should be cost-effective (small size), yet incorporate key information needed to associate genotypes with traits. A key design principle for most current platforms is to improve genome-wide imputation so that more SNPs not included in the array (imputed SNPs) can be predicted. However, current tag SNP selection methods mostly focus on imputation accuracy and coverage, but not the functional content of the array. It is those functional SNPs that are most likely associated with traits. Here, we propose LmTag, a novel method for tag SNP selection that not only improves imputation performance but also prioritizes highly functional SNP markers. We apply LmTag on a wide range of populations using both public and in-house whole-genome sequencing databases. Our results show that LmTag improved both functional marker prioritization and genome-wide imputation accuracy compared to existing methods. This novel approach could contribute to the next generation genotyping arrays that provide excellent imputation capability as well as facilitate array-based functional genetic studies. Such arrays are particularly suitable for under-represented populations in developing countries or non-model species, where little genomics data are available while investment in genome sequencing or high-density SNP arrays is limited. $\textrm{LmTag}$ is available at: https://github.com/datngu/LmTag.
Oxford University Press (OUP)
Title: LmTag: functional-enrichment and imputation-aware tag SNP selection for population-specific genotyping arrays
Description:
Abstract
Despite the rapid development of sequencing technology, single-nucleotide polymorphism (SNP) arrays are still the most cost-effective genotyping solutions for large-scale genomic research and applications.
Recent years have witnessed the rapid development of numerous genotyping platforms of different sizes and designs, but population-specific platforms are still lacking, especially for those in developing countries.
SNP arrays designed for these countries should be cost-effective (small size), yet incorporate key information needed to associate genotypes with traits.
A key design principle for most current platforms is to improve genome-wide imputation so that more SNPs not included in the array (imputed SNPs) can be predicted.
However, current tag SNP selection methods mostly focus on imputation accuracy and coverage, but not the functional content of the array.
It is those functional SNPs that are most likely associated with traits.
Here, we propose LmTag, a novel method for tag SNP selection that not only improves imputation performance but also prioritizes highly functional SNP markers.
We apply LmTag on a wide range of populations using both public and in-house whole-genome sequencing databases.
Our results show that LmTag improved both functional marker prioritization and genome-wide imputation accuracy compared to existing methods.
This novel approach could contribute to the next generation genotyping arrays that provide excellent imputation capability as well as facilitate array-based functional genetic studies.
Such arrays are particularly suitable for under-represented populations in developing countries or non-model species, where little genomics data are available while investment in genome sequencing or high-density SNP arrays is limited.
$\textrm{LmTag}$ is available at: https://github.
com/datngu/LmTag.
Related Results
Novel design of imputation-enabled SNP arrays for breeding and research applications supporting multi-species hybridisation
Novel design of imputation-enabled SNP arrays for breeding and research applications supporting multi-species hybridisation
AbstractArray-based SNP genotyping platforms have low genotype error and missing data rates compared to genotyping-by-sequencing technologies. However, design decisions used to cre...
Advanced methods for missing values imputation based on similarity learning
Advanced methods for missing values imputation based on similarity learning
The real-world data analysis and processing using data mining techniques often are facing observations that contain missing values. The main challenge of mining datasets is the exi...
Silhouette scores for assessment of SNP genotype clusters
Silhouette scores for assessment of SNP genotype clusters
Abstract
Background
High-throughput genotyping of single nucleotide polymorphisms (SNPs) generates large amounts of data. In many SN...
Imputation Accuracy Across Global Human Populations
Imputation Accuracy Across Global Human Populations
Abstract
Genotype imputation is now fundamental for genome-wide association studies but lacks fairness due to the underrepresentation of populations with non-Europe...
A comparative study of imputation techniques for missing values in healthcare diagnostic datasets
A comparative study of imputation techniques for missing values in healthcare diagnostic datasets
Abstract
Missing values are a common feature of real-world datasets, particularly in healthcare data. This can be challenging when applying machine learning algor...
A Coalescent Model for Genotype Imputation
A Coalescent Model for Genotype Imputation
AbstractThe potential for imputed genotypes to enhance an analysis of genetic data depends largely on the accuracy of imputation, which in turn depends on properties of the referen...
Evaluation of sequencing strategies for whole-genome imputation with hybrid peeling
Evaluation of sequencing strategies for whole-genome imputation with hybrid peeling
Abstract
Background
For assembling large whole-genome sequence datasets to be used routinely in research and breeding, the sequ...
Genotype Imputation
Genotype Imputation
Abstract
A missing data problem arises in genetic epidemiological studies when genotypes of particular markers are unavailable fo...

