Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Robust relationship inference in genome-wide association studies

View through CrossRef
Abstract Motivation: Genome-wide association studies (GWASs) have been widely used to map loci contributing to variation in complex traits and risk of diseases in humans. Accurate specification of familial relationships is crucial for family-based GWAS, as well as in population-based GWAS with unknown (or unrecognized) family structure. The family structure in a GWAS should be routinely investigated using the SNP data prior to the analysis of population structure or phenotype. Existing algorithms for relationship inference have a major weakness of estimating allele frequencies at each SNP from the entire sample, under a strong assumption of homogeneous population structure. This assumption is often untenable. Results: Here, we present a rapid algorithm for relationship inference using high-throughput genotype data typical of GWAS that allows the presence of unknown population substructure. The relationship of any pair of individuals can be precisely inferred by robust estimation of their kinship coefficient, independent of sample composition or population structure (sample invariance). We present simulation experiments to demonstrate that the algorithm has sufficient power to provide reliable inference on millions of unrelated pairs and thousands of relative pairs (up to 3rd-degree relationships). Application of our robust algorithm to HapMap and GWAS datasets demonstrates that it performs properly even under extreme population stratification, while algorithms assuming a homogeneous population give systematically biased results. Our extremely efficient implementation performs relationship inference on millions of pairs of individuals in a matter of minutes, dozens of times faster than the most efficient existing algorithm known to us. Availability: Our robust relationship inference algorithm is implemented in a freely available software package, KING, available for download at http://people.virginia.edu/∼wc9c/KING. Contact:  wmchen@virginia.edu Supplementary information:  Supplementary data are available at Bioinformatics online.
Title: Robust relationship inference in genome-wide association studies
Description:
Abstract Motivation: Genome-wide association studies (GWASs) have been widely used to map loci contributing to variation in complex traits and risk of diseases in humans.
Accurate specification of familial relationships is crucial for family-based GWAS, as well as in population-based GWAS with unknown (or unrecognized) family structure.
The family structure in a GWAS should be routinely investigated using the SNP data prior to the analysis of population structure or phenotype.
Existing algorithms for relationship inference have a major weakness of estimating allele frequencies at each SNP from the entire sample, under a strong assumption of homogeneous population structure.
This assumption is often untenable.
Results: Here, we present a rapid algorithm for relationship inference using high-throughput genotype data typical of GWAS that allows the presence of unknown population substructure.
The relationship of any pair of individuals can be precisely inferred by robust estimation of their kinship coefficient, independent of sample composition or population structure (sample invariance).
We present simulation experiments to demonstrate that the algorithm has sufficient power to provide reliable inference on millions of unrelated pairs and thousands of relative pairs (up to 3rd-degree relationships).
Application of our robust algorithm to HapMap and GWAS datasets demonstrates that it performs properly even under extreme population stratification, while algorithms assuming a homogeneous population give systematically biased results.
Our extremely efficient implementation performs relationship inference on millions of pairs of individuals in a matter of minutes, dozens of times faster than the most efficient existing algorithm known to us.
Availability: Our robust relationship inference algorithm is implemented in a freely available software package, KING, available for download at http://people.
virginia.
edu/∼wc9c/KING.
Contact:  wmchen@virginia.
edu Supplementary information:  Supplementary data are available at Bioinformatics online.

Related Results

Are Cervical Ribs Indicators of Childhood Cancer? A Narrative Review
Are Cervical Ribs Indicators of Childhood Cancer? A Narrative Review
Abstract A cervical rib (CR), also known as a supernumerary or extra rib, is an additional rib that forms above the first rib, resulting from the overgrowth of the transverse proce...
Whole Genome Resequencing and 1000 Genomes Project
Whole Genome Resequencing and 1000 Genomes Project
Abstract The recent advances in sequencing technologies have enabled the whole human genome to be sequenced within weeks. To date, several human...
Latency-Critical Inference Serving for Deep Learning
Latency-Critical Inference Serving for Deep Learning
Deep learning (DL) technology has made remarkable strides in terms of accuracy through the advancement of sophisticated and large deep neural networks (DNNs). Yet, its adoption in ...
Evolutionary Grammatical Inference
Evolutionary Grammatical Inference
Grammatical Inference (also known as grammar induction) is the problem of learning a grammar for a language from a set of examples. In a broad sense, some data is presented to the ...
Genome analysis ofElytrigia pycnanthaandThinopyrum junceiformeand of their putative natural hybrid using the GISH technique
Genome analysis ofElytrigia pycnanthaandThinopyrum junceiformeand of their putative natural hybrid using the GISH technique
Genomic in situ hybridization (GISH), using genomic DNA probes from Thinopyrum elongatum (Host) D.R. Dewey (E genome, 2n = 14), Th. bessarabicum (Savul. & Rayss) A. Löve (J gen...
Comparative genomics reveals insights into anuran genome size evolution
Comparative genomics reveals insights into anuran genome size evolution
Abstract Background Amphibians, particularly anurans, display an enormous variation in genome size. Due to the unavailability of whole genome datase...
Screening Deep Learning Inference Accelerators at the Production Lines
Screening Deep Learning Inference Accelerators at the Production Lines
Artificial Intelligence (AI) accelerators can be divided into two main buckets, one for training and another for inference over the trained models. Computation results of AI infere...
Next Generation Sequencing Technologies and Their Applications
Next Generation Sequencing Technologies and Their Applications
Abstract The advances in next generation sequencing (NGS) technologies have tremendous impacts on the studies of structural and f...

Back to Top