Javascript must be enabled to continue!
Robust relationship inference in genome-wide association studies
View through CrossRef
Abstract
Motivation: Genome-wide association studies (GWASs) have been widely used to map loci contributing to variation in complex traits and risk of diseases in humans. Accurate specification of familial relationships is crucial for family-based GWAS, as well as in population-based GWAS with unknown (or unrecognized) family structure. The family structure in a GWAS should be routinely investigated using the SNP data prior to the analysis of population structure or phenotype. Existing algorithms for relationship inference have a major weakness of estimating allele frequencies at each SNP from the entire sample, under a strong assumption of homogeneous population structure. This assumption is often untenable.
Results: Here, we present a rapid algorithm for relationship inference using high-throughput genotype data typical of GWAS that allows the presence of unknown population substructure. The relationship of any pair of individuals can be precisely inferred by robust estimation of their kinship coefficient, independent of sample composition or population structure (sample invariance). We present simulation experiments to demonstrate that the algorithm has sufficient power to provide reliable inference on millions of unrelated pairs and thousands of relative pairs (up to 3rd-degree relationships). Application of our robust algorithm to HapMap and GWAS datasets demonstrates that it performs properly even under extreme population stratification, while algorithms assuming a homogeneous population give systematically biased results. Our extremely efficient implementation performs relationship inference on millions of pairs of individuals in a matter of minutes, dozens of times faster than the most efficient existing algorithm known to us.
Availability: Our robust relationship inference algorithm is implemented in a freely available software package, KING, available for download at http://people.virginia.edu/∼wc9c/KING.
Contact: wmchen@virginia.edu
Supplementary information: Supplementary data are available at Bioinformatics online.
Oxford University Press (OUP)
Title: Robust relationship inference in genome-wide association studies
Description:
Abstract
Motivation: Genome-wide association studies (GWASs) have been widely used to map loci contributing to variation in complex traits and risk of diseases in humans.
Accurate specification of familial relationships is crucial for family-based GWAS, as well as in population-based GWAS with unknown (or unrecognized) family structure.
The family structure in a GWAS should be routinely investigated using the SNP data prior to the analysis of population structure or phenotype.
Existing algorithms for relationship inference have a major weakness of estimating allele frequencies at each SNP from the entire sample, under a strong assumption of homogeneous population structure.
This assumption is often untenable.
Results: Here, we present a rapid algorithm for relationship inference using high-throughput genotype data typical of GWAS that allows the presence of unknown population substructure.
The relationship of any pair of individuals can be precisely inferred by robust estimation of their kinship coefficient, independent of sample composition or population structure (sample invariance).
We present simulation experiments to demonstrate that the algorithm has sufficient power to provide reliable inference on millions of unrelated pairs and thousands of relative pairs (up to 3rd-degree relationships).
Application of our robust algorithm to HapMap and GWAS datasets demonstrates that it performs properly even under extreme population stratification, while algorithms assuming a homogeneous population give systematically biased results.
Our extremely efficient implementation performs relationship inference on millions of pairs of individuals in a matter of minutes, dozens of times faster than the most efficient existing algorithm known to us.
Availability: Our robust relationship inference algorithm is implemented in a freely available software package, KING, available for download at http://people.
virginia.
edu/∼wc9c/KING.
Contact: wmchen@virginia.
edu
Supplementary information: Supplementary data are available at Bioinformatics online.
Related Results
Are Cervical Ribs Indicators of Childhood Cancer? A Narrative Review
Are Cervical Ribs Indicators of Childhood Cancer? A Narrative Review
Abstract
A cervical rib (CR), also known as a supernumerary or extra rib, is an additional rib that forms above the first rib, resulting from the overgrowth of the transverse proce...
Whole Genome Resequencing and 1000 Genomes Project
Whole Genome Resequencing and 1000 Genomes Project
Abstract
The recent advances in sequencing technologies have enabled the whole human genome to be sequenced within weeks. To date, several human...
Latency-Critical Inference Serving for Deep Learning
Latency-Critical Inference Serving for Deep Learning
Deep learning (DL) technology has made remarkable strides in terms of accuracy through the advancement of sophisticated and large deep neural networks (DNNs). Yet, its adoption in ...
Evolutionary Grammatical Inference
Evolutionary Grammatical Inference
Grammatical Inference (also known as grammar induction) is the problem of learning a grammar for a language from a set of examples. In a broad sense, some data is presented to the ...
Genome analysis ofElytrigia pycnanthaandThinopyrum junceiformeand of their putative natural hybrid using the GISH technique
Genome analysis ofElytrigia pycnanthaandThinopyrum junceiformeand of their putative natural hybrid using the GISH technique
Genomic in situ hybridization (GISH), using genomic DNA probes from Thinopyrum elongatum (Host) D.R. Dewey (E genome, 2n = 14), Th. bessarabicum (Savul. & Rayss) A. Löve (J gen...
Comparative genomics reveals insights into anuran genome size evolution
Comparative genomics reveals insights into anuran genome size evolution
Abstract
Background
Amphibians, particularly anurans, display an enormous variation in genome size. Due to the unavailability of whole genome datase...
Screening Deep Learning Inference Accelerators at the Production Lines
Screening Deep Learning Inference Accelerators at the Production Lines
Artificial Intelligence (AI) accelerators can be divided into two main buckets, one for training and another for inference over the trained models. Computation results of AI infere...
Next Generation Sequencing Technologies and Their Applications
Next Generation Sequencing Technologies and Their Applications
Abstract
The advances in next generation sequencing (NGS) technologies have tremendous impacts on the studies of structural and f...

