Javascript must be enabled to continue!
Characterization of segmental duplications and large inversions using Linked-Reads
View through CrossRef
Abstract
Many algorithms aimed at characterizing genomic structural variation (SV) have been developed since the inception of high-throughput sequencing. However, the full spectrum of SVs in the human genome is not yet assessed. Most of the existing methods focus on discovery and genotyping of deletions, insertions, and mobile elements. Detection of balanced SVs with no gain or loss of genomic segments (e.g., inversions) is particularly a challenging task. Long read sequencing has been leveraged to find short inversions but there is still a need to develop methods to detect large genomic inversions. Furthermore, currently there are no algorithms to predict the insertion locus of large interspersed segmental duplications.
Here we propose novel algorithms to characterize large (>40Kbp) interspersed segmental duplications and (>80Kbp) inversions using Linked-Read sequencing data. Linked-Read sequencing provides long range information, where Illumina reads are tagged with barcodes that can be used to assign short reads to pools of larger (30-50 Kbp) molecules. Our methods rely on
split molecule
sequence signature that we have previously described [11]. Similar to the split read, split molecules refer to large segments of DNA that span an SV breakpoint. Therefore, when mapped to the reference genome, the mapping of these segments would be discontinuous. We redesign our earlier algorithm, VALOR, to specifically leverage Linked-Read sequencing data to discover large inversions and characterize interspersed segmental duplications. We implement our new algorithms in a new software package, called VALOR
2
.
Availability
VALOR
2
is available at
https://github.com/BilkentCompGen/valor
.
Title: Characterization of segmental duplications and large inversions using Linked-Reads
Description:
Abstract
Many algorithms aimed at characterizing genomic structural variation (SV) have been developed since the inception of high-throughput sequencing.
However, the full spectrum of SVs in the human genome is not yet assessed.
Most of the existing methods focus on discovery and genotyping of deletions, insertions, and mobile elements.
Detection of balanced SVs with no gain or loss of genomic segments (e.
g.
, inversions) is particularly a challenging task.
Long read sequencing has been leveraged to find short inversions but there is still a need to develop methods to detect large genomic inversions.
Furthermore, currently there are no algorithms to predict the insertion locus of large interspersed segmental duplications.
Here we propose novel algorithms to characterize large (>40Kbp) interspersed segmental duplications and (>80Kbp) inversions using Linked-Read sequencing data.
Linked-Read sequencing provides long range information, where Illumina reads are tagged with barcodes that can be used to assign short reads to pools of larger (30-50 Kbp) molecules.
Our methods rely on
split molecule
sequence signature that we have previously described [11].
Similar to the split read, split molecules refer to large segments of DNA that span an SV breakpoint.
Therefore, when mapped to the reference genome, the mapping of these segments would be discontinuous.
We redesign our earlier algorithm, VALOR, to specifically leverage Linked-Read sequencing data to discover large inversions and characterize interspersed segmental duplications.
We implement our new algorithms in a new software package, called VALOR
2
.
Availability
VALOR
2
is available at
https://github.
com/BilkentCompGen/valor
.
Related Results
Frequency and clinical significance of chromosomal inversions prenatally diagnosed by second trimester amniocentesis
Frequency and clinical significance of chromosomal inversions prenatally diagnosed by second trimester amniocentesis
AbstractTo compare the frequency and clinical significance of familial and de novo chromosomal inversions during prenatal diagnosis. This was a retrospective study of inversions di...
Inversions and Evolution of the Human Genome
Inversions and Evolution of the Human Genome
AbstractInversions are a type of naturally occurring DNA mutation in which the sequence of the DNA is reversed, resulting in it being read in the opposite direction to the wild typ...
Resolving the insertion sites of polymorphic duplications reveals a
HERC2
haplotype under selection
Resolving the insertion sites of polymorphic duplications reveals a
HERC2
haplotype under selection
ABSTRACT
Polymorphic duplications in humans have been shown to contribute to phenotypic diversity. However, the evolutionary forces that maintain...
L26/O-028 Paired trophectoderm and ICM analysis redefines segmental aneuploidy and embryo usability in poor-prognosis IVF patients
L26/O-028 Paired trophectoderm and ICM analysis redefines segmental aneuploidy and embryo usability in poor-prognosis IVF patients
Abstract
Study question
What is the parental origin, timing, and distribution of segmental aneuploidy diagnosed as non-mo...
Duplications and retrogenes are numerous and widespread in modern canine genomic assemblies
Duplications and retrogenes are numerous and widespread in modern canine genomic assemblies
Abstract
Recent years have seen a dramatic increase in the number of canine genome assemblies available. Duplications are an important source of evolutionary novelt...
Recombination Rate Predicts Inversion Size in Diptera
Recombination Rate Predicts Inversion Size in Diptera
AbstractMost species of the Drosophila genus and other Diptera are polymorphic for paracentric inversions. A common observation is that successful inversions are of intermediate si...
Myc super-enhancer dynamics in T-ALL
Myc super-enhancer dynamics in T-ALL
Abstract
Mutations in non-coding regulatory elements are increasingly recognized as critical drivers of cancer developme...
Benchmarking and improving imputation approaches for recurrent inversions in the human genome
Benchmarking and improving imputation approaches for recurrent inversions in the human genome
Inversions are a type of structural variant that are involved in phenotypic differences among individuals. Due to certain features, such as the fact of usually implying no loss or ...

