Javascript must be enabled to continue!
LJA: Assembling Long and Accurate Reads Using Multiplex de Bruijn Graphs
View through CrossRef
Abstract
Although most existing genome assemblers are based on the de Bruijn graphs, it remains unclear how to construct these graphs for large genomes and large
k
-mer sizes. This algorithmic challenge has become particularly important with the emergence of long high-fidelity (HiFi) reads that were recently utilized to generate a semi-manual telomere-to-telomere assembly of the human genome and to get a glimpse into biomedically important regions that evaded all previous attempts to sequence them. To enable automated assemblies of long and accurate reads, we developed a fast LJA algorithm that reduces the error rate in these reads by three orders of magnitude (making them nearly error-free) and constructs the de Bruijn graph for large genomes and large
k
-mer sizes. Since the de Bruijn graph constructed for a fixed
k
-mer size is typically either too tangled or too fragmented, LJA uses a new concept of a multiplex de Bruijn graph with varying
k
-mer sizes. We demonstrate that LJA improves on the state-of-the-art assemblers with respect to both accuracy and contiguity and enables automated telomere-to-telomere assemblies of entire human chromosomes.
Title: LJA: Assembling Long and Accurate Reads Using Multiplex de Bruijn Graphs
Description:
Abstract
Although most existing genome assemblers are based on the de Bruijn graphs, it remains unclear how to construct these graphs for large genomes and large
k
-mer sizes.
This algorithmic challenge has become particularly important with the emergence of long high-fidelity (HiFi) reads that were recently utilized to generate a semi-manual telomere-to-telomere assembly of the human genome and to get a glimpse into biomedically important regions that evaded all previous attempts to sequence them.
To enable automated assemblies of long and accurate reads, we developed a fast LJA algorithm that reduces the error rate in these reads by three orders of magnitude (making them nearly error-free) and constructs the de Bruijn graph for large genomes and large
k
-mer sizes.
Since the de Bruijn graph constructed for a fixed
k
-mer size is typically either too tangled or too fragmented, LJA uses a new concept of a multiplex de Bruijn graph with varying
k
-mer sizes.
We demonstrate that LJA improves on the state-of-the-art assemblers with respect to both accuracy and contiguity and enables automated telomere-to-telomere assemblies of entire human chromosomes.
Related Results
MBG: Minimizer-based Sparse de Bruijn Graph Construction
MBG: Minimizer-based Sparse de Bruijn Graph Construction
Motivation
De Bruijn graphs can be constructed from short reads efficiently and have been used for many purposes. Traditionally long read sequencing technologies ...
Multi de Bruijn Sequences and the Cross-Join Method
Multi de Bruijn Sequences and the Cross-Join Method
We show a method to construct binary multi de Bruijn sequences using the cross-join method. We extend the proof given by Alhakim for ordinary de Bruijn sequences to the case of mul...
Disentangled Long-Read De Bruijn Graphs via Optical Maps
Disentangled Long-Read De Bruijn Graphs via Optical Maps
Abstract
Pacific Biosciences (PacBio), the main third generation sequencing technology can produce scalable, high-throughput, unprecedented sequencing results throu...
Comparison of Three Molecular Methods for the Detection and Speciation of Five Human Plasmodium Species
Comparison of Three Molecular Methods for the Detection and Speciation of Five Human Plasmodium Species
In this study, three molecular assays (real-time multiplex polymerase chain reaction [PCR], merozoite surface antigen gene [MSP]-multiplex PCR, and the PlasmoNex Multiplex PCR Kit)...
Hybrid correction of highly noisy Oxford Nanopore long reads using a variable-order de Bruijn graph
Hybrid correction of highly noisy Oxford Nanopore long reads using a variable-order de Bruijn graph
Abstract
Motivation
The recent rise of long read sequencing technologies such as Pacific Biosciences and Oxford Nanopore allows...
Weakly Modular Graphs and Nonpositive Curvature
Weakly Modular Graphs and Nonpositive Curvature
This article investigates structural, geometrical, and topological characterizations and properties of weakly modular graphs and of cell complexes derived from them. The unifying t...
Phased Multi de Bruijn Sequences
Phased Multi de Bruijn Sequences
We introduce phased multi de Bruijn sequences, a generalization of de Bruijn sequences. A phased string is a string whose positions sequentially rotate through several alphabets; e...
Building Large Updatable Colored de Bruijn Graphs via Merging
Building Large Updatable Colored de Bruijn Graphs via Merging
MOTIVATION: There exists several massive genomic and metagenomic data collection efforts, including GenomeTrakr and MetaSub, which are routinely updated with new data. To analyze s...

