Javascript must be enabled to continue!
Genome assembly composition of the String “ACGT” array: a review of data structure accuracy and performance challenges
View through CrossRef
Background
The development of sequencing technology increases the number of genomes being sequenced. However, obtaining a quality genome sequence remains a challenge in genome assembly by assembling a massive number of short strings (reads) with the presence of repetitive sequences (repeats). Computer algorithms for genome assembly construct the entire genome from reads in two approaches. The
de novo
approach concatenates the reads based on the exact match between their suffix-prefix (overlapping). Reference-guided approach orders the reads based on their offsets in a well-known reference genome (reads alignment). The presence of repeats extends the technical ambiguity, making the algorithm unable to distinguish the reads resulting in misassembly and affecting the assembly approach accuracy. On the other hand, the massive number of reads causes a big assembly performance challenge.
Method
The repeat identification method was introduced for misassembly by prior identification of repetitive sequences, creating a repeat knowledge base to reduce ambiguity during the assembly process, thus enhancing the accuracy of the assembled genome. Also, hybridization between assembly approaches resulted in a lower misassembly degree with the aid of the reference genome. The assembly performance is optimized through data structure indexing and parallelization. This article’s primary aim and contribution are to support the researchers through an extensive review to ease other researchers’ search for genome assembly studies. The study also, highlighted the most recent developments and limitations in genome assembly accuracy and performance optimization.
Results
Our findings show the limitations of the repeat identification methods available, which only allow to detect of specific lengths of the repeat, and may not perform well when various types of repeats are present in a genome. We also found that most of the hybrid assembly approaches, either starting with
de novo
or reference-guided, have some limitations in handling repetitive sequences as it is more computationally costly and time intensive. Although the hybrid approach was found to outperform individual assembly approaches, optimizing its performance remains a challenge. Also, the usage of parallelization in overlapping and reads alignment for genome assembly is yet to be fully implemented in the hybrid assembly approach.
Conclusion
We suggest combining multiple repeat identification methods to enhance the accuracy of identifying the repeats as an initial step to the hybrid assembly approach and combining genome indexing with parallelization for better optimization of its performance.
Title: Genome assembly composition of the String “ACGT” array: a review of data structure accuracy and performance challenges
Description:
Background
The development of sequencing technology increases the number of genomes being sequenced.
However, obtaining a quality genome sequence remains a challenge in genome assembly by assembling a massive number of short strings (reads) with the presence of repetitive sequences (repeats).
Computer algorithms for genome assembly construct the entire genome from reads in two approaches.
The
de novo
approach concatenates the reads based on the exact match between their suffix-prefix (overlapping).
Reference-guided approach orders the reads based on their offsets in a well-known reference genome (reads alignment).
The presence of repeats extends the technical ambiguity, making the algorithm unable to distinguish the reads resulting in misassembly and affecting the assembly approach accuracy.
On the other hand, the massive number of reads causes a big assembly performance challenge.
Method
The repeat identification method was introduced for misassembly by prior identification of repetitive sequences, creating a repeat knowledge base to reduce ambiguity during the assembly process, thus enhancing the accuracy of the assembled genome.
Also, hybridization between assembly approaches resulted in a lower misassembly degree with the aid of the reference genome.
The assembly performance is optimized through data structure indexing and parallelization.
This article’s primary aim and contribution are to support the researchers through an extensive review to ease other researchers’ search for genome assembly studies.
The study also, highlighted the most recent developments and limitations in genome assembly accuracy and performance optimization.
Results
Our findings show the limitations of the repeat identification methods available, which only allow to detect of specific lengths of the repeat, and may not perform well when various types of repeats are present in a genome.
We also found that most of the hybrid assembly approaches, either starting with
de novo
or reference-guided, have some limitations in handling repetitive sequences as it is more computationally costly and time intensive.
Although the hybrid approach was found to outperform individual assembly approaches, optimizing its performance remains a challenge.
Also, the usage of parallelization in overlapping and reads alignment for genome assembly is yet to be fully implemented in the hybrid assembly approach.
Conclusion
We suggest combining multiple repeat identification methods to enhance the accuracy of identifying the repeats as an initial step to the hybrid assembly approach and combining genome indexing with parallelization for better optimization of its performance.
Related Results
Parameterized Strings: Algorithms and Applications
Parameterized Strings: Algorithms and Applications
The parameterized string (p-string), a generalization of the traditional string, is composed of constant and parameter symbols. A parameterized match (p-match) exists between two p...
Evaluating the Science to Inform the Physical Activity Guidelines for Americans Midcourse Report
Evaluating the Science to Inform the Physical Activity Guidelines for Americans Midcourse Report
Abstract
The Physical Activity Guidelines for Americans (Guidelines) advises older adults to be as active as possible. Yet, despite the well documented benefits of physical activi...
Haptic-enabled virtual planning and assessment of product assembly
Haptic-enabled virtual planning and assessment of product assembly
Purpose
This study aims to present a new haptic-enabled virtual assembly system for the automatic generation and objective assessment of assembly plans. The syste...
Sequencing for Super Seeds
Sequencing for Super Seeds
The year 2000 marked two important milestones in the science of whole genome sequencing (WGS): completion of the draft human genome and that of Arabidopsis, a model plant commonly ...
Combination of Synchronised Assembly Precedence Matrices
Combination of Synchronised Assembly Precedence Matrices
Abstract
The planning of assembly lines has always been a complex task, as a large number of factors that influence and disrupt production must be taken into account. Uncer...
Axial Excitation Tool String Modelling
Axial Excitation Tool String Modelling
Current types of axial excitation tool have been shown to produce beneficial results — in terms of load transfer to the bit, general reductions in string friction and reductions in...
Numerical analysis and Experimental Investigation of Lateral vibration on Drill String under Axial Load Constrained with Horizontal Pipe
Numerical analysis and Experimental Investigation of Lateral vibration on Drill String under Axial Load Constrained with Horizontal Pipe
Abstract
Horizontal well technology is an important means to improve drilling efficiency and oil and gas production, but it is easy to generate the lateral vibration...
18-3/4 in. FullBore Wellhead System
18-3/4 in. FullBore Wellhead System
Abstract
This paper describes the development of full-bore wellheads, a new 18-3/4 in.15,000 psi W.P. system, from conception to field installation. The wellheadw...

