Javascript must be enabled to continue!
The complexity landscape of viral genomes
View through CrossRef
Abstract
Background
Viruses are among the shortest yet highly abundant species that harbor minimal instructions to infect cells, adapt, multiply, and exist. However, with the current substantial availability of viral genome sequences, the scientific repertory lacks a complexity landscape that automatically enlights viral genomes’ organization, relation, and fundamental characteristics.
Results
This work provides a comprehensive landscape of the viral genome’s complexity (or quantity of information), identifying the most redundant and complex groups regarding their genome sequence while providing their distribution and characteristics at a large and local scale. Moreover, we identify and quantify inverted repeats abundance in viral genomes. For this purpose, we measure the sequence complexity of each available viral genome using data compression, demonstrating that adequate data compressors can efficiently quantify the complexity of viral genome sequences, including subsequences better represented by algorithmic sources (e.g., inverted repeats). Using a state-of-the-art genomic compressor on an extensive viral genomes database, we show that double-stranded DNA viruses are, on average, the most redundant viruses while single-stranded DNA viruses are the least. Contrarily, double-stranded RNA viruses show a lower redundancy relative to single-stranded RNA. Furthermore, we extend the ability of data compressors to quantify local complexity (or information content) in viral genomes using complexity profiles, unprecedently providing a direct complexity analysis of human herpesviruses. We also conceive a features-based classification methodology that can accurately distinguish viral genomes at different taxonomic levels without direct comparisons between sequences. This methodology combines data compression with simple measures such as GC-content percentage and sequence length, followed by machine learning classifiers.
Conclusions
This article presents methodologies and findings that are highly relevant for understanding the patterns of similarity and singularity between viral groups, opening new frontiers for studying viral genomes’ organization while depicting the complexity trends and classification components of these genomes at different taxonomic levels. The whole study is supported by an extensive website (https://asilab.github.io/canvas/) for comprehending the viral genome characterization using dynamic and interactive approaches.
Oxford University Press (OUP)
Title: The complexity landscape of viral genomes
Description:
Abstract
Background
Viruses are among the shortest yet highly abundant species that harbor minimal instructions to infect cells, adapt, multiply, and exist.
However, with the current substantial availability of viral genome sequences, the scientific repertory lacks a complexity landscape that automatically enlights viral genomes’ organization, relation, and fundamental characteristics.
Results
This work provides a comprehensive landscape of the viral genome’s complexity (or quantity of information), identifying the most redundant and complex groups regarding their genome sequence while providing their distribution and characteristics at a large and local scale.
Moreover, we identify and quantify inverted repeats abundance in viral genomes.
For this purpose, we measure the sequence complexity of each available viral genome using data compression, demonstrating that adequate data compressors can efficiently quantify the complexity of viral genome sequences, including subsequences better represented by algorithmic sources (e.
g.
, inverted repeats).
Using a state-of-the-art genomic compressor on an extensive viral genomes database, we show that double-stranded DNA viruses are, on average, the most redundant viruses while single-stranded DNA viruses are the least.
Contrarily, double-stranded RNA viruses show a lower redundancy relative to single-stranded RNA.
Furthermore, we extend the ability of data compressors to quantify local complexity (or information content) in viral genomes using complexity profiles, unprecedently providing a direct complexity analysis of human herpesviruses.
We also conceive a features-based classification methodology that can accurately distinguish viral genomes at different taxonomic levels without direct comparisons between sequences.
This methodology combines data compression with simple measures such as GC-content percentage and sequence length, followed by machine learning classifiers.
Conclusions
This article presents methodologies and findings that are highly relevant for understanding the patterns of similarity and singularity between viral groups, opening new frontiers for studying viral genomes’ organization while depicting the complexity trends and classification components of these genomes at different taxonomic levels.
The whole study is supported by an extensive website (https://asilab.
github.
io/canvas/) for comprehending the viral genome characterization using dynamic and interactive approaches.
Related Results
Viral Hijacking of Host RNA-Binding Proteins: Implications for Viral Replication and Pathogenesis
Viral Hijacking of Host RNA-Binding Proteins: Implications for Viral Replication and Pathogenesis
In the intricate dance between viruses and host cells, RNA-binding proteins (RBPs) serve as crucial orchestrators of gene expression and cellular processes. We will delve into the ...
Genomic characterization of the
C. tuberculostearicum
species complex, a ubiquitous member of the human skin microbiome
Genomic characterization of the
C. tuberculostearicum
species complex, a ubiquitous member of the human skin microbiome
ABSTRACT
Corynebacterium
is a predominant genus in the skin microbiome, yet its genetic diversity on skin is incompletely chara...
Statistique des comparaisons de génomes complets bactériens
Statistique des comparaisons de génomes complets bactériens
La génomique comparative est l'étude des relations structurales et fonctionnelles entre des génomes appartenant à différentes souches ou espèces. Cette discipline offre ainsi la po...
How chromosomal rearrangements shape genomes : a computational and mathematical study
How chromosomal rearrangements shape genomes : a computational and mathematical study
Comment les réarrangements chromosomiques façonnent les génomes : étude par modélisation et simulations
Les origines de la complexité des génomes, ainsi que les dét...
Triple lysine and nucleosome-binding motifs of the viral IE19 protein are required for human cytomegalovirus S-phase infections
Triple lysine and nucleosome-binding motifs of the viral IE19 protein are required for human cytomegalovirus S-phase infections
ABSTRACT
Herpesvirus genomes are maintained as extrachromosomal plasmids within the nuclei of infected cells. Some herpesviruses persist within ...
GIS-based landscape design research
GIS-based landscape design research
Landscape design research is important for cultivating spatial intelligence in landscape architecture. This study explores GIS (geographic information systems) as a tool for landsc...
vClean: assessing virus sequence contamination in viral genomes
vClean: assessing virus sequence contamination in viral genomes
Abstract
Recent advancements in viral metagenomics and single-virus genomics have improved our ability to obtain the draft genomes of environmental viruses. Howev...
Bioinformatics analysis and collection of protein post-translational modification sites in human viruses
Bioinformatics analysis and collection of protein post-translational modification sites in human viruses
AbstractIn viruses, post-translational modifications (PTMs) are essential for their life cycle. Recognizing viral PTMs is very important for better understanding the mechanism of v...

