Javascript must be enabled to continue!
Removing Contaminants from Metagenomic Databases
View through CrossRef
Abstract
Metagenomic sequencing of patient samples is a very promising method for the diagnosis of human infections. Sequencing has the ability to capture all the DNA or RNA from pathogenic organisms in a human sample. However, complete and accurate characterization of the sequence, including identification of any pathogens, depends on the availability and quality of genomes for comparison. Thousands of genomes are now available, and as these numbers grow, the power of metagenomic sequencing for diagnosis should increase. However, recent studies have exposed the presence of contamination in published genomes, which when used for diagnosis increases the risk of falsely identifying the wrong pathogen.
To address this problem, we have developed a bioinformatics system for eliminating contamination as well as low-complexity genomic sequences in the draft genomes of eukaryotic pathogens. We applied this software to identify and remove human, bacterial, archaeal, and viral sequences present in a comprehensive database of all sequenced eukaryotic pathogen genomes. We also removed low-complexity genomic sequences, another source of false positives. Using this pipeline, we have produced a database of “clean” eukaryotic pathogen genomes for use with bioinformatics classification and analysis tools. We demonstrate that when attempting to find eukaryotic pathogens in metagenomic samples, the new database provides better sensitivity than one using the original genomes while offering a dramatic reduction in false positives.
Title: Removing Contaminants from Metagenomic Databases
Description:
Abstract
Metagenomic sequencing of patient samples is a very promising method for the diagnosis of human infections.
Sequencing has the ability to capture all the DNA or RNA from pathogenic organisms in a human sample.
However, complete and accurate characterization of the sequence, including identification of any pathogens, depends on the availability and quality of genomes for comparison.
Thousands of genomes are now available, and as these numbers grow, the power of metagenomic sequencing for diagnosis should increase.
However, recent studies have exposed the presence of contamination in published genomes, which when used for diagnosis increases the risk of falsely identifying the wrong pathogen.
To address this problem, we have developed a bioinformatics system for eliminating contamination as well as low-complexity genomic sequences in the draft genomes of eukaryotic pathogens.
We applied this software to identify and remove human, bacterial, archaeal, and viral sequences present in a comprehensive database of all sequenced eukaryotic pathogen genomes.
We also removed low-complexity genomic sequences, another source of false positives.
Using this pipeline, we have produced a database of “clean” eukaryotic pathogen genomes for use with bioinformatics classification and analysis tools.
We demonstrate that when attempting to find eukaryotic pathogens in metagenomic samples, the new database provides better sensitivity than one using the original genomes while offering a dramatic reduction in false positives.
Related Results
Metagenomic Thermometer
Metagenomic Thermometer
Abstract
Various microorganisms exist in environments, and each of which has an optimal growth temperature (OGT). The relationship between genomi...
CAIM: Coverage-based Analysis for Identification of Microbiome
CAIM: Coverage-based Analysis for Identification of Microbiome
ABSTRACT
Accurate taxonomic profiling of microbial taxa in a metagenomic sample is vital to gain insights into microbial ecology. Recent advancements in sequencing ...
Devenir et processus de transfert des résidus pharmaceutiques et pesticides en amendement et compostage.
Devenir et processus de transfert des résidus pharmaceutiques et pesticides en amendement et compostage.
En France, le traitement de plus de 5 Gm3/an d’eaux résiduaires urbaines entraine une production de boues d’épuration dépassant 1.1 M de tonnes.an-1. Plus de 70% des boues produite...
Metagenomic Thermometer
Metagenomic Thermometer
Abstract
Various microorganisms exist in environments, and each of them has its optimal growth temperature (OGT). The relationship between genomic information and...
metaJAM: a Nextflow integrated metagenomic workflow for sedimentary ancient DNA
metaJAM: a Nextflow integrated metagenomic workflow for sedimentary ancient DNA
Abstract
The application of metagenomics in ancient DNA (aDNA) research is rapidly expanding, driven in particular by advances in sedimentary aDNA research and sequ...
LMAS: evaluating metagenomic short de novo assembly methods through defined communities
LMAS: evaluating metagenomic short de novo assembly methods through defined communities
Abstract
Background
The de novo assembly of raw sequence data is key in metagenomic analysis. It allows recovering draft genomes...
Can we clean up the earth?
Can we clean up the earth?
Introduction: Contamination causes undue risks to society, ecosystems, water and soil resources, and threatens the viability of many industries1,2. As well as affecting soil, surfa...
Electrocoagulation for the Treatment of Oil Sands Tailings Water
Electrocoagulation for the Treatment of Oil Sands Tailings Water
Title: Electrocoagulation for the treatment of oil sands tailings water
The Canadian oil sands industry has been receiving widespread International criticism on env...

