Javascript must be enabled to continue!
Gene Unprediction with Spurio: A tool to identify spurious protein sequences
View through CrossRef
We now have access to the sequences of tens of millions of proteins. These protein sequences are essential for modern molecular biology and computational biology. The vast majority of protein sequences are derived from gene prediction tools and have no experimental supporting evidence for their translation. Despite the increasing accuracy of gene prediction tools there likely exists a large number of spurious protein predictions in the sequence databases. We have developed the Spurio tool to help identify spurious protein predictions in prokaryotes. Spurio searches the query protein sequence against a prokaryotic nucleotide database using tblastn and identifies homologous sequences. The tblastn matches are used to score the query sequence’s likelihood of being a spurious protein prediction using a Gaussian process model. The most informative feature is the appearance of stop codons within the presumed translation of homologous DNA sequences. Benchmarking shows that the Spurio tool is able to distinguish spurious from true proteins. However, transposon proteins are prone to be predicted as spurious because of the frequency of degraded homologs found in the DNA sequence databases. Our initial experiments suggest that less than 1% of the proteins in the UniProtKB sequence database are likely to be spurious and that Spurio is able to identify over 60 times more spurious proteins than the AntiFam resource.
The Spurio software and source code is available under an MIT license at the following URL:
https://bitbucket.org/bateman-group/spurio
Title: Gene Unprediction with Spurio: A tool to identify spurious protein sequences
Description:
We now have access to the sequences of tens of millions of proteins.
These protein sequences are essential for modern molecular biology and computational biology.
The vast majority of protein sequences are derived from gene prediction tools and have no experimental supporting evidence for their translation.
Despite the increasing accuracy of gene prediction tools there likely exists a large number of spurious protein predictions in the sequence databases.
We have developed the Spurio tool to help identify spurious protein predictions in prokaryotes.
Spurio searches the query protein sequence against a prokaryotic nucleotide database using tblastn and identifies homologous sequences.
The tblastn matches are used to score the query sequence’s likelihood of being a spurious protein prediction using a Gaussian process model.
The most informative feature is the appearance of stop codons within the presumed translation of homologous DNA sequences.
Benchmarking shows that the Spurio tool is able to distinguish spurious from true proteins.
However, transposon proteins are prone to be predicted as spurious because of the frequency of degraded homologs found in the DNA sequence databases.
Our initial experiments suggest that less than 1% of the proteins in the UniProtKB sequence database are likely to be spurious and that Spurio is able to identify over 60 times more spurious proteins than the AntiFam resource.
The Spurio software and source code is available under an MIT license at the following URL:
https://bitbucket.
org/bateman-group/spurio.
Related Results
Optimising tool wear and workpiece condition monitoring via cyber-physical systems for smart manufacturing
Optimising tool wear and workpiece condition monitoring via cyber-physical systems for smart manufacturing
Smart manufacturing has been developed since the introduction of Industry 4.0. It consists of resource sharing and networking, predictive engineering, and material and data analyti...
Endothelial Protein C Receptor
Endothelial Protein C Receptor
IntroductionThe protein C anticoagulant pathway plays a critical role in the negative regulation of the blood clotting response. The pathway is triggered by thrombin, which allows ...
Expression and polymorphism of genes in gallstones
Expression and polymorphism of genes in gallstones
ABSTRACT
Through the method of clinical case control study, to explore the expression and genetic polymorphism of KLF14 gene (rs4731702 and rs972283) and SR-B1 gene...
Phylogenetic Classification of Feline Immunodeficiency Virus
Phylogenetic Classification of Feline Immunodeficiency Virus
Background: The feline immunodeficiency virus (FIV) is responsible for a retroviral disease that affects domestic and wild cats worldwide, causing Feline Acquired Immunodeficiency ...
TINGKAT PROTEIN DAN LISIN DALAM RANSUM TERHADAP EFISIENSI LISIN DAN PROTEIN NETTO PADA AYAM KAMPUNG UMUR 12 MINGGU
TINGKAT PROTEIN DAN LISIN DALAM RANSUM TERHADAP EFISIENSI LISIN DAN PROTEIN NETTO PADA AYAM KAMPUNG UMUR 12 MINGGU
Penelitian yang dilakukan ini dalam mencari pengaruh tingkat protein dan lisin terhadap efisiensi lisin dan penggunaan protein netto pada ayam kampung yang diperlihara sampai umur ...
Robot tool use: A survey
Robot tool use: A survey
Using human tools can significantly benefit robots in many application domains. Such ability would allow robots to solve problems that they were unable to without tools. However, r...
Translational regulation of the human cytomegalovirus pp28 (UL99) late gene
Translational regulation of the human cytomegalovirus pp28 (UL99) late gene
The pp28 (UL99) gene of human cytomegalovirus is expressed as a true late gene, in that DNA synthesis is absolutely required for mRNA expression. Our previous studies demonstrated ...
Charting Peptide Shared Sequences Between ‘Diabetes-Viruses’ and Human Pancreatic Proteins, Their Structural and Autoimmune Implications
Charting Peptide Shared Sequences Between ‘Diabetes-Viruses’ and Human Pancreatic Proteins, Their Structural and Autoimmune Implications
Diabetes mellitus (DM) is a metabolic syndrome characterized by hyperglycaemia, polydipsia, polyuria, and weight loss, among others. The pathophysiology for the disorders is comple...

