Javascript must be enabled to continue!
Automated annotation in UniProt
View through CrossRef
UniProt is a high quality, comprehensive protein resource in which the core activity is the expert review and annotation of proteins where the function has been experimentally investigated. At the same time, the UniProt database contains large numbers of proteins which are predicted to exist from gene models, but which do not have associated experimental evidence indicating their function. UniProt commits significant resources to developing computational methods for functional annotation of these predicted proteins based on the data in entries that have gone through the expert review process.
We will describe the two main automated annotation systems currently in use. First, UniRule, which is an established UniProt system in which curators manually develop rules for annotation. Second, ARBA (Association-Rule-Based Annotator), which is a multi-class learning system which uses rule mining techniques to generate concise annotation models. ARBA employs a data exclusion algorithm that censors data not suitable for computational annotation, and generates human-readable rules for each UniProt release. As part of our interest in engaging with the machine learning community, we will also introduce the contribution of ProtNLM (Protein Natural Language Model), from Google Research, which annotates proteins which have "uncharacterised" names.
We will also introduce UniFIRE, an open source software that enables researchers to annotate their own protein dataset by using the above mentioned annotation systems. In order to provide an easy and straightforward way to download and set up this tool we have containerised UniFIRE together with all its dependencies and the latest set of UniRule and ARBA rules. In this webinar, we will show how to create functional predictions for protein sequences by using this container image.
Title: Automated annotation in UniProt
Description:
UniProt is a high quality, comprehensive protein resource in which the core activity is the expert review and annotation of proteins where the function has been experimentally investigated.
At the same time, the UniProt database contains large numbers of proteins which are predicted to exist from gene models, but which do not have associated experimental evidence indicating their function.
UniProt commits significant resources to developing computational methods for functional annotation of these predicted proteins based on the data in entries that have gone through the expert review process.
We will describe the two main automated annotation systems currently in use.
First, UniRule, which is an established UniProt system in which curators manually develop rules for annotation.
Second, ARBA (Association-Rule-Based Annotator), which is a multi-class learning system which uses rule mining techniques to generate concise annotation models.
ARBA employs a data exclusion algorithm that censors data not suitable for computational annotation, and generates human-readable rules for each UniProt release.
As part of our interest in engaging with the machine learning community, we will also introduce the contribution of ProtNLM (Protein Natural Language Model), from Google Research, which annotates proteins which have "uncharacterised" names.
We will also introduce UniFIRE, an open source software that enables researchers to annotate their own protein dataset by using the above mentioned annotation systems.
In order to provide an easy and straightforward way to download and set up this tool we have containerised UniFIRE together with all its dependencies and the latest set of UniRule and ARBA rules.
In this webinar, we will show how to create functional predictions for protein sequences by using this container image.
Related Results
UniProt Tools: BLAST, Align, Peptide Search, and ID Mapping
UniProt Tools: BLAST, Align, Peptide Search, and ID Mapping
AbstractThe Universal Protein Resource (UniProt) is a comprehensive resource for protein sequence and annotation data (UniProt Consortium, 2023). The UniProt website receives about...
Principes et outils pour l’annotation des corpus
Principes et outils pour l’annotation des corpus
La linguistique de corpus, c’est à dire les recherches sur le langage portant sur un matériel linguistique écrit ou oral recueilli et conservé, s’est considérablement développée au...
Searching and Navigating UniProt Databases
Searching and Navigating UniProt Databases
AbstractThe Universal Protein Resource (UniProt) is a comprehensive resource for protein sequence and annotation data. The UniProt website receives about 800,000 unique visitors pe...
Searching and Navigating UniProt Databases
Searching and Navigating UniProt Databases
AbstractThe Universal Protein Resource (UniProt) is a comprehensive resource for protein sequence and annotation data. The UniProt Web site receives ∼400,000 unique visitors per mo...
UniProt Knowledgebase: a hub of integrated data
UniProt Knowledgebase: a hub of integrated data
AbstractData integration plays an increasingly important role in bringing together the large amounts of diverse information spread across disparate resources and presenting a compr...
UniProt Tools
UniProt Tools
AbstractThe Universal Protein Resource (UniProt) is a comprehensive resource for protein sequence and annotation data (UniProt Consortium, 2015). The UniProt Web site receives ∼400...
UniRule - Automatic Annotation In UniProtKB
UniRule - Automatic Annotation In UniProtKB
AbstractThe UniProt KnowledgeBase (UniProtKB) provides a stable, comprehensive, freely accessible, centralized resource on protein sequences and functional annotation. UniProtKB co...
Assessing text embedding models to assign UniProt classes to scientific literature
Assessing text embedding models to assign UniProt classes to scientific literature
Advances in biomedical sciences are increasingly dependent on knowledge encoded in curated biomedical databases. In particular, the Universal Protein Resource (UniProt) provides th...

