Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Assessing text embedding models to assign UniProt classes to scientific literature

View through CrossRef
Advances in biomedical sciences are increasingly dependent on knowledge encoded in curated biomedical databases. In particular, the Universal Protein Resource (UniProt) provides the scientific community with a comprehensive and accurately annotated protein sequence knowledgebase. The organization of UniProt protein entries and their associated publications in different topics, such as expression, function and interaction, helps users to find the information of interest in the knowledgebase. The current topic classification approach for computationally mapped bibliography, based solely on underlying sources, is limited. We investigate the use of (semi-) automated classifiers for helping UniProt to classify the scientific biomedical literature according to 11 topics. Publications annotated by UniProt curators are labeled using one or more of the classes above. As such, this is a multi-class multi-label classification problem. Our algorithm works as follows: given a text passage, such as an article abstract, it provides a ranked list of classes and the probabilities that the passage belongs to the available classes. We investigate a text embedding model, Doc2Vec, to compute document similarity and compare several machine learning methods, such as Naïve Bayes, kNN, Logistic regression, and MLP, for assigning class similarities. We use a collection of 100,000 documents classified by the UniProt team to train (99,000) and test (1,000) our algorithms. The validation of the classifier parameters was performed using 5% of the training collection. We compare the text embedded models with a baseline model based on bag of words, where divergence from randomness is used to compute document similarity and kNN to assign classes. In general, our algorithms achieved a high classification precision. The baseline model achieved a mean average precision (MAP) of 0.8270. Apart from the Naïve Bayes model (MAP 0.7163), all the models based on the text embedding approach outperformed the baseline method. Logistic regression achieved a MAP of 0.8376 (p=.32), MLP achieved a MAP of 0.8413 (p=.18), and kNN achieved a MAP of 0.8485 (p=.04). We believe that such classifiers could help improving the productivity in and reduce the cost of certain biocuration tasks. The next steps will be the application of the methodology to unclassified documents and to assess its effectiveness in assisting curators judge the relevance of articles for further curation.
Title: Assessing text embedding models to assign UniProt classes to scientific literature
Description:
Advances in biomedical sciences are increasingly dependent on knowledge encoded in curated biomedical databases.
In particular, the Universal Protein Resource (UniProt) provides the scientific community with a comprehensive and accurately annotated protein sequence knowledgebase.
The organization of UniProt protein entries and their associated publications in different topics, such as expression, function and interaction, helps users to find the information of interest in the knowledgebase.
The current topic classification approach for computationally mapped bibliography, based solely on underlying sources, is limited.
We investigate the use of (semi-) automated classifiers for helping UniProt to classify the scientific biomedical literature according to 11 topics.
Publications annotated by UniProt curators are labeled using one or more of the classes above.
As such, this is a multi-class multi-label classification problem.
Our algorithm works as follows: given a text passage, such as an article abstract, it provides a ranked list of classes and the probabilities that the passage belongs to the available classes.
We investigate a text embedding model, Doc2Vec, to compute document similarity and compare several machine learning methods, such as Naïve Bayes, kNN, Logistic regression, and MLP, for assigning class similarities.
We use a collection of 100,000 documents classified by the UniProt team to train (99,000) and test (1,000) our algorithms.
The validation of the classifier parameters was performed using 5% of the training collection.
We compare the text embedded models with a baseline model based on bag of words, where divergence from randomness is used to compute document similarity and kNN to assign classes.
In general, our algorithms achieved a high classification precision.
The baseline model achieved a mean average precision (MAP) of 0.
8270.
Apart from the Naïve Bayes model (MAP 0.
7163), all the models based on the text embedding approach outperformed the baseline method.
Logistic regression achieved a MAP of 0.
8376 (p=.
32), MLP achieved a MAP of 0.
8413 (p=.
18), and kNN achieved a MAP of 0.
8485 (p=.
04).
We believe that such classifiers could help improving the productivity in and reduce the cost of certain biocuration tasks.
The next steps will be the application of the methodology to unclassified documents and to assess its effectiveness in assisting curators judge the relevance of articles for further curation.

Related Results

UniProt Tools: BLAST, Align, Peptide Search, and ID Mapping
UniProt Tools: BLAST, Align, Peptide Search, and ID Mapping
AbstractThe Universal Protein Resource (UniProt) is a comprehensive resource for protein sequence and annotation data (UniProt Consortium, 2023). The UniProt website receives about...
Sleep Habits and Occurrence of Lowback Pain among Craftsmen
Sleep Habits and Occurrence of Lowback Pain among Craftsmen
<span style="color: #000000; font-family: Verdana, Arial, Helvetica, sans-serif; font-size: 10px; font-style: normal; font-variant-ligatures: normal; font-variant-caps: normal; ...
Sleep Habits and Occurrence of Lowback Pain among Craftsmen
Sleep Habits and Occurrence of Lowback Pain among Craftsmen
<span style="color: #000000; font-family: Verdana, Arial, Helvetica, sans-serif; font-size: 10px; font-style: normal; font-variant-ligatures: normal; font-variant-caps: normal; ...
Searching and Navigating UniProt Databases
Searching and Navigating UniProt Databases
AbstractThe Universal Protein Resource (UniProt) is a comprehensive resource for protein sequence and annotation data. The UniProt website receives about 800,000 unique visitors pe...
Searching and Navigating UniProt Databases
Searching and Navigating UniProt Databases
AbstractThe Universal Protein Resource (UniProt) is a comprehensive resource for protein sequence and annotation data. The UniProt Web site receives ∼400,000 unique visitors per mo...
Bounds on the sum of broadcast domination number and strong metric dimension of graphs
Bounds on the sum of broadcast domination number and strong metric dimension of graphs
Let [Formula: see text] be a connected graph of order at least two with vertex set [Formula: see text]. For [Formula: see text], let [Formula: see text] denote the length of an [Fo...
Evaluating the Science to Inform the Physical Activity Guidelines for Americans Midcourse Report
Evaluating the Science to Inform the Physical Activity Guidelines for Americans Midcourse Report
Abstract The Physical Activity Guidelines for Americans (Guidelines) advises older adults to be as active as possible. Yet, despite the well documented benefits of physical a...
ANALYSIS OF READING MATERIALS IN TEXTBOOK FOR GRADE XI SENIOR HIGH SCHOOL
ANALYSIS OF READING MATERIALS IN TEXTBOOK FOR GRADE XI SENIOR HIGH SCHOOL
This study aims to find out the GI and LD level, the text which has the highest GI and LD and what make the text has the highest GI and LD of Advanced Learning English 2 textbook. ...

Back to Top