Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Protein Embeddings and Local Alignments

View through CrossRef
Abstract Background The advent of protein embeddings has revolutionized bioinformatics by providing contextual representations that capture functional and evolutionary patterns. They have become, alongside sequence alignments, the cornerstone of bioinformatics. While embeddings cannot replace alignments, they can greatly help improving their quality. Our goal is replacing the BLOSUM matrices, the decades-old standard scoring system for protein alignment, with an embedding-based scoring method. Results We introduce a new scoring function and algorithm for local alignment of protein sequences, we offer a new comprehensive framework for evaluating local alignments. The score between two residues is given by the cosine similarity of their Ankh-embedding vectors and the algorithm uses dynamic programming with affine penalty. For the evaluation, we built multiple datasets, using both natural and inserted sequences, from the Conserved Domain Database, BAL-iBASE, and GPCRdb, designed a new algorithm for local alignment extraction, localization and quality evaluation, and employed five distance metrics to evaluate the similarity with the true alignment. We performed nearly one and a half million tests to compare the new algorithm with the best BLOSUM matrices, specialized GPCRtm matrices, and top programs, such as PEbA, ProtT5-score, DEDAL, vcMSA and pLM-BLAST. Regarding the protein embedding models, Ankh not only surpasses the best combination of ProtT5 and ESM2, but appears to better understand the “language” of proteins, as it behaves much better on natural sequences compared to artificial ones obtained by inserting domains in random protein sequences. Also, while ProtT5 and ESM2 combine to produce better results, Ankh does not combine well with other embeddings. Conclusions The new Ankh-score-based program is vastly superior to the BLOSUM matrices and clearly superior to all existing methods. New light shed on the protein embeddings can guide future improvements. In order to facilitate the use of the new method and protocol, they are freely available as a web server at e-score.csd.uwo.ca and as source code at github.com/lucian-ilie/E-score .
Title: Protein Embeddings and Local Alignments
Description:
Abstract Background The advent of protein embeddings has revolutionized bioinformatics by providing contextual representations that capture functional and evolutionary patterns.
They have become, alongside sequence alignments, the cornerstone of bioinformatics.
While embeddings cannot replace alignments, they can greatly help improving their quality.
Our goal is replacing the BLOSUM matrices, the decades-old standard scoring system for protein alignment, with an embedding-based scoring method.
Results We introduce a new scoring function and algorithm for local alignment of protein sequences, we offer a new comprehensive framework for evaluating local alignments.
The score between two residues is given by the cosine similarity of their Ankh-embedding vectors and the algorithm uses dynamic programming with affine penalty.
For the evaluation, we built multiple datasets, using both natural and inserted sequences, from the Conserved Domain Database, BAL-iBASE, and GPCRdb, designed a new algorithm for local alignment extraction, localization and quality evaluation, and employed five distance metrics to evaluate the similarity with the true alignment.
We performed nearly one and a half million tests to compare the new algorithm with the best BLOSUM matrices, specialized GPCRtm matrices, and top programs, such as PEbA, ProtT5-score, DEDAL, vcMSA and pLM-BLAST.
Regarding the protein embedding models, Ankh not only surpasses the best combination of ProtT5 and ESM2, but appears to better understand the “language” of proteins, as it behaves much better on natural sequences compared to artificial ones obtained by inserting domains in random protein sequences.
Also, while ProtT5 and ESM2 combine to produce better results, Ankh does not combine well with other embeddings.
Conclusions The new Ankh-score-based program is vastly superior to the BLOSUM matrices and clearly superior to all existing methods.
New light shed on the protein embeddings can guide future improvements.
In order to facilitate the use of the new method and protocol, they are freely available as a web server at e-score.
csd.
uwo.
ca and as source code at github.
com/lucian-ilie/E-score .

Related Results

COFFEE: an objective function for multiple sequence alignments.
COFFEE: an objective function for multiple sequence alignments.
Abstract MOTIVATION: In order to increase the accuracy of multiple sequence alignments, we designed a new strategy for optimizing multiple sequence alignments by gen...
Frequency of Common Chromosomal Abnormalities in Patients with Idiopathic Acquired Aplastic Anemia
Frequency of Common Chromosomal Abnormalities in Patients with Idiopathic Acquired Aplastic Anemia
Objective: To determine the frequency of common chromosomal aberrations in local population idiopathic determine the frequency of common chromosomal aberrations in local population...
SPACE: STRING proteins as complementary embeddings
SPACE: STRING proteins as complementary embeddings
Representation learning has revolutionized sequence-based prediction of protein function and subcellular localization. Protein networks are an important source of information compl...
Protein Embedding based Alignment
Protein Embedding based Alignment
Despite of the many progresses with alignment algorithms, aligning divergent protein sequences including those sharing less than 20-35% pairwise identity (so called “twilight zone”...
Endothelial Protein C Receptor
Endothelial Protein C Receptor
IntroductionThe protein C anticoagulant pathway plays a critical role in the negative regulation of the blood clotting response. The pathway is triggered by thrombin, which allows ...
Multiple Alignments of Data Objects and Generalized Center Star Algorithm
Multiple Alignments of Data Objects and Generalized Center Star Algorithm
Multiple alignments of strings have been extensively studied as an effective tool to study string-type data such as DNA. In this paper, we generalize the notion of multiple alignme...
Exploring Word Embeddings for Text Classification: A Comparative Analysis
Exploring Word Embeddings for Text Classification: A Comparative Analysis
For language tasks like text classification and sequence labeling, word embeddings are essential for providing input characteristics in deep models. There have been many word embed...
Truck stability on different types of horizontal curves combined with vertical alignments
Truck stability on different types of horizontal curves combined with vertical alignments
The combination of horizontal curves with vertical alignments is commonly used in different classifications of highways; either on highway mainstream or on highway interchange ramp...

Back to Top