Javascript must be enabled to continue!
FoldHSphere: deep hyperspherical embeddings for protein fold recognition
View through CrossRef
Abstract
Background
Current state-of-the-art deep learning approaches for protein fold recognition learn protein embeddings that improve prediction performance at the fold level. However, there still exists aperformance gap at the fold level and the (relatively easier) family level, suggesting that it might be possible to learn an embedding space that better represents the protein folds.
Results
In this paper, we propose the FoldHSphere method to learn a better fold embedding space through a two-stage training procedure. We first obtain prototype vectors for each fold class that are maximally separated in hyperspherical space. We then train a neural network by minimizing the angular large margin cosine loss to learn protein embeddings clustered around the corresponding hyperspherical fold prototypes. Our network architectures, ResCNN-GRU and ResCNN-BGRU, process the input protein sequences by applying several residual-convolutional blocks followed by a gated recurrent unit-based recurrent layer. Evaluation results on the LINDAHL dataset indicate that the use of our hyperspherical embeddings effectively bridges the performance gap at the family and fold levels. Furthermore, our FoldHSpherePro ensemble method yields an accuracy of 81.3% at the fold level, outperforming all the state-of-the-art methods.
Conclusions
Our methodology is efficient in learning discriminative and fold-representative embeddings for the protein domains. The proposed hyperspherical embeddings are effective at identifying the protein fold class by pairwise comparison, even when amino acid sequence similarities are low.
Springer Science and Business Media LLC
Title: FoldHSphere: deep hyperspherical embeddings for protein fold recognition
Description:
Abstract
Background
Current state-of-the-art deep learning approaches for protein fold recognition learn protein embeddings that improve prediction performance at the fold level.
However, there still exists aperformance gap at the fold level and the (relatively easier) family level, suggesting that it might be possible to learn an embedding space that better represents the protein folds.
Results
In this paper, we propose the FoldHSphere method to learn a better fold embedding space through a two-stage training procedure.
We first obtain prototype vectors for each fold class that are maximally separated in hyperspherical space.
We then train a neural network by minimizing the angular large margin cosine loss to learn protein embeddings clustered around the corresponding hyperspherical fold prototypes.
Our network architectures, ResCNN-GRU and ResCNN-BGRU, process the input protein sequences by applying several residual-convolutional blocks followed by a gated recurrent unit-based recurrent layer.
Evaluation results on the LINDAHL dataset indicate that the use of our hyperspherical embeddings effectively bridges the performance gap at the family and fold levels.
Furthermore, our FoldHSpherePro ensemble method yields an accuracy of 81.
3% at the fold level, outperforming all the state-of-the-art methods.
Conclusions
Our methodology is efficient in learning discriminative and fold-representative embeddings for the protein domains.
The proposed hyperspherical embeddings are effective at identifying the protein fold class by pairwise comparison, even when amino acid sequence similarities are low.
Related Results
Exploiting word embeddings for modeling bilexical relations
Exploiting word embeddings for modeling bilexical relations
There has been an exponential surge of text data in the recent years. As a consequence, unsupervised methods that make use of this data have been steadily growing in the field of n...
7
th
International Symposium on Enabling Technologies for Life Sciences (ETP)
7
th
International Symposium on Enabling Technologies for Life Sciences (ETP)
The seventh in the series of ETP Symposia (see
Rapid Communications in Mass Spectrometry
2012,
26
, ...
An analysis of protein language model embeddings for fold prediction
An analysis of protein language model embeddings for fold prediction
Abstract
The identification of the protein fold class is a challenging problem in structural biology. Recent computational methods for fold prediction leverage de...
An Analysis of Protein Language Model Embeddings for Fold Prediction
An Analysis of Protein Language Model Embeddings for Fold Prediction
Abstract
The identification of the protein fold class is a challenging problem in structural biology. Recent computational methods for fold prediction leverage deep...
SPACE: STRING proteins as complementary embeddings
SPACE: STRING proteins as complementary embeddings
Abstract
Motivation
Representation learning has revolutionized sequence-based prediction of protein function and subcellu...
SPACE: STRING proteins as complementary embeddings
SPACE: STRING proteins as complementary embeddings
Representation learning has revolutionized sequence-based prediction of protein function and subcellular localization. Protein networks are an important source of information compl...
Hyperspherical Mixed Prototypes Networks
Hyperspherical Mixed Prototypes Networks
This paper introduces hyperspherical mixed prototype networks. The key difference compared to hyperspherical prototype networks is mixing class prototypes and refined optimization ...
Role of Organic Agriculture in Enhancing Soil Health: Implications for Physico-Chemical and Biological Properties
Role of Organic Agriculture in Enhancing Soil Health: Implications for Physico-Chemical and Biological Properties
Soil health is fundamental to sustainable agriculture and food security. Organic agricultural practices have gained increasing recognition for its capacity to improving soil health...

