Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Elucidation of Genome-wide Understudied Proteins targeted by PROTAC-induced degradation using Interpretable Machine Learning

View through CrossRef
Abstract Proteolysis-targeting chimeras (PROTACs) are hetero-bifunctional molecules. They induce the degradation of a target protein by recruiting an E3 ligase to the target. The PROTAC can inactivate disease-related genes that are considered as understudied, thus has a great potential to be a new type of therapy for the treatment of incurable diseases. However, only hundreds of proteins have been experimentally tested if they are amenable to the PROTACs. It remains elusive what other proteins can be targeted by the PROTAC in the entire human genome. For the first time, we have developed an interpretable machine learning model PrePROTAC, which is based on a transformer-based protein sequence descriptor and random forest classification to predict genome-wide PROTAC-induced targets degradable by CRBN, one of the E3 ligases. In the benchmark studies, PrePROTAC achieved ROC-AUC of 0.81, PR-AUC of 0.84, and over 40% sensitivity at a false positive rate of 0.05, respectively. Furthermore, we developed an embedding SHapley Additive exPlanations (eSHAP) method to identify positions in the protein structure, which play key roles in the PROTAC activity. The key residues identified were consistent with our existing knowledge. We applied PrePROTAC to identify more than 600 novel understudied proteins that are potentially degradable by CRBN, and proposed PROTAC compounds for three novel drug targets associated with Alzheimer’s disease. Author Summary Many human diseases remain incurable because disease-causing genes cannot by selectively and effectively targeted by small molecules. Proteolysis-targeting chimera (PROTAC), an organic compound that binds to both a target and a degradation-mediating E3 ligase, has emerged as a promising approach to selectively target disease-driving genes that are not druggable by small molecules. Nevertheless, not all of proteins can be accommodated by E3 ligases, and be effectively degraded. Knowledge on the degradability of a protein will be crucial for the design of PROTACs. However, only hundreds of proteins have been experimentally tested if they are amenable to the PROTACs. It remains elusive what other proteins can be targeted by the PROTAC in the entire human genome. In this paper, we propose an intepretable machine learning model PrePROTAC that takes advantage of powerful protein language modeling. PrePROTAC achieves high accuracy when evaluated by an external dataset which comes from different gene families from the proteins in the training data, suggesting the generalizability of PrePROTAC. We apply PrePROTAC to the human genome, and identify more than 600 understudied proteins that are potentially responsive to the PROTAC. Furthermore, we design three PROTAC compounds for novel drug targets associated with Alzheimer’s disease.
Title: Elucidation of Genome-wide Understudied Proteins targeted by PROTAC-induced degradation using Interpretable Machine Learning
Description:
Abstract Proteolysis-targeting chimeras (PROTACs) are hetero-bifunctional molecules.
They induce the degradation of a target protein by recruiting an E3 ligase to the target.
The PROTAC can inactivate disease-related genes that are considered as understudied, thus has a great potential to be a new type of therapy for the treatment of incurable diseases.
However, only hundreds of proteins have been experimentally tested if they are amenable to the PROTACs.
It remains elusive what other proteins can be targeted by the PROTAC in the entire human genome.
For the first time, we have developed an interpretable machine learning model PrePROTAC, which is based on a transformer-based protein sequence descriptor and random forest classification to predict genome-wide PROTAC-induced targets degradable by CRBN, one of the E3 ligases.
In the benchmark studies, PrePROTAC achieved ROC-AUC of 0.
81, PR-AUC of 0.
84, and over 40% sensitivity at a false positive rate of 0.
05, respectively.
Furthermore, we developed an embedding SHapley Additive exPlanations (eSHAP) method to identify positions in the protein structure, which play key roles in the PROTAC activity.
The key residues identified were consistent with our existing knowledge.
We applied PrePROTAC to identify more than 600 novel understudied proteins that are potentially degradable by CRBN, and proposed PROTAC compounds for three novel drug targets associated with Alzheimer’s disease.
Author Summary Many human diseases remain incurable because disease-causing genes cannot by selectively and effectively targeted by small molecules.
Proteolysis-targeting chimera (PROTAC), an organic compound that binds to both a target and a degradation-mediating E3 ligase, has emerged as a promising approach to selectively target disease-driving genes that are not druggable by small molecules.
Nevertheless, not all of proteins can be accommodated by E3 ligases, and be effectively degraded.
Knowledge on the degradability of a protein will be crucial for the design of PROTACs.
However, only hundreds of proteins have been experimentally tested if they are amenable to the PROTACs.
It remains elusive what other proteins can be targeted by the PROTAC in the entire human genome.
In this paper, we propose an intepretable machine learning model PrePROTAC that takes advantage of powerful protein language modeling.
PrePROTAC achieves high accuracy when evaluated by an external dataset which comes from different gene families from the proteins in the training data, suggesting the generalizability of PrePROTAC.
We apply PrePROTAC to the human genome, and identify more than 600 understudied proteins that are potentially responsive to the PROTAC.
Furthermore, we design three PROTAC compounds for novel drug targets associated with Alzheimer’s disease.

Related Results

Interpretable PROTAC degradation prediction with structure-informed deep ternary attention framework
Interpretable PROTAC degradation prediction with structure-informed deep ternary attention framework
Proteolysis Targeting Chimeras (PROTACs) are heterobifunctional ligands that form ternary complexes with Protein Of Interests (POIs) and E3 ligases, exploiting the ubiquitin-protea...
LM-PROTAC: a language model-driven PROTAC generation pipeline with dual constraints of structure and property
LM-PROTAC: a language model-driven PROTAC generation pipeline with dual constraints of structure and property
Abstract The imperfect modeling of ternary complexes has limited the application of computer-aided drug discovery tools in PROTAC research and development. In this study, a...
Abstract 1685: Overcoming acquired resistance to PROTAC degraders
Abstract 1685: Overcoming acquired resistance to PROTAC degraders
Abstract Background: Proteolysis-targeting chimera (PROTAC) technology has been widely investigated for cancer treatment and there have been several PROTAC degrader-...
Data-driven Design of PROTAC Linkers to Improve PROTAC Cell Membrane Permeability
Data-driven Design of PROTAC Linkers to Improve PROTAC Cell Membrane Permeability
Proteolysis-targeting chimeras (PROTACs) are promising next-generation therapeutics for the degradation of disease-associated proteins. However, optimizing the physicochemical prop...
SE(3)-PROTACs: Geometric deep learning for PROTAC degradation prediction
SE(3)-PROTACs: Geometric deep learning for PROTAC degradation prediction
Abstract Proteolysis-targeting chimeras (PROTACs) are a valuable therapeutic method for degrading target proteins of interest. Their success depends on forming a ...
Benchmarking of PROTAC docking and virtual screening tools
Benchmarking of PROTAC docking and virtual screening tools
Abstract Proteolysis targeting chimeras (PROTACs) are bifunctional compounds that recruit an E3 ligase to a target protein to induce ubiquitination and degradation ...
Selection of Injectable Drug Product Composition using Machine Learning Models (Preprint)
Selection of Injectable Drug Product Composition using Machine Learning Models (Preprint)
BACKGROUND As of July 2020, a Web of Science search of “machine learning (ML)” nested within the search of “pharmacokinetics or pharmacodynamics” yielded over 100...
CREATING LEARNING MEDIA IN TEACHING ENGLISH AT SMP MUHAMMADIYAH 2 PAGELARAN ACADEMIC YEAR 2020/2021
CREATING LEARNING MEDIA IN TEACHING ENGLISH AT SMP MUHAMMADIYAH 2 PAGELARAN ACADEMIC YEAR 2020/2021
The pandemic Covid-19 currently demands teachers to be able to use technology in teaching and learning process. But in reality there are still many teachers who have not been able ...

Back to Top