Javascript must be enabled to continue!
Leveraging the Human Panproteome to Enhance Peptide and Protein Identification in Proteomics and Metaproteomics
View through CrossRef
Abstract
In this paper, we developed a novel approach to utilize the human pangenome to improve peptide and protein identification from proteomic data (MS/MS spectra). We propose a new data structure called panproteome graph (PPG), in which nodes are tryptic peptides, to represent the human pangenome. The PPG can be built in linear time and can be utilized via graph traversal using a depth-first search algorithm to generate potential peptides for peptide identification in proteomics. The PPG built using the 47 human proteomes from the Human Pangenome Reference Consortium (HPRC) coupled with UniProt human proteins resulted in more than 4.2M tryptic peptides, a 26% increase as compared to when only the UniProt proteins were included. Graph-based analysis of the PPG revealed a giant disconnected component with about 3M nodes, suggesting substantial sharing of tryptic peptides among proteins. We applied tryptic peptides derived from PPG to characterize three collections of human proteomic and metaproteomic datasets, and our results showed that by exploiting the human pangenome, we were able to increase the number of identified peptides on all datasets we tested (about 8% increase across all three collections). We also showed that using more complete human proteome would be useful for reducing potential misidentification of human peptides as microbial peptides, a problem that was previously studied but based on genomic sequencing data. Our tool for building PPG is available in a GitHub repo PPGpep, and PPG-derived tryptic peptides can be utilized by MetaProD, a pipeline for both human and bacterial peptide and protein identification from (meta)proteomics datasets.
Title: Leveraging the Human Panproteome to Enhance Peptide and Protein Identification in Proteomics and Metaproteomics
Description:
Abstract
In this paper, we developed a novel approach to utilize the human pangenome to improve peptide and protein identification from proteomic data (MS/MS spectra).
We propose a new data structure called panproteome graph (PPG), in which nodes are tryptic peptides, to represent the human pangenome.
The PPG can be built in linear time and can be utilized via graph traversal using a depth-first search algorithm to generate potential peptides for peptide identification in proteomics.
The PPG built using the 47 human proteomes from the Human Pangenome Reference Consortium (HPRC) coupled with UniProt human proteins resulted in more than 4.
2M tryptic peptides, a 26% increase as compared to when only the UniProt proteins were included.
Graph-based analysis of the PPG revealed a giant disconnected component with about 3M nodes, suggesting substantial sharing of tryptic peptides among proteins.
We applied tryptic peptides derived from PPG to characterize three collections of human proteomic and metaproteomic datasets, and our results showed that by exploiting the human pangenome, we were able to increase the number of identified peptides on all datasets we tested (about 8% increase across all three collections).
We also showed that using more complete human proteome would be useful for reducing potential misidentification of human peptides as microbial peptides, a problem that was previously studied but based on genomic sequencing data.
Our tool for building PPG is available in a GitHub repo PPGpep, and PPG-derived tryptic peptides can be utilized by MetaProD, a pipeline for both human and bacterial peptide and protein identification from (meta)proteomics datasets.
Related Results
Evaluation of Protein Reference Database Reduction and Its Impact on Peptide-Centric Metaproteomics
Evaluation of Protein Reference Database Reduction and Its Impact on Peptide-Centric Metaproteomics
Abstract
Introduction/Background
Recent large-scale restructurings of UniProtKB included removal of redundant entries, exclusio...
Protein-peptide Interaction Representation Learning with Pretrained Language Models
Protein-peptide Interaction Representation Learning with Pretrained Language Models
Abstract
Protein-peptide Interactions (PpIs) paly essential roles in diverse cellular processes, yet their systematic identification remains challenging due to the ...
Identification and applications of disease-associated differential human and bacterial proteins with metaproteomic evidence
Identification and applications of disease-associated differential human and bacterial proteins with metaproteomic evidence
Abstract
The gut microbiome plays a fundamental role in human health and disease. Individual variations in the microbiome and the corresponding functional implica...
MetaDIA: A Novel Database Reduction Strategy for DIA Human Gut Metaproteomics
MetaDIA: A Novel Database Reduction Strategy for DIA Human Gut Metaproteomics
Abstract
Background
Microbiomes, especially within the gut, are complex and may comprise hundreds of species. The identificatio...
Endothelial Protein C Receptor
Endothelial Protein C Receptor
IntroductionThe protein C anticoagulant pathway plays a critical role in the negative regulation of the blood clotting response. The pathway is triggered by thrombin, which allows ...
Anemia Is Inversely Associated with Serum C-Peptide Concentrations in Patients with Type 2 Diabetes
Anemia Is Inversely Associated with Serum C-Peptide Concentrations in Patients with Type 2 Diabetes
Results: The aim of the study was to investigate the relationship between anemia and serum C-peptide concentrations in Korean patients with type 2 diabetes. A total of 1,300 subjec...
ComplexBrowser: a tool for identification and quantification of protein complexes in large scale proteomics datasets
ComplexBrowser: a tool for identification and quantification of protein complexes in large scale proteomics datasets
Abstract
We have developed ComplexBrowser, an open source, online platform for supervised analysis of quantitative proteomics data that focuses on protein complexes...
Novel Approaches to Plastic Pollution: Leveraging Machine Learning and Metaproteomics for Advanced Plastic Degradation
Novel Approaches to Plastic Pollution: Leveraging Machine Learning and Metaproteomics for Advanced Plastic Degradation
This study addresses a pressing global issue—plastic waste—and explores technologies such as machine learning and metaproteomics as potential solutions. Current initiatives to redu...

