Javascript must be enabled to continue!
Leveraging the Human Panproteome to Enhance Peptide and Protein Identification in Proteomics and Metaproteomics
View through CrossRef
Abstract
In this paper, we developed a novel approach to utilize the human pangenome to improve peptide and protein identification from proteomic data (MS/MS spectra). We propose a new data structure called panproteome graph (PPG), in which nodes are tryptic peptides, to represent the human pangenome. The PPG can be built in linear time and can be utilized via graph traversal using a depth-first search algorithm to generate potential peptides for peptide identification in proteomics. The PPG built using the 47 human proteomes from the Human Pangenome Reference Consortium (HPRC) coupled with UniProt human proteins resulted in more than 4.2M tryptic peptides, a 26% increase as compared to when only the UniProt proteins were included. Graph-based analysis of the PPG revealed a giant disconnected component with about 3M nodes, suggesting substantial sharing of tryptic peptides among proteins. We applied tryptic peptides derived from PPG to characterize three collections of human proteomic and metaproteomic datasets, and our results showed that by exploiting the human pangenome, we were able to increase the number of identified peptides on all datasets we tested (about 8% increase across all three collections). We also showed that using more complete human proteome would be useful for reducing potential misidentification of human peptides as microbial peptides, a problem that was previously studied but based on genomic sequencing data. Our tool for building PPG is available in a GitHub repo PPGpep, and PPG-derived tryptic peptides can be utilized by MetaProD, a pipeline for both human and bacterial peptide and protein identification from (meta)proteomics datasets.
Title: Leveraging the Human Panproteome to Enhance Peptide and Protein Identification in Proteomics and Metaproteomics
Description:
Abstract
In this paper, we developed a novel approach to utilize the human pangenome to improve peptide and protein identification from proteomic data (MS/MS spectra).
We propose a new data structure called panproteome graph (PPG), in which nodes are tryptic peptides, to represent the human pangenome.
The PPG can be built in linear time and can be utilized via graph traversal using a depth-first search algorithm to generate potential peptides for peptide identification in proteomics.
The PPG built using the 47 human proteomes from the Human Pangenome Reference Consortium (HPRC) coupled with UniProt human proteins resulted in more than 4.
2M tryptic peptides, a 26% increase as compared to when only the UniProt proteins were included.
Graph-based analysis of the PPG revealed a giant disconnected component with about 3M nodes, suggesting substantial sharing of tryptic peptides among proteins.
We applied tryptic peptides derived from PPG to characterize three collections of human proteomic and metaproteomic datasets, and our results showed that by exploiting the human pangenome, we were able to increase the number of identified peptides on all datasets we tested (about 8% increase across all three collections).
We also showed that using more complete human proteome would be useful for reducing potential misidentification of human peptides as microbial peptides, a problem that was previously studied but based on genomic sequencing data.
Our tool for building PPG is available in a GitHub repo PPGpep, and PPG-derived tryptic peptides can be utilized by MetaProD, a pipeline for both human and bacterial peptide and protein identification from (meta)proteomics datasets.
Related Results
7
th
International Symposium on Enabling Technologies for Life Sciences (ETP)
7
th
International Symposium on Enabling Technologies for Life Sciences (ETP)
The seventh in the series of ETP Symposia (see
Rapid Communications in Mass Spectrometry
2012,
26
, ...
NovoLign: metaproteomics by sequence alignment
NovoLign: metaproteomics by sequence alignment
Abstract
Tremendous advances in mass spectrometric and bioinformatic approaches have expanded proteomics into the field of microbial ecology. The commonly used sp...
NovoLign: metaproteomics by sequence alignment
NovoLign: metaproteomics by sequence alignment
ABSTRACT
Tremendous advances in mass spectrometric and bioinformatic approaches have expanded proteomics into the field of microbial ecology. The...
Evaluation of Protein Reference Database Reduction and Its Impact on Peptide-Centric Metaproteomics
Evaluation of Protein Reference Database Reduction and Its Impact on Peptide-Centric Metaproteomics
Abstract
Introduction/Background
Recent large-scale restructurings of UniProtKB included removal of redundant entries, exclusio...
Identification and applications of disease-associated differential human and bacterial proteins with metaproteomic evidence
Identification and applications of disease-associated differential human and bacterial proteins with metaproteomic evidence
Abstract
The gut microbiome plays a fundamental role in human health and disease. Individual variations in the microbiome and the corresponding functional implica...
Protein-peptide Interaction Representation Learning with Pretrained Language Models
Protein-peptide Interaction Representation Learning with Pretrained Language Models
Abstract
Protein-peptide Interactions (PpIs) paly essential roles in diverse cellular processes, yet their systematic identification remains challenging due to the ...
MetaDIA: A Novel Database Reduction Strategy for DIA Human Gut Metaproteomics
MetaDIA: A Novel Database Reduction Strategy for DIA Human Gut Metaproteomics
Abstract
Background
Microbiomes, especially within the gut, are complex and may comprise hundreds of species. The identificatio...
Anemia Is Inversely Associated with Serum C-Peptide Concentrations in Patients with Type 2 Diabetes
Anemia Is Inversely Associated with Serum C-Peptide Concentrations in Patients with Type 2 Diabetes
Results: The aim of the study was to investigate the relationship between anemia and serum C-peptide concentrations in Korean patients with type 2 diabetes. A total of 1,300 subjec...

