Javascript must be enabled to continue!

Mining Mass Spectra for Peptide Facts

Abstract The current mainstream software for peptide-centric tandem mass spectrometry data analysis can be categorized as either database-driven, which rely on a library of mass spectra to identify the peptide associated with novel query spectra, or de novo sequencing-based, which aim to find the entire peptide sequence by relying only on the query mass spectrum. While the first paradigm currently produces state-of-the-art results in peptide identification tasks, it does not inherently make use of information present in the query mass spectrum itself to refine identifications. Meanwhile, de novo approaches attempt to solve a complex problem in one go, without any search space constraints in the general case, leading to comparatively poor results. In this paper, we decompose the de novo problem into putatively easier subproblems, and we show that peptide identification rates of database-driven methods may be improved in terms of peptide identification rate by solving one such subsproblem without requiring a solution for the complete de novo task. We demonstrate this using a de novo peptide length prediction task as the chosen subproblem. As a first prototype, we show that a deep learning-based length prediction model increases peptide identification rates in the ProteomeTools dataset as part of an Pepid-based identification pipeline. Using the predicted information to better rank the candidates, we show that combining ideas from the two paradigms produces clear benefits in this setting. We propose that the next generation of peptide-centric tandem mass spectrometry identification methods should combine elements of these paradigms by mining facts “de novo; about the peptide represented in a spectrum, while simultaneously limiting the search space with a peptide candidates database.

openRxiv

Jeremie Zumer Sebastien Lemieux

2023

Title: Mining Mass Spectra for Peptide Facts

Description:

While the first paradigm currently produces state-of-the-art results in peptide identification tasks, it does not inherently make use of information present in the query mass spectrum itself to refine identifications.

Meanwhile, de novo approaches attempt to solve a complex problem in one go, without any search space constraints in the general case, leading to comparatively poor results.

In this paper, we decompose the de novo problem into putatively easier subproblems, and we show that peptide identification rates of database-driven methods may be improved in terms of peptide identification rate by solving one such subsproblem without requiring a solution for the complete de novo task.

We demonstrate this using a de novo peptide length prediction task as the chosen subproblem.

As a first prototype, we show that a deep learning-based length prediction model increases peptide identification rates in the ProteomeTools dataset as part of an Pepid-based identification pipeline.

Using the predicted information to better rank the candidates, we show that combining ideas from the two paradigms produces clear benefits in this setting.

We propose that the next generation of peptide-centric tandem mass spectrometry identification methods should combine elements of these paradigms by mining facts “de novo; about the peptide represented in a spectrum, while simultaneously limiting the search space with a peptide candidates database.

Back

Literature—at least serious literature—is something that we work at. This is especially true within the academy. Literature departments are places where workers labour over texts c...

Light at the End of the Tunnel: Mining Justice and Health

The mining industry provides valuable mined commodities and financial support for communities worldwide. Mining has become safer for workers. Significant injustices, however, are c...

Anemia Is Inversely Associated with Serum C-Peptide Concentrations in Patients with Type 2 Diabetes

Results: The aim of the study was to investigate the relationship between anemia and serum C-peptide concentrations in Korean patients with type 2 diabetes. A total of 1,300 subjec...

Evolutionary Algorithms for Improving De Novo Peptide Sequencing

<p>De novo peptide sequencing algorithms have been developed for peptide identification in proteomics from tandem mass spectra (MS/MS), which can be used to identify and disc...

Peptide-Directed Supramolecular Self-Assembly of N-Substituted Perylene Imides

<p>Synthetic peptides offer enormous potential to encode the assembly of molecular electronic components, provided that the complex range of interactions is distilled into si...

Breast Carcinoma within Fibroadenoma: A Systematic Review

Abstract Introduction Fibroadenoma is the most common benign breast lesion; however, it carries a potential risk of malignant transformation. This systematic review provides an ove...

Impact of Mining on Socioeconomic Status in Puno, Peru

This study examines the direct and indirect effects of mining activities on key socioeconomic indicators such as per capita income, the Human Development Index (HDI), and education...

Expression of peptide YY in all four islet cell types in the developing mouse pancreas suggests a common peptide YY-producing progenitor

ABSTRACT The islets of Langerhans contain four distinct endocrine cell types producing the hormones glucagon, insulin, somatostatin and pancreatic polypeptide. These...

Email:
Password:

Email:

Mining Mass Spectra for Peptide Facts

Related Results