Javascript must be enabled to continue!

Evolutionary Algorithms for Improving De Novo Peptide Sequencing

<p>De novo peptide sequencing algorithms have been developed for peptide identification in proteomics from tandem mass spectra (MS/MS), which can be used to identify and discover novel peptides and proteins that do not have a database available. Despite improvements in MS instrumentation and de novo sequencing methods, a significant number of CID MS/MS spectra still remain unassigned with the current algorithms, often leading to low confidence of peptide assignments to the spectra. Moreover, current algorithms often fail to construct the completely matched sequences, and produce partial matches. Therefore, identification of full-length peptides remains challenging. Another major challenge is the existence of noise in MS/MS spectra which makes the data highly imbalanced. Also missing peaks, caused by incomplete MS fragmentation makes it more difficult to infer a full-length peptide sequence. In addition, the large search space of all possible amino acid sequences for each spectrum leads to a high false discovery rate. This thesis focuses on improving the performance of current methods by developing new algorithms corresponding to three steps of preprocessing, sequence optimisation and post-processing using machine learning for more comprehensive interrogation of MS/MS datasets. From the machine learning point of view, the three steps can be addressed by solving different tasks such as classification, optimisation, and symbolic regression. Since Evolutionary Algorithms (EAs), as effective global search techniques, have shown promising results in solving these problems, this thesis investigates the capability of EAs in improving the de novo peptide sequencing. In the preprocessing step, this thesis proposes an effective GP-based method for classification of signal and noise peaks in highly imbalanced MS/MS spectra with the purpose of having a positive influence on the reliability of the peptide identification. The results show that the proposed algorithm is the most stable classification method across various noise ratios, outperforming six other benchmark classification algorithms. The experimental results show a significant improvement in high confidence peptide assignments to MS/MS spectra when the data is preprocessed by the proposed GP method. Moreover, the first multi-objective GP approach for classification of peaks in MS/MS data, aiming at maximising the accuracy of the minority class (signal peaks) and the accuracy of the majority class (noise peaks) is also proposed in this thesis. The results show that the multi-objective GP method outperforms the single objective GP algorithm and a popular multi-objective approach in terms of retaining more signal peaks and removing more noise peaks. The multi-objective GP approach significantly improved the reliability of peptide identification. This thesis proposes a GA-based method to solve the complex optimisation task of de novo peptide sequencing, aiming at constructing full-length sequences. The proposed GA method benefits the GA capability of searching a large search space of potential amino acid sequences to find the most likely full-length sequence. The experimental results show that the proposed method outperforms the most commonly used de novo sequencing method at both amino acid level and peptide level. This thesis also proposes a novel method for re-scoring and re-ranking the peptide spectrum matches (PSMs) from the result of de novo peptide sequencing, aiming at minimising the false discovery rate as a post-processing approach. The proposed GP method evolves the computer programs to perform regression and classification simultaneously in order to generate an effective scoring function for finding the correct PSMs from many incorrect ones. The results show that the new GP-based PSM scoring function significantly improves the identification of full-length peptides when it is used to post-process the de novo sequencing results.</p>

Victoria University of Wellington Library

Samaneh Azari

2021

Title: Evolutionary Algorithms for Improving De Novo Peptide Sequencing

Description:

Despite improvements in MS instrumentation and de novo sequencing methods, a significant number of CID MS/MS spectra still remain unassigned with the current algorithms, often leading to low confidence of peptide assignments to the spectra.

Moreover, current algorithms often fail to construct the completely matched sequences, and produce partial matches.

Therefore, identification of full-length peptides remains challenging.

Another major challenge is the existence of noise in MS/MS spectra which makes the data highly imbalanced.

Also missing peaks, caused by incomplete MS fragmentation makes it more difficult to infer a full-length peptide sequence.

In addition, the large search space of all possible amino acid sequences for each spectrum leads to a high false discovery rate.

This thesis focuses on improving the performance of current methods by developing new algorithms corresponding to three steps of preprocessing, sequence optimisation and post-processing using machine learning for more comprehensive interrogation of MS/MS datasets.

From the machine learning point of view, the three steps can be addressed by solving different tasks such as classification, optimisation, and symbolic regression.

Since Evolutionary Algorithms (EAs), as effective global search techniques, have shown promising results in solving these problems, this thesis investigates the capability of EAs in improving the de novo peptide sequencing.

In the preprocessing step, this thesis proposes an effective GP-based method for classification of signal and noise peaks in highly imbalanced MS/MS spectra with the purpose of having a positive influence on the reliability of the peptide identification.

The results show that the proposed algorithm is the most stable classification method across various noise ratios, outperforming six other benchmark classification algorithms.

The experimental results show a significant improvement in high confidence peptide assignments to MS/MS spectra when the data is preprocessed by the proposed GP method.

Moreover, the first multi-objective GP approach for classification of peaks in MS/MS data, aiming at maximising the accuracy of the minority class (signal peaks) and the accuracy of the majority class (noise peaks) is also proposed in this thesis.

The results show that the multi-objective GP method outperforms the single objective GP algorithm and a popular multi-objective approach in terms of retaining more signal peaks and removing more noise peaks.

The multi-objective GP approach significantly improved the reliability of peptide identification.

This thesis proposes a GA-based method to solve the complex optimisation task of de novo peptide sequencing, aiming at constructing full-length sequences.

The proposed GA method benefits the GA capability of searching a large search space of potential amino acid sequences to find the most likely full-length sequence.

The experimental results show that the proposed method outperforms the most commonly used de novo sequencing method at both amino acid level and peptide level.

This thesis also proposes a novel method for re-scoring and re-ranking the peptide spectrum matches (PSMs) from the result of de novo peptide sequencing, aiming at minimising the false discovery rate as a post-processing approach.

The proposed GP method evolves the computer programs to perform regression and classification simultaneously in order to generate an effective scoring function for finding the correct PSMs from many incorrect ones.

The results show that the new GP-based PSM scoring function significantly improves the identification of full-length peptides when it is used to post-process the de novo sequencing results.

</p>.

Back

Results: The aim of the study was to investigate the relationship between anemia and serum C-peptide concentrations in Korean patients with type 2 diabetes. A total of 1,300 subjec...

Evolution and the cell

Genotype to phenotype, and back again Evolution is intimately linked to biology at the cellular scale- evolutionary processes act on the very genetic material that is carried and ...

Expression of peptide YY in all four islet cell types in the developing mouse pancreas suggests a common peptide YY-producing progenitor

ABSTRACT The islets of Langerhans contain four distinct endocrine cell types producing the hormones glucagon, insulin, somatostatin and pancreatic polypeptide. These...

12481 Evaluating The Device Handling And Preference Assessment Questionnaire For Growth Hormone Deficiency: Results From Content Validation Interviews

Abstract Disclosure: J. Neergaard: Employee; Self; Novo Nordisk. S. Akhtar: Employee; Self; Novo Nordisk. Stock Owner; Self; Novo Nordisk. B. Berg: Employee; Self; N...

1651-P: Elevated FGF21 Levels after Total Pancreatectomy and in Response to Single-Dose Glucagon Receptor Antagonism in Humans

Fibroblast growth factor 21 (FGF21) is a liver-secreted peptide hormone reportedly improving metabolic homeostasis, partly via reduced hunger for sugar, fat, protein and alcohol. E...

Robots Need Some Education

Evolutionary Robotics and Robot Learning are two fields in robotics that aim to automatically optimize robot designs. The key difference between them lies in what is being optimize...

MARS-seq2.0: an experimental and analytical pipeline for indexed sorting combined with single-cell RNA sequencing v1

Human tissues comprise trillions of cells that populate a complex space of molecular phenotypes and functions and that vary in abundance by 4–9 orders of magnitude. Relying solely ...

Evolutionary Biomechanics

Life has diversified on Earth in many stunning ways. Understanding how this diversity arose and has been maintained is a common interest for many evolutionary biologists. One appro...

Email:
Password:

Email:

Evolutionary Algorithms for Improving De Novo Peptide Sequencing

Related Results