Javascript must be enabled to continue!
Machine Learning Random Forest for predicting onco-somatic variants NGS analysis
View through CrossRef
Abstract
Motivation: Since 2017, we are using IonTorrent NGS platform in our hospital in order to diagnose cancer and treatment. Analysis variants at each run take us a longtime and we are still struggling with some variants which look correct on the first look at their metrics but found to be negative when we look further into them. Can any Machine Learning algorithm help us to classify NGS variant calling ? This has determined us to investigate which ML could fit to our NGS data and to develop a tool which can be implemented in Routine in order to help Biologists. Introduction: Nowadays, one of medicine challenges is processing a significant amount of data. It’s particularly true in molecular biology with the advantage of Next Generation Sequencing (NGS) for molecular tumor profile determination and treatment selection. In addition to bioinformatics pipelines, Artificial Intelligence (AI) can offer a very valuable help in analyzing. Generating sequencing data from patient DNA samples has become easy to perform in clinical trials. But analyzing the huge amount of genomic or transcriptomic data and extracting the key biomarkers associated with a clinical response to a specific therapy requires a formidable combination of scientific expertise, biomolecular skill and a panel of bioinformatics and biostatistics tools, in which artificial intelligence is now a success factor in developing future routine diagnostics. However, cancer genome complexity and technical artifacts make identification of real variants a challenge. We present a Machine Learning method to classify pathogenic Single Nucleotide Variants (SNVs), SNP (Single Nucleotide Polymorphism), MNVs (Multiple Nucleotide Variants), Insertion, Deletion detected by NGS from tumors specimens for Colorectal, Melanoma, Lung and Glioma cancer. Methods: We compared our NGS data to different machine learning algorithms using the 10-fold cross validation method and to neural networks (Deep Learning) in order to measure the performance of the different ML algorithms and determine which one is a valid model for confirming NGS variant calls in cancer diagnostic. We trained our Machine Learning with 70 % of our data samples, extracted from our local database (our data structure had 7 parameters: chromosome, position, exon, variant allele frequency, minor allele frequency, coverage and protein description) and validated it with 30 % remaining. The model offering the best accuracy was chosen and implemented in NGS analysis routine. The artificial intelligence was developed with R script language version 3.6.0. Results: We trained our model on 102011 variants. Our best error rate (0.22%) was found with Random Forest Machine Learning (ntree=500 and mtry=4) with an AUC of 0.99. Neural Networks achieved some good scores. The final trained model with Neural Network was able to achieve an accuracy of 98% and a ROC-AUC of 0.99 with validation data. We tested our RF model to interpret more than 2000 variants from our NGS database: 20 variants were misclassified (error rate <1%). The errors were nomenclature problems and false positive. After adding false positive in our training database and implementing our RF model in routine, our error rate was always < 0.5%. Conclusion: Our RF model shows excellent results for onco-somatic NGS interpretation and it could easily be implemented in other molecular biology laboratories. AI is taking an increasingly important place in molecular biomedical analysis and could be very helpful on processing of amount medical data. Neural Networks showed a good capacity in the classification of variants and in the future may be useful in the prediction of more complex variants.
Research Square Platform LLC
Title: Machine Learning Random Forest for predicting onco-somatic variants NGS analysis
Description:
Abstract
Motivation: Since 2017, we are using IonTorrent NGS platform in our hospital in order to diagnose cancer and treatment.
Analysis variants at each run take us a longtime and we are still struggling with some variants which look correct on the first look at their metrics but found to be negative when we look further into them.
Can any Machine Learning algorithm help us to classify NGS variant calling ? This has determined us to investigate which ML could fit to our NGS data and to develop a tool which can be implemented in Routine in order to help Biologists.
Introduction: Nowadays, one of medicine challenges is processing a significant amount of data.
It’s particularly true in molecular biology with the advantage of Next Generation Sequencing (NGS) for molecular tumor profile determination and treatment selection.
In addition to bioinformatics pipelines, Artificial Intelligence (AI) can offer a very valuable help in analyzing.
Generating sequencing data from patient DNA samples has become easy to perform in clinical trials.
But analyzing the huge amount of genomic or transcriptomic data and extracting the key biomarkers associated with a clinical response to a specific therapy requires a formidable combination of scientific expertise, biomolecular skill and a panel of bioinformatics and biostatistics tools, in which artificial intelligence is now a success factor in developing future routine diagnostics.
However, cancer genome complexity and technical artifacts make identification of real variants a challenge.
We present a Machine Learning method to classify pathogenic Single Nucleotide Variants (SNVs), SNP (Single Nucleotide Polymorphism), MNVs (Multiple Nucleotide Variants), Insertion, Deletion detected by NGS from tumors specimens for Colorectal, Melanoma, Lung and Glioma cancer.
Methods: We compared our NGS data to different machine learning algorithms using the 10-fold cross validation method and to neural networks (Deep Learning) in order to measure the performance of the different ML algorithms and determine which one is a valid model for confirming NGS variant calls in cancer diagnostic.
We trained our Machine Learning with 70 % of our data samples, extracted from our local database (our data structure had 7 parameters: chromosome, position, exon, variant allele frequency, minor allele frequency, coverage and protein description) and validated it with 30 % remaining.
The model offering the best accuracy was chosen and implemented in NGS analysis routine.
The artificial intelligence was developed with R script language version 3.
6.
Results: We trained our model on 102011 variants.
Our best error rate (0.
22%) was found with Random Forest Machine Learning (ntree=500 and mtry=4) with an AUC of 0.
99.
Neural Networks achieved some good scores.
The final trained model with Neural Network was able to achieve an accuracy of 98% and a ROC-AUC of 0.
99 with validation data.
We tested our RF model to interpret more than 2000 variants from our NGS database: 20 variants were misclassified (error rate <1%).
The errors were nomenclature problems and false positive.
After adding false positive in our training database and implementing our RF model in routine, our error rate was always < 0.
5%.
Conclusion: Our RF model shows excellent results for onco-somatic NGS interpretation and it could easily be implemented in other molecular biology laboratories.
AI is taking an increasingly important place in molecular biomedical analysis and could be very helpful on processing of amount medical data.
Neural Networks showed a good capacity in the classification of variants and in the future may be useful in the prediction of more complex variants.
Related Results
Machine learning random forest for predicting oncosomatic variant NGS analysis
Machine learning random forest for predicting oncosomatic variant NGS analysis
AbstractSince 2017, we have used IonTorrent NGS platform in our hospital to diagnose and treat cancer. Analyzing variants at each run requires considerable time, and we are still s...
Selection of Injectable Drug Product Composition using Machine Learning Models (Preprint)
Selection of Injectable Drug Product Composition using Machine Learning Models (Preprint)
BACKGROUND
As of July 2020, a Web of Science search of “machine learning (ML)” nested within the search of “pharmacokinetics or pharmacodynamics” yielded over 100...
Frequency of Common Chromosomal Abnormalities in Patients with Idiopathic Acquired Aplastic Anemia
Frequency of Common Chromosomal Abnormalities in Patients with Idiopathic Acquired Aplastic Anemia
Objective: To determine the frequency of common chromosomal aberrations in local population idiopathic determine the frequency of common chromosomal aberrations in local population...
Comparison of Three Assays for Identification of IDH Mutations in AML
Comparison of Three Assays for Identification of IDH Mutations in AML
Introduction
Isocitrate dehydrogenase (IDH) mutations are present in up to 20% of acute myeloid leukemia (AML) patients and lead to production of 2-hydroxyglutarate ...
Actionable insights: Liquid NGS in community oncology practice.
Actionable insights: Liquid NGS in community oncology practice.
e15081
Background:
Liquid-based next-generation sequencing (NGS) is an integral test alongside tissue biopsy for identifying actionabl...
Next-generation sequencing with emphasis on Illumina and Ion torrent platforms.
Next-generation sequencing with emphasis on Illumina and Ion torrent platforms.
Abstract
Background: Next-generation sequencing is a type of deep sequencing. In comparison to the previously used Sanger's method, ...
Distinct mutational landscapes when comparing germline and somatic cancer variants in forty tumor suppressor genes
Distinct mutational landscapes when comparing germline and somatic cancer variants in forty tumor suppressor genes
Abstract
Germline and somatic cancer variants in tumor suppressor genes (TSGs) share loss of function mechanisms with studies of a few genes (
...
Small Subclones Harboring NOTCH1, SF3B1 or BIRC3 Mutations Are Clinically Irrelevant in Chronic Lymphocytic Leukemia
Small Subclones Harboring NOTCH1, SF3B1 or BIRC3 Mutations Are Clinically Irrelevant in Chronic Lymphocytic Leukemia
Abstract
Introduction. Ultra-deep next generation sequencing (NGS) allows sensitive detection of mutations and estimation of their clonal abundance in tumor cell pop...

