Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Integrative Network Fusion: a multi-omics approach in molecular profiling

View through CrossRef
ABSTRACT Recent technological advances and international efforts, such as The Cancer Genome Atlas (TCGA), have made available several pan-cancer datasets encompassing multiple omics layers with detailed clinical information in large collection of samples. The need has thus arisen for the development of computational methods aimed at improving cancer subtyping and biomarker identification from multi-modal data. Here we apply the Integrative Network Fusion (INF) pipeline, which combines multiple omics layers exploiting Similarity Network Fusion (SNF) within a machine learning predictive framework. INF includes a feature ranking scheme (rSNF) on SNF-integrated features, used by a classifier over juxtaposed multi-omics features (juXT). In particular, we show instances of INF implementing Random Forest (RF) and linear Support Vector Machine (LSVM) as the classifier, and two baseline RF and LSVM models are also trained on juXT. A compact RF model, called rSNFi, trained on the intersection of top-ranked biomarkers from the two approaches juXT and rSNF is finally derived. All the classifiers are run in a 10×5-fold cross-validation schema to warrant reproducibility, following the guidelines for an unbiased Data Analysis Plan by the US FDA-led initiatives MAQC/SEQC. INF is demonstrated on four classification tasks on three multi-modal TCGA oncogenomics datasets. Gene expression, protein abundances and copy number variants are used to predict estrogen receptor status (BRCA-ER, N=381) and breast invasive carcinoma subtypes (BRCA-subtypes, N=305), while gene expression, miRNA expression and methylation data is used as predictor layers for acute myeloid leukemia and renal clear cell carcinoma survival (AML-OS, N=157; KIRC-OS, N=181). In test, INF achieved similar Matthews Correlation Coefficient (MCC) values and 97% to 83% smaller feature sizes (FS), compared with juXT for BRCA-ER (MCC: 0.83 vs 0.80; FS: 56 vs 1801) and BRCA-subtypes (0.84 vs 0.80; 302 vs 1801), improving KIRC-OS performance (0.38 vs 0.31; 111 vs 2319). INF predictions are generally more accurate in test than one-dimensional omics models, with smaller signatures too, where transcriptomics consistently play the leading role. Overall, the INF framework effectively integrates multiple data levels in oncogenomics classification tasks, improving over the performance of single layers alone and naive juxtaposition, and provides compact signature sizes 1 .
Title: Integrative Network Fusion: a multi-omics approach in molecular profiling
Description:
ABSTRACT Recent technological advances and international efforts, such as The Cancer Genome Atlas (TCGA), have made available several pan-cancer datasets encompassing multiple omics layers with detailed clinical information in large collection of samples.
The need has thus arisen for the development of computational methods aimed at improving cancer subtyping and biomarker identification from multi-modal data.
Here we apply the Integrative Network Fusion (INF) pipeline, which combines multiple omics layers exploiting Similarity Network Fusion (SNF) within a machine learning predictive framework.
INF includes a feature ranking scheme (rSNF) on SNF-integrated features, used by a classifier over juxtaposed multi-omics features (juXT).
In particular, we show instances of INF implementing Random Forest (RF) and linear Support Vector Machine (LSVM) as the classifier, and two baseline RF and LSVM models are also trained on juXT.
A compact RF model, called rSNFi, trained on the intersection of top-ranked biomarkers from the two approaches juXT and rSNF is finally derived.
All the classifiers are run in a 10×5-fold cross-validation schema to warrant reproducibility, following the guidelines for an unbiased Data Analysis Plan by the US FDA-led initiatives MAQC/SEQC.
INF is demonstrated on four classification tasks on three multi-modal TCGA oncogenomics datasets.
Gene expression, protein abundances and copy number variants are used to predict estrogen receptor status (BRCA-ER, N=381) and breast invasive carcinoma subtypes (BRCA-subtypes, N=305), while gene expression, miRNA expression and methylation data is used as predictor layers for acute myeloid leukemia and renal clear cell carcinoma survival (AML-OS, N=157; KIRC-OS, N=181).
In test, INF achieved similar Matthews Correlation Coefficient (MCC) values and 97% to 83% smaller feature sizes (FS), compared with juXT for BRCA-ER (MCC: 0.
83 vs 0.
80; FS: 56 vs 1801) and BRCA-subtypes (0.
84 vs 0.
80; 302 vs 1801), improving KIRC-OS performance (0.
38 vs 0.
31; 111 vs 2319).
INF predictions are generally more accurate in test than one-dimensional omics models, with smaller signatures too, where transcriptomics consistently play the leading role.
Overall, the INF framework effectively integrates multiple data levels in oncogenomics classification tasks, improving over the performance of single layers alone and naive juxtaposition, and provides compact signature sizes 1 .

Related Results

Why Pakistan Must Lead in Regional Multi-Omics Research for Precision Medicine
Why Pakistan Must Lead in Regional Multi-Omics Research for Precision Medicine
Precision medicine has emerged as one of the most transformative movements in global healthcare, shifting the clinical emphasis from generalized treatments to highly individualized...
The Nuclear Fusion Award
The Nuclear Fusion Award
The Nuclear Fusion Award ceremony for 2009 and 2010 award winners was held during the 23rd IAEA Fusion Energy Conference in Daejeon. This time, both 2009 and 2010 award winners w...
Benchmarking multi-omics integrative clustering methods for subtype identification in colorectal cancer
Benchmarking multi-omics integrative clustering methods for subtype identification in colorectal cancer
Abstract Background and objectives Colorectal cancer (CRC) represents a heterogeneous malignancy that has concerned global burden of incidence and mortality. The tradition...
Optimising primary molecular profiling in NSCLC
Optimising primary molecular profiling in NSCLC
Abstract Introduction Molecular profiling of NSCLC is essential for optimising treatment decisions, but often incomplete. We as...
Detection of gene communities in multi-networks reveals cancer drivers
Detection of gene communities in multi-networks reveals cancer drivers
In the past years the advent of high-throughput experimental technologies provided biologists with a flood of molecular data. This huge amount of information requires the design of...
Profiling Osteoporosis via Integrated Multi-Omics Technologies
Profiling Osteoporosis via Integrated Multi-Omics Technologies
Background: Osteoporosis is a complex disorder involving bone loss and muscle degeneration. Multi-omics technologies provide novel insights into its molecular mechanisms and may su...
Integration of multi-omics datasets enables molecular classification of COPD
Integration of multi-omics datasets enables molecular classification of COPD
Chronic obstructive pulmonary disease (COPD) is an umbrella diagnosis caused by a multitude of underlying mechanisms, and molecular sub-phenotyping is needed to develop molecular d...
Machine learning combining multi-omics data and network algorithms identifies adrenocortical carcinoma prognostic biomarkers
Machine learning combining multi-omics data and network algorithms identifies adrenocortical carcinoma prognostic biomarkers
Background: Rare endocrine cancers such as Adrenocortical Carcinoma (ACC) present a serious diagnostic and prognostication challenge. The knowledge about ACC pathogenesis is incomp...

Back to Top