Javascript must be enabled to continue!
NetMix: A network-structured mixture model for reduced-bias estimation of altered subnetworks
View through CrossRef
Abstract
A classic problem in computational biology is the identification of
altered subnetworks:
subnetworks of an interaction network that contain genes/proteins that are differentially expressed, highly mutated, or otherwise aberrant compared to other genes/proteins. Numerous methods have been developed to solve this problem under various assumptions, but the statistical properties of these methods are often unknown. For example, some widely-used methods are reported to output very large subnetworks that are difficult to interpret biologically. In this work, we formulate the identification of altered subnetworks as the problem of estimating the parameters of a class of probability distributions which we call the Altered Subset Distribution (ASD). We derive a connection between a popular method, jActiveModules, and the maximum likelihood estimator (MLE) of the ASD. We show that the MLE is
statistically biased
, explaining the large subnetworks output by jActiveModules. We introduce NetMix, an algorithm that uses Gaussian mixture models to obtain less biased estimates of the parameters of the ASD. We demonstrate that NetMix outperforms existing methods in identifying altered subnetworks on both simulated and real data, including the identification of differentially expressed genes from both microarray and RNA-seq experiments and the identification of cancer driver genes in somatic mutation data.
Availability
NetMix is available online at
https://github.com/raphael-group/netmix
.
Contact
braphael@princeton.edu
Title: NetMix: A network-structured mixture model for reduced-bias estimation of altered subnetworks
Description:
Abstract
A classic problem in computational biology is the identification of
altered subnetworks:
subnetworks of an interaction network that contain genes/proteins that are differentially expressed, highly mutated, or otherwise aberrant compared to other genes/proteins.
Numerous methods have been developed to solve this problem under various assumptions, but the statistical properties of these methods are often unknown.
For example, some widely-used methods are reported to output very large subnetworks that are difficult to interpret biologically.
In this work, we formulate the identification of altered subnetworks as the problem of estimating the parameters of a class of probability distributions which we call the Altered Subset Distribution (ASD).
We derive a connection between a popular method, jActiveModules, and the maximum likelihood estimator (MLE) of the ASD.
We show that the MLE is
statistically biased
, explaining the large subnetworks output by jActiveModules.
We introduce NetMix, an algorithm that uses Gaussian mixture models to obtain less biased estimates of the parameters of the ASD.
We demonstrate that NetMix outperforms existing methods in identifying altered subnetworks on both simulated and real data, including the identification of differentially expressed genes from both microarray and RNA-seq experiments and the identification of cancer driver genes in somatic mutation data.
Availability
NetMix is available online at
https://github.
com/raphael-group/netmix
.
Contact
braphael@princeton.
edu.
Related Results
NetMix2: Unifying network propagation and altered subnetworks
NetMix2: Unifying network propagation and altered subnetworks
Abstract
A standard paradigm in computational biology is to use interaction networks to analyze high-throughput biological data. Two common appro...
Chaos and Mixing in NETmix
Chaos and Mixing in NETmix
NETmix is a network of mixing chambers interconnected by channels, in
which the most efficient mixing occurs under chaotic flow regimes. Chaos
is highly influenced by the topology,...
SSGA and MSGA: two seed-growing algorithms for constructing collaborative subnetworks
SSGA and MSGA: two seed-growing algorithms for constructing collaborative subnetworks
AbstractThe establishment of a collaborative network of transcription factors (TFs) followed by decomposition and then construction of subnetworks is an effective way to obtain set...
Cement Concrete Mixture Performance Characterization
Cement Concrete Mixture Performance Characterization
The cementitious composite nature of concrete makes very diffi cult directly ascertaining each mixture-factors’ contribution to a given concrete mixture performance characteristics...
Demographic bias in machine learning: measuring transference from dataset bias to model predictions
Demographic bias in machine learning: measuring transference from dataset bias to model predictions
As artificial intelligence (AI) systems increasingly influence critical decisions in society, ensuring fairness and avoiding bias have become pressing challenges. This dissertation...
Tropical Indian Ocean Mixed Layer Bias in CMIP6 CGCMs Primarily Attributed tothe AGCM Surface Wind Bias
Tropical Indian Ocean Mixed Layer Bias in CMIP6 CGCMs Primarily Attributed tothe AGCM Surface Wind Bias
The relatively weak sea surface temperature bias in the tropical Indian Ocean (TIO) simulated in the coupledgeneral circulation model (CGCM) from the recently released CMIP6 has be...
Finite element method for solving asphalt mixture problem
Finite element method for solving asphalt mixture problem
Purpose
Asphalt mixture is widely used in road engineering, and its performance research is particularly important. But the study of asphalt mixture performance needs a lot of test...
Bias in Book Recommendation
Bias in Book Recommendation
Books occupy a significant cultural role in human societies, and libraries have long positioned themselves as institutions committed to equitable access to information and promotio...

