Javascript must be enabled to continue!
Robust Hierarchical Co-clustering to Explore Toxicogenomic Biomarkers and Their Regulatory Doses of Chemical Compounds
View through CrossRef
Abstract
Toxicogenomics combines high throughput molecular technologies with statistical and machine learning approaches to discover a similar group of doses of chemical compounds (DCCs) and genes to explore toxicogenomic biomarkers and their regulatory DCCs. This is also very important in the toxicity study of environmental stressors, synthetic chemicals and drug discovery and development process. Different clustering algorithms are concerned with the discovering of interesting clusters/groups of row or column entities of a dataset. Among those hierarchical clustering (HC) and logistic probabilistic hidden variable model (LPHVM) can identify toxicogenomic biomarkers and their regulatory DCCs forming co-cluster. However, the HC method is very sensitive to outlying observations. On the other hand, though LPHVM is a robust approach, it consumes more time for calculation since it is Expectation-Maximization (EM) based iterative approach. Additionally, the LPHVM creates artificiality problem taking absolute value of the data matrix. Therefore, to overcome these problems in this paper, we proposed a robust hierarchical co-clustering (RHCOC) algorithm to co-cluster genes and DCCs simultaneously with a view to explore toxicogenomic biomarkers and their regulatory DCCs. The performance of the proposed RHCOC algorithm over the conventional HC for clustering genes and DCCs of toxicogenomic data has been investigated based on the simulation study. The results of the simulation study have shown that the RHCOC approaches produce far lower clustering error rate (ER) than the conventional HC approaches in presence of outlying observations in the dataset. Otherwise they perform equally in absence of outlier in the dataset. To explore biomarker co-clusters consisting of toxicogenomic biomarker genes and their regulatory DCCs we used control chart for individual measurement (CCIM). We have also investigated the performance of the proposed approach in the case of the pathway level real life fold change gene expression (FCGE) toxicogenomic data analysis. The biomarker co-clusters consisting of toxicogenomic biomarker genes and their regulatory DCCs and biomarker genes explored by the proposed approaches have been validated by the literature and functional annotation. Our method is implemented in R package “rhcoclust” available on github (
https://github.com/mdbahadur/rhcoclust
).
Title: Robust Hierarchical Co-clustering to Explore Toxicogenomic Biomarkers and Their Regulatory Doses of Chemical Compounds
Description:
Abstract
Toxicogenomics combines high throughput molecular technologies with statistical and machine learning approaches to discover a similar group of doses of chemical compounds (DCCs) and genes to explore toxicogenomic biomarkers and their regulatory DCCs.
This is also very important in the toxicity study of environmental stressors, synthetic chemicals and drug discovery and development process.
Different clustering algorithms are concerned with the discovering of interesting clusters/groups of row or column entities of a dataset.
Among those hierarchical clustering (HC) and logistic probabilistic hidden variable model (LPHVM) can identify toxicogenomic biomarkers and their regulatory DCCs forming co-cluster.
However, the HC method is very sensitive to outlying observations.
On the other hand, though LPHVM is a robust approach, it consumes more time for calculation since it is Expectation-Maximization (EM) based iterative approach.
Additionally, the LPHVM creates artificiality problem taking absolute value of the data matrix.
Therefore, to overcome these problems in this paper, we proposed a robust hierarchical co-clustering (RHCOC) algorithm to co-cluster genes and DCCs simultaneously with a view to explore toxicogenomic biomarkers and their regulatory DCCs.
The performance of the proposed RHCOC algorithm over the conventional HC for clustering genes and DCCs of toxicogenomic data has been investigated based on the simulation study.
The results of the simulation study have shown that the RHCOC approaches produce far lower clustering error rate (ER) than the conventional HC approaches in presence of outlying observations in the dataset.
Otherwise they perform equally in absence of outlier in the dataset.
To explore biomarker co-clusters consisting of toxicogenomic biomarker genes and their regulatory DCCs we used control chart for individual measurement (CCIM).
We have also investigated the performance of the proposed approach in the case of the pathway level real life fold change gene expression (FCGE) toxicogenomic data analysis.
The biomarker co-clusters consisting of toxicogenomic biomarker genes and their regulatory DCCs and biomarker genes explored by the proposed approaches have been validated by the literature and functional annotation.
Our method is implemented in R package “rhcoclust” available on github (
https://github.
com/mdbahadur/rhcoclust
).
Related Results
The Kernel Rough K-Means Algorithm
The Kernel Rough K-Means Algorithm
Background:
Clustering is one of the most important data mining methods. The k-means
(c-means ) and its derivative methods are the hotspot in the field of clustering research in re...
7
th
International Symposium on Enabling Technologies for Life Sciences (ETP)
7
th
International Symposium on Enabling Technologies for Life Sciences (ETP)
The seventh in the series of ETP Symposia (see
Rapid Communications in Mass Spectrometry
2012,
26
, ...
Clustering Analysis of Data with High Dimensionality
Clustering Analysis of Data with High Dimensionality
Clustering analysis has been widely applied in diverse fields such as data mining, access structures, knowledge discovery, software engineering, organization of information systems...
Hierarchical Zeolites from Production Sand Waste as Catalysts for CO2 to Carbon Nanotubes CNTs: Exploration and Production Sustainability
Hierarchical Zeolites from Production Sand Waste as Catalysts for CO2 to Carbon Nanotubes CNTs: Exploration and Production Sustainability
Abstract
This project targets to convert sand waste from oil & gas production, which is typically disposed as landfill, to be the higher-value products, called "...
A COMPARATIVE ANALYSIS OF K-MEANS AND HIERARCHICAL CLUSTERING
A COMPARATIVE ANALYSIS OF K-MEANS AND HIERARCHICAL CLUSTERING
Clustering is the process of arranging comparable data elements into groups. One of the most frequent data mining analytical techniques is clustering analysis; the clustering algor...
Evaluating Clustering Algorithms: An Analysis using the EDAS Method
Evaluating Clustering Algorithms: An Analysis using the EDAS Method
Data clustering is frequently utilized in the early stages of analyzing big data. It enables the examination of massive datasets encompassing diverse types of data, with the aim of...
Image clustering using exponential discriminant analysis
Image clustering using exponential discriminant analysis
Local learning based image clustering models are usually employed to deal with images sampled from the non‐linear manifold. Recently, linear discriminant analysis (LDA) based vario...
Optimizing machine learning techniques for genomics clustering
Optimizing machine learning techniques for genomics clustering
Optimisation des techniques d’apprentissage automatique pour le clustering génomique
Dans le domaine de la bioinformatique, le clustering est une technique efficace...

