Javascript must be enabled to continue!
Weighted Deep Embedded Clustering for Robust Representation Learning in Noisy Data
View through CrossRef
Deep clustering methods effectively learn meaningful representations from unlabeled data. Typically, existing approaches treat all samples equally during training, which can lead to unstable clustering in the presence of noise and outliers. In this research, we propose a Reliability-Based Deep Embedded Clustering (RDEC) approach that improves clustering reliability through a novel distance-aware weighting strategy. Specifically, rather than assigning equal importance to all samples, RDEC estimates and exploits the typicality of each instance according to its distance from the cluster centers in the latent space. Accordingly, samples that lie closer to cluster centers are assigned higher weights, while distant, atypical and ambiguous samples are gradually down-weighted using an inverse polynomial function. Moreover, the proposed weighting approach is coupled with a novel Kullback–Leibler divergence objective function to focus on the most representative data instances and guide the clustering process. Furthermore, the resulting latent distances and reliability weights demonstrate that RDEC effectively distinguishes reliable samples from noisy ones, resulting in more stable clustering behavior. RDEC was evaluated on three benchmark image datasets, namely MNIST, USPS, and Fashion-MNIST, under both clean conditions and controlled image corruption settings, including semantic outliers and synthetic perturbations, to assess its robustness during clustering. The experimental results demonstrated that RDEC achieved competitive clustering performance compared with representative clustering methods evaluated under the same experimental setting. In particular, RDEC yielded the best overall performance on the clean USPS dataset, attaining an ACC, NMI, and ARI of 0.8079, 0.7714, and 0.7392, respectively. Moreover, on the clean Fashion-MNIST dataset, RDEC achieved the highest NMI and ARI while maintaining competitive clustering accuracy, demonstrating its effectiveness on a more challenging clustering benchmark. Under these controlled image corruption scenarios, RDEC consistently achieved strong clustering performance across multiple corruption levels on MNIST, USPS, and Fashion-MNIST datasets, demonstrating the effectiveness of the proposed reliability-aware weighting mechanism under diverse image characteristics. The present study focuses on robustness under controlled image-based corruption scenarios, with the evaluation limited to benchmark image datasets.
Title: Weighted Deep Embedded Clustering for Robust Representation Learning in Noisy Data
Description:
Deep clustering methods effectively learn meaningful representations from unlabeled data.
Typically, existing approaches treat all samples equally during training, which can lead to unstable clustering in the presence of noise and outliers.
In this research, we propose a Reliability-Based Deep Embedded Clustering (RDEC) approach that improves clustering reliability through a novel distance-aware weighting strategy.
Specifically, rather than assigning equal importance to all samples, RDEC estimates and exploits the typicality of each instance according to its distance from the cluster centers in the latent space.
Accordingly, samples that lie closer to cluster centers are assigned higher weights, while distant, atypical and ambiguous samples are gradually down-weighted using an inverse polynomial function.
Moreover, the proposed weighting approach is coupled with a novel Kullback–Leibler divergence objective function to focus on the most representative data instances and guide the clustering process.
Furthermore, the resulting latent distances and reliability weights demonstrate that RDEC effectively distinguishes reliable samples from noisy ones, resulting in more stable clustering behavior.
RDEC was evaluated on three benchmark image datasets, namely MNIST, USPS, and Fashion-MNIST, under both clean conditions and controlled image corruption settings, including semantic outliers and synthetic perturbations, to assess its robustness during clustering.
The experimental results demonstrated that RDEC achieved competitive clustering performance compared with representative clustering methods evaluated under the same experimental setting.
In particular, RDEC yielded the best overall performance on the clean USPS dataset, attaining an ACC, NMI, and ARI of 0.
8079, 0.
7714, and 0.
7392, respectively.
Moreover, on the clean Fashion-MNIST dataset, RDEC achieved the highest NMI and ARI while maintaining competitive clustering accuracy, demonstrating its effectiveness on a more challenging clustering benchmark.
Under these controlled image corruption scenarios, RDEC consistently achieved strong clustering performance across multiple corruption levels on MNIST, USPS, and Fashion-MNIST datasets, demonstrating the effectiveness of the proposed reliability-aware weighting mechanism under diverse image characteristics.
The present study focuses on robustness under controlled image-based corruption scenarios, with the evaluation limited to benchmark image datasets.
Related Results
CREATING LEARNING MEDIA IN TEACHING ENGLISH AT SMP MUHAMMADIYAH 2 PAGELARAN ACADEMIC YEAR 2020/2021
CREATING LEARNING MEDIA IN TEACHING ENGLISH AT SMP MUHAMMADIYAH 2 PAGELARAN ACADEMIC YEAR 2020/2021
The pandemic Covid-19 currently demands teachers to be able to use technology in teaching and learning process. But in reality there are still many teachers who have not been able ...
Selection of Injectable Drug Product Composition using Machine Learning Models (Preprint)
Selection of Injectable Drug Product Composition using Machine Learning Models (Preprint)
BACKGROUND
As of July 2020, a Web of Science search of “machine learning (ML)” nested within the search of “pharmacokinetics or pharmacodynamics” yielded over 100...
The Kernel Rough K-Means Algorithm
The Kernel Rough K-Means Algorithm
Background:
Clustering is one of the most important data mining methods. The k-means
(c-means ) and its derivative methods are the hotspot in the field of clustering research in re...
Image clustering using exponential discriminant analysis
Image clustering using exponential discriminant analysis
Local learning based image clustering models are usually employed to deal with images sampled from the non‐linear manifold. Recently, linear discriminant analysis (LDA) based vario...
Cluster evaluation on weighted networks
Cluster evaluation on weighted networks
(English) This thesis presents a systematic approach to validate the results of clustering methods on weighted networks, particularly for the cases where the existence of a communi...
Optimizing machine learning techniques for genomics clustering
Optimizing machine learning techniques for genomics clustering
Optimisation des techniques d’apprentissage automatique pour le clustering génomique
Dans le domaine de la bioinformatique, le clustering est une technique efficace...
RNN BASED FEATURE SELECTION MODEL WITH NOISY FEATURE REMOVAL
RNN BASED FEATURE SELECTION MODEL WITH NOISY FEATURE REMOVAL
The evolution of emerging technologies leads into the unbeatable data growth and this circumstance brings out the necessity of
reduction in data dimensionality. Feature Selection i...
How suitable are clustering methods for functional annotation of proteins?
How suitable are clustering methods for functional annotation of proteins?
Abstract
The advent of affordable high-throughput genome sequencing has drastically expanded protein sequence databases, necessitating the development of computatio...

