Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Rough set theory for document clustering: A review

View through CrossRef
Rough set theory is a mathematical framework that can be visualized as a soft computing tool dealing with the vagueness and uncertainty of data and is applied to pattern recognition, data mining, and knowledge discovery. Document clustering is another area of research with values which are a bag of words that describe contents within clusters. This work analyzes how rough set theory is used for document clustering to fix issues that clustering methods manage. In this survey, an exhaustive literature review of the concept of rough sets, as well as how the lower and upper approximation of a set can be used for document clustering, has been presented. Rough set clusters are shown to be useful for representing real-time applications such as biomedical inferences, network data handling, and citation analysis. The survey is done in phases, showing how machine learning algorithms have been incorporated for document clustering using rough set theory, as well as how rough set theory has been extended to adapt to document clustering with feature selection techniques and feature/dimensionality reduction and, finally, ending with a view of assorted clustering tasks where rough set theory is applied. The classification of rough set theory for document clustering is depicted and its applications presented in this paper. The rough set theory works with resolving ambiguity and uncertainty in data. To the best of our knowledge, a rough set clustering survey has not been done earlier in the literature reviewed and the survey ends with a critical analysis of rough set theory in each application of clustering.
Title: Rough set theory for document clustering: A review
Description:
Rough set theory is a mathematical framework that can be visualized as a soft computing tool dealing with the vagueness and uncertainty of data and is applied to pattern recognition, data mining, and knowledge discovery.
Document clustering is another area of research with values which are a bag of words that describe contents within clusters.
This work analyzes how rough set theory is used for document clustering to fix issues that clustering methods manage.
In this survey, an exhaustive literature review of the concept of rough sets, as well as how the lower and upper approximation of a set can be used for document clustering, has been presented.
Rough set clusters are shown to be useful for representing real-time applications such as biomedical inferences, network data handling, and citation analysis.
The survey is done in phases, showing how machine learning algorithms have been incorporated for document clustering using rough set theory, as well as how rough set theory has been extended to adapt to document clustering with feature selection techniques and feature/dimensionality reduction and, finally, ending with a view of assorted clustering tasks where rough set theory is applied.
The classification of rough set theory for document clustering is depicted and its applications presented in this paper.
The rough set theory works with resolving ambiguity and uncertainty in data.
To the best of our knowledge, a rough set clustering survey has not been done earlier in the literature reviewed and the survey ends with a critical analysis of rough set theory in each application of clustering.

Related Results

Theoretical study of laser-cooled SH<sup>–</sup> anion
Theoretical study of laser-cooled SH<sup>–</sup> anion
The potential energy curves, dipole moments, and transition dipole moments for the <inline-formula><tex-math id="M13">\begin{document}${{\rm{X}}^1}{\Sigma ^ + }$\end{do...
Ab initio study on the hydrogen desorption from $\rm {MH\text{&#x2013;}NH}_3$MH–NH3 (M = Li, Na, K) hydrogen storage systems
Ab initio study on the hydrogen desorption from $\rm {MH\text{&#x2013;}NH}_3$MH–NH3 (M = Li, Na, K) hydrogen storage systems
The hydrogen storage system LiH + \documentclass[12pt]{minimal}\begin{document}$\rm {NH}_3$\end{document} NH 3 ↔ \documentclass[12pt]{minimal}\begin{document}$\rm {LiNH}_2$\end{doc...
The Kernel Rough K-Means Algorithm
The Kernel Rough K-Means Algorithm
Background: Clustering is one of the most important data mining methods. The k-means (c-means ) and its derivative methods are the hotspot in the field of clustering research in re...
Revisiting near-threshold photoelectron interference in argon with a non-adiabatic semiclassical model
Revisiting near-threshold photoelectron interference in argon with a non-adiabatic semiclassical model
<sec> <b>Purpose:</b> The interaction of intense, ultrashort laser pulses with atoms gives rise to rich non-perturbative phenomena, which are encoded within th...
Decision Theoretic Evaluation of Rough Fuzzy Clustering
Decision Theoretic Evaluation of Rough Fuzzy Clustering
Clustering is the process of organizing dissimilar objects into natural groups in such a way objects in the same group is more similar than objects in the different groups. Since w...
Evaluating the Science to Inform the Physical Activity Guidelines for Americans Midcourse Report
Evaluating the Science to Inform the Physical Activity Guidelines for Americans Midcourse Report
Abstract The Physical Activity Guidelines for Americans (Guidelines) advises older adults to be as active as possible. Yet, despite the well documented benefits of physical a...
Image clustering using exponential discriminant analysis
Image clustering using exponential discriminant analysis
Local learning based image clustering models are usually employed to deal with images sampled from the non‐linear manifold. Recently, linear discriminant analysis (LDA) based vario...
Optimizing machine learning techniques for genomics clustering
Optimizing machine learning techniques for genomics clustering
Optimisation des techniques d’apprentissage automatique pour le clustering génomique Dans le domaine de la bioinformatique, le clustering est une technique efficace...

Back to Top