Javascript must be enabled to continue!
Analysis of Short Texts Using Intelligent Clustering Methods
View through CrossRef
This article presents a comprehensive review of short text clustering using state-of-the-art methods: Bidirectional Encoder Representations from Transformers (BERT), Term Frequency-Inverse Document Frequency (TF-IDF), and the novel hybrid method Latent Dirichlet Allocation+BERT+Autoencoder (LDA+BERT+AE). The article begins by outlining the theoretical foundation of each technique and their merits and limitations. BERT is critiqued for its capability to understand word dependence in text, while TF-IDF is lauded for its applicability in terms of importance assessment. The experimental section compares the efficacy of these methods in clustering short texts, with a specific focus on the hybrid LDA+BERT+AE approach. A detailed examination of the LDA-BERT model’s training and validation loss over 200 epochs shows that the loss values start above 1.2 and quickly decrease to around 0.8 within the first 25 epochs, eventually stabilizing at approximately 0.4. The close alignment of these curves suggests the model’s practical learning and generalization capabilities, with minimal overfitting. The study demonstrates that the hybrid LDA+BERT+AE method significantly enhances text clustering quality compared to individual methods. Based on the findings, the study recommends the optimum choice and use of clustering methods for different short texts and natural language processing operations. The applications of these methods in industrial and educational settings where successful text handling and categorization are critical are also addressed. The study ends by emphasizing the importance of holistic handling of short texts for deeper semantic comprehension and effective information retrieval.
Title: Analysis of Short Texts Using Intelligent Clustering Methods
Description:
This article presents a comprehensive review of short text clustering using state-of-the-art methods: Bidirectional Encoder Representations from Transformers (BERT), Term Frequency-Inverse Document Frequency (TF-IDF), and the novel hybrid method Latent Dirichlet Allocation+BERT+Autoencoder (LDA+BERT+AE).
The article begins by outlining the theoretical foundation of each technique and their merits and limitations.
BERT is critiqued for its capability to understand word dependence in text, while TF-IDF is lauded for its applicability in terms of importance assessment.
The experimental section compares the efficacy of these methods in clustering short texts, with a specific focus on the hybrid LDA+BERT+AE approach.
A detailed examination of the LDA-BERT model’s training and validation loss over 200 epochs shows that the loss values start above 1.
2 and quickly decrease to around 0.
8 within the first 25 epochs, eventually stabilizing at approximately 0.
4.
The close alignment of these curves suggests the model’s practical learning and generalization capabilities, with minimal overfitting.
The study demonstrates that the hybrid LDA+BERT+AE method significantly enhances text clustering quality compared to individual methods.
Based on the findings, the study recommends the optimum choice and use of clustering methods for different short texts and natural language processing operations.
The applications of these methods in industrial and educational settings where successful text handling and categorization are critical are also addressed.
The study ends by emphasizing the importance of holistic handling of short texts for deeper semantic comprehension and effective information retrieval.
Related Results
Žanrovska analiza pomorskopravnih tekstova i ostvarenje prijevodnih univerzalija u njihovim prijevodima s engleskoga jezika
Žanrovska analiza pomorskopravnih tekstova i ostvarenje prijevodnih univerzalija u njihovim prijevodima s engleskoga jezika
Genre implies formal and stylistic conventions of a particular text type, which inevitably affects the translation process. This „force of genre bias“ (Prieto Ramos, 2014) has been...
The Kernel Rough K-Means Algorithm
The Kernel Rough K-Means Algorithm
Background:
Clustering is one of the most important data mining methods. The k-means
(c-means ) and its derivative methods are the hotspot in the field of clustering research in re...
Image clustering using exponential discriminant analysis
Image clustering using exponential discriminant analysis
Local learning based image clustering models are usually employed to deal with images sampled from the non‐linear manifold. Recently, linear discriminant analysis (LDA) based vario...
Optimizing machine learning techniques for genomics clustering
Optimizing machine learning techniques for genomics clustering
Optimisation des techniques d’apprentissage automatique pour le clustering génomique
Dans le domaine de la bioinformatique, le clustering est une technique efficace...
How suitable are clustering methods for functional annotation of proteins?
How suitable are clustering methods for functional annotation of proteins?
Abstract
The advent of affordable high-throughput genome sequencing has drastically expanded protein sequence databases, necessitating the development of computatio...
An Ensemble Clustering Method Based on Several Different Clustering Methods
An Ensemble Clustering Method Based on Several Different Clustering Methods
Abstract
As an unsupervised learning method, clustering is done to find natural groupings of patterns, points, or objects. In clustering algorithms, an important problem is...
Intelligent clustering using moth flame optimizer for vehicular ad hoc networks
Intelligent clustering using moth flame optimizer for vehicular ad hoc networks
Vehicular ad hoc networks consist of access points for communication, transmission, and collecting information of nodes and environment for managing traffic loads. Clustering can b...
A Proposed Clustering Algorithm for Efficient Clustering of High-Dimensional Data
A Proposed Clustering Algorithm for Efficient Clustering of High-Dimensional Data
To partition transaction data values, clustering algorithms are used. To analyse the relationships between transactions, similarity measures are utilized. Similarity models based o...

