Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

IMPROVING ARABIC TEXT SUMMARIZATION USING ADVANCED PRE-TRAINED MODELS

View through CrossRef
The exponential growth of online content has made the task of locating specific information increasingly challenging, thus highlighting the necessity for automated text summarization. Deep learning techniques, particularly neural abstractive models such as Seq2Seq, have emerged as prominent solutions to this issue. The utilization of pre-trained models, including GPT, BERT, BART, and T5, has notably enhanced the quality of text summarization by addressing key aspects such as saliency, fluency, and semantic coherence. However, while these advancements have greatly benefited the English language, there remains a significant gap in support for low-resource languages. To bridge this gap, monolingual BERT and multilingual Seq2Seq models have been developed, enabling the application of state-of-the-art summarization techniques to languages like Arabic. Our research capitalizes on pre-trained Seq2Seq models to achieve superior results in Arabic text summarization tasks by leveraging datasets such as XLSum and Hindawi Books. Notably, the titles within these datasets serve as robust benchmarks for evaluating the effectiveness of our summarization techniques, underscoring the importance of high-quality input data. Additionally, one of our key contributions lies in the implementation of fine-tuning for the model using reinforcement learning. This innovative approach enhances the adaptability and performance of the model. Our findings indicate that monolingual BERT models outperform their multilingual counterparts, yielding a notable 2.4% increase in Rouge scores, further improving the quality of text summarization in Arabic. Our study encompasses cross-dataset evaluations, exploration of various text generation methodologies, and in-depth preprocessing analysis tailored specifically for Arabic text. By presenting a comprehensive approach to address the challenges in Arabic text summarization, our study contributes to the advancement of the field and underscores the significance of supporting low-resource languages in natural language processing tasks.
Title: IMPROVING ARABIC TEXT SUMMARIZATION USING ADVANCED PRE-TRAINED MODELS
Description:
The exponential growth of online content has made the task of locating specific information increasingly challenging, thus highlighting the necessity for automated text summarization.
Deep learning techniques, particularly neural abstractive models such as Seq2Seq, have emerged as prominent solutions to this issue.
The utilization of pre-trained models, including GPT, BERT, BART, and T5, has notably enhanced the quality of text summarization by addressing key aspects such as saliency, fluency, and semantic coherence.
However, while these advancements have greatly benefited the English language, there remains a significant gap in support for low-resource languages.
To bridge this gap, monolingual BERT and multilingual Seq2Seq models have been developed, enabling the application of state-of-the-art summarization techniques to languages like Arabic.
Our research capitalizes on pre-trained Seq2Seq models to achieve superior results in Arabic text summarization tasks by leveraging datasets such as XLSum and Hindawi Books.
Notably, the titles within these datasets serve as robust benchmarks for evaluating the effectiveness of our summarization techniques, underscoring the importance of high-quality input data.
Additionally, one of our key contributions lies in the implementation of fine-tuning for the model using reinforcement learning.
This innovative approach enhances the adaptability and performance of the model.
Our findings indicate that monolingual BERT models outperform their multilingual counterparts, yielding a notable 2.
4% increase in Rouge scores, further improving the quality of text summarization in Arabic.
Our study encompasses cross-dataset evaluations, exploration of various text generation methodologies, and in-depth preprocessing analysis tailored specifically for Arabic text.
By presenting a comprehensive approach to address the challenges in Arabic text summarization, our study contributes to the advancement of the field and underscores the significance of supporting low-resource languages in natural language processing tasks.

Related Results

A Comprehensive Study of Text Summarization with Advent of Large Language Models
A Comprehensive Study of Text Summarization with Advent of Large Language Models
Introduction: Communication is at the heart of the human race. With the growth of social media and other communication platforms, the globe is now connected at a single click. Peop...
Sleep Habits and Occurrence of Lowback Pain among Craftsmen
Sleep Habits and Occurrence of Lowback Pain among Craftsmen
<span style="color: #000000; font-family: Verdana, Arial, Helvetica, sans-serif; font-size: 10px; font-style: normal; font-variant-ligatures: normal; font-variant-caps: normal; ...
Sleep Habits and Occurrence of Lowback Pain among Craftsmen
Sleep Habits and Occurrence of Lowback Pain among Craftsmen
<span style="color: #000000; font-family: Verdana, Arial, Helvetica, sans-serif; font-size: 10px; font-style: normal; font-variant-ligatures: normal; font-variant-caps: normal; ...
Automatic Text Summarization Berdasarkan Pendekatan Statistika pada Dokumen Berbahasa Indonesia
Automatic Text Summarization Berdasarkan Pendekatan Statistika pada Dokumen Berbahasa Indonesia
Abstract—Propelled by the modern technological innovations data and text will be more abundant throughout the year. With this much text, automatic text summarization is needed now ...
Text Summarization Techniques Using Natural Language Processing: A Systematic Literature Review
Text Summarization Techniques Using Natural Language Processing: A Systematic Literature Review
In recent years, data has been growing rapidly in almost every domain. Due to this excessiveness of data, there is a need for an automatic text summarizer that summarizes long and ...
Advancements in Automatic Text Summarization using Natural Language Processing
Advancements in Automatic Text Summarization using Natural Language Processing
With the rapid expansion of data across various domains, the need for automated text summarization has become increasingly crucial. Given the overwhelming volu...
Bounds on the sum of broadcast domination number and strong metric dimension of graphs
Bounds on the sum of broadcast domination number and strong metric dimension of graphs
Let [Formula: see text] be a connected graph of order at least two with vertex set [Formula: see text]. For [Formula: see text], let [Formula: see text] denote the length of an [Fo...

Back to Top