Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

A Hybrid Arabic text summarization Approach based on Seq-to-seq and Transformer

View through CrossRef
Abstract Text summarization is essential in natural language processing as the data volume increases quickly. Therefore, the user needs to summarize that data into a meaningful text in a short time. There are two common methods of text summarization: extractive and abstractive. There are many efforts to summarize Latin texts. However, summarizing Arabic texts is challenging for many reasons, including the language’s complexity, structure, and morphology. Also, there is a need for benchmark data sources and a gold standard Arabic evaluation metrics summary. Thus, the contribution of this paper is multi-fold: First, the paper proposes a hybrid approach consisting of a Modified Sequence-To-Sequence (MSTS) model and a transformer-based model. The seq-to-seq-based model is modified by adding multi-layer encoders and a one-layer decoder to its structure. The output of the MSTS model is the extractive summarization. To generate the abstractive summarization, the extractive summarization is manipulated by a transformer-based model. Second, it introduces a new Arabic benchmark dataset, called the HASD, which includes 43k articles with their extractive and abstractive summaries. Third, this work modifies the well-known extractive EASC benchmarks by adding to each text its abstractive summarization. Finally, this paper proposes a new measure called the Arabic-rouge measure for the abstractive summary depending on structure and similarity between words. The proposed method is tested using the proposed HASD and Modified EASC benchmarks and evaluated using Rouge, Bleu, and Arabic Rouge. The experimental results show satisfactory results compared to state-of-the-art methods.
Title: A Hybrid Arabic text summarization Approach based on Seq-to-seq and Transformer
Description:
Abstract Text summarization is essential in natural language processing as the data volume increases quickly.
Therefore, the user needs to summarize that data into a meaningful text in a short time.
There are two common methods of text summarization: extractive and abstractive.
There are many efforts to summarize Latin texts.
However, summarizing Arabic texts is challenging for many reasons, including the language’s complexity, structure, and morphology.
Also, there is a need for benchmark data sources and a gold standard Arabic evaluation metrics summary.
Thus, the contribution of this paper is multi-fold: First, the paper proposes a hybrid approach consisting of a Modified Sequence-To-Sequence (MSTS) model and a transformer-based model.
The seq-to-seq-based model is modified by adding multi-layer encoders and a one-layer decoder to its structure.
The output of the MSTS model is the extractive summarization.
To generate the abstractive summarization, the extractive summarization is manipulated by a transformer-based model.
Second, it introduces a new Arabic benchmark dataset, called the HASD, which includes 43k articles with their extractive and abstractive summaries.
Third, this work modifies the well-known extractive EASC benchmarks by adding to each text its abstractive summarization.
Finally, this paper proposes a new measure called the Arabic-rouge measure for the abstractive summary depending on structure and similarity between words.
The proposed method is tested using the proposed HASD and Modified EASC benchmarks and evaluated using Rouge, Bleu, and Arabic Rouge.
The experimental results show satisfactory results compared to state-of-the-art methods.

Related Results

A Comprehensive Study of Text Summarization with Advent of Large Language Models
A Comprehensive Study of Text Summarization with Advent of Large Language Models
Introduction: Communication is at the heart of the human race. With the growth of social media and other communication platforms, the globe is now connected at a single click. Peop...
Automatic Text Summarization Berdasarkan Pendekatan Statistika pada Dokumen Berbahasa Indonesia
Automatic Text Summarization Berdasarkan Pendekatan Statistika pada Dokumen Berbahasa Indonesia
Abstract—Propelled by the modern technological innovations data and text will be more abundant throughout the year. With this much text, automatic text summarization is needed now ...
Sleep Habits and Occurrence of Lowback Pain among Craftsmen
Sleep Habits and Occurrence of Lowback Pain among Craftsmen
<span style="color: #000000; font-family: Verdana, Arial, Helvetica, sans-serif; font-size: 10px; font-style: normal; font-variant-ligatures: normal; font-variant-caps: normal; ...
Sleep Habits and Occurrence of Lowback Pain among Craftsmen
Sleep Habits and Occurrence of Lowback Pain among Craftsmen
<span style="color: #000000; font-family: Verdana, Arial, Helvetica, sans-serif; font-size: 10px; font-style: normal; font-variant-ligatures: normal; font-variant-caps: normal; ...
Automatic Load Sharing of Transformer
Automatic Load Sharing of Transformer
Transformer plays a major role in the power system. It works 24 hours a day and provides power to the load. The transformer is excessive full, its windings are overheated which lea...
IMPROVING ARABIC TEXT SUMMARIZATION USING ADVANCED PRE-TRAINED MODELS
IMPROVING ARABIC TEXT SUMMARIZATION USING ADVANCED PRE-TRAINED MODELS
The exponential growth of online content has made the task of locating specific information increasingly challenging, thus highlighting the necessity for automated text summarizati...
Text Summarization Techniques Using Natural Language Processing: A Systematic Literature Review
Text Summarization Techniques Using Natural Language Processing: A Systematic Literature Review
In recent years, data has been growing rapidly in almost every domain. Due to this excessiveness of data, there is a need for an automatic text summarizer that summarizes long and ...

Back to Top