Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Transformer Based Abstractive Text Summarization for Hausa: A Low Resource NLP Approach

View through CrossRef
The rapid proliferation of digital text across online platforms has created an urgent need for automated summarization systems capable of distilling essential information efficiently. For languages like Hausa, which occupies a prominent position among African languages yet remains significantly underrepresented in natural language processing (NLP) research, this need is especially pronounced. This paper presents an investigation into the application of the mT5_multilingual_XLSum transformer model for abstractive text summarization of Hausa language documents. Drawing on the GEM/xlsum Hausa dataset, this study constructs a comprehensive preprocessing and training pipeline that employs the T5Tokenizer and applies synonym-based data augmentation strategies to enrich the training corpus. Performance is evaluated through ROUGE metrics (ROUGE-1, ROUGE-2, ROUGE-L, and ROUGE-LSUM), with results demonstrating that the transformer-based approach substantially outperforms conventional deep learning baselines, including Long Short-Term Memory (LSTM) and Recurrent Neural Network (RNN) models. The mT5_multilingual_XLSum model achieves ROUGE-1, ROUGE-2, ROUGE-L, and ROUGE-LSUM scores of 63.73%, 41.54%, 51.79%, and 48.81%, respectively. These outcomes underscore the efficacy of transformer architectures in low-resource language summarization and open pathways for improved digital inclusion of Hausa speakers.
Title: Transformer Based Abstractive Text Summarization for Hausa: A Low Resource NLP Approach
Description:
The rapid proliferation of digital text across online platforms has created an urgent need for automated summarization systems capable of distilling essential information efficiently.
For languages like Hausa, which occupies a prominent position among African languages yet remains significantly underrepresented in natural language processing (NLP) research, this need is especially pronounced.
This paper presents an investigation into the application of the mT5_multilingual_XLSum transformer model for abstractive text summarization of Hausa language documents.
Drawing on the GEM/xlsum Hausa dataset, this study constructs a comprehensive preprocessing and training pipeline that employs the T5Tokenizer and applies synonym-based data augmentation strategies to enrich the training corpus.
Performance is evaluated through ROUGE metrics (ROUGE-1, ROUGE-2, ROUGE-L, and ROUGE-LSUM), with results demonstrating that the transformer-based approach substantially outperforms conventional deep learning baselines, including Long Short-Term Memory (LSTM) and Recurrent Neural Network (RNN) models.
The mT5_multilingual_XLSum model achieves ROUGE-1, ROUGE-2, ROUGE-L, and ROUGE-LSUM scores of 63.
73%, 41.
54%, 51.
79%, and 48.
81%, respectively.
These outcomes underscore the efficacy of transformer architectures in low-resource language summarization and open pathways for improved digital inclusion of Hausa speakers.

Related Results

A Comprehensive Study of Text Summarization with Advent of Large Language Models
A Comprehensive Study of Text Summarization with Advent of Large Language Models
Introduction: Communication is at the heart of the human race. With the growth of social media and other communication platforms, the globe is now connected at a single click. Peop...
Automatic summarization of Malayalam documents using clause identification method
Automatic summarization of Malayalam documents using clause identification method
<span>Text summarization is an active research area in the field of natural language processing. Huge amount of information in the internet necessitates the development of au...
Abstractive and Extractive Approaches for Summarizing Multi-document Travel Reviews
Abstractive and Extractive Approaches for Summarizing Multi-document Travel Reviews
Travel reviews offer insights into users' experiences at places they have visited, including hotels, restaurants, and tourist attractions. Reviews are a type of multidocument, wher...
Hausa
Hausa
With an estimated population of up to 50 million, Hausa make up one of the largest people groups practicing Islam. Despite settlement of today’s Hausaland in the central Sudan by t...
A Systematic Review and Experimental Evaluation of Classical and Transformer-Based Models for Urdu Abstractive Text Summarization
A Systematic Review and Experimental Evaluation of Classical and Transformer-Based Models for Urdu Abstractive Text Summarization
The rapid growth of digital content in Urdu has created an urgent need for effective automatic text summarization (ATS) systems. While extractive methods have been widely studied, ...
Transformer-Based Model for Named Entity Recognition in Hausa Language Texts
Transformer-Based Model for Named Entity Recognition in Hausa Language Texts
Named Entity Recognition (NER) is a fundamental problem in Natural Language Processing (NLP) that focuses on identifying and classifying textual entities such as people, places, or...
Evaluating Classical and Transformer-Based Models for Urdu Abstractive Text Summarization: A Systematic Review
Evaluating Classical and Transformer-Based Models for Urdu Abstractive Text Summarization: A Systematic Review
The rapid growth of digital content in Urdu has created an urgent need for effective automatic text summarization (ATS) systems. While extractive methods have been widely studied, ...
Abstractive text summarization of low-resourced languages using deep learning
Abstractive text summarization of low-resourced languages using deep learning
Background Humans must be able to cope with the huge amounts of information produced by the information technology revolution. As a result, automatic text summarizat...

Back to Top