Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

A PARALLEL CORPUS FOR ADVANCING ENGLISH–SANTALI NEURAL MACHINE TRANSLATION

View through CrossRef
Machine Translation (MT) poses a significant challenge in developing language corpora for low-resource languages due to their minimal digital availability. Building such corpora is essential for preserving and promoting these languages. Santali, for instance, has very limited representation across online resources, and no proper translation tools including Google Translate exist for it. Developing a translation framework under such constraints is particularly difficult, as issues like low translation accuracy and heavy computational requirements arise. To overcome these limitations, the proposed MT system employs EnSanCorp, an English-Santali parallel corpus designed to facilitate Neural Machine Translation (NMT). EnSanCorp is created using multiple approaches, such as web-based parallel data extraction and optical character recognition (OCR) applied to scanned documents. The OCR-based method also demonstrates its usefulness for building corpora of other low-resource languages lacking online data. EnSanCorp currently contains 5,930 aligned sentences, 39,646 English tokens, and 39,936 Santali tokens, making it the most extensive English-Santali corpus available for research and non-commercial purposes. Evaluation results show that the Bilingual Evaluation Understudy (BLEU) scores for Statistical Machine Translation (SMT) and NMT vary across word and sentence levels: for word pairs, the scores are 0.04 (SMT) and 1.10 (NMT); for sentence pairs, 1.15 (SMT) and 7.20 (NMT). The overall BLEU scores achieved are 0.05 for SMT and 3.10 for NMT.
Title: A PARALLEL CORPUS FOR ADVANCING ENGLISH–SANTALI NEURAL MACHINE TRANSLATION
Description:
Machine Translation (MT) poses a significant challenge in developing language corpora for low-resource languages due to their minimal digital availability.
Building such corpora is essential for preserving and promoting these languages.
Santali, for instance, has very limited representation across online resources, and no proper translation tools including Google Translate exist for it.
Developing a translation framework under such constraints is particularly difficult, as issues like low translation accuracy and heavy computational requirements arise.
To overcome these limitations, the proposed MT system employs EnSanCorp, an English-Santali parallel corpus designed to facilitate Neural Machine Translation (NMT).
EnSanCorp is created using multiple approaches, such as web-based parallel data extraction and optical character recognition (OCR) applied to scanned documents.
The OCR-based method also demonstrates its usefulness for building corpora of other low-resource languages lacking online data.
EnSanCorp currently contains 5,930 aligned sentences, 39,646 English tokens, and 39,936 Santali tokens, making it the most extensive English-Santali corpus available for research and non-commercial purposes.
Evaluation results show that the Bilingual Evaluation Understudy (BLEU) scores for Statistical Machine Translation (SMT) and NMT vary across word and sentence levels: for word pairs, the scores are 0.
04 (SMT) and 1.
10 (NMT); for sentence pairs, 1.
15 (SMT) and 7.
20 (NMT).
The overall BLEU scores achieved are 0.
05 for SMT and 3.
10 for NMT.

Related Results

Žanrovska analiza pomorskopravnih tekstova i ostvarenje prijevodnih univerzalija u njihovim prijevodima s engleskoga jezika
Žanrovska analiza pomorskopravnih tekstova i ostvarenje prijevodnih univerzalija u njihovim prijevodima s engleskoga jezika
Genre implies formal and stylistic conventions of a particular text type, which inevitably affects the translation process. This „force of genre bias“ (Prieto Ramos, 2014) has been...
Aviation English - A global perspective: analysis, teaching, assessment
Aviation English - A global perspective: analysis, teaching, assessment
This e-book brings together 13 chapters written by aviation English researchers and practitioners settled in six different countries, representing institutions and universities fro...
Evaluating IndicTrans2 and ByT5 for English-Santali Machine Translation Using the Ol Chiki Script
Evaluating IndicTrans2 and ByT5 for English-Santali Machine Translation Using the Ol Chiki Script
In this study, we examine and evaluate two multilingual NMT models, IndicTrans2 and ByT5, for English-Santali bidirectional translation using the Ol Chiki script. The models are tr...
Phrase Based Statistical Machine Translation Javanese-Indonesian
Phrase Based Statistical Machine Translation Javanese-Indonesian
This research aims to produce a statistical machine translation that can be implemented to perform Javanese-Indonesian translation and to know the influence of the main data source...
Japanese translation teaching corpus based on bilingual non parallel data model
Japanese translation teaching corpus based on bilingual non parallel data model
In recent years, with the development of Internet and intelligent technology, Japanese translation teaching has gradually explored a new teaching mode. Under the guidance of natura...
The neural machine translation models for the low-resource Kazakh–English language pair
The neural machine translation models for the low-resource Kazakh–English language pair
The development of the machine translation field was driven by people’s need to communicate with each other globally by automatically translating words, sentences, and texts from o...
Witchcraft and Indigeneity in Santali Writings
Witchcraft and Indigeneity in Santali Writings
Despite the dialectal variation of the Santali language, which is spoken across five different states (Bengal, Jharkhand, Odisha, Bihar, and Assam), many Santali speakers consider ...
Editorial Introduction: Translating the Future: Exploring the Impact of Technology and AI on Modern Translation Studies
Editorial Introduction: Translating the Future: Exploring the Impact of Technology and AI on Modern Translation Studies
We have entered into the era of artificial intelligence, neural machine translation, and especially large language models which have dramatically changed the landscape of human tra...

Back to Top