Javascript must be enabled to continue!
BiLSTMs and BPE for English to Telugu CLIR
View through CrossRef
A crucial component of Cross Lingual Information Retrieval (CLIR) is Neural Machine Translation (NMT). NMT performs a good job of transforming queries in the English language into Indian languages. This study focuses on the translation of English queries into Telugu. For translations, the NMT will make use of a parallel corpus. Due to a lack of resources in the Telugu language, it is exceedingly challenging to provide NMT with sizable parallel corpora. Thus, the NMT will encounter an issue known as Out of Vocabulary (OOV). Long Short-Term Memory (LSTM) with Byte Pair Encoding (BPE), which breaks up rare words into subwords and attempts to translate them to solve the OOV issue. Issues such as Named Entity Recognition (NER) continue to plague it. In sequence-to-sequence models, bidirectional LSTMs can solve certain NER challenges. Systems that need to be trained in both directions to recognize named entities can benefit from the use of Bidirectional LSTMs (BiLSTMs). The translation efficiency of NMT with BiLSTMs is significantly higher than normal LSTMs, as indicated by the accuracy metrics and Bilingual Evaluation Understudy (BLEU) score.
Science Research Society
Title: BiLSTMs and BPE for English to Telugu CLIR
Description:
A crucial component of Cross Lingual Information Retrieval (CLIR) is Neural Machine Translation (NMT).
NMT performs a good job of transforming queries in the English language into Indian languages.
This study focuses on the translation of English queries into Telugu.
For translations, the NMT will make use of a parallel corpus.
Due to a lack of resources in the Telugu language, it is exceedingly challenging to provide NMT with sizable parallel corpora.
Thus, the NMT will encounter an issue known as Out of Vocabulary (OOV).
Long Short-Term Memory (LSTM) with Byte Pair Encoding (BPE), which breaks up rare words into subwords and attempts to translate them to solve the OOV issue.
Issues such as Named Entity Recognition (NER) continue to plague it.
In sequence-to-sequence models, bidirectional LSTMs can solve certain NER challenges.
Systems that need to be trained in both directions to recognize named entities can benefit from the use of Bidirectional LSTMs (BiLSTMs).
The translation efficiency of NMT with BiLSTMs is significantly higher than normal LSTMs, as indicated by the accuracy metrics and Bilingual Evaluation Understudy (BLEU) score.
Related Results
An Efficient Deep Learning Model with Interrelated Tagging Prototype with Segmentation for Telugu Optical Character Recognition
An Efficient Deep Learning Model with Interrelated Tagging Prototype with Segmentation for Telugu Optical Character Recognition
More than 66 million people in India speak Telugu, a language that dates back thousands of years and is widely spoken in South India. There has not been much progress reported on t...
Aviation English - A global perspective: analysis, teaching, assessment
Aviation English - A global perspective: analysis, teaching, assessment
This e-book brings together 13 chapters written by aviation English researchers and practitioners settled in six different countries, representing institutions and universities fro...
IIST BCI Dataset-4 for Selected 100 Telugu words
IIST BCI Dataset-4 for Selected 100 Telugu words
To overcome the challenges faced by people with neurodegenerative diseases, Brain-Computer Interface (BCI) systems must make use of datasets relevant to patient's spoken languages....
IIST BCI Dataset-8 for Selected Common Telugu words
IIST BCI Dataset-8 for Selected Common Telugu words
Brain-Computer Interface (BCI) systems require the usage of datasets corresponding to the spoken languages of patients in order to serve them with neurodegenerative illnesses.There...
Telugu Dependency Treebank
Telugu Dependency Treebank
We discuss Telugu Language and Treebanks briefly in this work. Initially, we'll go over the Telugu language briefly. The paninian grammatical model utilized for Telugu dependency r...
Review of Prognostic Significance of Quantitative BPE Measurements
Review of Prognostic Significance of Quantitative BPE Measurements
Background/Objectives: Background parenchymal enhancement (BPE) on breast magnetic resonance imaging reflects hormonal and vascular activity of fibroglandular tissue and is studied...
Effective preprocessing based neural machine translation for English to Telugu cross-language information retrieval
Effective preprocessing based neural machine translation for English to Telugu cross-language information retrieval
<span id="docs-internal-guid-5b69f940-7fff-f443-1f09-a00e5e983714"><span>In cross-language information retrieval (CLIR), the neural machine translation (NMT) plays a vi...
A dependency parser for code-switched Telugu-English
A dependency parser for code-switched Telugu-English
This study addresses a notable gap in the availability of Code Switched Telugu-English dependency parsers. Conducting research into Code Switched data, especially Telugu-English co...

