Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Effective preprocessing based neural machine translation for English to Telugu cross-language information retrieval

View through CrossRef
<span id="docs-internal-guid-5b69f940-7fff-f443-1f09-a00e5e983714"><span>In cross-language information retrieval (CLIR), the neural machine translation (NMT) plays a vital role. CLIR retrieves the information written in a language which is different from the user's query language. In CLIR, the main concern is to translate the user query from the source language to the target language. NMT is useful for translating the data from one language to another. NMT has better accuracy for different languages like English to German and so-on. In this paper, NMT has applied for translating English to Indian languages, especially for Telugu. Besides NMT, an effort is also made to improve accuracy by applying effective preprocessing mechanism. The role of effective preprocessing in improving accuracy will be less but countable. Machine translation (MT) is a data-driven approach where parallel corpus will act as input in MT. NMT requires a massive amount of parallel corpus for performing the translation. Building an English - Telugu parallel corpus is costly because they are resource-poor languages. Different mechanisms are available for preparing the parallel corpus. The major issue in preparing parallel corpus is data replication that is handled during preprocessing. The other issue in machine translation is the out-of-vocabulary (OOV) problem. Earlier dictionaries are used to handle OOV problems. To overcome this problem the rare words are segmented into sequences of subwords during preprocessing. The parameters like accuracy, perplexity, cross-entropy and BLEU scores shows better translation quality for NMT with effective preprocessing.</span></span>
Title: Effective preprocessing based neural machine translation for English to Telugu cross-language information retrieval
Description:
<span id="docs-internal-guid-5b69f940-7fff-f443-1f09-a00e5e983714"><span>In cross-language information retrieval (CLIR), the neural machine translation (NMT) plays a vital role.
CLIR retrieves the information written in a language which is different from the user's query language.
In CLIR, the main concern is to translate the user query from the source language to the target language.
NMT is useful for translating the data from one language to another.
NMT has better accuracy for different languages like English to German and so-on.
In this paper, NMT has applied for translating English to Indian languages, especially for Telugu.
Besides NMT, an effort is also made to improve accuracy by applying effective preprocessing mechanism.
The role of effective preprocessing in improving accuracy will be less but countable.
Machine translation (MT) is a data-driven approach where parallel corpus will act as input in MT.
NMT requires a massive amount of parallel corpus for performing the translation.
Building an English - Telugu parallel corpus is costly because they are resource-poor languages.
Different mechanisms are available for preparing the parallel corpus.
The major issue in preparing parallel corpus is data replication that is handled during preprocessing.
The other issue in machine translation is the out-of-vocabulary (OOV) problem.
Earlier dictionaries are used to handle OOV problems.
To overcome this problem the rare words are segmented into sequences of subwords during preprocessing.
The parameters like accuracy, perplexity, cross-entropy and BLEU scores shows better translation quality for NMT with effective preprocessing.
</span></span>.

Related Results

Hubungan Perilaku Pola Makan dengan Kejadian Anak Obesitas
Hubungan Perilaku Pola Makan dengan Kejadian Anak Obesitas
<p><em><span style="font-size: 11.0pt; font-family: 'Times New Roman',serif; mso-fareast-font-family: 'Times New Roman'; mso-ansi-language: EN-US; mso-fareast-langua...
Aviation English - A global perspective: analysis, teaching, assessment
Aviation English - A global perspective: analysis, teaching, assessment
This e-book brings together 13 chapters written by aviation English researchers and practitioners settled in six different countries, representing institutions and universities fro...
Učinak poučavanja razrednomu jeziku u izobrazbi nastavnika njemačkoga
Učinak poučavanja razrednomu jeziku u izobrazbi nastavnika njemačkoga
The actual use of classroom language is principally limited to the classroom environment. As far as foreign language learning is concerned, the classroom often turns out to be the ...
An Efficient Deep Learning Model with Interrelated Tagging Prototype with Segmentation for Telugu Optical Character Recognition
An Efficient Deep Learning Model with Interrelated Tagging Prototype with Segmentation for Telugu Optical Character Recognition
More than 66 million people in India speak Telugu, a language that dates back thousands of years and is widely spoken in South India. There has not been much progress reported on t...
IIST BCI Dataset-4 for Selected 100 Telugu words
IIST BCI Dataset-4 for Selected 100 Telugu words
To overcome the challenges faced by people with neurodegenerative diseases, Brain-Computer Interface (BCI) systems must make use of datasets relevant to patient's spoken languages....
Telugu Dependency Treebank
Telugu Dependency Treebank
We discuss Telugu Language and Treebanks briefly in this work. Initially, we'll go over the Telugu language briefly. The paninian grammatical model utilized for Telugu dependency r...
IIST BCI Dataset-8 for Selected Common Telugu words
IIST BCI Dataset-8 for Selected Common Telugu words
Brain-Computer Interface (BCI) systems require the usage of datasets corresponding to the spoken languages of patients in order to serve them with neurodegenerative illnesses.There...
Žanrovska analiza pomorskopravnih tekstova i ostvarenje prijevodnih univerzalija u njihovim prijevodima s engleskoga jezika
Žanrovska analiza pomorskopravnih tekstova i ostvarenje prijevodnih univerzalija u njihovim prijevodima s engleskoga jezika
Genre implies formal and stylistic conventions of a particular text type, which inevitably affects the translation process. This „force of genre bias“ (Prieto Ramos, 2014) has been...

Back to Top