Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Chinese-Uyghur Bilingual Lexicon Extraction Based on Weak Supervision

View through CrossRef
Bilingual lexicon extraction is useful, especially for low-resource languages that can leverage from high-resource languages. The Uyghur language is a derivative language, and its language resources are scarce and noisy. Moreover, it is difficult to find a bilingual resource to utilize the linguistic knowledge of other large resource languages, such as Chinese or English. There is little related research on unsupervised extraction for the Chinese-Uyghur languages, and the existing methods mainly focus on term extraction methods based on translated parallel corpora. Accordingly, unsupervised knowledge extraction methods are effective, especially for the low-resource languages. This paper proposes a method to extract a Chinese-Uyghur bilingual dictionary by combining the inter-word relationship matrix mapped by the neural network cross-language word embedding vector. A seed dictionary is used as a weak supervision signal. A small Chinese-Uyghur parallel data resource is used to map the multilingual word vectors into a unified vector space. As the word-particles of these two languages are not well-coordinated, stems are used as the main linguistic particles. The strong inter-word semantic relationship of word vectors is used to associate Chinese-Uyghur semantic information. Two retrieval indicators, such as nearest neighbor retrieval and cross-domain similarity local scaling, are used to calculate similarity to extract bilingual dictionaries. The experimental results show that the accuracy of the Chinese-Uyghur bilingual dictionary extraction method proposed in this paper is improved to 65.06%. This method helps to improve Chinese-Uyghur machine translation, automatic knowledge extraction, and multilingual translations.
Title: Chinese-Uyghur Bilingual Lexicon Extraction Based on Weak Supervision
Description:
Bilingual lexicon extraction is useful, especially for low-resource languages that can leverage from high-resource languages.
The Uyghur language is a derivative language, and its language resources are scarce and noisy.
Moreover, it is difficult to find a bilingual resource to utilize the linguistic knowledge of other large resource languages, such as Chinese or English.
There is little related research on unsupervised extraction for the Chinese-Uyghur languages, and the existing methods mainly focus on term extraction methods based on translated parallel corpora.
Accordingly, unsupervised knowledge extraction methods are effective, especially for the low-resource languages.
This paper proposes a method to extract a Chinese-Uyghur bilingual dictionary by combining the inter-word relationship matrix mapped by the neural network cross-language word embedding vector.
A seed dictionary is used as a weak supervision signal.
A small Chinese-Uyghur parallel data resource is used to map the multilingual word vectors into a unified vector space.
As the word-particles of these two languages are not well-coordinated, stems are used as the main linguistic particles.
The strong inter-word semantic relationship of word vectors is used to associate Chinese-Uyghur semantic information.
Two retrieval indicators, such as nearest neighbor retrieval and cross-domain similarity local scaling, are used to calculate similarity to extract bilingual dictionaries.
The experimental results show that the accuracy of the Chinese-Uyghur bilingual dictionary extraction method proposed in this paper is improved to 65.
06%.
This method helps to improve Chinese-Uyghur machine translation, automatic knowledge extraction, and multilingual translations.

Related Results

Loanwords in Uyghur in a Historical and Socio-Cultural Perspective
Loanwords in Uyghur in a Historical and Socio-Cultural Perspective
Modern Uyghur is one of the Eastern Turkic languages which serves as the regional lingua franca and spoken by the Uyghur people living in the Xinjiang Uyghur Autonomous Region (XUA...
Računalno potpomognuto usmjeravanje kod dvojezičnih govornika
Računalno potpomognuto usmjeravanje kod dvojezičnih govornika
This thesis investigates whether modern computer models can confirm how people encounter words and then use these findings in didactics. In recent years, computers have been used i...
UYGUR LITERATURE IN INDEPENDENT KAZAKHSTAN: PAST AND PRESENT
UYGUR LITERATURE IN INDEPENDENT KAZAKHSTAN: PAST AND PRESENT
The article provides a brief overview of the socio-political periods of the development of Uyghur literature in Kazakhstan, some of its thematic and genre features. The main goal o...
Taranchis, Kashgaris, and the ‘Uyghur Question’ in Soviet Central Asia
Taranchis, Kashgaris, and the ‘Uyghur Question’ in Soviet Central Asia
AbstractUp till now, the problem of Uyghur identity construction has been studied from an almost exclusively anthropological perspective. Little Western research has been done on t...
Muqam transmission in the Xinjiang Uyghur autonomous region: negotiating artistic individuality in present-day Uyghur muqam performance practices
Muqam transmission in the Xinjiang Uyghur autonomous region: negotiating artistic individuality in present-day Uyghur muqam performance practices
The Uyghur Muqam of Xinjiang, also known as the Art of Chinese Xinjiang Uyghur Muqam (Ch. 中国新 疆维吾尔木卡姆艺术), was inscribed on the UNESCO Representative List of the Masterpieces of the...
Uyghur Historiography
Uyghur Historiography
Abstract The history of Uyghurs, the Turkic Muslim people indigenous to the Xinjiang Uyghur Autonomous Region of the People’s Republic of China, also known as Eas...
Experimental Study on the Acoustic Characteristics of “Similar” Vowels in Mandarin Learners
Experimental Study on the Acoustic Characteristics of “Similar” Vowels in Mandarin Learners
Abstract The rapid globalization of regions with different languages necessitates more advanced non-native tongue proficiency among individuals who need to commun...
William Colenso’s Māori-English Lexicon
William Colenso’s Māori-English Lexicon
<p>William Colenso, one of Victorian New Zealand’s most accomplished polymaths, is remembered best as a printer, a defrocked missionary, botanist, and politician. Up till now...

Back to Top