Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

A Robust Morpheme Sequence and Convolutional Neural Network-Based Uyghur and Kazakh Short Text Classification

View through CrossRef
In this paper, based on the multilingual morphological analyzer, we researched the similar low-resource languages, Uyghur and Kazakh, short text classification. Generally, the online linguistic resources of these languages are noisy. So a preprocessing is necessary and can significantly improve the accuracy. Uyghur and Kazakh are the languages with derivational morphology, in which words are coined by stems concatenated with suffixes. Usually, terms are used as the representation of text content while excluding functional parts as stop words in these languages. By extracting stems we can collect necessary terms and exclude stop words. Morpheme segmentation tool can split text into morphemes with 95% high reliability. After preparing both word- and morpheme-based training text corpora, we apply convolutional neural network (CNN) as a feature selection and text classification algorithm to perform text classification tasks. Experimental results show that the morpheme-based approach outperformed the word-based approach. Word embedding technique is frequently used in text representation both in the framework of neural networks and as a value expression, and can map language units into a sequential vector space based on context, and it is a natural way to extract and predict out-of-vocabulary (OOV) from context information. Multilingual morphological analysis has provided a convenient way for processing tasks of low resource languages like Uyghur and Kazakh.
Title: A Robust Morpheme Sequence and Convolutional Neural Network-Based Uyghur and Kazakh Short Text Classification
Description:
In this paper, based on the multilingual morphological analyzer, we researched the similar low-resource languages, Uyghur and Kazakh, short text classification.
Generally, the online linguistic resources of these languages are noisy.
So a preprocessing is necessary and can significantly improve the accuracy.
Uyghur and Kazakh are the languages with derivational morphology, in which words are coined by stems concatenated with suffixes.
Usually, terms are used as the representation of text content while excluding functional parts as stop words in these languages.
By extracting stems we can collect necessary terms and exclude stop words.
Morpheme segmentation tool can split text into morphemes with 95% high reliability.
After preparing both word- and morpheme-based training text corpora, we apply convolutional neural network (CNN) as a feature selection and text classification algorithm to perform text classification tasks.
Experimental results show that the morpheme-based approach outperformed the word-based approach.
Word embedding technique is frequently used in text representation both in the framework of neural networks and as a value expression, and can map language units into a sequential vector space based on context, and it is a natural way to extract and predict out-of-vocabulary (OOV) from context information.
Multilingual morphological analysis has provided a convenient way for processing tasks of low resource languages like Uyghur and Kazakh.

Related Results

Uyghur–Kazakh–Kirghiz Text Keyword Extraction Based on Morpheme Segmentation
Uyghur–Kazakh–Kirghiz Text Keyword Extraction Based on Morpheme Segmentation
In this study, based on a morpheme segmentation framework, we researched a text keyword extraction method for Uyghur, Kazakh and Kirghiz languages, which have similar grammatical a...
Loanwords in Uyghur in a Historical and Socio-Cultural Perspective
Loanwords in Uyghur in a Historical and Socio-Cultural Perspective
Modern Uyghur is one of the Eastern Turkic languages which serves as the regional lingua franca and spoken by the Uyghur people living in the Xinjiang Uyghur Autonomous Region (XUA...
UYGUR LITERATURE IN INDEPENDENT KAZAKHSTAN: PAST AND PRESENT
UYGUR LITERATURE IN INDEPENDENT KAZAKHSTAN: PAST AND PRESENT
The article provides a brief overview of the socio-political periods of the development of Uyghur literature in Kazakhstan, some of its thematic and genre features. The main goal o...
LEXICAL PARADIGMATICS OF OLD UYGHUR AND KAZAKH LANGUAGES
LEXICAL PARADIGMATICS OF OLD UYGHUR AND KAZAKH LANGUAGES
This article examines the lexical paradigms of the Old Uyghur and Kazakh languages. Homonyms, synonyms, and antonyms of the two related languages – ancient and modern – are compare...
Sleep Habits and Occurrence of Lowback Pain among Craftsmen
Sleep Habits and Occurrence of Lowback Pain among Craftsmen
<span style="color: #000000; font-family: Verdana, Arial, Helvetica, sans-serif; font-size: 10px; font-style: normal; font-variant-ligatures: normal; font-variant-caps: normal; ...
Sleep Habits and Occurrence of Lowback Pain among Craftsmen
Sleep Habits and Occurrence of Lowback Pain among Craftsmen
<span style="color: #000000; font-family: Verdana, Arial, Helvetica, sans-serif; font-size: 10px; font-style: normal; font-variant-ligatures: normal; font-variant-caps: normal; ...
DEIKSIS PERSONA DALAM BAHASA MUNA
DEIKSIS PERSONA DALAM BAHASA MUNA
Abstract : The purpose of this study is to describe the form and meaning of the word deiksis persona in the Muna language. This type of research is a qualitative descriptive. Quali...

Back to Top