Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Speech Representation System For The Karakalpak Automatic Speech Recognition

View through CrossRef
Abstract Recently, automatic speech recognition (ASR) has moved from deeplearning systems toward end-to-end (E2E) neural architectures. E2E models obtain strong results for high-resource languages, but training reliable systems remains difficult for low-resource languages because labeled speech is expensive to collect and verify. Self-supervised learning (SSL) has reduced this dependency by learning speech representations from large unlabeled corpora and then adapting them to a target language with a smaller amount of transcribed data. In this paper, we present a Karakalpak ASR system based on multilingual speech representations. The work follows the wav2vec 2.0 XLS-R approach, and the Karakalpak-specific acoustic model is trained and fine-tuned by the authors using the Karakalpak Speech Corpus (KSC). KSC is a publicly available benchmark speech-to-text dataset collected by the authors and released through Mendeley Data; it contains 50 hours of predominantly read speech from native Karakalpak speakers. The ASR model is fine-tuned with various KSC data setups, and the effect of the amount of labeled speech on recognition quality is discussed. In addition, word-level n-gram language models are used to improve decoding under low-resource conditions. Recognition error analysis is provided in terms of substitution, insertion, and deletion errors. The 50-hour KSC-trained wav2vec2 XLS-R model gives 21.10\% WER and 4.32\% CER on a held-out test set, and the experimental design treats this result as a reference point for further benchmarking.
Title: Speech Representation System For The Karakalpak Automatic Speech Recognition
Description:
Abstract Recently, automatic speech recognition (ASR) has moved from deeplearning systems toward end-to-end (E2E) neural architectures.
E2E models obtain strong results for high-resource languages, but training reliable systems remains difficult for low-resource languages because labeled speech is expensive to collect and verify.
Self-supervised learning (SSL) has reduced this dependency by learning speech representations from large unlabeled corpora and then adapting them to a target language with a smaller amount of transcribed data.
In this paper, we present a Karakalpak ASR system based on multilingual speech representations.
The work follows the wav2vec 2.
0 XLS-R approach, and the Karakalpak-specific acoustic model is trained and fine-tuned by the authors using the Karakalpak Speech Corpus (KSC).
KSC is a publicly available benchmark speech-to-text dataset collected by the authors and released through Mendeley Data; it contains 50 hours of predominantly read speech from native Karakalpak speakers.
The ASR model is fine-tuned with various KSC data setups, and the effect of the amount of labeled speech on recognition quality is discussed.
In addition, word-level n-gram language models are used to improve decoding under low-resource conditions.
Recognition error analysis is provided in terms of substitution, insertion, and deletion errors.
The 50-hour KSC-trained wav2vec2 XLS-R model gives 21.
10\% WER and 4.
32\% CER on a held-out test set, and the experimental design treats this result as a reference point for further benchmarking.

Related Results

PROCESSING OF CONSONANTS IN THE QARAQALPAQ LANGUAGE
PROCESSING OF CONSONANTS IN THE QARAQALPAQ LANGUAGE
One of the ancient Turkic peoples was the Karakalpak Turks. The Karakalpak language belongs to the Kipchak subgroup of the Turkic language family and is the official state language...
Fine-Tuning Whisper for Low-Resource Automatic Speech Recognition in the Karakalpak Language
Fine-Tuning Whisper for Low-Resource Automatic Speech Recognition in the Karakalpak Language
Abstract The Karakalpak language is still poorly represented in supervised speech resources, making reliable automatic speech recognition (ASR) difficult to build. ...
The Description of Nature in the Works of the Travel Genre: The Case of Karakalpak Writers
The Description of Nature in the Works of the Travel Genre: The Case of Karakalpak Writers
The subject of research in this work is the genre of travel in Karakalpak Literature. It analyzes the works of the travel genre of individual Karakalpak writers who have visited ab...
Pola Komunikasi Interpersonal Terapis Wicara pada Anak Telambat Bicara (Speech Delay)
Pola Komunikasi Interpersonal Terapis Wicara pada Anak Telambat Bicara (Speech Delay)
Abstract. The tittle of this study is Interpersonal Communication Patterns of Speech Therapists in Children with Speech Delay. Children with speech delay disorders are also social ...
Multimodal Emotion Recognition and Human Computer Interaction for AI-Driven Mental Health Support (Preprint)
Multimodal Emotion Recognition and Human Computer Interaction for AI-Driven Mental Health Support (Preprint)
BACKGROUND Mental health has become one of the most urgent global health issues of the twenty-first century. The World Health Organization (WHO) reports tha...
Automatic speech recognition in voice-speech rehabilitation effectiveness evaluation in patients after laryngectomy
Automatic speech recognition in voice-speech rehabilitation effectiveness evaluation in patients after laryngectomy
Introduction. Lost voice function compensation determines the personal and social life of laryngectomees. Automatic speech recognition and synthesis methods are...
Sociocultural Expression of National Identity of the Karakalpak People
Sociocultural Expression of National Identity of the Karakalpak People
This article briefly presents the history of the Karakalpak people, provides some historical data on the origin of the Karakalpak people, their socio-cultural characteristics, huma...

Back to Top