Javascript must be enabled to continue!
Speakerdependent automatic speech recognition in Buryat
View through CrossRef
Due to the dominance of Russian in education, formal communication, and mass media the usage of Buryat is gradually decreasing. Therefore, the problem of language preservation is of high concern for the native people. In
modern conditions, in resolving this question significant consideration is given to speech technologies based on
natural language processing, artificial intelligence and other methods to digitize the Buryat language and provide access to resources online. The analysis of current resources has found no Buryattargeted technologies that work
with audio data. Therefore, the present study covers the development of speech recognition model prototype in
166
Buryat for automatic speech processing. The model uses the DeepSpeech2 architecture allowing to combine the functions of convolutional neural networks to extract and compress acoustic information relevant for recognition,
as well as connections to identify time dependencies between phones in order to develop a model of a word in the Buryat language. The percentage of incorrectly recognized phones is 11%. This value significantly exceeds the values typical for modern models. The limitation of this model is the insufficient amount of training data and the dependence on the speaker's voice. Expanding the speech sample bank, annotating the data and optimizing the model can improve the accuracy of the system and its resistance to speech variability.
Amur State University
Title: Speakerdependent automatic speech recognition in Buryat
Description:
Due to the dominance of Russian in education, formal communication, and mass media the usage of Buryat is gradually decreasing.
Therefore, the problem of language preservation is of high concern for the native people.
In
modern conditions, in resolving this question significant consideration is given to speech technologies based on
natural language processing, artificial intelligence and other methods to digitize the Buryat language and provide access to resources online.
The analysis of current resources has found no Buryattargeted technologies that work
with audio data.
Therefore, the present study covers the development of speech recognition model prototype in
166
Buryat for automatic speech processing.
The model uses the DeepSpeech2 architecture allowing to combine the functions of convolutional neural networks to extract and compress acoustic information relevant for recognition,
as well as connections to identify time dependencies between phones in order to develop a model of a word in the Buryat language.
The percentage of incorrectly recognized phones is 11%.
This value significantly exceeds the values typical for modern models.
The limitation of this model is the insufficient amount of training data and the dependence on the speaker's voice.
Expanding the speech sample bank, annotating the data and optimizing the model can improve the accuracy of the system and its resistance to speech variability.
Related Results
Pola Komunikasi Interpersonal Terapis Wicara pada Anak Telambat Bicara (Speech Delay)
Pola Komunikasi Interpersonal Terapis Wicara pada Anak Telambat Bicara (Speech Delay)
Abstract. The tittle of this study is Interpersonal Communication Patterns of Speech Therapists in Children with Speech Delay. Children with speech delay disorders are also social ...
Tindak Tutur pada Tradisi Mappettuada Suku Bugis
Tindak Tutur pada Tradisi Mappettuada Suku Bugis
This study aims to The purpose of this study is to understand and describe locutionary, illocutionary, and perlocutionary speech acts in the Mappettuada tradition of the Bugis-Maka...
Buryat historical sources: digital infrastructure of machine translation
Buryat historical sources: digital infrastructure of machine translation
The study is dedicated to a vast yet still underexplored corpus of Buryat historical sources in the Old Written Mongolian language, preserved in academic and archival institutions ...
Automatic speech recognition in voice-speech rehabilitation effectiveness evaluation in patients after laryngectomy
Automatic speech recognition in voice-speech rehabilitation effectiveness evaluation in patients after laryngectomy
Introduction.
Lost voice function compensation determines the personal and social life of laryngectomees. Automatic speech recognition and synthesis methods are...
Cattle in Buryat Mythology and Ritual
Cattle in Buryat Mythology and Ritual
This study addresses, on the basis of ethnographic, folkloric, linguistic, and field data, the role of cattle in Buryat myths and rites, with reference to their economic significan...
Multimodal Emotion Recognition and Human Computer Interaction for AI-Driven Mental Health Support (Preprint)
Multimodal Emotion Recognition and Human Computer Interaction for AI-Driven Mental Health Support (Preprint)
BACKGROUND
Mental health has become one of the most urgent global health issues of the twenty-first century. The World Health Organization (WHO) reports tha...
Speech Recognition
Speech Recognition
AbstractThe ease at which humans use speech understates the complexity of speech recognition for machines. The fundamental difficulty with speech recognition is the overwhelming va...
Recent Advances in Robust Speech Recognition Technology
Recent Advances in Robust Speech Recognition Technology
This E-book is a collection of articles that describe advances in speech recognition technology. Robustness in speech recognition refers to the need to maintain high speech recogni...

