Javascript must be enabled to continue!
Enhancing dysarthric speech recognition through SepFormer and hierarchical attention network models with multistage transfer learning
View through CrossRef
AbstractDysarthria, a motor speech disorder that impacts articulation and speech clarity, presents significant challenges for Automatic Speech Recognition (ASR) systems. This study proposes a groundbreaking approach to enhance the accuracy of Dysarthric Speech Recognition (DSR). A primary innovation lies in the integration of the SepFormer-Speech Enhancement Generative Adversarial Network (S-SEGAN), an advanced generative adversarial network tailored for Dysarthric Speech Enhancement (DSE), as a front-end processing stage for DSR systems. The S-SEGAN integrates SEGAN’s adversarial learning with SepFormer speech separation capabilities, demonstrating significant improvements in performance. Furthermore, a multistage transfer learning approach is employed to assess the DSR models for both word-level and sentence-level DSR. These DSR models are first trained on a large speech dataset (LibriSpeech) and then fine-tuned on dysarthric speech data (both isolated and augmented). Evaluations demonstrate significant DSR accuracy improvements in DSE integration. The Dysarthric Speech (DS)-baseline models (without DSE), Transformer and Conformer achieved Word Recognition Accuracy (WRA) percentages of 68.60% and 69.87%, respectively. The introduction of Hierarchical Attention Network (HAN) with the Transformer and Conformer architectures resulted in improved performance, with T-HAN achieving a WRA of 71.07% and C-HAN reaching 73%. The Transformer model with DSE + DSR for isolated words achieves a WRA of 73.40%, while that of the Conformer model reaches 74.33%. Notably, the T-HAN and C-HAN models with DSE + DSR demonstrate even more substantial enhancements, with WRAs of 75.73% and 76.87%, respectively. Augmenting words further boosts model performance, with the Transformer and Conformer models achieving WRAs of 76.47% and 79.20%, respectively. Remarkably, the T-HAN and C-HAN models with DSE + DSR and augmented words exhibit WRAs of 82.13% and 84.07%, respectively, with C-HAN displaying the highest performance among all proposed models.
Springer Science and Business Media LLC
Title: Enhancing dysarthric speech recognition through SepFormer and hierarchical attention network models with multistage transfer learning
Description:
AbstractDysarthria, a motor speech disorder that impacts articulation and speech clarity, presents significant challenges for Automatic Speech Recognition (ASR) systems.
This study proposes a groundbreaking approach to enhance the accuracy of Dysarthric Speech Recognition (DSR).
A primary innovation lies in the integration of the SepFormer-Speech Enhancement Generative Adversarial Network (S-SEGAN), an advanced generative adversarial network tailored for Dysarthric Speech Enhancement (DSE), as a front-end processing stage for DSR systems.
The S-SEGAN integrates SEGAN’s adversarial learning with SepFormer speech separation capabilities, demonstrating significant improvements in performance.
Furthermore, a multistage transfer learning approach is employed to assess the DSR models for both word-level and sentence-level DSR.
These DSR models are first trained on a large speech dataset (LibriSpeech) and then fine-tuned on dysarthric speech data (both isolated and augmented).
Evaluations demonstrate significant DSR accuracy improvements in DSE integration.
The Dysarthric Speech (DS)-baseline models (without DSE), Transformer and Conformer achieved Word Recognition Accuracy (WRA) percentages of 68.
60% and 69.
87%, respectively.
The introduction of Hierarchical Attention Network (HAN) with the Transformer and Conformer architectures resulted in improved performance, with T-HAN achieving a WRA of 71.
07% and C-HAN reaching 73%.
The Transformer model with DSE + DSR for isolated words achieves a WRA of 73.
40%, while that of the Conformer model reaches 74.
33%.
Notably, the T-HAN and C-HAN models with DSE + DSR demonstrate even more substantial enhancements, with WRAs of 75.
73% and 76.
87%, respectively.
Augmenting words further boosts model performance, with the Transformer and Conformer models achieving WRAs of 76.
47% and 79.
20%, respectively.
Remarkably, the T-HAN and C-HAN models with DSE + DSR and augmented words exhibit WRAs of 82.
13% and 84.
07%, respectively, with C-HAN displaying the highest performance among all proposed models.
Related Results
A Survey of Automatic Speech Recognition for Dysarthric Speech
A Survey of Automatic Speech Recognition for Dysarthric Speech
Dysarthric speech has several pathological characteristics, such as discontinuous pronunciation, uncontrolled volume, slow speech, explosive pronunciation, improper pauses, excessi...
Empowering Dysarthric Speech: Leveraging Advanced LLMs for Accurate Speech Correction and Multimodal Emotion Analysis
Empowering Dysarthric Speech: Leveraging Advanced LLMs for Accurate Speech Correction and Multimodal Emotion Analysis
Dysarthria or Dysarthric speech as called is kind of a motor speech disorder which is caused by neurological damage that affects the muscles for speech production, which results in...
Recent Advances in Dysarthric Speech Recognition: Approaches and Datasets
Recent Advances in Dysarthric Speech Recognition: Approaches and Datasets
Dysarthria is a neuromotor speech disorder that results from physical disability and limits speech intelligibility. Dysarthric speakers can make use of speech recognition systems t...
Pola Komunikasi Interpersonal Terapis Wicara pada Anak Telambat Bicara (Speech Delay)
Pola Komunikasi Interpersonal Terapis Wicara pada Anak Telambat Bicara (Speech Delay)
Abstract. The tittle of this study is Interpersonal Communication Patterns of Speech Therapists in Children with Speech Delay. Children with speech delay disorders are also social ...
Multimodal Emotion Recognition and Human Computer Interaction for AI-Driven Mental Health Support (Preprint)
Multimodal Emotion Recognition and Human Computer Interaction for AI-Driven Mental Health Support (Preprint)
BACKGROUND
Mental health has become one of the most urgent global health issues of the twenty-first century. The World Health Organization (WHO) reports tha...
A Comprehensive Survey of Automatic Dysarthric Speech Recognition
A Comprehensive Survey of Automatic Dysarthric Speech Recognition
Automatic dysarthric speech recognition (DSR) is very crucial for many human computer interaction systems that enables the human to interact with machine in natural way. The object...
AFM signal model for dysarthric speech classification using speech biomarkers
AFM signal model for dysarthric speech classification using speech biomarkers
Neurological disorders include various conditions affecting the brain, spinal cord, and nervous system which results in reduced performance in different organs and muscles througho...
Comparative Analysis of Deep Learning Models for Dysarthric Speech Detection
Comparative Analysis of Deep Learning Models for Dysarthric Speech Detection
Abstract
Dysarthria is a speech communication disorder that is associated with neurological impairments. In order to detect this disorder from speech, we present an experim...

