Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Empowering Dysarthric Speech: Leveraging Advanced LLMs for Accurate Speech Correction and Multimodal Emotion Analysis

View through CrossRef
Dysarthria or Dysarthric speech as called is kind of a motor speech disorder which is caused by neurological damage that affects the muscles for speech production, which results in slurred, slow, or difficult-to-understand speech which has been affecting millions of people worldwide including those with conditions such as stroke, traumatic brain injury, cerebral palsy, Parkinson's disease, and multiple sclerosis, dysarthric presents a significant communication barrier, which impacts the quality of life and social interaction of individuals. The entire aim of this paper is to come up with a mechanism that can recognize and translate the speech of dysarthric users and empower their ability to communicate effectively. In this paper, we are proposing a novel approach leveraging Advanced Large Language models for accurate speech correction and multimodal emotion analysis. Our methodology involves converting dysarthric speech to text using OpenAI's whisper model, followed by fine-tuned open source models such as LlaMa 3.1 70B and Mistral 8x7B models on Groq AI accelerators to predict the intended sentences from distorted input speech accurately. The entire dataset used was made by combining TORGO dataset with Google speech data and then manually labeling emotional contexts. Our framework identifies emotions such as happiness, sadness, neutral, surprise, anger and fear highlighting the potential understanding of dysarthric speech. Our approach effectively reconsturcts the intended sentences and detects emotions with high accuracy.
Title: Empowering Dysarthric Speech: Leveraging Advanced LLMs for Accurate Speech Correction and Multimodal Emotion Analysis
Description:
Dysarthria or Dysarthric speech as called is kind of a motor speech disorder which is caused by neurological damage that affects the muscles for speech production, which results in slurred, slow, or difficult-to-understand speech which has been affecting millions of people worldwide including those with conditions such as stroke, traumatic brain injury, cerebral palsy, Parkinson's disease, and multiple sclerosis, dysarthric presents a significant communication barrier, which impacts the quality of life and social interaction of individuals.
The entire aim of this paper is to come up with a mechanism that can recognize and translate the speech of dysarthric users and empower their ability to communicate effectively.
In this paper, we are proposing a novel approach leveraging Advanced Large Language models for accurate speech correction and multimodal emotion analysis.
Our methodology involves converting dysarthric speech to text using OpenAI's whisper model, followed by fine-tuned open source models such as LlaMa 3.
1 70B and Mistral 8x7B models on Groq AI accelerators to predict the intended sentences from distorted input speech accurately.
The entire dataset used was made by combining TORGO dataset with Google speech data and then manually labeling emotional contexts.
Our framework identifies emotions such as happiness, sadness, neutral, surprise, anger and fear highlighting the potential understanding of dysarthric speech.
Our approach effectively reconsturcts the intended sentences and detects emotions with high accuracy.

Related Results

Multimodal Emotion Recognition and Human Computer Interaction for AI-Driven Mental Health Support (Preprint)
Multimodal Emotion Recognition and Human Computer Interaction for AI-Driven Mental Health Support (Preprint)
BACKGROUND Mental health has become one of the most urgent global health issues of the twenty-first century. The World Health Organization (WHO) reports tha...
A Survey of Automatic Speech Recognition for Dysarthric Speech
A Survey of Automatic Speech Recognition for Dysarthric Speech
Dysarthric speech has several pathological characteristics, such as discontinuous pronunciation, uncontrolled volume, slow speech, explosive pronunciation, improper pauses, excessi...
Exploring Large Language Models Integration in the Histopathologic Diagnosis of Skin Diseases: A Comparative Study
Exploring Large Language Models Integration in the Histopathologic Diagnosis of Skin Diseases: A Comparative Study
Abstract Introduction The exact manner in which large language models (LLMs) will be integrated into pathology is not yet fully comprehended. This study examines the accuracy, bene...
Enhancing dysarthric speech recognition through SepFormer and hierarchical attention network models with multistage transfer learning
Enhancing dysarthric speech recognition through SepFormer and hierarchical attention network models with multistage transfer learning
AbstractDysarthria, a motor speech disorder that impacts articulation and speech clarity, presents significant challenges for Automatic Speech Recognition (ASR) systems. This study...
Recent Advances in Dysarthric Speech Recognition: Approaches and Datasets
Recent Advances in Dysarthric Speech Recognition: Approaches and Datasets
Dysarthria is a neuromotor speech disorder that results from physical disability and limits speech intelligibility. Dysarthric speakers can make use of speech recognition systems t...
Perspectives and Experiences With Large Language Models in Health Care: Survey Study (Preprint)
Perspectives and Experiences With Large Language Models in Health Care: Survey Study (Preprint)
BACKGROUND Large language models (LLMs) are transforming how data is used, including within the health care sector. However, frameworks including the Unifie...
Perspectives and Experiences With Large Language Models in Health Care: Survey Study
Perspectives and Experiences With Large Language Models in Health Care: Survey Study
Background Large language models (LLMs) are transforming how data is used, including within the health care sector. However, frameworks including the Unified Th...
A Systematic Review of ChatGPT and Other Conversational Large Language Models in Healthcare
A Systematic Review of ChatGPT and Other Conversational Large Language Models in Healthcare
Abstract Background The launch of the Chat Generative Pre-trained Transformer (ChatGPT) in November 2022 has attracted public a...

Back to Top