Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Study on a Bimodal Emotion Recognition Algorithm Based on Deep Fusion of Speech and Images

View through CrossRef
Emotion recognition aims to identify affective categories by analyzing physiological signals and behavioral characteristics and is one of the key research directions in Artificial Intelligence (AI). Addressing the issues of limited unimodal representation capability and insufficiently established cross-modal deep association in existing emotion recognition algorithms, which lead to inadequate accuracy and robustness in complex scenarios, this paper proposes a bimodal emotion recognition algorithm based on deep fusion of speech and image to enhance the accuracy and robustness of emotion recognition by establishing an effective cross-modal interaction and adaptive fusion mechanism. For image modality, Bilinear Interpolation (BI) is used for image scale normalization, followed by an improved Convolutional Neural Network (CNN) to extract deep spatial features. An improved Sparse Autoencoder (SAE) is then employed to compress features, reduce redundancy, and enhance fine-grained details. Finally, an improved Multilayer Perceptron (MLP) performs emotion classification. For speech modality, prosodic features, Mel-Frequency Cepstral Coefficients (MFCCs), geometric features, and attribute features are fused to form a multidimensional acoustic representation, which is subsequently classified using an improved MLP. After unimodal recognition, the outputs of the two modalities are fused at the decision level using a dynamic adaptive weighting strategy to generate the final emotion category and its corresponding probability. The experimental results indicate that, under a unified test set, the accuracy of sentiment recognition in this paper is improved by 14% and 18% respectively compared to single-modality models. Compared to other fusion strategy models, the recognition accuracy is concentrated in the range of 65% to 70%. The deep fusion method proposed in this paper achieves an overall recognition accuracy of 81.18%, which is superior to other methods. This result verifies the effectiveness of the algorithm proposed in this paper in improving the accuracy of sentiment recognition and system robustness.
Title: Study on a Bimodal Emotion Recognition Algorithm Based on Deep Fusion of Speech and Images
Description:
Emotion recognition aims to identify affective categories by analyzing physiological signals and behavioral characteristics and is one of the key research directions in Artificial Intelligence (AI).
Addressing the issues of limited unimodal representation capability and insufficiently established cross-modal deep association in existing emotion recognition algorithms, which lead to inadequate accuracy and robustness in complex scenarios, this paper proposes a bimodal emotion recognition algorithm based on deep fusion of speech and image to enhance the accuracy and robustness of emotion recognition by establishing an effective cross-modal interaction and adaptive fusion mechanism.
For image modality, Bilinear Interpolation (BI) is used for image scale normalization, followed by an improved Convolutional Neural Network (CNN) to extract deep spatial features.
An improved Sparse Autoencoder (SAE) is then employed to compress features, reduce redundancy, and enhance fine-grained details.
Finally, an improved Multilayer Perceptron (MLP) performs emotion classification.
For speech modality, prosodic features, Mel-Frequency Cepstral Coefficients (MFCCs), geometric features, and attribute features are fused to form a multidimensional acoustic representation, which is subsequently classified using an improved MLP.
After unimodal recognition, the outputs of the two modalities are fused at the decision level using a dynamic adaptive weighting strategy to generate the final emotion category and its corresponding probability.
The experimental results indicate that, under a unified test set, the accuracy of sentiment recognition in this paper is improved by 14% and 18% respectively compared to single-modality models.
Compared to other fusion strategy models, the recognition accuracy is concentrated in the range of 65% to 70%.
The deep fusion method proposed in this paper achieves an overall recognition accuracy of 81.
18%, which is superior to other methods.
This result verifies the effectiveness of the algorithm proposed in this paper in improving the accuracy of sentiment recognition and system robustness.

Related Results

Multimodal Emotion Recognition and Human Computer Interaction for AI-Driven Mental Health Support (Preprint)
Multimodal Emotion Recognition and Human Computer Interaction for AI-Driven Mental Health Support (Preprint)
BACKGROUND Mental health has become one of the most urgent global health issues of the twenty-first century. The World Health Organization (WHO) reports tha...
The Nuclear Fusion Award
The Nuclear Fusion Award
The Nuclear Fusion Award ceremony for 2009 and 2010 award winners was held during the 23rd IAEA Fusion Energy Conference in Daejeon. This time, both 2009 and 2010 award winners w...
AI-Based Emotion Recognition in Education: Progress, Applications, and Open Challenges
AI-Based Emotion Recognition in Education: Progress, Applications, and Open Challenges
AI-based emotion recognition has emerged as a critical component of affect-aware educational technologies, particularly in online, large-scale, and technology-mediated learning env...
The impact of binge drinking on emotion recognition
The impact of binge drinking on emotion recognition
Binge drinking or heavy episodic drinking is variously defined but according to the World Health Organisation (WHO) it is the consumption of at least 60 grams or more of pure alcoh...
Exploring the Effect of Demographics Inclusion on Subject-independent Emotion Recognition
Exploring the Effect of Demographics Inclusion on Subject-independent Emotion Recognition
Electroencephalography (EEG) can capture electrical activity associated with human emotion processing from the scalp. The electrical activity can be processed using deep learning m...
Language Development in Children with Cochlear Implant using Bimodal Approach: SLP Perspective
Language Development in Children with Cochlear Implant using Bimodal Approach: SLP Perspective
Background: The development of language skills in children with cochlear implants is a vital area of research, particularly in understanding the impact of the bimodal approach. Thi...
Studies on visual emotion understanding
Studies on visual emotion understanding
As information explodes nowadays, visual data has become a crucial information carrier in various fields: social networks, e-commerce, online entertainment, etc. Visual emotion ana...
What about males? Exploring sex differences in the relationship between emotion difficulties and eating disorders
What about males? Exploring sex differences in the relationship between emotion difficulties and eating disorders
Abstract Objective: While eating disorders (ED) are more commonly diagnosed in females, there is growing awareness that men also experience ED and may do so in a different ...

Back to Top