Javascript must be enabled to continue!
Enhancing Embedded Space with Low–Level Features for Speech Emotion Recognition
View through CrossRef
This work proposes an approach that uses a feature space by combining the representation obtained in the unsupervised learning process and manually selected features defining the prosody of the utterances. In the experiments, we used two time-frequency representations (Mel and CQT spectrograms) and EmoDB and RAVDESS databases. As the results show, the proposed system improved the classification accuracy of both representations: 1.29% for CQT and 3.75% for Mel spectrogram compared to the typical CNN architecture for the EmoDB dataset and 3.02% for CQT and 0.63% for Mel spectrogram in the case of RAVDESS. Additionally, the results present a significant increase of around 14% in classification performance in the case of happiness and disgust emotions using Mel spectrograms and around 20% in happiness and disgust emotions for CQT in the case of best models trained on EmoDB. On the other hand, in the case of models that achieved the highest result for the RAVDESS database, the most significant improvement was observed in the classification of a neutral state, around 16%, using the Mel spectrogram. For CQT representation, the most significant improvement occurred for fear and surprise, around 9%. Additionally, the average results for all prepared models showed the positive impact of the method used on the quality of classification of most emotional states. For the EmoDB database, the highest average improvement was observed for happiness—14.6%. For other emotions, it ranged from 1.2% to 8.7%. The only exception was the emotion of sadness, for which the classification quality was average decreased by 1% when using the Mel spectrogram. In turn, for the RAVDESS database, the most significant improvement also occurred for happiness—7.5%, while for other emotions ranged from 0.2% to 7.1%, except disgust and calm, the classification of which deteriorated for the Mel spectrogram and the CQT representation, respectively.
Title: Enhancing Embedded Space with Low–Level Features for Speech Emotion Recognition
Description:
This work proposes an approach that uses a feature space by combining the representation obtained in the unsupervised learning process and manually selected features defining the prosody of the utterances.
In the experiments, we used two time-frequency representations (Mel and CQT spectrograms) and EmoDB and RAVDESS databases.
As the results show, the proposed system improved the classification accuracy of both representations: 1.
29% for CQT and 3.
75% for Mel spectrogram compared to the typical CNN architecture for the EmoDB dataset and 3.
02% for CQT and 0.
63% for Mel spectrogram in the case of RAVDESS.
Additionally, the results present a significant increase of around 14% in classification performance in the case of happiness and disgust emotions using Mel spectrograms and around 20% in happiness and disgust emotions for CQT in the case of best models trained on EmoDB.
On the other hand, in the case of models that achieved the highest result for the RAVDESS database, the most significant improvement was observed in the classification of a neutral state, around 16%, using the Mel spectrogram.
For CQT representation, the most significant improvement occurred for fear and surprise, around 9%.
Additionally, the average results for all prepared models showed the positive impact of the method used on the quality of classification of most emotional states.
For the EmoDB database, the highest average improvement was observed for happiness—14.
6%.
For other emotions, it ranged from 1.
2% to 8.
7%.
The only exception was the emotion of sadness, for which the classification quality was average decreased by 1% when using the Mel spectrogram.
In turn, for the RAVDESS database, the most significant improvement also occurred for happiness—7.
5%, while for other emotions ranged from 0.
2% to 7.
1%, except disgust and calm, the classification of which deteriorated for the Mel spectrogram and the CQT representation, respectively.
Related Results
Multimodal Emotion Recognition and Human Computer Interaction for AI-Driven Mental Health Support (Preprint)
Multimodal Emotion Recognition and Human Computer Interaction for AI-Driven Mental Health Support (Preprint)
BACKGROUND
Mental health has become one of the most urgent global health issues of the twenty-first century. The World Health Organization (WHO) reports tha...
Pola Komunikasi Interpersonal Terapis Wicara pada Anak Telambat Bicara (Speech Delay)
Pola Komunikasi Interpersonal Terapis Wicara pada Anak Telambat Bicara (Speech Delay)
Abstract. The tittle of this study is Interpersonal Communication Patterns of Speech Therapists in Children with Speech Delay. Children with speech delay disorders are also social ...
REVIEW AND ANALYSIS OF APPROACHES AND PRACTICAL APPLICATIONS OF HUMAN EMOTION RECOGNITION
REVIEW AND ANALYSIS OF APPROACHES AND PRACTICAL APPLICATIONS OF HUMAN EMOTION RECOGNITION
Human emotions are complex and multifaceted, making them difficult to quantify and analyze. However, as technology advances, researchers are exploring the artificial intelligence u...
Age-related Differences in Emotion Recognition Ability: Visual and Auditory Modalities
Age-related Differences in Emotion Recognition Ability: Visual and Auditory Modalities
Emotion recognition is an important aspect of social interaction. Deficits in emotion recognition have been tied to poor social competence, interpersonal functioning, and communica...
Tindak Tutur pada Tradisi Mappettuada Suku Bugis
Tindak Tutur pada Tradisi Mappettuada Suku Bugis
This study aims to The purpose of this study is to understand and describe locutionary, illocutionary, and perlocutionary speech acts in the Mappettuada tradition of the Bugis-Maka...
3477 Impaired emotion recognition accuracy after right-hemisphere stroke
3477 Impaired emotion recognition accuracy after right-hemisphere stroke
OBJECTIVES/SPECIFIC AIMS: Every year, approximately 800,000 Americans suffer a stroke. Supportive social environments are recognized as an important factor contributing to successf...
Acoustic emotion recognition using spectral and temporal features
Acoustic emotion recognition using spectral and temporal features
In this paper, utility of different low- level, spectral and temporal features is evaluated for the task of emotion recognition. The aim of an ideal speech emotion recognition syst...
The impact of binge drinking on emotion recognition
The impact of binge drinking on emotion recognition
Binge drinking or heavy episodic drinking is variously defined but according to the World Health Organisation (WHO) it is the consumption of at least 60 grams or more of pure alcoh...

