Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

eEVA: A Formative Pilot Evaluation of a Multimodal, Emotionally Responsive Virtual Agent for Mental Health Support (Preprint)

View through CrossRef
BACKGROUND Conversational agents are increasingly used for mental health support, but most operate on text alone and cannot perceive or respond to users’ emotional expressions. Prior work suggests that an agent’s ability to recognize affect and respond with congruent nonverbal behavior is associated with higher perceived empathy, a construct central to therapeutic rapport. OBJECTIVE This formative study describes the design and implementation of a multimodal emotion-responsive module for eEVA, an embodied virtual agent developed in the Virtual Intelligent Social Agents (VISAGE) research environment, and reports a small-scale pilot evaluation of the module’s perceived empathy using a validated computational-empathy questionnaire. METHODS he module fuses two input channels in real time: facial-expression analysis of webcam video and emotion analysis of speech transcribed to text, both obtained through a commercial affect-recognition API (Hume AI). A rule-based fusion procedure selects a dominant emotion per user utterance—agreement between channels resolves directly, disagreement resolves to the higher-intensity channel—and the agent displays a congruent facial expression when the dominant emotion’s intensity exceeds a 50% threshold. In a scripted, single-session laboratory interaction, 6 participants (male graduate computer science students, ages 18–24) interacted with the agent and completed the 78-item computational-empathy questionnaire of Brännström et al, which measures six dimensions: Perceive, Theory of Mind, Act, Manifest, Ratification, and Interpersonal. Scores were normalized (0–1) and summarized per dimension using the same procedure as the reference study, enabling a descriptive, cross-study comparison against published scores for two text-based mental health chatbots, Wysa and Replika. RESULTS Participants’ normalized ratings of eEVA were highest on the Act dimension (mean 0.69), followed by Perceive (0.55), Interpersonal (0.55), Theory of Mind (0.46), Manifest (0.43),and Ratification (0.42). In the descriptive cross-study comparison against the published day-7 scores for Wysa and Replika, eEVA scored higher on three of the six dimensions (Perceive, Theory of Mind, and Act) and lower on Manifest and Ratification. We attribute the Act advantage to the agent’s embodiment—its capacity to physically display congruent emotional responses— while dimensions that depend on sustained relationship formation favored the chatbots, which were rated after a week of repeated use. CONCLUSIONS A rule-based fusion of facial and linguistic affect signals is feasible in a working embodied agent and, in this small formative sample, was associated with favorable perceived empathy ratings on perception- and expression-related dimensions. Given the very small, homogeneous sample, the scripted single-session design, and the cross-study nature of the comparison, these results are preliminary feasibility findings rather than evidence of clinical efficacy. They motivate larger, controlled evaluations of multimodal emotion responsiveness in mental health support agents, including culturally and linguistically adapted settings.
Title: eEVA: A Formative Pilot Evaluation of a Multimodal, Emotionally Responsive Virtual Agent for Mental Health Support (Preprint)
Description:
BACKGROUND Conversational agents are increasingly used for mental health support, but most operate on text alone and cannot perceive or respond to users’ emotional expressions.
Prior work suggests that an agent’s ability to recognize affect and respond with congruent nonverbal behavior is associated with higher perceived empathy, a construct central to therapeutic rapport.
OBJECTIVE This formative study describes the design and implementation of a multimodal emotion-responsive module for eEVA, an embodied virtual agent developed in the Virtual Intelligent Social Agents (VISAGE) research environment, and reports a small-scale pilot evaluation of the module’s perceived empathy using a validated computational-empathy questionnaire.
METHODS he module fuses two input channels in real time: facial-expression analysis of webcam video and emotion analysis of speech transcribed to text, both obtained through a commercial affect-recognition API (Hume AI).
A rule-based fusion procedure selects a dominant emotion per user utterance—agreement between channels resolves directly, disagreement resolves to the higher-intensity channel—and the agent displays a congruent facial expression when the dominant emotion’s intensity exceeds a 50% threshold.
In a scripted, single-session laboratory interaction, 6 participants (male graduate computer science students, ages 18–24) interacted with the agent and completed the 78-item computational-empathy questionnaire of Brännström et al, which measures six dimensions: Perceive, Theory of Mind, Act, Manifest, Ratification, and Interpersonal.
Scores were normalized (0–1) and summarized per dimension using the same procedure as the reference study, enabling a descriptive, cross-study comparison against published scores for two text-based mental health chatbots, Wysa and Replika.
RESULTS Participants’ normalized ratings of eEVA were highest on the Act dimension (mean 0.
69), followed by Perceive (0.
55), Interpersonal (0.
55), Theory of Mind (0.
46), Manifest (0.
43),and Ratification (0.
42).
In the descriptive cross-study comparison against the published day-7 scores for Wysa and Replika, eEVA scored higher on three of the six dimensions (Perceive, Theory of Mind, and Act) and lower on Manifest and Ratification.
We attribute the Act advantage to the agent’s embodiment—its capacity to physically display congruent emotional responses— while dimensions that depend on sustained relationship formation favored the chatbots, which were rated after a week of repeated use.
CONCLUSIONS A rule-based fusion of facial and linguistic affect signals is feasible in a working embodied agent and, in this small formative sample, was associated with favorable perceived empathy ratings on perception- and expression-related dimensions.
Given the very small, homogeneous sample, the scripted single-session design, and the cross-study nature of the comparison, these results are preliminary feasibility findings rather than evidence of clinical efficacy.
They motivate larger, controlled evaluations of multimodal emotion responsiveness in mental health support agents, including culturally and linguistically adapted settings.

Related Results

Multimodal Emotion Recognition and Human Computer Interaction for AI-Driven Mental Health Support (Preprint)
Multimodal Emotion Recognition and Human Computer Interaction for AI-Driven Mental Health Support (Preprint)
BACKGROUND Mental health has become one of the most urgent global health issues of the twenty-first century. The World Health Organization (WHO) reports tha...
Organisatie van geestelijke gezondheidszorg voor mensen met een ernstige en persisterende mentale aandoening
Organisatie van geestelijke gezondheidszorg voor mensen met een ernstige en persisterende mentale aandoening
1 INTRODUCTION AND RESEARCH QUESTIONS 5 -- 2 GENERAL BACKGROUND: DEFINITIONS AND SCOPE OF THE STUDY 7 -- 2.1 CHRONIC AND COMPLEX MENTAL DISORDERS: DEFINITIONS AND SCOPE OF THE -- S...
Laboratory Model Study of Single Five-Spot and Single Injection Well Pilot Waterflooding
Laboratory Model Study of Single Five-Spot and Single Injection Well Pilot Waterflooding
Abstract Many full-scale waterflooding operations are preceded by pilot floods, one purpose of which is to provide an estimate of recoverable oil. A laboratory mo...
Digital Mental Health Landscaping in Low- and Middle-Income Countries 
Digital Mental Health Landscaping in Low- and Middle-Income Countries 
Introduction The aim of this project was to map the landscape of who is doing what and where in digital mental health, and to pr...

Back to Top