Javascript must be enabled to continue!

A Language Model–Powered Simulated Patient With Automated Feedback for History Taking: Prospective Study

Background Although history taking is fundamental for diagnosing medical conditions, teaching and providing feedback on the skill can be challenging due to resource constraints. Virtual simulated patients and web-based chatbots have thus emerged as educational tools, with recent advancements in artificial intelligence (AI) such as large language models (LLMs) enhancing their realism and potential to provide feedback. Objective In our study, we aimed to evaluate the effectiveness of a Generative Pretrained Transformer (GPT) 4 model to provide structured feedback on medical students’ performance in history taking with a simulated patient. Methods We conducted a prospective study involving medical students performing history taking with a GPT-powered chatbot. To that end, we designed a chatbot to simulate patients’ responses and provide immediate feedback on the comprehensiveness of the students’ history taking. Students’ interactions with the chatbot were analyzed, and feedback from the chatbot was compared with feedback from a human rater. We measured interrater reliability and performed a descriptive analysis to assess the quality of feedback. Results Most of the study’s participants were in their third year of medical school. A total of 1894 question-answer pairs from 106 conversations were included in our analysis. GPT-4’s role-play and responses were medically plausible in more than 99% of cases. Interrater reliability between GPT-4 and the human rater showed “almost perfect” agreement (Cohen κ=0.832). Less agreement (κ<0.6) detected for 8 out of 45 feedback categories highlighted topics about which the model’s assessments were overly specific or diverged from human judgement. Conclusions The GPT model was effective in providing structured feedback on history-taking dialogs provided by medical students. Although we unraveled some limitations regarding the specificity of feedback for certain feedback categories, the overall high agreement with human raters suggests that LLMs can be a valuable tool for medical education. Our findings, thus, advocate the careful integration of AI-driven feedback mechanisms in medical training and highlight important aspects when LLMs are used in that context.

JMIR Publications Inc.

Friederike Holderried Christian Stegemann-Philipps Anne Herrmann-Werner Teresa Festl-Wietek Martin Holderried Carsten Eickhoff Moritz Mahling

JMIR Medical Education

2024

Title: A Language Model–Powered Simulated Patient With Automated Feedback for History Taking: Prospective Study

Description:

Background Although history taking is fundamental for diagnosing medical conditions, teaching and providing feedback on the skill can be challenging due to resource constraints.

Virtual simulated patients and web-based chatbots have thus emerged as educational tools, with recent advancements in artificial intelligence (AI) such as large language models (LLMs) enhancing their realism and potential to provide feedback.

Objective In our study, we aimed to evaluate the effectiveness of a Generative Pretrained Transformer (GPT) 4 model to provide structured feedback on medical students’ performance in history taking with a simulated patient.

Methods We conducted a prospective study involving medical students performing history taking with a GPT-powered chatbot.

To that end, we designed a chatbot to simulate patients’ responses and provide immediate feedback on the comprehensiveness of the students’ history taking.

Students’ interactions with the chatbot were analyzed, and feedback from the chatbot was compared with feedback from a human rater.

We measured interrater reliability and performed a descriptive analysis to assess the quality of feedback.

Results Most of the study’s participants were in their third year of medical school.

A total of 1894 question-answer pairs from 106 conversations were included in our analysis.

GPT-4’s role-play and responses were medically plausible in more than 99% of cases.

Interrater reliability between GPT-4 and the human rater showed “almost perfect” agreement (Cohen κ=0.

832).

Less agreement (κ<0.

6) detected for 8 out of 45 feedback categories highlighted topics about which the model’s assessments were overly specific or diverged from human judgement.

Conclusions The GPT model was effective in providing structured feedback on history-taking dialogs provided by medical students.

Although we unraveled some limitations regarding the specificity of feedback for certain feedback categories, the overall high agreement with human raters suggests that LLMs can be a valuable tool for medical education.

Our findings, thus, advocate the careful integration of AI-driven feedback mechanisms in medical training and highlight important aspects when LLMs are used in that context.

Back

<p><em><span style="font-size: 11.0pt; font-family: 'Times New Roman',serif; mso-fareast-font-family: 'Times New Roman'; mso-ansi-language: EN-US; mso-fareast-langua...

Autonomy on Trial

Photo by CHUTTERSNAP on Unsplash Abstract This paper critically examines how US bioethics and health law conceptualize patient autonomy, contrasting the rights-based, individualist...

Increased life expectancy of heart failure patients in a rural center by a multidisciplinary program

Abstract Funding Acknowledgements Type of funding sources: None. INTRODUCTION Patients with heart failure (HF)...

A Language Model–Powered Simulated Patient With Automated Feedback for History Taking: Prospective Study (Preprint)

BACKGROUND Although history taking is fundamental for diagnosing medical conditions, teaching and providing feedback on the skill can be challenging due to ...

Written Feedback In Second Language Writing: Perceptions Of Vietnamese Teachers And Students

<p>Writing can be very challenging for ESL students since they need to overcome the changes associated with academic writing styles and their mechanics in order to improve th...

An empirical investigation of contemporary performance management systems

This dissertation provides a comprehensive empirical analysis of contemporary performance management systems (PMS), with a focus on how evolving feedback practices—particularly nar...

A Wideband mm-Wave Printed Dipole Antenna for 5G Applications

<span lang="EN-MY">In this paper, a wideband millimeter-wave (mm-Wave) printed dipole antenna is proposed to be used for fifth generation (5G) communications. The single elem...

TEACHERS’ FEEDBACK ON ESSAY WRITING. (c2020)

This study aimed at identifying English language teachers’ perspectives of middle-school L2 learners’ challenges when receiving feedback on their writing essays, and exploring cycl...

Email:
Password:

Email:

A Language Model–Powered Simulated Patient With Automated Feedback for History Taking: Prospective Study

Related Results