Javascript must be enabled to continue!
Speaker Role Identification in Clinical Conversations *
View through CrossRef
Patient-clinician communication research is crucial for understanding interaction dynamics and for predicting outcomes that are associated with clinical discourse. Traditionally, interaction analysis is conducted manually because of challenges such as Speaker Role Identification (SRI), which must reliably differentiate between doctors, medical assistants, patients, and other caregivers in the same room. Although automatic speech recognition with diarization can efficiently create a transcript with separate labels for each speaker, these systems are not able to assign roles to each person in the interaction. Previous SRI studies in task-oriented scenarios have directly predicted roles using linguistic features, bypassing diarization. However, to our knowledge nobody has investigated SRI in clinical settings. We explored whether Large Language Models (LLMs) such as BERT could accurately identify speaker roles in clinical transcripts, with and without diarization. We used veridical turn segmentation and diarization identifiers, fine-tuning each model at varying levels of identifier corruption to assess impact on performance. Our results demonstrate that BERT achieves high performance with linguistic signals alone (82% accuracy/82% F1-score), while incorporating accurate diarization identifiers further enhances accuracy (95%/95%). We conclude that fine-tuned LLMs are effective tools for SRI in clinical settings.
Title: Speaker Role Identification in Clinical Conversations
*
Description:
Patient-clinician communication research is crucial for understanding interaction dynamics and for predicting outcomes that are associated with clinical discourse.
Traditionally, interaction analysis is conducted manually because of challenges such as Speaker Role Identification (SRI), which must reliably differentiate between doctors, medical assistants, patients, and other caregivers in the same room.
Although automatic speech recognition with diarization can efficiently create a transcript with separate labels for each speaker, these systems are not able to assign roles to each person in the interaction.
Previous SRI studies in task-oriented scenarios have directly predicted roles using linguistic features, bypassing diarization.
However, to our knowledge nobody has investigated SRI in clinical settings.
We explored whether Large Language Models (LLMs) such as BERT could accurately identify speaker roles in clinical transcripts, with and without diarization.
We used veridical turn segmentation and diarization identifiers, fine-tuning each model at varying levels of identifier corruption to assess impact on performance.
Our results demonstrate that BERT achieves high performance with linguistic signals alone (82% accuracy/82% F1-score), while incorporating accurate diarization identifiers further enhances accuracy (95%/95%).
We conclude that fine-tuned LLMs are effective tools for SRI in clinical settings.
Related Results
Funkcije komunikacijski relevantne šutnje u njemačkome
Funkcije komunikacijski relevantne šutnje u njemačkome
Additionally, this chapter presents research of silence with review of main aspects of papers in the field of conversational analysis, ethnography of communication and metaphor of ...
Speaker Verification and Identification
Speaker Verification and Identification
A speaker recognition system verifies or identifies a speaker’s identity based on his/her voice. It is considered as one of the most convenient biometric characteristic for human m...
Quarantine Powers, Biodefense, and Andrew Speaker
Quarantine Powers, Biodefense, and Andrew Speaker
In January 2007, Andrew Speaker (Speaker) underwent a chest X-ray and CT scan, which revealed an abnormality in his lungs. However, tests results indicated that he did not ha...
Cometary Physics Laboratory: spectrophotometric experiments
Cometary Physics Laboratory: spectrophotometric experiments
<p><strong><span dir="ltr" role="presentation">1. Introduction</span></strong&...
Tiedon rajat ja vuorovaikutus. Toteamukseen tai vaihtoehtokysymykseen vastaavat VOI OLLA -rakenteet [On the limits of knowledge. Responding to an assertion or a polar question with VOI OLLA ‘(it) may be’ structures]
Tiedon rajat ja vuorovaikutus. Toteamukseen tai vaihtoehtokysymykseen vastaavat VOI OLLA -rakenteet [On the limits of knowledge. Responding to an assertion or a polar question with VOI OLLA ‘(it) may be’ structures]
Artikkeli tarkastelee toteamukseen tai vaihtoehtokysymykseen vastaavia VOI OLLA -rakenteita voi olla, se voi olla, voi se olla ja voihan se olla. Toteamuksella tarkoitetaan kannano...
Analyzing Noise Robustness of Cochleogram and Mel Spectrogram Features in Deep Learning Based Speaker Recognition
Analyzing Noise Robustness of Cochleogram and Mel Spectrogram Features in Deep Learning Based Speaker Recognition
The performance of speaker recognition systems is very well on the datasets without noise and mismatch. However, the performance gets degraded with the environmental noises, channe...
Samskapad implementering av proaktiva samtal inför vård och omsorg i livets sista tid på särskilt boende för äldre
Samskapad implementering av proaktiva samtal inför vård och omsorg i livets sista tid på särskilt boende för äldre
<p dir="ltr"><b>Background</b>: Older adults with extensive care needs often spend their final years, months, and days in residential care facilities. In this con...
Samskapad implementering av proaktiva samtal inför vård och omsorg i livets sista tid på särskilt boende för äldre
Samskapad implementering av proaktiva samtal inför vård och omsorg i livets sista tid på särskilt boende för äldre
<p dir="ltr"><b>Background</b>: Older adults with extensive care needs often spend their final years, months, and days in residential care facilities. In this con...

