Javascript must be enabled to continue!
Reducing Under-Triage Risk in Large Language Model Based Clinical Triage Using UMLS-CUI Augmentation
View through CrossRef
Abstract
Background
Public facing large language models (LLMs) are increasingly used for health guidance, including triage recommendations. We evaluated whether augmenting LLM prompts with standardized clinical concepts from the Unified Medical Language System (UMLS) could improve the safety and robustness of clinical triage recommendations.
Methods
We used a publicly available dataset comprising 60 clinician-authored clinical vignettes, each represented in 16 demographic and narrative variations, yielding 960 vignette-factor combinations. Clinical entities were extracted using a two-stage pipeline combining ClinicalBERT-based named entity recognition with rule-based identification of laboratory abnormalities. Extracted entities were mapped to UMLS Concept Unique Identifiers (CUIs).Negated concepts were excluded. A confidence-weighted CUI voting classifier was trained using empirical associations between CUIs and clinician-assigned triage categories. We compared five approaches: CUI-only classification, MedGemma 27B, MedGemma 27B augmented with CUIs, GPT-4o-mini, and GPT-4o-mini augmented with CUIs. Outcomes included overall accuracy, under-triage, over-triage, emergency-case accuracy, and sensitivity to anchoring statements.
Results
CUI augmentation decreased under-triage but increased over-triage in both models tested (GPT-4o-mini and MedGemma 27B). It improved high-acuity recognition while reducing recognition of low-acuity cases. CUI augmentation had mixed effects on overall triage accuracy; accuracy increased for MedGemma 27B but decreased for GPT-4o-mini. Emergency-case accuracy improved from 73.0% to 80.7% for GPT-4o-mini and from 60.5% to 68.5% for MedGemma 27B. CUI augmentation also reduced susceptibility to anchoring statements. These findings suggest that the principal value of CUI augmentation may be shifting model behavior toward safety-oriented behavior rather than uniformly improving overall accuracy.
Conclusions
Ontology-grounded prompt augmentation shifted LLM triage recommendations toward greater sensitivity to high-acuity presentations and reduced overall under-triage. These safety gains were accompanied by increased over-triage and mixed effects on overall accuracy. A hybrid architecture combining LLM-based language understanding with interpretable UMLS-derived clinical concepts may improve the safety and robustness of AI-assisted triage. Further evaluation using real-world patient communications and clinical outcomes is warranted.
Title: Reducing Under-Triage Risk in Large Language Model Based Clinical Triage Using UMLS-CUI Augmentation
Description:
Abstract
Background
Public facing large language models (LLMs) are increasingly used for health guidance, including triage recommendations.
We evaluated whether augmenting LLM prompts with standardized clinical concepts from the Unified Medical Language System (UMLS) could improve the safety and robustness of clinical triage recommendations.
Methods
We used a publicly available dataset comprising 60 clinician-authored clinical vignettes, each represented in 16 demographic and narrative variations, yielding 960 vignette-factor combinations.
Clinical entities were extracted using a two-stage pipeline combining ClinicalBERT-based named entity recognition with rule-based identification of laboratory abnormalities.
Extracted entities were mapped to UMLS Concept Unique Identifiers (CUIs).
Negated concepts were excluded.
A confidence-weighted CUI voting classifier was trained using empirical associations between CUIs and clinician-assigned triage categories.
We compared five approaches: CUI-only classification, MedGemma 27B, MedGemma 27B augmented with CUIs, GPT-4o-mini, and GPT-4o-mini augmented with CUIs.
Outcomes included overall accuracy, under-triage, over-triage, emergency-case accuracy, and sensitivity to anchoring statements.
Results
CUI augmentation decreased under-triage but increased over-triage in both models tested (GPT-4o-mini and MedGemma 27B).
It improved high-acuity recognition while reducing recognition of low-acuity cases.
CUI augmentation had mixed effects on overall triage accuracy; accuracy increased for MedGemma 27B but decreased for GPT-4o-mini.
Emergency-case accuracy improved from 73.
0% to 80.
7% for GPT-4o-mini and from 60.
5% to 68.
5% for MedGemma 27B.
CUI augmentation also reduced susceptibility to anchoring statements.
These findings suggest that the principal value of CUI augmentation may be shifting model behavior toward safety-oriented behavior rather than uniformly improving overall accuracy.
Conclusions
Ontology-grounded prompt augmentation shifted LLM triage recommendations toward greater sensitivity to high-acuity presentations and reduced overall under-triage.
These safety gains were accompanied by increased over-triage and mixed effects on overall accuracy.
A hybrid architecture combining LLM-based language understanding with interpretable UMLS-derived clinical concepts may improve the safety and robustness of AI-assisted triage.
Further evaluation using real-world patient communications and clinical outcomes is warranted.
Related Results
Hubungan Perilaku Pola Makan dengan Kejadian Anak Obesitas
Hubungan Perilaku Pola Makan dengan Kejadian Anak Obesitas
<p><em><span style="font-size: 11.0pt; font-family: 'Times New Roman',serif; mso-fareast-font-family: 'Times New Roman'; mso-ansi-language: EN-US; mso-fareast-langua...
Učinak poučavanja razrednomu jeziku u izobrazbi nastavnika njemačkoga
Učinak poučavanja razrednomu jeziku u izobrazbi nastavnika njemačkoga
The actual use of classroom language is principally limited to the classroom environment. As far as foreign language learning is concerned, the classroom often turns out to be the ...
Improving emergency department triage quality improvement project
Improving emergency department triage quality improvement project
Background: Healthcare presents challenges that require nursing professionals to continually evaluate their practice. Overcrowding in the emergency department (ED) has become a wor...
An Evaluation of UMLS as a Controlled Terminology for the Problem List Toolkit
An Evaluation of UMLS as a Controlled Terminology for the Problem List Toolkit
We are developing a set of software components—the Problem List Toolkit (PL-Tk) – to support operations on clinical problem labels. An adaptation of the Nationa...
Evaluating the Science to Inform the Physical Activity Guidelines for Americans Midcourse Report
Evaluating the Science to Inform the Physical Activity Guidelines for Americans Midcourse Report
Abstract
The Physical Activity Guidelines for Americans (Guidelines) advises older adults to be as active as possible. Yet, despite the well documented benefits of physical activi...
Referral to geriatric rehabilitation
Referral to geriatric rehabilitation
Summary
Older hospital patients are vulnerable to adverse outcomes of hospital stay. In aging societies,
post-acute care (PAC) programs were developed to support functional
recov...
Increased life expectancy of heart failure patients in a rural center by a multidisciplinary program
Increased life expectancy of heart failure patients in a rural center by a multidisciplinary program
Abstract
Funding Acknowledgements
Type of funding sources: None.
INTRODUCTION Patients with heart failure (HF)...
A Comparative Study on Concept Representation between the UMLS and the Clinical Terms in Korean Medical Records
A Comparative Study on Concept Representation between the UMLS and the Clinical Terms in Korean Medical Records
The Unified Medical Language System (UMLS) is a rich source of knowledge in the biomedical domain. In this paper, we evaluated the coverage of UMLS as compared with Korean medical ...

