Javascript must be enabled to continue!
Can ChatGPT and Gemini justify brain CT referrals? A comparative study with human experts and a custom prediction model
View through CrossRef
Abstract
Background
The poor uptake of imaging referral guidelines in Europe results in a substantial amount of inappropriate computed tomography (CT) scans. Publicly available chatbots, ChatGPT and Gemini, offer an alternative for justifying real-world referrals. Recent research reports high ChatGPT accuracy when analysing American College of Radiology Appropriateness Criteria variants. We compared the chatbots’ performance in interpreting, justifying, and suggesting alternative imaging for unstructured adult brain CT referrals in accordance with the European Society of Radiology iGuide. Our prediction model for automated iGuide categorisation of referrals was also compared against the chatbots.
Methods
The iGuide justification of 143 real-world CT brain referrals, used to evaluate a prediction model, was analysed by two radiographers and radiologists. ChatGPT-4’s and Gemini’s imaging recommendations and pathology suspicions were compared with those of humans, with respect to referral completeness. Inter-rater reliability with κ statistics determined the agreement between entities.
Results
Chatbots’ performance was limited (κ = 0.3) but improved for more complete referrals. The prediction model outperformed the chatbots in justification analysis (κ = 0.853). The chatbots’ interpretations of complete referrals were highly consistent (49/52, 94.2%). The agreement regarding alternative imaging was high for both complete and ambiguous referrals, with ChatGPT and Gemini correctly identifying imaging modality and anatomical region in 83/96 (86.5%) and 81/96 (84.4%) cases, respectively.
Conclusion
The chatbots’ ability to analyse the justification of adult brain CT referrals is limited to complete referrals, unlike our prediction model. Further research is needed to confirm these findings for other types of CT scans and modalities.
Relevance statement
ChatGPT and Gemini exhibit potential in justifying free text brain CT referrals; however, further improvements are required to handle real-world referrals of varying quality.
Key Points
Custom prediction model’s justification analysis strongly aligns with iGuide and surpasses chatbots.
Chatbots incorrectly justified almost one-half of all CT brain referrals.
Chatbots have limited performance in justifying ambiguous CT brain referrals.
Chatbot performance improved when referrals were detailed and included suspected pathology.
Graphical Abstract
Springer Science and Business Media LLC
Title: Can ChatGPT and Gemini justify brain CT referrals? A comparative study with human experts and a custom prediction model
Description:
Abstract
Background
The poor uptake of imaging referral guidelines in Europe results in a substantial amount of inappropriate computed tomography (CT) scans.
Publicly available chatbots, ChatGPT and Gemini, offer an alternative for justifying real-world referrals.
Recent research reports high ChatGPT accuracy when analysing American College of Radiology Appropriateness Criteria variants.
We compared the chatbots’ performance in interpreting, justifying, and suggesting alternative imaging for unstructured adult brain CT referrals in accordance with the European Society of Radiology iGuide.
Our prediction model for automated iGuide categorisation of referrals was also compared against the chatbots.
Methods
The iGuide justification of 143 real-world CT brain referrals, used to evaluate a prediction model, was analysed by two radiographers and radiologists.
ChatGPT-4’s and Gemini’s imaging recommendations and pathology suspicions were compared with those of humans, with respect to referral completeness.
Inter-rater reliability with κ statistics determined the agreement between entities.
Results
Chatbots’ performance was limited (κ = 0.
3) but improved for more complete referrals.
The prediction model outperformed the chatbots in justification analysis (κ = 0.
853).
The chatbots’ interpretations of complete referrals were highly consistent (49/52, 94.
2%).
The agreement regarding alternative imaging was high for both complete and ambiguous referrals, with ChatGPT and Gemini correctly identifying imaging modality and anatomical region in 83/96 (86.
5%) and 81/96 (84.
4%) cases, respectively.
Conclusion
The chatbots’ ability to analyse the justification of adult brain CT referrals is limited to complete referrals, unlike our prediction model.
Further research is needed to confirm these findings for other types of CT scans and modalities.
Relevance statement
ChatGPT and Gemini exhibit potential in justifying free text brain CT referrals; however, further improvements are required to handle real-world referrals of varying quality.
Key Points
Custom prediction model’s justification analysis strongly aligns with iGuide and surpasses chatbots.
Chatbots incorrectly justified almost one-half of all CT brain referrals.
Chatbots have limited performance in justifying ambiguous CT brain referrals.
Chatbot performance improved when referrals were detailed and included suspected pathology.
Graphical Abstract.
Related Results
Exploring Large Language Models Integration in the Histopathologic Diagnosis of Skin Diseases: A Comparative Study
Exploring Large Language Models Integration in the Histopathologic Diagnosis of Skin Diseases: A Comparative Study
Abstract
Introduction
The exact manner in which large language models (LLMs) will be integrated into pathology is not yet fully comprehended. This study examines the accuracy, bene...
Assessment of Chat-GPT, Gemini, and Perplexity in Principle of Research Publication: A Comparative Study
Assessment of Chat-GPT, Gemini, and Perplexity in Principle of Research Publication: A Comparative Study
Abstract
Introduction
Many researchers utilize artificial intelligence (AI) to aid their research endeavors. This study seeks to assess and contrast the performance of three sophis...
Brain Organoids, the Path Forward?
Brain Organoids, the Path Forward?
Photo by Maxim Berg on Unsplash
INTRODUCTION
The brain is one of the most foundational parts of being human, and we are still learning about what makes humans unique. Advancements ...
Performance of
AI
‐Chatbots to Common Temporomandibular Joint Disorders (
TMDs
) Patient Queries: Accuracy, Completeness, Reliability and Readability
Performance of
AI
‐Chatbots to Common Temporomandibular Joint Disorders (
TMDs
) Patient Queries: Accuracy, Completeness, Reliability and Readability
ABSTRACT
TMDs are a common group of conditions affecting the temporomandibular joint (TMJ) often resulting from factors like injury, stress or teeth grinding. Thi...
Unlocking Educational Potential: Exploring Students’ Satisfaction and Sustainable Engagement with ChatGPT Using the ECM Model
Unlocking Educational Potential: Exploring Students’ Satisfaction and Sustainable Engagement with ChatGPT Using the ECM Model
Aim/Purpose: The main goal of this study is to investigate the factors affecting students’ satisfaction and continuous usage of ChatGPT in an educational context, using the Expecta...
ChatGPT's Capabilities for Use in Anatomy Education and Anatomy Research
ChatGPT's Capabilities for Use in Anatomy Education and Anatomy Research
Dear Editors,
Recently, the discussion of an artificial intelligence (AI) - fueled platform in several articles in your journal has attracted the attention of many researchers [1, ...
Primerjalna književnost na prelomu tisočletja
Primerjalna književnost na prelomu tisočletja
In a comprehensive and at times critical manner, this volume seeks to shed light on the development of events in Western (i.e., European and North American) comparative literature ...
[RETRACTED] Gro-X Brain Reviews - Is Gro-X Brain A Scam? v1
[RETRACTED] Gro-X Brain Reviews - Is Gro-X Brain A Scam? v1
[RETRACTED]➢Item Name - Gro-X Brain➢ Creation - Natural Organic Compound➢ Incidental Effects - NA➢ Accessibility - Online➢ Rating - ⭐⭐⭐⭐⭐➢ Click Here To Visit - Official Website - ...

