Javascript must be enabled to continue!

Can ChatGPT and Gemini justify brain CT referrals? A comparative study with human experts and a custom prediction model

Abstract Background The poor uptake of imaging referral guidelines in Europe results in a substantial amount of inappropriate computed tomography (CT) scans. Publicly available chatbots, ChatGPT and Gemini, offer an alternative for justifying real-world referrals. Recent research reports high ChatGPT accuracy when analysing American College of Radiology Appropriateness Criteria variants. We compared the chatbots’ performance in interpreting, justifying, and suggesting alternative imaging for unstructured adult brain CT referrals in accordance with the European Society of Radiology iGuide. Our prediction model for automated iGuide categorisation of referrals was also compared against the chatbots. Methods The iGuide justification of 143 real-world CT brain referrals, used to evaluate a prediction model, was analysed by two radiographers and radiologists. ChatGPT-4’s and Gemini’s imaging recommendations and pathology suspicions were compared with those of humans, with respect to referral completeness. Inter-rater reliability with κ statistics determined the agreement between entities. Results Chatbots’ performance was limited (κ = 0.3) but improved for more complete referrals. The prediction model outperformed the chatbots in justification analysis (κ = 0.853). The chatbots’ interpretations of complete referrals were highly consistent (49/52, 94.2%). The agreement regarding alternative imaging was high for both complete and ambiguous referrals, with ChatGPT and Gemini correctly identifying imaging modality and anatomical region in 83/96 (86.5%) and 81/96 (84.4%) cases, respectively. Conclusion The chatbots’ ability to analyse the justification of adult brain CT referrals is limited to complete referrals, unlike our prediction model. Further research is needed to confirm these findings for other types of CT scans and modalities. Relevance statement ChatGPT and Gemini exhibit potential in justifying free text brain CT referrals; however, further improvements are required to handle real-world referrals of varying quality. Key Points Custom prediction model’s justification analysis strongly aligns with iGuide and surpasses chatbots. Chatbots incorrectly justified almost one-half of all CT brain referrals. Chatbots have limited performance in justifying ambiguous CT brain referrals. Chatbot performance improved when referrals were detailed and included suspected pathology. Graphical Abstract

Springer Science and Business Media LLC

Jaka Potočnik Edel Thomas Dearbhla Kearney Ronan P. Killeen Eric J. Heffernan Shane J. Foley

European Radiology Experimental

2025

Title: Can ChatGPT and Gemini justify brain CT referrals? A comparative study with human experts and a custom prediction model

Description:

Abstract Background The poor uptake of imaging referral guidelines in Europe results in a substantial amount of inappropriate computed tomography (CT) scans.

Publicly available chatbots, ChatGPT and Gemini, offer an alternative for justifying real-world referrals.

Recent research reports high ChatGPT accuracy when analysing American College of Radiology Appropriateness Criteria variants.

We compared the chatbots’ performance in interpreting, justifying, and suggesting alternative imaging for unstructured adult brain CT referrals in accordance with the European Society of Radiology iGuide.

Our prediction model for automated iGuide categorisation of referrals was also compared against the chatbots.

Methods The iGuide justification of 143 real-world CT brain referrals, used to evaluate a prediction model, was analysed by two radiographers and radiologists.

ChatGPT-4’s and Gemini’s imaging recommendations and pathology suspicions were compared with those of humans, with respect to referral completeness.

Inter-rater reliability with κ statistics determined the agreement between entities.

Results Chatbots’ performance was limited (κ = 0.

3) but improved for more complete referrals.

The prediction model outperformed the chatbots in justification analysis (κ = 0.

853).

The chatbots’ interpretations of complete referrals were highly consistent (49/52, 94.

2%).

The agreement regarding alternative imaging was high for both complete and ambiguous referrals, with ChatGPT and Gemini correctly identifying imaging modality and anatomical region in 83/96 (86.

5%) and 81/96 (84.

4%) cases, respectively.

Conclusion The chatbots’ ability to analyse the justification of adult brain CT referrals is limited to complete referrals, unlike our prediction model.

Further research is needed to confirm these findings for other types of CT scans and modalities.

Relevance statement ChatGPT and Gemini exhibit potential in justifying free text brain CT referrals; however, further improvements are required to handle real-world referrals of varying quality.

Key Points Custom prediction model’s justification analysis strongly aligns with iGuide and surpasses chatbots.

Chatbots incorrectly justified almost one-half of all CT brain referrals.

Chatbots have limited performance in justifying ambiguous CT brain referrals.

Chatbot performance improved when referrals were detailed and included suspected pathology.

Graphical Abstract.

Back

Abstract Introduction The exact manner in which large language models (LLMs) will be integrated into pathology is not yet fully comprehended. This study examines the accuracy, bene...

Assessment of Chat-GPT, Gemini, and Perplexity in Principle of Research Publication: A Comparative Study

Abstract Introduction Many researchers utilize artificial intelligence (AI) to aid their research endeavors. This study seeks to assess and contrast the performance of three sophis...

Brain Organoids, the Path Forward?

Photo by Maxim Berg on Unsplash INTRODUCTION The brain is one of the most foundational parts of being human, and we are still learning about what makes humans unique. Advancements ...

Performance of AI ‐Chatbots to Common Temporomandibular Joint Disorders ( TMDs ) Patient Queries: Accuracy, Completeness, Reliability and Readability

ABSTRACT TMDs are a common group of conditions affecting the temporomandibular joint (TMJ) often resulting from factors like injury, stress or teeth grinding. Thi...

Unlocking Educational Potential: Exploring Students’ Satisfaction and Sustainable Engagement with ChatGPT Using the ECM Model

Aim/Purpose: The main goal of this study is to investigate the factors affecting students’ satisfaction and continuous usage of ChatGPT in an educational context, using the Expecta...

ChatGPT's Capabilities for Use in Anatomy Education and Anatomy Research

Dear Editors, Recently, the discussion of an artificial intelligence (AI) - fueled platform in several articles in your journal has attracted the attention of many researchers [1, ...

Primerjalna književnost na prelomu tisočletja

In a comprehensive and at times critical manner, this volume seeks to shed light on the development of events in Western (i.e., European and North American) comparative literature ...

[RETRACTED] Gro-X Brain Reviews - Is Gro-X Brain A Scam? v1

[RETRACTED]➢Item Name - Gro-X Brain➢ Creation - Natural Organic Compound➢ Incidental Effects - NA➢ Accessibility - Online➢ Rating - ⭐⭐⭐⭐⭐➢ Click Here To Visit - Official Website - ...

Email:
Password:

Email:

Can ChatGPT and Gemini justify brain CT referrals? A comparative study with human experts and a custom prediction model

Related Results