Javascript must be enabled to continue!
Diagnostic accuracy of vision-language models on Japanese diagnostic radiology, nuclear medicine, and interventional radiology specialty board examinations
View through CrossRef
Abstract
Purpose
The performance of vision-language models (VLMs) with image interpretation capabilities, such as GPT-4 omni (GPT-4o), GPT-4 vision (GPT-4V), and Claude-3, has not been compared and remains unexplored in specialized radiological fields, including nuclear medicine and interventional radiology. This study aimed to evaluate and compare the diagnostic accuracy of various VLMs, including GPT-4 + GPT-4V, GPT-4o, Claude-3 Sonnet, and Claude-3 Opus, using Japanese diagnostic radiology, nuclear medicine, and interventional radiology (JDR, JNM, and JIR, respectively) board certification tests.
Materials and methods
In total, 383 questions from the JDR test (358 images), 300 from the JNM test (92 images), and 322 from the JIR test (96 images) from 2019 to 2023 were consecutively collected. The accuracy rates of the GPT-4 + GPT-4V, GPT-4o, Claude-3 Sonnet, and Claude-3 Opus were calculated for all questions or questions with images. The accuracy rates of the VLMs were compared using McNemar’s test.
Results
GPT-4o demonstrated the highest accuracy rates across all evaluations with the JDR (all questions, 49%; questions with images, 48%), JNM (all questions, 64%; questions with images, 59%), and JIR tests (all questions, 43%; questions with images, 34%), followed by Claude-3 Opus with the JDR (all questions, 40%; questions with images, 38%), JNM (all questions, 42%; questions with images, 43%), and JIR tests (all questions, 40%; questions with images, 30%). For all questions, McNemar’s test showed that GPT-4o significantly outperformed the other VLMs (all
P
< 0.007), except for Claude-3 Opus in the JIR test. For questions with images, GPT-4o outperformed the other VLMs in the JDR and JNM tests (all
P
< 0.001), except Claude-3 Opus in the JNM test.
Conclusion
The GPT-4o had the highest success rates for questions with images and all questions from the JDR, JNM, and JIR board certification tests.
Title: Diagnostic accuracy of vision-language models on Japanese diagnostic radiology, nuclear medicine, and interventional radiology specialty board examinations
Description:
Abstract
Purpose
The performance of vision-language models (VLMs) with image interpretation capabilities, such as GPT-4 omni (GPT-4o), GPT-4 vision (GPT-4V), and Claude-3, has not been compared and remains unexplored in specialized radiological fields, including nuclear medicine and interventional radiology.
This study aimed to evaluate and compare the diagnostic accuracy of various VLMs, including GPT-4 + GPT-4V, GPT-4o, Claude-3 Sonnet, and Claude-3 Opus, using Japanese diagnostic radiology, nuclear medicine, and interventional radiology (JDR, JNM, and JIR, respectively) board certification tests.
Materials and methods
In total, 383 questions from the JDR test (358 images), 300 from the JNM test (92 images), and 322 from the JIR test (96 images) from 2019 to 2023 were consecutively collected.
The accuracy rates of the GPT-4 + GPT-4V, GPT-4o, Claude-3 Sonnet, and Claude-3 Opus were calculated for all questions or questions with images.
The accuracy rates of the VLMs were compared using McNemar’s test.
Results
GPT-4o demonstrated the highest accuracy rates across all evaluations with the JDR (all questions, 49%; questions with images, 48%), JNM (all questions, 64%; questions with images, 59%), and JIR tests (all questions, 43%; questions with images, 34%), followed by Claude-3 Opus with the JDR (all questions, 40%; questions with images, 38%), JNM (all questions, 42%; questions with images, 43%), and JIR tests (all questions, 40%; questions with images, 30%).
For all questions, McNemar’s test showed that GPT-4o significantly outperformed the other VLMs (all
P
< 0.
007), except for Claude-3 Opus in the JIR test.
For questions with images, GPT-4o outperformed the other VLMs in the JDR and JNM tests (all
P
< 0.
001), except Claude-3 Opus in the JNM test.
Conclusion
The GPT-4o had the highest success rates for questions with images and all questions from the JDR, JNM, and JIR board certification tests.
Related Results
Hubungan Perilaku Pola Makan dengan Kejadian Anak Obesitas
Hubungan Perilaku Pola Makan dengan Kejadian Anak Obesitas
<p><em><span style="font-size: 11.0pt; font-family: 'Times New Roman',serif; mso-fareast-font-family: 'Times New Roman'; mso-ansi-language: EN-US; mso-fareast-langua...
Diagnostic Accuracy of Vision-Language Models on Japanese Diagnostic Radiology, Nuclear Medicine, and Interventional Radiology Specialty Board Examinations
Diagnostic Accuracy of Vision-Language Models on Japanese Diagnostic Radiology, Nuclear Medicine, and Interventional Radiology Specialty Board Examinations
Abstract
Purpose
The performance of vision-language models (VLMs) with image interpretation capabilities, such as GPT-4 omni (G...
Zero to hero
Zero to hero
Western images of Japan tell a seemingly incongruous story of love, sex and marriage – one full of contradictions and conflicting moral codes. We sometimes hear intriguing stories ...
Učinak poučavanja razrednomu jeziku u izobrazbi nastavnika njemačkoga
Učinak poučavanja razrednomu jeziku u izobrazbi nastavnika njemačkoga
The actual use of classroom language is principally limited to the classroom environment. As far as foreign language learning is concerned, the classroom often turns out to be the ...
Awareness of Interventional Radiology among Clinical Years’ Medical Students and Medical Interns at University of Hail
Awareness of Interventional Radiology among Clinical Years’ Medical Students and Medical Interns at University of Hail
Context:
One of the most important challenges facing the evolution of modern interventional radiology is its lack of awareness among medical students.
...
Increased life expectancy of heart failure patients in a rural center by a multidisciplinary program
Increased life expectancy of heart failure patients in a rural center by a multidisciplinary program
Abstract
Funding Acknowledgements
Type of funding sources: None.
INTRODUCTION Patients with heart failure (HF)...
Beyond Turf Wars: Redefining Boundaries and Building Consensus in Interventional Radiology
Beyond Turf Wars: Redefining Boundaries and Building Consensus in Interventional Radiology
Abstract
Interventional radiology (IR) has evolved from diagnostic angiography to a core therapeutic specialty addressing acute ischemic stroke, peripheral arteri...
AI and Incidental Findings
AI and Incidental Findings
Photo by Accuray on Unsplash
INTRODUCTION
Delayed and missed follow-up on incidental findings threatens patient health and is a major financial risk for healthcare systems. The hea...

