Javascript must be enabled to continue!
Performance of ChatGPT in Israeli Arabic-language OBGYN national medical licensure exam
View through CrossRef
Abstract
Background
Previous studies of ChatGPT performance in the field of medical exams have reached contradictory results. The performance of ChatGPT in languages other than English, including Arabic, which is the official language of medical education and practice in many countries, has yet to be explored. We aim to evaluate the performance of ChatGPT in Arabic-language Israeli OBGYN medical licensure exams for foreign university alumni.
Methods
We conducted a performance study using a consecutive sample of text-based multiple-choice questions, originated from authentic Arabic-language Israeli OBGYN medical licensure exams for foreign university alumni. ChatGPT-3.5 (using a newly created account) answered all questions in Arabic. We compared the performance of ChatGPT including in the different fields of the exam; Obstetrics, Reproductive medicine and Infertility, Gynecology and Gynecologic Oncology, and also compared ChatGPT Arabic performance vs. previously published English medical tests.
Results
Overall, 123 authentic questions were analyzed. ChatGPT correctly answered 54 questions (43.9%, 95% CI: 35.1% – 52.7%) and reached a score below 50%. There was no difference in ChatGPT performance in the four different subjects of the exam: Gynecologic Oncology (61.5%, 95% CI: 35.1% – 87.9%), Gynecology (44.0%, 95% CI: 24.5% – 63.5%), Obstetrics (42.3%, 95% CI: 28.9% – 55.7%), Reproductive medicine and Infertility (39.4%, 95% CI: 22.7% – 56.1%),
p
= .579. In a comparison to ChatGPT performance in 9,091 English language questions in the field of medicine, the performance of Arabic ChatGPT was lower (43.9% in Arabic vs. 60.7% in English,
p
< .001).
Conclusions
ChatGPT-3.5 answered correctly approximately 44% of Arabic OBGYN medical licensure exam questions. At the time of writing of this manuscript, considering the results of our analysis, ChatGPT-3.5 cannot be considered a reliable primary tool for exam preparation in Arabic. Further research and efforts should be made to improve ChatGPT performance in other languages besides English especially Arabic.
Springer Science and Business Media LLC
Title: Performance of ChatGPT in Israeli Arabic-language OBGYN national medical licensure exam
Description:
Abstract
Background
Previous studies of ChatGPT performance in the field of medical exams have reached contradictory results.
The performance of ChatGPT in languages other than English, including Arabic, which is the official language of medical education and practice in many countries, has yet to be explored.
We aim to evaluate the performance of ChatGPT in Arabic-language Israeli OBGYN medical licensure exams for foreign university alumni.
Methods
We conducted a performance study using a consecutive sample of text-based multiple-choice questions, originated from authentic Arabic-language Israeli OBGYN medical licensure exams for foreign university alumni.
ChatGPT-3.
5 (using a newly created account) answered all questions in Arabic.
We compared the performance of ChatGPT including in the different fields of the exam; Obstetrics, Reproductive medicine and Infertility, Gynecology and Gynecologic Oncology, and also compared ChatGPT Arabic performance vs.
previously published English medical tests.
Results
Overall, 123 authentic questions were analyzed.
ChatGPT correctly answered 54 questions (43.
9%, 95% CI: 35.
1% – 52.
7%) and reached a score below 50%.
There was no difference in ChatGPT performance in the four different subjects of the exam: Gynecologic Oncology (61.
5%, 95% CI: 35.
1% – 87.
9%), Gynecology (44.
0%, 95% CI: 24.
5% – 63.
5%), Obstetrics (42.
3%, 95% CI: 28.
9% – 55.
7%), Reproductive medicine and Infertility (39.
4%, 95% CI: 22.
7% – 56.
1%),
p
= .
579.
In a comparison to ChatGPT performance in 9,091 English language questions in the field of medicine, the performance of Arabic ChatGPT was lower (43.
9% in Arabic vs.
60.
7% in English,
p
< .
001).
Conclusions
ChatGPT-3.
5 answered correctly approximately 44% of Arabic OBGYN medical licensure exam questions.
At the time of writing of this manuscript, considering the results of our analysis, ChatGPT-3.
5 cannot be considered a reliable primary tool for exam preparation in Arabic.
Further research and efforts should be made to improve ChatGPT performance in other languages besides English especially Arabic.
Related Results
Hubungan Perilaku Pola Makan dengan Kejadian Anak Obesitas
Hubungan Perilaku Pola Makan dengan Kejadian Anak Obesitas
<p><em><span style="font-size: 11.0pt; font-family: 'Times New Roman',serif; mso-fareast-font-family: 'Times New Roman'; mso-ansi-language: EN-US; mso-fareast-langua...
Exploring Large Language Models Integration in the Histopathologic Diagnosis of Skin Diseases: A Comparative Study
Exploring Large Language Models Integration in the Histopathologic Diagnosis of Skin Diseases: A Comparative Study
Abstract
Introduction
The exact manner in which large language models (LLMs) will be integrated into pathology is not yet fully comprehended. This study examines the accuracy, bene...
Assessment of Chat-GPT, Gemini, and Perplexity in Principle of Research Publication: A Comparative Study
Assessment of Chat-GPT, Gemini, and Perplexity in Principle of Research Publication: A Comparative Study
Abstract
Introduction
Many researchers utilize artificial intelligence (AI) to aid their research endeavors. This study seeks to assess and contrast the performance of three sophis...
Unlocking Educational Potential: Exploring Students’ Satisfaction and Sustainable Engagement with ChatGPT Using the ECM Model
Unlocking Educational Potential: Exploring Students’ Satisfaction and Sustainable Engagement with ChatGPT Using the ECM Model
Aim/Purpose: The main goal of this study is to investigate the factors affecting students’ satisfaction and continuous usage of ChatGPT in an educational context, using the Expecta...
ChatGPT's Capabilities for Use in Anatomy Education and Anatomy Research
ChatGPT's Capabilities for Use in Anatomy Education and Anatomy Research
Dear Editors,
Recently, the discussion of an artificial intelligence (AI) - fueled platform in several articles in your journal has attracted the attention of many researchers [1, ...
ChatGPT takes the FCPS exam in Internal Medicine
ChatGPT takes the FCPS exam in Internal Medicine
ABSTRACT
Large language models (LLMs) have exhibited remarkable proficiency in clinical knowledge, encompassing diagnostic medicine, and have been tested on questio...
Učinak poučavanja razrednomu jeziku u izobrazbi nastavnika njemačkoga
Učinak poučavanja razrednomu jeziku u izobrazbi nastavnika njemačkoga
The actual use of classroom language is principally limited to the classroom environment. As far as foreign language learning is concerned, the classroom often turns out to be the ...
DISCOVERING THE EFFECTIVENESS OF TEACHING METHODS IN TEACHING COMMUNICATIVE ARABIC AT SULTAN SHARIF ALI ISLAMIC UNIVERSITY: FACULTY OF ARABIC LANGUAGE AS CASE STUDY
DISCOVERING THE EFFECTIVENESS OF TEACHING METHODS IN TEACHING COMMUNICATIVE ARABIC AT SULTAN SHARIF ALI ISLAMIC UNIVERSITY: FACULTY OF ARABIC LANGUAGE AS CASE STUDY
This research aims to identify the effectiveness of the objectives of teaching communicative Arabic at the Faculty of Arabic Language at Sultan Sharif Ali Islamic University in the...

