Javascript must be enabled to continue!
Performance of 5 AI Models on United States Medical Licensing Examination Step 1 Questions: Comparative Observational Study (Preprint)
View through CrossRef
BACKGROUND
Artificial intelligence (AI) models are increasingly being used in medical education. Although models like ChatGPT have previously demonstrated strong performance on United States Medical Licensing Examination (USMLE)–style questions, newer AI tools with enhanced capabilities are now available, necessitating comparative evaluations of their accuracy and reliability across different medical domains and question formats.
OBJECTIVE
This study aimed to evaluate and compare the performance of 5 publicly available AI models: Grok, ChatGPT-4, Copilot, Gemini, and DeepSeek, on the USMLE Step 1 free 120-question set, assessing their accuracy and consistency across question types and medical subjects.
METHODS
This cross-sectional observational study was conducted between February 10 and March 5, 2025. Each of the 119 USMLE-style questions (excluding 1 audio-based item) was presented to each AI model by using a standardized prompt cycle. Models answered each question 3 times to assess confidence and consistency. Questions were categorized as text-based or image-based and as case-based or information-based. Statistical analysis was performed using chi-square and Fisher exact tests, with Bonferroni adjustment for pairwise comparisons.
RESULTS
Grok achieved the highest score (109/119, 91.6%), followed by Copilot (101/119, 84.9%), Gemini (100/119, 84%), ChatGPT-4 (95/119, 79.8%), and DeepSeek (86/119, 72.3%). DeepSeek’s lower score was due to an inability to process visual media, resulting in 0% accuracy on image-based items. When limited to text-only questions (n=96), DeepSeek’s accuracy increased to 89.6% (86/96), matching Copilot. Grok showed the highest accuracy on image-based (21/23, 91.3%) and case-based questions (70/78, 89.7%), with statistically significant differences observed between Grok and DeepSeek on case-based items (<i>P</i>=.01). The models performed best in biostatistics and epidemiology (5.8/6, 96.7%) and worst in musculoskeletal, skin, and connective tissue (4.4/7, 62.9%). Grok maintained 100% consistency in responses, while Copilot demonstrated the most self-correction (112/119, 94.1% consistency), improving its accuracy to 89.9% (107/119) on the third attempt.
CONCLUSIONS
AI models showed varying strengths across domains, with Grok demonstrating the highest accuracy and consistency in this dataset, particularly for image-based and reasoning-heavy questions. Although ChatGPT-4 remains widely used, newer models like Grok and Copilot also performed competitively. Continuous evaluation is essential as AI tools rapidly evolve.
CLINICALTRIAL
JMIR Publications Inc.
Title: Performance of 5 AI Models on United States Medical Licensing Examination Step 1 Questions: Comparative Observational Study (Preprint)
Description:
BACKGROUND
Artificial intelligence (AI) models are increasingly being used in medical education.
Although models like ChatGPT have previously demonstrated strong performance on United States Medical Licensing Examination (USMLE)–style questions, newer AI tools with enhanced capabilities are now available, necessitating comparative evaluations of their accuracy and reliability across different medical domains and question formats.
OBJECTIVE
This study aimed to evaluate and compare the performance of 5 publicly available AI models: Grok, ChatGPT-4, Copilot, Gemini, and DeepSeek, on the USMLE Step 1 free 120-question set, assessing their accuracy and consistency across question types and medical subjects.
METHODS
This cross-sectional observational study was conducted between February 10 and March 5, 2025.
Each of the 119 USMLE-style questions (excluding 1 audio-based item) was presented to each AI model by using a standardized prompt cycle.
Models answered each question 3 times to assess confidence and consistency.
Questions were categorized as text-based or image-based and as case-based or information-based.
Statistical analysis was performed using chi-square and Fisher exact tests, with Bonferroni adjustment for pairwise comparisons.
RESULTS
Grok achieved the highest score (109/119, 91.
6%), followed by Copilot (101/119, 84.
9%), Gemini (100/119, 84%), ChatGPT-4 (95/119, 79.
8%), and DeepSeek (86/119, 72.
3%).
DeepSeek’s lower score was due to an inability to process visual media, resulting in 0% accuracy on image-based items.
When limited to text-only questions (n=96), DeepSeek’s accuracy increased to 89.
6% (86/96), matching Copilot.
Grok showed the highest accuracy on image-based (21/23, 91.
3%) and case-based questions (70/78, 89.
7%), with statistically significant differences observed between Grok and DeepSeek on case-based items (<i>P</i>=.
01).
The models performed best in biostatistics and epidemiology (5.
8/6, 96.
7%) and worst in musculoskeletal, skin, and connective tissue (4.
4/7, 62.
9%).
Grok maintained 100% consistency in responses, while Copilot demonstrated the most self-correction (112/119, 94.
1% consistency), improving its accuracy to 89.
9% (107/119) on the third attempt.
CONCLUSIONS
AI models showed varying strengths across domains, with Grok demonstrating the highest accuracy and consistency in this dataset, particularly for image-based and reasoning-heavy questions.
Although ChatGPT-4 remains widely used, newer models like Grok and Copilot also performed competitively.
Continuous evaluation is essential as AI tools rapidly evolve.
CLINICALTRIAL.
Related Results
Primerjalna književnost na prelomu tisočletja
Primerjalna književnost na prelomu tisočletja
In a comprehensive and at times critical manner, this volume seeks to shed light on the development of events in Western (i.e., European and North American) comparative literature ...
Health care managers’ perspectives on workforce licensing practice in Ethiopia: A qualitative study
Health care managers’ perspectives on workforce licensing practice in Ethiopia: A qualitative study
Abstract
Background: Active monitoring of entry into the workforce starts with the licensing of professionals before entering the workforce. The professional licensing bodi...
Exploring the implementation of public involvement in local alcohol availability policy: the case of alcohol licensing decision‐making in England
Exploring the implementation of public involvement in local alcohol availability policy: the case of alcohol licensing decision‐making in England
AbstractBackground and AimsIn 2003, the UK government passed the Licensing Act for England and Wales. The Act provides a framework for regulating alcohol sale, including four licen...
Envisioning Originalism Applied to Bioethics Cases
Envisioning Originalism Applied to Bioethics Cases
Photo ID 123697425 © Alexandersikov | Dreamstime.com
Abstract
Originalism is an increasingly prevalent method for interpreting provisions of the US Constitution. It requires strict...
Assessment of Chat-GPT, Gemini, and Perplexity in Principle of Research Publication: A Comparative Study
Assessment of Chat-GPT, Gemini, and Perplexity in Principle of Research Publication: A Comparative Study
Abstract
Introduction
Many researchers utilize artificial intelligence (AI) to aid their research endeavors. This study seeks to assess and contrast the performance of three sophis...
Evaluating the Science to Inform the Physical Activity Guidelines for Americans Midcourse Report
Evaluating the Science to Inform the Physical Activity Guidelines for Americans Midcourse Report
Abstract
The Physical Activity Guidelines for Americans (Guidelines) advises older adults to be as active as possible. Yet, despite the well documented benefits of physical activi...
History of the medical licensing examination (<i>uieop</i>) in Korea’s Goryeo Dynasty (918-1392)
History of the medical licensing examination (<i>uieop</i>) in Korea’s Goryeo Dynasty (918-1392)
This article aims to describe the training and medical licensing system (uieop) for becoming a physician officer (uigwan) during Korea’s Goryeo Dynasty (918-1392). In the Goryeo Dy...
Business Licensing and Financial Performance of Small and Medium Enterprises in the Manufacturing Sector in Bungoma County, Kenya
Business Licensing and Financial Performance of Small and Medium Enterprises in the Manufacturing Sector in Bungoma County, Kenya
Small and Medium-Sized Enterprises (SMEs) play a pivotal role in Kenya's economic growth; however, their financial performance is often constrained by regulatory complexities, part...

