Javascript must be enabled to continue!
How does DeepSeek-R1 perform on USMLE?
View through CrossRef
AbstractDeepSeek, a Chinese artificial intelligence company, released its first free chatbot app based on its DeepSeek-R1 model. DeepSeek provides its models, algorithms, and training details to ensure transparency and reproducibility. Their new model is trained with reinforcement learning, allowing it to learn through interactions and feedback rather than relying solely on supervised learning. Reports showcase that DeepSeek’s model shows competitive performances against established large language models (LLMs) such as Anthropic’s Claude and OpenAI’s GPT-4o on established benchmarks in language understanding, mathematics (AIME 2024) and programming (Codeforces) while trained at a fraction of the costs. Additionally, running inference shows significantly lower costs, leading to DeepSeek surpassing ChatGPT as the most downloaded free app on the American iOS App Store. This development contributed to a nearly 17% drop in Nvidia’s share price, resulting in the most significant one-day loss in U.S. history, amounting to nearly $600 billion. The open-source models also bring a significant shift in the healthcare system, allowing cost-efficient medical LLMs to be deployed within hospital networks. To understand its performance in the healthcare sector, we analyse the new DeepSeek-R1 model on the United States Medical Licensing Examination (USMLE) and compare it to ChatGPT.
Title: How does DeepSeek-R1 perform on USMLE?
Description:
AbstractDeepSeek, a Chinese artificial intelligence company, released its first free chatbot app based on its DeepSeek-R1 model.
DeepSeek provides its models, algorithms, and training details to ensure transparency and reproducibility.
Their new model is trained with reinforcement learning, allowing it to learn through interactions and feedback rather than relying solely on supervised learning.
Reports showcase that DeepSeek’s model shows competitive performances against established large language models (LLMs) such as Anthropic’s Claude and OpenAI’s GPT-4o on established benchmarks in language understanding, mathematics (AIME 2024) and programming (Codeforces) while trained at a fraction of the costs.
Additionally, running inference shows significantly lower costs, leading to DeepSeek surpassing ChatGPT as the most downloaded free app on the American iOS App Store.
This development contributed to a nearly 17% drop in Nvidia’s share price, resulting in the most significant one-day loss in U.
S.
history, amounting to nearly $600 billion.
The open-source models also bring a significant shift in the healthcare system, allowing cost-efficient medical LLMs to be deployed within hospital networks.
To understand its performance in the healthcare sector, we analyse the new DeepSeek-R1 model on the United States Medical Licensing Examination (USMLE) and compare it to ChatGPT.
Related Results
Educators’ Perspectives on DeepSeek in ELT: A Qualitative Case Study of Pedagogical Potentials and Pitfalls in Chinese Higher Education
Educators’ Perspectives on DeepSeek in ELT: A Qualitative Case Study of Pedagogical Potentials and Pitfalls in Chinese Higher Education
Aim/Purpose: This study aimed to investigate the perspectives of English Language Teaching (ELT) educators on DeepSeek, emphasizing its pedagogical value, practical challenges, and...
DeepSeek for Pathology Report Understanding: A Benchmark Study of Cancer Type Extraction, AJCC Staging, and Prognosis Prediction
DeepSeek for Pathology Report Understanding: A Benchmark Study of Cancer Type Extraction, AJCC Staging, and Prognosis Prediction
Abstract
Background:
Pathology reports contain clinically important information for cancer diagnosis, staging, and outcome research, but their unstructured format l...
Are USMLE Scores Valid Measures for Chief Resident Selection?
Are USMLE Scores Valid Measures for Chief Resident Selection?
ABSTRACT
Background
The US Medical Licensing Examination (USMLE) Step 1 and Step 2 scores are often used to inform a vari...
A Survey of DeepSeek Models
A Survey of DeepSeek Models
Advances in artificial intelligence (AI) rely on systems capable of human-like reasoning, a limitation for conventional Large Language Models (LLMs), which struggle with multi-step...
Factors Associated with Infectious Diseases Fellowship Academic Success
Factors Associated with Infectious Diseases Fellowship Academic Success
Abstract
Background: A multitude of factors are considered in an infectious diseases (ID) training program’s meticulous selection process of ID fellows but their correlatio...
Performance of 5 AI Models on United States Medical Licensing Examination Step 1 Questions: Comparative Observational Study (Preprint)
Performance of 5 AI Models on United States Medical Licensing Examination Step 1 Questions: Comparative Observational Study (Preprint)
BACKGROUND
Artificial intelligence (AI) models are increasingly being used in medical education. Although models like ChatGPT have previously demonstrated strong ...
Head-to-head evaluation of ChatGPT, DeepSeek, and Perplexity on acid–base disorder case clinical management and drug treatment: Accuracy, domain performance, and response consistency assessment
Head-to-head evaluation of ChatGPT, DeepSeek, and Perplexity on acid–base disorder case clinical management and drug treatment: Accuracy, domain performance, and response consistency assessment
Background
Large language models (LLMs) are increasingly used in medical education, but their performance and reliability on mechanistically demanding topics li...
Factors That Predict Success in the General Surgery Match With a Pass/Fail USMLE Step 1 Exam
Factors That Predict Success in the General Surgery Match With a Pass/Fail USMLE Step 1 Exam
Introduction
Matching into a residency program is an intricate process, including holistic application review, selection for interview, and ranking of interview...

