Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

How does DeepSeek-R1 perform on USMLE?

View through CrossRef
AbstractDeepSeek, a Chinese artificial intelligence company, released its first free chatbot app based on its DeepSeek-R1 model. DeepSeek provides its models, algorithms, and training details to ensure transparency and reproducibility. Their new model is trained with reinforcement learning, allowing it to learn through interactions and feedback rather than relying solely on supervised learning. Reports showcase that DeepSeek’s model shows competitive performances against established large language models (LLMs) such as Anthropic’s Claude and OpenAI’s GPT-4o on established benchmarks in language understanding, mathematics (AIME 2024) and programming (Codeforces) while trained at a fraction of the costs. Additionally, running inference shows significantly lower costs, leading to DeepSeek surpassing ChatGPT as the most downloaded free app on the American iOS App Store. This development contributed to a nearly 17% drop in Nvidia’s share price, resulting in the most significant one-day loss in U.S. history, amounting to nearly $600 billion. The open-source models also bring a significant shift in the healthcare system, allowing cost-efficient medical LLMs to be deployed within hospital networks. To understand its performance in the healthcare sector, we analyse the new DeepSeek-R1 model on the United States Medical Licensing Examination (USMLE) and compare it to ChatGPT.
Title: How does DeepSeek-R1 perform on USMLE?
Description:
AbstractDeepSeek, a Chinese artificial intelligence company, released its first free chatbot app based on its DeepSeek-R1 model.
DeepSeek provides its models, algorithms, and training details to ensure transparency and reproducibility.
Their new model is trained with reinforcement learning, allowing it to learn through interactions and feedback rather than relying solely on supervised learning.
Reports showcase that DeepSeek’s model shows competitive performances against established large language models (LLMs) such as Anthropic’s Claude and OpenAI’s GPT-4o on established benchmarks in language understanding, mathematics (AIME 2024) and programming (Codeforces) while trained at a fraction of the costs.
Additionally, running inference shows significantly lower costs, leading to DeepSeek surpassing ChatGPT as the most downloaded free app on the American iOS App Store.
This development contributed to a nearly 17% drop in Nvidia’s share price, resulting in the most significant one-day loss in U.
S.
history, amounting to nearly $600 billion.
The open-source models also bring a significant shift in the healthcare system, allowing cost-efficient medical LLMs to be deployed within hospital networks.
To understand its performance in the healthcare sector, we analyse the new DeepSeek-R1 model on the United States Medical Licensing Examination (USMLE) and compare it to ChatGPT.

Related Results

Educators’ Perspectives on DeepSeek in ELT: A Qualitative Case Study of Pedagogical Potentials and Pitfalls in Chinese Higher Education
Educators’ Perspectives on DeepSeek in ELT: A Qualitative Case Study of Pedagogical Potentials and Pitfalls in Chinese Higher Education
Aim/Purpose: This study aimed to investigate the perspectives of English Language Teaching (ELT) educators on DeepSeek, emphasizing its pedagogical value, practical challenges, and...
Are USMLE Scores Valid Measures for Chief Resident Selection?
Are USMLE Scores Valid Measures for Chief Resident Selection?
ABSTRACT Background The US Medical Licensing Examination (USMLE) Step 1 and Step 2 scores are often used to inform a vari...
A Survey of DeepSeek Models
A Survey of DeepSeek Models
Advances in artificial intelligence (AI) rely on systems capable of human-like reasoning, a limitation for conventional Large Language Models (LLMs), which struggle with multi-step...
Factors Associated with Infectious Diseases Fellowship Academic Success
Factors Associated with Infectious Diseases Fellowship Academic Success
Abstract Background: A multitude of factors are considered in an infectious diseases (ID) training program’s meticulous selection process of ID fellows but their correlatio...
Factors That Predict Success in the General Surgery Match With a Pass/Fail USMLE Step 1 Exam
Factors That Predict Success in the General Surgery Match With a Pass/Fail USMLE Step 1 Exam
Introduction Matching into a residency program is an intricate process, including holistic application review, selection for interview, and ranking of interview...

Back to Top