Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Head-to-head evaluation of ChatGPT, DeepSeek, and Perplexity on acid–base disorder case clinical management and drug treatment: Accuracy, domain performance, and response consistency assessment

View through CrossRef
Background Large language models (LLMs) are increasingly used in medical education, but their performance and reliability on mechanistically demanding topics like acid–base interpretation are unclear. Methods We conducted a cross-sectional head-to-head evaluation of three LLMs (ChatGPT, DeepSeek, Perplexity) using 75 textbook acid–base cases related to management and drug treatment. One obsolete case was excluded, leaving 74 vignettes (29 metabolic, 10 respiratory, 35 mixed) that generated 510 multiple-choice questions (MCQs). Each MCQ was mapped to case category, one of seven cognitive domains, and one of four distractor-based difficulty bands. All questions were posed via public web interfaces in two independent phases (Phase I and II). Primary outcome was accuracy (proportion correct); consistency outcomes were identical answers across phases and reproducibly correct answers (correct in both phases). Results Overall accuracy was 59.8% (305/510) for ChatGPT, 58.8% (300/510) for DeepSeek, and 53.7% (274/510) for Perplexity (Cochran’s Q p = 0.16); ChatGPT and DeepSeek each outperformed Perplexity (p = 0.0086 and p = 0.0374). In metabolic cases (164 items), Perplexity scored 50.6% versus 60.4% and 62.8% for ChatGPT and DeepSeek. Accuracy was highest for compensation/expected response (116 items; 70.7% for ChatGPT and DeepSeek, 62.1% for Perplexity) and lowest for metabolic acidosis subtypes (59 items; 52.5%, 42.4%, 32.2%). Accuracy declined with difficulty, with Perplexity significantly lower than the other models in moderate–hard and hard bands. Identical answers across phases occurred in 81.7% of items for ChatGPT, 81.1% for DeepSeek, and 76.1% for Perplexity, but reproducibly correct answers were 52.1%, 48.8%, and 43.2%. Conclusions ChatGPT and DeepSeek showed moderate performance and generally higher accuracy than Perplexity. Stable repeated errors despite only moderate reproducibly correct responses support cautious interpretation of LLM outputs and argue against their use as stand-alone tools for acid–base learning or clinical decision making of management and drug treatment .
Title: Head-to-head evaluation of ChatGPT, DeepSeek, and Perplexity on acid–base disorder case clinical management and drug treatment: Accuracy, domain performance, and response consistency assessment
Description:
Background Large language models (LLMs) are increasingly used in medical education, but their performance and reliability on mechanistically demanding topics like acid–base interpretation are unclear.
Methods We conducted a cross-sectional head-to-head evaluation of three LLMs (ChatGPT, DeepSeek, Perplexity) using 75 textbook acid–base cases related to management and drug treatment.
One obsolete case was excluded, leaving 74 vignettes (29 metabolic, 10 respiratory, 35 mixed) that generated 510 multiple-choice questions (MCQs).
Each MCQ was mapped to case category, one of seven cognitive domains, and one of four distractor-based difficulty bands.
All questions were posed via public web interfaces in two independent phases (Phase I and II).
Primary outcome was accuracy (proportion correct); consistency outcomes were identical answers across phases and reproducibly correct answers (correct in both phases).
Results Overall accuracy was 59.
8% (305/510) for ChatGPT, 58.
8% (300/510) for DeepSeek, and 53.
7% (274/510) for Perplexity (Cochran’s Q p = 0.
16); ChatGPT and DeepSeek each outperformed Perplexity (p = 0.
0086 and p = 0.
0374).
In metabolic cases (164 items), Perplexity scored 50.
6% versus 60.
4% and 62.
8% for ChatGPT and DeepSeek.
Accuracy was highest for compensation/expected response (116 items; 70.
7% for ChatGPT and DeepSeek, 62.
1% for Perplexity) and lowest for metabolic acidosis subtypes (59 items; 52.
5%, 42.
4%, 32.
2%).
Accuracy declined with difficulty, with Perplexity significantly lower than the other models in moderate–hard and hard bands.
Identical answers across phases occurred in 81.
7% of items for ChatGPT, 81.
1% for DeepSeek, and 76.
1% for Perplexity, but reproducibly correct answers were 52.
1%, 48.
8%, and 43.
2%.
Conclusions ChatGPT and DeepSeek showed moderate performance and generally higher accuracy than Perplexity.
Stable repeated errors despite only moderate reproducibly correct responses support cautious interpretation of LLM outputs and argue against their use as stand-alone tools for acid–base learning or clinical decision making of management and drug treatment .

Related Results

Assessment of Chat-GPT, Gemini, and Perplexity in Principle of Research Publication: A Comparative Study
Assessment of Chat-GPT, Gemini, and Perplexity in Principle of Research Publication: A Comparative Study
Abstract Introduction Many researchers utilize artificial intelligence (AI) to aid their research endeavors. This study seeks to assess and contrast the performance of three sophis...
Exploring Large Language Models Integration in the Histopathologic Diagnosis of Skin Diseases: A Comparative Study
Exploring Large Language Models Integration in the Histopathologic Diagnosis of Skin Diseases: A Comparative Study
Abstract Introduction The exact manner in which large language models (LLMs) will be integrated into pathology is not yet fully comprehended. This study examines the accuracy, bene...
Educators’ Perspectives on DeepSeek in ELT: A Qualitative Case Study of Pedagogical Potentials and Pitfalls in Chinese Higher Education
Educators’ Perspectives on DeepSeek in ELT: A Qualitative Case Study of Pedagogical Potentials and Pitfalls in Chinese Higher Education
Aim/Purpose: This study aimed to investigate the perspectives of English Language Teaching (ELT) educators on DeepSeek, emphasizing its pedagogical value, practical challenges, and...
Hydatid Disease of The Brain Parenchyma: A Systematic Review
Hydatid Disease of The Brain Parenchyma: A Systematic Review
Abstarct Introduction Isolated brain hydatid disease (BHD) is an extremely rare form of echinococcosis. A prompt and timely diagnosis is a crucial step in disease management. This ...
ChatGPT's Capabilities for Use in Anatomy Education and Anatomy Research
ChatGPT's Capabilities for Use in Anatomy Education and Anatomy Research
Dear Editors, Recently, the discussion of an artificial intelligence (AI) - fueled platform in several articles in your journal has attracted the attention of many researchers [1, ...
Unlocking Educational Potential: Exploring Students’ Satisfaction and Sustainable Engagement with ChatGPT Using the ECM Model
Unlocking Educational Potential: Exploring Students’ Satisfaction and Sustainable Engagement with ChatGPT Using the ECM Model
Aim/Purpose: The main goal of this study is to investigate the factors affecting students’ satisfaction and continuous usage of ChatGPT in an educational context, using the Expecta...
Small Cell Lung Cancer and Tarlatamab: A Meta-Analysis of Clinical Trials
Small Cell Lung Cancer and Tarlatamab: A Meta-Analysis of Clinical Trials
Abstract Introduction Tarlatamab is a Delta-like ligand 3 (DLL3) -directed bispecific T-cell engager recently approved for use in patients with advanced small cell lung cancer (SCL...

Back to Top