Javascript must be enabled to continue!
ChatGPT takes the FCPS exam in Internal Medicine
View through CrossRef
ABSTRACT
Large language models (LLMs) have exhibited remarkable proficiency in clinical knowledge, encompassing diagnostic medicine, and have been tested on questions related to medical licensing examinations. ChatGPT has recently gained popularity because of its ability to generate human-like responses when presented with exam questions. It has been tested on multiple undergraduate and subspecialty exams and the results have been mixed. We aim to test ChatGPT on questions mirroring the standards of the FCPS exam, the highest medical qualification in Pakistan.
We used 111 randomly chosen MCQs of internal medicine of FCPS level in the form of a text prompt, thrice on 3 consecutive days. The average of the three answers was taken as the final response. The responses were recorded and compared to the answers given by subject experts. Agreement between the two was assessed using the Chi-square test and Cohen’s Kappa with 0.75 Kappa as an acceptable agreement. Univariate regression analysis was done for the effect of subspeciality, word count, and case scenarios in the success of ChatGPT.. Post-risk stratification chi-square and kappa statistics were applied.
ChatGPT 4.0 scored 73% (69%-74%). Although close to the passing criteria, it could not clear the FCPS exam. Question characteristics and subspecialties did not affect the ChatGPT responses statistically. ChatGPT shows a high concordance between its responses indicating sound knowledge and a high reliability.
This study’s findings underline the necessity for caution in over-reliance on AI for critical clinical decisions without human oversight. Creating specialized models tailored for medical education could provide a viable solution to this problem.
Author Summary
Artificial intelligence is the future of the world. Since the launch of ChatGPT in 2014, it become one of the most widely used application for people in all fields of life. A wave of excitement was felt among the medical community when the chatbot was announced to have cleared the USMLE exams. Here, we have tested ChatGPT on MCQs mirroring the standard of FCPS exam questions. The FCPS is the highest medical qualification in Pakistan. We found that with a vast data base, ChatGPT could not clear the exam in all of the three attempts taken by it. ChatGPT, however, scored a near passing score indicating a relatively sound knowledge.
We found ChatGPT to be a consistent LLM for complex medical scenarios faced by doctors in their daily lives irrespective of the subspecialty, length or word count of the questions. Although ChatGPT did not pass the FCPS exam, its answers displayed a high level of consistency, indicating a solid understanding of internal medicine. This demonstrates the potential of AI to support and improve medical education and healthcare services in near future.
Title: ChatGPT takes the FCPS exam in Internal Medicine
Description:
ABSTRACT
Large language models (LLMs) have exhibited remarkable proficiency in clinical knowledge, encompassing diagnostic medicine, and have been tested on questions related to medical licensing examinations.
ChatGPT has recently gained popularity because of its ability to generate human-like responses when presented with exam questions.
It has been tested on multiple undergraduate and subspecialty exams and the results have been mixed.
We aim to test ChatGPT on questions mirroring the standards of the FCPS exam, the highest medical qualification in Pakistan.
We used 111 randomly chosen MCQs of internal medicine of FCPS level in the form of a text prompt, thrice on 3 consecutive days.
The average of the three answers was taken as the final response.
The responses were recorded and compared to the answers given by subject experts.
Agreement between the two was assessed using the Chi-square test and Cohen’s Kappa with 0.
75 Kappa as an acceptable agreement.
Univariate regression analysis was done for the effect of subspeciality, word count, and case scenarios in the success of ChatGPT.
Post-risk stratification chi-square and kappa statistics were applied.
ChatGPT 4.
0 scored 73% (69%-74%).
Although close to the passing criteria, it could not clear the FCPS exam.
Question characteristics and subspecialties did not affect the ChatGPT responses statistically.
ChatGPT shows a high concordance between its responses indicating sound knowledge and a high reliability.
This study’s findings underline the necessity for caution in over-reliance on AI for critical clinical decisions without human oversight.
Creating specialized models tailored for medical education could provide a viable solution to this problem.
Author Summary
Artificial intelligence is the future of the world.
Since the launch of ChatGPT in 2014, it become one of the most widely used application for people in all fields of life.
A wave of excitement was felt among the medical community when the chatbot was announced to have cleared the USMLE exams.
Here, we have tested ChatGPT on MCQs mirroring the standard of FCPS exam questions.
The FCPS is the highest medical qualification in Pakistan.
We found that with a vast data base, ChatGPT could not clear the exam in all of the three attempts taken by it.
ChatGPT, however, scored a near passing score indicating a relatively sound knowledge.
We found ChatGPT to be a consistent LLM for complex medical scenarios faced by doctors in their daily lives irrespective of the subspecialty, length or word count of the questions.
Although ChatGPT did not pass the FCPS exam, its answers displayed a high level of consistency, indicating a solid understanding of internal medicine.
This demonstrates the potential of AI to support and improve medical education and healthcare services in near future.
Related Results
Exploring Large Language Models Integration in the Histopathologic Diagnosis of Skin Diseases: A Comparative Study
Exploring Large Language Models Integration in the Histopathologic Diagnosis of Skin Diseases: A Comparative Study
Abstract
Introduction
The exact manner in which large language models (LLMs) will be integrated into pathology is not yet fully comprehended. This study examines the accuracy, bene...
Assessment of Chat-GPT, Gemini, and Perplexity in Principle of Research Publication: A Comparative Study
Assessment of Chat-GPT, Gemini, and Perplexity in Principle of Research Publication: A Comparative Study
Abstract
Introduction
Many researchers utilize artificial intelligence (AI) to aid their research endeavors. This study seeks to assess and contrast the performance of three sophis...
Unlocking Educational Potential: Exploring Students’ Satisfaction and Sustainable Engagement with ChatGPT Using the ECM Model
Unlocking Educational Potential: Exploring Students’ Satisfaction and Sustainable Engagement with ChatGPT Using the ECM Model
Aim/Purpose: The main goal of this study is to investigate the factors affecting students’ satisfaction and continuous usage of ChatGPT in an educational context, using the Expecta...
ChatGPT's Capabilities for Use in Anatomy Education and Anatomy Research
ChatGPT's Capabilities for Use in Anatomy Education and Anatomy Research
Dear Editors,
Recently, the discussion of an artificial intelligence (AI) - fueled platform in several articles in your journal has attracted the attention of many researchers [1, ...
Assessment of Artificial Intelligence Chatbot Performance on the Canadian Otolaryngology and Head and Neck Surgery In-Training Exam: Insights from a Comparative Analysis (Preprint)
Assessment of Artificial Intelligence Chatbot Performance on the Canadian Otolaryngology and Head and Neck Surgery In-Training Exam: Insights from a Comparative Analysis (Preprint)
BACKGROUND
The introduction of large language models (LLM) has rapidly transformed the field of healthcare. Its performance, often compared to that of physi...
Appearance of ChatGPT and English Study
Appearance of ChatGPT and English Study
The purpose of this study is to examine the definition and characteristics of ChatGPT in order to present the direction of self-directed learning to learners, and to explore the po...
User Intentions to Use ChatGPT for Self-Diagnosis and Health-Related Purposes: Cross-sectional Survey Study (Preprint)
User Intentions to Use ChatGPT for Self-Diagnosis and Health-Related Purposes: Cross-sectional Survey Study (Preprint)
BACKGROUND
With the rapid advancement of artificial intelligence (AI) technologies, AI-powered chatbots, such as Chat Generative Pretrained Transformer (Cha...
Performance of
AI
‐Chatbots to Common Temporomandibular Joint Disorders (
TMDs
) Patient Queries: Accuracy, Completeness, Reliability and Readability
Performance of
AI
‐Chatbots to Common Temporomandibular Joint Disorders (
TMDs
) Patient Queries: Accuracy, Completeness, Reliability and Readability
ABSTRACT
TMDs are a common group of conditions affecting the temporomandibular joint (TMJ) often resulting from factors like injury, stress or teeth grinding. Thi...

