Javascript must be enabled to continue!
GPT-4 shows comparable performance to human examiners in ranking open-text answers
View through CrossRef
Abstract
Can GPT-4 replace human examiners? To address this question we explore the performance of GPT-4 as an examiner of answers to open-text questions. We formulate questions and sample solutions in the field of macroeconomics and collect answers from cohorts of undergraduate students. We then conduct a fair competition between GPT-4 and human experts, employing their expertise to assess the quality of the answers. We observe that the substitution of GPT-4 for a human examiner does not decrease inter-rater reliability on tasks that rank the quality of answers. We run checks on potential biases (whether GPT-4 prefers AI-generated or lengthy answers). We find no consistent evidence of such biases. Our findings are robust to tilting the competition to one side’s advantage, by using inferior or advanced prompting strategies. Our results are more attenuated on tasks where GPT-4 assigns points to student answers. Here, GPT-4 shows a bias towards longer answers. Overall, our study cautiously supports the utilization of GPT-4 as an assistant for automated grading systems, particularly those where answers are ranked according to their quality.
Springer Science and Business Media LLC
Title: GPT-4 shows comparable performance to human examiners in ranking open-text answers
Description:
Abstract
Can GPT-4 replace human examiners? To address this question we explore the performance of GPT-4 as an examiner of answers to open-text questions.
We formulate questions and sample solutions in the field of macroeconomics and collect answers from cohorts of undergraduate students.
We then conduct a fair competition between GPT-4 and human experts, employing their expertise to assess the quality of the answers.
We observe that the substitution of GPT-4 for a human examiner does not decrease inter-rater reliability on tasks that rank the quality of answers.
We run checks on potential biases (whether GPT-4 prefers AI-generated or lengthy answers).
We find no consistent evidence of such biases.
Our findings are robust to tilting the competition to one side’s advantage, by using inferior or advanced prompting strategies.
Our results are more attenuated on tasks where GPT-4 assigns points to student answers.
Here, GPT-4 shows a bias towards longer answers.
Overall, our study cautiously supports the utilization of GPT-4 as an assistant for automated grading systems, particularly those where answers are ranked according to their quality.
Related Results
Performance of Novel GPT-4 in Otolaryngology Knowledge Assessment
Performance of Novel GPT-4 in Otolaryngology Knowledge Assessment
Abstract
Purpose
GPT-4, recently released by OpenAI, improves upon GPT-3.5 with increased reliability and expanded capabilities, including user-spec...
Sleep Habits and Occurrence of Lowback Pain among Craftsmen
Sleep Habits and Occurrence of Lowback Pain among Craftsmen
<span style="color: #000000; font-family: Verdana, Arial, Helvetica, sans-serif; font-size: 10px; font-style: normal; font-variant-ligatures: normal; font-variant-caps: normal; ...
Sleep Habits and Occurrence of Lowback Pain among Craftsmen
Sleep Habits and Occurrence of Lowback Pain among Craftsmen
<span style="color: #000000; font-family: Verdana, Arial, Helvetica, sans-serif; font-size: 10px; font-style: normal; font-variant-ligatures: normal; font-variant-caps: normal; ...
Diagnostic Accuracy of Vision-Language Models on Japanese Diagnostic Radiology, Nuclear Medicine, and Interventional Radiology Specialty Board Examinations
Diagnostic Accuracy of Vision-Language Models on Japanese Diagnostic Radiology, Nuclear Medicine, and Interventional Radiology Specialty Board Examinations
Abstract
Purpose
The performance of vision-language models (VLMs) with image interpretation capabilities, such as GPT-4 omni (G...
GPT-agents based on medical guidelines can improve the responsiveness and explainability of outcomes for traumatic brain injury rehabilitation
GPT-agents based on medical guidelines can improve the responsiveness and explainability of outcomes for traumatic brain injury rehabilitation
AbstractThis study explored the application of generative pre-trained transformer (GPT) agents based on medical guidelines using large language model (LLM) technology for traumatic...
Analisis Penggunaan GPT dalam Pembelajaran Klinik Optik I di ARO Gapopin
Analisis Penggunaan GPT dalam Pembelajaran Klinik Optik I di ARO Gapopin
Perkembangan teknologi kecerdasan buatan (Artificial Intelligence/AI), khususnya model bahasa besar seperti Generative Pre-trained Transformer (GPT), telah membawa transformasi bes...
Diagnostic accuracy of vision-language models on Japanese diagnostic radiology, nuclear medicine, and interventional radiology specialty board examinations
Diagnostic accuracy of vision-language models on Japanese diagnostic radiology, nuclear medicine, and interventional radiology specialty board examinations
Abstract
Purpose
The performance of vision-language models (VLMs) with image interpretation capabilities, such as GPT-4 o...
Performance of ChatGPT in Ophthalmic Registration and Clinical Diagnosis: Cross-Sectional Study (Preprint)
Performance of ChatGPT in Ophthalmic Registration and Clinical Diagnosis: Cross-Sectional Study (Preprint)
BACKGROUND
Artificial intelligence (AI) chatbots such as ChatGPT are expected to impact vision health care significantly. Their potential to optimize the co...

