Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

GPT-4 shows comparable performance to human examiners in ranking open-text answers

View through CrossRef
Abstract Can GPT-4 replace human examiners? To address this question we explore the performance of GPT-4 as an examiner of answers to open-text questions. We formulate questions and sample solutions in the field of macroeconomics and collect answers from cohorts of undergraduate students. We then conduct a fair competition between GPT-4 and human experts, employing their expertise to assess the quality of the answers. We observe that the substitution of GPT-4 for a human examiner does not decrease inter-rater reliability on tasks that rank the quality of answers. We run checks on potential biases (whether GPT-4 prefers AI-generated or lengthy answers). We find no consistent evidence of such biases. Our findings are robust to tilting the competition to one side’s advantage, by using inferior or advanced prompting strategies. Our results are more attenuated on tasks where GPT-4 assigns points to student answers. Here, GPT-4 shows a bias towards longer answers. Overall, our study cautiously supports the utilization of GPT-4 as an assistant for automated grading systems, particularly those where answers are ranked according to their quality.
Title: GPT-4 shows comparable performance to human examiners in ranking open-text answers
Description:
Abstract Can GPT-4 replace human examiners? To address this question we explore the performance of GPT-4 as an examiner of answers to open-text questions.
We formulate questions and sample solutions in the field of macroeconomics and collect answers from cohorts of undergraduate students.
We then conduct a fair competition between GPT-4 and human experts, employing their expertise to assess the quality of the answers.
We observe that the substitution of GPT-4 for a human examiner does not decrease inter-rater reliability on tasks that rank the quality of answers.
We run checks on potential biases (whether GPT-4 prefers AI-generated or lengthy answers).
We find no consistent evidence of such biases.
Our findings are robust to tilting the competition to one side’s advantage, by using inferior or advanced prompting strategies.
Our results are more attenuated on tasks where GPT-4 assigns points to student answers.
Here, GPT-4 shows a bias towards longer answers.
Overall, our study cautiously supports the utilization of GPT-4 as an assistant for automated grading systems, particularly those where answers are ranked according to their quality.

Related Results

Performance of Novel GPT-4 in Otolaryngology Knowledge Assessment
Performance of Novel GPT-4 in Otolaryngology Knowledge Assessment
Abstract Purpose GPT-4, recently released by OpenAI, improves upon GPT-3.5 with increased reliability and expanded capabilities, including user-spec...
Sleep Habits and Occurrence of Lowback Pain among Craftsmen
Sleep Habits and Occurrence of Lowback Pain among Craftsmen
<span style="color: #000000; font-family: Verdana, Arial, Helvetica, sans-serif; font-size: 10px; font-style: normal; font-variant-ligatures: normal; font-variant-caps: normal; ...
Sleep Habits and Occurrence of Lowback Pain among Craftsmen
Sleep Habits and Occurrence of Lowback Pain among Craftsmen
<span style="color: #000000; font-family: Verdana, Arial, Helvetica, sans-serif; font-size: 10px; font-style: normal; font-variant-ligatures: normal; font-variant-caps: normal; ...
Analisis Penggunaan GPT dalam Pembelajaran Klinik Optik I di ARO Gapopin
Analisis Penggunaan GPT dalam Pembelajaran Klinik Optik I di ARO Gapopin
Perkembangan teknologi kecerdasan buatan (Artificial Intelligence/AI), khususnya model bahasa besar seperti Generative Pre-trained Transformer (GPT), telah membawa transformasi bes...
Performance of ChatGPT in Ophthalmic Registration and Clinical Diagnosis: Cross-Sectional Study (Preprint)
Performance of ChatGPT in Ophthalmic Registration and Clinical Diagnosis: Cross-Sectional Study (Preprint)
BACKGROUND Artificial intelligence (AI) chatbots such as ChatGPT are expected to impact vision health care significantly. Their potential to optimize the co...

Back to Top