Javascript must be enabled to continue!
Human Versus Artificial Intelligence: Comparing Cochrane Authors' and ChatGPT's Risk of Bias Assessments
View through CrossRef
ABSTRACT
Introduction
Systematic reviews and meta‐analyses synthesize randomized trial data to guide clinical decisions but require significant time and resources. Artificial intelligence (AI) offers a promising solution to streamline evidence synthesis, aiding study selection, data extraction, and risk of bias assessment. This study aims to evaluate the performance of ChatGPT‐4o in assessing the risk of bias in randomised controlled trials (RCTs) using the Risk of Bias 2 (RoB 2) tool, comparing its results with those conducted by human reviewers in Cochrane Reviews.
Methods
A sample of Cochrane Reviews utilizing the RoB 2 tool was identified through the Cochrane Database of Systematic Reviews (CDSR). Protocols, qualitative systematic reviews, and reviews employing alternative risk of bias assessment tools were excluded. The study utilized ChatGPT‐4o to assess the risk of bias using a structured set of prompts corresponding to the RoB 2 domains. The agreement between ChatGPT‐4o and consensus‐based human reviewer assessments was evaluated using weighted kappa statistics. Additionally, accuracy, sensitivity, specificity, positive predictive value, and negative predictive value were calculated. All analyses were performed using R Studio (version 4.3.0).
Results
A total of 42 Cochrane Reviews were screened, yielding a final sample of eight eligible reviews comprising 84 RCTs. The primary outcome of each included review was selected for risk of bias assessment. ChatGPT‐4o demonstrated moderate agreement with human reviewers for the overall risk of bias judgments (weighted kappa = 0.51, 95% CI: 0.36–0.66). Agreement varied across domains, ranging from fair (
κ
= 0.20 for selection of the reported results) to moderate (
κ
= 0.59 for measurement of outcomes). ChatGPT‐4o exhibited a sensitivity of 53% for identifying high‐risk studies and a specificity of 99% for classifying low‐risk studies.
Conclusion
This study shows that ChatGPT‐4o can perform risk of bias assessments using RoB 2 with fair to moderate agreement with human reviewers. While AI‐assisted risk of bias assessment remains imperfect, advancements in prompt engineering and model refinement may enhance performance. Future research should explore standardised prompts and investigate interrater reliability among human reviewers to provide a more robust comparison.
Title: Human Versus Artificial Intelligence: Comparing Cochrane Authors' and ChatGPT's Risk of Bias Assessments
Description:
ABSTRACT
Introduction
Systematic reviews and meta‐analyses synthesize randomized trial data to guide clinical decisions but require significant time and resources.
Artificial intelligence (AI) offers a promising solution to streamline evidence synthesis, aiding study selection, data extraction, and risk of bias assessment.
This study aims to evaluate the performance of ChatGPT‐4o in assessing the risk of bias in randomised controlled trials (RCTs) using the Risk of Bias 2 (RoB 2) tool, comparing its results with those conducted by human reviewers in Cochrane Reviews.
Methods
A sample of Cochrane Reviews utilizing the RoB 2 tool was identified through the Cochrane Database of Systematic Reviews (CDSR).
Protocols, qualitative systematic reviews, and reviews employing alternative risk of bias assessment tools were excluded.
The study utilized ChatGPT‐4o to assess the risk of bias using a structured set of prompts corresponding to the RoB 2 domains.
The agreement between ChatGPT‐4o and consensus‐based human reviewer assessments was evaluated using weighted kappa statistics.
Additionally, accuracy, sensitivity, specificity, positive predictive value, and negative predictive value were calculated.
All analyses were performed using R Studio (version 4.
3.
0).
Results
A total of 42 Cochrane Reviews were screened, yielding a final sample of eight eligible reviews comprising 84 RCTs.
The primary outcome of each included review was selected for risk of bias assessment.
ChatGPT‐4o demonstrated moderate agreement with human reviewers for the overall risk of bias judgments (weighted kappa = 0.
51, 95% CI: 0.
36–0.
66).
Agreement varied across domains, ranging from fair (
κ
= 0.
20 for selection of the reported results) to moderate (
κ
= 0.
59 for measurement of outcomes).
ChatGPT‐4o exhibited a sensitivity of 53% for identifying high‐risk studies and a specificity of 99% for classifying low‐risk studies.
Conclusion
This study shows that ChatGPT‐4o can perform risk of bias assessments using RoB 2 with fair to moderate agreement with human reviewers.
While AI‐assisted risk of bias assessment remains imperfect, advancements in prompt engineering and model refinement may enhance performance.
Future research should explore standardised prompts and investigate interrater reliability among human reviewers to provide a more robust comparison.
Related Results
Exploring Large Language Models Integration in the Histopathologic Diagnosis of Skin Diseases: A Comparative Study
Exploring Large Language Models Integration in the Histopathologic Diagnosis of Skin Diseases: A Comparative Study
Abstract
Introduction
The exact manner in which large language models (LLMs) will be integrated into pathology is not yet fully comprehended. This study examines the accuracy, bene...
Assessment of Chat-GPT, Gemini, and Perplexity in Principle of Research Publication: A Comparative Study
Assessment of Chat-GPT, Gemini, and Perplexity in Principle of Research Publication: A Comparative Study
Abstract
Introduction
Many researchers utilize artificial intelligence (AI) to aid their research endeavors. This study seeks to assess and contrast the performance of three sophis...
ChatGPT's Capabilities for Use in Anatomy Education and Anatomy Research
ChatGPT's Capabilities for Use in Anatomy Education and Anatomy Research
Dear Editors,
Recently, the discussion of an artificial intelligence (AI) - fueled platform in several articles in your journal has attracted the attention of many researchers [1, ...
Unlocking Educational Potential: Exploring Students’ Satisfaction and Sustainable Engagement with ChatGPT Using the ECM Model
Unlocking Educational Potential: Exploring Students’ Satisfaction and Sustainable Engagement with ChatGPT Using the ECM Model
Aim/Purpose: The main goal of this study is to investigate the factors affecting students’ satisfaction and continuous usage of ChatGPT in an educational context, using the Expecta...
Citation of updated and co-published Cochrane Methodology Reviews
Citation of updated and co-published Cochrane Methodology Reviews
Abstract
Background To evaluate the number of citations for Cochrane Methodology Reviews after they have been updated or co-published in another journal.
Methods We identif...
Appearance of ChatGPT and English Study
Appearance of ChatGPT and English Study
The purpose of this study is to examine the definition and characteristics of ChatGPT in order to present the direction of self-directed learning to learners, and to explore the po...
User Intentions to Use ChatGPT for Self-Diagnosis and Health-Related Purposes: Cross-sectional Survey Study (Preprint)
User Intentions to Use ChatGPT for Self-Diagnosis and Health-Related Purposes: Cross-sectional Survey Study (Preprint)
BACKGROUND
With the rapid advancement of artificial intelligence (AI) technologies, AI-powered chatbots, such as Chat Generative Pretrained Transformer (Cha...
ChatGPT: "To be or not to be" ... in academic research. The human mind's analytical rigor and capacity to discriminate between AI bots' truths and hallucinations
ChatGPT: "To be or not to be" ... in academic research. The human mind's analytical rigor and capacity to discriminate between AI bots' truths and hallucinations
Background. ChatGPT can generate increasingly realistic language, but the correctness and integrity of implementing these models in scientific papers remain unknown.
Recently publ...

