Javascript must be enabled to continue!
Religious Bias Benchmarks for ChatGPT
View through CrossRef
The objectives of this study are to: 1) estimate the frequency of six types of biases in ChatGPT’s responses to religious belief-specific morality and ethics questions, 2) assess how those biases vary by religious belief and ChatGPT model version and 3) determine how model engineering techniques affect these biases. ChatGPT responses were collected from a set of 112 general morality and ethics questions, each individually tailored to five different belief systems: Zen Buddhism, Catholicism, Sunni Islam, Orthodox Judaism and secular humanism. The resulting questions were then posed ten times to various baseline and derivative ChatGPT version 3 and version 4 models with and without the application of prompt engineering. The final ChatGPT response dataset contained 45,920 query responses and over 11.4 million words of text. Analyses of this dataset showed that this dataset contained explicit biases, anthropomorphic biases, statement biases, framing biases and coverage biases, often in favor of Buddhism or secular humanism and/or against the Abrahamic religions. Three of the biases (explicit, coverage and framing bias) were mitigated by the more advanced GPT-4 models, but two biases (anthropomorphic, statement) were higher with GPT-4. Analysis of the sixth bias, information bias, was inconclusive, although a potential link was found between responses that contain unsafe speech and ChatGPT hallucinations and multi-lingual response errors. None of the model engineering approaches tested, persona assumption, N-shot engineering, model fine tuning or research assistants, was successful at eliminating all biases.
Title: Religious Bias Benchmarks for ChatGPT
Description:
The objectives of this study are to: 1) estimate the frequency of six types of biases in ChatGPT’s responses to religious belief-specific morality and ethics questions, 2) assess how those biases vary by religious belief and ChatGPT model version and 3) determine how model engineering techniques affect these biases.
ChatGPT responses were collected from a set of 112 general morality and ethics questions, each individually tailored to five different belief systems: Zen Buddhism, Catholicism, Sunni Islam, Orthodox Judaism and secular humanism.
The resulting questions were then posed ten times to various baseline and derivative ChatGPT version 3 and version 4 models with and without the application of prompt engineering.
The final ChatGPT response dataset contained 45,920 query responses and over 11.
4 million words of text.
Analyses of this dataset showed that this dataset contained explicit biases, anthropomorphic biases, statement biases, framing biases and coverage biases, often in favor of Buddhism or secular humanism and/or against the Abrahamic religions.
Three of the biases (explicit, coverage and framing bias) were mitigated by the more advanced GPT-4 models, but two biases (anthropomorphic, statement) were higher with GPT-4.
Analysis of the sixth bias, information bias, was inconclusive, although a potential link was found between responses that contain unsafe speech and ChatGPT hallucinations and multi-lingual response errors.
None of the model engineering approaches tested, persona assumption, N-shot engineering, model fine tuning or research assistants, was successful at eliminating all biases.
Related Results
Exploring Large Language Models Integration in the Histopathologic Diagnosis of Skin Diseases: A Comparative Study
Exploring Large Language Models Integration in the Histopathologic Diagnosis of Skin Diseases: A Comparative Study
Abstract
Introduction
The exact manner in which large language models (LLMs) will be integrated into pathology is not yet fully comprehended. This study examines the accuracy, bene...
Assessment of Chat-GPT, Gemini, and Perplexity in Principle of Research Publication: A Comparative Study
Assessment of Chat-GPT, Gemini, and Perplexity in Principle of Research Publication: A Comparative Study
Abstract
Introduction
Many researchers utilize artificial intelligence (AI) to aid their research endeavors. This study seeks to assess and contrast the performance of three sophis...
ChatGPT's Capabilities for Use in Anatomy Education and Anatomy Research
ChatGPT's Capabilities for Use in Anatomy Education and Anatomy Research
Dear Editors,
Recently, the discussion of an artificial intelligence (AI) - fueled platform in several articles in your journal has attracted the attention of many researchers [1, ...
Unlocking Educational Potential: Exploring Students’ Satisfaction and Sustainable Engagement with ChatGPT Using the ECM Model
Unlocking Educational Potential: Exploring Students’ Satisfaction and Sustainable Engagement with ChatGPT Using the ECM Model
Aim/Purpose: The main goal of this study is to investigate the factors affecting students’ satisfaction and continuous usage of ChatGPT in an educational context, using the Expecta...
Appearance of ChatGPT and English Study
Appearance of ChatGPT and English Study
The purpose of this study is to examine the definition and characteristics of ChatGPT in order to present the direction of self-directed learning to learners, and to explore the po...
Leveraging on Chatgpt, an Artificial Intelligence (AI) Tool to Transform Examination Writing in Higher Education
Leveraging on Chatgpt, an Artificial Intelligence (AI) Tool to Transform Examination Writing in Higher Education
Abstract
Purpose
The study explored how ChatGPT could transform examination writing in higher education. The research question was: How can the AI tool ChatGPT help transfo...
User Intentions to Use ChatGPT for Self-Diagnosis and Health-Related Purposes: Cross-sectional Survey Study (Preprint)
User Intentions to Use ChatGPT for Self-Diagnosis and Health-Related Purposes: Cross-sectional Survey Study (Preprint)
BACKGROUND
With the rapid advancement of artificial intelligence (AI) technologies, AI-powered chatbots, such as Chat Generative Pretrained Transformer (Cha...
Performance of
AI
‐Chatbots to Common Temporomandibular Joint Disorders (
TMDs
) Patient Queries: Accuracy, Completeness, Reliability and Readability
Performance of
AI
‐Chatbots to Common Temporomandibular Joint Disorders (
TMDs
) Patient Queries: Accuracy, Completeness, Reliability and Readability
ABSTRACT
TMDs are a common group of conditions affecting the temporomandibular joint (TMJ) often resulting from factors like injury, stress or teeth grinding. Thi...

