Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Efficiency vs. Depth: A Comparative Study of AI-generated and Human-synthesized Qualitative Analyses of Ugandan Women’s Experiences of Obstetric Fistula (Preprint)

View through CrossRef
BACKGROUND Limited recent literature has evaluated the use of large language models (LLMs) in the qualitative analysis of health data; more research is needed to expand the generalizability of LLM use and to evaluate potential ethical considerations. OBJECTIVE Our research sought to (1) describe the process of using artificial intelligence (AI) to analyze qualitative in-depth interview data, (2) identify similarities and differences between the human and AI-generated analyses to compare the quality and rigor of the two techniques and describe the strengths and weaknesses of each approach, and (3) make recommendations regarding the bounds of ethics and the role of researcher bias in AI-assisted qualitative research. METHODS Nested within a larger mixed-methods study, 17 Ugandan women recovering from female genital fistula surgery participated in hour-long semi-structured interviews exploring their mental, physical, and overall health trajectories. Each interview lasted about an hour and was audio recorded. Following translation and transcription, the data underwent human analysis and analysis using Versa, a University of California San Francisco (UCSF) developed LLM powered by ChatGPT-4o. The AI analysis was conducted using two strategies: inductively (without a codebook) and deductively (using the human-developed codebook). Finally, the outputs were compared to evaluate the code frequency and alignment, thematic depth, analysis quality, and efficiency of each method. RESULTS A comparative analysis revealed significant thematic overlap between the human-synthesized and AI-generated outputs, though notable differences in granularity and efficiency emerged. When inductively coding, Versa identified 39 codes, whereas human researchers utilized a more expansive set of 54 codes. While Versa’s thematic analyses were generally accurate, human-synthesized themes were more robust. The disparity in efficiency was stark: the human analysis took approximately 15 hours to complete, while Versa produced the analysis in about 2.5 hours. CONCLUSIONS While Versa significantly expedited the initial coding phase, it still relied heavily on human researchers to create appropriate prompts and input all the data into the chat. Versa’s inability to replicate the narrative depth and description of human synthesis suggests that LLMs currently lack the interpretive sensitivity required to capture the lived experiences present in qualitative data. Further, the misrepresentation of participant excerpts as direct quotes by Versa presents a significant threat to research integrity and ethics. While Versa and similar LLM models can serve as a powerful assistant for efficiency, human-driven analysis is essential to maintain ethical and interpretive rigor. Using Versa alone does not currently yield a high-quality analysis; significant human engagement is needed. CLINICALTRIAL ClinicalTrials.gov NCT05437939; https://clinicaltrials.gov/study/NCT05437939
Title: Efficiency vs. Depth: A Comparative Study of AI-generated and Human-synthesized Qualitative Analyses of Ugandan Women’s Experiences of Obstetric Fistula (Preprint)
Description:
BACKGROUND Limited recent literature has evaluated the use of large language models (LLMs) in the qualitative analysis of health data; more research is needed to expand the generalizability of LLM use and to evaluate potential ethical considerations.
OBJECTIVE Our research sought to (1) describe the process of using artificial intelligence (AI) to analyze qualitative in-depth interview data, (2) identify similarities and differences between the human and AI-generated analyses to compare the quality and rigor of the two techniques and describe the strengths and weaknesses of each approach, and (3) make recommendations regarding the bounds of ethics and the role of researcher bias in AI-assisted qualitative research.
METHODS Nested within a larger mixed-methods study, 17 Ugandan women recovering from female genital fistula surgery participated in hour-long semi-structured interviews exploring their mental, physical, and overall health trajectories.
Each interview lasted about an hour and was audio recorded.
Following translation and transcription, the data underwent human analysis and analysis using Versa, a University of California San Francisco (UCSF) developed LLM powered by ChatGPT-4o.
The AI analysis was conducted using two strategies: inductively (without a codebook) and deductively (using the human-developed codebook).
Finally, the outputs were compared to evaluate the code frequency and alignment, thematic depth, analysis quality, and efficiency of each method.
RESULTS A comparative analysis revealed significant thematic overlap between the human-synthesized and AI-generated outputs, though notable differences in granularity and efficiency emerged.
When inductively coding, Versa identified 39 codes, whereas human researchers utilized a more expansive set of 54 codes.
While Versa’s thematic analyses were generally accurate, human-synthesized themes were more robust.
The disparity in efficiency was stark: the human analysis took approximately 15 hours to complete, while Versa produced the analysis in about 2.
5 hours.
CONCLUSIONS While Versa significantly expedited the initial coding phase, it still relied heavily on human researchers to create appropriate prompts and input all the data into the chat.
Versa’s inability to replicate the narrative depth and description of human synthesis suggests that LLMs currently lack the interpretive sensitivity required to capture the lived experiences present in qualitative data.
Further, the misrepresentation of participant excerpts as direct quotes by Versa presents a significant threat to research integrity and ethics.
While Versa and similar LLM models can serve as a powerful assistant for efficiency, human-driven analysis is essential to maintain ethical and interpretive rigor.
Using Versa alone does not currently yield a high-quality analysis; significant human engagement is needed.
CLINICALTRIAL ClinicalTrials.
gov NCT05437939; https://clinicaltrials.
gov/study/NCT05437939.

Related Results

Women’s knowledge of symptoms of obstetric fistula, experiences, and associated factors in Sierra Leone
Women’s knowledge of symptoms of obstetric fistula, experiences, and associated factors in Sierra Leone
Background Obstetric fistula is a devastating childbirth condition that results from prolonged obstructed labour without timely medical intervention, leading to a tear between the ...
Building a Country-Wide Fistula Treatment Network in Kenya: Results From the First Six Years (2014-2020)
Building a Country-Wide Fistula Treatment Network in Kenya: Results From the First Six Years (2014-2020)
Abstract It is estimated that one million women worldwide live with untreated fistula, a devastating injury primarily caused by prolonged obstructed labor when women do not...
Primerjalna književnost na prelomu tisočletja
Primerjalna književnost na prelomu tisočletja
In a comprehensive and at times critical manner, this volume seeks to shed light on the development of events in Western (i.e., European and North American) comparative literature ...
Pregnant Prisoners in Shackles
Pregnant Prisoners in Shackles
Photo by niu niu on Unsplash ABSTRACT Shackling prisoners has been implemented as standard procedure when transporting prisoners in labor and during childbirth. This procedure ensu...

Back to Top