Javascript must be enabled to continue!
Efficiency vs. Depth: A Comparative Study of AI-generated and Human-synthesized Qualitative Analyses of Ugandan Women’s Experiences of Obstetric Fistula (Preprint)
View through CrossRef
BACKGROUND
Limited recent literature has evaluated the use of large language models (LLMs) in the qualitative analysis of health data; more research is needed to expand the generalizability of LLM use and to evaluate potential ethical considerations.
OBJECTIVE
Our research sought to (1) describe the process of using artificial intelligence (AI) to analyze qualitative in-depth interview data, (2) identify similarities and differences between the human and AI-generated analyses to compare the quality and rigor of the two techniques and describe the strengths and weaknesses of each approach, and (3) make recommendations regarding the bounds of ethics and the role of researcher bias in AI-assisted qualitative research.
METHODS
Nested within a larger mixed-methods study, 17 Ugandan women recovering from female genital fistula surgery participated in hour-long semi-structured interviews exploring their mental, physical, and overall health trajectories. Each interview lasted about an hour and was audio recorded. Following translation and transcription, the data underwent human analysis and analysis using Versa, a University of California San Francisco (UCSF) developed LLM powered by ChatGPT-4o. The AI analysis was conducted using two strategies: inductively (without a codebook) and deductively (using the human-developed codebook). Finally, the outputs were compared to evaluate the code frequency and alignment, thematic depth, analysis quality, and efficiency of each method.
RESULTS
A comparative analysis revealed significant thematic overlap between the human-synthesized and AI-generated outputs, though notable differences in granularity and efficiency emerged. When inductively coding, Versa identified 39 codes, whereas human researchers utilized a more expansive set of 54 codes. While Versa’s thematic analyses were generally accurate, human-synthesized themes were more robust. The disparity in efficiency was stark: the human analysis took approximately 15 hours to complete, while Versa produced the analysis in about 2.5 hours.
CONCLUSIONS
While Versa significantly expedited the initial coding phase, it still relied heavily on human researchers to create appropriate prompts and input all the data into the chat. Versa’s inability to replicate the narrative depth and description of human synthesis suggests that LLMs currently lack the interpretive sensitivity required to capture the lived experiences present in qualitative data. Further, the misrepresentation of participant excerpts as direct quotes by Versa presents a significant threat to research integrity and ethics. While Versa and similar LLM models can serve as a powerful assistant for efficiency, human-driven analysis is essential to maintain ethical and interpretive rigor. Using Versa alone does not currently yield a high-quality analysis; significant human engagement is needed.
CLINICALTRIAL
ClinicalTrials.gov NCT05437939; https://clinicaltrials.gov/study/NCT05437939
JMIR Publications Inc.
Title: Efficiency vs. Depth: A Comparative Study of AI-generated and Human-synthesized Qualitative Analyses of Ugandan Women’s Experiences of Obstetric Fistula (Preprint)
Description:
BACKGROUND
Limited recent literature has evaluated the use of large language models (LLMs) in the qualitative analysis of health data; more research is needed to expand the generalizability of LLM use and to evaluate potential ethical considerations.
OBJECTIVE
Our research sought to (1) describe the process of using artificial intelligence (AI) to analyze qualitative in-depth interview data, (2) identify similarities and differences between the human and AI-generated analyses to compare the quality and rigor of the two techniques and describe the strengths and weaknesses of each approach, and (3) make recommendations regarding the bounds of ethics and the role of researcher bias in AI-assisted qualitative research.
METHODS
Nested within a larger mixed-methods study, 17 Ugandan women recovering from female genital fistula surgery participated in hour-long semi-structured interviews exploring their mental, physical, and overall health trajectories.
Each interview lasted about an hour and was audio recorded.
Following translation and transcription, the data underwent human analysis and analysis using Versa, a University of California San Francisco (UCSF) developed LLM powered by ChatGPT-4o.
The AI analysis was conducted using two strategies: inductively (without a codebook) and deductively (using the human-developed codebook).
Finally, the outputs were compared to evaluate the code frequency and alignment, thematic depth, analysis quality, and efficiency of each method.
RESULTS
A comparative analysis revealed significant thematic overlap between the human-synthesized and AI-generated outputs, though notable differences in granularity and efficiency emerged.
When inductively coding, Versa identified 39 codes, whereas human researchers utilized a more expansive set of 54 codes.
While Versa’s thematic analyses were generally accurate, human-synthesized themes were more robust.
The disparity in efficiency was stark: the human analysis took approximately 15 hours to complete, while Versa produced the analysis in about 2.
5 hours.
CONCLUSIONS
While Versa significantly expedited the initial coding phase, it still relied heavily on human researchers to create appropriate prompts and input all the data into the chat.
Versa’s inability to replicate the narrative depth and description of human synthesis suggests that LLMs currently lack the interpretive sensitivity required to capture the lived experiences present in qualitative data.
Further, the misrepresentation of participant excerpts as direct quotes by Versa presents a significant threat to research integrity and ethics.
While Versa and similar LLM models can serve as a powerful assistant for efficiency, human-driven analysis is essential to maintain ethical and interpretive rigor.
Using Versa alone does not currently yield a high-quality analysis; significant human engagement is needed.
CLINICALTRIAL
ClinicalTrials.
gov NCT05437939; https://clinicaltrials.
gov/study/NCT05437939.
Related Results
Women’s knowledge of symptoms of obstetric fistula, experiences, and associated factors in Sierra Leone
Women’s knowledge of symptoms of obstetric fistula, experiences, and associated factors in Sierra Leone
Background
Obstetric fistula is a devastating childbirth condition that results from prolonged obstructed labour without timely medical intervention, leading to a tear between the ...
Obstetric fistula repair failure and its associated factors among women underwent repair in Yirgalem Hamlin fistula center, Sidama Regional State, Southern Ethiopia, 2021: a retrospective cross sectional study
Obstetric fistula repair failure and its associated factors among women underwent repair in Yirgalem Hamlin fistula center, Sidama Regional State, Southern Ethiopia, 2021: a retrospective cross sectional study
Abstract
Background
Obstetric fistula repair failure is a combination of unsuccessful fistula closure and/or incontinence following a successful clo...
Prevalence and associated risk factors for failed obstetric fistula repair in East African countries: A systematic review and meta-analysis
Prevalence and associated risk factors for failed obstetric fistula repair in East African countries: A systematic review and meta-analysis
Objective: Obstetric fistula repair failure is a combination of unsuccessful fistula closure and/or incontinence following a successful closure. There is an inconsistent finding on...
Knowledge of obstetric fistula and its associated factors among women of reproductive age in Northwestern Ethiopia: a community-based cross-sectional study
Knowledge of obstetric fistula and its associated factors among women of reproductive age in Northwestern Ethiopia: a community-based cross-sectional study
Abstract
Background
Obstetric fistula has been a major maternal health challenges in low and middle-income countries, especially in Ethiopia, due to...
Building a Country-Wide Fistula Treatment Network in Kenya: Results From the First Six Years (2014-2020)
Building a Country-Wide Fistula Treatment Network in Kenya: Results From the First Six Years (2014-2020)
Abstract
It is estimated that one million women worldwide live with untreated fistula, a devastating injury primarily caused by prolonged obstructed labor when women do not...
Primerjalna književnost na prelomu tisočletja
Primerjalna književnost na prelomu tisočletja
In a comprehensive and at times critical manner, this volume seeks to shed light on the development of events in Western (i.e., European and North American) comparative literature ...
Pregnant Prisoners in Shackles
Pregnant Prisoners in Shackles
Photo by niu niu on Unsplash
ABSTRACT
Shackling prisoners has been implemented as standard procedure when transporting prisoners in labor and during childbirth. This procedure ensu...
Awareness on presentation of obstetric fistula and associated factors among reproductive age women in south eastern zone of tigray,ethiopia,2020.cross sectional study
Awareness on presentation of obstetric fistula and associated factors among reproductive age women in south eastern zone of tigray,ethiopia,2020.cross sectional study
Abstract
Background
Worldwide, around one million girls and women are currently living with fistula. Less than 20,000 women with obstetric fistula are treated each year. L...

