Javascript must be enabled to continue!
The Transformative Potential of Large Language Models in Mining Electronic Health Records Data: Content Analysis
View through CrossRef
Background
In this study, we evaluate the accuracy, efficiency, and cost-effectiveness of large language models in extracting and structuring information from free-text clinical reports, particularly in identifying and classifying patient comorbidities within oncology electronic health records. We specifically compare the performance of gpt-3.5-turbo-1106 and gpt-4-1106-preview models against that of specialized human evaluators.
Objective
We specifically compare the performance of gpt-3.5-turbo-1106 and gpt-4-1106-preview models against that of specialized human evaluators.
Methods
We implemented a script using the OpenAI application programming interface to extract structured information in JavaScript object notation format from comorbidities reported in 250 personal history reports. These reports were manually reviewed in batches of 50 by 5 specialists in radiation oncology. We compared the results using metrics such as sensitivity, specificity, precision, accuracy, F-value, κ index, and the McNemar test, in addition to examining the common causes of errors in both humans and generative pretrained transformer (GPT) models.
Results
The GPT-3.5 model exhibited slightly lower performance compared to physicians across all metrics, though the differences were not statistically significant (McNemar test, P=.79). GPT-4 demonstrated clear superiority in several key metrics (McNemar test, P<.001). Notably, it achieved a sensitivity of 96.8%, compared to 88.2% for GPT-3.5 and 88.8% for physicians. However, physicians marginally outperformed GPT-4 in precision (97.7% vs 96.8%). GPT-4 showed greater consistency, replicating the exact same results in 76% of the reports across 10 repeated analyses, compared to 59% for GPT-3.5, indicating more stable and reliable performance. Physicians were more likely to miss explicit comorbidities, while the GPT models more frequently inferred nonexplicit comorbidities, sometimes correctly, though this also resulted in more false positives.
Conclusions
This study demonstrates that, with well-designed prompts, the large language models examined can match or even surpass medical specialists in extracting information from complex clinical reports. Their superior efficiency in time and costs, along with easy integration with databases, makes them a valuable tool for large-scale data mining and real-world evidence generation.
Title: The Transformative Potential of Large Language Models in Mining Electronic Health Records Data: Content Analysis
Description:
Background
In this study, we evaluate the accuracy, efficiency, and cost-effectiveness of large language models in extracting and structuring information from free-text clinical reports, particularly in identifying and classifying patient comorbidities within oncology electronic health records.
We specifically compare the performance of gpt-3.
5-turbo-1106 and gpt-4-1106-preview models against that of specialized human evaluators.
Objective
We specifically compare the performance of gpt-3.
5-turbo-1106 and gpt-4-1106-preview models against that of specialized human evaluators.
Methods
We implemented a script using the OpenAI application programming interface to extract structured information in JavaScript object notation format from comorbidities reported in 250 personal history reports.
These reports were manually reviewed in batches of 50 by 5 specialists in radiation oncology.
We compared the results using metrics such as sensitivity, specificity, precision, accuracy, F-value, κ index, and the McNemar test, in addition to examining the common causes of errors in both humans and generative pretrained transformer (GPT) models.
Results
The GPT-3.
5 model exhibited slightly lower performance compared to physicians across all metrics, though the differences were not statistically significant (McNemar test, P=.
79).
GPT-4 demonstrated clear superiority in several key metrics (McNemar test, P<.
001).
Notably, it achieved a sensitivity of 96.
8%, compared to 88.
2% for GPT-3.
5 and 88.
8% for physicians.
However, physicians marginally outperformed GPT-4 in precision (97.
7% vs 96.
8%).
GPT-4 showed greater consistency, replicating the exact same results in 76% of the reports across 10 repeated analyses, compared to 59% for GPT-3.
5, indicating more stable and reliable performance.
Physicians were more likely to miss explicit comorbidities, while the GPT models more frequently inferred nonexplicit comorbidities, sometimes correctly, though this also resulted in more false positives.
Conclusions
This study demonstrates that, with well-designed prompts, the large language models examined can match or even surpass medical specialists in extracting information from complex clinical reports.
Their superior efficiency in time and costs, along with easy integration with databases, makes them a valuable tool for large-scale data mining and real-world evidence generation.
Related Results
Hubungan Perilaku Pola Makan dengan Kejadian Anak Obesitas
Hubungan Perilaku Pola Makan dengan Kejadian Anak Obesitas
<p><em><span style="font-size: 11.0pt; font-family: 'Times New Roman',serif; mso-fareast-font-family: 'Times New Roman'; mso-ansi-language: EN-US; mso-fareast-langua...
Učinak poučavanja razrednomu jeziku u izobrazbi nastavnika njemačkoga
Učinak poučavanja razrednomu jeziku u izobrazbi nastavnika njemačkoga
The actual use of classroom language is principally limited to the classroom environment. As far as foreign language learning is concerned, the classroom often turns out to be the ...
Light at the End of the Tunnel: Mining Justice and Health
Light at the End of the Tunnel: Mining Justice and Health
The mining industry provides valuable mined commodities and financial support for communities worldwide. Mining has become safer for workers. Significant injustices, however, are c...
Increased life expectancy of heart failure patients in a rural center by a multidisciplinary program
Increased life expectancy of heart failure patients in a rural center by a multidisciplinary program
Abstract
Funding Acknowledgements
Type of funding sources: None.
INTRODUCTION Patients with heart failure (HF)...
The Significance of Text Mining in Research: A Comprehensive Review
The Significance of Text Mining in Research: A Comprehensive Review
Text mining has emerged as a pivotal tool in various domains of research, revolutionizing the way scholars and scientists extract valuable insights from vast volumes of textual dat...
Platonic Relations
Platonic Relations
The loop is one of the primary means of structuration for electronic music from mainstream to avant-garde styles. Indeed, during forums at the recent 2002 AD Analogue 2 Digital eve...
Social Media Use in Neurology: An Analysis of Alzheimer's Information on TikTok with Emphasis on Role of Healthcare Professionals
Social Media Use in Neurology: An Analysis of Alzheimer's Information on TikTok with Emphasis on Role of Healthcare Professionals
Abstract
Introduction
Alzheimer's disease (AD) is the most common neurodegenerative cause of dementia. Social media has become a major source of information for patients and famili...
AN ASSESSMENT OF E-RECORDS READINESS AT THE MINISTRY OF LABOUR AND HOME AFFAIRS, GABORONE, BOTSWANA
AN ASSESSMENT OF E-RECORDS READINESS AT THE MINISTRY OF LABOUR AND HOME AFFAIRS, GABORONE, BOTSWANA
This study sought to assess electronic records (e-records) readiness at the Ministry of Labour and Home Affairs (MLHA), Gaborone, Botswana, within the purview of the implementation...

