Javascript must be enabled to continue!
Natural Language Processing for Clinical Laboratory Data Repository Systems: Implementation and Evaluation for Respiratory Viruses
View through CrossRef
Abstract
Background
With the growing volume and complexity of laboratory repositories, it has become tedious to parse unstructured data into structured and tabulated formats for secondary uses such as decision support, quality assurance, and outcome analysis. However, advances in Natural Language Processing (NLP) approaches have enabled efficient and automated extraction of clinically meaningful medical concepts from unstructured reports.
Objective
In this study, we aimed to determine the feasibility of using the NLP model for information extraction as an alternative approach to a time-consuming and operationally resource-intensive handcrafted rule-based tool. Therefore, we sought to develop and evaluate a deep learning-based NLP model to derive knowledge and extract information from text-based laboratory reports sourced from a provincial laboratory repository system.
Methods
The NLP model, a hierarchical multi-label classifier, was trained on a corpus of laboratory reports covering testing for 14 different respiratory viruses and viral subtypes. The corpus included 85
k
unique laboratory reports annotated by eight Subject Matter Experts (SME). The model’s performance stability and variation were analyzed across fine-grained and coarse-grained classes. Moreover, the model’s generalizability was also evaluated internally and externally on various test sets.
Results
The NLP model was trained several times with random initialization on the development corpus, and the results of the top ten best-performing models are presented in this paper. Overall, the NLP model performed well on internal, out-of-time (pre-COVID-19), and external (different laboratories) test sets with micro-averaged F1 scores >94% across all classes. Higher Precision and Recall scores with less variability were observed for the internal and pre-COVID-19 test sets. As expected, the model’s performance varied across categories and virus types due to the imbalanced nature of the corpus and sample sizes per class. There were intrinsically fewer classes of viruses being
detected
than those
tested
; therefore, the model’s performance (lowest F1-score of 57%) was noticeably lower in the “
detected
” cases.
Conclusions
We demonstrated that deep learning-based NLP models are promising solutions for information extraction from text-based laboratory reports. These approaches enable scalable, timely, and practical access to high-quality and encoded laboratory data if integrated into laboratory information system repositories.
Title: Natural Language Processing for Clinical Laboratory Data Repository Systems: Implementation and Evaluation for Respiratory Viruses
Description:
Abstract
Background
With the growing volume and complexity of laboratory repositories, it has become tedious to parse unstructured data into structured and tabulated formats for secondary uses such as decision support, quality assurance, and outcome analysis.
However, advances in Natural Language Processing (NLP) approaches have enabled efficient and automated extraction of clinically meaningful medical concepts from unstructured reports.
Objective
In this study, we aimed to determine the feasibility of using the NLP model for information extraction as an alternative approach to a time-consuming and operationally resource-intensive handcrafted rule-based tool.
Therefore, we sought to develop and evaluate a deep learning-based NLP model to derive knowledge and extract information from text-based laboratory reports sourced from a provincial laboratory repository system.
Methods
The NLP model, a hierarchical multi-label classifier, was trained on a corpus of laboratory reports covering testing for 14 different respiratory viruses and viral subtypes.
The corpus included 85
k
unique laboratory reports annotated by eight Subject Matter Experts (SME).
The model’s performance stability and variation were analyzed across fine-grained and coarse-grained classes.
Moreover, the model’s generalizability was also evaluated internally and externally on various test sets.
Results
The NLP model was trained several times with random initialization on the development corpus, and the results of the top ten best-performing models are presented in this paper.
Overall, the NLP model performed well on internal, out-of-time (pre-COVID-19), and external (different laboratories) test sets with micro-averaged F1 scores >94% across all classes.
Higher Precision and Recall scores with less variability were observed for the internal and pre-COVID-19 test sets.
As expected, the model’s performance varied across categories and virus types due to the imbalanced nature of the corpus and sample sizes per class.
There were intrinsically fewer classes of viruses being
detected
than those
tested
; therefore, the model’s performance (lowest F1-score of 57%) was noticeably lower in the “
detected
” cases.
Conclusions
We demonstrated that deep learning-based NLP models are promising solutions for information extraction from text-based laboratory reports.
These approaches enable scalable, timely, and practical access to high-quality and encoded laboratory data if integrated into laboratory information system repositories.
Related Results
Hubungan Perilaku Pola Makan dengan Kejadian Anak Obesitas
Hubungan Perilaku Pola Makan dengan Kejadian Anak Obesitas
<p><em><span style="font-size: 11.0pt; font-family: 'Times New Roman',serif; mso-fareast-font-family: 'Times New Roman'; mso-ansi-language: EN-US; mso-fareast-langua...
Učinak poučavanja razrednomu jeziku u izobrazbi nastavnika njemačkoga
Učinak poučavanja razrednomu jeziku u izobrazbi nastavnika njemačkoga
The actual use of classroom language is principally limited to the classroom environment. As far as foreign language learning is concerned, the classroom often turns out to be the ...
Globally Findable Planetary Data: The Interdisciplinary TRR170-DB Repository
Globally Findable Planetary Data: The Interdisciplinary TRR170-DB Repository
Introduction: The TRR170-DB data repository (https://planetary-data-portal.org/) manages the research data from the collaborative research center ‘Late Accretion onto Ter...
Increased life expectancy of heart failure patients in a rural center by a multidisciplinary program
Increased life expectancy of heart failure patients in a rural center by a multidisciplinary program
Abstract
Funding Acknowledgements
Type of funding sources: None.
INTRODUCTION Patients with heart failure (HF)...
Non-Recommended Publishing Lists: Strategies for Detecting Deceitful Journals
Non-Recommended Publishing Lists: Strategies for Detecting Deceitful Journals
Abstract
The rapid growth of open access publishing (OAP) has significantly improved the accessibility and dissemination of scientific knowledge. However, this expansion has also c...
INCIDENCE OF OTHER RESPIRATORY VIRUSES IN BRAZIL DURING SARS-COV2 PANDEMIC
INCIDENCE OF OTHER RESPIRATORY VIRUSES IN BRAZIL DURING SARS-COV2 PANDEMIC
OBJETIVE: Viruses are commonly associated with respiratory infections. Pandemics caused by respiratory viruses have affected humans considerably throughout history. We are currentl...
Phylogenomic analysis of Uganda influenza type-A viruses to assess their relatedness to the vaccine strains and other Africa viruses: a molecular epidemiology study
Phylogenomic analysis of Uganda influenza type-A viruses to assess their relatedness to the vaccine strains and other Africa viruses: a molecular epidemiology study
ABSTRACT
Background
Genetic characterisation of circulating influenza viruses is essential for vaccine selection and mitigation...
Pilarowski–Bjornsson Syndrome with Congenital Heart Defect: A Case Report and Literature Review
Pilarowski–Bjornsson Syndrome with Congenital Heart Defect: A Case Report and Literature Review
Abstract
Introduction
Pilarowski–Bjornsson syndrome (PILBOS) is a rare autosomal dominant neurodevelopmental disorder caused by heterozygous variants in chromodomain helicase DNA-b...

