Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Automatic Detection and Extraction of Key Resources from Tables in Biomedical Papers

View through CrossRef
Abstract Tables are useful information artifacts that allow easy detection of data “missingness” by humans and have been deployed by several publishers to improve the amount of information present for key resources and reagents such as antibodies, cell lines, and other tools that constitute the inputs to a study. The STAR*Methods tables, specifically, have increased the “findability” of these key resources, but they have not been commonly available outside of the Cell Press journal family. To improve the availability of these tables in the broader biomedical literature, we have attempted to automatically process BioRxiv preprints to create tables from text or to recognize tables already created by authors and structure them for later use by publishers and search systems, to improve “findability” of resources in a larger amount of the scientific literature. The extraction of key resource tables in PDF files by the best in class tools resulted in Grid Table Similarity (GriTS) score of 0.12, so we have created several multimodal pipelines employing machine learning approaches for key resource table page identification, Table Transformer models for table detection and table structure recognition and a new table-specific language model for row over-segmentation to improve the extraction of text in tables created by biomedical authors and published on BioRxiv to around GriTS score of 0.90 enabling the deployment of automated research resource extraction tools onto BioRxiv. Author summary Tables are useful information artifacts that allow for easy detection of data “missingness” by humans and have been implemented by several publishers to improve the amount of information present for key resources and reagents such as antibodies, cell lines, and other tools that constitute the inputs to a study. To improve the availability of these tables in the broader biomedical literature, we introduced four pipelines for key resource table extraction from biomedical documents in PDF format. Our approach reconstructs key resource tables using image level table detection and structure detection generated table boundary, column (and row) bounding box information together with PDF text alignment. To remedy row over-segmentation resulting from overflowing table cell contents, we introduced a language modeling (LM) based row merging solution where a character-level generative pre-trained transformer (GPT) model was pre-trained on more than 11 million scientific table contents from PubMed Central Open Access Subset (PMC OAS). All introduced pipelines significantly outperformed GROBID baseline while our Table LM based row merging based pipeline, significantly outperformed all other pipelines including our OCR based pipeline.
Title: Automatic Detection and Extraction of Key Resources from Tables in Biomedical Papers
Description:
Abstract Tables are useful information artifacts that allow easy detection of data “missingness” by humans and have been deployed by several publishers to improve the amount of information present for key resources and reagents such as antibodies, cell lines, and other tools that constitute the inputs to a study.
The STAR*Methods tables, specifically, have increased the “findability” of these key resources, but they have not been commonly available outside of the Cell Press journal family.
To improve the availability of these tables in the broader biomedical literature, we have attempted to automatically process BioRxiv preprints to create tables from text or to recognize tables already created by authors and structure them for later use by publishers and search systems, to improve “findability” of resources in a larger amount of the scientific literature.
The extraction of key resource tables in PDF files by the best in class tools resulted in Grid Table Similarity (GriTS) score of 0.
12, so we have created several multimodal pipelines employing machine learning approaches for key resource table page identification, Table Transformer models for table detection and table structure recognition and a new table-specific language model for row over-segmentation to improve the extraction of text in tables created by biomedical authors and published on BioRxiv to around GriTS score of 0.
90 enabling the deployment of automated research resource extraction tools onto BioRxiv.
Author summary Tables are useful information artifacts that allow for easy detection of data “missingness” by humans and have been implemented by several publishers to improve the amount of information present for key resources and reagents such as antibodies, cell lines, and other tools that constitute the inputs to a study.
To improve the availability of these tables in the broader biomedical literature, we introduced four pipelines for key resource table extraction from biomedical documents in PDF format.
Our approach reconstructs key resource tables using image level table detection and structure detection generated table boundary, column (and row) bounding box information together with PDF text alignment.
To remedy row over-segmentation resulting from overflowing table cell contents, we introduced a language modeling (LM) based row merging solution where a character-level generative pre-trained transformer (GPT) model was pre-trained on more than 11 million scientific table contents from PubMed Central Open Access Subset (PMC OAS).
All introduced pipelines significantly outperformed GROBID baseline while our Table LM based row merging based pipeline, significantly outperformed all other pipelines including our OCR based pipeline.

Related Results

Utilizing Large Language Models for Geoscience Literature Information Extraction
Utilizing Large Language Models for Geoscience Literature Information Extraction
Extracting information from unstructured and semi-structured geoscience literature is a crucial step in conducting geological research. The traditional machine learning extraction ...
Investigation on Mechanical Properties of X80 Pipeline Girth Weld Welded by Semi-Automatic and Automatic Welding
Investigation on Mechanical Properties of X80 Pipeline Girth Weld Welded by Semi-Automatic and Automatic Welding
Abstract The traditional manual welding in pipeline construction is being gradually replaced by semi-automatic and automatic welding in China. Semi-automatic welding...
Incremental prognostic value of fully automatic LVEF measured at stress using machine learning
Incremental prognostic value of fully automatic LVEF measured at stress using machine learning
Abstract Background Cardiovascular magnetic resonance (CMR) is the gold standard to measure left ventricular ejection fraction (...
Biomedical Engineering International joins the Family of Platinum Open Access Journals
Biomedical Engineering International joins the Family of Platinum Open Access Journals
We are delightfully announcing the launch of Biomedical Engineering International, a new interdisciplinary international scholarly open-access journal dedicated to publishing origi...
ADJUSTMENT OF EXTRACTION PARAMETERS OF URTICA DIOCIA USING ADVANCED EXTRACTION METHODS
ADJUSTMENT OF EXTRACTION PARAMETERS OF URTICA DIOCIA USING ADVANCED EXTRACTION METHODS
The pharmaceutical raw material Urtica diocia requires proper preparation so that the extraction process is conducted without obstacles and enables reaching the highest extraction ...
The Management of Transalevolar Surgery Teeth with Pulpal Polyps Condition
The Management of Transalevolar Surgery Teeth with Pulpal Polyps Condition
Background: The failure of intra-alveolar tooth extraction in cases of tooth extraction with complications is generally resolved by performing transalveolar tooth extraction. The o...

Back to Top