Javascript must be enabled to continue!
Automatic Detection and Extraction of Key Resources from Tables in Biomedical Papers
View through CrossRef
Abstract
Tables are useful information artifacts that allow easy detection of data “missingness” by humans and have been deployed by several publishers to improve the amount of information present for key resources and reagents such as antibodies, cell lines, and other tools that constitute the inputs to a study. The STAR*Methods tables, specifically, have increased the “findability” of these key resources, but they have not been commonly available outside of the Cell Press journal family. To improve the availability of these tables in the broader biomedical literature, we have attempted to automatically process BioRxiv preprints to create tables from text or to recognize tables already created by authors and structure them for later use by publishers and search systems, to improve “findability” of resources in a larger amount of the scientific literature. The extraction of key resource tables in PDF files by the best in class tools resulted in Grid Table Similarity (GriTS) score of 0.12, so we have created several multimodal pipelines employing machine learning approaches for key resource table page identification, Table Transformer models for table detection and table structure recognition and a new table-specific language model for row over-segmentation to improve the extraction of text in tables created by biomedical authors and published on BioRxiv to around GriTS score of 0.90 enabling the deployment of automated research resource extraction tools onto BioRxiv.
Author summary
Tables are useful information artifacts that allow for easy detection of data “missingness” by humans and have been implemented by several publishers to improve the amount of information present for key resources and reagents such as antibodies, cell lines, and other tools that constitute the inputs to a study. To improve the availability of these tables in the broader biomedical literature, we introduced four pipelines for key resource table extraction from biomedical documents in PDF format. Our approach reconstructs key resource tables using image level table detection and structure detection generated table boundary, column (and row) bounding box information together with PDF text alignment. To remedy row over-segmentation resulting from overflowing table cell contents, we introduced a language modeling (LM) based row merging solution where a character-level generative pre-trained transformer (GPT) model was pre-trained on more than 11 million scientific table contents from PubMed Central Open Access Subset (PMC OAS). All introduced pipelines significantly outperformed GROBID baseline while our Table LM based row merging based pipeline, significantly outperformed all other pipelines including our OCR based pipeline.
Title: Automatic Detection and Extraction of Key Resources from Tables in Biomedical Papers
Description:
Abstract
Tables are useful information artifacts that allow easy detection of data “missingness” by humans and have been deployed by several publishers to improve the amount of information present for key resources and reagents such as antibodies, cell lines, and other tools that constitute the inputs to a study.
The STAR*Methods tables, specifically, have increased the “findability” of these key resources, but they have not been commonly available outside of the Cell Press journal family.
To improve the availability of these tables in the broader biomedical literature, we have attempted to automatically process BioRxiv preprints to create tables from text or to recognize tables already created by authors and structure them for later use by publishers and search systems, to improve “findability” of resources in a larger amount of the scientific literature.
The extraction of key resource tables in PDF files by the best in class tools resulted in Grid Table Similarity (GriTS) score of 0.
12, so we have created several multimodal pipelines employing machine learning approaches for key resource table page identification, Table Transformer models for table detection and table structure recognition and a new table-specific language model for row over-segmentation to improve the extraction of text in tables created by biomedical authors and published on BioRxiv to around GriTS score of 0.
90 enabling the deployment of automated research resource extraction tools onto BioRxiv.
Author summary
Tables are useful information artifacts that allow for easy detection of data “missingness” by humans and have been implemented by several publishers to improve the amount of information present for key resources and reagents such as antibodies, cell lines, and other tools that constitute the inputs to a study.
To improve the availability of these tables in the broader biomedical literature, we introduced four pipelines for key resource table extraction from biomedical documents in PDF format.
Our approach reconstructs key resource tables using image level table detection and structure detection generated table boundary, column (and row) bounding box information together with PDF text alignment.
To remedy row over-segmentation resulting from overflowing table cell contents, we introduced a language modeling (LM) based row merging solution where a character-level generative pre-trained transformer (GPT) model was pre-trained on more than 11 million scientific table contents from PubMed Central Open Access Subset (PMC OAS).
All introduced pipelines significantly outperformed GROBID baseline while our Table LM based row merging based pipeline, significantly outperformed all other pipelines including our OCR based pipeline.
Related Results
Water Trash Collector
Water Trash Collector
In today day to day life, approximately 71% of the Earth's surface is covered by Without affecting significant role that technology plays in our modern world, environmental and wat...
Investigation on Mechanical Properties of X80 Pipeline Girth Weld Welded by Semi-Automatic and Automatic Welding
Investigation on Mechanical Properties of X80 Pipeline Girth Weld Welded by Semi-Automatic and Automatic Welding
Abstract
The traditional manual welding in pipeline construction is being gradually replaced by semi-automatic and automatic welding in China. Semi-automatic welding...
Pengaruh Tri-n-Oktil Posfin Oksida dan Tingkat Ekstraksi pada Pemurnian Konsentrat Thorium
Pengaruh Tri-n-Oktil Posfin Oksida dan Tingkat Ekstraksi pada Pemurnian Konsentrat Thorium
Telah dilakukan ekstraksi konsentrat thorium oksalat hasil olah monasit memakai ekstraktan Tri – n - Oktil Posfin Oksida (TOPO). Pengotor yang paling banyak terkandung dalam kon...
Utilizing Large Language Models for Geoscience Literature Information Extraction
Utilizing Large Language Models for Geoscience Literature Information Extraction
Extracting information from unstructured and semi-structured geoscience literature is a crucial step in conducting geological research. The traditional machine learning extraction ...
Incremental prognostic value of fully automatic LVEF measured at stress using machine learning
Incremental prognostic value of fully automatic LVEF measured at stress using machine learning
Abstract
Background
Cardiovascular magnetic resonance (CMR) is the gold standard to measure left ventricular ejection fraction (...
Optimization of ultrasonic extraction of
Lycium barbarum
polysaccharides using response surface methodology
Optimization of ultrasonic extraction of
Lycium barbarum
polysaccharides using response surface methodology
Abstract
Ultrasonic extraction was a new development method to achieve high-efficiency extraction of
Lycium barbarum
...
Biomedical Engineering International joins the Family of Platinum Open Access Journals
Biomedical Engineering International joins the Family of Platinum Open Access Journals
We are delightfully announcing the launch of Biomedical Engineering International, a new interdisciplinary international scholarly open-access journal dedicated to publishing origi...
Extraction of astaxanthin from Haematococcus pluvialis
Extraction of astaxanthin from Haematococcus pluvialis
The astaxanthin extraction work comprised two main objectives. The first objective described the fundamental information of solubility of astaxanthin in supercritical carbon dioxid...

