Javascript must be enabled to continue!
A Pre-Training Technique to Localize Medical BERT and to Enhance Biomedical BERT
View through CrossRef
Abstract
Background: Pre-training large-scale neural language models on raw texts has been shown to make a significant contribution to a strategy for transfer learning in natural language processing (NLP). With the introduction of transformer-based language models, such as Bidirectional Encoder Representations from Transformers (BERT), the performance of information extraction from free text by NLP has significantly improved for both the general domain and the medical domain; however, it is difficult for languages in which there are few publicly available medical databases with a high quality and a large size to train medical BERT models that perform well.Method: We introduce a method to train a BERT model using a small medical corpus both in English and in Japanese. Our proposed method consists of two interventions: simultaneous pre-training, which is intended to encourage masked language modeling and next-sentence prediction on the small medical corpus, and amplified vocabulary, which helps with suiting the small corpus when building the customized corpus by byte-pair encoding. Moreover, we used whole PubMed abstracts and developed a high-performance BERT model, Bidirectional Encoder Representations from Transformers for Biomedical Text Mining by Osaka University (ouBioBERT), in English via our method. We then evaluated the performance of our BERT models and publicly available baselines and compared them.Results: We confirmed that our Japanese medical BERT outperforms conventional baselines and the other BERT models in terms of the medical-document classification task and that our English BERT pre-trained using both the general and medical domain corpora performs sufficiently for practical use in terms of the biomedical language understanding evaluation (BLUE) benchmark. Moreover, ouBioBERT shows that the total score of the BLUE benchmark is 1.1 points above that of BioBERT and 0.3 points above that of the ablation model trained without our proposed method.Conclusions: Our proposed method makes it feasible to construct a practical medical BERT model in both Japanese and English, and it has a potential to produce higher performing models for biomedical shared tasks.
Springer Science and Business Media LLC
Title: A Pre-Training Technique to Localize Medical BERT and to Enhance Biomedical BERT
Description:
Abstract
Background: Pre-training large-scale neural language models on raw texts has been shown to make a significant contribution to a strategy for transfer learning in natural language processing (NLP).
With the introduction of transformer-based language models, such as Bidirectional Encoder Representations from Transformers (BERT), the performance of information extraction from free text by NLP has significantly improved for both the general domain and the medical domain; however, it is difficult for languages in which there are few publicly available medical databases with a high quality and a large size to train medical BERT models that perform well.
Method: We introduce a method to train a BERT model using a small medical corpus both in English and in Japanese.
Our proposed method consists of two interventions: simultaneous pre-training, which is intended to encourage masked language modeling and next-sentence prediction on the small medical corpus, and amplified vocabulary, which helps with suiting the small corpus when building the customized corpus by byte-pair encoding.
Moreover, we used whole PubMed abstracts and developed a high-performance BERT model, Bidirectional Encoder Representations from Transformers for Biomedical Text Mining by Osaka University (ouBioBERT), in English via our method.
We then evaluated the performance of our BERT models and publicly available baselines and compared them.
Results: We confirmed that our Japanese medical BERT outperforms conventional baselines and the other BERT models in terms of the medical-document classification task and that our English BERT pre-trained using both the general and medical domain corpora performs sufficiently for practical use in terms of the biomedical language understanding evaluation (BLUE) benchmark.
Moreover, ouBioBERT shows that the total score of the BLUE benchmark is 1.
1 points above that of BioBERT and 0.
3 points above that of the ablation model trained without our proposed method.
Conclusions: Our proposed method makes it feasible to construct a practical medical BERT model in both Japanese and English, and it has a potential to produce higher performing models for biomedical shared tasks.
Related Results
Over-Sampling Effect in Pre-Training for Bidirectional Encoder Representations from Transformers (BERT) to Localize Medical BERT and Enhance Biomedical BERT (Preprint)
Over-Sampling Effect in Pre-Training for Bidirectional Encoder Representations from Transformers (BERT) to Localize Medical BERT and Enhance Biomedical BERT (Preprint)
BACKGROUND
Pre-training large-scale neural language models on raw texts has made a significant contribution to improving transfer learning in natural langua...
Investigation of Improving The Pre-Training And Fine-Tuning of BERT Model For Biomedical Relation Extraction
Investigation of Improving The Pre-Training And Fine-Tuning of BERT Model For Biomedical Relation Extraction
Abstract
Background: Recently, automatically extracting biomedical relations has been a significant subject in biomedical research due to the rapid growth of biomedical lit...
Bridging the Knowledge Gap: Improving BERT models for answering MCQs by using Ontology-generated synthetic MCQA Dataset
Bridging the Knowledge Gap: Improving BERT models for answering MCQs by using Ontology-generated synthetic MCQA Dataset
BERT-based models possess impressive language understanding capabilities but often lack domain-specific knowledge, limiting their performance on specialised tasks such as medical m...
On the effectiveness of small, discriminatively pre-trained language representation models for biomedical text mining
On the effectiveness of small, discriminatively pre-trained language representation models for biomedical text mining
Abstract
Neural language representation models such as BERT [1] have recently shown state of the art performance in downstream NLP tasks and bio-medical domain adap...
Big Data as the Foundation of a Novel Training Platform for Biomedical Researchers in Qatar
Big Data as the Foundation of a Novel Training Platform for Biomedical Researchers in Qatar
BackgroundTechnological breakthroughs witnessed over the past decade have led to an explosive increase in molecular profiling capabilities. This has ushered a new “data-rich era” f...
Improving chinese hate speech detection with bert-fasttext fusion and BERT-BiLSTM fusion
Improving chinese hate speech detection with bert-fasttext fusion and BERT-BiLSTM fusion
Hate speech detection is an essential technique in the online environment, especially on social media platforms. This technique helps to create a safer space and reduce the risk of...
Trooping the (School) Colour
Trooping the (School) Colour
Introduction
Throughout the early and mid-twentieth century, cadet training was a feature of many secondary schools and educational establishments across Australia, with countless ...
Hate Speech and Reality Check Analysis of Disaster Tweets using BERT Deep Learning Model
Hate Speech and Reality Check Analysis of Disaster Tweets using BERT Deep Learning Model
This study looks into the application of the BERT model for a reality check analysis of disaster-related tweets. The project intends to tackle the problem of authenticating informa...

