Javascript must be enabled to continue!
Calculating Semantic Frequency of GSL Words Using a BERT Model in Large Corpora
View through CrossRef
There has always been a pressing need to provide semantic information for words in high-frequency word lists, but technical limitations have hindered this goal. This study addresses this challenge by leveraging a large language model, such as BERT, to semantically annotate large corpora and identify the high-frequency senses of headwords from the General Service List (GSL). We aim to explore three key questions: (1) Can BERT automatically annotate large corpora and accurately calculate sense frequencies? (2) What are the high-frequency senses of GSL words? (3) Can this approach be verified? Using a BERT-based framework, we annotated 1,891 GSL headwords (10,925 senses) in the 100-million-word British National Corpus (BNC), representing each sense with a 1,024-dimensional vector. From this, we identified 3,695 high-frequency senses for the GSL words. Three main conclusions are drawn from this study. First, BERT demonstrates high accuracy in sense annotation, achieving 92% precision when disambiguating the senses of GSL words. Second, a relatively small number of high-frequency senses account for a significant portion of corpus coverage. Specifically, these high-frequency senses (33.8% of the total) cover approximately 60% of all GSL word occurrences in the BNC. Third, the high-frequency senses selected via this method can be verified by their consistent coverage across different corpora. This study illustrates a pioneering method for semantic annotation in large corpora, which can be easily applied to calculate semantic frequencies for other word lists.
Title: Calculating Semantic Frequency of GSL Words Using a BERT Model in Large Corpora
Description:
There has always been a pressing need to provide semantic information for words in high-frequency word lists, but technical limitations have hindered this goal.
This study addresses this challenge by leveraging a large language model, such as BERT, to semantically annotate large corpora and identify the high-frequency senses of headwords from the General Service List (GSL).
We aim to explore three key questions: (1) Can BERT automatically annotate large corpora and accurately calculate sense frequencies? (2) What are the high-frequency senses of GSL words? (3) Can this approach be verified? Using a BERT-based framework, we annotated 1,891 GSL headwords (10,925 senses) in the 100-million-word British National Corpus (BNC), representing each sense with a 1,024-dimensional vector.
From this, we identified 3,695 high-frequency senses for the GSL words.
Three main conclusions are drawn from this study.
First, BERT demonstrates high accuracy in sense annotation, achieving 92% precision when disambiguating the senses of GSL words.
Second, a relatively small number of high-frequency senses account for a significant portion of corpus coverage.
Specifically, these high-frequency senses (33.
8% of the total) cover approximately 60% of all GSL word occurrences in the BNC.
Third, the high-frequency senses selected via this method can be verified by their consistent coverage across different corpora.
This study illustrates a pioneering method for semantic annotation in large corpora, which can be easily applied to calculate semantic frequencies for other word lists.
Related Results
Računalno potpomognuto usmjeravanje kod dvojezičnih govornika
Računalno potpomognuto usmjeravanje kod dvojezičnih govornika
This thesis investigates whether modern computer models can confirm how people encounter words and then use these findings in didactics. In recent years, computers have been used i...
Interaktion von Shiga Toxin mit primären humanen intestinalen und renalen Epithelzellen
Interaktion von Shiga Toxin mit primären humanen intestinalen und renalen Epithelzellen
ZusammenfassungInfektionen durch enterohämorrhagische Escherichia coli (EHEC)‐Bakterien können beim Menschen wässrige und blutige Durchfälle verursachen und im schlimmsten Fall das...
Frequency of Common Chromosomal Abnormalities in Patients with Idiopathic Acquired Aplastic Anemia
Frequency of Common Chromosomal Abnormalities in Patients with Idiopathic Acquired Aplastic Anemia
Objective: To determine the frequency of common chromosomal aberrations in local population idiopathic determine the frequency of common chromosomal aberrations in local population...
Goniosynechialysis under a microscope alone and under direct gonioscopy for chronic angle-closure glaucoma patients coexisted with cataract
Goniosynechialysis under a microscope alone and under direct gonioscopy for chronic angle-closure glaucoma patients coexisted with cataract
AIM: To compare the efficacy of goniosynechialysis (GSL) under a microscope alone (GM) and under direct gonioscopy (GG) for chronic angle-closure glaucoma (CACG) coexisted with cat...
Nouvelles approches pour la valorisation des graines de moutarde riches en glucosinolates dans un concept de bioraffinerie
Nouvelles approches pour la valorisation des graines de moutarde riches en glucosinolates dans un concept de bioraffinerie
Ce travail de thèse est dédié à la mise en place d’une stratégie de valorisation optimale et raisonnée des cultures intermédiaires piège à nitrates de B. juncea (moutarde brune) da...
Over-Sampling Effect in Pre-Training for Bidirectional Encoder Representations from Transformers (BERT) to Localize Medical BERT and Enhance Biomedical BERT (Preprint)
Over-Sampling Effect in Pre-Training for Bidirectional Encoder Representations from Transformers (BERT) to Localize Medical BERT and Enhance Biomedical BERT (Preprint)
BACKGROUND
Pre-training large-scale neural language models on raw texts has made a significant contribution to improving transfer learning in natural langua...
A Semantic Orthogonal Mapping Method Through Deep-Learning for Semantic Computing
A Semantic Orthogonal Mapping Method Through Deep-Learning for Semantic Computing
In order to realize an artificial intelligent system, a basic mechanism should be provided for expressing and processing the semantic. We have presented semantic computing models i...
Semantic Similarity Caculating based on BERT
Semantic Similarity Caculating based on BERT
The exploration of semantic similarity is a fundamental aspect of natural language processing, as it aids in comprehending the significance and usage of vocabulary present in a lan...

