Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

A Conditional Random Field Approach for Named Entity Recognition in Bengali and Hindi

View through CrossRef
This paper describes the development of Named Entity Recognition (NER) systems for two leading Indian languages, namely Bengali and Hindi, using the Conditional Random Field (CRF) framework. The system makes use of different types of contextual information along with a variety of features that are helpful in predicting the different named entity (NE) classes. This set of features includes language independent as well as language dependent components. We have used the annotated corpora of 122,467 tokens for Bengali and 502,974 tokens for Hindi tagged with a tag set of twelve different NE classes, defined as part of the IJCNLP-08 NER Shared Task for South and South East Asian Languages (SSEAL). We have considered only the tags that denote person names, location names, organization names, number expressions, time expressions and measurement expressions. A number of experiments have been carried out in order to find out the most suitable features for NER in Bengali and Hindi. The system has been tested with the gold standard test sets of 35K for Bengali and 50K tokens for Hindi. Evaluation results in overall f-score values of 81.15% for Bengali and 78.29% for Hindi for the test sets. 10-fold cross validation tests yield f-score values of 83.89% for Bengali and 80.93% for Hindi. ANOVA analysis is performed to show that the performance improvement due to the use of language dependent features is statistically significant.
Title: A Conditional Random Field Approach for Named Entity Recognition in Bengali and Hindi
Description:
This paper describes the development of Named Entity Recognition (NER) systems for two leading Indian languages, namely Bengali and Hindi, using the Conditional Random Field (CRF) framework.
The system makes use of different types of contextual information along with a variety of features that are helpful in predicting the different named entity (NE) classes.
This set of features includes language independent as well as language dependent components.
We have used the annotated corpora of 122,467 tokens for Bengali and 502,974 tokens for Hindi tagged with a tag set of twelve different NE classes, defined as part of the IJCNLP-08 NER Shared Task for South and South East Asian Languages (SSEAL).
We have considered only the tags that denote person names, location names, organization names, number expressions, time expressions and measurement expressions.
A number of experiments have been carried out in order to find out the most suitable features for NER in Bengali and Hindi.
The system has been tested with the gold standard test sets of 35K for Bengali and 50K tokens for Hindi.
Evaluation results in overall f-score values of 81.
15% for Bengali and 78.
29% for Hindi for the test sets.
10-fold cross validation tests yield f-score values of 83.
89% for Bengali and 80.
93% for Hindi.
ANOVA analysis is performed to show that the performance improvement due to the use of language dependent features is statistically significant.

Related Results

Neural Machine Translation from Bengali Language to English language and vice-versa
Neural Machine Translation from Bengali Language to English language and vice-versa
Bengali ranks among the first ten spoken languages in the world with a native speaker numbering about 230 million people.  With UNESCO declaring 21st February as International Moth...
A Web-based Intelligent Handwriting Education System for Autonomous Learning of Bengali Characters
A Web-based Intelligent Handwriting Education System for Autonomous Learning of Bengali Characters
In this paper, we describe a prototype of web-based intelligent handwriting education system for autonomous learning of Bengali characters. Bengali language is used by more than 21...
Hindi/Bengali Sentiment Analysis using Transfer Learning and Joint Dual Input Learning with Self Attention
Hindi/Bengali Sentiment Analysis using Transfer Learning and Joint Dual Input Learning with Self Attention
Sentiment Analysis typically refers to using natural language processing, text analysis, and computational linguistics to extract effect and emotion-based information from text dat...
Graph-based entity-oriented search
Graph-based entity-oriented search
Entity-oriented search has revolutionized search engines. In the era of Google Knowledge Graph and Microsoft Satori, users demand an effortless process of search. Whether they expr...
Reform and Change in Early 20th Century Bengali Society: A Study of Chattopadhyay's Novel Nishkriti
Reform and Change in Early 20th Century Bengali Society: A Study of Chattopadhyay's Novel Nishkriti
The goal of this research is to examine the societal reforms and modifications that took place in early 20th-century Bengal as a result of the flourishing Bengali Renaissance, as p...
Hindi/Bengali sentiment analysis using transfer learningand joint dual input learning with self-attention
Hindi/Bengali sentiment analysis using transfer learningand joint dual input learning with self-attention
Sentiment analysis typically refers to using natural language processing, text analysis, and computationallinguistics to extract effect and emotion-based information from text data...

Back to Top