Javascript must be enabled to continue!
Improving Brill's tagger lexical and transformation rule for Afaan Oromo language
View through CrossRef
Natural Language Processing (NLP) refers to Human-like language processing which reveals that it is a discipline within the field of Artificial Intelligence (AI). However, the ultimate goal of research on Natural Language Processing is to parse and understand language, which is not fully achieved yet. For this reason, much research in NLP has focused on intermediate tasks that make sense of some of the structure inherent in language without requiring complete understanding. One such task is part-of-speech tagging, or simply tagging. Lack of standard part of speech tagger for Afaan Oromo will be the main obstacle for researchers in the area of machine translation, spell checkers, dictionary compilation and automatic sentence parsing and constructions. Even though several works have been done in POS tagging for Afaan Oromo, the performance of the tagger is not sufficiently improved yet. Hence,the aim of this thesis is to improve Brill’s tagger lexical and transformation rule for Afaan Oromo POS tagging with sufficiently large training corpus. Accordingly, Afaan Oromo literatures on grammar and morphology are reviewed to understand nature of the language and also to identify possible tagsets. As a result, 26 broad tagsets were identified and 17,473 words from around 1100 sentences containing 6750 distinct words were tagged for training and testing purpose. From which 258 sentences are taken from the previous work. Since there is only a few ready made standard corpuses, the manual tagging process to prepare corpus for this work was challenging and hence, it is recommended that a standard corpus is prepared. Transformation-based Error driven learning are adapted for Afaan Oromo part of speech tagging. Different experiments are conducted for the rule based approach taking 20% of the whole data for testing. A comparison with the previously adapted Brill’s Tagger made. The previously adapted Brill’s Tagger shows an accuracy of 80.08% whereas the improved Brill’s Tagger result shows an accuracy of 95.6% which has an improvement of 15.52%. Hence, it is found that the size of the training corpus, the rule generating system in the lexical rule learner, and moreover, using Afaan Oromo HMM tagger as initial state tagger have a significant effect on the improvement of the tagger.
Title: Improving Brill's tagger lexical and transformation rule for Afaan Oromo language
Description:
Natural Language Processing (NLP) refers to Human-like language processing which reveals that it is a discipline within the field of Artificial Intelligence (AI).
However, the ultimate goal of research on Natural Language Processing is to parse and understand language, which is not fully achieved yet.
For this reason, much research in NLP has focused on intermediate tasks that make sense of some of the structure inherent in language without requiring complete understanding.
One such task is part-of-speech tagging, or simply tagging.
Lack of standard part of speech tagger for Afaan Oromo will be the main obstacle for researchers in the area of machine translation, spell checkers, dictionary compilation and automatic sentence parsing and constructions.
Even though several works have been done in POS tagging for Afaan Oromo, the performance of the tagger is not sufficiently improved yet.
Hence,the aim of this thesis is to improve Brill’s tagger lexical and transformation rule for Afaan Oromo POS tagging with sufficiently large training corpus.
Accordingly, Afaan Oromo literatures on grammar and morphology are reviewed to understand nature of the language and also to identify possible tagsets.
As a result, 26 broad tagsets were identified and 17,473 words from around 1100 sentences containing 6750 distinct words were tagged for training and testing purpose.
From which 258 sentences are taken from the previous work.
Since there is only a few ready made standard corpuses, the manual tagging process to prepare corpus for this work was challenging and hence, it is recommended that a standard corpus is prepared.
Transformation-based Error driven learning are adapted for Afaan Oromo part of speech tagging.
Different experiments are conducted for the rule based approach taking 20% of the whole data for testing.
A comparison with the previously adapted Brill’s Tagger made.
The previously adapted Brill’s Tagger shows an accuracy of 80.
08% whereas the improved Brill’s Tagger result shows an accuracy of 95.
6% which has an improvement of 15.
52%.
Hence, it is found that the size of the training corpus, the rule generating system in the lexical rule learner, and moreover, using Afaan Oromo HMM tagger as initial state tagger have a significant effect on the improvement of the tagger.
Related Results
Hubungan Perilaku Pola Makan dengan Kejadian Anak Obesitas
Hubungan Perilaku Pola Makan dengan Kejadian Anak Obesitas
<p><em><span style="font-size: 11.0pt; font-family: 'Times New Roman',serif; mso-fareast-font-family: 'Times New Roman'; mso-ansi-language: EN-US; mso-fareast-langua...
Tartiiba fufiilee yaasaafi hormaata gochimoota Afaan Oromoo
Tartiiba fufiilee yaasaafi hormaata gochimoota Afaan Oromoo
Qorannoon kun tartiiba fufiilee yaasaafi hormaata gochimoota Afaan Oromoo addeessuufi xiinxaluu irratti xiyyeeffate. Afaan Oromoo afaan xiinjecha gochimaa gabbataafi walxaxaa qabu ...
Afaan Oromo Multi-Label News Text Classification Using Deep Learning Approach
Afaan Oromo Multi-Label News Text Classification Using Deep Learning Approach
Abstract
Classification is a technique for categorizing textual data into a form of predefined categories. Due to its major consequences in regard to critical tasks such as...
Comparative Analysis of Deep Learning Based Afaan Oromo Hate Speech Detection
Comparative Analysis of Deep Learning Based Afaan Oromo Hate Speech Detection
Abstract
The network and openness of online media stages permit individuals to communicate their thoughts and offer encounters without any problem. Nonetheless, the Interne...
The Oromo national memories
The Oromo national memories
The author defines nation as a territorial community of nativity and attributes significance to the biological fact of birth into the historically evolving territorial structure of...
The morphosyntactic integration of English words into Afaan Oromoo
The morphosyntactic integration of English words into Afaan Oromoo
The present study investigates the morphosyntactic integration of English lexical items into Afaan Oromoo within multilingual conversations recorded in Dambi Dollo, Oromia regional...
Adeemsa dhamsagaa daangaa dhamjecha Afaan Oromoo irratti
Adeemsa dhamsagaa daangaa dhamjecha Afaan Oromoo irratti
Kaayyoon qorannoo kanaa adeemsa jijjiirama amala dhamsagaa daangaa dhamjechootaa irratti mul’atu xiinxaluudha. Adeemsi dhamsagaa jechoota Afaan Oromoo keessatti mul’atan kanaan...
Generational Wisdom: Lesson from the Oromo People
Generational Wisdom: Lesson from the Oromo People
This review explores the foundational elements of Oromo generational wisdom, focusing on how their rich cultural heritage, particularly the Gadaa system, is passed down through gen...

