Javascript must be enabled to continue!
COMPARISON OF SPELL CORRECTION IN BAHASA INDONESIA: PETER NORVIG, LSTM, AND N-GRAM
View through CrossRef
This study conducts a comprehensive comparison of spell-checking methods in Bahasa Indonesia, specifically focusing on three approaches: Peter Norvig's method, Long Short-Term Memory (LSTM), and N-gram. The primary metric for evaluation is the accuracy in correcting spelling errors. Notably, Peter Norvig's method outperforms the others, with N-gram following closely, and LSTM trailing behind. The study draws valuable insights that contribute to the enhancement of spelling correction accuracy in the Bahasa Indonesia language. To carry out the evaluation, the research employs SPECIL data (Spell Error Corpus for Indonesian Language), which includes documents with various error types such as insertion, deletion, transposition, and substitution. The testing dataset consists of 150 words, aligning with the 150-word corpus references from the 'Leipzig Corpora Collection' used for Peter Norvig's and N-gram methods. It's noteworthy that the LSTM method utilizes a reference dataset from SPECIL, comprising 150 data points and specifically focusing on insertion errors for the test data. This research provides valuable insights for researchers, developers, and language technology enthusiasts seeking to refine spell-checking techniques for the Bahasa Indonesia language. By leveraging diverse error types and a standardized testing dataset, the study aims to contribute to the continual improvement of spell-checking tools
LPPM Universitas Khairun
Title: COMPARISON OF SPELL CORRECTION IN BAHASA INDONESIA: PETER NORVIG, LSTM, AND N-GRAM
Description:
This study conducts a comprehensive comparison of spell-checking methods in Bahasa Indonesia, specifically focusing on three approaches: Peter Norvig's method, Long Short-Term Memory (LSTM), and N-gram.
The primary metric for evaluation is the accuracy in correcting spelling errors.
Notably, Peter Norvig's method outperforms the others, with N-gram following closely, and LSTM trailing behind.
The study draws valuable insights that contribute to the enhancement of spelling correction accuracy in the Bahasa Indonesia language.
To carry out the evaluation, the research employs SPECIL data (Spell Error Corpus for Indonesian Language), which includes documents with various error types such as insertion, deletion, transposition, and substitution.
The testing dataset consists of 150 words, aligning with the 150-word corpus references from the 'Leipzig Corpora Collection' used for Peter Norvig's and N-gram methods.
It's noteworthy that the LSTM method utilizes a reference dataset from SPECIL, comprising 150 data points and specifically focusing on insertion errors for the test data.
This research provides valuable insights for researchers, developers, and language technology enthusiasts seeking to refine spell-checking techniques for the Bahasa Indonesia language.
By leveraging diverse error types and a standardized testing dataset, the study aims to contribute to the continual improvement of spell-checking tools.
Related Results
Evolution of Antimicrobial Resistance in Community vs. Hospital-Acquired Infections
Evolution of Antimicrobial Resistance in Community vs. Hospital-Acquired Infections
Abstract
Introduction
Hospitals are high-risk environments for infections. Despite the global recognition of these pathogens, few studies compare microorganisms from community-acqu...
Spelling Corrector Bahasa Indonesia dengan Kombinasi Metode Peter Norvig dan N-Gram
Spelling Corrector Bahasa Indonesia dengan Kombinasi Metode Peter Norvig dan N-Gram
Abstrak - Kesalahan pengetikan dalam suatu dokumen merupakan human error yang sulit dihindari, akibatnya pesan yang ingin disampaikan tidak maksimal. Menggunakan fitur Spelling Cor...
KEUNIKAN BAHASA MANTRA BANJAR: PANAH ARJUNA (The Uniqueness of the Expression of Banjarese Spell:Panah Arjuna)
KEUNIKAN BAHASA MANTRA BANJAR: PANAH ARJUNA (The Uniqueness of the Expression of Banjarese Spell:Panah Arjuna)
Mantra Panah Arjuna adalah salah satu mantra Banjar berupa mantra cinta untuk menundukkan hati seseorang yang dicintai. Lazimnya mantra ini dipergunakan oleh laki-laki untuk menak...
PERSPEKTIF AKULTURASI NILAI BILINGUALISME BAHASA DI SITUBONDO
PERSPEKTIF AKULTURASI NILAI BILINGUALISME BAHASA DI SITUBONDO
Abstrak, Indonesia sebagai sebuah bangsa memiliki keragaman budaya dan bahasa yang sangat tinggi. Tingkat kemajemukan yang sangat tinggi ini tercermin dalamjumlahbahasa daerah yang...
Development of a Custom Spell-Checker for Emergency Department Data
Development of a Custom Spell-Checker for Emergency Department Data
ObjectiveTo share progress on a custom spell-checker for emergency department chief complaint free-text data and demonstrate a spell-checker validation Shiny application.Introducti...
PEMETAAN LANSKAP LINGUISTIK DI UNIVERSITAS AIRLANGGA SURABAYA
PEMETAAN LANSKAP LINGUISTIK DI UNIVERSITAS AIRLANGGA SURABAYA
Lanskap Linguistik ( LL) merujuk pada objek penggunaan bahasa di ruang publik. Menurut Landry and Bourhis (1997) yang termasuk dalam LL adalah bahasa di ruang-ruang publik seperti ...
Mass Conserving LSTM with Dual States for Improved Streamflow Prediction through Quickflow and Slow Storage Separation
Mass Conserving LSTM with Dual States for Improved Streamflow Prediction through Quickflow and Slow Storage Separation
Long-Short Term Memory (LSTM) shows exceptional performance for rainfall-runoff modelling, but lacks physical realism. Efforts to integrate mass conserving into the model architect...
PENGAJARAN BAHASA DAN SASTRA SASAK DI SEKOLAH (Hambatan dan Alternatif Pemecahannya)
PENGAJARAN BAHASA DAN SASTRA SASAK DI SEKOLAH (Hambatan dan Alternatif Pemecahannya)
Di dunia saat ini, terdapat tidak kurang dari 6000 bahasa. Separuh dari bahasa-bahasa tersebut terancam punah. Dari jumlah 6000 bahasa tersebut, sebanyak 746 bahasa berada di Indon...

