Javascript must be enabled to continue!
Fine-tuning Whisper model for end-to-end Kannada to English speech transcription
View through CrossRef
Abstract
Building reliable automatic speech recognition (ASR) systems for low-resource languages such as Kannada is challenging due to the scarcity of training data and the cross-lingual translation requirements. This study investigates fine-tuning Whisper models – Tiny, Small, and Medium – to achieve end-to-end Kannada-to-English speech transcription. Using a dataset of Kannada speech, the multilingual Whisper models are fine-tuned to improve transcription and translation performance. The accuracy of the fine-tuned models was evaluated using word error rate (WER) and character error rate (CER), revealing clear improvements across the different model sizes. Based on the analysis of the experimental results, the Whisper-Medium model was found to perform the best, with significant reductions in WER and CER compared to the zero-shot baselines. The results highlight the feasibility of lightweight and midsize Whisper architectures for cross-lingual ASR in low-resource Indian languages and indicate that these models are more suitable for real-world applications in multilingual speech systems.
Title: Fine-tuning Whisper model for end-to-end Kannada to English speech transcription
Description:
Abstract
Building reliable automatic speech recognition (ASR) systems for low-resource languages such as Kannada is challenging due to the scarcity of training data and the cross-lingual translation requirements.
This study investigates fine-tuning Whisper models – Tiny, Small, and Medium – to achieve end-to-end Kannada-to-English speech transcription.
Using a dataset of Kannada speech, the multilingual Whisper models are fine-tuned to improve transcription and translation performance.
The accuracy of the fine-tuned models was evaluated using word error rate (WER) and character error rate (CER), revealing clear improvements across the different model sizes.
Based on the analysis of the experimental results, the Whisper-Medium model was found to perform the best, with significant reductions in WER and CER compared to the zero-shot baselines.
The results highlight the feasibility of lightweight and midsize Whisper architectures for cross-lingual ASR in low-resource Indian languages and indicate that these models are more suitable for real-world applications in multilingual speech systems.
Related Results
Aviation English - A global perspective: analysis, teaching, assessment
Aviation English - A global perspective: analysis, teaching, assessment
This e-book brings together 13 chapters written by aviation English researchers and practitioners settled in six different countries, representing institutions and universities fro...
Pola Komunikasi Interpersonal Terapis Wicara pada Anak Telambat Bicara (Speech Delay)
Pola Komunikasi Interpersonal Terapis Wicara pada Anak Telambat Bicara (Speech Delay)
Abstract. The tittle of this study is Interpersonal Communication Patterns of Speech Therapists in Children with Speech Delay. Children with speech delay disorders are also social ...
Electric field tuning characteristic of multiple optical parametric oscillator based on MgO:QPLN
Electric field tuning characteristic of multiple optical parametric oscillator based on MgO:QPLN
The quasi-phase matching optical parametric oscillator tuning methods, i.e. grating period tuning, temperature tuning, pumping wavelength tuning, and angle tuning are more simple a...
 Automatic detection of the electron density from de WHISPER instrument onboard CLUSTER II
 Automatic detection of the electron density from de WHISPER instrument onboard CLUSTER II
The Waves of HIgh frequency and Sounder for Probing Electron density by Relaxation(WHISPER) instrument, is part of the Wave Experiment Consortium (WEC) of the ESACLUSTER II mission...
Validation of 12-item general health questionnaire into Kannada language
Validation of 12-item general health questionnaire into Kannada language
17077 Background: The cancer load in India is enormous and majority of the cases present in an advanced stage. There is no valid translation of 12-item General Health Questionnair...
Tindak Tutur pada Tradisi Mappettuada Suku Bugis
Tindak Tutur pada Tradisi Mappettuada Suku Bugis
This study aims to The purpose of this study is to understand and describe locutionary, illocutionary, and perlocutionary speech acts in the Mappettuada tradition of the Bugis-Maka...
Indo-Anglian: Connotations and Denotations
Indo-Anglian: Connotations and Denotations
A different name than English literature, ‘Anglo-Indian Literature’, was given to the body of literature in English that emerged on account of the British interaction with India un...
Fine-Tuning Whisper for Low-Resource Automatic Speech Recognition in the Karakalpak Language
Fine-Tuning Whisper for Low-Resource Automatic Speech Recognition in the Karakalpak Language
Abstract
The Karakalpak language is still poorly represented in supervised speech resources, making reliable automatic speech recognition (ASR) difficult to build. ...

