Javascript must be enabled to continue!
Clinera ASR Benchmark: Evaluating Medical Code-Switching Automatic Speech Recognition for Arabic-English (ARZ-EN) (Preprint)
View through CrossRef
BACKGROUND
Code-switching between Egyptian Arabic (ARZ) and English is ubiquitous in clinical settings across the Arab world, yet no dedicated benchmark exists for evaluating Automatic Speech Recognition (ASR) systems under these conditions. Existing Arabic ASR benchmarks evaluate models on Modern Standard Arabic or single-dialect speech without medical vocabulary or code-switching.
OBJECTIVE
To introduce the Clinera ASR Benchmark, the first benchmark targeting medical ARZ-EN code-switched speech, and to evaluate commercial and open-source ASR systems on accurate recognition of medical terminology using both standard and novel medical-term-aware metrics
METHODS
We curated 683 utterances of Egyptian Arabic-English medical speech with dense medical-term annotations. We evaluated six ASR systems: the proprietary ElevenLabs Scribe v1, domain-adapted ArAZN-Whisper-S (244M parameters), Whisper-Large-V3, SeamlessM4T v2, Wav2Vec2 Large AR, and Whisper-Medium. We computed standard ASR metrics (WER, CER, MER) and proposed novel medical-term metrics (MT-Precision, MT-Recall, MT-F1, MT-wF1 weighted by inverse document frequency). Statistical comparisons used bootstrap confidence intervals and Wilcoxon signed-rank tests with Bonferroni correction.
RESULTS
ElevenLabs Scribe v1 achieved the lowest WER (27.8%) and highest MT-wF1 (60.5%), outperforming all open models by more than six times its size on medical-term recognition. General-purpose multilingual models achieved higher WER while almost entirely failing to recognize medical terms. Among open models, ArAZN-Whisper-S dominated on medical-term efficiency (26.0 MT-wF1 per 100M parameters). Even the best commercial system produced 11 fatal or high-risk medical errors affecting 1.6% of the corpus.
CONCLUSIONS
Current ASR systems exhibit a critical gap between general transcription accuracy and medical-term recognition in code-switched Arabic-English speech. Our benchmark and novel MT-metrics provide the first standardized evaluation framework for this clinically important setting, revealing that even top-performing systems pose patient safety risks through medical terminology errors.
Title: Clinera ASR Benchmark: Evaluating Medical Code-Switching Automatic Speech Recognition for Arabic-English (ARZ-EN) (Preprint)
Description:
BACKGROUND
Code-switching between Egyptian Arabic (ARZ) and English is ubiquitous in clinical settings across the Arab world, yet no dedicated benchmark exists for evaluating Automatic Speech Recognition (ASR) systems under these conditions.
Existing Arabic ASR benchmarks evaluate models on Modern Standard Arabic or single-dialect speech without medical vocabulary or code-switching.
OBJECTIVE
To introduce the Clinera ASR Benchmark, the first benchmark targeting medical ARZ-EN code-switched speech, and to evaluate commercial and open-source ASR systems on accurate recognition of medical terminology using both standard and novel medical-term-aware metrics
METHODS
We curated 683 utterances of Egyptian Arabic-English medical speech with dense medical-term annotations.
We evaluated six ASR systems: the proprietary ElevenLabs Scribe v1, domain-adapted ArAZN-Whisper-S (244M parameters), Whisper-Large-V3, SeamlessM4T v2, Wav2Vec2 Large AR, and Whisper-Medium.
We computed standard ASR metrics (WER, CER, MER) and proposed novel medical-term metrics (MT-Precision, MT-Recall, MT-F1, MT-wF1 weighted by inverse document frequency).
Statistical comparisons used bootstrap confidence intervals and Wilcoxon signed-rank tests with Bonferroni correction.
RESULTS
ElevenLabs Scribe v1 achieved the lowest WER (27.
8%) and highest MT-wF1 (60.
5%), outperforming all open models by more than six times its size on medical-term recognition.
General-purpose multilingual models achieved higher WER while almost entirely failing to recognize medical terms.
Among open models, ArAZN-Whisper-S dominated on medical-term efficiency (26.
0 MT-wF1 per 100M parameters).
Even the best commercial system produced 11 fatal or high-risk medical errors affecting 1.
6% of the corpus.
CONCLUSIONS
Current ASR systems exhibit a critical gap between general transcription accuracy and medical-term recognition in code-switched Arabic-English speech.
Our benchmark and novel MT-metrics provide the first standardized evaluation framework for this clinically important setting, revealing that even top-performing systems pose patient safety risks through medical terminology errors.
Related Results
ALIH KODE DALAM DIALOG NOVEL SURGA YANG TAK DIRINDUKAN KARYA ASMA NADIA
ALIH KODE DALAM DIALOG NOVEL SURGA YANG TAK DIRINDUKAN KARYA ASMA NADIA
<p><em>The objectives of this research are to explain: (1) the forms of code switching in a dialogue of novel Surga yang Tak Dirindukan, (2) the factors influencing of ...
Aviation English - A global perspective: analysis, teaching, assessment
Aviation English - A global perspective: analysis, teaching, assessment
This e-book brings together 13 chapters written by aviation English researchers and practitioners settled in six different countries, representing institutions and universities fro...
The Effectiveness of Using Code Switching in Teaching English on Higher Education
The Effectiveness of Using Code Switching in Teaching English on Higher Education
The occurence of using code switching in classroom appears for lecturer to explain topic, it also uses English-Indonesia, the students are difficult to understand when lecturer spe...
Pola Komunikasi Interpersonal Terapis Wicara pada Anak Telambat Bicara (Speech Delay)
Pola Komunikasi Interpersonal Terapis Wicara pada Anak Telambat Bicara (Speech Delay)
Abstract. The tittle of this study is Interpersonal Communication Patterns of Speech Therapists in Children with Speech Delay. Children with speech delay disorders are also social ...
Code Switching and Code Mixing in the Communication of Arabic Language Education Master Students UIN Malang: A Sociolinguistic Study
Code Switching and Code Mixing in the Communication of Arabic Language Education Master Students UIN Malang: A Sociolinguistic Study
Code switching is the phenomenon of changing the use of language or language variants in a conversation by a speaker, either between languages (e.g., from Indonesian to English) or...
Alih Kode Dan Campur Kode Dalam Interaksi Masyarakat Terminal Motabuik Kota Atambua
Alih Kode Dan Campur Kode Dalam Interaksi Masyarakat Terminal Motabuik Kota Atambua
This research aims to describe the use of language in community interactions at the Motabuik terminal, Atambua City. The use of language in question is the form and function of cod...
CODE-SWITCHING BETWEEN ARABIC AND ENGLISH IN SELECTED DISCOURSES
CODE-SWITCHING BETWEEN ARABIC AND ENGLISH IN SELECTED DISCOURSES
Code-switching is a linguistic phenomenon that has been studied on both written and spoken discourses in the recent
years. This study is an attempt to analyze code-switching in sel...
ANALISIS ALIH KODE DAN CAMPUR KODE PADA FILM “SANG PRAWIRA EPISODE I DAN EPISODE II” KARYA ONET ADITHIA RIZLAN
ANALISIS ALIH KODE DAN CAMPUR KODE PADA FILM “SANG PRAWIRA EPISODE I DAN EPISODE II” KARYA ONET ADITHIA RIZLAN
This study of code switching and code mixing analysis in the film "Sang Prawira Episode I and Episode II" by Onet Adithia Rizlan aims to determine code switching and code mixing se...

