Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Bangla Diarizz: Domain-Adapted Speaker Diarization for Bengali Long-Form Audio

View through CrossRef
Speaker diarization-the task of automatically determining who spoke when in a multi-speaker recording-remains a notable open challenge for low-resource languages, and Bengali (Bangla) is no exception. Despite being among the most widely spoken languages globally, Bengali has historically lacked largescale, publicly available diarization corpora and purpose-built models. This paper presents Bangla Diarizz, a domain-adapted speaker diarization system developed for the DL Sprint 4.0 Bengali Speaker Diarization competition (BUET CSE Fest 2026). Our system surgically replaces both the segmentation and embedding modules of the pretrained Pyannote-Community-1 pipeline: the segmentation model is fine-tuned on the competition dataset using the diarizers library, and the speaker embedding model is replaced with WeSpeaker ResNet34-LM. We furthermore introduce a post-processing stitching mechanism that merges short intra-speaker gaps and reduces fragmentation. Our approach achieves a competitive Diarization Error Rate (DER) on the Bengali-Loop benchmark [1] while reducing inference time from 1 hour 22 minutes to approximately 36 minutes on 14 test audio files, representing a 56% wall-clock speed improvement. This work demonstrates that targeted domain adaptation through fine-tuning on limited in-domain data can yield substantial improvements over language-agnostic baselines for under-resourced speech.
Title: Bangla Diarizz: Domain-Adapted Speaker Diarization for Bengali Long-Form Audio
Description:
Speaker diarization-the task of automatically determining who spoke when in a multi-speaker recording-remains a notable open challenge for low-resource languages, and Bengali (Bangla) is no exception.
Despite being among the most widely spoken languages globally, Bengali has historically lacked largescale, publicly available diarization corpora and purpose-built models.
This paper presents Bangla Diarizz, a domain-adapted speaker diarization system developed for the DL Sprint 4.
0 Bengali Speaker Diarization competition (BUET CSE Fest 2026).
Our system surgically replaces both the segmentation and embedding modules of the pretrained Pyannote-Community-1 pipeline: the segmentation model is fine-tuned on the competition dataset using the diarizers library, and the speaker embedding model is replaced with WeSpeaker ResNet34-LM.
We furthermore introduce a post-processing stitching mechanism that merges short intra-speaker gaps and reduces fragmentation.
Our approach achieves a competitive Diarization Error Rate (DER) on the Bengali-Loop benchmark [1] while reducing inference time from 1 hour 22 minutes to approximately 36 minutes on 14 test audio files, representing a 56% wall-clock speed improvement.
This work demonstrates that targeted domain adaptation through fine-tuning on limited in-domain data can yield substantial improvements over language-agnostic baselines for under-resourced speech.

Related Results

Investigating the Role of Speaker Counter in Handling Overlapping Speeches in Speaker Diarization Systems
Investigating the Role of Speaker Counter in Handling Overlapping Speeches in Speaker Diarization Systems
In real-life conversations, meetings, or debates, there are often situations where many people speak at the same time, leading to overlapping speech segments. Such overlapping spee...
Metaheuristic adapted convolutional neural network for Telugu speaker diarization
Metaheuristic adapted convolutional neural network for Telugu speaker diarization
In speech technology, a pivotal role is being played by the Speaker diarization mechanism. In general, speaker diarization is the mechanism of partitioning the input audio stream i...
Multimodal Speaker Diarization Using a Pre-Trained Audio-Visual Synchronization Model
Multimodal Speaker Diarization Using a Pre-Trained Audio-Visual Synchronization Model
Speaker diarization systems aim to find ‘who spoke when?’ in multi-speaker recordings. The dataset usually consists of meetings, TV/talk shows, telephone and multi-party interactio...
No Sudden Audio Switch – Preventing discontinuous POI audio playing in LBS
No Sudden Audio Switch – Preventing discontinuous POI audio playing in LBS
Abstract. Many LBS applications provide automatic audio playing functions for introducing POI’s. Appropriate automatic audio playing can improve users’ expressions during traveling...
Robust speaker diarization for meetings
Robust speaker diarization for meetings
Aquesta tesi doctoral mostra la recerca feta en l'àrea de la diarització de locutor per a sales de reunions. En la present s'estudien els algorismes i la implementació d'un sistema...
Development of Supervised Speaker Diarization System Based on the PyAnnote Audio Processing Library
Development of Supervised Speaker Diarization System Based on the PyAnnote Audio Processing Library
Diarization is an important task when work with audiodata is executed, as it provides a solution to the problem related to the need of dividing one analyzed call recording into sev...
Neural Machine Translation from Bengali Language to English language and vice-versa
Neural Machine Translation from Bengali Language to English language and vice-versa
Bengali ranks among the first ten spoken languages in the world with a native speaker numbering about 230 million people.  With UNESCO declaring 21st February as International Moth...
Bangla Linguistics and Abul Mansur Ahmad
Bangla Linguistics and Abul Mansur Ahmad
There was a time when it was a matter of great debate whether Urdu or Bangla was the mother tongue of the Bengali Muslim. In the field of Bengali Language and Literature another de...

Back to Top