Javascript must be enabled to continue!
Improving Speaker Diarization for Overlapped Speech with Texture-Aware Feature Fusion
View through CrossRef
Speaker diarization (SD), which aims to address the “who spoke when” problem, is a key technology in speech processing. Although end-to-end neural speaker diarization methods have simplified the traditional multi-stage pipeline, their capability to extract discriminative speaker-specific features remains constrained, particularly in overlapping speech segments. To address this limitation, we propose EEND-ECB-CGA, an enhanced neural network built upon the EEND-VC framework. Our approach introduces a texture-aware fusion module that integrates an Edge-oriented Convolution Block (ECB) with Content-Guided Attention (CGA). The ECB extracts complementary texture and edge features from spectrograms, capturing speaker-specific structural patterns that are often overlooked by energy-based features, thereby improving the detection of speaker change points. The CGA module then dynamically weights the texture-enhanced features based on their importance, emphasizing speaker-dominant regions while suppressing noise and overlap interference. Evaluations on the LibriSpeech_mini and LibriSpeech datasets demonstrate that our EEND-ECB-CGA method significantly reduces the diarization error rate (DER) compared to the baseline. Furthermore, it outperforms several mainstream end-to-end clustering-based approaches. These results validate the robustness of our method in complex, multi-speaker environments, particularly in challenging scenarios with overlapping speech.
Title: Improving Speaker Diarization for Overlapped Speech with Texture-Aware Feature Fusion
Description:
Speaker diarization (SD), which aims to address the “who spoke when” problem, is a key technology in speech processing.
Although end-to-end neural speaker diarization methods have simplified the traditional multi-stage pipeline, their capability to extract discriminative speaker-specific features remains constrained, particularly in overlapping speech segments.
To address this limitation, we propose EEND-ECB-CGA, an enhanced neural network built upon the EEND-VC framework.
Our approach introduces a texture-aware fusion module that integrates an Edge-oriented Convolution Block (ECB) with Content-Guided Attention (CGA).
The ECB extracts complementary texture and edge features from spectrograms, capturing speaker-specific structural patterns that are often overlooked by energy-based features, thereby improving the detection of speaker change points.
The CGA module then dynamically weights the texture-enhanced features based on their importance, emphasizing speaker-dominant regions while suppressing noise and overlap interference.
Evaluations on the LibriSpeech_mini and LibriSpeech datasets demonstrate that our EEND-ECB-CGA method significantly reduces the diarization error rate (DER) compared to the baseline.
Furthermore, it outperforms several mainstream end-to-end clustering-based approaches.
These results validate the robustness of our method in complex, multi-speaker environments, particularly in challenging scenarios with overlapping speech.
Related Results
Investigating the Role of Speaker Counter in Handling Overlapping Speeches in Speaker Diarization Systems
Investigating the Role of Speaker Counter in Handling Overlapping Speeches in Speaker Diarization Systems
In real-life conversations, meetings, or debates, there are often
situations where many people speak at the same time, leading to
overlapping speech segments. Such overlapping spee...
Metaheuristic adapted convolutional neural network for Telugu speaker diarization
Metaheuristic adapted convolutional neural network for Telugu speaker diarization
In speech technology, a pivotal role is being played by the Speaker diarization mechanism. In general, speaker diarization is the mechanism of partitioning the input audio stream i...
Robust speaker diarization for meetings
Robust speaker diarization for meetings
Aquesta tesi doctoral mostra la recerca feta en l'àrea de la diarització de locutor per a sales de reunions. En la present s'estudien els algorismes i la implementació d'un sistema...
Bangla Diarizz: Domain-Adapted Speaker Diarization for Bengali Long-Form Audio
Bangla Diarizz: Domain-Adapted Speaker Diarization for Bengali Long-Form Audio
Speaker diarization-the task of automatically determining who spoke when in a multi-speaker recording-remains a notable open challenge for low-resource languages, and Bengali (Bang...
Multimodal Speaker Diarization Using a Pre-Trained Audio-Visual Synchronization Model
Multimodal Speaker Diarization Using a Pre-Trained Audio-Visual Synchronization Model
Speaker diarization systems aim to find ‘who spoke when?’ in multi-speaker recordings. The dataset usually consists of meetings, TV/talk shows, telephone and multi-party interactio...
The Nuclear Fusion Award
The Nuclear Fusion Award
The Nuclear Fusion Award ceremony for 2009 and 2010 award winners was held during the 23rd IAEA Fusion Energy Conference in Daejeon. This time, both 2009 and 2010 award winners w...
Speaker Verification and Identification
Speaker Verification and Identification
A speaker recognition system verifies or identifies a speaker’s identity based on his/her voice. It is considered as one of the most convenient biometric characteristic for human m...
Development of Supervised Speaker Diarization System Based on the PyAnnote Audio Processing Library
Development of Supervised Speaker Diarization System Based on the PyAnnote Audio Processing Library
Diarization is an important task when work with audiodata is executed, as it provides a solution to the problem related to the need of dividing one analyzed call recording into sev...

