Javascript must be enabled to continue!
Investigating the Role of Speaker Counter in Handling Overlapping Speeches in Speaker Diarization Systems
View through CrossRef
In real-life conversations, meetings, or debates, there are often
situations where many people speak at the same time, leading to
overlapping speech segments. Such overlapping speech is an extremely
challenging problem for the speaker diarization task. The widely used
clustering-based diarization approaches perform quite poorly under such
situations due to their limited capabilities in handling overlapping
speeches. This paper investigates a speaker diarization framework in
which a new building block, called speaker count, is integrated. Such
speaker counter predicts the number of active speakers in each analyzing
audio window, then its output is used in the conventional
re-segmentation step of the diarization pipelines in order to better
label the active speakers in each considered segment. We also
investigate the effect of the analyzing audio window size on diarization
performance by theoretical analysis. We claim that the speaker count
block ensures a lower diarization error rate when the analyzing window
size is small enough. Experiment results obtained from two
state-of-the-art diarization systems with different settings on two
benchmark datasets, AMI Headset mix and DIHARD III, confirmed the
effectiveness of the proposed approach.
Title: Investigating the Role of Speaker Counter in Handling Overlapping Speeches in Speaker Diarization Systems
Description:
In real-life conversations, meetings, or debates, there are often
situations where many people speak at the same time, leading to
overlapping speech segments.
Such overlapping speech is an extremely
challenging problem for the speaker diarization task.
The widely used
clustering-based diarization approaches perform quite poorly under such
situations due to their limited capabilities in handling overlapping
speeches.
This paper investigates a speaker diarization framework in
which a new building block, called speaker count, is integrated.
Such
speaker counter predicts the number of active speakers in each analyzing
audio window, then its output is used in the conventional
re-segmentation step of the diarization pipelines in order to better
label the active speakers in each considered segment.
We also
investigate the effect of the analyzing audio window size on diarization
performance by theoretical analysis.
We claim that the speaker count
block ensures a lower diarization error rate when the analyzing window
size is small enough.
Experiment results obtained from two
state-of-the-art diarization systems with different settings on two
benchmark datasets, AMI Headset mix and DIHARD III, confirmed the
effectiveness of the proposed approach.
Related Results
Metaheuristic adapted convolutional neural network for Telugu speaker diarization
Metaheuristic adapted convolutional neural network for Telugu speaker diarization
In speech technology, a pivotal role is being played by the Speaker diarization mechanism. In general, speaker diarization is the mechanism of partitioning the input audio stream i...
Robust speaker diarization for meetings
Robust speaker diarization for meetings
Aquesta tesi doctoral mostra la recerca feta en l'àrea de la diarització de locutor per a sales de reunions. En la present s'estudien els algorismes i la implementació d'un sistema...
Multimodal Speaker Diarization Using a Pre-Trained Audio-Visual Synchronization Model
Multimodal Speaker Diarization Using a Pre-Trained Audio-Visual Synchronization Model
Speaker diarization systems aim to find ‘who spoke when?’ in multi-speaker recordings. The dataset usually consists of meetings, TV/talk shows, telephone and multi-party interactio...
Bangla Diarizz: Domain-Adapted Speaker Diarization for Bengali Long-Form Audio
Bangla Diarizz: Domain-Adapted Speaker Diarization for Bengali Long-Form Audio
Speaker diarization-the task of automatically determining who spoke when in a multi-speaker recording-remains a notable open challenge for low-resource languages, and Bengali (Bang...
Development of Supervised Speaker Diarization System Based on the PyAnnote Audio Processing Library
Development of Supervised Speaker Diarization System Based on the PyAnnote Audio Processing Library
Diarization is an important task when work with audiodata is executed, as it provides a solution to the problem related to the need of dividing one analyzed call recording into sev...
Improving Speaker Diarization for Overlapped Speech with Texture-Aware Feature Fusion
Improving Speaker Diarization for Overlapped Speech with Texture-Aware Feature Fusion
Speaker diarization (SD), which aims to address the “who spoke when” problem, is a key technology in speech processing. Although end-to-end neural speaker diarization methods have ...
Cometary Physics Laboratory: spectrophotometric experiments
Cometary Physics Laboratory: spectrophotometric experiments
<p><strong><span dir="ltr" role="presentation">1. Introduction</span></strong&...
Speaker Role Identification in Clinical Conversations
*
Speaker Role Identification in Clinical Conversations
*
Patient-clinician communication research is crucial for understanding interaction dynamics and for predicting outcomes that are associated with clinical discourse. Traditionally, i...

