Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Investigating the Role of Speaker Counter in Handling Overlapping Speeches in Speaker Diarization Systems

View through CrossRef
In real-life conversations, meetings, or debates, there are often situations where many people speak at the same time, leading to overlapping speech segments. Such overlapping speech is an extremely challenging problem for the speaker diarization task. The widely used clustering-based diarization approaches perform quite poorly under such situations due to their limited capabilities in handling overlapping speeches. This paper investigates a speaker diarization framework in which a new building block, called speaker count, is integrated. Such speaker counter predicts the number of active speakers in each analyzing audio window, then its output is used in the conventional re-segmentation step of the diarization pipelines in order to better label the active speakers in each considered segment. We also investigate the effect of the analyzing audio window size on diarization performance by theoretical analysis. We claim that the speaker count block ensures a lower diarization error rate when the analyzing window size is small enough. Experiment results obtained from two state-of-the-art diarization systems with different settings on two benchmark datasets, AMI Headset mix and DIHARD III, confirmed the effectiveness of the proposed approach.
Title: Investigating the Role of Speaker Counter in Handling Overlapping Speeches in Speaker Diarization Systems
Description:
In real-life conversations, meetings, or debates, there are often situations where many people speak at the same time, leading to overlapping speech segments.
Such overlapping speech is an extremely challenging problem for the speaker diarization task.
The widely used clustering-based diarization approaches perform quite poorly under such situations due to their limited capabilities in handling overlapping speeches.
This paper investigates a speaker diarization framework in which a new building block, called speaker count, is integrated.
Such speaker counter predicts the number of active speakers in each analyzing audio window, then its output is used in the conventional re-segmentation step of the diarization pipelines in order to better label the active speakers in each considered segment.
We also investigate the effect of the analyzing audio window size on diarization performance by theoretical analysis.
We claim that the speaker count block ensures a lower diarization error rate when the analyzing window size is small enough.
Experiment results obtained from two state-of-the-art diarization systems with different settings on two benchmark datasets, AMI Headset mix and DIHARD III, confirmed the effectiveness of the proposed approach.

Related Results

Metaheuristic adapted convolutional neural network for Telugu speaker diarization
Metaheuristic adapted convolutional neural network for Telugu speaker diarization
In speech technology, a pivotal role is being played by the Speaker diarization mechanism. In general, speaker diarization is the mechanism of partitioning the input audio stream i...
Robust speaker diarization for meetings
Robust speaker diarization for meetings
Aquesta tesi doctoral mostra la recerca feta en l'àrea de la diarització de locutor per a sales de reunions. En la present s'estudien els algorismes i la implementació d'un sistema...
Multimodal Speaker Diarization Using a Pre-Trained Audio-Visual Synchronization Model
Multimodal Speaker Diarization Using a Pre-Trained Audio-Visual Synchronization Model
Speaker diarization systems aim to find ‘who spoke when?’ in multi-speaker recordings. The dataset usually consists of meetings, TV/talk shows, telephone and multi-party interactio...
Bangla Diarizz: Domain-Adapted Speaker Diarization for Bengali Long-Form Audio
Bangla Diarizz: Domain-Adapted Speaker Diarization for Bengali Long-Form Audio
Speaker diarization-the task of automatically determining who spoke when in a multi-speaker recording-remains a notable open challenge for low-resource languages, and Bengali (Bang...
Development of Supervised Speaker Diarization System Based on the PyAnnote Audio Processing Library
Development of Supervised Speaker Diarization System Based on the PyAnnote Audio Processing Library
Diarization is an important task when work with audiodata is executed, as it provides a solution to the problem related to the need of dividing one analyzed call recording into sev...
Improving Speaker Diarization for Overlapped Speech with Texture-Aware Feature Fusion
Improving Speaker Diarization for Overlapped Speech with Texture-Aware Feature Fusion
Speaker diarization (SD), which aims to address the “who spoke when” problem, is a key technology in speech processing. Although end-to-end neural speaker diarization methods have ...
Cometary Physics Laboratory: spectrophotometric experiments
Cometary Physics Laboratory: spectrophotometric experiments
<p><strong><span dir="ltr" role="presentation">1. Introduction</span></strong&...
Speaker Role Identification in Clinical Conversations *
Speaker Role Identification in Clinical Conversations *
Patient-clinician communication research is crucial for understanding interaction dynamics and for predicting outcomes that are associated with clinical discourse. Traditionally, i...

Back to Top