Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Multimodal Speaker Diarization Using a Pre-Trained Audio-Visual Synchronization Model

View through CrossRef
Speaker diarization systems aim to find ‘who spoke when?’ in multi-speaker recordings. The dataset usually consists of meetings, TV/talk shows, telephone and multi-party interaction recordings. In this paper, we propose a novel multimodal speaker diarization technique, which finds the active speaker through audio-visual synchronization model for diarization. A pre-trained audio-visual synchronization model is used to find the synchronization between a visible person and the respective audio. For that purpose, short video segments comprised of face-only regions are acquired using a face detection technique and are then fed to the pre-trained model. This model is a two streamed network which matches audio frames with their respective visual input segments. On the basis of high confidence video segments inferred by the model, the respective audio frames are used to train Gaussian mixture model (GMM)-based clusters. This method helps in generating speaker specific clusters with high probability. We tested our approach on a popular subset of AMI meeting corpus consisting of 5.4 h of recordings for audio and 5.8 h of different set of multimodal recordings. A significant improvement is noticed with the proposed method in term of DER when compared to conventional and fully supervised audio based speaker diarization. The results of the proposed technique are very close to the complex state-of-the art multimodal diarization which shows significance of such simple yet effective technique.
Title: Multimodal Speaker Diarization Using a Pre-Trained Audio-Visual Synchronization Model
Description:
Speaker diarization systems aim to find ‘who spoke when?’ in multi-speaker recordings.
The dataset usually consists of meetings, TV/talk shows, telephone and multi-party interaction recordings.
In this paper, we propose a novel multimodal speaker diarization technique, which finds the active speaker through audio-visual synchronization model for diarization.
A pre-trained audio-visual synchronization model is used to find the synchronization between a visible person and the respective audio.
For that purpose, short video segments comprised of face-only regions are acquired using a face detection technique and are then fed to the pre-trained model.
This model is a two streamed network which matches audio frames with their respective visual input segments.
On the basis of high confidence video segments inferred by the model, the respective audio frames are used to train Gaussian mixture model (GMM)-based clusters.
This method helps in generating speaker specific clusters with high probability.
We tested our approach on a popular subset of AMI meeting corpus consisting of 5.
4 h of recordings for audio and 5.
8 h of different set of multimodal recordings.
A significant improvement is noticed with the proposed method in term of DER when compared to conventional and fully supervised audio based speaker diarization.
The results of the proposed technique are very close to the complex state-of-the art multimodal diarization which shows significance of such simple yet effective technique.

Related Results

Investigating the Role of Speaker Counter in Handling Overlapping Speeches in Speaker Diarization Systems
Investigating the Role of Speaker Counter in Handling Overlapping Speeches in Speaker Diarization Systems
In real-life conversations, meetings, or debates, there are often situations where many people speak at the same time, leading to overlapping speech segments. Such overlapping spee...
Metaheuristic adapted convolutional neural network for Telugu speaker diarization
Metaheuristic adapted convolutional neural network for Telugu speaker diarization
In speech technology, a pivotal role is being played by the Speaker diarization mechanism. In general, speaker diarization is the mechanism of partitioning the input audio stream i...
No Sudden Audio Switch – Preventing discontinuous POI audio playing in LBS
No Sudden Audio Switch – Preventing discontinuous POI audio playing in LBS
Abstract. Many LBS applications provide automatic audio playing functions for introducing POI’s. Appropriate automatic audio playing can improve users’ expressions during traveling...
Robust speaker diarization for meetings
Robust speaker diarization for meetings
Aquesta tesi doctoral mostra la recerca feta en l'àrea de la diarització de locutor per a sales de reunions. En la present s'estudien els algorismes i la implementació d'un sistema...
Bangla Diarizz: Domain-Adapted Speaker Diarization for Bengali Long-Form Audio
Bangla Diarizz: Domain-Adapted Speaker Diarization for Bengali Long-Form Audio
Speaker diarization-the task of automatically determining who spoke when in a multi-speaker recording-remains a notable open challenge for low-resource languages, and Bengali (Bang...
Development of Supervised Speaker Diarization System Based on the PyAnnote Audio Processing Library
Development of Supervised Speaker Diarization System Based on the PyAnnote Audio Processing Library
Diarization is an important task when work with audiodata is executed, as it provides a solution to the problem related to the need of dividing one analyzed call recording into sev...
Synchronization transition with coexistence of attractors in coupled discontinuous system
Synchronization transition with coexistence of attractors in coupled discontinuous system
The studies of extended dynamics systems are relevant to the understanding of spatiotemporal patterns observed in diverse fields. One of the well-established models for such comple...
Feature selection for multimodal: acoustic event detection
Feature selection for multimodal: acoustic event detection
The detection of the Acoustic Events (AEs) naturally produced in a meeting room may help to describe the human and social activity. The automatic description of interactions betwee...

Back to Top