Javascript must be enabled to continue!
Development of Supervised Speaker Diarization System Based on the PyAnnote Audio Processing Library
View through CrossRef
Diarization is an important task when work with audiodata is executed, as it provides a solution to the problem related to the need of dividing one analyzed call recording into several speech recordings, each of which belongs to one speaker. Diarization systems segment audio recordings by defining the time boundaries of utterances, and typically use unsupervised methods to group utterances belonging to individual speakers, but do not answer the question “who is speaking?” On the other hand, there are biometric systems that identify individuals on the basis of their voices, but such systems are designed with the prerequisite that only one speaker is present in the analyzed audio recording. However, some applications involve the need to identify multiple speakers that interact freely in an audio recording. This paper proposes two architectures of speaker identification systems based on a combination of diarization and identification methods, which operate on the basis of segment-level or group-level classification. The open-source PyAnnote framework was used to develop the system. The performance of the speaker identification system was verified through the application of the AMI Corpus open-source audio database, which contains 100 h of annotated and transcribed audio and video data. The research method consisted of four experiments to select the best-performing supervised diarization algorithms on the basis of PyAnnote. The first experiment was designed to investigate how the selection of the distance function between vector embedding affects the reliability of identification of a speaker’s utterance in a segment-level classification architecture. The second experiment examines the architecture of cluster-centroid (group-level) classification, i.e., the selection of the best clustering and classification methods. The third experiment investigates the impact of different segmentation algorithms on the accuracy of identifying speaker utterances, and the fourth examines embedding window sizes. Experimental results demonstrated that the group-level approach offered better identification results were compared to the segment-level approach, and the latter had the advantage of real-time processing.
Title: Development of Supervised Speaker Diarization System Based on the PyAnnote Audio Processing Library
Description:
Diarization is an important task when work with audiodata is executed, as it provides a solution to the problem related to the need of dividing one analyzed call recording into several speech recordings, each of which belongs to one speaker.
Diarization systems segment audio recordings by defining the time boundaries of utterances, and typically use unsupervised methods to group utterances belonging to individual speakers, but do not answer the question “who is speaking?” On the other hand, there are biometric systems that identify individuals on the basis of their voices, but such systems are designed with the prerequisite that only one speaker is present in the analyzed audio recording.
However, some applications involve the need to identify multiple speakers that interact freely in an audio recording.
This paper proposes two architectures of speaker identification systems based on a combination of diarization and identification methods, which operate on the basis of segment-level or group-level classification.
The open-source PyAnnote framework was used to develop the system.
The performance of the speaker identification system was verified through the application of the AMI Corpus open-source audio database, which contains 100 h of annotated and transcribed audio and video data.
The research method consisted of four experiments to select the best-performing supervised diarization algorithms on the basis of PyAnnote.
The first experiment was designed to investigate how the selection of the distance function between vector embedding affects the reliability of identification of a speaker’s utterance in a segment-level classification architecture.
The second experiment examines the architecture of cluster-centroid (group-level) classification, i.
e.
, the selection of the best clustering and classification methods.
The third experiment investigates the impact of different segmentation algorithms on the accuracy of identifying speaker utterances, and the fourth examines embedding window sizes.
Experimental results demonstrated that the group-level approach offered better identification results were compared to the segment-level approach, and the latter had the advantage of real-time processing.
Related Results
Investigating the Role of Speaker Counter in Handling Overlapping Speeches in Speaker Diarization Systems
Investigating the Role of Speaker Counter in Handling Overlapping Speeches in Speaker Diarization Systems
In real-life conversations, meetings, or debates, there are often
situations where many people speak at the same time, leading to
overlapping speech segments. Such overlapping spee...
Multimodal Speaker Diarization Using a Pre-Trained Audio-Visual Synchronization Model
Multimodal Speaker Diarization Using a Pre-Trained Audio-Visual Synchronization Model
Speaker diarization systems aim to find ‘who spoke when?’ in multi-speaker recordings. The dataset usually consists of meetings, TV/talk shows, telephone and multi-party interactio...
Metaheuristic adapted convolutional neural network for Telugu speaker diarization
Metaheuristic adapted convolutional neural network for Telugu speaker diarization
In speech technology, a pivotal role is being played by the Speaker diarization mechanism. In general, speaker diarization is the mechanism of partitioning the input audio stream i...
Bangla Diarizz: Domain-Adapted Speaker Diarization for Bengali Long-Form Audio
Bangla Diarizz: Domain-Adapted Speaker Diarization for Bengali Long-Form Audio
Speaker diarization-the task of automatically determining who spoke when in a multi-speaker recording-remains a notable open challenge for low-resource languages, and Bengali (Bang...
No Sudden Audio Switch – Preventing discontinuous POI audio playing in LBS
No Sudden Audio Switch – Preventing discontinuous POI audio playing in LBS
Abstract. Many LBS applications provide automatic audio playing functions for introducing POI’s. Appropriate automatic audio playing can improve users’ expressions during traveling...
Robust speaker diarization for meetings
Robust speaker diarization for meetings
Aquesta tesi doctoral mostra la recerca feta en l'àrea de la diarització de locutor per a sales de reunions. En la present s'estudien els algorismes i la implementació d'un sistema...
Speaker Verification and Identification
Speaker Verification and Identification
A speaker recognition system verifies or identifies a speaker’s identity based on his/her voice. It is considered as one of the most convenient biometric characteristic for human m...
Improving Speaker Diarization for Overlapped Speech with Texture-Aware Feature Fusion
Improving Speaker Diarization for Overlapped Speech with Texture-Aware Feature Fusion
Speaker diarization (SD), which aims to address the “who spoke when” problem, is a key technology in speech processing. Although end-to-end neural speaker diarization methods have ...

