Javascript must be enabled to continue!
GSNet: A Hybrid Graph Siamese Network for Robust Deepfake Audio Detection
View through CrossRef
As neural vocoders increasingly produce perceptually seamless deepfakes, conventional detectors relying on fragile, localized artifacts fail against real-world degradations such as lossy compression. To address this, we propose the Hybrid Graph Siamese Network (GSNet), which models the global structural integrity of audio signals rather than relying on local anomalies. Operating within a Siamese metric learning framework, GSNet integrates a CNN backbone for feature extraction with a Graph Neural Network (GNN) refiner to capture long-range, non-Euclidean acoustic dependencies. Empirical validation across multiple benchmarks demonstrates superior robustness: GSNet achieves an Equal Error Rate (EER) of 0.929% on the ASVspoof 2021 (DF) task and 0.536% on the ASVspoof 5 (Track 1) closed condition. Furthermore, the model demonstrates strong cross-domain generalization, attaining 0.048% EER on the WaveFake datasets. These results confirm that prioritizing structural consistency over artifact detection significantly enhances resilience against compression and unseen generative attacks.
Institute of Electrical and Electronics Engineers (IEEE)
Title: GSNet: A Hybrid Graph Siamese Network for Robust Deepfake Audio Detection
Description:
As neural vocoders increasingly produce perceptually seamless deepfakes, conventional detectors relying on fragile, localized artifacts fail against real-world degradations such as lossy compression.
To address this, we propose the Hybrid Graph Siamese Network (GSNet), which models the global structural integrity of audio signals rather than relying on local anomalies.
Operating within a Siamese metric learning framework, GSNet integrates a CNN backbone for feature extraction with a Graph Neural Network (GNN) refiner to capture long-range, non-Euclidean acoustic dependencies.
Empirical validation across multiple benchmarks demonstrates superior robustness: GSNet achieves an Equal Error Rate (EER) of 0.
929% on the ASVspoof 2021 (DF) task and 0.
536% on the ASVspoof 5 (Track 1) closed condition.
Furthermore, the model demonstrates strong cross-domain generalization, attaining 0.
048% EER on the WaveFake datasets.
These results confirm that prioritizing structural consistency over artifact detection significantly enhances resilience against compression and unseen generative attacks.
Related Results
No Sudden Audio Switch – Preventing discontinuous POI audio playing in LBS
No Sudden Audio Switch – Preventing discontinuous POI audio playing in LBS
Abstract. Many LBS applications provide automatic audio playing functions for introducing POI’s. Appropriate automatic audio playing can improve users’ expressions during traveling...
Evaluating the Threshold of Authenticity in Deepfake Audio and Its Implications Within Criminal Justice
Evaluating the Threshold of Authenticity in Deepfake Audio and Its Implications Within Criminal Justice
Deepfake technology has come a long way in recent years and the world has already seen cases where it has been used maliciously. After a deepfake of UK independent financial adviso...
How Frequency and Harmonic Profiling of a ‘Voice’ Can Inform Authentication of Deepfake Audio: An Efficiency Investigation
How Frequency and Harmonic Profiling of a ‘Voice’ Can Inform Authentication of Deepfake Audio: An Efficiency Investigation
As life in the digital era becomes more complex, the capacity for criminal activity within the digital realm becomes even more widespread. More recently, the development of deepfak...
Graph convolutional neural networks for 3D data analysis
Graph convolutional neural networks for 3D data analysis
(English) Deep Learning allows the extraction of complex features directly from raw input data, eliminating the need for hand-crafted features from the classical Machine Learning p...
Feature selection for multimodal: acoustic event detection
Feature selection for multimodal: acoustic event detection
The detection of the Acoustic Events (AEs) naturally produced in a meeting room may help to describe the human and social activity. The automatic description of interactions betwee...
Deepfake attack prevention using steganography GANs
Deepfake attack prevention using steganography GANs
Background
Deepfakes are fake images or videos generated by deep learning algorithms. Ongoing progress in deep learning techniques like auto-encoders and generative adversarial net...
Analysis of deepfake crime trends using BIGKinds
Analysis of deepfake crime trends using BIGKinds
This study is significant for analyzing criminal trends using deepfake technology based on media reports. A total of 478 articles related to crimes using deepfake technology were e...
Deepfake Detection using Deep Learning with InceptionV3
Deepfake Detection using Deep Learning with InceptionV3
Deepfake technology has rapidly evolved, making it increasingly difficult to distinguish between real and manipulated videos. This poses serious risks, including misinformation, id...

