Journal of Signal Processing
Online ISSN : 1880-1013
Print ISSN : 1342-6230
ISSN-L : 1342-6230
Enhancing Model Robustness for Deepfake Singing Voice Detection through Data Augmentation
Islam J. A. M. SamiulKhalid ZamanKai LiAnuwat ChaiwongyenShogo OkadaMasashi Unoki
Author information
JOURNAL FREE ACCESS

2025 Volume 29 Issue 6 Pages 221-225

Details
Abstract

The rapid advancements in artificial intelligence (AI) technologies for singing voice synthesis have revolutionized music production but introduced significant challenges, including the misuse of AI-generated voices for deepfake purposes. This study proposes a framework for detecting deepfake singing voices using data augmentation techniques, the Synthetic Minority Over- sampling Technique (SMOTE), time stretching, and time shifting, combined with RawNet-based waveform feature extraction integrated into the Audio Anti-Spoofing Using Integrated Spectro-Temporal Graph Attention Networks (AASIST) model. The method effectively captures subtle temporal and spectral artifacts, achieving an equal error rate (EER) of 10.39% on the SVDD SingFake 2024 dataset, outperforming baseline models. This work provides a robust solution to safeguard the authenticity of AI-generated singing voices in digital media.

Content from these authors
© 2025 Research Institute of Signal Processing, Japan
Previous article
feedback
Top