2025 Volume 29 Issue 6 Pages 221-225
The rapid advancements in artificial intelligence (AI) technologies for singing voice synthesis have revolutionized music production but introduced significant challenges, including the misuse of AI-generated voices for deepfake purposes. This study proposes a framework for detecting deepfake singing voices using data augmentation techniques, the Synthetic Minority Over- sampling Technique (SMOTE), time stretching, and time shifting, combined with RawNet-based waveform feature extraction integrated into the Audio Anti-Spoofing Using Integrated Spectro-Temporal Graph Attention Networks (AASIST) model. The method effectively captures subtle temporal and spectral artifacts, achieving an equal error rate (EER) of 10.39% on the SVDD SingFake 2024 dataset, outperforming baseline models. This work provides a robust solution to safeguard the authenticity of AI-generated singing voices in digital media.