Journal of Signal Processing
Online ISSN : 1880-1013
Print ISSN : 1342-6230
ISSN-L : 1342-6230
Enhancing Model Robustness for Deepfake Singing Voice Detection through Data Augmentation
Islam J. A. M. SamiulKhalid ZamanKai LiAnuwat ChaiwongyenShogo OkadaMasashi Unoki
著者情報
キーワード: singing voice, deepfake, SMOTE, AASIST
ジャーナル フリー

2025 年 29 巻 6 号 p. 221-225

詳細
抄録

The rapid advancements in artificial intelligence (AI) technologies for singing voice synthesis have revolutionized music production but introduced significant challenges, including the misuse of AI-generated voices for deepfake purposes. This study proposes a framework for detecting deepfake singing voices using data augmentation techniques, the Synthetic Minority Over- sampling Technique (SMOTE), time stretching, and time shifting, combined with RawNet-based waveform feature extraction integrated into the Audio Anti-Spoofing Using Integrated Spectro-Temporal Graph Attention Networks (AASIST) model. The method effectively captures subtle temporal and spectral artifacts, achieving an equal error rate (EER) of 10.39% on the SVDD SingFake 2024 dataset, outperforming baseline models. This work provides a robust solution to safeguard the authenticity of AI-generated singing voices in digital media.

著者関連情報
© 2025 Research Institute of Signal Processing, Japan
前の記事
feedback
Top