2026 年 30 巻 4 号 p. 989-1003
This study addresses the problems of imbalanced voice sequence, insufficient training stability, and symbol-audio modality mismatch in multi-track music generation. To this end, a gradient penalty-constrained multi-track adversarial training framework and a Wasserstein distance-driven cross-modal distribution alignment mechanism are studied and designed to optimize voice coordination accuracy and auditory perception authenticity. The core innovation lies in building a dual-module collaborative optimization architecture, pioneering gradient penalty constraints to eliminate multi-track training oscillations, and proposing Wasserstein feature space mapping mechanism to eradicate cross-modal perception mismatch, establishing a theoretical paradigm and technical path for joint optimization of voice and sound effects. Experimental verification: the voice conflict rate reaches 0.60%, the gradient norm variance is 1.13×10-3, and the training stability is controlled. The cross-modal distribution distance is 0.28, achieving precise alignment. The feature alignment error of 0.17 exceeds the technical limit, and the style fidelity is 91.7% (Baroque 94% / Jazz Blues 95%) to restore artistic expression. The dynamic expressive power is 4.66 points, approaching human creativity, and the auditory similarity is 0.89, establishing perceptual authenticity. The conflict rate of voice parts in the ablation test decreases by 52%, and cross-modal mismatch compression is reduced by 49%. Parameter sensitivity analysis shows that when the hidden space dimension is 64, the structural entropy is 0.810 and the perceptual similarity is 0.87, reaching the global optimum. This model significantly improves the coordination of multi-track structures and cross-modal perception quality, providing core technical support solutions for industrial-grade artificial intelligence (AI) music creation platforms such as film and television music composition and digital composition.
この記事は最新の被引用情報を取得できません。