Fast end-to-end non-parallel voice conversion based on speaker-adaptive neural vocoder with cycle-consistent learning

Shuhei Imai; Aoi Kanagaki; Takashi Nose; Shogo Fukawa; Akinori Ito

doi:10.1250/ast.e24.46

ACOUSTICAL LETTERS

Fast end-to-end non-parallel voice conversion based on speaker-adaptive neural vocoder with cycle-consistent learning

Shuhei Imai, Aoi Kanagaki, Takashi Nose, Shogo Fukawa, Akinori Ito

著者情報

キーワード: Voice conversion (VC), End-to-end VC, Non-parallel VC, Neural vocoder, Cycle-consistent learning

ジャーナルオープンアクセス

2025 年 46 巻 1 号 p. 116-119

DOI https://doi.org/10.1250/ast.e24.46

Browse “Advance online publication” version

詳細

抄録

This paper proposes a fast end-to-end non-parallel voice conversion (VC) named Tachylone. In Thachylone, speaker conversion and waveform generation is performed by a single vocoder network. In the training of Tachylone, a pre-trained universal neural vocoder is used as the initial model, and the model parameters are updated using source and target speakers' non-parallel data based on cycle-consistent learning in an end-to-end manner. We compare Tachylone to conventional CycleGAN-based VC with objective and subjective measures and discuss the results.

責任著者(Corresponding author)

J-STAGEへの登録はこちら（無料）