Exemplar-Based Voice Conversion Using Sparse Representation in Noisy Environments

Ryoichi TAKASHIMA; Tetsuya TAKIGUCHI; Yasuo ARIKI

doi:10.1587/transfun.E96.A.1946

Abstract

This paper presents a voice conversion (VC) technique for noisy environments, where parallel exemplars are introduced to encode the source speech signal and synthesize the target speech signal. The parallel exemplars (dictionary) consist of the source exemplars and target exemplars, having the same texts uttered by the source and target speakers. The input source signal is decomposed into the source exemplars, noise exemplars and their weights (activities). Then, by using the weights of the source exemplars, the converted signal is constructed from the target exemplars. We carried out speaker conversion tasks using clean speech data and noise-added speech data. The effectiveness of this method was confirmed by comparing its effectiveness with that of a conventional Gaussian Mixture Model (GMM)-based method.

Content from these authors

Favorites & Alerts

Corresponding author

Register with J-STAGE for free!