Paper
Comming soon
Authors
Takaaki Saeki,
Yuki Saito,
Shinnosuke Takamichi,
Hiroshi Saruwatari
(The University of Tokyo, Japan)
Abstract
This paper proposes two high-fidelity and computationally efficient neural voice conversion (VC) methods based on a direct waveform modification using spectral differentials.
The conventional spectral-differential VC method with a minimum-phase filter achieves high-quality conversion for narrow-band (16 kHz-sampled) VC but requires heavy computational cost in filtering.
This is because the minimum phase obtained using a fixed lifter of the Hilbert transform often results in a long-tap filter.
Furthermore, when we extend the method to full-band (48 kHz-sampled) VC, the computational cost is heavy due to increased sampling points, and the converted-speech quality degrades due to large fluctuations in the high-frequency band.
To construct a short-tap filter, we propose a lifter-training method for data-driven phase reconstruction that trains a lifter of the Hilbert transform by taking into account filter truncation.
We also propose a frequency-band-wise modeling method based on sub-band multi-rate signal processing (sub-band modeling method) for full-band VC.
It enhances the computational efficiency by reducing sampling points of signals converted with filtering and improves converted-speech quality by modeling only the low-frequency band.
We conducted several objective and subjective evaluations to investigate the effectiveness of the proposed methods through implementation of the real-time, online, full-band VC system we developed, which is based on the proposed methods.
The results indicate that 1) the proposed lifter-training method for narrow-band VC can shorten the tap length to 1/16 without degrading the converted-speech quality, and 2) the proposed sub-band modeling method for full-band VC can improve the converted-speech quality while reducing the computational cost, and 3) our real-time, online, full-band VC system can convert 48 kHz-sampled speech in real time attaining the converted speech with a 3.6 out of 5.0 mean opinion score of naturalness.
Source | Target | [Arakawa+, 2019] | Benchmark | Narrow-band+ | Full-band | Full-band+ | |
---|---|---|---|---|---|---|---|
Sample 1 | |||||||
Sample 2 | |||||||
Sample 3 |
NOTE: We only publish speech samples used in experimental evaluations for only female-to-female conversion due to licensing issues with 48 kHz-sampled JVS corpus [2].