Hongqiang Du

dblp:258/8437 · DBLP profile ↗
← Back
8ranked-venue papers
6as first author
6since 2021 · last 2022
0000-0002-4168-9655ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 3 since 2021
YearPublicationVenuePosition
2022 One-Shot Voice Conversion For Style Transfer Based On Speaker Adaptation
abstract
One-shot style transfer is a challenging task, since training on one utterance makes model extremely easy to over-fit to training data and causes low speaker similarity and lack of expressiveness. In this paper, we build on the recognition-synthesis framework and propose a one-shot voice conversion approach for style transfer based on speaker adaptation. First, a speaker normalization module is adopted to remove speaker-related information in bottleneck features extracted by ASR. Second, we adopt weight regularization in the adaptation process to prevent over-fitting caused by using only one utterance from target speaker as training data. Finally, to comprehensively decouple the speech factors, i.e., content, speaker, style, and transfer source style to the target, a prosody module is used to extract prosody representation. Experiments show that our approach is superior to the state-of-the-art one-shot VC systems in terms of style and speaker similarity; additionally, our approach also maintains good speech quality.
Zhichao Wang 0002, Qicong Xie, Tao Li 0051, Hongqiang Du, Lei Xie 0001, Pengcheng Zhu 0004, Mengxiao Bi
ICASSP4
2022 Noise-robust voice conversion with domain adversarial training
Hongqiang Du, Lei Xie 0001, Haizhou Li 0001
Neural Networks1
2021 Improving Robustness of One-Shot Voice Conversion with Deep Discriminative Speaker Encoder
abstract
One-shot voice conversion has received significant attention since only one utterance from source speaker and target speaker respectively is required.Moreover, source speaker and target speaker do not need to be seen during training.However, available one-shot voice conversion approaches are not stable for unseen speakers as the speaker embedding extracted from one utterance of an unseen speaker is not reliable.In this paper, we propose a deep discriminative speaker encoder to extract speaker embedding from one utterance more effectively.Specifically, the speaker encoder first integrates residual network and squeeze-and-excitation network to extract discriminative speaker information in frame level by modeling framewise and channel-wise interdependence in features.Then attention mechanism is introduced to further emphasize speaker related information via assigning different weights to frame level speaker information.Finally a statistic pooling layer is used to aggregate weighted frame level speaker information to form utterance level speaker embedding.The experimental results demonstrate that our proposed speaker encoder can improve the robustness of one-shot voice conversion for unseen speakers and outperforms baseline systems in terms of speech quality and speaker similarity.
Hongqiang Du
Interspeech1
2021 Enriching Source Style Transfer in Recognition-Synthesis Based Non-Parallel Voice Conversion
abstract
Current voice conversion (VC) methods can successfully convert timbre of the audio.As modeling source audio's prosody effectively is a challenging task, there are still limitations of transferring source style to the converted speech.This study proposes a source style transfer method based on recognitionsynthesis framework.Previously in speech generation task, prosody can be modeled explicitly with prosodic features or implicitly with a latent prosody extractor.In this paper, taking advantages of both, we model the prosody in a hybrid manner, which effectively combines explicit and implicit methods in a proposed prosody module.Specifically, prosodic features are used to explicit model prosody, while VAE and reference encoder are used to implicitly model prosody, which take Mel spectrum and bottleneck feature as input respectively.Furthermore, adversarial training is introduced to remove speakerrelated information from the VAE outputs, avoiding leaking source speaker information while transferring style.Finally, we use a modified self-attention based encoder to extract sentential context from bottleneck features, which also implicitly aggregates the prosodic aspects of source speech from the layered representations.Experiments show that our approach is superior to the baseline and a competitive system in terms of style transfer; meanwhile, the speech quality and speaker similarity are well maintained.
Zhichao Wang 0002, Xinyong Zhou, Fengyu Yang 0002, Tao Li 0051, Hongqiang Du, Lei Xie 0001, Wendong Gan
Interspeech5
2021 Optimizing Voice Conversion Network with Cycle Consistency Loss of Speaker Identity
abstract
We propose a novel training scheme to optimize voice conversion network with a speaker identity loss function. The training scheme not only minimizes frame-level spectral loss, but also speaker identity loss. We introduce a cycle consistency loss that constrains the converted speech to maintain the same speaker identity as reference speech at utterance level. While the proposed training scheme is applicable to any voice conversion networks, we formulate the study under the average model voice conversion framework in this paper. Experiments conducted on CMU-ARCTIC and CSTR-VCTK corpus confirm that the proposed method outperforms baseline methods in terms of speaker similarity.
Hongqiang Du, Xiaohai Tian, Lei Xie 0001, Haizhou Li 0001
SLT1
2021 Factorized WaveNet for voice conversion with limited data
Hongqiang Du, Xiaohai Tian, Lei Xie 0001, Haizhou Li 0001
Speech Commun.1
2020 Effective Wavenet Adaptation for Voice Conversion with Limited Data
abstract
WaveNet has shown its great potential as a direct conversion model in voice conversion. However, due to the model complexity, WaveNet always requires a large amount of training data, which has limited its applications in voice conversion, where training data is scarce. In this paper, we propose a WaveNet adaptation method that effectively reduces the need of adaptation data. We first train a speaker independent WaveNet conversion model with multi-speaker dataset. Adaptation is then applied with limited target speaker’s data. Specifically, singular value decomposition (SVD) is applied to dilated convolution layers of WaveNet to reduce the number of parameters, which makes adaptation more effective with limited data. Experiments conducted on CMU-ARCTIC and CSTR-VCTK corpus show that the proposed method outperforms baseline methods in terms of both quality and similarity.
Hongqiang Du, Xiaohai Tian, Lei Xie 0001, Haizhou Li 0001
ICASSP1
2019 WaveNet Factorization with Singular Value Decomposition for Voice Conversion
abstract
WaveNet vocoder has seen its great advantage over traditional vocoders in voice quality. However, it usually requires a relatively large amount of speech data to train a speaker-dependent WaveNet vocoder. Therefore, it remains a challenge to build a high-quality WaveNet vocoder for low resource tasks, e.g. voice conversion, where speech samples are limited in real applications. We propose to use singular value decomposition (SVD) to reduce WaveNet parameters while maintaining its output voice quality. Specifically, we apply SVD on dilated convolution layers, and impose semi-orthogonal constraint to improve the performance. Experiments conducted on CMU-ARCTIC database show that as compared with the original WaveNet vocoder, the proposed method maintains similar performance, in terms of both quality and similarity, while using much less training data.
Hongqiang Du, Xiaohai Tian, Lei Xie 0001, Haizhou Li 0001
ASRU1