EDBT 2026 Demo / reviewers in the wild / expert
Pengcheng Zhu 0004
dblp:37/5521-4
· DBLP profile ↗
16ranked-venue papers
1as first author
11since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 1 first-author · 11 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | REF-VC: Robust, Expressive and Fast Zero-Shot Voice Conversion with Diffusion TransformersabstractIn real-world voice conversion applications, environmental noise in source speech and user demands for expressive output pose critical challenges. Traditional ASR-based methods ensure noise robustness but suppress prosody richness, while SSL-based models improve expressiveness but suffer from timbre leakage and noise sensitivity. This paper proposes REF-VC, a noise-robust expressive voice conversion system. Key innovations include: (1) A random erasing strategy to mitigate the information redundancy inherent in SSL features, enhancing noise robustness and expressiveness; (2) Implicit alignment inspired by E2TTS to suppress non-essential feature reconstruction; (3) Integration of Shortcut Models to accelerate flow matching inference, significantly reducing to 4 steps. Experimental results demonstrate that REF-VC outperforms baselines such as Seed-VC in zero-shot scenarios on the noisy set, while also performing comparably to Seed-VC on the clean set. In addition, REF-VC can be compatible with singing voice conversion within one model. The samples can be found at: https://rxyj.github.io/asru2025/ Yuepeng Jiang, Ziqian Ning, Shuai Wang 0016, Chengjia Wang, Mengxiao Bi, Pengcheng Zhu 0004, Zhong-Hua Fu, Lei Xie 0001 |
ASRU | 6 |
| 2025 | MacST: Multi-Accent Speech Synthesis via Text Transliteration for Accent ConversionabstractIn accented voice conversion or accent conversion, we seek to convert the accent in speech from one another while preserving speaker identity and semantic content. In this study, we formulate a novel method for creating multi-accented speech samples, thus pairs of accented speech samples by the same speaker, through text transliteration for training accent conversion systems. We begin by generating transliterated text with Large Language Models (LLMs), which is then fed into multilingual TTS models to synthesize accented English speech. As a reference system, we built a sequence-to-sequence model on the synthetic parallel corpus for accent conversion. We validated the proposed method for both native and non-native English speakers. Subjective and objective evaluations further validate our dataset’s effectiveness in accent conversion studies. Sho Inoue, Shuai Wang 0016, Wanxing Wang, Pengcheng Zhu 0004, Mengxiao Bi, Haizhou Li 0001 |
ICASSP | 4 |
| 2025 | E1 TTS: Simple and Fast Non-Autoregressive TTSabstractThis paper introduces Easy One-Step Text-to-Speech (E1 TTS), an efficient non-autoregressive zero-shot text-to-speech system based on denoising diffusion pretraining and distribution matching distillation. The training of E1 TTS is straightforward; it does not require explicit monotonic alignment between the text and audio pairs. The inference of E1 TTS is efficient, requiring only one neural network evaluation for each utterance. Despite its sampling efficiency, E1 TTS achieves naturalness and speaker similarity comparable to various strong baseline models. Audio samples are available at e1tts.github.io. Shuai Wang 0016, Pengcheng Zhu 0004, Mengxiao Bi, Haizhou Li 0001 |
ICASSP | 3 |
| 2024 | Dualvc 2: Dynamic Masked Convolution for Unified Streaming and Non-Streaming Voice ConversionabstractVoice conversion is becoming increasingly popular, and a growing number of application scenarios require models with streaming inference capabilities. The recently proposed DualVC attempts to achieve this objective through streaming model architecture design and intra-model knowledge distillation along with hybrid predictive coding to compensate for the lack of future information. However, DualVC encounters several problems that limit its performance. First, the autoregressive decoder has error accumulation in its nature and limits the inference speed as well. Second, the causal convolution enables streaming capability but cannot sufficiently use future information within chunks. Third, the model is unable to effectively address the noise in the unvoiced segments, lowering the sound quality. In this paper, we propose DualVC 2 to address these issues. Specifically, the model backbone is migrated to a Conformer-based architecture, empowering parallel inference. Causal convolution is replaced by non-causal convolution with a dynamic chunk mask to make better use of within-chunk future information. Also, quiet attention is introduced to enhance the model’s noise robustness. Experiments show that DualVC 2 outperforms DualVC and other baseline systems in both subjective and objective metrics, with only 186.4 ms latency. Our audio samples are made publicly available1. Ziqian Ning, Yuepeng Jiang, Pengcheng Zhu 0004, Shuai Wang 0016, Jixun Yao, Lei Xie 0001, Mengxiao Bi |
ICASSP | 3 |
| 2024 | DualVC 3: Leveraging Language Model Generated Pseudo Context for End-to-end Low Latency Streaming Voice Conversion
Ziqian Ning, Shuai Wang 0016, Pengcheng Zhu 0004, Zhichao Wang 0002, Jixun Yao, Lei Xie 0001, Mengxiao Bi |
INTERSPEECH | 3 |
| 2023 | Expressive-VC: Highly Expressive Voice Conversion with Attention Fusion of Bottleneck and Perturbation FeaturesabstractVoice conversion for highly expressive speech is challenging. Current approaches struggle with the balance between speaker similarity, intelligibility, and expressiveness. To address this problem, we propose Expressive-VC, a novel end-to-end voice conversion framework that leverages advantages from both the neural bottleneck feature (BNF) approach and the information perturbation approach. Specifically, we use a BNF encoder and a Perturbed-Wav encoder to form a content extractor to learn linguistic and para-linguistic features respectively, where BNFs come from a robust pre-trained ASR model and the perturbed wave becomes speaker-irrelevant after signal perturbation. We further fuse the linguistic and para-linguistic features through an attention mechanism, where speaker-dependent prosody features are used as the attention query, which results from a prosody encoder with target speaker embedding and normalized pitch and energy of source speech as input. Finally, the decoder consumes the integrated features and the speaker-dependent prosody feature to generate the converted speech. Experiments show that Expressive-VC is superior to several popular systems, achieving both high expressiveness captured from the source speech and high speaker similarity with the target speaker; meanwhile intelligibility is well maintained. Ziqian Ning, Qicong Xie, Pengcheng Zhu 0004, Zhichao Wang 0002, Liumeng Xue, Jixun Yao, Lei Xie 0001, Mengxiao Bi |
ICASSP | 3 |
| 2023 | DualVC: Dual-mode Voice Conversion using Intra-model Knowledge Distillation and Hybrid Predictive Coding
Ziqian Ning, Yuepeng Jiang, Pengcheng Zhu 0004, Jixun Yao, Shuai Wang 0016, Lei Xie 0001, Mengxiao Bi |
INTERSPEECH | 3 |
| 2022 | One-Shot Voice Conversion For Style Transfer Based On Speaker AdaptationabstractOne-shot style transfer is a challenging task, since training on one utterance makes model extremely easy to over-fit to training data and causes low speaker similarity and lack of expressiveness. In this paper, we build on the recognition-synthesis framework and propose a one-shot voice conversion approach for style transfer based on speaker adaptation. First, a speaker normalization module is adopted to remove speaker-related information in bottleneck features extracted by ASR. Second, we adopt weight regularization in the adaptation process to prevent over-fitting caused by using only one utterance from target speaker as training data. Finally, to comprehensively decouple the speech factors, i.e., content, speaker, style, and transfer source style to the target, a prosody module is used to extract prosody representation. Experiments show that our approach is superior to the state-of-the-art one-shot VC systems in terms of style and speaker similarity; additionally, our approach also maintains good speech quality. Zhichao Wang 0002, Qicong Xie, Tao Li 0051, Hongqiang Du, Lei Xie 0001, Pengcheng Zhu 0004, Mengxiao Bi |
ICASSP | 6 |
| 2022 | VISinger: Variational Inference with Adversarial Learning for End-to-End Singing Voice SynthesisabstractIn this paper, we propose VISinger, a complete end-to-end high-quality singing voice synthesis (SVS) system that directly generates singing audio from lyrics and musical score. Our approach is inspired by VITS [1], an end-to-end speech generation model which adopts VAE-based posterior encoder augmented with normalizing flow based prior encoder and adversarial decoder. VISinger follows the main architecture of VITS, but makes substantial improvements to the prior encoder according to the characteristics of singing. First, instead of using phoneme-level mean and variance of acoustic features, we introduce a length regulator and a frame prior network to get the frame-level mean and variance on acoustic features, modeling the rich acoustic variation in singing. Second, we further introduce an F0 predictor to guide the frame prior network, leading to stabler singing performance. Finally, to improve the singing rhythm, we modify the duration predictor to specifically predict the phoneme to note duration ratio, helped with singing note normalization. Experiments on a professional Mandarin singing corpus show that VISinger significantly outperforms FastSpeech+Neural-Vocoder two-stage approach and the oracle VITS; ablation study demonstrates the effectiveness of different contributions. Yongmao Zhang, Jian Cong, Heyang Xue, Lei Xie 0001, Pengcheng Zhu 0004, Mengxiao Bi |
ICASSP | 5 |
| 2022 | Opencpop: A High-Quality Open Source Chinese Popular Song Corpus for Singing Voice SynthesisabstractThis paper introduces Opencpop, a publicly available highquality Mandarin singing corpus designed for singing voice synthesis (SVS).The corpus consists of 100 popular Mandarin songs performed by a female professional singer.Audio files are recorded with studio quality at a sampling rate of 44,100 Hz and the corresponding lyrics and musical scores are provided.All singing recordings have been phonetically annotated with phoneme boundaries and syllable (note) boundaries.To demonstrate the reliability of the released data and to provide a baseline for future research, we built baseline deep neural network-based SVS models and evaluated them with both objective metrics and subjective mean opinion score (MOS) measure.Experimental results show that the best SVS model trained on our database achieves 3.70 MOS, indicating the reliability of the provided corpus.Opencpop is released to the open-source community WeNet 1 , and the corpus, as well as synthesized demos, can be found on the project homepage 2 . Pengcheng Zhu 0004, Jie Wu 0017, Hanzhao Li, Heyang Xue, Yongmao Zhang, Lei Xie 0001, Mengxiao Bi |
INTERSPEECH | 3 |
| 2022 | Learn2Sing 2.0: Diffusion and Mutual Information-Based Target Speaker SVS by Learning from Singing TeacherabstractBuilding a high-quality singing corpus for a person who is not good at singing is non-trivial, thus making it challenging to create a singing voice synthesizer for this person.Learn2Sing is dedicated to synthesizing the singing voice of a speaker without his or her singing data by learning from data recorded by others, i.e., the singing teacher.Inspired by the fact that pitch is the key style factor to distinguish singing from speaking voice, the proposed Learn2Sing 2.0 first generates the preliminary acoustic feature with averaged pitch value in the phone level, which allows the training of this process for different styles, i.e., speaking or singing, share same conditions except for the speaker information.Then, conditioned on the specific style, a diffusion decoder, which is accelerated by a fast sampling algorithm during the inference stage, is adopted to gradually restore the final acoustic feature.During the training, to avoid the information confusion of the speaker embedding and the style embedding, mutual information is employed to restrain the learning of speaker embedding and style embedding.Experiments show that the proposed approach is capable of synthesizing highquality singing voice for the target speaker without singing data with 10 decoding steps. Heyang Xue, Yongmao Zhang, Lei Xie 0001, Pengcheng Zhu 0004, Mengxiao Bi |
INTERSPEECH | 5 |
| 2019 | Improving Mandarin End-to-End Speech Synthesis by Self-Attention and Learnable Gaussian BiasabstractCompared to conventional speech synthesis, end-to-end speech synthesis has achieved much better naturalness with more simplified system building pipeline. End-to-end framework can generate natural speech directly from characters for English. But for other languages like Chinese, recent studies have indicated that extra engineering features are still needed for model robustness and naturalness, e.g, word boundaries and prosody boundaries, which makes the front-end pipeline as complicated as the traditional approach. To maintain the naturalness of generated speech and discard language-specific expertise as much as possible, in Mandarin TTS, we introduce a novel self-attention based encoder with learnable Gaussian bias in Tacotron. We evaluate different systems with and without complex prosody information and results show that the proposed approach has the ability to generate stable and natural speech with minimum language-dependent front-end modules. Fengyu Yang 0001, Shan Yang 0001, Pengcheng Zhu 0004, Pengju Yan, Lei Xie 0001 |
ASRU | 3 |
| 2015 | BLSTM neural networks for speech driven head motion synthesis
Chuang Ding, Pengcheng Zhu 0004, Lei Xie 0001 |
INTERSPEECH | 2 |
| 2015 | Articulatory movement prediction using deep bidirectional long short-term memory based recurrent neural networks and word/phone embeddingsabstractAutomatic prediction of articulatory movements from speech or text can be beneficial for many applications such as speech recognition and synthesis. A recent approach has reported stateof-the-art performance in speech-to-articulatory prediction using feed forward neural networks. In this paper, we investigate the feasibility of using bidirectional long short-term memory based recurrent neural networks (BLSTM-RNNs) in articulatory movement prediction because they have long-context trajectory modeling ability. We show on the MNGU0 dataset that BLSTM-RNN apparently outperforms feed forward networks and pushes the state-of-the-art RMSE from 0.885 mm to 0.565 mm. On the other hand, predicting articulatory information from text heavily relies on handcrafted linguistic and prosodic features, e.g., POS and TOBI labels. In this paper, we propose to use word and phone embeddings to substitute these manual features. Word/phone embedding features are automatically learned from unlabeled text data by a neural network language model. We show that word and phone embeddings can achieve comparable performance without using POS and TOBI features. More promisingly, combining the conventional full feature set with phone embedding, the lowest RMSE is achieved. Pengcheng Zhu 0004, Lei Xie 0001, Yunlin Chen |
INTERSPEECH | 1 |
| 2015 | Head motion synthesis from speech using deep neural networks
Chuang Ding, Lei Xie 0001, Pengcheng Zhu 0004 |
Multim. Tools Appl. | 3 |
| 2014 | Speech-driven head motion synthesis using neural networksabstractThis paper presents a neural network approach for speech-driven head motion synthesis, which can automatically predict a speaker’s head movement from his/her speech. Specifically, we realize speech-to-head-motion mapping by learning a multi-layer perceptron from audio-visual broadcast news data. First, we show that a generatively pre-trained neural network significantly outperforms a randomly initialized network and the hidden Markov model (HMM) approach. Second, we demonstrate that the feature combination of log Mel-scale filter-bank (FBank), energy and fundamental frequency (F0) performs best in head motion prediction. Third, we discover that using long context acoustic information can further improve the performance. Finally, extra unlabeled training data used in the pre-training stage can achieve more performance gain. The proposed speech-driven head motion synthesis approach increases the CCA from 0.299 (the HMM approach) to 0.565 and it can be effectively used in expressive talking avatar animation. Index Terms: head motion synthesis, neural network, deep neural network, talking avatar Chuang Ding, Pengcheng Zhu 0004, Lei Xie 0001, Dongmei Jiang, Zhong-Hua Fu |
INTERSPEECH | 2 |