VLDB 2026 Research / reviewers in the wild / expert
Shifeng Pan
dblp:73/9231
· DBLP profile ↗
12ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0002-9338-7800ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Drop the Beat! Freestyler for Accompaniment Conditioned Rapping Voice GenerationabstractRap, a prominent genre of vocal performance, remains underexplored in vocal generation. General vocal synthesis depends on precise note and duration inputs, requiring users to have related musical knowledge, which limits flexibility. In contrast, rap typically features simpler melodies, with a core focus on a strong rhythmic sense that harmonizes with accompanying beats. In this paper, we propose Freestyler, the first system that generates rapping vocals directly from lyrics and accompaniment inputs. Freestyler utilizes language model-based token generation, followed by a conditional flow matching model to produce spectrograms and a neural vocoder to restore audio. It allows a 3-second prompt to enable zero-shot timbre control. Due to the scarcity of publicly available rap datasets, we also present RapBank, a rap song dataset collected from the internet, alongside a meticulously designed processing pipeline. Experimental results show that Freestyler produces high-quality rapping voice generation with enhanced naturalness and strong alignment with accompanying beats, both stylistically and rhythmically. Ziqian Ning, Shuai Wang 0016, Yuepeng Jiang, Jixun Yao, Lei He 0005, Shifeng Pan, Lei Xie 0001 |
AAAI | 6 |
| 2022 | Infergrad: Improving Diffusion Models for Vocoder by Considering Inference in TrainingabstractDenoising diffusion probabilistic models (diffusion models for short) require a large number of iterations in inference to achieve the generation quality that matches or surpasses the state-of-the-art generative models, which invariably results in slow inference speed. Previous approaches aim to optimize the choice of inference schedule over a few iterations to speed up inference. However, this results in reduced generation quality, mainly because the inference process is optimized separately, without jointly optimizing with the training process. In this paper, we propose InferGrad, a diffusion model for vocoder that incorporates inference process into training, to reduce the inference iterations while maintaining high generation quality. More specifically, during training, we generate data from random noise through a reverse process under inference schedules with a few iterations, and impose a loss to minimize the gap between the generated and ground-truth data samples. Then, unlike existing approaches, the training of InferGrad considers the inference process. The advantages of InferGrad are demonstrated through experiments on the LJSpeech dataset showing that InferGrad achieves better voice quality than the baseline WaveGrad under same conditions while maintaining the same voice quality as the baseline but with 3x speedup (2 iterations for InferGrad vs 6 iterations for WaveGrad). Zehua Chen 0005, Xu Tan 0003, Shifeng Pan, Danilo P. Mandic, Lei He 0005, Sheng Zhao 0002 |
ICASSP | 4 |
| 2022 | Prosodyspeech: Towards Advanced Prosody Model for Neural Text-to-SpeechabstractThis paper proposes ProsodySpeech, a novel prosody model to enhance encoder-decoder neural Text-To-Speech (TTS), to generate high expressive and personalized speech even with very limited training data. First, we use a Prosody Extractor built from a large speech corpus with various speakers to generate a set of prosody exemplars from multiple reference speeches, in which Mutual Information based Style content separation (MIST) is adopted to alleviate "content leakage" problem. Second, we use a Prosody Distributor to make a soft selection of appropriate prosody exemplars in phone-level with the help of an attention mechanism. The resulting prosody feature is then aggregated into the output of text encoder, together with additional phone-level pitch feature to enrich the prosody. We apply this method into two tasks: highly expressive multi style/emotion TTS and few-shot personalized TTS. The experiments show the proposed model outperforms baseline FastSpeech 2 + GST with significant improvements in terms of similarity and style expression. Yuanhao Yi, Lei He 0005, Shifeng Pan, Xi Wang 0016, Yujia Xiao |
ICASSP | 3 |
| 2022 | SoftSpeech: Unsupervised Duration Model in FastSpeech 2
Yuanhao Yi, Lei He 0005, Shifeng Pan, Xi Wang 0016 |
INTERSPEECH | 3 |
| 2021 | Cross-Speaker Style Transfer with Prosody Bottleneck in Neural Speech SynthesisabstractCross-speaker style transfer is crucial to the applications of multi-style and expressive speech synthesis at scale. It does not require the target speakers to be experts in expressing all styles and to collect corresponding recordings for model training. However, the performances of existing style transfer methods are still far behind real application needs. The root causes are mainly twofold. Firstly, the style embedding extracted from single reference speech can hardly provide fine-grained and appropriate prosody information for arbitrary text to synthesize. Secondly, in these models the content/text, prosody, and speaker timbre are usually highly entangled, it's therefore not realistic to expect a satisfied result when freely combining these components, such as to transfer speaking style between speakers. In this paper, we propose a cross-speaker style transfer text-to-speech (TTS) model with explicit prosody bottleneck. The prosody bottleneck builds up the kernels accounting for speaking style robustly, and disentangles the prosody from content and speaker timbre, therefore guarantees high quality cross-speaker style transfer. Evaluation result shows the proposed method even achieves on-par performance with source speaker's speaker-dependent (SD) model in objective measurement of prosody, and significantly outperforms the cycle consistency and GMVAE-based baselines in objective and subjective evaluations. Shifeng Pan, Lei He 0005 |
Interspeech | 1 |
| 2021 | Cycle consistent network for end-to-end style transfer TTS training
Liumeng Xue, Shifeng Pan, Lei He 0005, Lei Xie 0001, Frank K. Soong |
Neural Networks | 2 |
| 2019 | Learning Latent Representations for Style Control and Transfer in End-to-end Speech SynthesisabstractIn this paper, we introduce the Variational Autoencoder (VAE) to an end-to-end speech synthesis model, to learn the latent representation of speaking styles in an unsupervised manner. The style representation learned through VAE shows good properties such as disentangling, scaling, and combination, which makes it easy for style control. Style transfer can be achieved in this framework by first inferring style representation through the recognition network of VAE, then feeding it into TTS network to guide the style in synthesizing speech. To avoid Kullback-Leibler (KL) divergence collapse in training, several techniques are adopted. Finally, the proposed model shows good performance of style control and outperforms Global Style Token (GST) model in ABX preference tests on style transfer. Shifeng Pan, Lei He 0005, Zhen-Hua Ling |
ICASSP | 2 |
| 2019 | Forward-Backward Decoding for Regularizing End-to-End TTSabstractNeural end-to-end TTS can generate very high-quality synthesized speech, and even close to human recording within similar domain text. However, it performs unsatisfactory when scaling it to challenging test sets. One concern is that the encoder-decoder with attention-based network adopts autoregressive generative sequence model with the limitation of exposure bias To address this issue, we propose two novel methods, which learn to predict future by improving agreement between forward and backward decoding sequence. The first one is achieved by introducing divergence regularization terms into model training objective to reduce the mismatch between two directional models, namely L2R and R2L (which generates targets from left-to-right and right-to-left, respectively). While the second one operates on decoder-level and exploits the future information during decoding. In addition, we employ a joint training strategy to allow forward and backward decoding to improve each other in an interactive process. Experimental results show our proposed methods especially the second one (bidirectional decoder regularization), leads a significantly improvement on both robustness and overall naturalness, as outperforming baseline (the revised version of Tacotron2) with a MOS gap of 0.14 in a challenging test, and achieving close to human quality (4.42 vs. 4.49 in MOS) on general test. Yibin Zheng, Xi Wang 0016, Lei He 0005, Shifeng Pan, Frank K. Soong, Zhengqi Wen, Jianhua Tao 0001 |
INTERSPEECH | 4 |
| 2011 | The CASIA Audio Emotion Recognition Method for Audio/Visual Emotion Challenge 2011
Shifeng Pan, Jianhua Tao 0001, Ya Li 0001 |
ACII (2) | 1 |
| 2011 | Global variance modeling on frequency domain delta LSP for HMM-based speech synthesisabstractThe speech parameter generation algorithm considering global variance (GV) for HMM-based speech synthesis proved to be effective against the over-smoothing problem. However, the correlation between dimensions of parameter vector is not sufficiently considered in the current GV model. For some parameters, e.g., Line Spectral Pairs (LSP), the difference of adjacent LSPs has strong influence on the spectral envelope. Considering this important feature, the paper proposes a GV modeling on the difference of adjacent LSPs, i.e., GV on frequency domain delta LSP. By improving the GV likelihood on frequency domain delta LSP, the over-smoothing effect of generated parameter trajectory is better alleviated than conventional one. The result of a perceptual evaluation shows the proposed method outperforms the conventional one, and the naturalness of synthetic speech is improved. Shifeng Pan, Yoshihiko Nankaku, Keiichi Tokuda, Jianhua Tao 0001 |
ICASSP | 1 |
| 2010 | Text-based unstressed syllable prediction in MandarinabstractRecently, an increasing attention has been paid to Mandarin word stress which is important for improving the naturalness of speech synthesis. Most of the research on Mandarin speech synthesis focuses on three stress levels: stressed, regular and unstressed. This paper emphasizes the unstressed syllable prediction because the unstressed syllable is also important to the intelligibility of the synthetic speech. Similar as the prosodic structure, it is not easy to detect stress from text analysis due to the complicated context information. A method based on Classification and Regression Tree (CART) model has been proposed to predict the unstressed syllables with the high accuracy of 85%. The method has been finally applied into the TTS system. The experiment shows that the MOS score of synthetic speech has been improved by 0.35; the pitch contour of the new synthesized speech is also closer to natural speech. Index Terms: Text-to-Speech, stress, unstressed syllable, prosody Ya Li 0001, Jianhua Tao 0001, Shifeng Pan, Xiaoying Xu |
INTERSPEECH | 4 |
| 2010 | A novel hybrid approach for Mandarin speech synthesisabstractThe paper investigates a new method to solve concatenation problems of Mandarin speech synthesis which is based on the hybrid approach of HMM-based speech synthesis and unit selection. Unlike other works which use only boundary F0 errors as concatenation cost, a CART based F0 dependency model which considers much context information is trained to measure smoothness of F0. Instead of phoneme-sized units, the basic units of our HUS system are syllables, which has been proved to be better for the prosody stability in Mandarin. The experiments show that the proposed method achieves better performance than conventional hybrid system and unit selection system. Index Terms: Speech synthesis, hidden Markov model, unit selection, hybrid Shifeng Pan, Jianhua Tao 0001 |
INTERSPEECH | 1 |