VLDB 2026 Research / reviewers in the wild / expert
Fenglong Xie
dblp:270/4593 · also Feng-Long Xie
· DBLP profile ↗
14ranked-venue papers
6as first author
7since 2021 · last 2025
0000-0002-1206-3696ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 6 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Speaking from Coarse to Fine: Improving Neural Codec Language Model via Multi-Scale Speech Coding and GenerationabstractThe neural codec language model (CLM) has demonstrated remarkable performance in text-to-speech (TTS) synthesis. However, troubled by "recency bias", CLM lacks sufficient attention to coarse-grained information at a higher temporal scale, often producing unnatural or even unintelligible speech. This work proposes CoFi-Speech, a coarse-to-fine CLM-TTS approach, employing multi-scale speech coding and generation to address this issue. We train a multi-scale neural codec, CoFi-Codec, to encode speech into a multi-scale discrete representation, comprising multiple token sequences with different time resolutions. Then, we propose CoFi-LM that can generate this representation in two modes: the single-LM-based chain-of-scale generation and the multiple-LM-based stack-of-scale generation. In experiments, CoFi-Speech significantly outperforms single-scale baseline systems on naturalness and speaker similarity in zero-shot TTS. The analysis of multi-scale coding demonstrates the effectiveness of CoFi-Codec in learning multi-scale discrete speech representations while keeping high-quality speech reconstruction. The coarse-to-fine multi-scale generation, especially for the stack-of-scale approach, is also validated as a crucial approach in pursuing a high-quality neural codec language model for TTS. Haohan Guo, Fenglong Xie, Dongchao Yang, Xixin Wu, Helen M. Meng |
ICASSP | 2 |
| 2024 | SoCodec: A Semantic-Ordered Multi-Stream Speech Codec For Efficient Language Model Based Text-to-Speech SynthesisabstractThe long speech sequence has been troubling language models (LM) based TTS approaches in terms of modeling complexity and efficiency. This work proposes SoCodec, a semantic-ordered multi-stream speech codec, to address this issue. It compresses speech into a shorter, multi-stream discrete semantic sequence with multiple tokens at each frame. Meanwhile, the ordered product quantization is proposed to constrain this sequence into an ordered representation. It can be applied with a multi-stream delayed LM to achieve better autoregressive generation along both time and stream axes in TTS. The experimental result strongly demonstrates the effectiveness of the proposed approach, achieving superior performance over baseline systems even if compressing the frameshift of speech from 20 ms to 240 ms (12 x). The ablation studies further validate the importance of learning the proposed ordered multi-stream semantic representation in pursuing shorter speech sequences for efficient LM-based TTS. Haohan Guo, Fenglong Xie, Dongchao Yang, Dake Guo, Xixin Wu, Helen M. Meng |
SLT | 2 |
| 2024 | Addressing Index Collapse of Large-Codebook Speech Tokenizer With Dual-Decoding Product-Quantized Variational Auto-EncoderabstractVQ-VAE, as a mainstream approach of speech tokenizer, has been troubled by “index collapse”, where only a small number of codewords are activated in large codebooks. This work proposes product-quantized (PQ) VAE with more codebooks but fewer codewords to address this problem and build large-codebook speech tokenizers. It encodes speech features into multiple VQ subspaces and composes them into codewords in a larger codebook. Besides, to utilize each VQ subspace well, we also enhance PQ-VAE via a dual-decoding training strategy with the encoding and quantized sequences. The experimental results demonstrate that PQ-VAE addresses “index collapse” effectively, especially for larger codebooks. The model with the proposed training strategy further improves codebook perplexity and reconstruction quality, outperforming other multi-codebook VQ approaches. Finally, PQ-VAE demonstrates its effectiveness in language-model-based TTS, supporting higher-quality speech generation with larger codebooks. Haohan Guo, Fenglong Xie, Dongchao Yang, Xixin Wu, Helen M. Meng |
SLT | 2 |
| 2023 | MSMC-TTS: Multi-Stage Multi-Codebook VQ-VAE Based Neural TTSabstractThis paper aims to improve neural TTS with vector-quantized, compact speech representations. We propose a Vector-Quantized Variational AutoEncoder (VQ-VAE) based feature analyzer to encode acoustic features into sequences with different time resolutions, and quantize them with multiple VQ codebooks to form the Multi-Stage Multi-Codebook Representation (MSMCR). The TTS system, MSMC-TTS, is proposed to predict better speech via this representation. In prediction, the multi-stage predictor is trained to map the input text sequence to MSMCRs in stages, by minimizing Euclidean distance and “triplet loss”. In synthesis, the neural vocoder converts ground-truth or predicted MSMCRs into speech waveforms. The proposed system is trained with single-speaker TTS datasets and tested in various scenarios for comprehensive evaluation. In TTS evaluation, MSMC-TTS obtains MOS of 4.34 and 4.10 on English and Chinese datasets, which significantly outperforms VITS with scores of 3.78 and 3.90. Meanwhile, compared with Mel-Spectrograms, the domain discrepancy between prediction and ground truth is lower in MSMCRs with the higher Domain-classification Error Rate (DER). Furthermore, this system shows lower modeling complexity and data size requirements, preserving excellent performance even with fewer model parameters or training data. The noticeable improvement in analysis-synthesis and TTS from multiple codebooks and stages also validate them as vital components in seeking a more profitable speech representation and building high-performance neural TTS. Haohan Guo, Fenglong Xie, Xixin Wu, Frank K. Soong, Helen M. Meng |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | A Multi-Stage Multi-Codebook VQ-VAE Approach to High-Performance Neural TTSabstractWe propose a Multi-Stage, Multi-Codebook (MSMC) approach to high performance neural TTS synthesis.A vector-quantized, variational autoencoder (VQ-VAE) based feature analyzer is used to encode Mel spectrograms of speech training data by down-sampling progressively in multiple stages into MSMC Representations (MSMCRs) with different time resolutions, and quantizing them with multiple VQ codebooks, respectively.Multi-stage predictors are trained to map the input text sequence to MSMCRs progressively by minimizing a combined loss of the reconstruction Mean Square Error (MSE) and "triplet loss".In synthesis, the neural vocoder converts the predicted MSM-CRs into final speech waveforms.The proposed approach is trained and tested with an English TTS database of 16 hours by a female speaker.The proposed TTS achieves an MOS score of 4.41, which outperforms the baseline with an MOS of 3.62.Compact versions of the proposed TTS with much less parameters can still preserve high MOS scores.Ablation studies show that both multiple stages and multiple codebooks are effective for achieving high TTS performance. Haohan Guo, Fenglong Xie, Frank K. Soong, Xixin Wu, Helen M. Meng |
INTERSPEECH | 2 |
| 2021 | A New High Quality Trajectory Tiling Based Hybrid TTS In Real TimeabstractA trajectory tiling based, hybrid TTS is revisited in this study for improving its synthesis performance. A combination of Transformer encoder and RNN based decoder architecture where two-level, at both word and Chinese phonetic alphabet letter levels, linguistic representation is exploited to generate a cogent and smooth speech parameter trajectory. And then a segment candidate lattice is constructed by minimizing the log spectral distortion of mel-spectrograms and RMSE of F0 between the generated trajectory and candidates. Normalized cross-correlation is used to find the best sequence of "wave-form tiles" in the lattice for synthesizing the final speech waveforms. Subjective A/B preference tests show that the new hybrid system outperforms our earlier trajectory-tiling hybrid baseline TTS (67% vs 11%) and the state-of-the-art, real-time TTS system constructed with Tacotron 2 and LPC-Net (56% vs 27%). Fenglong Xie, Wen-Chao Su, Frank K. Soong |
ICASSP | 1 |
| 2021 | Triple M: A Practical Text-to-Speech Synthesis System with Multi-Guidance Attention and Multi-Band Multi-Time LPCNetabstractIn this work, a robust and efficient text-to-speech (TTS) synthesis system named Triple M is proposed for large-scale online application.The key components of Triple M are: 1) A sequence-to-sequence model adopts a novel multi-guidance attention to transfer complementary advantages from guiding attention mechanisms to the basic attention mechanism without in-domain performance loss and online service modification.Compared with single attention mechanism, multi-guidance attention not only brings better naturalness to long sentence synthesis, but also reduces the word error rate by 26.8%.2) A new efficient multi-band multi-time vocoder framework, which reduces the computational complexity from 2.8 to 1.0 GFLOP and speeds up LPCNet by 2.75x on a single CPU. Shilun Lin, Fenglong Xie |
Interspeech | 2 |
| 2020 | An Improved Frame-Unit-Selection Based Voice Conversion System Without Parallel Training DataabstractA frame-unit-selection based voice conversion system proposed earlier by us is revisited here to enhance its performance in both speech naturalness and speaker similarity. Speaker independent, bilingual (Mandarin Chinese and American English) deep neural net (DNN) acoustic model’s output, frame-level phone posterior probability (PPP), is used to represent the phonetic information. The corresponding frame-level F0 is used as the prosodic information. Kullback-Leibler divergence (KLD) between source and target PPPs (phonetic distortion) and the absolute difference between normalized source and target F0 (prosodic distortion) are used for selecting target frame candidates to construct a search lattice. The optimal target unit trajectory is obtained by Viterbi algorithm which tries to minimize the dynamic acoustic difference between the acoustic trajectory of the source speech and target candidates. The obtained spectral trajectory together with the enhanced pitch period and pitch correlation trajectory are sent to LPCNet vocoder to synthesize the converted waveforms. Compared with the top-rank system in Voice Conversion Challenge 2018, our new system can achieve on-par performance on studio to studio American English VC test, and better performance on non-studio to studio Mandarin Chinese VC test, in both speech naturalness MOS and speaker similarity DMOS. Fenglong Xie, Yibin Zheng, Frank K. Soong |
ICASSP | 1 |
| 2020 | Improving End-to-End Speech Synthesis with Local Recurrent Neural Network Enhanced TransformerabstractAlthough Transformer based neural end-to-end TTS model has demonstrated extreme effectiveness in capturing long-term dependencies and achieved state-of-the-art performance, it still suffers from two problems. 1) limited ability to model sequential and local structures in sequences; 2) heavily rely on position embeddings that have limited effect but require an amount of design efforts. In this paper, we introduce local recurrent neural network (Local-RNN) into Transformer to make full use of the advantages of both RNN and Transformer while mitigating their drawbacks. The sequential and local structures could be effectively modeled by Local-RNN, while the long-term dependencies could be captured by Transformer without any use of position embeddings. Subjective evaluation results show our proposed model outperforms baseline (Transformer) with a gap of 0.12 in MOS and achieves close to human quality (4.34 vs. 4.45 in MOS) on general test. Case level intelligibility test also show an absolute improvement of 6.5% in case level intelligibility rate over the baseline on a challenging test. Yibin Zheng, Fenglong Xie |
ICASSP | 3 |
| 2019 | Voice conversion with SI-DNN and KL divergence based mapping without parallel training data
Fenglong Xie, Frank K. Soong, Haifeng Li 0001 |
Speech Commun. | 1 |
| 2016 | A KL divergence and DNN approach to cross-lingual TTSabstractWe propose a Kullback-Leibler divergence (KLD) and deep neural net (DNN) based approach to cross-lingual TTS (CL-TTS) training. A speaker independent DNN (SI-DNN) ASR is used to equalize the speaker difference between a source speaker in L1 and a reference speaker in L2. Two speaker dependent GMM-HMM parametric TTS systems are first trained in the respective languages. The senones sets of the two TTS are matched in the SI-DNN ASR in terms of their output posteriors distributions in KLD. The minimum KLD criterion is used to transform the senones in the source speaker's TTS (L1) to the corresponding "closest" senones in the target language (L2). The new CL-TTS thus trained has been shown to achieve high speaker similarity to the source speaker in L1 while high intelligibility and naturalness are preserved. For untranscribed source speaker's recordings, say, conversational speech, a frame mapping, instead of "senone mapping" is also proposed to achieve a high but slightly inferior CL-TTS. Fenglong Xie, Frank K. Soong, Haifeng Li 0001 |
ICASSP | 1 |
| 2016 | A KL Divergence and DNN-Based Approach to Voice Conversion without Parallel Training Sentences
Fenglong Xie, Frank K. Soong, Haifeng Li 0001 |
INTERSPEECH | 1 |
| 2014 | TTS synthesis with bidirectional LSTM based recurrent neural networksabstractFeed-forward, Deep neural networks (DNN)-based text-tospeech (TTS) systems have been recently shown to outperform decision-tree clustered context-dependent HMM TTS systems [1, 4]. However, the long time span contextual effect in a speech utterance is still not easy to accommodate, due to the intrinsic, feed-forward nature in DNN-based modeling. Also, to synthesize a smooth speech trajectory, the dynamic features are commonly used to constrain speech parameter trajectory generation in HMM-based TTS [2]. In this paper, Recurrent Neural Networks (RNNs) with Bidirectional Long Short Term Memory (BLSTM) cells are adopted to capture the correlation or co-occurrence information between any two instants in a speech utterance for parametric TTS synthesis. Experimental results show that a hybrid system of DNN and BLSTM-RNN, i.e., lower hidden layers with a feed-forward structure which is cascaded with upper hidden layers with a bidirectional RNN structure of LSTM, can outperform either the conventional, decision tree-based HMM, or a DNN TTS system, both objectively and subjectively. The speech trajectory generated by the BLSTM-RNN TTS is fairly smooth and no dynamic constraints are needed. Yuchen Fan 0001, Yao Qian, Fenglong Xie, Frank K. Soong |
INTERSPEECH | 3 |
| 2014 | Sequence error (SE) minimization training of neural network for voice conversionabstractNeural network (NN) based voice conversion, which employs a nonlinear function to map the features from a source to a target speaker, has been shown to outperform GMM-based voice conversion approach [4-7]. However, there are still limitations to be overcome in NN-based voice conversion, e.g. NN is trained on a Frame Error (FE) minimization criterion and the corresponding weights are adjusted to minimize the error squares over the whole source-target, stereo training data set. In this paper, we use the idea of sentence optimization based, minimum generation error (MGE) training in HMM-based TTS synthesis, and modify the FE minimization to Sequence Error (SE) minimization in NN training for voice conversion. The conversion error over a training sentence from a source speaker to a target speaker is minimized via a gradient descent-based, back propagation (BP) procedure. Experimental results show that the speech converted by the NN, which is first trained with frame error minimization and then refined with sequence error minimization, sounds subjectively better than the converted speech by NN trained with frame error minimization only. Scores on both naturalness and similarity to the target speaker are improved. Index Terms: voice conversion, neural network, pre-training, sequence error minimization Fenglong Xie, Yao Qian, Yuchen Fan 0001, Frank K. Soong, Haifeng Li 0001 |
INTERSPEECH | 1 |