EDBT 2026 Demo / reviewers in the wild / expert
Chenfeng Miao
dblp:270/4712
· DBLP profile ↗
13ranked-venue papers
6as first author
12since 2021 · last 2025
0009-0006-1986-3035ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 8 · 5 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LEF-TTS: Lightweight and Efficient End-to-End Text-to-Speech Synthesis With Multi-Stream GeneratorabstractRecently, the field of Text-to-speech synthesis has been predominantly characterized by end-to-end models, with the quality of speech generated by these models becoming increasingly comparable to that of human speech. In this work, we propose a Lightweight and Efficient Text-to-speech model, a fast end-to-end framework based on EfficientTTS 2 with fully differentiable. We utilize Fast Linear Attention with a Single Head instead of the standard stacked Transformer, which decreases computational complexity and reduces parameters. Additionally, we improve a network architecture ConvWaveNet to further decrease model parameters, and accelerate inference through a multi-stream inverse short-time Fourier Transform generator. These improvements significantly reduce model parameters and increase inference speed, thereby achieving the objectives of faster inference and lightweight modeling. Experimental results show that the proposed model achieves speech quality comparable to that of the baseline models, while also offering improved inference speed and reduced model size. Minchuan Chen, Chenfeng Miao, Ning Cheng 0001, Jing Xiao 0006 |
ICASSP | 4 |
| 2025 | EffectiveASR: A Single-Step Non-Autoregressive Mandarin Speech Recognition Architecture with High Accuracy and Inference SpeedabstractNon-autoregressive (NAR) automatic speech recognition (ASR) models predict tokens independently and simultaneously, bringing high inference speed. However, there is still a gap in the accuracy of the NAR models compared to the autoregressive (AR) models. In this paper, we propose a single-step NAR ASR architecture with high accuracy and inference speed, called EffectiveASR. It uses an Index Mapping Vector (IMV) based alignment generator to generate alignments during training, and an alignment predictor to learn the alignments for inference. It can be trained end-to-end (E2E) with cross-entropy loss combined with alignment loss. The proposed EffectiveASR achieves competitive results on the AISHELL-1 and AISHELL-2 Mandarin benchmarks compared to the leading models. Specifically, it achieves character error rates (CER) of 4.26%/4.62% on the AISHELL-1 dev/test dataset, which outperforms the AR Conformer with about 30x inference speedup. Ziyang Zhuang, Chenfeng Miao, Tao Wei 0003, Ning Cheng 0001, Jing Xiao 0006 |
ICASSP | 2 |
| 2024 | Improving Attention-Based End-to-End Speech Recognition by Monotonic Alignment Attention Matrix ReconstructionabstractIn automatic speech recognition (ASR) task, the output sequence should correspond to a linear transcription of the input sequence. Lots of works have been done to learn the monotonic alignment in end-to-end (E2E) ASR model, but their methods mainly focus on streaming propose and usually result in a decline in ASR performance. On the contrary, some studies have shown that for non-streaming attention-based models, monotonic alignment is beneficial to model performance. Based on this motivation, we propose the enhanced Gaussian Monotonic Alignment (e-GMA), which reduces the difficulty of learning monotonic alignment, and the reconstructed attention matrix leads to an improved accuracy in ASR tasks. Experiments on the LibriSpeech dataset demonstrate the effectiveness of the proposed approach. Comparing with a strong baseline obtained from WeNet, the proposed model yields 12.2% relative WER reduction on test-clean benchmark and 9.9% on test-other. Ziyang Zhuang, Chenfeng Miao, Tao Wei 0003, Jing Xiao 0006 |
ICASSP | 3 |
| 2024 | DFlow: A Generative Model Combining Denoising AutoEncoder and Normalizing Flow for High Fidelity Waveform GenerationabstractIn this work, we present DFlow, a novel generative framework that combines Normalizing Flow (NF) with a Denoising AutoEncoder (DAE), for high-fidelity waveform generation. With a tactfully designed structure, DFlow seamlessly integrates the capabilities of both NF and DAE, resulting in a significantly improved performance compared to the standard NF models. Experimental results showcase DFlow’s superiority, achieving the highest MOS score among the existing methods on commonly used datasets and the fastest synthesis speed among all likelihood models. We further demonstrate the generalization ability of DFlow by generating high-quality out-of-distribution audio samples, such as singing and music audio. Additionally, we extend the model capacity of DFlow by scaling up both the model size and training set size. Our large-scale universal vocoder, DFlow-XL, achieves highly competitive performance against the best universal vocoder, BigVGAN. Chenfeng Miao, Qingying Zhu, Minchuan Chen, Jing Xiao 0006 |
ICML | 1 |
| 2024 | E-Paraformer: A Faster and Better Parallel Transformer for Non-autoregressive End-to-End Mandarin Speech Recognition
Fengyun Tan, Ziyang Zhuang, Chenfeng Miao, Tao Wei 0003, Shaodan Zhai, Jing Xiao 0006 |
INTERSPEECH | 4 |
| 2024 | EfficientTTS 2: Variational End-to-End Text-to-Speech Synthesis and Voice ConversionabstractRecently, the field of Text-to-Speech (TTS) has been dominated by one-stage text-to-waveform models which have significantly improved speech quality compared to two-stage models. In this work, we propose EfficientTTS 2 (EFTS2), a one-stage high-quality end-to-end TTS framework that is fully differentiable and highly efficient. Our method adopts an adversarial training process, with a differentiable aligner and a hierarchical-VAE-based waveform generator. These design choices free the model from the use of external aligners, invertible structures, and complex training procedures as most previous TTS works have. Moreover, we extend EFTS2 to the voice conversion (VC) task and propose EFTS2-VC, an end-to-end VC model that allows high-quality speech-to-speech conversion. Experimental results suggest that the two proposed models achieve better or at least comparable speech quality compared to baseline models, while also providing faster inference speeds and smaller model sizes. Chenfeng Miao, Qingying Zhu, Minchuan Chen, Jun Ma 0018, Jing Xiao 0006 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2023 | Exploring multi-task learning and data augmentation in dementia detection with self-supervised pretrained models
Minchuan Chen, Chenfeng Miao, Jun Ma 0018, Jing Xiao 0006 |
INTERSPEECH | 2 |
| 2022 | A compact transformer-based GAN vocoder
Chenfeng Miao, Minchuan Chen, Jun Ma 0018, Jing Xiao 0006 |
INTERSPEECH | 1 |
| 2022 | Towards Efficiently Learning Monotonic Alignments for Attention-based End-to-End Speech Recognition
Chenfeng Miao, Ziyang Zhuang, Tao Wei 0003, Jun Ma 0018, Jing Xiao 0006 |
INTERSPEECH | 1 |
| 2021 | Unsupervised Learning for Multi-Style Speech Synthesis with Limited DataabstractExisting multi-style speech synthesis methods require either style labels or large amounts of unlabeled training data, making data acquisition difficult. In this paper, we present an unsupervised multi-style speech synthesis method that can be trained with limited data. We leverage instance discriminator to guide a style encoder to learn meaningful style representations from a multi-style dataset. Furthermore, we employ information bottleneck to filter out style-irrelevant information in the representations, which can improve speech quality and style similarity. Our method is able to produce desirable speech using a fairly small dataset, where the baseline GST-Tacotron fails. ABX tests show that our model significantly outperforms GST-Tacotron in both emotional speech synthesis task and multi-speaker speech synthesis task. In addition, we demonstrate that our method is able to learn meaningful style features with only 50 training samples per style. Chenfeng Miao, Minchuan Chen, Jun Ma 0018, Jing Xiao 0006 |
ICASSP | 2 |
| 2021 | EfficientTTS: An Efficient and High-Quality Text-to-Speech ArchitectureabstractIn this work, we address the Text-to-Speech (TTS) task by proposing a non-autoregressive architecture called EfficientTTS. Unlike the dominant non-autoregressive TTS models, which are trained with the need of external aligners, EfficientTTS optimizes all its parameters with a stable, end-to-end training procedure, allowing for synthesizing high quality speech in a fast and efficient manner. EfficientTTS is motivated by a new monotonic alignment modeling approach, which specifies monotonic constraints to the sequence alignment with almost no increase of computation. By combining EfficientTTS with different feed-forward network structures, we develop a family of TTS models, including both text-to-melspectrogram and text-to-waveform networks. We experimentally show that the proposed models significantly outperform counterpart models such as Tacotron 2 and Glow-TTS in terms of speech quality, training efficiency and synthesis speed, while still producing the speeches of strong robustness and great diversity. In addition, we demonstrate that proposed approach can be easily extended to autoregressive models such as Tacotron 2. Chenfeng Miao, Zhengchen Liu, Minchuan Chen, Jun Ma 0018, Jing Xiao 0006 |
ICML | 1 |
| 2021 | EfficientSing: A Chinese Singing Voice Synthesis System Using Duration-Free Acoustic Model and HiFi-GAN Vocoder
Zhengchen Liu, Chenfeng Miao, Qingying Zhu, Minchuan Chen, Jun Ma 0018, Jing Xiao 0006 |
Interspeech | 2 |
| 2020 | Flow-TTS: A Non-Autoregressive Network for Text to Speech Based on FlowabstractIn this work, we propose Flow-TTS, a non-autoregressive end-to-end neural TTS model based on generative flow. Unlike other non-autoregressive models, Flow-TTS can achieve high-quality speech generation by using a single feed-forward network. To our knowledge, Flow-TTS is the first TTS model utilizing flow in spectrogram generation network and the first non-autoregssive model which jointly learns the alignment and spectrogram generation through a single network. Experiments on LJSpeech show that the speech quality of Flow-TTS heavily approaches that of human and is even better than that of autoregressive model Tacotron 2 (outperforms Tacotron 2 with a gap of 0.09 in MOS). Meanwhile, the inference speed of Flow-TTS is about 23 times speed-up over Tacotron 2, which is comparable to FastSpeech.1 Chenfeng Miao, Minchuan Chen, Jun Ma 0018, Jing Xiao 0006 |
ICASSP | 1 |