VLDB 2026 Research / reviewers in the wild / expert
Wenrui Liu 0003
dblp:156/8975-3
· DBLP profile ↗
11ranked-venue papers
3as first author
11since 2021 · last 2026
0009-0000-5940-5369ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Say More with Less: Variable-Frame-Rate Speech Tokenization via Adaptive Clustering and Implicit Duration CodingabstractExisting speech tokenizers typically assign a fixed number of tokens per second, regardless of the varying information density or temporal fluctuations in the speech signal. This uniform token allocation mismatches the intrinsic structure of speech, where information is distributed unevenly over time. To address this, we propose VARSTok, a VAriable-frame-Rate Speech Tokenizer that adapts token allocation based on local feature similarity. VARSTok introduces two key innovations: (1) a temporal-aware density peak clustering algorithm that adaptively segments speech into variable-length units, and (2) a novel implicit duration coding scheme that embeds both content and temporal span into a single token index, eliminating the need for auxiliary duration predictors. Extensive experiments show that VARSTok significantly outperforms strong fixed-rate baselines. Notably, it achieves superior reconstruction naturalness while using up to 23% fewer tokens than a 40 Hz fixed-frame-rate baseline. VARSTok further yields lower word error rates and improved naturalness in zero-shot text-to-speech synthesis. To the best of our knowledge, this is the first work to demonstrate that a fully dynamic, variable-frame-rate acoustic speech tokenizer can be seamlessly integrated into downstream speech language models. Rui-Chen Zheng, Wenrui Liu 0003, Hui-Peng Du, Chong Deng, Qian Chen 0003, Wen Wang 0001, Yang Ai, Zhen-Hua Ling |
AAAI | 2 |
| 2025 | CART: A Generative Cross-Modal Retrieval Framework With Coarse-To-Fine Semantic ModelingabstractMinghui Fang, Shengpeng Ji, Jialong Zuo, Hai Huang, Yan Xia, Jieming Zhu, Xize Cheng, Xiaoda Yang, Wenrui Liu, Gang Wang, Zhenhua Dong, Zhou Zhao. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Minghui Fang 0002, Shengpeng Ji, Jialong Zuo, Hai Huang 0013, Yan Xia 0006, Jieming Zhu, Xize Cheng, Xiaoda Yang, Wenrui Liu 0003, Zhenhua Dong, Zhou Zhao 0001 |
ACL (1) | 9 |
| 2025 | Analyzing and Mitigating Inconsistency in Discrete Speech Tokens for Neural Codec Language ModelsabstractBuilding upon advancements in Large Language Models (LLMs), the field of audio processing has seen increased interest in training speech generation tasks with discrete speech token sequences.However, directly discretizing speech by neural audio codecs often results in sequences that fundamentally differ from text sequences.Unlike text, where text token sequences are deterministic, discrete speech tokens can exhibit significant variability based on contextual factors, while still producing perceptually identical audio segments.We refer to this phenomenon as Discrete Representation Inconsistency (DRI).This inconsistency can lead to a single speech segment being represented by multiple divergent sequences, which creates confusion in neural codec language models and results in poor generated speech.In this paper, we quantitatively analyze the DRI phenomenon within popular audio tokenizers such as En-Codec.Our approach effectively mitigates the DRI phenomenon of the neural audio codec.Furthermore, extensive experiments on the neural codec language model over LibriTTS and large-scale MLS dataset (44,000 hours) demonstrate the effectiveness and generality of our method.The demo of audio samples is available online 1 . Wenrui Liu 0003, Zhifang Guo, Jin Xu 0010, Yuanjun Lv, Yunfei Chu, Junyang Lin |
ACL (1) | 1 |
| 2025 | VoxpopuliTTS: a large-scale multilingual TTS corpus for zero-shot speech generationabstractIn recent years, speech generation fields have achieved significant advancements, primarily due to improvements in large TTS (text-to-speech) systems and scalable TTS datasets. However, there is still a lack of large-scale multilingual TTS datasets, which limits the development of cross-language and multilingual TTS systems. Hence, we refine Voxpopuli dataset and propose VoxpopuliTTS dataset. This dataset comprises 30,000 hours of high-quality speech data, across 3 languages with multiple speakers and styles, suitable for various speech tasks such as TTS and ASR. To enhance the quality of speech data from Voxpopuli, we improve the existing processing pipeline by: 1) filtering out low-quality speech-text pairs based on ASR confidence scores, and 2) concatenating short transcripts by checking semantic information completeness to generate the long transcript. Experimental results demonstrate the effectiveness of the VoxpopuliTTS dataset and the proposed processing pipeline. Wenrui Liu 0003, Jionghao Bai, Xize Cheng, Jialong Zuo, Ziyue Jiang 0001, Shengpeng Ji, Minghui Fang 0002, Xiaoda Yang, Qian Yang 0006, Zhou Zhao 0001 |
COLING | 1 |
| 2025 | Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching ModelabstractThis paper introduces PFlow-VC, a conditional flow matching voice conversion model that leverages fine-grained discrete pitch tokens and target speaker prompt information for expressive voice conversion (VC). Previous VC works primarily focus on speaker conversion, with further exploration needed in enhancing expressiveness (such as prosody and emotion) for timbre conversion. Unlike previous methods, we adopt a simple and efficient approach to enhance the style expressiveness of voice conversion models. Specifically, we pretrain a self-supervised pitch VQVAE model to discretize speaker-irrelevant pitch information and leverage a masked pitch-conditioned flow matching model for Mel-spectrogram synthesis, which provides in-context pitch modeling capabilities for the speaker conversion model, effectively improving the voice style transfer capacity. Additionally, we improve timbre similarity by combining global timbre embeddings with time-varying timbre tokens. Experiments on unseen LibriTTS test-clean and emotional speech dataset ESD show the superiority of the PFlow-VC model in both timbre conversion and style transfer. Audio samples are available on the demo page https://speechai-demo.github.io/PFlow-VC/. Jialong Zuo, Shengpeng Ji, Minghui Fang 0002, Ziyue Jiang 0001, Xize Cheng, Qian Yang 0006, Wenrui Liu 0003, Guangyan Zhang, Zehai Tu, Yiwen Guo, Zhou Zhao 0001 |
ICASSP | 7 |
| 2025 | GTA: Towards Generative Text-To-Audio Retrieval via Multi-Scale Tokenizer
Minghui Fang 0002, Shengpeng Ji, Jialong Zuo, Xize Cheng, Wenrui Liu 0003, Xiaoda Yang, Ruofan Hu 0002, Jieming Zhu, Zhou Zhao 0001 |
INTERSPEECH | 5 |
| 2025 | MelRe: Vision-Based Mel-Spectrogram Restoration
Kaixuan Luan, Xiaoda Yang, Shile Cai, Ruofan Hu 0002, Minghui Fang 0002, Wenrui Liu 0003, Jialong Zuo, Jiaqi Duan |
INTERSPEECH | 6 |
| 2025 | Speech Token Prediction via Compressed-to-fine Language Modeling for Speech GenerationabstractNeural audio codecs, used as speech tokenizers, have demonstrated remarkable potential in the field of speech generation. However, to ensure high-fidelity audio reconstruction, neural audio codecs typically encode audio into long sequences of speech tokens, posing a significant challenge for downstream language models in long-context modeling. We observe that speech token sequences exhibit short-range dependency: due to the monotonic alignment between text and speech in text-to-speech (TTS) tasks, the prediction of the current token primarily relies on its local context, while long-range tokens contribute less to the current token prediction and often contain redundant information. Inspired by this observation, we propose a compressed-to-fine language modeling approach to address the challenge of long sequence speech tokens within neural codec language models: (1) Fine-grained Initial and Short-range Information: Our approach retains the prompt and local tokens during prediction to ensure text alignment and the integrity of paralinguistic information; (2) Compressed Long-range Context: Our approach compresses long-range token spans into compact representations to reduce redundant information while preserving essential semantics. Extensive experiments on various neural audio codecs and downstream language models validate the effectiveness and generalizability of the proposed approach, highlighting the importance of token compression in improving speech generation within neural codec language models. The demo of audio samples will be available at https://anonymous.4open.science/r/SpeechTokenPredictionViaCompressedToFinedLM. Wenrui Liu 0003, Qian Chen 0003, Wen Wang 0019, Guanrou Yang, Minghui Fang 0002, Jialong Zuo, Xiaoda Yang, Tao Jin 0004, Jin Xu 0010, Yafeng Chen, Jionghao Bai, Zhifang Guo |
ACM Multimedia | 1 |
| 2025 | EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text PromptingabstractHuman speech goes beyond the mere transfer of information; it is a profound exchange of emotions and a connection between individuals. While Text-to-Speech (TTS) models have made huge progress, they still face challenges in controlling the emotional expression in the generated speech. In this work, we propose EmoVoice, a novel emotion-controllable TTS model that exploits large language models (LLMs) to enable fine-grained freestyle natural language emotion control, and a phoneme boost variant design that makes the model output phoneme tokens and audio tokens in parallel to enhance content consistency, inspired by chain-of-thought (CoT) and modality-of-thought (CoM) techniques. Besides, we introduce EmoVoice-DB, a high-quality 40-hour English emotion dataset featuring expressive speech and fine-grained emotion labels with natural language descriptions. EmoVoice achieves state-of-the-art performance on the English EmoVoice-DB test set using only synthetic training data, and on the Chinese Secap test set using our in-house data. We further investigate the reliability of existing emotion evaluation metrics and their alignment with human perceptual preferences, and explore using SOTA multimodal LLMs GPT-4o-audio and Gemini to assess emotional speech. Dataset, code, checkpoints and demo samples are available at https://github.com/yanghaha0908/EmoVoice. Guanrou Yang, Qian Chen 0003, Ziyang Ma 0001, Wen Wang 0019, Tianrui Wang, Yifan Yang 0005, Zhikang Niu, Wenrui Liu 0003, Fan Yu 0002, Zhihao Du, Zhifu Gao, Shiliang Zhang, Xie Chen 0001 |
ACM Multimedia | 10 |
| 2024 | AIR-Bench: Benchmarking Large Audio-Language Models via Generative ComprehensionabstractQian Yang, Jin Xu, Wenrui Liu, Yunfei Chu, Ziyue Jiang, Xiaohuan Zhou, Yichong Leng, Yuanjun Lv, Zhou Zhao, Chang Zhou, Jingren Zhou. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Qian Yang 0006, Jin Xu 0010, Wenrui Liu 0003, Yunfei Chu, Ziyue Jiang 0001, Xiaohuan Zhou, Yichong Leng, Yuanjun Lv, Zhou Zhao 0001, Chang Zhou 0005, Jingren Zhou 0001 |
ACL (1) | 3 |
| 2024 | Prompt-Singer: Controllable Singing-Voice-Synthesis with Natural Language PromptabstractYongqi Wang, Ruofan Hu, Rongjie Huang, Zhiqing Hong, Ruiqi Li, Wenrui Liu, Fuming You, Tao Jin, Zhou Zhao. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Ruofan Hu 0002, Rongjie Huang 0001, Zhiqing Hong, Ruiqi Li 0002, Wenrui Liu 0003, Fuming You, Tao Jin 0004, Zhou Zhao 0001 |
NAACL-HLT | 6 |