VLDB 2026 Research / reviewers in the wild / expert
Ruiqi Li 0002
dblp:157/8924-2
· DBLP profile ↗
11ranked-venue papers
1as first author
11since 2021 · last 2025
0009-0005-3798-6568ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 1 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TechSinger: Technique Controllable Multilingual Singing Voice Synthesis via Flow MatchingabstractSinging voice synthesis has made remarkable progress in generating natural and high-quality voices. However, existing methods rarely provide precise control over vocal techniques such as intensity, mixed voice, falsetto, bubble, and breathy tones, thus limiting the expressive potential of synthetic voices. We introduce TechSinger, an advanced system for controllable singing voice synthesis that supports five languages and seven vocal techniques. TechSinger leverages a flow-matching-based generative model to produce singing voices with enhanced expressive control over various techniques. To enhance the diversity of training data, we develop a technique detection model that automatically annotates datasets with phoneme-level technique labels. Additionally, our prompt-based technique prediction model enables users to specify desired vocal attributes through natural language, offering fine-grained control over the synthesized singing. Experimental results demonstrate that TechSinger significantly enhances the expressiveness and realism of synthetic singing voices, outperforming existing methods in terms of audio quality and technique-specific control. Wenxiang Guo, Yu Zhang 0126, Changhao Pan, Rongjie Huang 0001, Ruiqi Li 0002, Zhiqing Hong, Zhou Zhao 0001 |
AAAI | 6 |
| 2025 | WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language ModelingabstractLanguage models have been effectively applied to modeling natural signals, such as images, video, speech, and audio. A crucial component of these models is the codec tokenizer, which compresses high-dimensional natural signals into lower-dimensional discrete tokens. In this paper, we introduce WavTokenizer, which offers several advantages over previous SOTA acoustic codec models in the audio domain: 1) extreme compression. By compressing the layers of quantizers and the temporal dimension of the discrete codec, one-second audio of 24kHz sampling rate requires only a single quantizer with 40 or 75 tokens. 2) improved subjective quality. Despite the reduced number of tokens, WavTokenizer achieves state-of-the-art reconstruction quality with outstanding UTMOS scores and inherently contains richer semantic information. Specifically, we achieve these results by designing a broader VQ space, extended contextual windows, and improved attention networks, as well as introducing a powerful multi-scale discriminator and an inverse Fourier transform structure. We conducted extensive reconstruction experiments in the domains of speech, audio, and music. WavTokenizer exhibited strong performance across various objective and subjective metrics compared to state-of-the-art models. We also tested semantic information, VQ utilization, and adaptability to generative models. Comprehensive ablation studies confirm the necessity of each module in WavTokenizer. The code is available at https://github.com/jishengpeng/WavTokenizer. Shengpeng Ji, Ziyue Jiang 0001, Wen Wang 0001, Minghui Fang 0002, Jialong Zuo, Qian Yang 0006, Xize Cheng, Zehan Wang 0001, Ruiqi Li 0002, Xiaoda Yang, Rongjie Huang 0001, Yidi Jiang, Qian Chen 0003, Zhou Zhao 0001 |
ICLR | 10 |
| 2024 | StyleSinger: Style Transfer for Out-of-Domain Singing Voice SynthesisabstractStyle transfer for out-of-domain (OOD) singing voice synthesis (SVS) focuses on generating high-quality singing voices with unseen styles (such as timbre, emotion, pronunciation, and articulation skills) derived from reference singing voice samples. However, the endeavor to model the intricate nuances of singing voice styles is an arduous task, as singing voices possess a remarkable degree of expressiveness. Moreover, existing SVS methods encounter a decline in the quality of synthesized singing voices in OOD scenarios, as they rest upon the assumption that the target vocal attributes are discernible during the training phase. To overcome these challenges, we propose StyleSinger, the first singing voice synthesis model for zero-shot style transfer of out-of-domain reference singing voice samples. StyleSinger incorporates two critical approaches for enhanced effectiveness: 1) the Residual Style Adaptor (RSA) which employs a residual quantization module to capture diverse style characteristics in singing voices, and 2) the Uncertainty Modeling Layer Normalization (UMLN) to perturb the style attributes within the content representation during the training phase and thus improve the model generalization. Our extensive evaluations in zero-shot style transfer undeniably establish that StyleSinger outperforms baseline models in both audio quality and similarity to the reference singing voice samples. Access to singing voice samples can be found at https://stylesinger.github.io/. Yu Zhang 0126, Rongjie Huang 0001, Ruiqi Li 0002, Jinzheng He, Yan Xia 0006, Feiyang Chen 0001, Xinyu Duan, Baoxing Huai, Zhou Zhao 0001 |
AAAI | 3 |
| 2024 | Text-to-Song: Towards Controllable Music Generation Incorporating Vocal and AccompanimentabstractZhiqing Hong, Rongjie Huang, Xize Cheng, Yongqi Wang, Ruiqi Li, Fuming You, Zhou Zhao, Zhimeng Zhang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zhiqing Hong, Rongjie Huang 0001, Xize Cheng, Ruiqi Li 0002, Fuming You, Zhou Zhao 0001 |
ACL (1) | 5 |
| 2024 | Robust Singing Voice Transcription Serves SynthesisabstractNote-level Automatic Singing Voice Transcription (AST) converts singing recordings into note sequences, facilitating the automatic annotation of singing datasets for Singing Voice Synthesis (SVS) applications.Current AST methods, however, struggle with accuracy and robustness when used for practical annotation.This paper presents ROSVOT, the first robust AST model that serves SVS, incorporating a multi-scale framework that effectively captures coarse-grained note information and ensures fine-grained frame-level segmentation, coupled with an attention-based pitch decoder for reliable pitch prediction.We also established a comprehensive annotation-and-training pipeline for SVS to test the model in realworld settings.Experimental findings reveal that ROSVOT achieves state-of-the-art transcription accuracy with either clean or noisy inputs.Moreover, when trained on enlarged, automatically annotated datasets, the SVS model outperforms its baseline, affirming the capability for practical application.Audio samples are available at https://rosvot.github.io.Codes can be found at https://github.com/RickyL- 2000/ROSVOT. Ruiqi Li 0002, Yu Zhang 0126, Zhiqing Hong, Rongjie Huang 0001, Zhou Zhao 0001 |
ACL (1) | 1 |
| 2024 | TCSinger: Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style ControlabstractZero-shot singing voice synthesis (SVS) with style transfer and style control aims to generate high-quality singing voices with unseen timbres and styles (including singing method, emotion, rhythm, technique, and pronunciation) from audio and text prompts.However, the multifaceted nature of singing styles poses a significant challenge for effective modeling, transfer, and control.Furthermore, current SVS models often fail to generate singing voices rich in stylistic nuances for unseen singers.To address these challenges, we introduce TCSinger, the first zero-shot SVS model for style transfer across cross-lingual speech and singing styles, along with multi-level style control.Specifically, TCSinger proposes three primary modules: 1) the clustering style encoder employs a clustering vector quantization model to stably condense style information into a compact latent space; 2) the Style and Duration Language Model (S&D-LM) concurrently predicts style information and phoneme duration, which benefits both; 3) the style adaptive decoder uses a novel mel-style adaptive normalization method to generate singing voices with enhanced details.Experimental results show that TCSinger outperforms all baseline models in synthesis quality, singer similarity, and style controllability across various tasks, including zero-shot style transfer, multi-level style control, crosslingual style transfer, and speech-to-singing style transfer.Singing voice samples can be accessed at https://tcsinger.github.io/. Yu Zhang 0126, Ziyue Jiang 0001, Ruiqi Li 0002, Changhao Pan, Jinzheng He, Rongjie Huang 0001, Chuxin Wang, Zhou Zhao 0001 |
EMNLP | 3 |
| 2024 | Prompt-Singer: Controllable Singing-Voice-Synthesis with Natural Language PromptabstractYongqi Wang, Ruofan Hu, Rongjie Huang, Zhiqing Hong, Ruiqi Li, Wenrui Liu, Fuming You, Tao Jin, Zhou Zhao. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Ruofan Hu 0002, Rongjie Huang 0001, Zhiqing Hong, Ruiqi Li 0002, Wenrui Liu 0003, Fuming You, Tao Jin 0004, Zhou Zhao 0001 |
NAACL-HLT | 5 |
| 2024 | Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow MatchingabstractVideo-to-audio (V2A) generation aims to synthesize content-matching audio from silent video, and it remains challenging to build V2A models with high generation quality, efficiency, and visual-audio temporal synchrony.
We propose Frieren, a V2A model based on rectified flow matching. Frieren regresses the conditional transport vector field from noise to spectrogram latent with straight paths and conducts sampling by solving ODE, outperforming autoregressive and score-based models in terms of audio quality. By employing a non-autoregressive vector field estimator based on a feed-forward transformer and channel-level cross-modal feature fusion with strong temporal alignment, our model generates audio that is highly synchronized with the input video. Furthermore, through reflow and one-step distillation with guided vector field, our model can generate decent audio in a few, or even only one sampling step. Experiments indicate that Frieren achieves state-of-the-art performance in both generation quality and temporal alignment on VGGSound, with alignment accuracy reaching 97.22\%, and 6.2\% improvement in inception score over the strong diffusion-based baseline. Audio samples and code are available at http://frieren-v2a.github.io. Wenxiang Guo, Rongjie Huang 0001, Jiawei Huang 0008, Zehan Wang 0001, Fuming You, Ruiqi Li 0002, Zhou Zhao 0001 |
NeurIPS | 7 |
| 2024 | GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing TasksabstractThe scarcity of high-quality and multi-task singing datasets significantly hinders the development of diverse controllable and personalized singing tasks, as existing singing datasets suffer from low quality, limited diversity of languages and singers, absence of multi-technique information and realistic music scores, and poor task suitability.To tackle these problems, we present GTSinger, a large Global, multi-Technique, free-to-use, high-quality singing corpus with realistic music scores, designed for all singing tasks, along with its benchmarks.Particularly,(1) we collect 80.59 hours of high-quality singing voices, forming the largest recorded singing dataset;(2) 20 professional singers across nine widely spoken languages offer diverse timbres and styles;(3) we provide controlled comparison and phoneme-level annotations of six commonly used singing techniques, helping technique modeling and control;(4) GTSinger offers realistic music scores, assisting real-world musical composition;(5) singing voices are accompanied by manual phoneme-to-audio alignments, global style labels, and 16.16 hours of paired speech for various singing tasks.Moreover, to facilitate the use of GTSinger, we conduct four benchmark experiments: technique-controllable singing voice synthesis, technique recognition, style transfer, and speech-to-singing conversion. Yu Zhang 0126, Changhao Pan, Wenxiang Guo, Ruiqi Li 0002, Jingyu Lu 0001, Zhiqing Hong, Chuxin Wang, Jinzheng He, Ziyue Jiang 0001, Jiecheng Zhou, Zhou Zhao 0001 |
NeurIPS | 4 |
| 2023 | DisCover: Disentangled Music Representation Learning for Cover Song IdentificationabstractIn the field of music information retrieval (MIR), cover song identification (CSI) is a challenging task that aims to identify cover versions of a query song from a massive collection. Existing works still suffer from high intra-song variances and inter-song correlations, due to the entangled nature of version-specific and version-invariant factors in their modeling. In this work, we set the goal of disentangling version-specific and version-invariant factors, which could make it easier for the model to learn invariant music representations for unseen query songs. We analyze the CSI task in a disentanglement view with the causal graph technique, and identify the intra-version and inter-version effects biasing the invariant learning. To block these effects, we propose the disentangled music representation learning framework (DisCover) for CSI. DisCover consists of two critical components: (1) Knowledge-guided Disentanglement Module (KDM) and (2) Gradient-based Adversarial Disentanglement Module (GADM), which block intra-version and inter-version biased effects, respectively. KDM minimizes the mutual information between the learned representations and version-variant factors that are identified with prior domain knowledge. GADM identifies version-variant factors by simulating the representation transitions between intra-song versions, and exploits adversarial distillation for effect blocking. Extensive comparisons with best-performing methods and in-depth analysis demonstrate the effectiveness of DisCover and the and necessity of disentanglement for CSI. Jiahao Xun, Shengyu Zhang 0001, Yanting Yang, Jieming Zhu, Liqun Deng, Zhou Zhao 0001, Zhenhua Dong, Ruiqi Li 0002, Fei Wu 0001 |
SIGIR | 8 |
| 2022 | M4Singer: A Multi-Style, Multi-Singer and Musical Score Provided Mandarin Singing CorpusabstractThe lack of publicly available high-quality and accurately labeled datasets has long been a major bottleneck for singing voice synthesis (SVS). To tackle this problem, we present M4Singer, a free-to-use Multi-style, Multi-singer Mandarin singing collection with elaborately annotated Musical scores as well as its benchmarks. Specifically, 1) we construct and release a large high-quality Chinese singing voice corpus, which is recorded by 20 professional singers, covering 700 Chinese pop songs as well as all the four SATB types (i.e., soprano, alto, tenor, and bass); 2) we take extensive efforts to manually compose the musical scores for each recorded song, which are necessary to the study of the prosody modeling for SVS. 3) To facilitate the use and demonstrate the quality of M4Singer, we conduct four different benchmark experiments: score-based SVS, controllable singing voice (CSV), singing voice conversion (SVC) and automatic music transcription (AMT). Ruiqi Li 0002, Shoutong Wang, Liqun Deng, Jinglin Liu, Yi Ren 0006, Jinzheng He, Rongjie Huang 0001, Jieming Zhu, Zhou Zhao 0001 |
NeurIPS | 2 |