EDBT 2026 Demo / reviewers in the wild / expert
Jinzheng He
dblp:272/8857
· DBLP profile ↗
17ranked-venue papers
1as first author
16since 2021 · last 2025
0009-0003-3024-9624ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 1 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue ModelsabstractYifu Chen, Shengpeng Ji, Haoxiao Wang, Ziqing Wang, Siyu Chen, Jinzheng He, Jin Xu, Zhou Zhao. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Shengpeng Ji, Haoxiao Wang, Jinzheng He |
ACL (1) | 6 |
| 2024 | StyleSinger: Style Transfer for Out-of-Domain Singing Voice SynthesisabstractStyle transfer for out-of-domain (OOD) singing voice synthesis (SVS) focuses on generating high-quality singing voices with unseen styles (such as timbre, emotion, pronunciation, and articulation skills) derived from reference singing voice samples. However, the endeavor to model the intricate nuances of singing voice styles is an arduous task, as singing voices possess a remarkable degree of expressiveness. Moreover, existing SVS methods encounter a decline in the quality of synthesized singing voices in OOD scenarios, as they rest upon the assumption that the target vocal attributes are discernible during the training phase. To overcome these challenges, we propose StyleSinger, the first singing voice synthesis model for zero-shot style transfer of out-of-domain reference singing voice samples. StyleSinger incorporates two critical approaches for enhanced effectiveness: 1) the Residual Style Adaptor (RSA) which employs a residual quantization module to capture diverse style characteristics in singing voices, and 2) the Uncertainty Modeling Layer Normalization (UMLN) to perturb the style attributes within the content representation during the training phase and thus improve the model generalization. Our extensive evaluations in zero-shot style transfer undeniably establish that StyleSinger outperforms baseline models in both audio quality and similarity to the reference singing voice samples. Access to singing voice samples can be found at https://stylesinger.github.io/. Yu Zhang 0126, Rongjie Huang 0001, Ruiqi Li 0002, Jinzheng He, Yan Xia 0006, Feiyang Chen 0001, Xinyu Duan, Baoxing Huai, Zhou Zhao 0001 |
AAAI | 4 |
| 2024 | TCSinger: Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style ControlabstractZero-shot singing voice synthesis (SVS) with style transfer and style control aims to generate high-quality singing voices with unseen timbres and styles (including singing method, emotion, rhythm, technique, and pronunciation) from audio and text prompts.However, the multifaceted nature of singing styles poses a significant challenge for effective modeling, transfer, and control.Furthermore, current SVS models often fail to generate singing voices rich in stylistic nuances for unseen singers.To address these challenges, we introduce TCSinger, the first zero-shot SVS model for style transfer across cross-lingual speech and singing styles, along with multi-level style control.Specifically, TCSinger proposes three primary modules: 1) the clustering style encoder employs a clustering vector quantization model to stably condense style information into a compact latent space; 2) the Style and Duration Language Model (S&D-LM) concurrently predicts style information and phoneme duration, which benefits both; 3) the style adaptive decoder uses a novel mel-style adaptive normalization method to generate singing voices with enhanced details.Experimental results show that TCSinger outperforms all baseline models in synthesis quality, singer similarity, and style controllability across various tasks, including zero-shot style transfer, multi-level style control, crosslingual style transfer, and speech-to-singing style transfer.Singing voice samples can be accessed at https://tcsinger.github.io/. Yu Zhang 0126, Ziyue Jiang 0001, Ruiqi Li 0002, Changhao Pan, Jinzheng He, Rongjie Huang 0001, Chuxin Wang, Zhou Zhao 0001 |
EMNLP | 5 |
| 2024 | Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech SynthesisabstractZero-shot text-to-speech (TTS) aims to synthesize voices with unseen speech prompts, which significantly reduces the data and computation requirements for voice cloning by skipping the fine-tuning process. However, the prompting mechanisms of zero-shot TTS still face challenges in the following aspects: 1) previous works of zero-shot TTS are typically trained with single-sentence prompts, which significantly restricts their performance when the data is relatively sufficient during the inference stage. 2) The prosodic information in prompts is highly coupled with timbre, making it untransferable to each other.
This paper introduces Mega-TTS 2, a generic prompting mechanism for zero-shot TTS, to tackle the aforementioned challenges. Specifically, we design a powerful acoustic autoencoder that separately encodes the prosody and timbre information into the compressed latent space while providing high-quality reconstructions. Then, we propose a multi-reference timbre encoder and a prosody latent language model (P-LLM) to extract useful information from multi-sentence prompts. We further leverage the probabilities derived from multiple P-LLM outputs to produce transferable and controllable prosody.
Experimental results demonstrate that Mega-TTS 2 could not only synthesize identity-preserving speech with a short prompt of an unseen speaker from arbitrary sources but consistently outperform the fine-tuning method when the volume of data ranges from 10 seconds to 5 minutes. Furthermore, our method enables to transfer various speaking styles to the target timbre in a fine-grained and controlled manner. Audio samples can be found in https://boostprompt.github.io/boostprompt/. Ziyue Jiang 0001, Jinglin Liu, Yi Ren 0006, Jinzheng He, Zhenhui Ye, Shengpeng Ji, Qian Yang 0006, Chen Zhang 0020, Pengfei Wei 0001, Xiang Yin 0006, Zejun Ma 0001, Zhou Zhao 0001 |
ICLR | 4 |
| 2024 | Real3D-Portrait: One-shot Realistic 3D Talking Portrait SynthesisabstractOne-shot 3D talking portrait generation aims to reconstruct a 3D avatar from an unseen image, and then animate it with a reference video or audio to generate a talking portrait video. The existing methods fail to simultaneously achieve the goals of accurate 3D avatar reconstruction and stable talking face animation. Besides, while the existing works mainly focus on synthesizing the head part, it is also vital to generate natural torso and background segments to obtain a realistic talking portrait video. To address these limitations, we present Real3D-Potrait, a framework that (1) improves the one-shot 3D reconstruction power with a large image-to-plane model that distills 3D prior knowledge from a 3D face generative model; (2) facilitates accurate motion-conditioned animation with an efficient motion adapter; (3) synthesizes realistic video with natural torso movement and switchable background using a head-torso-background super-resolution model; and (4) supports one-shot audio-driven talking face generation with a generalizable audio-to-motion model. Extensive experiments show that Real3D-Portrait generalizes well to unseen identities and generates more realistic talking portrait videos compared to previous methods. Video samples are available at https://real3dportrait.github.io. Zhenhui Ye, Tianyun Zhong, Yi Ren 0006, Jiaqi Yang 0008, Weichuang Li, Jiawei Huang 0008, Ziyue Jiang 0001, Jinzheng He, Rongjie Huang 0001, Jinglin Liu, Chen Zhang 0020, Xiang Yin 0006, Zejun Ma 0001, Zhou Zhao 0001 |
ICLR | 8 |
| 2024 | MimicTalk: Mimicking a personalized and expressive 3D talking face in minutesabstractTalking face generation (TFG) aims to animate a target identity's face to create realistic talking videos. Personalized TFG is a variant that emphasizes the perceptual identity similarity of the synthesized result (from the perspective of appearance and talking style). While previous works typically solve this problem by learning an individual neural radiance field (NeRF) for each identity to implicitly store its static and dynamic information, we find it inefficient and non-generalized due to the per-identity-per-training framework and the limited training data. To this end, we propose MimicTalk, the first attempt that exploits the rich knowledge from a NeRF-based person-agnostic generic model for improving the efficiency and robustness of personalized TFG. To be specific, (1) we first come up with a person-agnostic 3D TFG model as the base model and propose to adapt it into a specific identity; (2) we propose a static-dynamic-hybrid adaptation pipeline to help the model learn the personalized static appearance and facial dynamic features; (3) To generate the facial motion of the personalized talking style, we propose an in-context stylized audio-to-motion model that mimics the implicit talking style provided in the reference video without information loss by an explicit style representation. The adaptation process to an unseen identity can be performed in 15 minutes, which is 47 times faster than previous person-dependent methods. Experiments show that our MimicTalk surpasses previous baselines regarding video quality, efficiency, and expressiveness. Video samples are available at https://mimictalk.github.io . Zhenhui Ye, Tianyun Zhong, Yi Ren 0006, Ziyue Jiang 0001, Jiawei Huang 0008, Rongjie Huang 0001, Jinglin Liu, Jinzheng He, Chen Zhang 0020, Zehan Wang 0001, Xize Cheng, Xiang Yin 0006, Zhou Zhao 0001 |
NeurIPS | 8 |
| 2024 | GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing TasksabstractThe scarcity of high-quality and multi-task singing datasets significantly hinders the development of diverse controllable and personalized singing tasks, as existing singing datasets suffer from low quality, limited diversity of languages and singers, absence of multi-technique information and realistic music scores, and poor task suitability.To tackle these problems, we present GTSinger, a large Global, multi-Technique, free-to-use, high-quality singing corpus with realistic music scores, designed for all singing tasks, along with its benchmarks.Particularly,(1) we collect 80.59 hours of high-quality singing voices, forming the largest recorded singing dataset;(2) 20 professional singers across nine widely spoken languages offer diverse timbres and styles;(3) we provide controlled comparison and phoneme-level annotations of six commonly used singing techniques, helping technique modeling and control;(4) GTSinger offers realistic music scores, assisting real-world musical composition;(5) singing voices are accompanied by manual phoneme-to-audio alignments, global style labels, and 16.16 hours of paired speech for various singing tasks.Moreover, to facilitate the use of GTSinger, we conduct four benchmark experiments: technique-controllable singing voice synthesis, technique recognition, style transfer, and speech-to-singing conversion. Yu Zhang 0126, Changhao Pan, Wenxiang Guo, Ruiqi Li 0002, Jingyu Lu 0001, Zhiqing Hong, Chuxin Wang, Jinzheng He, Ziyue Jiang 0001, Jiecheng Zhou, Zhou Zhao 0001 |
NeurIPS | 12 |
| 2023 | AV-TranSpeech: Audio-Visual Robust Speech-to-Speech TranslationabstractRongjie Huang, Huadai Liu, Xize Cheng, Yi Ren, Linjun Li, Zhenhui Ye, Jinzheng He, Lichao Zhang, Jinglin Liu, Xiang Yin, Zhou Zhao. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Rongjie Huang 0001, Huadai Liu, Xize Cheng, Yi Ren 0006, Linjun Li, Zhenhui Ye, Jinzheng He, Jinglin Liu, Xiang Yin 0006, Zhou Zhao 0001 |
ACL (1) | 7 |
| 2023 | CLAPSpeech: Learning Prosody from Text Context with Contrastive Language-Audio Pre-TrainingabstractZhenhui Ye, Rongjie Huang, Yi Ren, Ziyue Jiang, Jinglin Liu, Jinzheng He, Xiang Yin, Zhou Zhao. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Zhenhui Ye, Rongjie Huang 0001, Yi Ren 0006, Ziyue Jiang 0001, Jinglin Liu, Jinzheng He, Xiang Yin 0006, Zhou Zhao 0001 |
ACL (1) | 6 |
| 2023 | ViT-TTS: Visual Text-to-Speech with Scalable Diffusion TransformerabstractText-to-speech(TTS) has undergone remarkable improvements in performance, particularly with the advent of Denoising Diffusion Probabilistic Models (DDPMs).However, the perceived quality of audio depends not solely on its content, pitch, rhythm, and energy, but also on the physical environment.In this work, we propose ViT-TTS, the first visual TTS model with scalable diffusion transformers.ViT-TTS complement the phoneme sequence with the visual information to generate high-perceived audio, opening up new avenues for practical applications of AR and VR to allow a more immersive and realistic audio experience.To mitigate the data scarcity in learning visual acoustic information, we 1) introduce a selfsupervised learning framework to enhance both the visual-text encoder and denoiser decoder; 2) leverage the diffusion transformer scalable in terms of parameters and capacity to learn visual scene information.Experimental results demonstrate that ViT-TTS achieves new stateof-the-art results, outperforming cascaded systems and other baselines regardless of the visibility of the scene.With low-resource data (1h, 2h, 5h), ViT-TTS achieves comparative results with rich-resource baselines.1 2 Huadai Liu, Rongjie Huang 0001, Xuan Lin, Maozong Zheng, Hong Chen 0002, Jinzheng He, Zhou Zhao 0001 |
EMNLP | 7 |
| 2023 | TranSpeech: Speech-to-Speech Translation With Bilateral Perturbation
Rongjie Huang 0001, Jinglin Liu, Huadai Liu, Yi Ren 0006, Jinzheng He, Zhou Zhao 0001 |
ICLR | 6 |
| 2023 | GeneFace: Generalized and High-Fidelity Audio-Driven 3D Talking Face Synthesis
Zhenhui Ye, Ziyue Jiang 0001, Yi Ren 0006, Jinglin Liu, Jinzheng He, Zhou Zhao 0001 |
ICLR | 5 |
| 2023 | UniSinger: Unified End-to-End Singing Voice Synthesis With Cross-Modality Information MatchingabstractThough previous works have shown remarkable achievements in singing voice generation, most existing models focus on one specific application and there is a lack of unified singing voice synthesis models. In addition to low relevance among tasks, different input modalities are one of the most intractable hindrances. Current methods suffer from information confusion and they can not perform precise control. In this work, we propose UniSinger, a unified end-to-end singing voice synthesizer, which integrates three abilities related to singing voice generation: singing voice synthesis (SVS), singing voice conversion (SVC), and singing voice editing (SVE) into a single framework. Specifically, we perform representation disentanglement for controlling different attributes of the singing voice. We further propose a cross-modality information matching method to close the distribution gap between multi-modal inputs and achieve end-to-end training. The experiments conducted on the OpenSinger dataset demonstrate that UniSinger achieves state-of-the-art results in three applications. Further extensive experiments verify the capability of representation disentanglement and information matching, reflecting that UniSinger enjoys great superiority in sample quality, timbre similarity, and multi-task compatibility. Audio samples can be found in https://unisinger.github.io/Samples/. Zhiqing Hong, Chenye Cui, Rongjie Huang 0001, Jinglin Liu, Jinzheng He, Zhou Zhao 0001 |
ACM Multimedia | 6 |
| 2022 | Flow-Based Unconstrained Lip to Speech GenerationabstractUnconstrained lip-to-speech aims to generate corresponding speeches based on silent facial videos with no restriction to head pose or vocabulary. It is desirable to generate intelligible and natural speech with a fast speed in unconstrained settings. Currently, to handle the more complicated scenarios, most existing methods adopt the autoregressive architecture, which is optimized with the MSE loss. Although these methods have achieved promising performance, they are prone to bring issues including high inference latency and mel-spectrogram over-smoothness. To tackle these problems, we propose a novel flow-based non-autoregressive lip-to-speech model (GlowLTS) to break autoregressive constraints and achieve faster inference. Concretely, we adopt a flow-based decoder which is optimized by maximizing the likelihood of the training data and is capable of more natural and fast speech generation. Moreover, we devise a condition module to improve the intelligibility of generated speech. We demonstrate the superiority of our proposed method through objective and subjective evaluation on Lip2Wav-Chemistry-Lectures and Lip2Wav-Chess-Analysis datasets. Our demo video can be found at https://glowlts.github.io/. Jinzheng He, Zhou Zhao 0001, Yi Ren 0006, Jinglin Liu, Baoxing Huai, Nicholas Jing Yuan |
AAAI | 1 |
| 2022 | DualSign: Semi-Supervised Sign Language Production with Balanced Multi-Modal Multi-Task Dual TransformationabstractSign Language Production (SLP) aims to translate a spoken language description to its corresponding continuous sign language sequence. A prevailing solution for this problem is in a two-staged manner: it formulates SLP as two sub-tasks, i.e., Text to Gloss (T2G) translation and Gloss to Pose (G2P) animation, with gloss annotations as pivots. Although two-staged approaches achieve better performance than their direct translation counterparts, the requirement of gloss intermediaries causes a parallel data bottleneck. In this paper, to reduce reliance on gloss annotations in two-staged approaches, we propose DualSign, a semi-supervised two-staged SLP framework, which can effectively utilize partially gloss-annotated text-pose pairs and monolingual gloss data. The key component of DualSign is a novel Balanced Multi-Modal Multi-Task Dual Transformation (BM3T-DT) method, where two well-designed models, i.e., a Multi-Modal T2G model (MM-T2G) and a Multi-Task G2P model (MT-G2P), are jointly trained by leveraging their task duality and unlabeled data. After applying BM3T-DT, we derive the expected uni-modal T2G model from the well-trained MM-T2G with knowledge distillation. Considering that the MM-T2G may suffer from modality imbalance when decoding with multiple input modalities, we devise a cross-modal balancing loss, further boosting the system's overall performance. Extensive experiments conducted on the PHOENIX14T dataset show the effectiveness of our approach in the semi-supervised setting. By training with additionally collected unlabeled data, DualSign substantially improves previous state-of-the-art SLP methods. Wencan Huang, Zhou Zhao 0001, Jinzheng He, Mingmin Zhang 0001 |
ACM Multimedia | 3 |
| 2022 | M4Singer: A Multi-Style, Multi-Singer and Musical Score Provided Mandarin Singing CorpusabstractThe lack of publicly available high-quality and accurately labeled datasets has long been a major bottleneck for singing voice synthesis (SVS). To tackle this problem, we present M4Singer, a free-to-use Multi-style, Multi-singer Mandarin singing collection with elaborately annotated Musical scores as well as its benchmarks. Specifically, 1) we construct and release a large high-quality Chinese singing voice corpus, which is recorded by 20 professional singers, covering 700 Chinese pop songs as well as all the four SATB types (i.e., soprano, alto, tenor, and bass); 2) we take extensive efforts to manually compose the musical scores for each recorded song, which are necessary to the study of the prosody modeling for SVS. 3) To facilitate the use and demonstrate the quality of M4Singer, we conduct four different benchmark experiments: score-based SVS, controllable singing voice (CSV), singing voice conversion (SVC) and automatic music transcription (AMT). Ruiqi Li 0002, Shoutong Wang, Liqun Deng, Jinglin Liu, Yi Ren 0006, Jinzheng He, Rongjie Huang 0001, Jieming Zhu, Zhou Zhao 0001 |
NeurIPS | 7 |
| 2020 | PopMAG: Pop Music Accompaniment GenerationabstractIn pop music, accompaniments are usually played by multiple instruments (tracks) such as drum, bass, string and guitar, and can make a song more expressive and contagious by arranging together with its melody. Previous works usually generate multiple tracks separately and the music notes from different tracks not explicitly depend on each other, which hurts the harmony modeling. To improve harmony, in this paper, we propose a novel MUlti-track MIDI representation (MuMIDI), which enables simultaneous multi-track generation in a single sequence and explicitly models the dependency of the notes from different tracks. While this greatly improves harmony, unfortunately, it enlarges the sequence length and brings the new challenge of long-term music modeling. We further introduce two new techniques to address this challenge: 1) We model multiple note attributes (e.g., pitch, duration, velocity) of a musical note in one step instead of multiple steps, which can shorten the length of a MuMIDI sequence. 2) We introduce extra long-context as memory to capture long-term dependency in music. We call our system for pop music accompaniment generation as PopMAG. We evaluate PopMAG on multiple datasets (LMD, FreeMidi and CPMD, a private dataset of Chinese pop songs) with both subjective and objective metrics. The results demonstrate the effectiveness of PopMAG for multi-track harmony modeling and long-term context modeling. Specifically, PopMAG wins 42%/38%/40% votes when comparing with ground truth musical pieces on LMD, FreeMidi and CPMD datasets respectively and largely outperforms other state-of-the-art music accompaniment generation models and multi-track MIDI representations in terms of subjective and objective metrics. Yi Ren 0006, Jinzheng He, Xu Tan 0003, Tao Qin 0001, Zhou Zhao 0001, Tie-Yan Liu |
ACM Multimedia | 2 |