Huadai Liu

dblp:321/0749 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
15since 2021 · last 2025
0009-0004-5782-5641ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021
YearPublicationVenuePosition
2025 FlashAudio: Rectified Flow for Fast and High-Fidelity Text-to-Audio Generation
abstract
Recent advancements in latent diffusion models (LDMs) have markedly enhanced text-to-audio generation, yet their iterative sampling processes impose substantial computational demands, limiting practical deployment. While recent methods utilizing consistency-based distillation aim to achieve few-step or single-step inference, their one-step performance is constrained by curved trajectories, preventing them from surpassing traditional diffusion models. In this work, we introduce FlashAudio with rectified flows to learn straight flow for fast simulation. To alleviate the inefficient timesteps allocation and suboptimal distribution of noise, FlashAudio optimizes the time distribution of rectified flow with Bifocal Samplers and proposes immiscible flow to minimize the total distance of data-noise pairs in a batch vias assignment. Furthermore, to address the amplified accumulation error caused by the classifier-free guidance (CFG), we propose Anchored Optimization, which refines the guidance scale by anchoring it to a reference trajectory. Experimental results on text-to-audio generation demonstrate that FlashAudio’s one-step generation performance surpasses the diffusion-based models with hundreds of sampling steps on audio quality and enables a sampling speed of 400x faster than real-time on a single NVIDIA 4090Ti GPU. Code will be available at https://github.com/liuhuadai/FlashAudio. Audio Samples are available at https://FlashAudio-TTA.github.io/.
Huadai Liu, Rongjie Huang 0001, Yang Liu 0278, Zhou Zhao 0001, Wei Xue 0002
ACL (1)1
2025 Fast Adaptation of Pretrained Speaker Verification System for Source Speaker Tracking
abstract
Traditional speaker verification system aims at distinguish speaker identity in real world audio, and has achieved satisfying performance in many scenarios. However, it is also very vulnerable, and can be easily attacked by voice anonymization system. In this report, we describe how to fast adapt a pretrained speaker verification model to source speaker tracking task with pretrained feature and Lora [1] technique. It significantly reduce EER on voice anonymization system, as well as keep its performance in real world audio intact. Experiment on Attacker Challenge [2] shows that our system successfully reduce baseline EER by 32% in average, and achieve lowest EER in all voice anonymization system except T8-5.
Yuxuan Wang 0014, Huadai Liu
ICASSP4
2025 Build LLM-Based Zero-Shot Streaming TTS System with Cosyvoice
abstract
LLM-based text-to-speech(TTS) system has becoming the new trend and SOTA due to its high naturalness and zero-shot capability. However, it relies heavily on training data, usually requires at least thousands hours of labeled audio. In this report, we describe how to use pretrained CosyVoice model, to develop a streaming TTS system which supports Indian English and Indian languages. Though the pretrained CosyVoice model has never seen such data, it shows good performance in both specific speaker TTS and zero-shot voice clone after finetuning with merely 280 hours data. Experiment on LIMMITS25 challenge shows that our system achieves 4.46/4.19/4.55 naturalness, and 4.29/4.34/4.27 similarity in track1/track2/track3 respectively, which ranked 1st in all tracks.
Yuxuan Wang 0014, Hao Wang 0199, Huadai Liu, Zhihao Du
ICASSP5
2025 Both Ears Wide Open: Towards Language-Driven Spatial Audio Generation
abstract
Recently, diffusion models have achieved great success in mono-channel audio generation. However, when it comes to stereo audio generation, the soundscapes often have a complex scene of multiple objects and directions. Controlling stereo audio with spatial contexts remains challenging due to high data costs and unstable generative models. To the best of our knowledge, this work represents the first attempt to address these issues. We first construct a large-scale, simulation-based, and GPT-assisted dataset, BEWO-1M, with abundant soundscapes and descriptions even including moving and multiple sources. Beyond text modality, we have also acquired a set of images and rationally paired stereo audios through retrieval to advance multimodal generation. Existing audio generation models tend to generate rather random and indistinct spatial audio. To provide accurate guidance for Latent Diffusion Models, we introduce the SpatialSonic model utilizing spatial-aware encoders and azimuth state matrices to reveal reasonable spatial guidance. By leveraging spatial guidance, our model not only achieves the objective of generating immersive and controllable spatial audio from text but also extends to other modalities as the pioneer attempt. Finally, under fair settings, we conduct subjective and objective evaluations on simulated and real-world data to compare our approach with prevailing methods. The results demonstrate the effectiveness of our method, highlighting its capability to generate spatial audio that adheres to physical rules.
Peiwen Sun, Sitong Cheng, Xiangtai Li, Zhen Ye 0006, Huadai Liu, Honggang Zhang 0002, Wei Xue 0002, Yike Guo
ICLR5
2025 OmniAudio: Generating Spatial Audio from 360-Degree Video
abstract
Traditional video-to-audio generation techniques primarily focus on perspective video and non-spatial audio, often missing the spatial cues necessary for accurately representing sound sources in 3D environments. To address this limitation, we introduce a novel task, 360V2SA, to generate spatial audio from 360-degree videos, specifically producing First-order Ambisonics (FOA) audio - a standard format for representing 3D spatial audio that captures sound directionality and enables realistic 3D audio reproduction. We first create Sphere360, a novel dataset tailored for this task that is curated from real-world data. We also design an efficient semi-automated pipeline for collecting and cleaning paired video-audio data. To generate spatial audio from 360-degree video, we propose a novel framework OmniAudio, which leverages self-supervised pre-training using both spatial audio data (in FOA format) and large-scale non-spatial data. Furthermore, OmniAudio features a dual-branch framework that utilizes both panoramic and perspective video inputs to capture comprehensive local and global information from 360-degree videos. Experimental results demonstrate that OmniAudio achieves state-of-the-art performance across both objective and subjective metrics on Sphere360. Code and datasets are available at https://github.com/liuhuadai/OmniAudio. The project website is available at https://OmniAudio-360V2SA.github.io.
Huadai Liu, Tianyi Luo, Kaicheng Luo, Qikai Jiang, Peiwen Sun, Rongjie Huang 0001, Qian Chen 0003, Wen Wang 0001, Xiangtai Li, Shiliang Zhang, Zhijie Yan, Zhou Zhao 0001, Wei Xue 0002
ICML1
2025 MelodyEdit: Zero-shot Music Editing with Disentangled Inversion Control
abstract
Text-guided diffusion models revolutionize audio generation by adapting source audio to specific text prompts. However, existing zero-shot audio editing methods such as DDIM inversion accumulate errors across diffusion steps, reducing the effectiveness. Moreover, existing editing methods struggle with conducting complex non-rigid music edits while maintaining content integrity and high fidelity. To address these challenges, we propose MelodyEdit, a novel zero-shot music editing system based on innovative Disentangled Inversion Control (DIC) technique, which comprises Harmonized Attention Control and Disentangled Inversion. Disentangled Inversion disentangles the diffusion process into triple branches to rectify the deviated path of the source branch caused by DDIM inversion. Harmonized Attention Control unifies the mutual self-attention control and the cross-attention control with an intermediate Harmonic Branch to progressively generate the desired harmonic and melodic information in the target music. We also introduce ZoME-Bench, a comprehensive music editing benchmark with 1,100 samples covering ten distinct editing categories. ZoME-Bench facilitates both zero-shot and instruction-based music editing tasks. Our method outperforms state-of-the-art inversion techniques in editing fidelity and content preservation.
Huadai Liu, Xiangtai Li, Wen Wang 0019, Qian Chen 0003, Rongjie Huang 0001, Zhou Zhao 0001, Wei Xue 0002
ACM Multimedia1
2025 ThinkSound: Chain-of-Thought Reasoning in Multimodal LLMs for Audio Generation and Editing
abstract
While end-to-end video-to-audio generation has greatly improved, producing high-fidelity audio that authentically captures the nuances of visual content remains challenging. Like professionals in the creative industries, this generation requires sophisticated reasoning about items such as visual dynamics, acoustic environments, and temporal relationships. We present **ThinkSound**, a novel framework that leverages Chain-of-Thought (CoT) reasoning to enable stepwise, interactive audio generation and editing for videos. Our approach decomposes the process into three complementary stages: foundational foley generation that creates semantically coherent soundscapes, interactive object-centric refinement through precise user interactions, and targeted editing guided by natural language instructions. At each stage, a multimodal large language model generates contextually aligned CoT reasoning that guides a unified audio foundation model. Furthermore, we introduce **AudioCoT**, a comprehensive dataset with structured reasoning annotations that establishes connections between visual content, textual descriptions, and sound synthesis. Experiments demonstrate that ThinkSound achieves state-of-the-art performance in video-to-audio generation across both audio metrics and CoT metrics, and excels in the out-of-distribution Movie Gen Audio benchmark. The project page is available at https://ThinkSound-Project.github.io.
Huadai Liu, Kaicheng Luo, Wen Wang 0001, Qian Chen 0003, Zhou Zhao 0001, Wei Xue 0002
NeurIPS1
2024 AntCritic: Argument Mining for Free-Form and Visually-Rich Financial Comments
abstract
Argument mining aims to detect all possible argumentative components and identify their relationships automatically. As a thriving task in natural language processing, there has been a large amount of corpus for academic study and application development in this field. However, the research in this area is still constrained by the inherent limitations of existing datasets. Specifically, all the publicly available datasets are relatively small in scale, and few of them provide information from other modalities to facilitate the learning process. Moreover, the statements and expressions in these corpora are usually in a compact form, which restricts the generalization ability of models. To this end, we collect a novel dataset AntCritic to serve as a helpful complement to this area, which consists of about 10k free-form and visually-rich financial comments and supports both argument component detection and argument relation prediction tasks. Besides, to cope with the challenges brought by scenario expansion, we thoroughly explore the fine-grained relation prediction and structure reconstruction scheme and discuss the encoding mechanism for visual styles and layouts. On this basis, we design two simple but effective model architectures and conduct various experiments on this dataset to provide benchmark performances as a reference and verify the practicability of our proposed architecture. We release our data and code in this link, and this dataset follows CC BY-NC-ND 4.0 license.
Huadai Liu, Xuan Lin, Jingjing Huo
LREC/COLING1
2024 AudioLCM: Efficient and High-Quality Text-to-Audio Generation with Minimal Inference Steps
Huadai Liu, Rongjie Huang 0001, Yang Liu 0278, Hengyuan Cao, Xize Cheng, Zhou Zhao 0001
ACM Multimedia1
2024 Extending Multi-modal Contrastive Representations
abstract
Multi-modal contrastive representation (MCR) of more than three modalities is critical in multi-modal learning. Although recent methods showcase impressive achievements, the high dependence on large-scale, high-quality paired data and the expensive training costs limit their further development. Inspired by recent C-MCR, this paper proposes $\textbf{Ex}$tending $\textbf{M}$ultimodal $\textbf{C}$ontrastive $\textbf{R}$epresentation (Ex-MCR), a training-efficient and paired-data-free method to build unified contrastive representation for many modalities. Since C-MCR is designed to learn a new latent space for the two non-overlapping modalities and projects them onto this space, a significant amount of information from their original spaces is lost in the projection process. To address this issue, Ex-MCR proposes to extend one modality's space into the other's, rather than mapping both modalities onto a completely new space. This method effectively preserves semantic alignment in the original space. Experimentally, we extend pre-trained audio-text and 3D-image representations to the existing vision-text space. Without using paired data, Ex-MCR achieves comparable performance to advanced methods on a series of audio-image-text and 3D-image-text tasks and achieves superior performance when used in parallel with data-driven methods. Moreover, semantic alignment also emerges between the extended modalities (e.g., audio and 3D).
Zehan Wang 0001, Luping Liu, Rongjie Huang 0001, Xize Cheng, Zhenhui Ye, Huadai Liu, Haifeng Huang 0001, Yang Zhao 0022, Tao Jin 0004, Zhou Zhao 0001
NeurIPS8
2023 AV-TranSpeech: Audio-Visual Robust Speech-to-Speech Translation
abstract
Rongjie Huang, Huadai Liu, Xize Cheng, Yi Ren, Linjun Li, Zhenhui Ye, Jinzheng He, Lichao Zhang, Jinglin Liu, Xiang Yin, Zhou Zhao. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Rongjie Huang 0001, Huadai Liu, Xize Cheng, Yi Ren 0006, Linjun Li, Zhenhui Ye, Jinzheng He, Jinglin Liu, Xiang Yin 0006, Zhou Zhao 0001
ACL (1)2
2023 ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
abstract
Text-to-speech(TTS) has undergone remarkable improvements in performance, particularly with the advent of Denoising Diffusion Probabilistic Models (DDPMs).However, the perceived quality of audio depends not solely on its content, pitch, rhythm, and energy, but also on the physical environment.In this work, we propose ViT-TTS, the first visual TTS model with scalable diffusion transformers.ViT-TTS complement the phoneme sequence with the visual information to generate high-perceived audio, opening up new avenues for practical applications of AR and VR to allow a more immersive and realistic audio experience.To mitigate the data scarcity in learning visual acoustic information, we 1) introduce a selfsupervised learning framework to enhance both the visual-text encoder and denoiser decoder; 2) leverage the diffusion transformer scalable in terms of parameters and capacity to learn visual scene information.Experimental results demonstrate that ViT-TTS achieves new stateof-the-art results, outperforming cascaded systems and other baselines regardless of the visibility of the scene.With low-resource data (1h, 2h, 5h), ViT-TTS achieves comparative results with rich-resource baselines.1 2
Huadai Liu, Rongjie Huang 0001, Xuan Lin, Maozong Zheng, Hong Chen 0002, Jinzheng He, Zhou Zhao 0001
EMNLP1
2023 MixSpeech: Cross-Modality Self-Learning with Audio-Visual Stream Mixup for Visual Speech Translation and Recognition
abstract
Multi-media communications facilitate global interaction among people. However, despite researchers exploring cross-lingual translation techniques such as machine translation and audio speech translation to overcome language barriers, there is still a shortage of cross-lingual studies on visual speech. This lack of research is mainly due to the absence of datasets containing visual speech and translated text pairs. In this paper, we present AVMuST-TED, the first dataset for Audio-Visual Multilingual Speech Translation, derived from TED talks. Nonetheless, visual speech is not as distinguishable as audio speech, making it difficult to develop a mapping from source speech phonemes to the target language text. To address this issue, we propose MixSpeech, a cross-modality self-learning framework that utilizes audio speech to regularize the training of visual speech tasks. To further minimize the cross-modality gap and its impact on knowledge transfer, we suggest adopting mixed speech, which is created by interpolating audio and visual streams, along with a curriculum learning strategy to adjust the mixing ratio as needed. MixSpeech enhances speech translation in noisy environments, improving BLEU scores for four languages on AVMuST-TED by +1.4 to +4.2. Moreover, it achieves state-of-the-art performance in lip reading on CMLR (11.1%), LRS2 (25.5%), and LRS3 (28.0%).
Xize Cheng, Tao Jin 0004, Rongjie Huang 0001, Linjun Li, Zehan Wang 0001, Ye Wang 0018, Huadai Liu, Aoxiong Yin, Zhou Zhao 0001
ICCV8
2023 TranSpeech: Speech-to-Speech Translation With Bilateral Perturbation
Rongjie Huang 0001, Jinglin Liu, Huadai Liu, Yi Ren 0006, Jinzheng He, Zhou Zhao 0001
ICLR3
2022 ProDiff: Progressive Fast Diffusion Model for High-Quality Text-to-Speech
abstract
Denoising diffusion probabilistic models (DDPMs) have recently achieved leading performances in many generative tasks. However, the inherited iterative sampling process costs hinder their applications to text-to-speech deployment. Through the preliminary study on diffusion model parameterization, we find that previous gradient-based TTS models require hundreds or thousands of iterations to guarantee high sample quality, which poses a challenge for accelerating sampling. In this work, we propose ProDiff, on progressive fast diffusion model for high-quality text-to-speech. Unlike previous work estimating the gradient for data density, ProDiff parameterizes the denoising model by directly predicting clean data to avoid distinct quality degradation in accelerating sampling. To tackle the model convergence challenge with decreased diffusion iterations, ProDiff reduces the data variance in the target site via knowledge distillation. Specifically, the denoising model uses the generated mel-spectrogram from an N-step DDIM teacher as the training target and distills the behavior into a new model with N/2 steps. As such, it allows the TTS model to make sharp predictions and further reduces the sampling time by orders of magnitude. Our evaluation demonstrates that ProDiff needs only 2 iterations to synthesize high-fidelity mel-spectrograms, while it maintains sample quality and diversity competitive with state-of-the-art models using hundreds of steps. ProDiff enables a sampling speed of 24x faster than real-time on a single NVIDIA 2080Ti GPU, making diffusion models practically applicable to text-to-speech synthesis deployment for the first time. Our extensive ablation studies demonstrate that each design in ProDiff is effective, and we further show that ProDiff can be easily extended to the multi-speaker setting.
Rongjie Huang 0001, Zhou Zhao 0001, Huadai Liu, Jinglin Liu, Chenye Cui, Yi Ren 0006
ACM Multimedia3