Yuyue Wang 0003

dblp:267/2206-3 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2025
0009-0005-6987-1028ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Animate and Sound an Image
abstract
This paper addresses a promising yet underexplored task, Image-to-Sounding-Video (I2SV) generation, which animates a static image and generates synchronized sound simultaneously. Despite advances in video and audio generation models, challenges remain to develop a unified model for generating naturally sounding videos. In this work, we propose a novel approach that leverages two separate pretrained diffusion models and makes vision and audio influence each other during generation based on the Diffusion Transformer (DiT) architecture. First, the individual video and audio pretrained generation models are decomposed into input, output, and expert sub-modules. We propose using a unified joint DiT block to integrate the expert sub-modules to effectively model the interaction between the two modalities, resulting in high-quality I2SV generation. Then, we introduce a joint classifier-free guidance technique to boost the performance during joint generation. Finally, we conduct extensive experiments on three popular benchmark datasets, and in both objective and subjective evaluation our method surpass all the baseline methods in almost all metrics. Case studies show our generated sounding videos are high quality and synchronized between video and audio.
Xihua Wang 0002, Ruihua Song, Chongxuan Li, Xin Cheng 0008, Yihan Wu 0008, Yuyue Wang 0003, Hongteng Xu
CVPR7
2025 LoVA: Long-form Video-to-Audio Generation
abstract
Video-to-audio (V2A) generation is important for video editing and post-processing, enabling the creation of semantics-aligned audio for silent video. However, most existing methods focus on generating short-form audio for short video segment (less than 10 seconds), while giving little attention to the scenario of long-form video inputs. For current UNet-based diffusion V2A models, an inevitable problem when handling long-form audio generation is the inconsistencies within the final concatenated audio. In this paper, we first highlight the importance of long-form V2A problem. Besides, we propose LoVA, a novel model for Long-form Video-to-Audio generation. Based on the Diffusion Transformer (DiT) architecture, LoVA proves to be more effective at generating long-form audio compared to existing autoregressive models and UNet-based diffusion models. Extensive objective and subjective experiments demonstrate that LoVA achieves comparable performance on 10second V2A benchmark and outperforms all other baselines on a benchmark with long-form video input.
Xin Cheng 0008, Xihua Wang 0002, Yihan Wu 0008, Yuyue Wang 0003, Ruihua Song
ICASSP4
2025 VAFlow: Video-to-Audio Generation with Cross-Modality Flow Matching
Xihua Wang 0002, Xin Cheng 0008, Yuyue Wang 0003, Ruihua Song
ICCV3
2025 A Visual Speech Language Model for Visual Text-to-Speech Task
abstract
The task of Visual Text-to-Speech (VisualTTS), also known as video dubbing, aims to generate speech synchronized with the lip movements in an input video, in additional to being consistent with the content of input text and cloning the timbre of a reference speech. Existing VisualTTS models typically adopt lightweight architectures and design specialized modules to achieve the above goals respectively, yet the speech quality is not satisfied due to the model capacity and the limited data in VisualTTS. Recently, speech large language models (SpeechLLM) show the robust ability to generate high-quality speech. But few work has been done to well leverage temporal cues from video input in generating lip-synchronized speech. To generate both high-quality and lip-synchronized speech in VisualTTS tasks, we propose a novel Visual Speech Language Model called VSpeechLM based upon a SpeechLLM. To capture the synchronization relationship between text and video, we propose a text-video aligner. It first learns fine-grained alignment between phonemes and lip movements, and then outputs an expanded phoneme sequence containing lip-synchronization cues. Next, our proposed SpeechLLM based decoders take the expanded phoneme sequence as input and learns to generate lip-synchronized speech. Extensive experiments demonstrate that our VSpeechLM significantly outperforms previous VisualTTS methods in terms of overall quality, speaker similarity, and synchronization metrics.
Yuyue Wang 0003, Xin Cheng 0008, Yihan Wu 0008, Xihua Wang 0002, Jinchuan Tian, Ruihua Song
MMAsia1
2024 TiVA: Time-Aligned Video-to-Audio Generation
Xihua Wang 0002, Yuyue Wang 0003, Yihan Wu 0008, Ruihua Song, Xu Tan 0003, Zehua Chen 0005, Hongteng Xu, Guodong Sui
ACM Multimedia2
2023 ComedicSpeech: Text To Speech For Stand-up Comedies in Low-Resource Scenarios
abstract
Text to Speech (TTS) models can generate natural and highquality speech, but it is not expressive enough when synthesizing speech with dramatic expressiveness, such as stand-up comedies.Considering comedians have diverse personal speech styles, including personal prosody, rhythm, and fillers, it requires real-world datasets and strong speech style modeling capabilities, which brings challenges.In this paper, we construct a new dataset and develop ComedicSpeech, a TTS system tailored for the stand-up comedy synthesis in low-resource scenarios.First, we extract prosody representation by the prosody encoder and condition it to the TTS model in a flexible way.Second, we enhance the personal rhythm modeling by a conditional duration predictor.Third, we model the personal fillers by introducing comedian-related special tokens.Experiments show that ComedicSpeech achieves better expressiveness than baselines with only ten-minute training data for each comedian.The audio samples are available at https://xh621.github.io/stand-up-comedy-demo/
Yuyue Wang 0003, Yihan Wu 0008, Ruihua Song
INTERSPEECH1