Junxian Wu 0003

dblp:318/3989-3 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2025
0009-0005-8447-2121ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
3 papers
Audio and music processing · 81% Multimedia analysis and retrieval · 12% Multimedia systems and quality of experience · 4%
Artificial intelligence
1 paper
Generative modeling · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Audio and music processing
music generation
2.632025
Spatial-Temporal Decomposition and Alignment in Controllable Video-to-Music Generation · ACM Multimedia 2025
Controllable Video-to-Music Generation with Multiple Time-Varying Conditions · ACM Multimedia 2025
GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions · AAAI 2025
Audio and music processing › music generation
video-to-music generation
2.632025
Spatial-Temporal Decomposition and Alignment in Controllable Video-to-Music Generation · ACM Multimedia 2025
Controllable Video-to-Music Generation with Multiple Time-Varying Conditions · ACM Multimedia 2025
GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions · AAAI 2025
Machine learning › Generative modeling
flow matching
0.912025
Spatial-Temporal Decomposition and Alignment in Controllable Video-to-Music Generation · ACM Multimedia 2025
Audio and music processing › music generation
controllable music generation
0.912025
Controllable Video-to-Music Generation with Multiple Time-Varying Conditions · ACM Multimedia 2025
Multimedia analysis and retrieval › audio-visual analysis
video-music correspondence
0.912025
GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions · AAAI 2025
Multimedia systems and quality of experience › multimedia synchronization
audio-visual synchronization
0.312025
Controllable Video-to-Music Generation with Multiple Time-Varying Conditions · ACM Multimedia 2025
Image and video processing › image sequence processing
temporal alignment
0.312025
Controllable Video-to-Music Generation with Multiple Time-Varying Conditions · ACM Multimedia 2025

Methods — techniques the papers use, named apart from their topics

flow-matching alignment · 1.7feature-free guidance · 1.7feature alignment · 1.7zero-shot generation · 0.9two-stage training · 0.9temporal alignment attention · 0.9spatial-temporal decomposition · 0.9hierarchical attention · 0.9feature selection · 0.9control-guided decoding · 0.9conditional fusion · 0.9
YearPublicationVenuePosition
2025 GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions
abstract
Composing music for video is essential yet challenging, leading to a growing interest in automating music generation for video applications. Existing approaches often struggle to achieve robust music-video correspondence and generative diversity, primarily due to inadequate feature alignment methods and insufficient datasets. In this study, we present General Video-to-Music Generation model (GVMGen), designed for generating high-related music to the video input. Our model employs hierarchical attentions to extract and align video features with music in both spatial and temporal dimensions, ensuring the preservation of pertinent features while minimizing redundancy. Remarkably, our method is versatile, capable of generating multi-style music from different video inputs, even in zero-shot scenarios. We also propose an evaluation model along with two novel objective metrics for assessing video-music alignment. Additionally, we have compiled a large-scale dataset comprising diverse types of video-music pairs. Experimental results demonstrate that GVMGen surpasses previous models in terms of music-video correspondence, music quality generative diversity, and application universality.
Heda Zuo, Weitao You, Junxian Wu 0003, Shihong Ren, Pei Chen 0005, Mingxu Zhou, Yujia Lu, Lingyun Sun
AAAI3
2025 Controllable Video-to-Music Generation with Multiple Time-Varying Conditions
abstract
Music enhances video narratives and emotions, driving demand for automatic video-to-music (V2M) generation. However, existing V2M methods relying solely on visual features or supplementary textual inputs generate music in a black-box manner, often failing to meet user expectations. To address this challenge, we propose a novel multi-condition guided V2M generation framework that incorporates multiple time-varying conditions for enhanced control over music generation. Our method uses a two-stage training strategy that enables learning of V2M fundamentals and audiovisual temporal synchronization while meeting users' needs for multi-condition control. In the first stage, we introduce a fine-grained feature selection module and a progressive temporal alignment attention mechanism to ensure flexible feature alignment. For the second stage, we develop a dynamic conditional fusion module and a control-guided decoder module to integrate multiple conditions and accurately guide the music composition process. Extensive experiments demonstrate that our method outperforms existing V2M pipelines in both subjective and objective evaluations, significantly enhancing control and alignment with user expectations.
Junxian Wu 0003, Weitao You, Heda Zuo, Dengming Zhang, Pei Chen 0005, Lingyun Sun
ACM Multimedia1
2025 Spatial-Temporal Decomposition and Alignment in Controllable Video-to-Music Generation
abstract
Achieving high-quality output alongside enhanced controllability is crucial in video-to-music generation, especially for optimizing user experience in real-life application scenarios. Most existing studies emphasize generative quality, but often overlooking the vital aspect of controllability. Therefore, the generated music cannot be easily fine-tuned or modified to meet users' expectations. In this paper, we delve into the spatial-temporal decomposition and alignment in controllable video-to-music generation. We first introduce a novel video-music decomposition and transformation approach in both spatial and temporal domain, and enhance the cross-modal correspondence through feature alignment and flow-matching based alignment. Furthermore, our method attains unsupervised controllability during training via feature-free guidance. Experimental results demonstrate that our model achieves state-of-the-art results in overall generative quality. Moreover, its controllability significantly outperforms existing models, making it exceptionally well-suited to accommodate users' flexible and diverse control requirements.
Weitao You, Heda Zuo, Junxian Wu 0003, Dengming Zhang, Zhibin Zhou 0002, Lingyun Sun
ACM Multimedia3