VLDB 2026 Research / reviewers in the wild / expert
Jiaxing Yu
dblp:212/6000
· DBLP profile ↗
8ranked-venue papers
5as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
2 papers |
Audio and music processing · 93% Multimedia systems and quality of experience · 7% | |
| Artificial intelligence
1 paper |
Generative modeling · 100% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Audio and music processing
music generation |
1.9 | 2 | 2026 | Diff-V2M: A Hierarchical Conditional Diffusion Model with Explicit Rhythmic Modeling for Video-to-Music Generation · AAAI 2026 SongGLM: Lyric-to-Melody Generation with 2D Alignment Encoding and Multi-Task Pre-Training · AAAI 2025 |
Machine learning › Generative modeling › diffusion model
conditional diffusion model |
1.0 | 1 | 2026 | Diff-V2M: A Hierarchical Conditional Diffusion Model with Explicit Rhythmic Modeling for Video-to-Music Generation · AAAI 2026 |
Machine learning › Generative modeling
diffusion model |
1.0 | 1 | 2026 | Diff-V2M: A Hierarchical Conditional Diffusion Model with Explicit Rhythmic Modeling for Video-to-Music Generation · AAAI 2026 |
Audio and music processing › music generation
video-to-music generation |
1.0 | 1 | 2026 | Diff-V2M: A Hierarchical Conditional Diffusion Model with Explicit Rhythmic Modeling for Video-to-Music Generation · AAAI 2026 |
Audio and music processing › music generation
lyric-to-melody generation |
0.9 | 1 | 2025 | SongGLM: Lyric-to-Melody Generation with 2D Alignment Encoding and Multi-Task Pre-Training · AAAI 2025 |
Multimedia systems and quality of experience › multimedia synchronization
audio-visual synchronization |
0.3 | 1 | 2026 | Diff-V2M: A Hierarchical Conditional Diffusion Model with Explicit Rhythmic Modeling for Video-to-Music Generation · AAAI 2026 |
Audio and music processing › music technology › computer music
symbolic music representation |
0.3 | 1 | 2025 | SongGLM: Lyric-to-Melody Generation with 2D Alignment Encoding and Multi-Task Pre-Training · AAAI 2025 |
Methods — techniques the papers use, named apart from their topics
rhythm modeling · 2.0diffusion model · 2.0cross-attention · 2.0multi-task pre-training · 0.9general language model · 0.9blank infilling · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Diff-V2M: A Hierarchical Conditional Diffusion Model with Explicit Rhythmic Modeling for Video-to-Music GenerationabstractVideo-to-music (V2M) generation aims to create music that aligns with visual content. However, two main challenges persist in existing methods: (1) the lack of explicit rhythm modeling hinders audiovisual temporal alignments; (2) effectively integrating various visual features to condition music generation remains non-trivial. To address these issues, we propose Diff-V2M, a general V2M framework based on a hierarchical conditional diffusion model, comprising two core components: visual feature extraction and conditional music generation. For rhythm modeling, we begin by evaluating several rhythmic representations, including low-resolution mel-spectrograms, tempograms, and onset detection functions (ODF), and devise a rhythmic predictor to infer them directly from videos. To ensure contextual and affective coherence, we also extract semantic and emotional features. All features are incorporated into the generator via a hierarchical cross-attention mechanism, where emotional features shape the affective tone via the first layer, while semantic and rhythmic features are fused in the second cross-attention layer. To enhance feature integration, we introduce timestep-aware fusion strategies, including feature-wise linear modulation (FiLM) and weighted fusion, allowing the model to adaptively balance semantic and rhythmic cues throughout the diffusion process. Extensive experiments identify low-resolution ODF as a more effective signal for modeling musical rhythm and demonstrate that Diff-V2M outperforms existing models on both in-domain and out-of-domain datasets, achieving state-of-the-art performance in terms of objective metrics and subjective comparisons. Shulei Ji, Jiaxing Yu, Xiangyuan Yang, Songruoyao Wu |
AAAI | 3 |
| 2026 | Dynamic spatiotemporal air quality modeling with local and sparse spatial fusion and structured temporal modeling
Jiaxing Yu, Kaibing Zhang, Dinghua Xue, Pengfang Li, Minna Xiao, Shuyun Yang |
Inf. Sci. | 2 |
| 2025 | SongGLM: Lyric-to-Melody Generation with 2D Alignment Encoding and Multi-Task Pre-TrainingabstractLyric-to-melody generation aims to automatically create melodies based on given lyrics, requiring the capture of complex and subtle correlations between them. However, previous works usually suffer from two main challenges: 1) lyric-melody alignment modeling, which is often simplified to one-syllable/word-to-one-note alignment, while others have the problem of low alignment accuracy; 2) lyric-melody harmony modeling, which usually relies heavily on intermediates or strict rules, limiting model's capabilities and generative diversity. In this paper, we propose SongGLM, a lyric-to-melody generation system that leverages 2D alignment encoding and multi-task pre-training based on the General Language Model (GLM) to guarantee the alignment and harmony between lyrics and melodies. Specifically, 1) we introduce a unified symbolic song representation for lyrics and melodies with word-level and phrase-level (2D) alignment encoding to capture the lyric-melody alignment; 2) we design a multi-task pre-training framework with hierarchical blank infilling objectives (n-gram, phrase, and long span), and incorporate lyric-melody relationships into the extraction of harmonized n-grams to ensure the lyric-melody harmony. We also construct a large-scale lyric-melody paired dataset comprising over 200,000 English song pieces for pre-training and fine-tuning. The objective and subjective results indicate that SongGLM can generate melodies from lyrics with significant improvements in both alignment and harmony, outperforming all the previous baseline methods. Jiaxing Yu, Xinda Wu, Tieyao Zhang, Songruoyao Wu, Le Ma 0002 |
AAAI | 1 |
| 2025 | Enhancing Image Super-Resolution with Dual Compression Transformer
Jiaxing Yu, Zheng Chen 0014, Jingkai Wang 0003, Linghe Kong, Jiajie Yan |
Vis. Comput. | 1 |
| 2024 | RDT-RRT: Real-time double-tree rapidly-exploring random tree path planning for autonomous vehicles
Jiaxing Yu, Ci Chen 0003, Aliasghar Arab, Jingang Yi, Xiaofei Pei, Xuexun Guo |
Expert Syst. Appl. | 1 |
| 2024 | Suno: potential, prospects, and trends
Jiaxing Yu, Songruoyao Wu, Guanting Lu, Li Zhou 0016 |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2024 | Motion Planning and Control of Autonomous Aggressive Vehicle ManeuversabstractAggressive vehicle maneuvers such as those performed by professional racing drivers achieve high agility motion at the edge of handling limits. These aggressive maneuvers can be used to design human-inspired active safety features for next-generation “accident-free” vehicles. We present a motion planning and control design for autonomous aggressive vehicle maneuvers. The motion planner takes advantages of the sparse stable trees and the enhanced rapidly exploring random tree (RRT*) algorithms. The use of the sparsity property helps to reduce the computational cost of the RRT* method by removing non-useful nodes in each iteration and therefore to rapidly converge to the optimal solution. The proposed motion control design allows the vehicle to operate outside the stability region to accomplish a safe, agile maneuver. A safety region is computed to augment the stability region and the motion control is built on a modified nonlinear model predictive control method. We implement the proposed planner and controller and demonstrate the autonomous aggressive maneuvers on a 1/7-scale racing vehicle platform. Comparison with human expert driver and other existing methods is also presented to demonstrate the performance and robustness.Note to Practitioners—Motion planning and control of human driver-inspired aggressive vehicle maneuvers is a challenging task because of high-agility, unstable fast vehicle motions. This paper is motivated by addressing this challenge in autonomous driving technologies. Instead of restricting vehicle motions within a stability region that is taken by existing methods, we augment the conservative stability region to a safety region with guaranteed performance. To improve the computational efficiency of sampling-based motion planners, we take advantage of sparsity and also integration of a nonlinear predictive control method to compute feasible vehicle motion in searching space. The stability of the vehicle motion controller and sub-optimality of the motion planner are analyzed and guaranteed. Using a scaled vehicle testbed, we validate and compare the proposed motion planning and control design with other existing methods and human expert driver. The experimental results demonstrate the superior performance than the other methods and comparable with human expert driving skills. Aliasghar Arab, Kaiyan Yu, Jiaxing Yu, Jingang Yi |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2023 | Hierarchical framework integrating rapidly-exploring random tree with deep reinforcement learning for autonomous vehicle
Jiaxing Yu, Aliasghar Arab, Jingang Yi, Xiaofei Pei, Xuexun Guo |
Appl. Intell. | 1 |