Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Xuanchen Wang

dblp:377/7151 · DBLP profile ↗
← Back
3ranked-venue papers
3as first author
3since 2021 · last 2025
0009-0006-1655-5942ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Generative modeling · 100%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 77% Computer animation and physical simulation · 23%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
0.912025
ChoreoMuse: Robust Music-to-Dance Video Generation with Style Transfer and Beat-Adherent Motion · ACM Multimedia 2025
Machine learning › Generative modeling › motion generation
music-to-dance generation
0.912025
ChoreoMuse: Robust Music-to-Dance Video Generation with Style Transfer and Beat-Adherent Motion · ACM Multimedia 2025
Visual content generation and editing › video generation
dance video synthesis
0.912025
ChoreoMuse: Robust Music-to-Dance Video Generation with Style Transfer and Beat-Adherent Motion · ACM Multimedia 2025
Computer animation and physical simulation
motion synthesis
0.312025
ChoreoMuse: Robust Music-to-Dance Video Generation with Style Transfer and Beat-Adherent Motion · ACM Multimedia 2025

Methods — techniques the papers use, named apart from their topics

style transfer · 1.7music encoder · 1.7diffusion · 1.7
YearPublicationVenuePosition
2025 ChoreoMuse: Robust Music-to-Dance Video Generation with Style Transfer and Beat-Adherent Motion
abstract
Modern artistic productions increasingly demand automated choreography generation that adapts to diverse musical styles and individual dancer characteristics. Existing approaches often fail to produce high-quality dance videos that harmonize with both musical rhythm and user-defined choreography styles, limiting their applicability in real-world creative contexts. To address this gap, we introduce ChoreoMuse, a diffusion-based framework that uses SMPL format parameters and their variation version as intermediaries between music and video generation, thereby overcoming the usual constraints imposed by video resolution. Critically, ChoreoMuse supports style-controllable, high-fidelity dance video generation across diverse musical genres and individual dancer characteristics, including the flexibility to handle any reference individual at any resolution. Our method employs a novel music encoder MotionTune to capture motion cues from audio, ensuring that the generated choreography closely follows the beat and expressive qualities of the input music. To quantitatively evaluate how well the generated dances match both musical and choreographic styles, we introduce two new metrics that measure alignment with the intended stylistic cues. Extensive experiments confirm that ChoreoMuse achieves state-of-the-art performance across multiple dimensions, including video quality, beat alignment, dance diversity, and style adherence, demonstrating its potential as a robust solution for a wide range of creative applications. Video results can be found on our project page: https://choreomuse.github.io.
Xuanchen Wang, Heng Wang 0007, Tom Weidong Cai
ACM Multimedia1
2025 KeyRegionPose: Region-Aware Feature Interaction and Multi-Scale Token Pruning for Efficient Human Pose Estimation
abstract
The primary challenge in deploying Human Pose Estimation (HPE) methods in real-world applications lies in balancing computational speed, model compactness, and prediction accuracy. Existing methods achieve strong performance in one or two aspects, but usually at the expense of the remaining one. To overcome this trade-off, we propose KeyRegionPose, a novel framework that achieves high accuracy while reducing model size and computational cost. Central to our design is the Region Focus Mechanism, which enables the model to concentrate on keypoint-relevant regions rather than the entire image. During training, we generate intermediate keypoint proposals to estimate keypoint-specific areas, from which the model learns region-focused features and refines predictions. To ensure accurate keypoint localization and enhance final pose estimation performance, we introduce a Cross-Representation Consistency Loss (CRC Loss) that enforces alignment between the predicted heatmaps and the regressed keypoint coordinates. Additionally, we propose Progressive Multi-Scale Token Pruning (PMTP), a strategy that prunes irrelevant tokens across multiple feature scales to accelerate inference. KeyRegionPose achieves 76.0 AP on the COCO validation set and 75.4 AP on the test-dev set, with only 20.0 million parameters and 8.6 GFLOPs—representing a 27.3% reduction in parameter count, 21.8% decrease in GFLOPs, and a competitive result (+0.2%) over state-of-the-art lightweight HPE models.
Xuanchen Wang, Heng Wang 0007, Dongnan Liu, Tom Weidong Cai
MMAsia1
2025 Dance any Beat: Blending Beats with Visuals in Dance Video Generation
abstract
Generating dance from music is crucial for advancing automated choreography. Current methods typically produce skeleton keypoint sequences instead of dance videos and lack the capability to make specific individuals dance, which reduces their real-world applicability. These methods also require precise keypoint annotations, complicating data collection and limiting the use of self-collected video datasets. To overcome these challenges, we introduce a novel task: generating dance videos directly from images of individuals guided by music. This task enables the dance generation of specific individuals without requiring keypoint annotations, making it more versatile and applicable to various situations. Our solution, the Dance Any Beat Diffusion model (DabFusion), utilizes a reference image and a music piece to generate dance videos featuring various dance types and choreographies. The music is analyzed by our specially designed music encoder, which identifies essential features including dance style, movement, and rhythm. DabFusion excels in generating dance videos not only for individuals in the training dataset but also for any previously unseen person. This versatility stems from its approach of generating latent optical flow, which contains all necessary motion information to animate any person in the image. We evaluate DabFusion's performance using the AIST + + dataset, focusing on video quality, audio-video synchronization, and motion-music alignment. We propose a 2D Motion-Music Alignment Score (2D-MM Align), which builds on the Beat Alignment Score to more effectively evaluate motion-music alignment for this new task. Experiments show that our DabFusion establishes a solid baseline for this innovative task. Video results can be found on our project page: https://DabFusion.github.io.
Xuanchen Wang, Heng Wang 0007, Dongnan Liu, Tom Weidong Cai
WACV1