EDBT 2026 Demo / reviewers in the wild / expert
Zixuan Wang 0026
dblp:05/10698-26
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2025
0000-0001-7291-6198ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Minimal Impact ControlNet: Advancing Multi-ControlNet IntegrationabstractWith the advancement of diffusion models, there is a growing demand for high-quality, controllable image generation, particularly through methods that utilize one or multiple control signals based on ControlNet. However, in current ControlNet training, each control is designed to influence all areas of an image, which can lead to conflicts when different control signals are expected to manage different parts of the image in practical applications. This issue is especially pronounced with edge-type control conditions, where regions lacking boundary information often represent low-frequency signals, referred to as silent control signals. When combining multiple ControlNets, these silent control signals can suppress the generation of textures in related areas, resulting in suboptimal outcomes. To address this problem, we propose Minimal Impact ControlNet. Our approach mitigates conflicts through three key strategies: constructing a balanced dataset, combining and injecting feature signals in a balanced manner, and addressing the asymmetry in the score function’s Jacobian matrix induced by ControlNet. These improvements enhance the compatibility of control signals, allowing for freer and more harmonious generation in areas with silent control signals. Shikun Sun, Zixuan Wang 0026, Xubin Li, Tiezheng Ge, Zijie Ye, Xiaoyu Qin 0001, Junliang Xing, Bo Zheng 0007, Jia Jia 0001 |
ICLR | 3 |
| 2025 | DEPO: Enhancing E-commerce Image Background Generation with Short Trajectory Direct Expected Preference Optimization
Shikun Sun, Zixuan Wang 0026, Xiaoyu Qin 0001, Tiezheng Ge, Bo Zheng 0007, Jia Jia 0001 |
ACM Multimedia | 4 |
| 2025 | PP-Motion: Physical-Perceptual Fidelity Evaluation for Human Motion Generation
Sihan Zhao, Zixuan Wang 0026, Tianyu Luan, Jia Jia 0001, Wentao Zhu 0004, Jiebo Luo 0001, Junsong Yuan 0001, Nan Xi |
ACM Multimedia | 2 |
| 2025 | HarmoniVox: Painting Voices to Match the Avatar's SoulabstractImagine James Bond speaking like Mr. Bean---such a mismatch would create a jarring dissonance and break the viewer's immersion. Current research on virtual avatar animation has focused on modeling 3D geometry, appearance, motion generation, however, neglecting the harmony between speech prosody and the avatar's visual presentation and contextual environment. In this paper, we seek to bridge this gap by firstly identifying and defining the key elements necessary for achieving audiovisual harmony, such as appearance, expression, body posture, backgrounds and colors. Subsequently, we propose a method that jointly models semantic consistency in avatar animation, named HarmoniVox, specifically on crafting prosodic speech consistent with the avatar's essence from given visual image. To achieve this, we implement a technical framework with a mutual modal contrastive learning strategy, enhancing multimodal alignment in a coarse-to-fine fashion. To support this method, we establish a experimental dataset HarAvaSpeech comprising 28,929 image-audio pairs, designed to encompass expressive speech prosody and rich avatar visual presentations across a wide range of contexts. Leveraging this dataset, our experiments demonstrate that the proposed method outperforms the baselines in manipulating the nuanced tone and harmonious rhythm of speech with the avatar visual presentations, and reveal generalizability on out-of-domain cases. Demo would be provided in https://harmonivox.github.io/harmonivox/. Songtao Zhou, Xiaoyu Qin 0001, Yixuan Zhou 0002, Qixin Wang 0002, Zeyu Jin, Zixuan Wang 0026, Zhiyong Wu 0001, Jia Jia 0001 |
ACM Multimedia | 6 |
| 2024 | DanceCamera3D: 3D Camera Movement Synthesis with Music and DanceabstractChoreographers determine what the dances look like, while cameramen determine the final presentation of dances. Recently, various methods and datasets have show-cased the feasibility of dance synthesis. However, camera movement synthesis with music and dance remains an un-solved challenging problem due to the scarcity of paired data. Thus, we present DCM, a new multi-modal 3D dataset, which for the first time combines camera movement with dance motion and music audio. This dataset encom-passes 108 dance sequences (3.2 hours) of paired dance-camera-music data from the anime community, covering 4 music genres. With this dataset, we uncover that dance camera movement is multifaceted and human-centric, and possesses multiple influencing factors, making dance camera synthesis a more challenging task compared to camera or dance synthesis alone. To overcome these difficulties, we propose DanceCamera3D, a transformer-based diffusion model that incorporates a novel body attention loss and a condition separation strategy. For evaluation, we devise new metrics measuring camera movement quality, diversity, and dancer fidelity. Utilizing these metrics, we conduct extensive experiments on our DCM dataset, providing both quantitative and qualitative evidence showcasing the effectiveness of our DanceCamera3D model. Code and video demos are available at https://github.com/Carmenw1203/DanceCamera3D-Official. Zixuan Wang 0026, Jia Jia 0001, Shikun Sun, Haozhe Wu, Jiaqing Zhou, Jiebo Luo 0001 |
CVPR | 1 |
| 2024 | Inner Classifier-Free Guidance and Its Taylor Expansion for Diffusion ModelsabstractClassifier-free guidance (CFG) is a pivotal technique for balancing the diversity and fidelity of samples in conditional diffusion models. This approach involves utilizing a single model to jointly optimize the conditional score predictor and unconditional score predictor, eliminating the need for additional classifiers. It delivers impressive results and can be employed for continuous and discrete condition representations. However, when the condition is continuous, it prompts the question of whether the trade-off can be further enhanced. Our proposed inner classifier-free guidance (ICFG) provides an alternative perspective on the CFG method when the condition has a specific structure, demonstrating that CFG represents a first-order case of ICFG. Additionally, we offer a second-order implementation, highlighting that even without altering the training policy, our second-order approach can introduce new valuable information and achieve an improved balance between fidelity and diversity for Stable Diffusion. Shikun Sun, Longhui Wei, Zhicai Wang, Zixuan Wang 0026, Junliang Xing, Jia Jia 0001, Qi Tian 0001 |
ICLR | 4 |
| 2024 | PlacidDreamer: Advancing Harmony in Text-to-3D GenerationabstractRecently, text-to-3D generation has attracted significant attention, resulting in notable performance enhancements. Previous methods utilize end-to-end 3D generation models to initialize 3D Gaussians, multi-view diffusion models to enforce multi-view consistency, and text-to-image diffusion models to refine details with score distillation algorithms. However, these methods exhibit two limitations. Firstly, they encounter conflicts in generation directions since different models aim to produce diverse 3D assets. Secondly, the issue of over-saturation in score distillation has not been thoroughly investigated and solved. To address these limitations, we propose PlacidDreamer, a text-to-3D framework that harmonizes initialization, multi-view generation, and text-conditioned generation with a single multi-view diffusion model, while simultaneously employing a novel score distillation algorithm to achieve balanced saturation. To unify the generation direction, we introduce the Latent-Plane module, a training-friendly plug-in extension that enables multi-view diffusion models to provide fast geometry reconstruction for initialization and enhanced multi-view images to personalize the text-to-image diffusion model. To address the over-saturation problem, we propose to view score distillation as a multi-objective optimization problem and introduce the Balanced Score Distillation algorithm, which offers a Pareto Optimal solution that achieves both rich details and balanced saturation. Extensive experiments validate the outstanding capabilities of our PlacidDreamer. The code is available at https://github.com/HansenHuang0823/PlacidDreamer. Shuo Huang 0005, Shikun Sun, Zixuan Wang 0026, Xiaoyu Qin 0001, Yanmin Xiong, Yuan Zhang 0020, Pengfei Wan 0001, Di Zhang 0026, Jia Jia 0001 |
ACM Multimedia | 3 |
| 2024 | DanceCamAnimator: Keyframe-Based Controllable 3D Dance Camera SynthesisabstractSynthesizing camera movements from music and dance is highly challenging due to the contradicting requirements and complexities of dance cinematography. Unlike human movements, which are always continuous, dance camera movements involve both continuous sequences of variable lengths and sudden drastic changes to simulate the switching of multiple cameras. However, in previous works, every camera frame is equally treated and this causes jittering and unavoidable smoothing in post-processing. To solve these problems, we propose to integrate animator dance cinematography knowledge by formulating this task as a three-stage process: keyframe detection, keyframe synthesis, and tween function prediction. Following this formulation, we design a novel end-to-end dance camera synthesis framework DanceCamAnimator, which imitates human animation procedures and shows powerful keyframe-based controllability with variable lengths. Extensive experiments on the DCM dataset demonstrate that our method surpasses previous baselines quantitatively and qualitatively. Code will be available at https://github.com/Carmenw1203/DanceCamAnimator-Official. Zixuan Wang 0026, Xiaoyu Qin 0001, Shikun Sun, Songtao Zhou, Jia Jia 0001, Jiebo Luo 0001 |
ACM Multimedia | 1 |
| 2023 | HoloSinger: Semantics and Music Driven Motion Generation with Octahedral Holographic ProjectionabstractLyrics and music are both significant for a singer to perform a song. Therefore, it is important in singer's motion generation to model both semantic and acoustic correlation with motions at the same time. In this paper, we propose HoloSinger, a novel comprehensive system that synthesizes singing motions according to the given song. Additionally, we present singing avatar with octahedral holographic projection. For singing motion generation, we introduce a Transformer-VAE generative model to decompose lyrics and music, then fuse their impacts to synthesize singer's motions. Extensive experiments and user studies show that our method automatically generates realistic motions that adhere to musical choreography and reflect the lyric semantics appropriately. Furthermore, we design a desktop-level holographic projection device with an octahedral structure. It achieves high-definition holographic projection effects with smaller volume, larger imaging area ratio, and the ability of real-time AI interaction. Zeyu Jin, Zixuan Wang 0026, Qixin Wang 0002, Jia Jia 0001, Ye Bai 0001, Yi Zhao 0006, Hao Li 0078 |
ACM Multimedia | 2 |
| 2022 | GroupDancer: Music to Multi-People Dance Synthesis with Style CollaborationabstractDifferent people dance in different styles. So when multiple people dance together, the phenomenon of style collaboration occurs: people need to seek common points while reserving differences in various dancing periods. Thus, we introduce a novel Music-driven Group Dance Synthesis task. Compared with single-people dance synthesis explored by most previous works, modeling the style collaboration phenomenon and choreographing for multiple people are more complicated and challenging. Moreover, the lack of sufficient records for conducting multi-people choreography in prior datasets further aggravates this problem. To address these issues, we construct a rich-annotated 3D Multi-Dancer Choreography dataset (MDC) and newly devise a metric SCEU for style collaboration evaluation. To our best knowledge, MDC is the first 3D dance dataset that collects both individual and collaborated music-dance pairs. Based on MDC, we present a novel framework, GroupDancer, consisting of three stages: Dancer Collaboration, Motion Choreography and Motion Transition. The Dancer Collaboration stage determines when and which dancers should collaborate their dancing styles from music. Afterward, the Motion Choreography stage produces a motion sequence for each dancer. Finally, the Motion Transition stage fills the gaps between the motions to achieve fluent and natural group dance. To make GroupDancer trainable from end to end and able to synthesize group dance with style collaboration, we propose mixed training and selective updating strategies. Comprehensive evaluations on the MDC dataset demonstrate that the proposed GroupDancer model can synthesize quite satisfactory group dance synthesis results with style collaboration. Zixuan Wang 0026, Jia Jia 0001, Haozhe Wu, Junliang Xing, Jinghe Cai, Guowen Chen |
ACM Multimedia | 1 |