EDBT 2026 Demo / reviewers in the wild / expert
Zunnan Xu
dblp:352/5335
· DBLP profile ↗
18ranked-venue papers
4as first author
18since 2021 · last 2026
0000-0001-5586-4971ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 14 since 2021Artificial intelligence and machine learning · 13 · 4 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Zo3T: Zero-Shot 3D-Aware Trajectory-Guided Image-to-Video Generation via Test-Time TrainingabstractTrajectory-Guided image-to-video (I2V) generation aims to synthesize videos that adhere to user-specified motion instructions. Existing methods typically rely on computationally expensive fine-tuning on scarce annotated datasets. Although some zero-shot methods attempt to trajectory control in the latent space, they may yield unrealistic motion by neglecting 3D perspective and creating a misalignment between the manipulated latents and the network's noise predictions. To address these challenges, we introduce Zo3T, a novel zero-shot test-time-training framework for trajectory-guided generation with three core innovations: First, we incorporate a 3D-Aware Kinematic Projection, leveraging inferring scene depth to derive perspective-correct affine transformations for target regions. Second, we introduce Trajectory-Guided Test-Time LoRA, a mechanism that dynamically injects and optimizes ephemeral LoRA adapters into the denoising network alongside the latent state. Driven by a regional feature consistency loss, this co-adaptation effectively enforces motion constraints while allowing the pre-trained model to locally adapt its internal representations to the manipulated latent, thereby ensuring generative fidelity and on-manifold adherence. Finally, we develop Guidance Field Rectification, which refines the denoising evolutionary path by optimizing the conditional guidance field through a one-step lookahead strategy, ensuring efficient generative progression towards the target trajectory. Zo3T significantly enhances 3D realism and motion accuracy in trajectory-controlled I2V generation, demonstrating superior performance over existing training-based and zero-shot approaches. Ruicheng Zhang, Zunnan Xu, Zihao Liu 0006, Jiehui Huang, Xiu Li 0001 |
AAAI | 3 |
| 2025 | Densely Connected Parameter-Efficient Tuning for Referring Image SegmentationabstractIn the domain of computer vision, Parameter-Efficient Tuning (PET) is increasingly replacing the traditional paradigm of pre-training followed by full fine-tuning. PET is particularly favored for its effectiveness in large foundation models, as it streamlines transfer learning costs and optimizes hardware utilization. However, the current PET methods are mainly designed for single-modal optimization. While some pioneering studies have undertaken preliminary explorations, they still remain at the level of aligned encoders (e.g., CLIP) and lack exploration of misaligned encoders. These methods show sub-optimal performance with misaligned encoders, as they fail to effectively align the multimodal features during fine-tuning. In this paper, we introduce DETRIS, a parameter-efficient tuning framework designed to enhance low-rank visual feature propagation by establishing dense interconnections between each layer and all preceding layers, which enables effective cross-modal feature interaction and adaptation to misaligned encoders. We also suggest using text adapters to improve textual features. Our simple yet efficient approach greatly surpasses state-of-the-art methods with 0.9% to 1.8% backbone parameter updates, evaluated on challenging benchmarks. Jiaqi Huang 0003, Zunnan Xu, Haonan Han, Kehong Yuan, Xiu Li 0001 |
AAAI | 2 |
| 2025 | TGS-LGP: Text-Guided Medical Image Segmentation Via Local-Global PerceptionabstractIn this paper, we propose TGS-LGP, a framework for text-guided medical image segmentation that synergistically combines cross-modal understanding with local-global features perception. Our method employs a local-global text encoder that simultaneously processes clinical language inputs through two complementary pathways: one leverages contrastive language-image pretraining to capture global semantic context, and another employes domain-specific linguistic analysis to extract fine-grained lesion characteristics. These multiscale textual representations are progressively fused with refined visual features through an attention-based decoder that dynamically fuse cross-modal features. This architectural design enables context-aware integration of diagnostic text guidance with radiological image patterns. Additionally, we introduce vision refiner, a novel component that effectively integrates global features of medical images with minimal computational cost, compensating for the information loss caused by the local modeling bias of CNNs in medical imaging. This enhancement strengthens visual representation, thereby facilitating more effective fusion with textual features. Experimental results on two widely used benchmarks demonstrate that our method achieves the best accuracy compared to state-of-the-art methods. Jiaqi Huang 0003, Zunnan Xu, Mingwen Ou, Sen Zeng, Kehong Yuan |
BIBM | 3 |
| 2025 | AToM: Aligning Text-to-Motion Model at Event-Level with GPT-4Vision RewardabstractRecently, text-to-motion models have opened new possibilities for creating realistic human motion with greater efficiency and flexibility. However, aligning motion generation with event-level textual descriptions presents unique challenges due to the complex relationship between textual prompts and desired motion outcomes. To address this, we introduce AToM, a framework that enhances the alignment between generated motion and text prompts by leveraging reward from GPT-4Vision. AToM comprises three main stages: Firstly, we construct a dataset MotionPreferthat pairs three types of event-level textual prompts with generated motions, which cover the integrity, temporal relationship and frequency of motion. Secondly, we design a paradigm that utilizes GPT-4Vision for detailed motion annotation, including visual data formatting, task-specific instructions and scoring rules for each sub-task. Finally, we fine-tune an existing text-to-motion model using reinforcement learning guided by this paradigm. Experimental results demonstrate that AToM significantly improves the event-level alignment quality of text-to-motion generation. Project page is available at https://atom-motion.github.io/. Haonan Han, Xiangzuo Wu, Huan Liao, Zunnan Xu, Zhongyuan Hu, Ronghui Li, Yachao Zhang 0001, Xiu Li 0001 |
CVPR | 4 |
| 2025 | HunyuanPortrait: Implicit Condition Control for Enhanced Portrait AnimationabstractWe introduce HunyuanPortrait, a diffusion-based condition control method that employs implicit representations for highly controllable and lifelike portrait animation. Given a single portrait image as an appearance reference and video clips as driving templates, HunyuanPortrait can animate the character in the reference image by the facial expression and head pose of the driving videos. In our framework, we utilize pre-trained encoders to achieve the decoupling of portrait motion information and identity in videos. To do so, implicit representation is adopted to encode motion information and is employed as control signals in the animation phase. By leveraging the power of stable video diffusion as the main building block, we carefully design adapter layers to inject control signals into the denoising unet through attention mechanisms. These bring spatial richness of details and temporal consistency. HunyuanPortrait also exhibits strong generalization performance, which can effectively disentangle appearance and motion under different image styles. Our framework outperforms existing methods, demonstrating superior temporal consistency and controllability. Our project is available at HunyuanPortrait. Zunnan Xu, Zhentao Yu, Xiaoyu Jin, Fa-Ting Hong, Xiaozhong Ji, Chengfei Cai, Shiyu Tang, Qin Lin 0003, Xiu Li 0001, Qinglin Lu |
CVPR | 1 |
| 2025 | FireEdit: Fine-grained Instruction-based Image Editing via Region-aware Vision Language ModelabstractCurrently, instruction-based image editing methods have made significant progress by leveraging the powerful cross-modal understanding capabilities of vision language models (VLMs). However, they still face challenges in three key areas: 1) complex scenarios; 2) semantic consistency; and 3) fine-grained editing. To address these issues, we propose FireEdit, an innovative Fine-grained Instruction-based image editing framework that exploits a REgion-aware VLM. FireEdit is designed to accurately comprehend user instructions and ensure effective control over the editing process. Specifically, we enhance the fine-grained visual perception capabilities of the VLM by introducing additional region tokens. Relying solely on the output of the LLM to guide the diffusion model may lead to suboptimal editing results. Therefore, we propose a Time-Aware Target Injection module and a Hybrid Visual Cross Attention module. The former dynamically adjusts the guidance strength at various denoising stages by integrating timestep embeddings with the text embeddings. The latter enhances visual details for image editing, thereby preserving semantic consistency between the edited result and the source image. By combining the VLM enhanced with fine-grained region tokens and the time-dependent diffusion model, FireEdit demonstrates significant advantages in comprehending editing instructions and maintaining high semantic consistency. Extensive experiments indicate that our approach surpasses the state-of-the-art instruction-based image editing methods. Jiahao Li 0001, Zunnan Xu, Yiji Cheng, Fa-Ting Hong, Qin Lin 0003, Qinglin Lu, Xiaodan Liang |
CVPR | 3 |
| 2025 | REPARO: Compositional 3D Assets Generation with Differentiable 3D Layout AlignmentabstractTraditional image-to-3D models often struggle with scenes containing multiple objects due to biases and occlusion complexities. To address this challenge, we present REPARO, a novel approach for compositional 3D asset generation from single images. REPARO employs a two-step process: first, it extracts individual objects from the scene and reconstructs their 3D meshes using off-the-shelf image-to-3D models; then, it optimizes the layout of these meshes through differentiable rendering techniques, ensuring coherent scene composition. By integrating optimal transport-based long-range appearance loss term and high-level semantic loss term in the differentiable rendering, REPARO can effectively recover the layout of 3D assets. The proposed method can significantly enhance object independence, detail accuracy, and overall scene coherence. Extensive evaluation of multi-object scenes demonstrates that our REPARO offers a comprehensive approach to address the complexities of multi-object 3D scene generation from single images. Haonan Han, Rui Yang 0010, Huan Liao, Jiankai Xing, Zunnan Xu, Xiaoming Yu, Junwei Zha, Xiu Li 0001, Wanhua Li 0001 |
ICCV | 5 |
| 2025 | Audio-Visual Controlled Video Diffusion with Masked Selective State Spaces Modeling for Natural Talking Head GenerationabstractTalking head synthesis is vital for virtual avatars and human-computer interaction. However, most existing methods are typically limited to accepting control from a single primary modality, restricting their practical utility. To this end, we introduce \textbf{ACTalker}, an end-to-end video diffusion framework that supports both multi-signals control and single-signal control for talking head video generation. For multiple control, we design a parallel mamba structure with multiple branches, each utilizing a separate driving signal to control specific facial regions. A gate mechanism is applied across all branches, providing flexible control over video generation. To ensure natural coordination of the controlled video both temporally and spatially, we employ the mamba structure, which enables driving signals to manipulate feature tokens across both dimensions in each branch. Additionally, we introduce a mask-drop strategy that allows each driving signal to independently control its corresponding facial region within the mamba structure, preventing control conflicts. Experimental results demonstrate that our method produces natural-looking facial videos driven by diverse signals and that the mamba layer seamlessly integrates multiple driving modalities without conflict. The project website can be found at https://harlanhong.github.io/publications/actalker/index.html. Fa-Ting Hong, Zunnan Xu, Xiu Li 0001, Qin Lin 0003, Qinglin Lu, Dan Xu 0002 |
ICCV | 2 |
| 2025 | InterAnimate: Taming Region-Aware Diffusion Model for Realistic Human Interaction Animation
Yukang Lin, Yan Hong 0001, Zunnan Xu, Xindi Li, Chuanbiao Song, Ronghui Li, Haoxing Chen, Jun Lan 0001, Huijia Zhu, Weiqiang Wang 0002, Jianfu Zhang 0003, Xiu Li 0001 |
ACM Multimedia | 3 |
| 2025 | Separate to Collaborate: Dual-Stream Diffusion Model for Coordinated Piano Hand Motion Synthesis
Zihao Liu 0006, Mingwen Ou, Zunnan Xu, Jiaqi Huang 0003, Haonan Han, Ronghui Li, Xiu Li 0001 |
ACM Multimedia | 3 |
| 2025 | SAM-R1: Leveraging SAM for Reward Feedback in Multimodal Segmentation via Reinforcement LearningabstractLeveraging multimodal large models for image segmentation has become a prominent research direction. However, existing approaches typically rely heavily on manually annotated datasets that include explicit reasoning processes, which are costly and time-consuming to produce. Recent advances suggest that reinforcement learning (RL) can endow large models with reasoning capabilities without requiring such reasoning-annotated data. In this paper, we propose SAM-R1, a novel framework that enables multimodal large models to perform fine-grained reasoning in image understanding tasks. Our approach is the first to incorporate fine-grained segmentation settings during the training of multimodal reasoning models. By integrating task-specific, fine-grained rewards with a tailored optimization objective, we further enhance the model's reasoning and segmentation alignment. We also leverage the Segment Anything Model (SAM) as a strong and flexible reward provider to guide the learning process. With only 3k training samples, SAM-R1 achieves strong performance across multiple benchmarks, demonstrating the effectiveness of reinforcement learning in equipping multimodal models with segmentation-oriented reasoning capabilities. Jiaqi Huang 0003, Zunnan Xu, Yicheng Xiao, Mingwen Ou, Xiu Li 0001, Kehong Yuan |
NeurIPS | 2 |
| 2024 | Chain of Generation: Multi-Modal Gesture Synthesis via Cascaded Conditional ControlabstractThis study aims to improve the generation of 3D gestures by utilizing multimodal information from human speech. Previous studies have focused on incorporating additional modalities to enhance the quality of generated gestures. However, these methods perform poorly when certain modalities are missing during inference. To address this problem, we suggest using speech-derived multimodal priors to improve gesture generation. We introduce a novel method that separates priors from speech and employs multimodal priors as constraints for generating gestures. Our approach utilizes a chain-like modeling method to generate facial blendshapes, body movements, and hand gestures sequentially. Specifically, we incorporate rhythm cues derived from facial deformation and stylization prior based on speech emotions, into the process of generating gestures. By incorporating multimodal priors, our method improves the quality of generated gestures and eliminate the need for expensive setup preparation during inference. Extensive experiments and user studies confirm that our proposed approach achieves state-of-the-art performance. Zunnan Xu, Yachao Zhang 0001, Ronghui Li, Xiu Li 0001 |
AAAI | 1 |
| 2024 | MaPPER: Multimodal Prior-guided Parameter Efficient Tuning for Referring Expression ComprehensionabstractReferring Expression Comprehension (REC), which aims to ground a local visual region via natural language, is a task that heavily relies on multimodal alignment.Most existing methods utilize powerful pre-trained models to transfer visual/linguistic knowledge by full fine-tuning.However, full fine-tuning the entire backbone not only breaks the rich prior knowledge embedded in the pre-training, but also incurs significant computational costs.Motivated by the recent emergence of Parameter-Efficient Transfer Learning (PETL) methods, we aim to solve the REC task in an effective and efficient manner.Directly applying these PETL methods to the REC task is inappropriate, as they lack the specific-domain abilities for precise local visual perception and visual-language alignment.Therefore, we propose a novel framework of Multimodal Prior-guided Parameter Efficient Tuning, namely MaPPER.Specifically, MaPPER comprises Dynamic Prior Adapters guided by an aligned prior, and Local Convolution Adapters to extract precise local semantics for better visual perception.Moreover, the Prior-Guided Text module is proposed to further utilize the prior for facilitating the cross-modal alignment.Experimental results on three widely-used benchmarks demonstrate that MaPPER achieves the best accuracy compared to the full fine-tuning and other PETL methods with only 1.41% tunable backbone parameters.Our code is available at https://github.com/liuting20/MaPPER. Zunnan Xu, Liangtao Shi, Quanjun Yin |
EMNLP | 2 |
| 2024 | FreeTalker: Controllable Speech and Text-Driven Gesture Generation Based on Diffusion Models for Enhanced Speaker NaturalnessabstractCurrent talking avatars mostly generate co-speech gestures based on audio and text of the utterance, without considering the non-speaking motion of the speaker. Furthermore, previous works on co-speech gesture generation have designed network structures based on individual gesture datasets, which results in limited data volume, compromised generalizability, and restricted speaker movements. To tackle these issues, we introduce FreeTalker, which, to the best of our knowledge, is the first framework for the generation of both spontaneous (e.g., co-speech gesture) and non-spontaneous (e.g., moving around the podium) speaker motions. Specifically, we train a diffusion-based model for speaker motion generation that employs unified representations of both speech-driven gestures and text-driven motions, utilizing heterogeneous data sourced from various motion datasets. During inference, we utilize classifier-free guidance to highly control the style in the clips. Additionally, to create smooth transitions between clips, we utilize DoubleTake, a method that leverages a generative prior and ensures seamless motion blending. Extensive experiments show that our method generates natural and controllable speaker movements. Our code, model, and demo are are available at https://youngseng.github.io/FreeTalker/. Zunnan Xu, Haiwei Xue, Yongkang Cheng, Shaoli Huang, Mingming Gong, Zhiyong Wu 0001 |
ICASSP | 2 |
| 2024 | BATON: Aligning Text-to-Audio Model Using Human Preference Feedback
Huan Liao, Haonan Han, Kai Yang 0050, Tianjiao Du, Rui Yang 0010, Zunnan Xu, Jiasheng Lu, Xiu Li 0001 |
IJCAI | 7 |
| 2024 | Consistent123: One Image to Highly Consistent 3D Asset Using Case-Aware Diffusion PriorsabstractReconstructing 3D objects from a single image guided by pretrained diffusion models has demonstrated promising outcomes. However, due to utilizing the case-agnostic rigid strategy, their generalization ability to arbitrary cases and the 3D consistency of reconstruction are still poor. In this work, we propose Consistent123, a case-aware two-stage method for highly consistent 3D asset reconstruction from one image with both 2D and 3D diffusion priors. In the first stage, Consistent123 utilizes only 3D structural priors for sufficient geometry exploitation, with a CLIP-based case-aware adaptive detection mechanism embedded within this process. In the second stage, 2D texture priors are introduced and progressively take on a dominant guiding role, delicately sculpting the details of the 3D model. Consistent123 aligns more closely with the evolving trends in guidance requirements, adaptively providing adequate 3D geometric initialization and suitable 2D texture refinement for different objects. Consistent123 can obtain highly 3D-consistent reconstruction and exhibits strong generalization ability across various objects. Qualitative and quantitative experiments show that our method significantly outperforms state-of-the-art image-to-3D methods. Yukang Lin, Haonan Han, Chaoqun Gong, Zunnan Xu, Yachao Zhang 0001, Xiu Li 0001 |
ACM Multimedia | 4 |
| 2024 | MambaTalk: Efficient Holistic Gesture Synthesis with Selective State Space ModelsabstractGesture synthesis is a vital realm of human-computer interaction, with wide-ranging applications across various fields like film, robotics, and virtual reality.
Recent advancements have utilized the diffusion model to improve gesture synthesis.
However, the high computational complexity of these techniques limits the application in reality.
In this study, we explore the potential of state space models (SSMs).
Direct application of SSMs in gesture synthesis encounters difficulties, which stem primarily from the diverse movement dynamics of various body parts.
The generated gestures may also exhibit unnatural jittering issues.
To address these, we implement a two-stage modeling strategy with discrete motion priors to enhance the quality of gestures.
Built upon the selective scan mechanism, we introduce MambaTalk, which integrates hybrid fusion modules, local and global scans to refine latent space representations.
Subjective and objective experiments demonstrate that our method surpasses the performance of state-of-the-art models. Our project is publicly available at~\url{https://kkakkkka.github.io/MambaTalk/}. Zunnan Xu, Yukang Lin, Haonan Han, Ronghui Li, Yachao Zhang 0001, Xiu Li 0001 |
NeurIPS | 1 |
| 2023 | Bridging Vision and Language Encoders: Parameter-Efficient Tuning for Referring Image SegmentationabstractParameter Efficient Tuning (PET) has gained attention for reducing the number of parameters while maintaining performance and providing better hardware resource savings, but few studies investigate dense prediction tasks and interaction between modalities. In this paper, we do an investigation of efficient tuning problems on referring image segmentation. We propose a novel adapter called Bridger to facilitate cross-modal information exchange and inject task-specific information into the pre-trained model. We also design a lightweight decoder for image segmentation. Our approach achieves comparable or superior performance with only 1.61% to 3.38% backbone parameter updates, evaluated on challenging benchmarks. The code is available at https://github.com/kkakkkka/ETRIS. Zunnan Xu, Yong Zhang 0034, Yibing Song, Guanbin Li |
ICCV | 1 |