VLDB 2026 Research / reviewers in the wild / expert
Chi Zhang 0067
dblp:91/195-67
· DBLP profile ↗
9ranked-venue papers
0as first author
7since 2021 · last 2025
0009-0002-3514-2490ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | StyleStudio: Text-Driven Style Transfer with Selective Control of Style ElementsabstractText-driven style transfer aims to merge the style of a reference image with content described by a text prompt. Recent advancements in text-to-image models have improved the nuance of style transformations, yet significant challenges remain, particularly with overfitting to reference styles, limiting stylistic control, and misaligning with textual content. In this paper, we propose three complementary strategies to address these issues. First, we introduce a cross-modal Adaptive Instance Normalization (AdaIN) mechanism for better integration of style and text features, enhancing alignment. Second, we develop a Style-based Classifier-Free Guidance (SCFG) approach that enables selective control over stylistic elements, reducing irrelevant influences. Finally, we incorporate a teacher model during early generation stages to stabilize spatial layouts and mitigate artifacts. Our extensive evaluations demonstrate significant improvements in style transfer quality and alignment with textual prompts. Furthermore, our approach can be integrated into existing style transfer frameworks without fine-tuning. Mingkun Lei, Beier Zhu, Hao Wang 0094, Chi Zhang 0067 |
CVPR | 5 |
| 2025 | Project-Probe-Aggregate: Efficient Fine-Tuning for Group RobustnessabstractWhile image-text foundation models have succeeded across diverse downstream tasks, they still face challenges in the presence of spurious correlations between the input and label. To address this issue, we propose a simple three-step approach-Project-Probe-Aggregate (PPA)-that enables parameter-efficient fine-tuning for foundation models without relying on group annotations. Building upon the failure-based debiasing scheme, our method, PPA, improves its two key components: minority samples identification and the robust training algorithm. Specifically, we first train biased classifiers by projecting image features onto the nullspace of class proxies from text encoders. Next, we infer group labels using the biased classifier and probe group targets with prior correction. Finally, we aggregate group weights of each class to produce the debiased classifier Our theoretical analysis shows that our PPA enhances minority group identification and is Bayes optimal for minimizing the balanced group error, mitigating spurious correlations. Extensive experimental results confirm the effectiveness of our PPA: it outperforms the state-of-the-art by an average worst-group accuracy while requiring less than 0.01% tunable parameters without training group labels. Beier Zhu, Jiequan Cui, Hanwang Zhang, Chi Zhang 0067 |
CVPR | 4 |
| 2025 | Distilling Parallel Gradients for Fast ODE Solvers of Diffusion ModelsabstractDiffusion models (DMs) have achieved state-of-the-art generative performance but suffer from high sampling latency due to their sequential denoising nature. Existing solver-based acceleration methods often face image quality degradation under a low-latency budget. In this paper, we propose the Ensemble Parallel Direction solver (dubbed as \ours), a novel ODE solver that mitigates truncation errors by incorporating multiple parallel gradient evaluations in each ODE step. Importantly, since the additional gradient computations are independent, they can be fully parallelized, preserving low-latency sampling. Our method optimizes a small set of learnable parameters in a distillation fashion, ensuring minimal training overhead. In addition, our method can serve as a plugin to improve existing ODE samplers. Extensive experiments on various image synthesis benchmarks demonstrate the effectiveness of our \ours~in achieving high-quality and low-latency sampling. For example, at the same latency level of 5 NFE, EPD achieves an FID of 4.47 on CIFAR-10, 7.97 on FFHQ, 8.17 on ImageNet, and 8.26 on LSUN Bedroom, surpassing existing learning-based solvers by a significant margin. Codes are available in https://github.com/BeierZhu/EPD. Beier Zhu, Hanwang Zhang, Chi Zhang 0067 |
ICCV | 5 |
| 2025 | EMMA: Your Text-to-Image Diffusion Model Can Secretly Accept Multi-Modal PromptsabstractRecent advancements in image generation have enabled the creation of high-quality images from diverse conditions, such as text and images. However, existing methods struggle to balance multiple conditions effectively, typically favoring certain modalities over others. To address this challenge, we introduce EMMA, a novel multi-modal image generation model that integrates both text and additional modalities, which we refer to as "text + X", to guide the generation of images. EMMA incorporates an innovative Multi-modal Feature Connector design, which effectively integrates textual and supplementary modal information through a special attention mechanism. Additionally, we propose a method for the assembly of existing modules to produce images conditioned on multiple modalities at the same time, eliminating the need for additional training. This modular nature also facilitates easy adaptation to different existing frameworks, making EMMA a flexible and effective tool for producing personalized and context-aware images. Extensive experiments demonstrate the effectiveness of EMMA in maintaining high fidelity and detail in generated images, showcasing its potential as a robust solution for advanced multi-modal conditional image generation tasks. Yucheng Han, Rui Wang 0099, Chi Zhang 0067, Hanwang Zhang |
IJCNN | 3 |
| 2025 | Uni-Inter: Unifying 3D Human Motion Synthesis Across Diverse Interaction ContextsabstractWe present Uni-Inter, a unified framework for human motion generation that supports a wide range of interaction scenarios: including human-human, human-object, and human-scene—within a single, task-agnostic architecture. In contrast to existing methods that rely on task-specific designs and exhibit limited generalization, Uni-Inter introduces the Unified Interactive Volume (UIV), a volumetric representation that encodes heterogeneous interactive entities into a shared spatial field. This enables consistent relational reasoning and compound interaction modeling. Motion generation is formulated as joint-wise probabilistic prediction over the UIV, allowing the model to capture fine-grained spatial dependencies and produce coherent, context-aware behaviors. Experiments across three representative interaction tasks demonstrate that Uni-Inter achieves competitive performance and generalizes well to novel combinations of entities. These results suggest that unified modeling of compound interactions offers a promising direction for scalable motion synthesis in complex environments. Sheng Liu 0013, Yuanzhi Liang, Jiepeng Wang 0005, Sidan Du, Chi Zhang 0067, Xuelong Li 0001 |
SIGGRAPH Asia | 5 |
| 2024 | StreamMOTP: Streaming and Unified Framework for Joint 3D Multi-Object Tracking and Trajectory Prediction
Jiaheng Zhuang, Guoan Wang, Siyu Zhang 0002, Xiyang Wang 0002, Hangning Zhou, Ziyao Xu 0003, Chi Zhang 0067, Zhiheng Li 0001 |
ACCV (2) | 7 |
| 2024 | Lever LM: Configuring In-Context Sequence to Lever Large Vision Language ModelsabstractAs Archimedes famously said, ``Give me a lever long enough and a fulcrum on which to place it, and I shall move the world'', in this study, we propose to use a tiny Language Model (LM), \eg, a Transformer with 67M parameters, to lever much larger Vision-Language Models (LVLMs) with 9B parameters. Specifically, we use this tiny \textbf{Lever-LM} to configure effective in-context demonstration (ICD) sequences to improve the In-Context Learinng (ICL) performance of LVLMs. Previous studies show that diverse ICD configurations like the selection and ordering of the demonstrations heavily affect the ICL performance, highlighting the significance of configuring effective ICD sequences. Motivated by this and by re-considering the the process of configuring ICD sequence, we find this is a mirror process of human sentence composition and further assume that effective ICD configurations may contain internal statistical patterns that can be captured by Lever-LM. Then a dataset with effective ICD sequences is constructed to train Lever-LM. After training, given novel queries, new ICD sequences are configured by the trained Lever-LM to solve vision-language tasks through ICL. Experiments show that these ICD sequences can improve the ICL performance of two LVLMs compared with some strong baselines in Visual Question Answering and Image Captioning, validating that Lever-LM can really capture the statistical patterns for levering LVLMs. The code is available at \url{https://anonymous.4open.science/r/Lever-LM-604A/}. Xu Yang 0021, Yingzhe Peng, Haoxuan Ma, Chi Zhang 0067, Yucheng Han, Hanwang Zhang |
NeurIPS | 5 |
| 2020 | Iterative Distance-Aware Similarity Matrix Convolution with Mutual-Supervised Point Elimination for Efficient Point Cloud Registration
Changhao Zhang, Ziyao Xu 0003, Hangning Zhou, Chi Zhang 0067 |
ECCV (24) | 5 |
| 2018 | Learning Unmanned Aerial Vehicle Control for Autonomous Target FollowingabstractWhile deep reinforcement learning (RL) methods have achieved unprecedented successes in a range of challenging problems, their applicability has been mainly limited to simulation or game domains due to the high sample complexity of the trial-and-error learning process. However, real-world robotic applications often need a data-efficient learning process with safety-critical constraints. In this paper, we consider the challenging problem of learning unmanned aerial vehicle (UAV) control for tracking a moving target. To acquire a strategy that combines perception and control, we represent the policy by a convolutional neural network. We develop a hierarchical approach that combines a model-free policy gradient method with a conventional feedback proportional-integral-derivative (PID) controller to enable stable learning without catastrophic failure. The neural network is trained by a combination of supervised learning from raw images and reinforcement learning from games of self-play. We show that the proposed approach can learn a target following policy in a simulator efficiently and the learned behavior can be successfully transferred to the DJI quadrotor platform for real-world UAV control. Tianbo Liu 0001, Chi Zhang 0067, Dit-Yan Yeung, Shaojie Shen |
IJCAI | 3 |