EDBT 2026 Demo / reviewers in the wild / expert
Zhicai Wang
dblp:250/1975
· DBLP profile ↗
15ranked-venue papers
4as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 2 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Res-Bench: Benchmarking the Robustness of Multimodal Large Language Models to Dynamic Resolution InputabstractMultimodal Large Language Models (MLLMs) increasingly support dynamic image resolutions. However, current evaluation paradigms primarily assess semantic performance, overlooking the critical question of resolution robustness - whether performance remains stable across varying input resolutions. To address this gap, we introduce Res-Bench, a comprehensive benchmark comprising 14,400 samples across 12 resolution levels and six core capability dimensions. We designed a novel evaluation framework that goes beyond traditional accuracy metrics to capture performance stability. This framework introduces multiple robustness metrics: Spearman's correlation for assessing resolution-performance trends, and Absolute/Relative Continuous Error (ACE/RCE) for measuring performance volatility. Using these metrics, we conducted a large-scale evaluation of leading MLLMs. Our analysis encompasses: (1) model-centric and task-centric robustness examination, (2) investigation of preprocessing strategies including padding and super-resolution, and (3) exploration of fine-tuning for stability enhancement. Chenxu Li, Zhicai Wang, Yuan Sheng, Yanbin Hao, Xiang Wang 0010 |
AAAI | 2 |
| 2026 | Causal-HalBench: Uncovering LVLMs Object Hallucinations Through Causal InterventionabstractLarge Vision-Language Models (LVLMs) often suffer from object hallucination, making erroneous judgments about the presence of objects in images. We propose this primarily stems from spurious correlations arising when models strongly associate highly co-occurring objects during training, leading to hallucinated objects influenced by visual context. Current benchmarks mainly focus on hallucination detection but lack a formal characterization and quantitative evaluation of spurious correlations in LVLMs. To address this, we introduce causal analysis into the object recognition scenario of LVLMs, establishing a Structural Causal Model (SCM). Utilizing the language of causality, we formally define spurious correlations arising from co-occurrence bias. To quantify the influence induced by these spurious correlations, we develop Causal-HalBench, a benchmark specifically constructed with counterfactual samples and integrated with comprehensive causal metrics designed to assess model robustness against spurious correlations. Concurrently, we propose an extensible pipeline for the construction of these counterfactual samples, leveraging the capabilities of proprietary LVLMs and Text-to-Image (T2I) models for their generation. Our evaluations on mainstream LVLMs using Causal-HalBench demonstrate these models exhibit susceptibility to spurious correlations, albeit to varying extents. Zhicai Wang, Junkang Wu, Jinda Lu, Xiang Wang 0010 |
AAAI | 2 |
| 2026 | Mitigating Hallucinations in Large Vision-Language Models without Performance DegradationabstractXingyu Zhu, Junfeng Fang, Shuo Wang, Beier Zhu, Zhicai Wang, Yonghui Yang, Xiangnan He. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Junfeng Fang, Shuo Wang 0008, Beier Zhu, Zhicai Wang, Yonghui Yang 0001, Xiangnan He 0001 |
ACL (1) | 5 |
| 2025 | BACON: Improving Clarity of Image Captions via Bag-of-Concept GraphsabstractAdvancements in large Vision-Language Models have brought precise, accurate image captioning, vital for advancing multi-modal image understanding and processing. Yet these captions often carry lengthy, intertwined contexts that are difficult to parse and frequently overlook essential cues, posing a great barrier for models like GroundingDINO and SDXL, which lack the strong text encoding and syntax analysis needed to fully leverage dense captions. To address this, we propose BACON, a prompting method that breaks down VLM-generated captions into disentangled, structured elements such as objects, relationships, styles, and themes. This approach not only minimizes confusion from handling complex contexts but also allows for efficient transfer into a JSON dictionary, enabling models without linguistic processing capabilities to easily access key information. We annotated 100,000 image-caption pairs using BACON with GPT-4V and trained an LLaVA captioner on this dataset, enabling it to produce BACON-style captions without relying on costly GPT-4V. Evaluations of overall quality, precision, and recall—as well as user studies—demonstrate that the resulting caption model consistently outperforms other SOTA VLM models in generating high-quality captions. Besides, we show that BACON-style captions exhibit better clarity when applied to various models, enabling them to accomplish previously unattainable tasks or surpass existing SOTA solutions without training. For example, BACON-style captions help GroundingDINO achieve 1.51× higher recall scores on open-vocabulary object detection tasks compared to leading methods. Zhantao Yang, Ruili Feng, Huangji Wang, Zhicai Wang, Shangwen Zhu, Han Zhang 0010, Jie Xiao 0002, Pingyu Wu, Kai Zhu 0004, Jixuan Chen, Chen-Wei Xie, Hongyang Zhang 0001, Yu Liu 0063, Fan Cheng 0002 |
CVPR | 5 |
| 2025 | Dynamic Multimodal Prototype Learning in Vision-Language ModelsabstractWith the increasing attention to pre-trained vision-language models (VLMs), \eg, CLIP, substantial efforts have been devoted to many downstream tasks, especially in test-time adaptation (TTA). However, previous works focus on learning prototypes only in the textual modality while overlooking the ambiguous semantics in class names. These ambiguities lead to textual prototypes that are insufficient to capture visual concepts, resulting in limited performance. To address this issue, we introduce \textbf{ProtoMM}, a training-free framework that constructs multimodal prototypes to adapt VLMs during the test time. By viewing the prototype as a discrete distribution over the textual descriptions and visual particles, ProtoMM has the ability to combine the multimodal features for comprehensive prototype learning. More importantly, the visual particles are dynamically updated as the testing stream flows. This allows our multimodal prototypes to continually learn from the data, enhancing their generalizability in unseen scenarios. In addition, we quantify the importance of the prototypes and test images by formulating their semantic distance as an optimal transport problem. Extensive experiments on 15 zero-shot benchmarks demonstrate the effectiveness of our method, achieving a 1.03\% average accuracy improvement over state-of-the-art methods on ImageNet and its variant datasets. Shuo Wang 0008, Beier Zhu, Miaoge Li, Junfeng Fang, Zhicai Wang, Dongsheng Wang 0003, Hanwang Zhang |
ICCV | 7 |
| 2025 | UniSVG: A Unified Dataset for Vector Graphic Understanding and Generation with Multimodal Large Language ModelsabstractUnlike bitmap images, scalable vector graphics (SVG) maintain quality when scaled, frequently employed in computer vision and artistic design in the representation of SVG code. In this era of proliferating AI-powered systems, enabling AI to understand and generate SVG has become increasingly urgent. However, AI-driven SVG understanding and generation (U&G) remain significant challenges. SVG code, equivalent to a set of curves and lines controlled by floating-point parameters, demands high precision in SVG U&G. Besides, SVG generation operates under diverse conditional constraints, including textual prompts and visual references, which requires powerful multi-modal processing for condition-to-SVG transformation. Recently, the rapid growth of Multi-modal Large Language Models (MLLMs) have demonstrated capabilities to process multi-modal inputs and generate complex vector controlling parameters, suggesting the potential to address SVG U&G tasks within a unified model. To unlock MLLM's capabilities in the SVG area, we propose an SVG-centric dataset called UniSVG, comprising 525k data items, tailored for MLLM training and evaluation. To our best knowledge, it is the first comprehensive dataset designed for unified SVG generation (from textual prompts and images) and SVG understanding (color, category, usage, etc.). As expected, learning on the proposed dataset boosts open-source MLLMs' performance on various SVG U&G tasks, surpassing SOTA close-source MLLMs like GPT-4V. We release dataset, benchmark, weights, codes and experiment details on https://ryanlijinke.github.io/. Jiarui Yu, Chenxing Wei, Hande Dong, Liangjing Yang, Zhicai Wang, Yanbin Hao |
ACM Multimedia | 7 |
| 2025 | The Matrix: Infinite-Horizon World Generation with Real-Time Moving ControlabstractWe present The Matrix, a foundational realistic world simulator capable of generating infinitely long 720p high-fidelity real-scene video streams with real-time, responsive control in both first- and third-person perspectives. Trained on limited supervised data from video games like Forza Horizon 5 and Cyberpunk 2077, complemented by large-scale unsupervised footage from real-world settings like Tokyo streets, The Matrix allows users to traverse diverse terrains—deserts, grasslands, water bodies, and urban landscapes—in continuous, uncut hour-long sequences. With speeds of up to 16 FPS, the system supports real-time interactivity and demonstrates zero-shot generalization, translating virtual game environments to real-world contexts where collecting continuous movement data is often infeasible. For example, The Matrix can simulate a BMW X3 driving through an office setting—an environment present in neither gaming data nor real-world sources. This approach showcases the potential of game data to advance robust world models, bridging the gap between simulations and real-world applications in scenarios with limited data. Ruili Feng, Han Zhang 0010, Zhilei Shu, Zhantao Yang, Longxiang Tang, Zhicai Wang, Andy Zheng, Jie Xiao 0002, Ruihang Chu, Yu Liu 0063, Hongyang Zhang 0001 |
NeurIPS | 6 |
| 2025 | Prior Preserved Text-to-Image Personalization Without Image RegularizationabstractThe current state-of-the-art text-to-image (T2I) models have found numerous applications, driven by their ability to produce photorealistic images. Concept learning, as one notable application, aims to enable T2I models to generate personalized content and better enable users to create images according to their interests. Nevertheless, the process of concept learning often involves model fine-tuning, which in turn brings the potential risk of overfitting. Such overfitting causes the T2I model to have reduced output diversity and results in poor editability. To mitigate the overfitting problem, we introduce two simple yet effective designs, namely masked textual inversion (MaskTI) and text regularization (TextReg). MaskTI is a variant of vanilla textual inversion that forces the learnable identifier to only attend to the class descriptor. This modification can effectively reduce the overfitting to those uninterested backgrounds. TextReg regulates the fine-tuning of cross-attention modules with simple text prompts without identifiers, which avoids the usage of real images as the regularization prior. Our extensive experiments demonstrate that not only does our approach effectively protect prior knowledge but also has high editability for the personalized model. Zhicai Wang, Ouxiang Li, Longhui Wei, Yanbin Hao, Xiang Wang 0010, Qi Tian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Enhance Image Classification via Inter-Class Image Mixup with Diffusion ModelabstractText-to-image (T2I) generative models have recently emerged as a powerful tool, enabling the creation of photo-realistic images and giving rise to a multitude of applications. However, the effective integration of T2I models into fundamental image classification tasks remains an open question. A prevalent strategy to bolster image classification performance is through augmenting the training set with synthetic images generated by T2I models. In this study, we scrutinize the shortcomings of both current gener-ative and conventional data augmentation techniques. Our analysis reveals that these methods struggle to produce images that are both faithful (in terms of foreground objects) and diverse (in terms of background contexts) for domain-specific concepts. To tackle this challenge, we introduce an innovative inter-class data augmentation method known as Diff-Mix11https://github.com/Zhicaiwww/Diff-Mix, which enriches the dataset by performing image translations between classes. Our empirical results demonstrate that Diff-Mix achieves a better balance between faith-fulness and diversity, leading to a marked improvement in performance across diverse image classification scenarios, including few-shot, conventional, and long-tail classifications for domain-specific datasets. Zhicai Wang, Longhui Wei, Heyu Chen, Yanbin Hao, Xiang Wang 0010, Xiangnan He 0001, Qi Tian 0001 |
CVPR | 1 |
| 2024 | Inner Classifier-Free Guidance and Its Taylor Expansion for Diffusion ModelsabstractClassifier-free guidance (CFG) is a pivotal technique for balancing the diversity and fidelity of samples in conditional diffusion models. This approach involves utilizing a single model to jointly optimize the conditional score predictor and unconditional score predictor, eliminating the need for additional classifiers. It delivers impressive results and can be employed for continuous and discrete condition representations. However, when the condition is continuous, it prompts the question of whether the trade-off can be further enhanced. Our proposed inner classifier-free guidance (ICFG) provides an alternative perspective on the CFG method when the condition has a specific structure, demonstrating that CFG represents a first-order case of ICFG. Additionally, we offer a second-order implementation, highlighting that even without altering the training policy, our second-order approach can introduce new valuable information and achieve an improved balance between fidelity and diversity for Stable Diffusion. Shikun Sun, Longhui Wei, Zhicai Wang, Zixuan Wang 0026, Junliang Xing, Jia Jia 0001, Qi Tian 0001 |
ICLR | 3 |
| 2024 | PosMLP-Video: Spatial and Temporal Relative Position Encoding for Efficient Video Recognition
Yanbin Hao, Diansong Zhou, Zhicai Wang, Chong-Wah Ngo, Meng Wang 0001 |
Int. J. Comput. Vis. | 3 |
| 2023 | Bi-Directional Distribution Alignment for Transductive Zero-Shot LearningabstractZero-shot learning (ZSL) suffers intensely from the domain shift issue, i.e., the mismatch (or misalignment) between the true and learned data distributions for classes without training data (unseen classes). By learning additionally from unlabelled data collected for the unseen classes, transductive ZSL (TZSL) could reduce the shift but only to a certain extent. To improve TZSL, we propose a novel approach Bi-VAEGAN which strengthens the distribution alignment between the visual space and an auxiliary space. As a result, it can reduce largely the domain shift. The proposed key designs include (1) a bi-directional distribution alignment, (2) a simple but effective L2-norm based feature normalization approach, and (3) a more sophisticated unseen class prior estimation. Evaluated by four benchmark datasets, Bi-VAEGAN11Code is available at https://github.com/Zhicaiwww/Bi-VAEGAN achieves the new state of the art under both the standard and generalized TZSL settings. Zhicai Wang, Yanbin Hao, Tingting Mu, Ouxiang Li, Shuo Wang 0008, Xiangnan He 0001 |
CVPR | 1 |
| 2023 | Backdoor Defense via Deconfounded Representation LearningabstractDeep neural networks (DNNs) are recently shown to be vulnerable to backdoor attacks, where attackers embed hidden backdoors in the DNN model by injecting a few poisoned examples into the training dataset. While extensive efforts have been made to detect and remove backdoors from backdoored DNNs, it is still not clear whether a backdoor-free clean model can be directly obtained from poisoned datasets. In this paper, we first construct a causal graph to model the generation process of poisoned data and find that the backdoor attack acts as the confounder, which brings spurious associations between the input images and target labels, making the model predictions less reliable. Inspired by the causal understanding, we propose the Causality-inspired Backdoor Defense (CBD), to learn deconfounded representations for reliable classification. Specifically, a backdoored model is intentionally trained to capture the confounding effects. The other clean model dedicates to capturing the desired causal effects by minimizing the mutual information with the confounding representations from the backdoored model and employing a sample-wise re-weighting scheme. Extensive experiments on multiple benchmark datasets against 6 state-of-the-art attacks verify that our proposed defense method is effective in reducing backdoor threats while maintaining high accuracy in predicting benign samples. Further analysis shows that CBD can also resist potential adaptive attacks. The code is available at https://github.com/zaixizhang/CBD. Zaixi Zhang, Qi Liu 0003, Zhicai Wang, Zepu Lu, Qingyong Hu |
CVPR | 3 |
| 2023 | Generate What You Prefer: Reshaping Sequential Recommendation via Guided DiffusionabstractSequential recommendation aims to recommend the next item that matches a user’s
interest, based on the sequence of items he/she interacted with before. Scrutinizing
previous studies, we can summarize a common learning-to-classify paradigm—
given a positive item, a recommender model performs negative sampling to add
negative items and learns to classify whether the user prefers them or not, based on
his/her historical interaction sequence. Although effective, we reveal two inherent
limitations: (1) it may differ from human behavior in that a user could imagine
an oracle item in mind and select potential items matching the oracle; and (2)
the classification is limited in the candidate pool with noisy or easy supervision
from negative samples, which dilutes the preference signals towards the oracle
item. Yet, generating the oracle item from the historical interaction sequence is
mostly unexplored. To bridge the gap, we reshape sequential recommendation
as a learning-to-generate paradigm, which is achieved via a guided diffusion
model, termed DreamRec. Specifically, for a sequence of historical items, it
applies a Transformer encoder to create guidance representations. Noising target
items explores the underlying distribution of item space; then, with the guidance of
historical interactions, the denoising process generates an oracle item to recover
the positive item, so as to cast off negative sampling and depict the true preference
of the user directly. We evaluate the effectiveness of DreamRec through extensive
experiments and comparisons with existing methods. Codes and data are open-sourced
at https://github.com/YangZhengyi98/DreamRec. Zhengyi Yang 0007, Jiancan Wu, Zhicai Wang, Xiang Wang 0010, Yancheng Yuan, Xiangnan He 0001 |
NeurIPS | 3 |
| 2022 | Parameterization of Cross-token Relations with Relative Positional Encoding for Vision MLPabstractVision multi-layer perceptrons (MLPs) have shown promising performance in computer vision tasks, and become the main competitor of CNNs and vision Transformers. They use token-mixing layers to capture cross-token interactions, as opposed to the multi-head self-attention mechanism used by Transformers. However, the heavily parameterized token-mixing layers naturally lack mechanisms to capture local information and multi-granular non-local relations, thus their discriminative power is restrained. To tackle this issue, we propose a new positional spacial gating unit (PoSGU). It exploits the attention formulations used in the classical relative positional encoding (RPE), to efficiently encode the cross-token relations for token mixing. It can successfully reduce the current quadratic parameter complexity O(N2) of vision MLPs to $O(N)$ and O(1). We experiment with two RPE mechanisms, and further propose a group-wise extension to improve their expressive power with the accomplishment of multi-granular contexts. These then serve as the key building blocks of a new type of vision MLP, referred to as PosMLP. We evaluate the effectiveness of the proposed approach by conducting thorough experiments, demonstrating an improved or comparable performance with reduced parameter complexity. For instance, for a model trained on ImageNet1K, we achieve a performance improvement from 72.14% to 74.02% and a learnable parameter reduction from 19.4M to 18.2M. Code could be found at https://github.com/Zhicaiwww/PosMLP https://github.com/Zhicaiwww/PosMLP. Zhicai Wang, Yanbin Hao, Xingyu Gao 0001, Hao Zhang 0047, Shuo Wang 0008, Tingting Mu, Xiangnan He 0001 |
ACM Multimedia | 1 |