VLDB 2026 Research / reviewers in the wild / expert
Huanjia Zhu
dblp:389/2704
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BayesVQA: Energy-Guided Bayesian Debiasing for Language-Bias-Robust Visual Question AnsweringabstractNumerous studies have demonstrated that Visual Question Answering (VQA) models are vulnerable to language priors and dataset biases, often leading to spurious correlations between questions and answers. As a result, these models excessively rely on linguistic cues, neglecting essential visual information and causing representational distortions. To address this challenge, we propose a novel Bayesian debiasing framework termed BayesVQA, which integrates three carefully designed mechanisms: Energy-guided Prior Variance (EPV), Energy-guided Posterior Sampling (EPS), and Energy-guided Likelihood Reweighting (ELR). Specifically, we explicitly decompose each sample's latent representation into a biased feature and a stochastic corrective perturbation δ. Using a Bayesian formulation, we model the posterior distribution of the perturbation δ conditioned on the predictive uncertainty, quantified via calibrated energy scores. To mitigate language bias, the posterior is optimized through energy-driven variational inference with an uncertainty-adaptive prior and sampling strategy. Moreover, the ELR mechanism incorporates an energy-based weighting of the reconstruction objective and enforces an energy-coherence constraint to emphasize challenging, high-uncertainty instances and align model confidence before and after debiasing. Extensive experiments conducted across multiple standard VQA benchmarks consistently validate the superior performance of our BayesVQA method over state-of-the-art competitors under distributional shifts and challenging bias conditions. Huanjia Zhu, Xiangwen Deng, Qinghao Zhong, Bingzhi Chen |
AAAI | 2 |
| 2025 | Towards Differential Optimization: Rehearsal-Free Class-Incremental Learning with Slow Learners and Fast AdaptersabstractClass-incremental learning (CIL) enables models to learn new tasks without forgetting previously acquired knowledge. However, existing CIL approaches often struggle with inadequate adaptation to task-specific feature spaces and catastrophic forgetting of previously-acquired knowledge, compromising the models’ plasticity and stability. To address these challenges, this paper proposes a novel differential optimization paradigm called DO-CIL, which incorporates task-agnostic slow learner (TSL) with task-specific fast adapter (TFA) for rehearsal-free CIL. Specifically, TSL aims to effectively capture shared knowledge with low learning rates for robust generalization, while TFA allows pre-trained models to adapt to new task-specific feature spaces. Benefitting from the classifier retraining strategy, a learnable semantic shift network is also proposed to align prototypes with the evolving model representation, facilitating the retraining of task-specific classifiers based on these updated prototypes. Extensive experiments on multiple benchmark datasets consistently demonstrate the superiority and effectiveness of our DO-CIL approach compared to state-of-the-art baselines. Yinghong Chen, Huanjia Zhu, Jieyi Cai, Jun Liang 0002, Bingzhi Chen |
ICASSP | 2 |
| 2025 | Advancing Few-Shot Class-Incremental Learning with Virtual Prototype Guidance PromptingabstractFew-Shot Class-Incremental Learning (FSCIL) aims to incrementally learn new class knowledge from limited samples while preserving previously knowledge from encountered classes. However, existing FSCIL methods encounter two primary challenges: (1) inadequate adaptation, where overfitting to new classes compromises the model’s adaptability, and (2) catastrophic forgetting, where previously learned knowledge is not well preserved. In this paper, we propose the Virtual Prototype Guidance Prompting (VPGP) paradigm, integrating the Multi-Grained Prompt (MGP) and Virtual-Prototype Guidance (VPG) strategies. Specifically, MGP enhances adaptation and prevents overfitting by introducing domain-general and fine-grained prompts, expanding the embedding space to capture core feature representations of novel classes. Meanwhile, VPG mitigates catastrophic forgetting by employing a dynamic fusion strategy to retrieve robust old class knowledge and generate virtual prototype, guiding the model to maintain learned knowledge across different sessions. Extensive experiments on multiple benchmark datasets demonstrate the superiority of our proposed VPGP framework. Huanjia Zhu, Xiaocheng Fang, Jun Liang 0002, Bingzhi Chen |
ICASSP | 2 |
| 2025 | Cause-Effect Driven Optimization for Robust Medical Visual Question Answering with Language BiasesabstractExisting Medical Visual Question Answering (Med-VQA) models often suffer from language biases, where spurious correlations between question types and answer categories are inadvertently established. To address these issues, we propose a novel Cause-Effect Driven Optimization framework called CEDO, that incorporates three well-established mechanisms, i.e., Modality-driven Heterogeneous Optimization (MHO), Gradient-guided Modality Synergy (GMS), and Distribution-adapted Loss Rescaling (DLR), for comprehensively mitigating language biases from both causal and effectual perspectives. Specifically, MHO employs adaptive learning rates for specific modalities to achieve heterogeneous optimization, thus enhancing robust reasoning capabilities. Additionally, GMS leverages the Pareto optimization method to foster synergistic interactions between modalities and enforce gradient orthogonality to eliminate bias updates, thereby mitigating language biases from the effect side, i.e., shortcut bias. Furthermore, DLR is designed to assign adaptive weights to individual losses to ensure balanced learning across all answer categories, effectively alleviating language biases from the cause side, i.e., imbalance biases within datasets. Extensive experiments on multiple traditional and bias-sensitive benchmarks consistently demonstrate the robustness of CEDO over state-of-the-art competitors. Huanjia Zhu, Yishu Liu 0001, Xiaozhao Fang, Guangming Lu 0002, Bingzhi Chen |
IJCAI | 1 |
| 2025 | Med-BiasX: Robust Medical Visual Question Answering with Language Biases
Huanjia Zhu, Yishu Liu 0001, Chengju Zhou, Guangming Lu 0002, Bingzhi Chen |
MICCAI (14) | 1 |
| 2025 | Language-Bias-Resilient Visual Question Answering via Adaptive Multi-Margin Collaborative DebiasingabstractLanguage bias in Visual Question Answering (VQA) arises when models exploit spurious statistical correlations between question templates and answers, particularly in out-of-distribution scenarios, thereby neglecting essential visual cues and compromising genuine multimodal reasoning. Despite numerous efforts to enhance the robustness of VQA models, a principled understanding of how such bias originates and influences model behavior remains underdeveloped. In this paper, we address this gap through a comprehensive empirical and theoretical analysis, revealing that modality-specific gradient imbalances, which originate from the inherent heterogeneity of multimodal data, lead to skewed feature fusion and biased classifier weights. To alleviate these issues, we propose a novel Multi-Margin Collaborative Debiasing (MMCD) framework that adaptively integrates frequency-, confidence-, and difficulty-aware angular margins with a dynamic difficulty-aware contrastive learning mechanism, to dynamically reshape decision boundaries. Extensive experiments across multiple challenging VQA benchmarks confirm the consistent superiority of our proposed MMCD over state-of-the-art baselines in combating language bias. Huanjia Zhu, Shuyuan Zheng, Yishu Liu 0001, Sudong Cai, Bingzhi Chen |
NeurIPS | 1 |
| 2025 | LBF-VQA: Towards Language Bias-Free Visual Question Answering With Multi-Space Collaborative Debiasing
Yishu Liu 0001, Huanjia Zhu, Bingzhi Chen, Xiaozhao Fang, Guangming Lu 0002, Shengli Xie 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2024 | Combating Visual Question Answering Hallucinations via Robust Multi-Space Co-Debias Learning
Yishu Liu 0001, Huanjia Zhu, Yuncheng Jiang 0004, Zheng Zhang 0006, Bingzhi Chen |
ACM Multimedia | 3 |