VLDB 2026 Research / reviewers in the wild / expert
Erdun Gao
dblp:246/5884
· DBLP profile ↗
15ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0003-1736-2764ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Distraction is All You Need for Multimodal Large Language Model JailbreakingabstractMultimodal Large Language Models (MLLMs) bridge the gap between visual and textual data, enabling a range of advanced applications. However, complex internal interactions among visual elements and their alignment with text can introduce vulnerabilities, which may be exploited to bypass safety mechanisms. To address this, we analyze the relationship between image content and task and find that the complexity of subimages, rather than their content, is key. Building on this insight, we propose the Distraction Hypothesis, followed by a novel framework called Contrasting Subimage Distraction Jailbreaking (CS-DJ), to achieve jailbreaking by disrupting MLLMs alignment through multi-level distraction strategies. CS-DJ consists of two components: structured distraction, achieved through query decomposition that induces a distributional shift by fragmenting harmful prompts into sub-queries, and visual-enhanced distraction, realized by constructing contrasting subimages to disrupt the interactions among visual elements within the model. This dual strategy disperses the model’s attention, reducing its ability to detect and mitigate harmful content. Extensive experiments across five representative scenarios and four popular closed-source MLLMs, including GPT-4o-mini, GPT-4o, GPT-4V, and Gemini-1.5-Flash, demonstrate that CS-DJ achieves average success rates of 52.40% for the attack success rate and 74.10% for the ensemble attack success rate. These results reveal the potential of distraction-based approaches to exploit and bypass MLLMs’ defenses, offering new insights for attack strategies. Our code is available at https://github.com/TeamPigeonLab/CS-DJ.Warning: This paper contains unfiltered content generated by MLLMs that may be offensive to readers Zuopeng Yang, Jiluan Fan, Anli Yan, Erdun Gao, Kanghua Mo, Changyu Dong |
CVPR | 4 |
| 2025 | MissScore: High-Order Score Estimation in the Presence of Missing DataabstractScore-based generative models are essential in various machine learning applications, with strong capabilities in generation quality. In particular, high-order derivatives (scores) of data density offer deep insights into data distributions, building on the proven effectiveness of first-order scores for modeling and generating synthetic data, unlocking new possibilities for applications. However, learning them typically requires complete data, which is often unavailable in domains such as healthcare and finance due to data corruption, acquisition constraints, or incomplete records. To tackle this challenge, we introduce MissScore, a novel framework for estimating high-order scores in the presence of missing data. We derive objective functions for estimating high-order scores under different missing data mechanisms and propose a new algorithm specifically designed to handle missing data effectively. Our empirical results demonstrate that MissScore accurately and efficiently learns the high-order scores from incomplete data and generates high-quality samples, resulting in strong performance across a range of downstream tasks. Wenqin Liu, Haoze Hou, Erdun Gao, Biwei Huang, Qiuhong Ke, Howard D. Bondell, Mingming Gong |
ICML | 3 |
| 2025 | Causality-aligned Prompt Learning via Diffusion-based Counterfactual GenerationabstractPrompt learning has garnered attention for its efficiency over traditional model training and fine-tuning. However, existing methods, constrained by inadequate theoretical foundations, encounter difficulties in achieving causally invariant prompts, ultimately falling short of capturing robust features that generalize effectively across categories. To address these challenges, we introduce the DiCap model, a theoretically grounded Diffusion-based Counterfactual prompt learning framework, which leverages a diffusion process to iteratively sample gradients from the marginal and conditional distributions of the causal model, guiding the generation of counterfactuals that satisfy the minimal sufficiency criterion. Grounded in rigorous theoretical derivations, this approach guarantees the identifiability of counterfactual outcomes while imposing strict bounds on estimation errors. We further employ a contrastive learning framework that leverages the generated counterfactuals, thereby enabling the refined extraction of prompts that are precisely aligned with the causal features of the data. Extensive experimental results demonstrate that our method performs excellently across tasks such as image classification, image-text retrieval, and visual question answering, with particularly strong advantages in unseen categories. Xinshu Li 0001, Ruoyu Wang 0038, Erdun Gao, Mingming Gong, Lina Yao 0001 |
ACM Multimedia | 3 |
| 2025 | On the Value of Cross-Modal Misalignment in Multimodal Representation LearningabstractMultimodal representation learning, exemplified by multimodal contrastive learning (MMCL) using image-text pairs, aims to learn powerful representations by aligning cues across modalities. This approach relies on the core assumption that the exemplar image-text pairs constitute two representations of an identical concept. However, recent research has revealed that real-world datasets often exhibit cross-modal misalignment. There are two distinct viewpoints on how to address this issue: one suggests mitigating the misalignment, and the other leveraging it. We seek here to reconcile these seemingly opposing perspectives, and to provide a practical guide for practitioners. Using latent variable models we thus formalize cross-modal misalignment by introducing two specific mechanisms: Selection bias, where some semantic variables are absent in the text, and perturbation bias, where semantic variables are altered—both leading to misalignment in data pairs. Our theoretical analysis demonstrates that, under mild assumptions, the representations learned by MMCL capture exactly the information related to the subset of the semantic variables invariant to selection and perturbation biases. This provides a unified perspective for understanding misalignment. Based on this, we further offer actionable insights into how misalignment should inform the design of real-world ML systems. We validate our theoretical findings via extensive empirical studies on both synthetic data and real image-text datasets, shedding light on the nuanced impact of cross-modal misalignment on multimodal representation learning. Yichao Cai 0001, Yuhang Liu 0002, Erdun Gao, Tianjiao Jiang, Zhen Zhang 0008, Anton van den Hengel, Qinfeng Shi |
NeurIPS | 3 |
| 2024 | A Variational Framework for Estimating Continuous Treatment Effects with Measurement ErrorabstractEstimating treatment effects has numerous real-world applications in various fields, such as epidemiology and political science. While much attention has been devoted to addressing the challenge using fully observational data, there has been comparatively limited exploration of this issue in cases when the treatment is not directly observed. In this paper, we tackle this problem by developing a general variational framework, which is flexible to integrate with advanced neural network-based approaches, to identify the average dose-response function (ADRF) with the continuously valued error-contaminated treatment. Our approach begins with the formulation of a probabilistic data generation model, treating the unobserved treatment as a latent variable. In this model, we leverage a learnable density estimation neural network to derive its prior distribution conditioned on covariates. This module also doubles as a generalized propensity score estimator, effectively mitigating selection bias arising from observed confounding variables. Subsequently, we calculate the posterior distribution of the treatment, taking into account the observed measurement and outcome. To mitigate the impact of treatment error, we introduce a re-parametrized treatment value, replacing the error-affected one, to make more accurate predictions regarding the outcome. To demonstrate the adaptability of our framework, we incorporate two state-of-the-art ADRF estimation methods and rigorously assess its efficacy through extensive simulations and experiments using semi-synthetic data. Erdun Gao, Howard D. Bondell, Mingming Gong |
ICLR | 1 |
| 2024 | Eliminating Contextual Prior Bias for Semantic Image Editing via Dual-Cycle DiffusionabstractThe recent success of text-to-image generation diffusion models has also revolutionized semantic image editing, enabling the manipulation of images based on query/target texts. Despite these advancements, a significant challenge lies in the potential introduction of contextual prior bias in pre-trained models during image editing, e.g., making unexpected modifications to inappropriate regions. To address this issue, we present a novel approach called Dual-Cycle Diffusion, which generates an unbiased mask to guide image editing. The proposed model incorporates a Bias Elimination Cycle that consists of both a forward path and an inverted path, each featuring a Structural Consistency Cycle to ensure the preservation of image content during the editing process. The forward path utilizes the pre-trained model to produce the edited image, while the inverted path converts the result back to the source image. The unbiased mask is generated by comparing differences between the processed source image and the edited image to ensure that both conform to the same distribution. Our experiments demonstrate the effectiveness of the proposed method, as it significantly improves the D-CLIP score from 0.272 to 0.283. The code will be available athttps://github.com/JohnDreamer/DualCycleDiffsion. Zuopeng Yang, Erdun Gao, Daqing Liu, Jie Yang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | A novel two-level embedding pattern for grayscale-invariant reversible data hiding
Zhibin Pan, Erdun Gao, Xinyi Gao 0001, Guojun Fan |
Multim. Tools Appl. | 4 |
| 2022 | MissDAG: Causal Discovery in the Presence of Missing Data with Continuous Additive Noise ModelsabstractState-of-the-art causal discovery methods usually assume that the observational data is complete. However, the missing data problem is pervasive in many practical scenarios such as clinical trials, economics, and biology. One straightforward way to address the missing data problem is first to impute the data using off-the-shelf imputation methods and then apply existing causal discovery methods. However, such a two-step method may suffer from suboptimality, as the imputation algorithm may introduce bias for modeling the underlying data distribution. In this paper, we develop a general method, which we call MissDAG, to perform causal discovery from data with incomplete observations. Focusing mainly on the assumptions of ignorable missingness and the identifiable additive noise models (ANMs), MissDAG maximizes the expected likelihood of the visible part of observations under the expectation-maximization (EM) framework. In the E-step, in cases where computing the posterior distributions of parameters in closed-form is not feasible, Monte Carlo EM is leveraged to approximate the likelihood. In the M-step, MissDAG leverages the density transformation to model the noise distributions with simpler and specific formulations by virtue of the ANMs and uses a likelihood-based causal discovery algorithm with directed acyclic graph constraint. We demonstrate the flexibility of MissDAG for incorporating various causal discovery algorithms and its efficacy through extensive simulations and real data experiments. Erdun Gao, Ignavier Ng, Mingming Gong, Li Shen 0008, Tongliang Liu, Kun Zhang 0001, Howard D. Bondell |
NeurIPS | 1 |
| 2021 | Reversible data hiding method based on combining IPVO with bias-added Rhombus predictor by multi-predictor mechanism
Guojun Fan, Zhibin Pan, Erdun Gao, Xinyi Gao 0001 |
Signal Process. | 3 |
| 2020 | Effective reversible data hiding using dynamic neighboring pixels prediction based on prediction-error histogram
Zhibin Pan, Xinyi Gao 0001, Erdun Gao |
Multim. Tools Appl. | 4 |
| 2020 | Reversible data hiding for high dynamic range images using two-dimensional prediction-error histogram of the second time prediction
Xinyi Gao 0001, Zhibin Pan, Erdun Gao, Guojun Fan |
Signal Process. | 3 |
| 2020 | Adaptive Complexity for Pixel-Value-Ordering Based Reversible Data HidingabstractPixel-value-ordering (PVO) is a widely used reversible data hiding (RDH) framework which aims to achieve the high quality of stego-image under low capacity. In this letter, we propose a general location-based adaptive complexity for PVO. Different from the block-based complexity in the previous PVO-based methods, our proposed method adaptively selects context pixels from the perspective of the relative locations of predicted pixel and prediction pixel. Consequently, different number of high correlation context pixels can be adaptively selected and the context pixels can break the limitation of the current block. Moreover, instead of sharing the same block complexity by two predicted pixels in the current block, each predicted pixel can be utilized independently according to its own corresponding complexity. Our proposed adaptive complexity can combine with any PVO-based methods and the experimental results show that our proposed method achieves a significant improvement in prediction accuracy and embedding performance. Zhibin Pan, Xinyi Gao 0001, Erdun Gao, Guojun Fan |
IEEE Signal Process. Lett. | 3 |
| 2019 | Reversible data hiding based on novel pairwise PVO and annular merging strategy
Erdun Gao, Zhibin Pan, Xinyi Gao 0001 |
Inf. Sci. | 1 |
| 2019 | Reversible data hiding based on novel embedding structure PVO and adaptive block-merging strategy
Zhibin Pan, Erdun Gao |
Multim. Tools Appl. | 2 |
| 2019 | A low bit-rate SOC-based reversible data hiding algorithm by using new encoding strategies
Zhibin Pan, Erdun Gao, Ruoxin Zhu |
Multim. Tools Appl. | 2 |