Erdun Gao

dblp:246/5884 · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0003-1736-2764ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2025 Distraction is All You Need for Multimodal Large Language Model Jailbreaking
abstract
Multimodal Large Language Models (MLLMs) bridge the gap between visual and textual data, enabling a range of advanced applications. However, complex internal interactions among visual elements and their alignment with text can introduce vulnerabilities, which may be exploited to bypass safety mechanisms. To address this, we analyze the relationship between image content and task and find that the complexity of subimages, rather than their content, is key. Building on this insight, we propose the Distraction Hypothesis, followed by a novel framework called Contrasting Subimage Distraction Jailbreaking (CS-DJ), to achieve jailbreaking by disrupting MLLMs alignment through multi-level distraction strategies. CS-DJ consists of two components: structured distraction, achieved through query decomposition that induces a distributional shift by fragmenting harmful prompts into sub-queries, and visual-enhanced distraction, realized by constructing contrasting subimages to disrupt the interactions among visual elements within the model. This dual strategy disperses the model’s attention, reducing its ability to detect and mitigate harmful content. Extensive experiments across five representative scenarios and four popular closed-source MLLMs, including GPT-4o-mini, GPT-4o, GPT-4V, and Gemini-1.5-Flash, demonstrate that CS-DJ achieves average success rates of 52.40% for the attack success rate and 74.10% for the ensemble attack success rate. These results reveal the potential of distraction-based approaches to exploit and bypass MLLMs’ defenses, offering new insights for attack strategies. Our code is available at https://github.com/TeamPigeonLab/CS-DJ.Warning: This paper contains unfiltered content generated by MLLMs that may be offensive to readers
Zuopeng Yang, Jiluan Fan, Anli Yan, Erdun Gao, Kanghua Mo, Changyu Dong
CVPR4
2025 MissScore: High-Order Score Estimation in the Presence of Missing Data
abstract
Score-based generative models are essential in various machine learning applications, with strong capabilities in generation quality. In particular, high-order derivatives (scores) of data density offer deep insights into data distributions, building on the proven effectiveness of first-order scores for modeling and generating synthetic data, unlocking new possibilities for applications. However, learning them typically requires complete data, which is often unavailable in domains such as healthcare and finance due to data corruption, acquisition constraints, or incomplete records. To tackle this challenge, we introduce MissScore, a novel framework for estimating high-order scores in the presence of missing data. We derive objective functions for estimating high-order scores under different missing data mechanisms and propose a new algorithm specifically designed to handle missing data effectively. Our empirical results demonstrate that MissScore accurately and efficiently learns the high-order scores from incomplete data and generates high-quality samples, resulting in strong performance across a range of downstream tasks.
Wenqin Liu, Haoze Hou, Erdun Gao, Biwei Huang, Qiuhong Ke, Howard D. Bondell, Mingming Gong
ICML3
2025 Causality-aligned Prompt Learning via Diffusion-based Counterfactual Generation
abstract
Prompt learning has garnered attention for its efficiency over traditional model training and fine-tuning. However, existing methods, constrained by inadequate theoretical foundations, encounter difficulties in achieving causally invariant prompts, ultimately falling short of capturing robust features that generalize effectively across categories. To address these challenges, we introduce the DiCap model, a theoretically grounded Diffusion-based Counterfactual prompt learning framework, which leverages a diffusion process to iteratively sample gradients from the marginal and conditional distributions of the causal model, guiding the generation of counterfactuals that satisfy the minimal sufficiency criterion. Grounded in rigorous theoretical derivations, this approach guarantees the identifiability of counterfactual outcomes while imposing strict bounds on estimation errors. We further employ a contrastive learning framework that leverages the generated counterfactuals, thereby enabling the refined extraction of prompts that are precisely aligned with the causal features of the data. Extensive experimental results demonstrate that our method performs excellently across tasks such as image classification, image-text retrieval, and visual question answering, with particularly strong advantages in unseen categories.
Xinshu Li 0001, Ruoyu Wang 0038, Erdun Gao, Mingming Gong, Lina Yao 0001
ACM Multimedia3
2025 On the Value of Cross-Modal Misalignment in Multimodal Representation Learning
abstract
Multimodal representation learning, exemplified by multimodal contrastive learning (MMCL) using image-text pairs, aims to learn powerful representations by aligning cues across modalities. This approach relies on the core assumption that the exemplar image-text pairs constitute two representations of an identical concept. However, recent research has revealed that real-world datasets often exhibit cross-modal misalignment. There are two distinct viewpoints on how to address this issue: one suggests mitigating the misalignment, and the other leveraging it. We seek here to reconcile these seemingly opposing perspectives, and to provide a practical guide for practitioners. Using latent variable models we thus formalize cross-modal misalignment by introducing two specific mechanisms: Selection bias, where some semantic variables are absent in the text, and perturbation bias, where semantic variables are altered—both leading to misalignment in data pairs. Our theoretical analysis demonstrates that, under mild assumptions, the representations learned by MMCL capture exactly the information related to the subset of the semantic variables invariant to selection and perturbation biases. This provides a unified perspective for understanding misalignment. Based on this, we further offer actionable insights into how misalignment should inform the design of real-world ML systems. We validate our theoretical findings via extensive empirical studies on both synthetic data and real image-text datasets, shedding light on the nuanced impact of cross-modal misalignment on multimodal representation learning.
Yichao Cai 0001, Yuhang Liu 0002, Erdun Gao, Tianjiao Jiang, Zhen Zhang 0008, Anton van den Hengel, Qinfeng Shi
NeurIPS3
2024 A Variational Framework for Estimating Continuous Treatment Effects with Measurement Error
abstract
Estimating treatment effects has numerous real-world applications in various fields, such as epidemiology and political science. While much attention has been devoted to addressing the challenge using fully observational data, there has been comparatively limited exploration of this issue in cases when the treatment is not directly observed. In this paper, we tackle this problem by developing a general variational framework, which is flexible to integrate with advanced neural network-based approaches, to identify the average dose-response function (ADRF) with the continuously valued error-contaminated treatment. Our approach begins with the formulation of a probabilistic data generation model, treating the unobserved treatment as a latent variable. In this model, we leverage a learnable density estimation neural network to derive its prior distribution conditioned on covariates. This module also doubles as a generalized propensity score estimator, effectively mitigating selection bias arising from observed confounding variables. Subsequently, we calculate the posterior distribution of the treatment, taking into account the observed measurement and outcome. To mitigate the impact of treatment error, we introduce a re-parametrized treatment value, replacing the error-affected one, to make more accurate predictions regarding the outcome. To demonstrate the adaptability of our framework, we incorporate two state-of-the-art ADRF estimation methods and rigorously assess its efficacy through extensive simulations and experiments using semi-synthetic data.
Erdun Gao, Howard D. Bondell, Mingming Gong
ICLR1
2024 Eliminating Contextual Prior Bias for Semantic Image Editing via Dual-Cycle Diffusion
abstract
The recent success of text-to-image generation diffusion models has also revolutionized semantic image editing, enabling the manipulation of images based on query/target texts. Despite these advancements, a significant challenge lies in the potential introduction of contextual prior bias in pre-trained models during image editing, e.g., making unexpected modifications to inappropriate regions. To address this issue, we present a novel approach called Dual-Cycle Diffusion, which generates an unbiased mask to guide image editing. The proposed model incorporates a Bias Elimination Cycle that consists of both a forward path and an inverted path, each featuring a Structural Consistency Cycle to ensure the preservation of image content during the editing process. The forward path utilizes the pre-trained model to produce the edited image, while the inverted path converts the result back to the source image. The unbiased mask is generated by comparing differences between the processed source image and the edited image to ensure that both conform to the same distribution. Our experiments demonstrate the effectiveness of the proposed method, as it significantly improves the D-CLIP score from 0.272 to 0.283. The code will be available athttps://github.com/JohnDreamer/DualCycleDiffsion.
Zuopeng Yang, Erdun Gao, Daqing Liu, Jie Yang 0002
IEEE Trans. Circuits Syst. Video Technol.4
2023 A novel two-level embedding pattern for grayscale-invariant reversible data hiding
Zhibin Pan, Erdun Gao, Xinyi Gao 0001, Guojun Fan
Multim. Tools Appl.4
2022 MissDAG: Causal Discovery in the Presence of Missing Data with Continuous Additive Noise Models
abstract
State-of-the-art causal discovery methods usually assume that the observational data is complete. However, the missing data problem is pervasive in many practical scenarios such as clinical trials, economics, and biology. One straightforward way to address the missing data problem is first to impute the data using off-the-shelf imputation methods and then apply existing causal discovery methods. However, such a two-step method may suffer from suboptimality, as the imputation algorithm may introduce bias for modeling the underlying data distribution. In this paper, we develop a general method, which we call MissDAG, to perform causal discovery from data with incomplete observations. Focusing mainly on the assumptions of ignorable missingness and the identifiable additive noise models (ANMs), MissDAG maximizes the expected likelihood of the visible part of observations under the expectation-maximization (EM) framework. In the E-step, in cases where computing the posterior distributions of parameters in closed-form is not feasible, Monte Carlo EM is leveraged to approximate the likelihood. In the M-step, MissDAG leverages the density transformation to model the noise distributions with simpler and specific formulations by virtue of the ANMs and uses a likelihood-based causal discovery algorithm with directed acyclic graph constraint. We demonstrate the flexibility of MissDAG for incorporating various causal discovery algorithms and its efficacy through extensive simulations and real data experiments.
Erdun Gao, Ignavier Ng, Mingming Gong, Li Shen 0008, Tongliang Liu, Kun Zhang 0001, Howard D. Bondell
NeurIPS1
2021 Reversible data hiding method based on combining IPVO with bias-added Rhombus predictor by multi-predictor mechanism
Guojun Fan, Zhibin Pan, Erdun Gao, Xinyi Gao 0001
Signal Process.3
2020 Effective reversible data hiding using dynamic neighboring pixels prediction based on prediction-error histogram
Zhibin Pan, Xinyi Gao 0001, Erdun Gao
Multim. Tools Appl.4
2020 Reversible data hiding for high dynamic range images using two-dimensional prediction-error histogram of the second time prediction
Xinyi Gao 0001, Zhibin Pan, Erdun Gao, Guojun Fan
Signal Process.3
2020 Adaptive Complexity for Pixel-Value-Ordering Based Reversible Data Hiding
abstract
Pixel-value-ordering (PVO) is a widely used reversible data hiding (RDH) framework which aims to achieve the high quality of stego-image under low capacity. In this letter, we propose a general location-based adaptive complexity for PVO. Different from the block-based complexity in the previous PVO-based methods, our proposed method adaptively selects context pixels from the perspective of the relative locations of predicted pixel and prediction pixel. Consequently, different number of high correlation context pixels can be adaptively selected and the context pixels can break the limitation of the current block. Moreover, instead of sharing the same block complexity by two predicted pixels in the current block, each predicted pixel can be utilized independently according to its own corresponding complexity. Our proposed adaptive complexity can combine with any PVO-based methods and the experimental results show that our proposed method achieves a significant improvement in prediction accuracy and embedding performance.
Zhibin Pan, Xinyi Gao 0001, Erdun Gao, Guojun Fan
IEEE Signal Process. Lett.3
2019 Reversible data hiding based on novel pairwise PVO and annular merging strategy
Erdun Gao, Zhibin Pan, Xinyi Gao 0001
Inf. Sci.1
2019 Reversible data hiding based on novel embedding structure PVO and adaptive block-merging strategy
Zhibin Pan, Erdun Gao
Multim. Tools Appl.2
2019 A low bit-rate SOC-based reversible data hiding algorithm by using new encoding strategies
Zhibin Pan, Erdun Gao, Ruoxin Zhu
Multim. Tools Appl.2