VLDB 2026 Research / reviewers in the wild / expert
Bochuan Cao
dblp:334/3881
· DBLP profile ↗
13ranked-venue papers
4as first author
13since 2021 · last 2025
0009-0007-1973-8186ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Security and privacy · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | JoPA: Explaining Large Language Model's Generation via Joint Prompt AttributionabstractLarge Language Models (LLMs) have demonstrated impressive performances in complex text generation tasks.However, the contribution of the input prompt to the generated content still remains obscure to humans, underscoring the necessity of understanding the causality between input and output pairs.Existing works for providing prompt-specific explanation often confine model output to be classification or next-word prediction.Few initial attempts aiming to explain the entire language generation often treat input prompt texts independently, ignoring their combinatorial effects on the followup generation.In this study, we introduce a counterfactual explanation framework based on Joint Prompt Attribution, JoPA, which aims to explain how a few prompt texts collaboratively influences the LLM's complete generation.Particularly, we formulate the task of prompt attribution for generation interpretation as a combinatorial optimization problem, and introduce a probabilistic algorithm to search for the casual input combination in the discrete space.We define and utilize multiple metrics to evaluate the produced explanations, demonstrating both the faithfulness and efficiency of our framework. Yurui Chang, Bochuan Cao, Lu Lin 0001 |
ACL (1) | 2 |
| 2025 | You Can't Steal Nothing: Mitigating Prompt Leakages in LLMs via System VectorsabstractLarge language models (LLMs) have been widely adopted across various applications, leveraging customized system prompts for diverse tasks. Facing potential system prompt leakage risks, model developers have implemented strategies to prevent leakage, primarily by disabling LLMs from repeating their context when encountering known attack patterns. However, it remains vulnerable to new and unforeseen prompt-leaking techniques. In this paper, we first introduce a simple yet effective prompt leaking attack to reveal such risks. Our attack is capable of extracting system prompts from various LLM-based application, even from SOTA LLM models such as GPT-4o or Claude 3.5 Sonnet. Our findings further inspire us to search for a fundamental solution to the problems by having no system prompt in the context. To this end, we propose SysVec, a novel method that encodes system prompts as internal representation vectors rather than raw text. By doing so, SysVec minimizes the risk of unauthorized disclosure while preserving the LLM's core language capabilities. Remarkably, this approach not only enhances security but also improves the model's general instruction-following abilities. Experimental results demonstrate that SysVec effectively mitigates prompt leakage attacks, preserves the LLM's functional integrity, and helps alleviate the forgetting issue in long-context scenarios. Bochuan Cao, Changjiang Li, Yuanpu Cao, Yameng Ge, Ting Wang 0006 |
CCS | 1 |
| 2025 | TruthFlow: Truthful LLM Generation via Representation Flow CorrectionabstractLarge language models (LLMs) are known to struggle with consistently generating truthful responses. While various representation intervention techniques have been proposed, these methods typically apply a universal representation correction vector to all input queries, limiting their effectiveness against diverse queries in practice. In this study, we introduce TruthFlow, a novel method that leverages the Flow Matching technique for query-specific truthful representation correction. Specifically, TruthFlow first uses a flow model to learn query-specific correction vectors that transition representations from hallucinated to truthful states. Then, during inference, the trained flow model generates these correction vectors to enhance the truthfulness of LLM outputs. Experimental results demonstrate that TruthFlow significantly improves performance on open-ended generation tasks across various advanced LLMs evaluated on TruthfulQA. Moreover, the trained TruthFlow model exhibits strong transferability, performing effectively on other unseen hallucination benchmarks. Bochuan Cao, Yuanpu Cao |
ICML | 2 |
| 2025 | AdvI2I: Adversarial Image Attack on Image-to-Image Diffusion ModelsabstractRecent advances in diffusion models have significantly enhanced the quality of image synthesis, yet they have also introduced serious safety concerns, particularly the generation of Not Safe for Work (NSFW) content. Previous research has demonstrated that adversarial prompts can be used to generate NSFW content. However, such adversarial text prompts are often easily detectable by text-based filters, limiting their efficacy. In this paper, we expose a previously overlooked vulnerability: adversarial image attacks targeting Image-to-Image (I2I) diffusion models. We propose AdvI2I, a novel framework that manipulates input images to induce diffusion models to generate NSFW content. By optimizing a generator to craft adversarial images, AdvI2I circumvents existing defense mechanisms, such as Safe Latent Diffusion (SLD), without altering the text prompts. Furthermore, we introduce AdvI2I-Adaptive, an enhanced version that adapts to potential countermeasures and minimizes the resemblance between adversarial images and NSFW concept embeddings, making the attack more resilient against defenses. Through extensive experiments, we demonstrate that both AdvI2I and AdvI2I-Adaptive can effectively bypass current safeguards, highlighting the urgent need for stronger security measures to address the misuse of I2I diffusion models. Yaopei Zeng, Yuanpu Cao, Bochuan Cao, Yurui Chang, Lu Lin 0001 |
ICML | 3 |
| 2025 | Watch the Watchers! On the Security Risks of Robustness-Enhancing Diffusion Models
Changjiang Li, Ren Pang, Bochuan Cao, Fenglong Ma, Shouling Ji, Ting Wang 0006 |
USENIX Security Symposium | 3 |
| 2024 | Defending Against Alignment-Breaking Attacks via Robustly Aligned LLMabstractRecently, Large Language Models (LLMs) have made significant advancements and are now widely used across various domains.Unfortunately, there has been a rising concern that LLMs can be misused to generate harmful or malicious content.Though a line of research has focused on aligning LLMs with human values and preventing them from producing inappropriate content, such alignments are usually vulnerable and can be bypassed by alignmentbreaking attacks via adversarially optimized or handcrafted jailbreaking prompts.In this work, we introduce a Robustly Aligned LLM (RA-LLM) to defend against potential alignmentbreaking attacks.RA-LLM can be directly constructed upon an existing aligned LLM with a robust alignment checking function, without requiring any expensive retraining or fine-tuning process of the original LLM.Furthermore, we also provide a theoretical analysis for RA-LLM to verify its effectiveness in defending against alignment-breaking attacks.Through real-world experiments on open-source large language models, we demonstrate that RA-LLM can successfully defend against both state-of-the-art adversarial prompts and popular handcrafted jailbreaking prompts by reducing their attack success rates from nearly 100% to around 10% or less.WARNING: This paper contains unsafe model responses.Reader discretion is advised. Bochuan Cao, Yuanpu Cao, Lu Lin 0001 |
ACL (1) | 1 |
| 2024 | Jailbreak Open-Sourced Large Language Models via Enforced DecodingabstractHangfan Zhang, Zhimeng Guo, Huaisheng Zhu, Bochuan Cao, Lu Lin, Jinyuan Jia, Jinghui Chen, Dinghao Wu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Hangfan Zhang, Zhimeng Guo, Huaisheng Zhu, Bochuan Cao, Lu Lin 0001, Jinyuan Jia 0001, Dinghao Wu |
ACL (1) | 4 |
| 2024 | Stealthy and Persistent Unalignment on Large Language Models via Backdoor InjectionsabstractYuanpu Cao, Bochuan Cao, Jinghui Chen. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Yuanpu Cao, Bochuan Cao |
NAACL-HLT | 2 |
| 2024 | Data Free Backdoor AttacksabstractBackdoor attacks aim to inject a backdoor into a classifier such that it predicts any input with an attacker-chosen backdoor trigger as an attacker-chosen target class.
Existing backdoor attacks require either retraining the classifier with some clean data or modifying the model's architecture.
As a result, they are 1) not applicable when clean data is unavailable, 2) less efficient when the model is large, and 3) less stealthy due to architecture changes.
In this work, we propose DFBA, a novel retraining-free and data-free backdoor attack without changing the model architecture.
Technically, our proposed method modifies a few parameters of a classifier to inject a backdoor.
Through theoretical analysis, we verify that our injected backdoor is provably undetectable and unremovable by various state-of-the-art defenses under mild assumptions.
Our evaluation on multiple datasets further demonstrates that our injected backdoor: 1) incurs negligible classification loss, 2) achieves 100\% attack success rates, and 3) bypasses six existing state-of-the-art defenses.
Moreover, our comparison with a state-of-the-art non-data-free backdoor attack shows our attack is more stealthy and effective against various defenses while achieving less classification accuracy loss.
We will release our code upon paper acceptance. Bochuan Cao, Jinyuan Jia 0001, Chuxuan Hu, Wenbo Guo 0002, Zhen Xiang, Bo Li 0026, Dawn Song |
NeurIPS | 1 |
| 2024 | Personalized Steering of Large Language Models: Versatile Steering Vectors Through Bi-directional Preference OptimizationabstractResearchers have been studying approaches to steer the behavior of Large Language Models (LLMs) and build personalized LLMs tailored for various applications. While fine-tuning seems to be a direct solution, it requires substantial computational resources and may significantly affect the utility of the original LLM.
Recent endeavors have introduced more lightweight strategies, focusing on extracting ``steering vectors'' to guide the model's output toward desired behaviors by adjusting activations within specific layers of the LLM's transformer architecture. However, such steering vectors are directly extracted from the activations of human preference data and thus often lead to suboptimal results and occasional failures, especially in alignment-related scenarios.
In this work, we propose an innovative approach that could produce more effective steering vectors through bi-directional preference optimization.
Our method is designed to allow steering vectors to directly influence the generation probability of contrastive human preference data pairs, thereby offering a more precise representation of the target behavior. By carefully adjusting the direction and magnitude of the steering vector, we enabled personalized control over the desired behavior across a spectrum of intensities.
Extensive experimentation across various open-ended generation tasks, particularly focusing on steering AI personas, has validated the efficacy of our approach.
Moreover, we comprehensively investigate critical alignment-concerning scenarios, such as managing truthfulness, mitigating hallucination, and addressing jailbreaking attacks alongside their respective defenses. Remarkably, our method can still demonstrate outstanding steering effectiveness across these scenarios. Furthermore, we showcase the transferability of our steering vectors across different models/LoRAs and highlight the synergistic benefits of applying multiple vectors simultaneously. These findings significantly broaden the practicality and versatility of our proposed method. Yuanpu Cao, Tianrong Zhang, Bochuan Cao, Ziyi Yin 0003, Lu Lin 0001, Fenglong Ma |
NeurIPS | 3 |
| 2024 | On the Difficulty of Defending Contrastive Learning against Backdoor Attacks
Changjiang Li, Ren Pang, Bochuan Cao, Zhaohan Xi, Shouling Ji, Ting Wang 0006 |
USENIX Security Symposium | 3 |
| 2023 | IMPRESS: Evaluating the Resilience of Imperceptible Perturbations Against Unauthorized Data Usage in Diffusion-Based Generative AIabstractDiffusion-based image generation models, such as Stable Diffusion or DALL·E 2, are able to learn from given images and generate high-quality samples following the guidance from prompts. For instance, they can be used to create artistic images that mimic the style of an artist based on his/her original artworks or to maliciously edit the original images for fake content. However, such ability also brings serious ethical issues without proper authorization from the owner of the original images. In response, several attempts have been made to protect the original images from such unauthorized data usage by adding imperceptible perturbations, which are designed to mislead the diffusion model and make it unable to properly generate new samples. In this work, we introduce a perturbation purification platform, named IMPRESS, to evaluate the effectiveness of imperceptible perturbations as a protective measure.
IMPRESS is based on the key observation that imperceptible perturbations could lead to a perceptible inconsistency between the original image and the diffusion-reconstructed image, which can be used to devise a new optimization strategy for purifying the image, which may weaken the protection of the original image from unauthorized data usage (e.g., style mimicking, malicious editing).
The proposed IMPRESS platform offers a comprehensive evaluation of several contemporary protection methods, and can be used as an evaluation platform for future protection methods. Bochuan Cao, Changjiang Li, Ting Wang 0006, Jinyuan Jia 0001, Bo Li 0026 |
NeurIPS | 1 |
| 2022 | Wild-Time: A Benchmark of in-the-Wild Distribution Shift over TimeabstractDistribution shifts occur when the test distribution differs from the training distribution, and can considerably degrade performance of machine learning models deployed in the real world. While recent works have studied robustness to distribution shifts, distribution shifts arising from the passage of time have the additional structure of timestamp metadata. Real-world examples of such shifts are underexplored, and it is unclear whether existing models can leverage trends in past distribution shifts to reliably extrapolate into the future. To address this gap, we curate Wild-Time, a benchmark of 5 datasets that reflect temporal distribution shifts arising in a variety of real-world applications, including drug discovery, patient prognosis, and news classification. On these datasets, we systematically benchmark 13 approaches with various inductive biases. We evaluate methods in domain-generalization, continual learning, self-supervised learning, and ensemble learning, which leverage timestamps to extract the common structure of the distribution shifts. We extend several domain-generalization methods to the temporal distribution shift setting by treating windows of time as different domains. Finally, we propose two evaluation strategies to evaluate model performance under temporal distribution shifts---evaluation with a fixed time split (Eval-Fix) and evaluation with a data stream (Eval-Stream). Eval-Fix, our primary evaluation strategy, aims to provide a simple evaluation protocol for the broader machine learning community, while Eval-Stream serves as a complementary benchmark for continual learning approaches. Our experiments demonstrate that existing methods are limited in tackling temporal distribution shift: across all settings, we observe an average performance drop of 20% from in-distribution to out-of-distribution data. Huaxiu Yao, Caroline Choi, Bochuan Cao, Yoonho Lee 0001, Pang Wei Koh, Chelsea Finn |
NeurIPS | 3 |