EDBT 2026 Demo / reviewers in the wild / expert
Xiaojun Jia
dblp:57/5656
· DBLP profile ↗
66ranked-venue papers
11as first author
62since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 40 · 7 first-author · 39 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 5 first-author · 26 since 2021Security and privacy · 13 · 2 first-author · 13 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PhysPatch: A Physically Realizable and Transferable Adversarial Patch Attack for Multimodal Large Language Models-based Autonomous Driving SystemsabstractMultimodal Large Language Models (MLLMs) are becoming integral to autonomous driving (AD) systems due to their strong vision-language reasoning capabilities. However, MLLMs are vulnerable to adversarial attacks—particularly adversarial patch attacks—which can pose serious threats in real-world scenarios. Existing patch-based attack methods are primarily designed for object detection models. Due to the more complex architectures and strong reasoning capabilities of MLLMs, these approaches perform poorly when transferred to MLLM-based systems. To address these limitations, we propose PhysPatch, a physically realizable and transferable adversarial patch framework tailored for MLLM-based AD systems. PhysPatch jointly optimizes patch location, shape, and content to enhance attack effectiveness and real-world applicability. It introduces a semantic-based mask initialization strategy for realistic placement, an SVD-based local alignment loss with patch-guided crop-resize to improve transferability, and a potential field-based mask refinement method. Extensive experiments across open-source, commercial, and reasoning-capable MLLMs demonstrate that PhysPatch significantly outperforms state-of-the-art (SOTA) methods in steering MLLM-based AD systems toward target-aligned perception and planning outputs. Moreover, PhysPatch consistently places adversarial patches in physically feasible regions of AD scenes, ensuring strong real-world applicability and deployability. Qi Guo 0008, Xiaojun Jia, Shanmin Pang, Simeng Qin, Lin Wang 0026, Ju Jia, Yang Liu 0003, Qing Guo 0005 |
AAAI | 2 |
| 2026 | GeoShield: Safeguarding Geolocation Privacy from Vision-Language Models via Adversarial PerturbationsabstractVision-Language Models (VLMs) such as GPT-4o now demonstrate a remarkable ability to infer users' locations from public shared images, posing a substantial risk to geoprivacy. Although adversarial perturbations offer a potential defense, current methods are ill-suited for this scenario: they often perform poorly on high-resolution images and low perturbation budgets, and may introduce irrelevant semantic content. To address these limitations, we propose GeoShield, a novel adversarial framework designed for robust geoprivacy protection in real-world scenarios. GeoShield comprises three key modules: a feature disentanglement module that separates geographical and non-geographical information, an exposure element identification module that pinpoints geo-revealing regions within an image, and a scale-adaptive enhancement module that jointly optimizes perturbations at both global and local levels to ensure effectiveness across resolutions. Extensive experiments on challenging benchmarks show that GeoShield consistently surpasses prior methods in black-box settings, achieving strong privacy protection with minimal impact on visual or semantic quality. To our knowledge, this work is the first to explore adversarial perturbations for defending against geolocation inference by advanced VLMs, providing a practical solution to escalating privacy concerns. Xiaojun Jia, Yuan Xun, Simeng Qin, Xiaochun Cao |
AAAI | 2 |
| 2026 | The Emotional Baby Is Truly Deadly: Does Your Multimodal Large Reasoning Model Have Emotional Flattery Towards Humans?abstractMultimodal large reasoning models (MLRMs) have advanced visual-textual integration, enabling sophisticated human-AI interaction. While prior work has exposed MLRMs to visual jailbreaks, it remains underexplored how their reasoning capabilities reshape the security landscape under adversarial inputs. To fill this gap, we conduct a systematic security assessment of MLRMs and uncover a security-reasoning paradox: although deeper reasoning boosts cross‑modal risk recognition, it also creates cognitive blind spots that adversaries can exploit. We observe that MLRMs oriented toward human-centric service are highly susceptible to users' emotional cues during the deep-thinking stage, often overriding safety protocols or built‑in safety checks under high emotional intensity. Inspired by this key insight, we propose EmoAgent, an autonomous adversarial emotion-agent that orchestrates exaggerated affective prompts to hijack reasoning pathways. Even when visual risks are correctly identified, models can still produce harmful completions through emotional misalignment. We further identify persistent high-risk failure modes in transparent deep-thinking scenarios, such as MLRMs generating harmful reasoning masked behind seemingly safe responses. These failures expose misalignments between internal inference and surface-level behavior, eluding existing content-based safeguards. To quantify these risks, we introduce three metrics: (1) Risk-Reasoning Stealth Score (RRSS) for harmful reasoning beneath benign outputs; (2) Risk-Visual Neglect Rate (RVNR) for unsafe completions despite visual risk recognition; and (3) Refusal Attitude Inconsistency (RAIC) for evaluating refusal unstability under prompt variants. Extensive experiments on advanced MLRMs demonstrate the effectiveness of EmoAgent and reveal deeper emotional cognitive misalignments in model safety. Yuan Xun, Xiaojun Jia, Simeng Qin, Hua Zhang 0008 |
AAAI | 2 |
| 2026 | AsFT: Anchoring Safety During LLM Fine-Tuning Within Narrow Safety BasinabstractFine-tuning large language models (LLMs) improves performance but introduces critical safety vulnerabilities: even minimal harmful data can severely compromise safety measures. We observe that perturbations orthogonal to the alignment direction—defined by weight differences between aligned (safe) and unaligned models—rapidly compromise model safety. In contrast, updates along the alignment direction largely preserve it, revealing the parameter space as a "narrow safety basin". To address this, we propose AsFT (Anchoring Safety in Fine-Tuning) to maintain safety by explicitly constraining update directions during fine-tuning. By penalizing updates orthogonal to the alignment direction, AsFT effectively constrains the model within the "narrow safety basin," thus preserving its inherent safety. Extensive experiments on multiple datasets and models show that AsFT reduces harmful behaviors by up to 7.60%, improves task performance by 3.44%, and consistently outperforms existing methods across multiple tasks. Qihui Zhang, Yue Huang 0001, Xiaojun Jia, Kun-Peng Ning, Jia-Yu Yao, Jigang Wang, Hailiang Dai, Yibing Song, Li Yuan 0007 |
AAAI | 5 |
| 2026 | MPAS: Breaking Sequential Constraints of Multi-Agent Communication Topologies via Individual-Epistemic Message PropagationabstractLarge language model (LLM)-driven agents are designed to handle a wide range of tasks autonomously. As tasks become increasingly composite, the integration of multiple agents into a graph-structured system offers a promising solution. Recent advances mainly architect the communication order among agents into a specified directed acyclic graph, from which a one-by-one execution can be determined by topological sort. However, sequential architectures restrict the diversity of the information flow, hinder parallel computation, and exhibit vulnerabilities to potential backdoor threats. To overcome underlying shortcomings of sequential structures, we propose a node-wise multi-agent scheme, named message passing agent system (MPAS). Specifically, to parallelize the communication across agents, we extend the message propagation mechanism in graph representation learning to multi-agent scenarios and introduce our individual-epistemic message propagation. To further enhance expressiveness and robustness, we investigate three self-driven message aggregators. To achieve desired working flows, collaborative connections can be optimized without constraints. The experimental results reveal that compared to state-of-the-art sequential designs, MPAS could architect more advanced algorithms in 93.8% of the evaluations, reduce the average communication time from 84.6 seconds to 14.2 seconds per round on AQuA, and improve resilience against backdoor misinformation injection in 94.4% tests. Jingxuan Yu, Ju Jia, Simeng Qin, Xiaojun Jia, Siqi Ma 0001, Yihao Huang 0001, Yali Yuan, Guang Cheng 0001 |
AAAI | 4 |
| 2026 | GAMBIT: A Gamified Jailbreak Framework for Multimodal Large Language ModelsabstractMultimodal Large Language Models (MLLMs) have become widely deployed, yet their safety alignment remains fragile under adversarial inputs. Previous work has shown that increasing inference steps can disrupt safety mechanisms and lead MLLMs to generate attacker-desired harmful content. However, most existing attacks focus on increasing the complexity of the modified visual task itself and do not explicitly leverage the model’s own reasoning incentives. This leads to them underperforming on reasoning models (Models with Chain-of-Thoughts) compared to non-reasoning ones (Models without Chain-of-Thoughts). If a model can think like a human, can we influence its cognitive-stage decisions so that it proactively completes a jailbreak? To validate this idea, we propose GAMBIT (Gamified Adversarial Multimodal Breakout via Instructional Traps), a novel multimodal jailbreak framework that decomposes and reassembles harmful visual semantics, then constructs a gamified scene that drives the model to explore, reconstruct intent, and answer as part of winning the game. The resulting structured reasoning chain increases task complexity in both vision and text, positioning the model as a participant whose goal pursuit reduces safety attention and induces it to answer the reconstructed malicious query. Extensive experiments on popular reasoning and non-reasoning MLLMs demonstrate that GAMBIT achieves high Attack Success Rates (ASR), reaching 92.13% on Gemini 2.5 Flash, 91.20% on QvQ-MAX, and 85.87% on GPT-4o, significantly outperforming baselines. Warning: This paper contains unsafe and offensive examples. Xiangdong Hu, Yangyang Jiang, Qin Hu 0001, Xiaojun Jia |
ACL (1) | 4 |
| 2026 | Buster: Implanting Semantic Backdoor Into Text Encoder to Mitigate NSFW Content Generation
Xiaojun Chen 0004, Yuexin Xuan, Zhendong Zhao, Xinfeng Li, Xiaojun Jia, Xiaofeng Wang 0001 |
DASFAA (5) | 6 |
| 2026 | EmoRAG: Evaluating RAG Robustness to Symbolic PerturbationsabstractRetrieval-Augmented Generation (RAG) systems are increasingly central to robust AI, enhancing large language model (LLM) faithfulness by incorporating external knowledge. However, our study unveils a critical, overlooked vulnerability: their profound susceptibility to subtle symbolic perturbations, particularly through near-imperceptible emotional icons (e.g., "(@_@)") that can catastrophically mislead retrieval, termed EmoRAG. We demonstrate that injecting a single emoticon into a query makes it nearly 100% likely to retrieve semantically unrelated texts, which contain a matching emoticon. Our extensive experiment across general question-answering and code domains, using a range of state-of-the-art retrievers and generators, reveals three key findings: (I) Single-Emoticon Disaster: Minimal emoticon injections cause maximal disruptions, with a single emoticon almost 100% dominating RAG output. (II) Positional Sensitivity: Placing an emoticon at the beginning of a query can cause severe perturbation, with F1-Scores exceeding 0.92 across all datasets. (III) Parameter-Scale Vulnerability: Counterintuitively, models with larger parameters exhibit greater vulnerability to the interference. We provide an in-depth analysis to uncover the underlying mechanisms of these phenomena. Furthermore, we raise a critical concern regarding the robustness assumption of current RAG systems, envisioning a threat scenario where an adversary exploits this vulnerability to manipulate the RAG system. We evaluate standard defenses and find them insufficient against EmoRAG. To address this, we propose targeted defenses, analyzing their strengths and limitations in mitigating emoticon-based perturbations. Finally, we outline future directions for building robust RAG systems. Xinyun Zhou, Xinfeng Li, Yinan Peng, Ming Xu 0006, Xuanwang Zhang, Yidong Wang 0003, Xiaojun Jia, Kun Wang 0056, Qingsong Wen, XiaoFeng Wang 0001, Wei Dong 0007 |
KDD (1) | 8 |
| 2026 | Odysseus: Jailbreaking Commercial Multimodal LLM-integrated Systems via Dual Steganography
Jiameng Cheng, Yiming Li 0004, Xiaojun Jia, Dacheng Tao |
NDSS | 4 |
| 2026 | WebCloak: Characterizing and Mitigating Threats From LLM-Driven Web Agents as Intelligent Scrapers
Xinfeng Li, Tianze Qiu, Yingbin Jin, Lixu Wang, Hanqing Guo, Xiaojun Jia, Xiaofeng Wang 0001, Wei Dong 0007 |
SP | 6 |
| 2026 | Boosting adversarial transferability of vision-language pre-trained models via optimal transport
Simeng Qin, Sensen Gao, Dongchen Han, Xiaojun Jia, Yang Bai 0011, Jindong Gu, Xiaochun Cao |
Pattern Recognit. | 5 |
| 2026 | AudioJailbreak: Jailbreak Attacks Against End-to-End Large Audio-Language ModelsabstractJailbreak attacks to Large audio-language models (LALMs) are studied recently, but they exclusively focused on the attack scenario where the adversary can fully manipulate user prompts (named strong adversary) and limited in effectiveness, applicability, and practicability. In this work, we first conduct an extensive evaluation showing that advanced text jailbreak attacks cannot be easily ported to end-to-end LALMs via text-to-speech (TTS) techniques. We then propose AUDIOJAILBREAK, a novel audio jailbreak attack, featuring (1) asynchrony: the jailbreak audios do not need to align with user prompts in the time axis by crafting suffixal jailbreak audios; (2) universality: a single jailbreak perturbation is effective for different prompts by incorporating multiple prompts into the perturbation generation; (3) stealthiness: the malicious intent of jailbreak audios is concealed by proposing various intent concealment strategies; and (4) over-the-air robustness: the jailbreak audios remain effective when being played over the air by incorporating reverberation into the perturbation generation. In contrast, all prior audio jailbreak attacks cannot offer asynchrony, universality, stealthiness, and/or over-the-air robustness. Moreover, AUDIOJAILBREAK is also applicable to a more practical and broader attack scenario where the adversary cannot fully manipulate user prompts (named weak adversary). Extensive experiments with thus far the most LALMs demonstrate the high effectiveness of AUDIOJAILBREAK, in particular, it can jailbreak openAI's GPT-4o-Audio and bypass Meta's Llama-Guard-3 safeguard, in the weak adversary scenario. We highlight that our work peeks into the security implications of audio jailbreak attacks against LALMs, and realistically fosters improving their robustness, especially for the newly proposed weak adversary. Guangke Chen, Fu Song, Zhe Zhao 0007, Xiaojun Jia, Yang Liu 0003, Yanchen Qiao, Weizhe Zhang, Weiping Tu, Yuhong Yang 0001, Bo Du 0001 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2026 | SMInject: Specious Malignant Injection Attacks With Semantically-Enhanced Tokens in Cross-Modal RetrievalabstractThe pre-training multimodal models have achieved remarkable success with powerful cross-modal understanding capabilities, while easily being affected by deliberate injection attacks. Although the deceptive injection attacks are harmful, they are valuable in revealing the vulnerability and improving the robustness for multimodal models. Unfortunately, the existing multimodal injection attacks pay less attention to the complicated roles of different modality-related causal correlation, which results in such attacks being susceptible to detection and defense. To alleviate this issue, we propose a novel specious malignant injection attack framework, calledSMInject, which exploits both the irrationality and causal correlation across diverse modalities to stealthily manipulate the space of output. To enhance the stealthiness, we generate deceptive injections to assemble the concepts by analyzing causal correlation under four types of attacks. To further boost the effectiveness, the malignant injections are guided to penetrate in the encoded embedding space by designing the premise-hypothesis consensus alignment. Extensive experiments on representative multimodal models demonstrate that ourSMInjectachieves over 14% higher attack success rate and 6% higher Hit@5 metric than state-of-the-art methods while preserving the overall utility of models. Moreover, we highlight that theSMInjectalso exhibits the desired transferability by investigating the impact of contextual factors, such as similar attack profiles, imperceptible noise perturbations,etc. Our code is available athttps://anonymous.4open.science/r/SMInject-0DBC. Ju Jia, Jiabao Guo, Xiaojun Jia, Siqi Ma 0001, Jie Gui, Robert H. Deng |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2026 | PersGuard: Preventing Malicious Personalization in Text-to-Image Diffusion Models via Model BackdoorsabstractDiffusion models (DMs) have made significant advances in high-quality text-to-image (T2I) synthesis, yet their personalization capabilities raise serious privacy and copyright concerns. Malicious actors can misuse these models to generate unauthorized content, such as realistic portraits or artistic style replicas. Existing proactive defenses primarily rely on applying adversarial perturbations to user-provided reference images to disrupt training. However, these approaches face some limitations in real-world scenarios: they rely on the assumption that all training images are pre-perturbed, and they are prone to failure when datasets contain unperturbed images or undergo minor data transformations. In this paper, we introduce PersGuard, a novel backdoor-based framework designed to prevent unauthorized personalization of pre-trained T2I diffusion models. Unlike perturbation-based methods, we assume protectors can control and embed protective backdoors into the models before their release. This mechanism ensures that if a downstream user fine-tunes the model on protected images, the model retains the backdoor and generates predefined protective outputs; conversely, for unprotected images, the backdoor is effectively removed during fine-tuning to ensure normal model utility. We formulate the backdoor injection as a unified optimization problem incorporating three objectives: a backdoor behavior loss to activate protection, a prior preservation loss to maintain standard generation capabilities, and a novel backdoor retention loss. The retention loss is specifically designed to mirror personalization loss, ensuring the backdoor remains robust during the downstream fine-tuning process. Extensive experiments across various scenarios, including gray-box and black-box settings, multi-object protection, and facial identity protection, demonstrate that PersGuard provides superior privacy protection compared to existing perturbation-based methods. Xiaojun Jia, Yuan Xun, Hua Zhang 0008, Xiaochun Cao |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2026 | PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image ModelsabstractRecent text-to-image (T2I) models have exhibited remarkable performance in generating high-quality images from text descriptions. However, these models are vulnerable to misuse, particularly generating not-safe-for-work (NSFW) content, such as sexually explicit, violent, political, and disturbing images, raising serious ethical concerns. In this work, we present PromptGuard, a novel content moderation technique that draws inspiration from the system prompt mechanism in large language models (LLMs) for safety alignment. Unlike LLMs, T2I models lack a direct interface for enforcing behavioral guidelines. Our key idea is to optimize a safety soft prompt that functions as an implicit system prompt within the T2I model’s textual embedding space. This universal soft prompt (P∗) directly moderates NSFW inputs, enabling safe yet realistic image generation without altering the inference efficiency or requiring proxy models.We further enhance its reliability and helpfulness through a divide-and-conquer strategy, which optimizes category-specific soft prompts and combines them into holistic safety guidance. Extensive experiments across five datasets demonstrate that PromptGuard effectively mitigates NSFW content generation while preserving high-quality benign outputs. PromptGuard achieves 3.8 times faster than prior content moderation methods, surpassing eight state-of-the-art defenses. Rigorous evaluation using both multi-head classifiers and VLM-based guardrails confirms its robustness, achieving an optimal average unsafe ratios down to 5.84% and 6.18%, respectively. Our code and dataset are available at https://t2ipromptguard. github.io/. Lingzhi Yuan, Xinfeng Li, Chejian Xu, Guanhong Tao 0001, Xiaojun Jia, Yihao Huang 0001, Wei Dong 0007, Yang Liu 0003, Bo Li 0026 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2026 | JailGuard: A Universal Detection Framework for Prompt-based Attacks on LLM SystemsabstractThe systems and software powered by Large Language Models (LLMs) and Multi-Modal Large Language Models (MLLMs) have played a critical role in numerous scenarios. However, current LLM systems are vulnerable to prompt-based attacks, with jailbreaking attacks enabling the LLM system to generate harmful content, while hijacking attacks manipulate the LLM system to perform attacker-desired tasks, underscoring the necessity for detection tools. Unfortunately, existing detecting approaches are usually tailored to specific attacks, resulting in poor generalization in detecting various attacks across different modalities. To address it, we propose JailGuard , a universal detection framework deployed on top of LLM systems for prompt-based attacks across text and image modalities. JailGuard operates on the principle that attacks are inherently less robust than benign ones. Specifically, JailGuard mutates untrusted inputs to generate variants and leverages the discrepancy of the variants’ responses on the target model to distinguish attack samples from benign samples. We implement 18 mutators for text and image inputs and design a mutator combination policy to further improve detection generalization. The evaluation on the dataset containing 15 known attack types suggests that JailGuard achieves the best detection accuracy of 86.14%/82.90% on text and image inputs, outperforming state-of-the-art methods by 11.81–25.73% and 12.20–21.40%. Xiaoyu Zhang 0013, Cen Zhang, Tianlin Li, Yihao Huang 0001, Xiaojun Jia, Ming Hu 0003, Jie Zhang 0073, Yang Liu 0003, Shiqing Ma, Chao Shen 0001 |
ACM Trans. Softw. Eng. Methodol. | 5 |
| 2025 | Perception-Guided Jailbreak Against Text-to-Image ModelsabstractIn recent years, Text-to-Image (T2I) models have garnered significant attention due to their remarkable advancements. However, security concerns have emerged due to their potential to generate inappropriate or Not-Safe-For-Work (NSFW) images. In this paper, inspired by the observation that texts with different semantics can lead to similar human perceptions, we propose an LLM-driven perception-guided jailbreak method, termed PGJ. It is a black-box jailbreak method that requires no specific T2I model (model-free) and generates highly natural attack prompts. Specifically, we propose identifying a safe phrase that is similar in human perception yet inconsistent in text semantics with the target unsafe word and using it as a substitution. The experiments conducted on six open-source models and commercial online services with thousands of prompts have verified the effectiveness of PGJ. Yihao Huang 0001, Le Liang, Tianlin Li, Xiaojun Jia, Run Wang 0001, Weikai Miao, Geguang Pu, Yang Liu 0003 |
AAAI | 4 |
| 2025 | Efficient Universal Goal Hijacking with Semantics-guided Prompt OrganizationabstractYihao Huang, Chong Wang, Xiaojun Jia, Qing Guo, Felix Juefei-Xu, Jian Zhang, Yang Liu, Geguang Pu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yihao Huang 0001, Chong Wang 0013, Xiaojun Jia, Qing Guo 0005, Felix Juefei-Xu, Jian Zhang 0087, Yang Liu 0003, Geguang Pu |
ACL (1) | 3 |
| 2025 | PBI-Attack: Prior-Guided Bimodal Interactive Black-Box Jailbreak Attack for Toxicity MaximizationabstractRuoxi Cheng, Yizhong Ding, Shuirong Cao, Ranjie Duan, Xiaoshuang Jia, Shaowei Yuan, Simeng Qin, Zhiqiang Wang, Xiaojun Jia. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Ruoxi Cheng, Yizhong Ding, Shuirong Cao, Ranjie Duan, Xiaoshuang Jia, Shaowei Yuan, Simeng Qin, Xiaojun Jia |
EMNLP | 9 |
| 2025 | AutoPrompt: Automated Red-Teaming of Text-to-Image Models via LLM-Driven Adversarial Prompts
Yufan Liu 0002, Wanqian Zhang, Huashan Chen, Lin Wang 0108, Xiaojun Jia, Zheng Lin 0001, Weiping Wang 0005 |
ICCV | 5 |
| 2025 | 3D Gaussian Splatting Driven Multi-View Robust Physical Adversarial Camouflage Generation
Tianrui Lou, Xiaojun Jia, Siyuan Liang 0004, Xiaochun Cao |
ICCV | 2 |
| 2025 | Heuristic-Induced Multimodal Risk Distribution Jailbreak Attack for Multimodal Large Language Models
Xiaojun Jia, Ranjie Duan, Xinfeng Li, Yihao Huang 0001, Xiaoshuang Jia, Zhixuan Chu, Wenqi Ren |
ICCV | 2 |
| 2025 | Accelerate 3D Object Detection Models via Zero-Shot Attention Key PruningabstractQuery-based methods with dense features have demonstrated remarkable success in 3D object detection tasks. However, the computational demands of these models, particularly with large image sizes and multiple transformer layers, pose significant challenges for efficient running on edge devices. Existing pruning and distillation methods either need retraining or are designed for ViT models, which are hard to migrate to 3D detectors. To address this issue, we propose a zero-shot runtime pruning method for transformer decoders in 3D object detection models. The method, termed tgGBC (trim keys gradually Guided By Classification scores), systematically trims keys in transformer modules based on their importance. We expand the classification score to multiply it with the attention map to get the importance score of each key and then prune certain keys after each transformer layer according to their importance scores. Our method achieves a 1.99x speedup in the transformer decoder of the latest ToC3D model, with only a minimal performance loss of less than 1%. Interestingly, for certain models, our method even enhances their performance. Moreover, we deploy 3D detectors with tgGBC on an edge device, further validating the effectiveness of our method. The code can be found at https://github.com/iseri27/tg_gbc. Lizhen Xu, Xiuxiu Bai, Xiaojun Jia, Jianwu Fang, Shanmin Pang |
ICCV | 3 |
| 2025 | Improved Techniques for Optimization-Based Jailbreaking on Large Language ModelsabstractLarge language models (LLMs) are being rapidly developed, and a key component of their widespread deployment is their safety-related alignment. Many red-teaming efforts aim to jailbreak LLMs, where among these efforts, the Greedy Coordinate Gradient (GCG) attack's success has led to a growing interest in the study of optimization-based jailbreaking techniques. Although GCG is a significant milestone, its attacking efficiency remains unsatisfactory. In this paper, we present several improved (empirical) techniques for optimization-based jailbreaks like GCG. We first observe that the single target template of ”Sure'' largely limits the attacking performance of GCG; given this, we propose to apply diverse target templates containing harmful self-suggestion and/or guidance to mislead LLMs. Besides, from the optimization aspects, we propose an automatic multi-coordinate updating strategy in GCG (i.e., adaptively deciding how many tokens to replace in each step) to accelerate convergence, as well as tricks like easy-to-hard initialization. Then, we combine these improved technologies to develop an efficient jailbreak method, dubbed $\mathcal{I}$-GCG. In our experiments, we evaluate our $\mathcal{I}$-GCG on a series of benchmarks (such as NeurIPS 2023 Red Teaming Track). The results demonstrate that our improved techniques can help GCG outperform state-of-the-art jailbreaking attacks and achieve a nearly 100\% attack success rate.
The code is released at https://github.com/jiaxiaojunQAQ/I-GCG. Xiaojun Jia, Tianyu Pang, Yihao Huang 0001, Jindong Gu, Yang Liu 0003, Xiaochun Cao |
ICLR | 1 |
| 2025 | DAMA: Data- and Model-aware Alignment of Multi-modal LLMsabstractDirect Preference Optimization (DPO) has shown effectiveness in aligning multi-modal large language models (MLLM) with human preferences. However, existing methods exhibit an imbalanced responsiveness to the data of varying hardness, tending to overfit on the easy-to-distinguish data while underfitting on the hard-to-distinguish data. In this paper, we propose Data- and Model-aware DPO (DAMA) to dynamically adjust the optimization process from two key aspects: (1) a data-aware strategy that incorporates data hardness, and (2) a model-aware strategy that integrates real-time model responses. By combining the two strategies, DAMA enables the model to effectively adapt to data with varying levels of hardness. Extensive experiments on five benchmarks demonstrate that DAMA not only significantly enhances the trustworthiness, but also improves the effectiveness over general tasks. For instance, on the Object HalBench, our DAMA-7B reduces response-level and mentioned-level hallucination by 90.0% and 95.3%, respectively, surpassing the performance of GPT-4V. Jinda Lu, Junkang Wu, Jinghan Li, Xiaojun Jia, Shuo Wang 0008, Yifan Zhang 0004, Junfeng Fang, Xiang Wang 0010, Xiangnan He 0001 |
ICML | 4 |
| 2025 | Cannot See the Forest for the Trees: Invoking Heuristics and Biases to Elicit Irrational Choices of LLMsabstractDespite the remarkable performance of Large Language Models (LLMs), they remain vulnerable to jailbreak attacks, which can compromise their safety mechanisms. Existing studies often rely on brute-force optimization or manual design, failing to uncover potential risks in real-world scenarios. To address this, we propose a novel jailbreak attack framework, ICRT, inspired by heuristics and biases in human cognition. Leveraging the simplicity effect, we employ cognitive decomposition to reduce the complexity of malicious prompts. Simultaneously, relevance bias is utilized to reorganize prompts, enhancing semantic alignment and inducing harmful outputs effectively. Furthermore, we introduce a ranking-based harmfulness evaluation metric that surpasses the traditional binary success-or-failure paradigm by employing ranking aggregation methods such as Elo, HodgeRank, and Rank Centrality to comprehensively quantify the harmfulness of generated content. Experimental results show that our approach consistently bypasses mainstream LLMs’ safety mechanisms and generates high-risk content. Haoming Yang, Ke Ma 0001, Xiaojun Jia, Yingfei Sun, Qianqian Xu 0001, Qingming Huang |
ICML | 3 |
| 2025 | The Eye of Sherlock Holmes: Uncovering User Private Attribute Profiling via Vision-Language Model Agentic Framework
Feiran Liu, Xinyi Huang 0015, Yinan Peng, Xinfeng Li, Lixu Wang, Yutong Shen, Ranjie Duan, Simeng Qin, Xiaojun Jia, Qingsong Wen, Wei Dong 0007 |
ACM Multimedia | 10 |
| 2025 | Adversarial Attacks against Closed-Source MLLMs via Feature Optimal AlignmentabstractMultimodal large language models (MLLMs) remain vulnerable to transferable adversarial examples. While existing methods typically achieve targeted attacks by aligning global features—such as CLIP’s [CLS] token—between adversarial and target samples, they often overlook the rich local information encoded in patch tokens. This leads to suboptimal alignment and limited transferability, particularly for closed-source models. To address this limitation, we propose a targeted transferable adversarial attack method based on feature optimal alignment, called FOA-Attack, to improve adversarial transfer capability. Specifically, at the global level, we introduce a global feature loss based on cosine similarity to align the coarse-grained features of adversarial samples with those of target samples. At the local level, given the rich local representations within Transformers, we leverage clustering techniques to extract compact local patterns to alleviate redundant local features. We then formulate local feature alignment between adversarial and target samples as an optimal transport (OT) problem and propose a local clustering optimal transport loss to refine fine-grained feature alignment. Additionally, we propose a dynamic ensemble model weighting strategy to adaptively balance the influence of multiple models during adversarial example generation, thereby further improving transferability. Extensive experiments across various models demonstrate the superiority of the proposed method, outperforming state-of-the-art methods, especially in transferring to closed-source MLLMs. Xiaojun Jia, Sensen Gao, Simeng Qin, Tianyu Pang, Yihao Huang 0001, Xinfeng Li, Yiming Li 0004, Bo Li 0026, Yang Liu 0003 |
NeurIPS | 1 |
| 2025 | SeCon-RAG: A Two-Stage Semantic Filtering and Conflict-Free Framework for Trustworthy RAGabstractRetrieval-augmented generation (RAG) systems enhance large language models (LLMs) with external knowledge but are vulnerable to corpus poisoning and contamination attacks, which can compromise output integrity. Existing defenses often apply aggressive filtering, leading to unnecessary loss of valuable information and reduced reliability in generation.
To address this problem, we propose a two-stage semantic filtering and conflict-free framework for trustworthy RAG.
In the first stage, we perform a joint filter with semantic and cluster-based filtering which is guided by the Entity-intent-relation extractor (EIRE). EIRE extracts entities, latent objectives, and entity relations from both the user query and filtered documents, scores their semantic relevance, and selectively adds valuable documents into the clean retrieval database.
In the second stage, we proposed an EIRE-guided conflict-aware filtering module, which analyzes semantic consistency between the query, candidate answers, and retrieved knowledge before final answer generation, filtering out internal and external contradictions that could mislead the model.
Through this two-stage process, SeCon-RAG effectively preserves useful knowledge while mitigating conflict contamination, achieving significant improvements in both generation robustness and output trustworthiness.
Extensive experiments across various LLMs and datasets demonstrate that the proposed SeCon-RAG markedly outperforms state-of-the-art defense methods. Xiaonan Si, Meilin Zhu, Simeng Qin, Lijia Yu, Shuaitong Liu, Xinfeng Li, Ranjie Duan, Xiaojun Jia |
NeurIPS | 10 |
| 2025 | PATFinger: Prompt-Adapted Transferable Fingerprinting against Unauthorized Multimodal Dataset UsageabstractThe multimodal datasets can be leveraged to pre-train large-scale vision-language models by providing cross-modal semantics. Current endeavors for determining the usage of datasets mainly focus on single-modal dataset ownership verification through intrusive methods and non-intrusive techniques, while cross-modal approaches remain under-explored. Intrusive methods can adapt to multimodal datasets but degrade model accuracy, while non-intrusive methods rely on label-driven decision boundaries that fail to guarantee stable behaviors for verification. To address these issues, we propose a novel prompt-adapted transferable fingerprinting scheme from a training-free perspective, called PATFinger, which incorporates the global optimal perturbation (GOP) and the adaptive prompts to capture dataset-specific distribution characteristics. Our scheme utilizes inherent dataset attributes as fingerprints instead of compelling the model to learn triggers. The GOP is derived from the sample distribution to maximize embedding drifts between different modalities. Subsequently, our PATFinger re-aligns the adaptive prompt with GOP samples to capture the cross-modal interactions on the carefully crafted surrogate model. This allows the dataset owner to check the usage of datasets by observing specific prediction behaviors linked to the PATFinger during retrieval queries. Extensive experiments demonstrate the effectiveness of our scheme against unauthorized multimodal dataset usage on various cross-modal retrieval architectures by 30% over state-of-the-art baselines. Ju Jia, Xiaojun Jia, Yihao Huang 0001, Xinfeng Li, Cong Wu 0003, Lina Wang 0001 |
SIGIR | 3 |
| 2025 | Semantic-Aligned Adversarial Evolution Triangle for High-Transferability Vision-Language AttackabstractVision-language pre-training (VLP) models excel at interpreting both images and text but remain vulnerable to multimodal adversarial examples (AEs). Advancing the generation of transferable AEs, which succeed across unseen models, is key to developing more robust and practical VLP models. Previous approaches augment image-text pairs to enhance diversity within the adversarial example generation process, aiming to improve transferability by expanding the contrast space of image-text features. However, these methods focus solely on diversity around the current AEs, yielding limited gains in transferability. To address this issue, we propose to increase the diversity of AEs by leveraging the intersection regions along the adversarial trajectory during optimization. Specifically, we propose sampling from adversarial evolution triangles composed of clean, historical, and current adversarial examples to enhance adversarial diversity. We provide a theoretical analysis to demonstrate the effectiveness of the proposed adversarial evolution triangle. Moreover, we find that redundant inactive dimensions can dominate similarity calculations, distorting feature matching and making AEs model-dependent with reduced transferability. Hence, we propose to generate AEs in the semantic image-text feature contrast space, which can project the original feature space into a semantic corpus subspace. The proposed semantic-aligned subspace can reduce the image feature redundancy, thereby improving adversarial transferability. Extensive experiments across different datasets and models demonstrate that the proposed method can effectively improve adversarial transferability and outperform state-of-the-art adversarial attack methods. Xiaojun Jia, Sensen Gao, Qing Guo 0005, Simeng Qin, Ke Ma 0001, Yihao Huang 0001, Yang Liu 0003, Ivor W. Tsang, Xiaochun Cao |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | MOVE: Effective and Harmless Ownership Verification via Embedded External FeaturesabstractCurrently, deep neural networks (DNNs) are widely adopted in different applications. Despite its commercial values, training a well-performing DNN is resource-consuming. Accordingly, the well-trained model is valuable intellectual property for its owner. However, recent studies revealed the threats of model stealing, where the adversaries can obtain a function-similar copy of the victim model, even when they can only query the model. In this paper, we propose an effective and harmless model ownership verification (MOVE) to defend against different types of model stealing simultaneously, without introducing new security risks. In general, we conduct the ownership verification by verifying whether a suspicious model contains the knowledge of defender-specified external features. Specifically, we embed the external features by modifying a few training samples with style transfer. We then train a meta-classifier to determine whether a model is stolen from the victim. This approach is inspired by the understanding that the stolen models should contain the knowledge of features learned by the victim model. In particular, we develop our MOVE method under both glass-boxand closed-box settings and analyze its theoretical foundation to provide comprehensive model protection. Extensive experiments on benchmark datasets verify the effectiveness of our method and its resistance to potential adaptive attacks. Yiming Li 0004, Linghui Zhu, Xiaojun Jia, Yang Bai 0011, Yong Jiang 0001, Shutao Xia, Xiaochun Cao, Kui Ren 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Efficient Generation of Targeted and Transferable Adversarial Examples for Vision-Language Models via Diffusion ModelsabstractAdversarial attacks, particularly targeted transfer-based attacks, can be used to assess the adversarial robustness of large visual-language models (VLMs), allowing for a more thorough examination of potential security flaws before deployment. However, previous transfer-based adversarial attacks incur high costs due to high iteration counts and complex method structure. Furthermore, due to the unnaturalness of adversarial semantics, the generated adversarial examples have low transferability. These issues limit the utility of existing methods for assessing robustness. To address these issues, we propose AdvDiffVLM, which uses diffusion models to generate natural, unrestricted and targeted adversarial examples via score matching. Specifically, AdvDiffVLM uses Adaptive Ensemble Gradient Estimation (AEGE) to modify the score during the diffusion model’s reverse generation process, ensuring that the produced adversarial examples have natural adversarial targeted semantics, which improves their transferability. Simultaneously, to improve the quality of adversarial examples, we use the GradCAM-guided Mask Generation (GCMG) to disperse adversarial semantics throughout the image rather than concentrating them in a single area. Finally, AdvDiffVLM embeds more target semantics into adversarial examples after multiple iterations. Experimental results show that our method generates adversarial examples 5x to 10x faster than state-of-the-art (SOTA) transfer-based adversarial attacks while maintaining higher quality adversarial examples. Furthermore, compared to previous transfer-based adversarial attacks, the adversarial examples generated by our method have better transferability. Notably, AdvDiffVLM can successfully attack a variety of commercial VLMs in a black-box environment, including GPT-4V. The code is available athttps://github.com/gq-max/AdvDiffVLM Qi Guo 0008, Shanmin Pang, Xiaojun Jia, Yang Liu 0003, Qing Guo 0005 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Scale-Invariant Adversarial Attack Against Arbitrary-Scale Super-ResolutionabstractThe advent of local continuous image function (LIIF) has garnered significant attention for arbitrary-scale super-resolution (SR) techniques. However, while the vulnerabilities of fixed-scale SR have been assessed, the robustness of continuous representation-based arbitrary-scale SR against adversarial attacks remains an area warranting further exploration. The elaborately designed adversarial attacks for fixed-scale SR are scale-dependent, which will cause time-consuming and memory-consuming problems when applied to arbitrary-scale SR. To address this concern, we propose a simple yet effective “scale-invariant” SR adversarial attack method with good transferability, termed SIAGT. Specifically, we propose to construct resource-saving attacks by exploiting finite discrete points of continuous representation. In addition, we formulate a coordinate-dependent loss to enhance the cross-model transferability of the attack. The attack can significantly deteriorate the SR images while introducing imperceptible distortion to the targeted low-resolution (LR) images. Experiments carried out on three popular LIIF-based SR approaches and four classical SR datasets show remarkable attack performance and transferability of SIAGT. Yihao Huang 0001, Qing Guo 0005, Felix Juefei-Xu, Xiaojun Jia, Weikai Miao, Geguang Pu, Yang Liu 0003 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | CleanerCLIP: Fine-Grained Counterfactual Semantic Augmentation for Backdoor Defense in Contrastive LearningabstractWith the rise of the open-source community, multimodal pre-trained models such as CLIP have become increasingly vulnerable to backdoor attacks. Backdoor triggers can manipulate model outputs during inference, posing a significant threat to downstream users. While post-training defenses based on fine-tuning have made some progress, their effectiveness remains limited due to two key challenges: (1) They fail to weaken the connection between the backdoor trigger and the target text space, making it possible for the model to still rely on the backdoor pattern for prediction. (2) Although batch-level fine-tuning expands the data distribution, it lacks precise guidance for vision-language alignment. To address these limitations, we propose a fine-grained counterfactual text-driven sample-level fine-tuning defense. By generating counterfactual sub-texts, we explicitly guide the text space to shift towards a more discriminative and robust representation, thereby indirectly weakening the association between backdoor triggers and target semantics. Furthermore, we introduce intra-sample contrastive learning with hard negative sub-texts, which enforces a more precise gradient direction to enhance vision-language fine-grained alignment. We evaluate our approach against six different backdoor attack methods and conduct a comprehensive zero-shot classification study on ImageNet-1K. Experimental results demonstrate that our method surpasses SoTA CleanCLIP in defending against BadCLIP attacks, reducing the attack success rate (ASR) in Top-1 classification by 52.02% and Top-10 classification by 63.88%. We aim to enhance the robustness of multimodal models against backdoor threats, fostering safer deployment in real-world applications. Yuan Xun, Siyuan Liang 0004, Xiaojun Jia, Jun Chen 0035, Xiaochun Cao |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | NSB-H2GAN: "Negative Sample"-Boosted Hierarchical Heterogeneous Graph Attention Network for Interpretable Classification of Whole-Slide ImagesabstractGigapixel whole-slide image (WSI) prediction and region-of-interest localization present considerable challenges due to the diverse range of features both across different slides and within individual slides. Most current methods rely on weakly supervised learning using homogeneous graphs to establish context-aware relevance within slides, often neglecting the rich diversity of heterogeneous information inherent in pathology images. Inspired by the negative sampling strategy of the Determinantal Point Process (DPP) and the hierarchical structure of pathology slides, we introduce the Negative Sample Boosted Hierarchical Heterogeneous Graph Attention Network (NSB-H2GAN). This model addresses the over-smoothing issue typically encountered in classical Graph Convolutional Networks (GCNs) when applied to pathology slides. By incorporating "negative samples" at multiple scales and utilizing hierarchical, heterogeneous feature discrimination, NSB-H2GAN more effectively captures the unique features of each patch, leading to an improved representation of gigapixel WSIs. We evaluated the performance of NSB-H2GAN on three publicly available datasets: CAMELYON16, TCGA-NSCLC and TCGA-COAD. The results show that NSB-H2GAN significantly outperforms existing state-of-the-art methods in both qualitative and quantitative evaluations. Moreover, NSB-H2GAN generates more detailed and interpretable heatmaps, allowing for precise localization of tiny lesions as small as $200\mu m\times 200\mu m$ that are often missed by the human eye. The robust performance of NSB-H2GAN offers a new paradigm for computer-aided pathology diagnosis and holds great potential for advancing the clinical applications of computational pathology. Meiyan Liang, Shupeng Zhang, Xikai Wang, Muhammad Hamza Javed, Xiaojun Jia, Lin Wang 0110 |
IEEE Trans. Image Process. | 6 |
| 2024 | Does Few-Shot Learning Suffer from Backdoor Attacks?abstractThe field of few-shot learning (FSL) has shown promising results in scenarios where training data is limited, but its vulnerability to backdoor attacks remains largely unexplored. We first explore this topic by first evaluating the performance of the existing backdoor attack methods on few-shot learning scenarios. Unlike in standard supervised learning, existing backdoor attack methods failed to perform an effective attack in FSL due to two main issues. Firstly, the model tends to overfit to either benign features or trigger features, causing a tough trade-off between attack success rate and benign accuracy. Secondly, due to the small number of training samples, the dirty label or visible trigger in the support set can be easily detected by victims, which reduces the stealthiness of attacks. It seemed that FSL could survive from backdoor attacks. However, in this paper, we propose the Few-shot Learning Backdoor Attack (FLBA) to show that FSL can still be vulnerable to backdoor attacks. Specifically, we first generate a trigger to maximize the gap between poisoned and benign features. It enables the model to learn both benign and trigger features, which solves the problem of overfitting. To make it more stealthy, we hide the trigger by optimizing two types of imperceptible perturbation, namely attractive and repulsive perturbation, instead of attaching the trigger directly. Once we obtain the perturbations, we can poison all samples in the benign support set into a hidden poisoned support set and fine-tune the model on it. Our method demonstrates a high Attack Success Rate (ASR) in FSL tasks with different few-shot learning paradigms while preserving clean accuracy and maintaining stealthiness. This study reveals that few-shot learning still suffers from backdoor attacks, and its security should be given attention. Xiaojun Jia, Jindong Gu, Yuan Xun, Siyuan Liang 0004, Xiaochun Cao |
AAAI | 2 |
| 2024 | Hide in Thicket: Generating Imperceptible and Rational Adversarial Perturbations on 3D Point CloudsabstractAdversarial attack methods based on point manipulation for 3D point cloud classification have revealed the fragility of 3D models, yet the adversarial examples they produce are easily perceived or defended against. The tradeoff between the imperceptibility and adversarial strength leads most point attack methods to inevitably introduce easily detectable outlier points upon a successful attack. An-other promising strategy, shape-based attack, can effectively eliminate outliers, but existing methods often suffer significant reductions in imperceptibility due to irrational deformations. We find that concealing deformation perturbations in areas insensitive to human eyes can achieve a better tradeoff between imperceptibility and adversarial strength, specifically in parts of the object surface that are complex and exhibit drastic curvature changes. Therefore, we propose a novel shape-based adversarial attack method, HiT-ADV, which initially conducts a two-stage search for attack regions based on saliency and imperceptibility scores, and then adds deformation perturbations in each attack region using Gaussian kernel functions. Additionally, HiT-ADV is extendable to physical attack. We propose that by employing benign resampling and benign rigid transformations, we can further enhance physical adversarial strength with little sacrifice to imperceptibility. Extensive experiments have validated the superiority of our method in terms of adversarial and imperceptible properties in both digital and physical spaces. Our code is avaliable at: https://github.com/TRLou/HiT-ADV. Tianrui Lou, Xiaojun Jia, Jindong Gu, Li Liu 0002, Siyuan Liang 0004, Bangyan He, Xiaochun Cao |
CVPR | 2 |
| 2024 | Boosting Transferability in Vision-Language Attacks via Diversification Along the Intersection Region of Adversarial Trajectory
Sensen Gao, Xiaojun Jia, Xuhong Ren, Ivor W. Tsang, Qing Guo 0005 |
ECCV (57) | 2 |
| 2024 | EAT-Face: Emotion-Controllable Audio-Driven Talking Face Generation via Diffusion ModelabstractAudio-driven talking face generation is a promising task with a lot of attention. Despite abundant efforts are devoted to video quality and lip synchronization, most existing works do not take the unignorable aspect of facial emotional expression into account during generation. In this paper, we propose an Emotion-controllable Audio-driven Talking Face generation framework called EAT-Face that enables us to control multiple types of emotions. Specifically, the proposed method consists of a Talking Face Reconstructor (TFR) and a Facial Emotion Controller (FEC), utilizing fused multimodal information including audio signals, visual images, and textual emotions for synthesis. Firstly, TFR predicts face images synchronized with given audios from random noises, leveraging external guidances comprised of audio features, character references, and face masks as conditions. Then, FEC further manipulates the facial emotions based on TFR, leveraging the emotion embeddings extracted from emotion texts. However, a semantic misalignment problem lies in the emotion-texts and character images. To tackle this issue, we additionally propose a strategy called joint Emotion-Visual Embedding (EVE) to mitigate the misalignment. In this way, the proposed EAT-Face is captive to control emotion more precisely. Extensive experiments involving both objective evaluations and subjective investigations demonstrate the effectiveness of our framework in synthesizing high-fidelity and emotional talking face videos. Haodi Wang, Xiaojun Jia, Xiaochun Cao |
FG | 2 |
| 2024 | Poisoned Forgery Face: Towards Backdoor Attacks on Face Forgery DetectionabstractThe proliferation of face forgery techniques has raised significant concerns within society, thereby motivating the development of face forgery detection methods. These methods aim to distinguish forged faces from genuine ones and have proven effective in practical applications. However, this paper introduces a novel and previously unrecognized threat in face forgery detection scenarios caused by backdoor attack. By embedding backdoors into models and incorporating specific trigger patterns into the input, attackers can deceive detectors into producing erroneous predictions for forged faces. To achieve this goal, this paper proposes \emph{Poisoned Forgery Face} framework, which enables clean-label backdoor attacks on face forgery detectors. Our approach involves constructing a scalable trigger generator and utilizing a novel convolving process to generate translation-sensitive trigger patterns. Moreover, we employ a relative embedding method based on landmark-based regions to enhance the stealthiness of the poisoned samples. Consequently, detectors trained on our poisoned samples are embedded with backdoors. Notably, our approach surpasses SoTA backdoor baselines with a significant improvement in attack success rate (+16.39\% BD-AUC) and reduction in visibility (-12.65\% $L_\infty$). Furthermore, our attack exhibits promising performance against backdoor defenses. We anticipate that this paper will draw greater attention to the potential threats posed by backdoor attacks in face forgery detection scenarios. Our codes will be made available at \url{https://github.com/JWLiang007/PFF}. Siyuan Liang 0004, Aishan Liu, Xiaojun Jia, Junhao Kuang, Xiaochun Cao |
ICLR | 4 |
| 2024 | Multimodal Unlearnable Examples: Protecting Data against Multimodal Contrastive LearningabstractMultimodal contrastive learning (MCL) has shown remarkable advances in zero-shot classification by learning from millions of image-caption pairs crawled from the Internet. However, this reliance poses privacy risks, as hackers may unauthorizedly exploit image-text data for model training, potentially including personal and privacy-sensitive information. Recent works propose generating unlearnable examples by adding imperceptible perturbations to training images to build shortcuts for protection. However, they are designed for unimodal classification, which remains largely unexplored in MCL. We first explore this context by evaluating the performance of existing methods on image-caption pairs, and they do not generalize effectively to multimodal data and exhibit limited impact to build shortcuts due to the lack of labels and the dispersion of pairs in MCL. In this paper, we propose Multi-step Error Minimization (MEM), a novel optimization process for generating multimodal unlearnable examples. It extends the Error-Minimization (EM) framework to optimize both image noise and an additional text trigger, thereby enlarging the optimized space and effectively misleading the model to learn the shortcut between the noise features and the text trigger. Specifically, we adopt projected gradient descent to solve the noise minimization problem and use HotFlip to approximate the gradient and replace words to find the optimal text trigger. Extensive experiments demonstrate the effectiveness of MEM, with post-protection retrieval results nearly half of random guessing, and its high transferability across different models. Our code is available on the https://github.com/thinwayliu/Multimodal-Unlearnable-Examples Xiaojun Jia, Yuan Xun, Siyuan Liang 0004, Xiaochun Cao |
ACM Multimedia | 2 |
| 2024 | Context-Aware Robust Fine-Tuning
Xiaofeng Mao, Xiaojun Jia, Rong Zhang 0006, Hui Xue 0001, Zhao Li 0007 |
Int. J. Comput. Vis. | 3 |
| 2024 | Improving Fast Adversarial Training With Prior-Guided KnowledgeabstractFast adversarial training (FAT) is an efficient method to improve robustness in white-box attack scenarios. However, the original FAT suffers from catastrophic overfitting, which dramatically and suddenly reduces robustness after a few training epochs. Although various FAT variants have been proposed to prevent overfitting, they require high training time. In this paper, we investigate the relationship between adversarial example quality and catastrophic overfitting by comparing the training processes of standard adversarial training and FAT. We find that catastrophic overfitting occurs when the attack success rate of adversarial examples becomes worse. Based on this observation, we propose a positive prior-guided adversarial initialization to prevent overfitting by improving adversarial example quality without extra training time. This initialization is generated by using high-quality adversarial perturbations from the historical training process. We provide theoretical analysis for the proposed initialization and propose a prior-guided regularization method that boosts the smoothness of the loss function. Additionally, we design a prior-guided ensemble FAT method that averages the different model weights of historical models using different decay rates. Our proposed method, called FGSM-PGK, assembles the prior-guided knowledge, i.e., the prior-guided initialization and model weights, acquired during the historical training process. The proposed method can effectively improve the model's adversarial robustness in white-box attack scenarios. Evaluations of four datasets demonstrate the superiority of the proposed method. Xiaojun Jia, Yong Zhang 0034, Xingxing Wei 0001, Baoyuan Wu, Ke Ma 0001, Jue Wang 0001, Xiaochun Cao |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Texture Re-Scalable Universal Adversarial PerturbationabstractUniversal adversarial perturbation (UAP), also known as image-agnostic perturbation, is a fixed perturbation map that can fool the classifier with high probabilities on arbitrary images, making it more practical for attacking deep models in the real world. Previous UAP methods generate a scale-fixed and texture-fixed perturbation map for all images, which ignores the multi-scale objects in images and usually results in a low fooling ratio. Since the widely used convolution neural networks tend to classify objects according to semantic information stored in local textures, it seems a reasonable and intuitive way to improve the UAP from the perspective of utilizing local contents effectively. In this work, we find that the fooling ratios significantly increase when we add a constraint to encourage a small-scale UAP map and repeat it vertically and horizontally to fill the whole image domain. To this end, we propose texture scale-constrained UAP (TSC-UAP), a simple yet effective UAP enhancement method that automatically generates UAPs with category-specific local textures that can fool deep models more easily. Through a low-cost operation that restricts the texture scale, TSC-UAP achieves a considerable improvement in the fooling ratio and attack transferability for both data-dependent and data-free UAP methods. Experiments conducted on two state-of-the-art UAP methods, eight popular CNN models and four classical datasets show the remarkable performance of TSC-UAP. Yihao Huang 0001, Qing Guo 0005, Felix Juefei-Xu, Ming Hu 0003, Xiaojun Jia, Xiaochun Cao, Geguang Pu, Yang Liu 0003 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | Revisiting and Exploring Efficient Fast Adversarial Training via LAW: Lipschitz Regularization and Auto Weight AveragingabstractFast Adversarial Training (FAT) not only improves the model robustness but also reduces the training cost of standard adversarial training. However, fast adversarial training often suffers from Catastrophic Overfitting (CO), which results in poor robustness performance. Catastrophic Overfitting describes the phenomenon of a sudden and significant decrease in robust accuracy during the training of fast adversarial training. Many effective techniques have been developed to prevent Catastrophic Overfitting and improve the model robustness from different perspectives. However, these techniques adopt inconsistent training settings and require different training costs, i.e., training time and memory costs, leading to unfair comparisons. In this paper, we conduct a comprehensive study of over 10 fast adversarial training methods in terms of adversarial robustness and training costs. We revisit the effectiveness and efficiency of fast adversarial training techniques in preventing Catastrophic Overfitting from the perspective of model local nonlinearity and propose an effective Lipschitz regularization method for fast adversarial training. Furthermore, we explore the effect of data augmentation and weight averaging in fast adversarial training and propose a simple yet effective auto weight averaging method to improve robustness further. By assembling these techniques, we propose an effective FGSM-based fast adversarial training method equipped with Lipschitz regularization and Auto Weight averaging, abbreviated as FGSM-LAW. Experimental evaluations on four benchmark databases demonstrate the superiority of our method over state-of-the-art fast adversarial training methods and the advanced standard adversarial training methods. Xiaojun Jia, Yuefeng Chen, Xiaofeng Mao, Ranjie Duan, Jindong Gu, Rong Zhang 0006, Hui Xue 0001, Yang Liu 0003, Xiaochun Cao |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2024 | Fast Propagation Is Better: Accelerating Single-Step Adversarial Training via Sampling SubnetworksabstractAdversarial training has shown promise in building robust models against adversarial examples. A major drawback of adversarial training is the computational overhead introduced by the generation of adversarial examples. To overcome this limitation, adversarial training based on single-step attacks has been explored. Previous work improves the single-step adversarial training from different perspectives, e.g., sample initialization, loss regularization, and training strategy. Almost all of them treat the underlying model as a black box. In this work, we propose to exploit the interior building blocks of the model to improve efficiency. Specifically, we propose to dynamically sample lightweight subnetworks as a surrogate model during training. By doing this, both the forward and backward passes can be accelerated for efficient adversarial training. Besides, we provide theoretical analysis to show the model robustness can be improved by the single-step adversarial training with sampled subnetworks. Furthermore, we propose a novel sampling strategy where the sampling varies from layer to layer and from iteration to iteration. Compared with previous methods, our method not only reduces the training cost but also achieves better model robustness. Evaluations on a series of popular datasets demonstrate the effectiveness of the proposed FB-Better. Our code has been released at https://github.com/jiaxiaojunQAQ/FP-Better. Xiaojun Jia, Jianshu Li, Jindong Gu, Yang Bai 0011, Xiaochun Cao |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2024 | Minimalism is King! High-Frequency Energy-Based Screening for Data-Efficient Backdoor AttacksabstractGiven the effectiveness of deep neural networks in various fields, the security of neural networks has received great attention. The backdoor attack, which induces malicious behaviors of models by poisoning part of the training set, still remains a challenging problem. Many recent efforts have proposed different ways of embedding backdoors to improve the stealthiness of backdoor attacks. Yet, lowering the percentage of poisoned samples is one of the most direct ways to increase stealthiness. A recent study (Filtering-and-Updating strategy, FUS) has revealed that the sample selection for poisoning is also crucial, as different samples contribute differently to the final decision boundary of the network. Concretely, they utilize each sample’s forgetting events during the training stage to identify which samples will contribute more to the network’s prediction. The training phase of their search method, however, is computationally expensive and slow. To overcome this, in this paper, we propose an efficient sample selection strategy based on the high-frequency energy (HFE) of training samples with a global screening and updating strategy, which can not only achieve a higher backdoor-attack success rate but also reduce the searching time by a factor of 4320 compared to FUS (12 hours vs 10 seconds). The extensive experiment results on CIFAR-10, CIFAR-100, and ImageNet-10 have shown that our proposed method is much simpler, faster, and more efficient. Yuan Xun, Xiaojun Jia, Jindong Gu, Qing Guo 0005, Xiaochun Cao |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2023 | Generating Transferable 3D Adversarial Point Cloud via Random Perturbation FactorizationabstractRecent studies have demonstrated that existing deep neural networks (DNNs) on 3D point clouds are vulnerable to adversarial examples, especially under the white-box settings where the adversaries have access to model parameters. However, adversarial 3D point clouds generated by existing white-box methods have limited transferability across different DNN architectures. They have only minor threats in real-world scenarios under the black-box settings where the adversaries can only query the deployed victim model. In this paper, we revisit the transferability of adversarial 3D point clouds. We observe that an adversarial perturbation can be randomly factorized into two sub-perturbations, which are also likely to be adversarial perturbations. It motivates us to consider the effects of the perturbation and its sub-perturbations simultaneously to increase the transferability for sub-perturbations also contain helpful information. In this paper, we propose a simple yet effective attack method to generate more transferable adversarial 3D point clouds. Specifically, rather than simply optimizing the loss of perturbation alone, we combine it with its random factorization. We conduct experiments on benchmark dataset, verifying our method's effectiveness in increasing transferability while preserving high efficiency. Bangyan He, Jian Liu 0012, Yiming Li 0004, Siyuan Liang 0004, Jingzhi Li 0002, Xiaojun Jia, Xiaochun Cao |
AAAI | 6 |
| 2023 | Inequality phenomenon in l∞-adversarial training, and its unrealized threats
Ranjie Duan, Yuefeng Chen, Yao Zhu 0003, Xiaojun Jia, Rong Zhang 0006, Hui Xue 0001 |
ICLR | 4 |
| 2023 | Robust Automatic Speech Recognition via WavAugment Guided Phoneme Adversarial Training
Gege Qi, Yuefeng Chen, Xiaofeng Mao, Xiaojun Jia, Ranjie Duan, Rong Zhang 0006, Hui Xue 0001 |
INTERSPEECH | 4 |
| 2023 | Hi-SIGIR: Hierachical Semantic-Guided Image-to-image Retrieval via Scene GraphabstractImage-to-image retrieval, a fundamental task, aims at matching similar images based on a query image. Existing methods with convolutional neural networks are usually sensitive to low-level visual features, and ignore high-level semantic relationship information. This makes retrieving complicated images with multiple objects and various relationships a significant challenge. Although some works introduce the scene graph to capture the global semantic features of the objects and their relations, they ignore the local visual representations. In addition, due to the fragility of individual modal representations, poisoning attacks in adversarial scenarios are easily achieved, hurting the robustness of the visual-guided foundation image retrieval model. To overcome these issues, we propose a novel hierarchical semantic-guided image-to-image retrieval method via scene graph, called Hi-SIGIR. Specifically, to begin with, our proposed method generates the scene graph of an image. Then, our model extracts and learns both the visual and semantic features of the nodes and relations within the scene graphs. Next, these features are fused to obtain local information and sent to the graph neural network to obtain global information. Using these information, the similarity between the scene graphs of several images is calculated at both the local and global levels to perform image retrieval. Finally, we introduce a surrogate that calculates relevance in a cross-modal manner to understand image content better. Experimental evaluations on several wildly-used benchmarks demonstrate the superiority of the proposed method. Yulu Wang, Pengwen Dai, Xiaojun Jia, Zhitao Zeng, Rui Li 0109, Xiaochun Cao |
ACM Multimedia | 3 |
| 2023 | Interpretable Inference and Classification of Tissue Types in Histological Colorectal Cancer Slides Based on Ensembles Adaptive Boosting Prototype TreeabstractDigital pathology images are treated as the "gold standard" for the diagnosis of colorectal lesions, especially colon cancer. Real-time, objective and accurate inspection results will assist clinicians to choose symptomatic treatment in a timely manner, which is of great significance in clinical medicine. However, Manual methods suffers from long inspection cycle and serious reliance on subjective interpretation. It is also a challenging task for existing computer-aided diagnosis methods to obtain models that are both accurate and interpretable. Models that exhibit high accuracy are always more complex and opaque, while interpretable models may lack the necessary accuracy. Therefore, the framework of ensemble adaptive boosting prototype tree is proposed to predict the colorectal pathology images and provide interpretable inference by visualizing the decision-making process in each base learner. The results showed that the proposed method could effectively address the "accuracy-interpretability trade-off" issue by ensemble of m adaptive boosting neural prototype trees. The superior performance of the framework provides a novel paradigm for interpretable inference and high-precision prediction of pathology image patches in computational pathology. Meiyan Liang, Jianan Liang, Lin Wang 0110, Xiaojun Jia, Qinghui Chen, Cunlin Zhang |
IEEE J. Biomed. Health Informatics | 6 |
| 2022 | Defending against Model Stealing via Verifying Embedded External FeaturesabstractObtaining a well-trained model involves expensive data collection and training procedures, therefore the model is a valuable intellectual property. Recent studies revealed that adversaries can `steal' deployed models even when they have no training samples and can not get access to the model parameters or structures. Currently, there were some defense methods to alleviate this threat, mostly by increasing the cost of model stealing. In this paper, we explore the defense from another angle by verifying whether a suspicious model contains the knowledge of defender-specified external features. Specifically, we embed the external features by tempering a few training samples with style transfer. We then train a meta-classifier to determine whether a model is stolen from the victim. This approach is inspired by the understanding that the stolen models should contain the knowledge of features learned by the victim model. We examine our method on both CIFAR-10 and ImageNet datasets. Experimental results demonstrate that our method is effective in detecting different types of model stealing simultaneously, even if the stolen model is obtained via a multi-stage stealing process. The codes for reproducing main results are available at Github (https://github.com/zlh-thu/StealingVerification). Yiming Li 0004, Linghui Zhu, Xiaojun Jia, Yong Jiang 0001, Shutao Xia, Xiaochun Cao |
AAAI | 3 |
| 2022 | LAS-AT: Adversarial Training with Learnable Attack StrategyabstractAdversarial training (AT) is always formulated as a minimax problem, of which the performance depends on the inner optimization that involves the generation of adver-sarial examples (AEs). Most previous methods adopt Projected Gradient Decent (PGD) with manually specifying attack parameters for AE generation. A combination of the attack parameters can be referred to as an attack strategy. Several works have revealed that using a fixed attack strategy to generate AEs during the whole training phase limits the model robustness and propose to exploit different attack strategies at different training stages to improve robustness. But those multi-stage handcrafted attack strategies need much domain expertise, and the robustness improvement is limited. In this paper, we propose a novel framework for adversarial training by introducing the concept of “learnable attack strategy”, dubbed LAS-AT, which learns to automatically produce attack strategies to improve the model robustness. Our framework is composed of a target network that uses AEs for training to improve robustness, and a strategy network that produces attack strategies to control the AE generation. Experimental evaluations on three benchmark databases demonstrate the superiority of the proposed method. The code is released at https://github.com/jiaxiaojunQAQ/LAS-AT. Xiaojun Jia, Yong Zhang 0034, Baoyuan Wu, Ke Ma 0001, Jue Wang 0001, Xiaochun Cao |
CVPR | 1 |
| 2022 | Prior-Guided Adversarial Initialization for Fast Adversarial Training
Xiaojun Jia, Yong Zhang 0034, Xingxing Wei 0001, Baoyuan Wu, Ke Ma 0001, Jue Wang 0001, Xiaochun Cao |
ECCV (4) | 1 |
| 2022 | A Large-Scale Multiple-objective Method for Black-box Attack Against Object Detection
Siyuan Liang 0004, Longkang Li, Yanbo Fan, Xiaojun Jia, Jingzhi Li 0002, Baoyuan Wu, Xiaochun Cao |
ECCV (4) | 4 |
| 2022 | Watermark Vaccine: Adversarial Attacks to Prevent Watermark Removal
Jian Liu 0012, Yang Bai 0011, Jindong Gu, Xiaojun Jia, Xiaochun Cao |
ECCV (14) | 6 |
| 2022 | Boosting Fast Adversarial Training With Learnable Adversarial InitializationabstractAdversarial training (AT) has been demonstrated to be effective in improving model robustness by leveraging adversarial examples for training. However, most AT methods are in face of expensive time and computational cost for calculating gradients at multiple steps in generating adversarial examples. To boost training efficiency, fast gradient sign method (FGSM) is adopted in fast AT methods by calculating gradient only once. Unfortunately, the robustness is far from satisfactory. One reason may arise from the initialization fashion. Existing fast AT generally uses a random sample-agnostic initialization, which facilitates the efficiency yet hinders a further robustness improvement. Up to now, the initialization in fast AT is still not extensively explored. In this paper, focusing on image classification, we boost fast AT with a sample-dependent adversarial initialization, i.e., an output from a generative network conditioned on a benign image and its gradient information from the target network. As the generative network and the target network are optimized jointly in the training phase, the former can adaptively generate an effective initialization with respect to the latter, which motivates gradually improved robustness. Experimental evaluations on four benchmark databases demonstrate the superiority of our proposed method over state-of-the-art fast AT methods, as well as comparable robustness to advanced multi-step AT methods. The code is released at https://github.com//jiaxiaojunQAQ//FGSM-SDI. Xiaojun Jia, Yong Zhang 0034, Baoyuan Wu, Jue Wang 0001, Xiaochun Cao |
IEEE Trans. Image Process. | 1 |
| 2021 | Applying BERT to analyze investor sentiment in stock market
Menggang Li, Xiaojun Jia, Guangwei Rui |
Neural Comput. Appl. | 4 |
| 2021 | Multi-source data fusion for economic data analysis
Menggang Li, Fang Wang 0021, Xiaojun Jia, Ting Li 0013, Guangwei Rui |
Neural Comput. Appl. | 3 |
| 2021 | A novel dual-biological-community swarm intelligence algorithm with a commensal evolution strategy for multimodal problems
Xiaochen Shen, Xiaojun Jia |
J. Supercomput. | 3 |
| 2020 | Adv-watermark: A Novel Watermark Perturbation for Adversarial ExamplesabstractRecent research has demonstrated that adding some imperceptible perturbations to original images can fool deep learning models. However, the current adversarial perturbations are usually shown in the form of noises, and thus have no practical meaning. Image watermark is a technique widely used for copyright protection. We can regard image watermark as a king of meaningful noises and adding it to the original image will not affect people's understanding of the image content, and will not arouse people's suspicion. Therefore, it will be interesting to generate adversarial examples using watermarks. In this paper, we propose a novel watermark perturbation for adversarial examples (Adv-watermark) which combines image watermarking techniques and adversarial example algorithms. Adding a meaningful watermark to the clean images can attack the DNN models. Specifically, we propose a novel optimization algorithm, which is called Basin Hopping Evolution (BHE), to generate adversarial watermarks in the black-box attack mode. Thanks to the BHE, Adv-watermark only requires a few queries from the threat models to finish the attacks. A series of experiments conducted on ImageNet and CASIA-WebFace datasets show that the proposed method can efficiently generate adversarial examples, and outperforms the state-of-the-art attack methods. Moreover, Adv-watermark is more robust against image transformation defense methods. Xiaojun Jia, Xingxing Wei 0001, Xiaochun Cao, Xiaoguang Han 0001 |
ACM Multimedia | 1 |
| 2020 | Quantum network based on non-classical light
Xiaolong Su, Meihong Wang, Zhihui Yan, Xiaojun Jia, Changde Xie, Kunchi Peng |
Sci. China Inf. Sci. | 4 |
| 2019 | ComDefend: An Efficient Image Compression Model to Defend Adversarial ExamplesabstractDeep neural networks (DNNs) have been demonstrated to be vulnerable to adversarial examples. Specifically, adding imperceptible perturbations to clean images can fool the well trained deep neural networks. In this paper, we propose an end-to-end image compression model to defend adversarial examples: ComDefend. The proposed model consists of a compression convolutional neural network (ComCNN) and a reconstruction convolutional neural network (ResCNN). The ComCNN is used to maintain the structure information of the original image and purify adversarial perturbations. And the ResCNN is used to reconstruct the original image with high quality. In other words, ComDefend can transform the adversarial image to its clean version, which is then fed to the trained classifier. Our method is a pre-processing module, and does not modify the classifier's structure during the whole process. Therefore it can be combined with other model-specific defense models to jointly improve the classifier's robustness. A series of experiments conducted on MNIST, CIFAR10 and ImageNet show that the proposed method outperforms the state-of-the-art defense methods, and is consistently effective to protect classifiers against adversarial attacks. Xiaojun Jia, Xingxing Wei 0001, Xiaochun Cao, Hassan Foroosh |
CVPR | 1 |
| 2005 | An Encoded Mini-grid Structured Light Pattern for Dynamic Scenes
Qingcang Yu, Xiaojun Jia |
ICIC (1) | 2 |