EDBT 2026 Demo / reviewers in the wild / expert
Siyuan Liang 0004
dblp:205/8767-4
· DBLP profile ↗
59ranked-venue papers
8as first author
57since 2021 · last 2026
0000-0002-6154-0233ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 6 first-author · 34 since 2021Graphics, computer vision, multimedia, augmented reality and games · 31 · 6 first-author · 29 since 2021Security and privacy · 11 · 1 first-author · 11 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adversarial Generation and Collaborative Evolution of Safety-Critical Scenarios for Autonomous VehiclesabstractThe generation of safety-critical scenarios in simulation has become increasingly crucial for safety evaluation in autonomous vehicles (AV) prior to road deployment in society. However, current approaches largely rely on predefined threat patterns or rule-based strategies, which limit their ability to expose diverse and unforeseen failure modes. To overcome these, we propose ScenGE, a framework that can generate plentiful safety-critical scenarios by reasoning novel adversarial cases and then amplifying them with complex traffic flows. Given a simple prompt of a benign scene, it first performs Meta-Scenario Generation, where a large language model (LLM), grounded in structured driving knowledge (e.g., traffic regulations, real-world accident records), infers an adversarial agent whose behavior poses a threat that is both plausible and deliberately challenging. This meta-scenario is then specified in executable code for precise in-simulator control. Subsequently, Complex Scenario Evolution uses background vehicles to amplify the core threat introduced by Meta-Scenario. It builds an adversarial collaborator graph to identify key agent trajectories for optimization. These perturbations are designed to simultaneously reduce the ego vehicle's maneuvering space and create critical occlusions. Extensive experiments conducted on multiple reinforcement learning (RL) based AV models show that ScenGE uncovers more severe collision cases (+31.96%) on average than SoTA baselines. Additionally, our ScenGE can be applied to large model based AV systems and deployed on different simulators; we further observe that adversarial training on our scenarios improves the model robustness. We hope our paper can build up a critical step towards building public trust and ensuring their safe deployment. Jiangfan Liu 0001, Yongkang Guo, Fangzhi Zhong, Tianyuan Zhang 0004, Zonglei Jing, Siyuan Liang 0004, Jiakai Wang, Mingchuan Zhang, Aishan Liu, Xianglong Liu 0001 |
AAAI | 6 |
| 2026 | SRD: Reinforcement-Learned Semantic Perturbation for Backdoor Defense in VLMsabstractVisual language models (VLMs) have made significant progress in image captioning tasks, yet recent studies have found they are vulnerable to backdoor attacks. Attackers can inject undetectable perturbations into the data during inference, triggering abnormal behavior and generating malicious captions. These attacks are particularly challenging to detect and defend against due to the stealthiness and cross-modal propagation of the trigger signals. In this paper, we identify two key vulnerabilities by analyzing existing attack patterns: (1) the model exhibits abnormal attention concentration on certain regions of the input image, and (2) backdoor attacks often induce semantic drift and sentence incoherence. Based on these insights, we propose Semantic Reward Defense (SRD), a reinforcement learning framework that mitigates backdoor behavior without requiring any prior knowledge of trigger patterns. SRD learns to apply discrete perturbations to sensitive contextual regions of image inputs via a deep Q-network policy, aiming to confuse attention and disrupt the activation of malicious paths. To guide policy optimization, we design a reward signal named semantic fidelity score, which jointly assesses the semantic consistency and linguistic fluency of the generated captions, encouraging the agent to achieve a robust yet faithful output. SRD offers a trigger-agnostic, policy-interpretable defense paradigm that effectively mitigates local (TrojVLM) and global (Shadowcast) backdoor attacks, reducing ASR to 3.6% and 5.6% respectively, with less than 15% average CIDEr drop on the clean inputs. Shuhan Xu, Siyuan Liang 0004, Hongling Zheng, Aishan Liu, Xinbiao Wang, Yong Luo 0002, Leszek Rutkowski, Dacheng Tao |
AAAI | 2 |
| 2026 | Controllable Contamination Detection for Reliable LLM Evaluation with Statistical GuaranteesabstractZheng Zhang, Qi Liu, Siyuan Liang, Ning Li, Zirui Hu, Weibo Gao, Rui Li, Zhenya Huang, Leszek Rutkowski, Baosheng Yu, Dacheng Tao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zheng Zhang 0048, Qi Liu 0003, Siyuan Liang 0004, Ning Li 0055, Zirui Hu, Weibo Gao, Rui Li 0093, Zhenya Huang, Leszek Rutkowski, Baosheng Yu, Dacheng Tao |
ACL (1) | 3 |
| 2026 | T2VShield: Model-Agnostic Jailbreak Defense for Text-to-Video Models
Siyuan Liang 0004, Jiecheng Zhai, Tianmeng Fang, Rongcheng Tu, Aishan Liu, Xiaochun Cao, Dacheng Tao |
Int. J. Comput. Vis. | 1 |
| 2026 | SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models
Zonghao Ying, Aishan Liu, Siyuan Liang 0004, Lei Huang 0015, Jinyang Guo 0002, Wenbo Zhou 0004, Xianglong Liu 0001, Dacheng Tao |
Int. J. Comput. Vis. | 3 |
| 2026 | WFCAT: Augmenting Website Fingerprinting With Channel-Wise Attention on Timing FeaturesabstractWebsite Fingerprinting (WF) aims to deanonymize users on the Tor network by analyzing encrypted network traffic. Recent deep-learning-based attacks show high accuracy on undefended traces. However, they struggle against modern defenses that use tactics like injecting dummy packets and delaying real packets, which significantly degrade classification performance. Our analysis reveals that current attacks inadequately leverage the timing information inherent in traffic traces, which persists as a source of leakage even under robust defenses. Addressing this shortfall, we introduce a novel feature representation named the Inter-Arrival Time (IAT) histogram, which quantifies the frequencies of packet inter-arrival times across predetermined time slots. Complementing this feature, we propose a new CNN-based attack, WFCAT, enhanced with two architectural blocks designed to effectively extract and utilize timing information. The model employs convolutional kernels of varying sizes to capture multi-scale temporal features, which are then integrated through a weighted combination across feature channels. This channel-wise attention mechanism enables the model to adaptively emphasize informative patterns while suppressing noise, thereby improving its robustness against timing obfuscation. Our experiments validate that WFCAT substantially outperforms existing methods on defended traces in both closed- and open-world scenarios. Notably, WFCAT achieves over 59% accuracy against Surakav, a recently developed robust defense, marking an improvement of over 28% and 48% against the state-of-the-art attacks RF and Tik-Tok, respectively, in the closed-world scenario. Jiajun Gong, Siyuan Liang 0004, Tao Wang 0012, Ee-Chien Chang |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2026 | CogMorph: Cognitive Morphing Attacks for Text-to-Image ModelsabstractThe development of text-to-image (T2I) generative models, that enable the creation of high-quality synthetic images from textual prompts, has opened new frontiers in creative design and content generation. However, this paper reveals a significant and previously unrecognized ethical risk inherent in this technology and introduces a novel method, termed the Cognitive Morphing Attack(CogMorph), which manipulates T2I models to generate images that retain the original core subjects but embeds toxic or harmful contextual elements. This nuanced manipulation exploits the cognitive principle that human perception of concepts is shaped by the entire visual scene and its context, producing images that amplify emotional harm far beyond attacks that merely preserve the original semantics. To address this, we first construct an imagery toxicity taxonomy spanning 10 major and 48 sub-categories, aligned with human cognitive-perceptual dimensions, and further build a toxicity risk matrix resulting in 1,176 high-quality T2I toxic prompts. Based on this, ourCogMorphfirst introduces Cognitive Toxicity Augmentation, which develops a cognitive toxicity knowledge base with rich external toxic representations for humans (e.g., fine-grained visual features) that can be utilized to further guide the optimization of adversarial prompts. In addition, we present Contextual Hierarchical Morphing, which hierarchically extracts critical parts of the original prompt (e.g., scenes, subjects, and body parts), and then iteratively retrieves and fuses toxic features to inject harmful contexts. Extensive experiments on multiple open-source T2I models and black-box commercial APIs (e.g., DALL$\cdot$E-3) demonstrate the efficacy ofCogMorphwhich significantly outperforms other baselines by large margins (+20.62% on average). Our codes are available athttps://github.com/raykr/CogMorph.Warning: This paper contains harmful imagery that might be offensive to some readers. Zonglei Jing, Zonghao Ying, Le Wang 0014, Siyuan Liang 0004, Mingchuan Zhang, Aishan Liu, Xianglong Liu 0001, Dacheng Tao |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2026 | TrapFlow: Controllable Website Fingerprinting Defense via Dynamic Backdoor LearningabstractWebsite fingerprinting (WF) attacks, which covertly monitor user communications to identify the web pages they visit, pose a serious threat to user privacy. Existing WF defenses attempt to reduce attack accuracy by disrupting traffic patterns, but attackers can retrain their models to adapt, making these defenses ineffective. Meanwhile, their high overhead limits deployability. To overcome these limitations, we introduce a novel controllable website fingerprinting defense called TrapFlow based on backdoor learning. TrapFlow exploits the tendency of neural networks to memorize subtle patterns by injecting crafted trigger sequences into targeted website traffic, causing the attacker’s model to build incorrect associations during training. If the attacker attempts to adapt by training on such noisy data, TrapFlow ensures that the model internalizes the trigger as a dominant feature, leading to widespread misclassification across unrelated websites. Conversely, if the attacker ignores these patterns and trains only on clean data, the trigger behaves as an adversarial patch at inference time, causing model misclassification. To achieve this dual effect, we optimize the trigger using the Fast Levenshtein-like distance to maximize both its learnability and distinctiveness from normal traffic. Experiments show that TrapFlow significantly reduces the accuracy of the RF attack from 99% to 6% with 74% data overhead. This compares favorably against two SOTA defenses: FRONT reduces accuracy by only 2% at a similar overhead, while Palette achieves 32% accuracy, but with 48% more overhead. We further validate the practicality of our method in a real Tor network environment. Siyuan Liang 0004, Jiajun Gong, Tianmeng Fang, Aishan Liu, Tao Wang 0012, Xiaochun Cao, Dacheng Tao, Ee-Chien Chang |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2026 | SafeSteer: Adaptive Subspace Steering for Efficient Jailbreak Defense in Vision Language Models
Shuchao Pang, Xiyu Zeng, Siyuan Liang 0004, Chuanting Zhang, Enguang Liu, Basem Shihada, Yongbin Zhou, Minhui Xue 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | ELBA-Bench: An Efficient Learning Backdoor Attacks Benchmark for Large Language ModelsabstractXuxu Liu, Siyuan Liang, Mengya Han, Yong Luo, Aishan Liu, Xiantao Cai, Zheng He, Dacheng Tao. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xuxu Liu, Siyuan Liang 0004, Mengya Han, Yong Luo 0002, Aishan Liu, Xiantao Cai, Zheng He 0001, Dacheng Tao |
ACL (1) | 2 |
| 2025 | Interpreting Object-level Foundation Models via Visual Precision SearchabstractAdvances in multimodal pre-training have propelled object-level foundation models, such as Grounding DINO and Florence-2, in tasks like visual grounding and object detection. However, interpreting these models’ decisions has grown increasingly challenging. Existing interpretable attribution methods for object-level task interpretation have notable limitations: (1) gradient-based methods lack precise localization due to visual-textual fusion in foundation models, and (2) perturbation-based methods produce noisy saliency maps, limiting fine-grained interpretability. To address these, we propose a Visual Precision Search method that generates accurate attribution maps with fewer regions. Our method bypasses internal model parameters to overcome attribution issues from multimodal fusion, dividing inputs into sparse sub-regions and using consistency and collaboration scores to accurately identify critical decision-making regions. We also conducted a theoretical analysis of the boundary guarantees and scope of applicability of our method. Experiments on RefCOCO, MS COCO, and LVIS show our approach enhances object-level task interpretability over SOTA for Grounding DINO and Florence-2 across various evaluation metrics, with faithfulness gains of 23.7%, 31.6%, and 20.1% on MS COCO, LVIS, and RefCOCO for Grounding DINO, and 50.7% and 66.9% on MS COCO and RefCOCO for Florence-2. Additionally, our method can interpret failures in visual grounding and object detection tasks, surpassing existing methods across multiple evaluation metrics. The code is released at https://github.com/RuoyuChen10/VPS. Ruoyu Chen 0001, Siyuan Liang 0004, Jingzhi Li 0002, Shiming Liu, Maosen Li, Zhen Huang 0006, Hua Zhang 0008, Xiaochun Cao |
CVPR | 2 |
| 2025 | Revisiting Backdoor Attacks against Large Vision-Language Models from Domain ShiftabstractInstruction tuning enhances large vision-language models (LVLMs) but increases their vulnerability to backdoor attacks due to their open design. Unlike prior studies in static settings, this paper explores backdoor attacks in LVLM instruction tuning across mismatched training and testing domains. We introduce a new evaluation dimension, backdoor domain generalization, to assess attack robustness under visual and text domain shifts. Our findings reveal two insights: (1) backdoor generalizability improves when distinctive trigger patterns are independent of specific data domains or model architectures, and (2) the competitive interaction between trigger patterns and clean semantic regions, where guiding the model to predict triggers enhances attack generalizability. Based on these insights, we propose a multimodal attribution backdoor attack (MABA) that injects domain-agnostic triggers into critical areas using attributional interpretation. Experiments with OpenFlamingo, Blip-2, and Otter show that MABA significantly boosts the attack success rate of generalization by 36.4% over the unimodal attack, achieving a 97% success rate at a 0.2% poisoning rate. This study reveals limitations in current evaluations and highlights how enhanced backdoor generalizability poses a security threat to LVLMs, even without test data access. Our codes are available online1. Siyuan Liang 0004, Tianyu Pang, Aishan Liu, Mingli Zhu, Xiaochun Cao, Dacheng Tao |
CVPR | 1 |
| 2025 | CopyrightShield: Enhancing Diffusion Model Security Against Copyright Infringement Attacks
Zhixiang Guo, Siyuan Liang 0004, Aishan Liu, Dacheng Tao |
ICCV | 2 |
| 2025 | Gradient-Reweighted Adversarial Camouflage for Physical Object Detection Evasion
Siyuan Liang 0004, Tianrui Lou, Wenjin Li, Dunqiu Fan, Xiaochun Cao |
ICCV | 2 |
| 2025 | 3D Gaussian Splatting Driven Multi-View Robust Physical Adversarial Camouflage Generation
Tianrui Lou, Xiaojun Jia, Siyuan Liang 0004, Xiaochun Cao |
ICCV | 3 |
| 2025 | Towards a 3D Transfer-Based Black-Box Attack via Critical Feature GuidanceabstractDeep neural networks for 3D point clouds have been demonstrated to be vulnerable to adversarial examples. Previous 3D adversarial attack methods often exploit certain information about the target models, such as model parameters or outputs, to generate adversarial point clouds. However, in realistic scenarios, it is challenging to obtain any information about the target models under conditions of absolute security. Therefore, we focus on transfer-based attacks, where generating adversarial point clouds does not require any information about the target models. Based on our observation that the critical features used for point cloud classification are consistent across different DNN architectures, we propose CFG, a novel transfer-based black-box attack method that improves the transferability of adversarial point clouds via the proposed Critical Feature Guidance. Specifically, our method regularizes the search of adversarial point clouds by computing the importance of the extracted features, prioritizing the corruption of critical features that are likely to be adopted by diverse architectures. Further, we explicitly constrain the maximum deviation extent of the generated adversarial point clouds in the loss function to ensure their imperceptibility. Extensive experiments conducted on the ModelNet40 and ScanObjectNN benchmark datasets demonstrate that the proposed CFG outperforms the state-of-the-art attack methods by a large margin. Shuchao Pang, Zhenghan Chen, Siyuan Liang 0004, Anan Du, Yongbin Zhou |
ICCV | 5 |
| 2025 | NoVo: Norm Voting off Hallucinations with Attention Heads in Large Language ModelsabstractHallucinations in Large Language Models (LLMs) remain a major obstacle, particularly in high-stakes applications where factual accuracy is critical. While representation editing and reading methods have made strides in reducing hallucinations, their heavy reliance on specialised tools and training on in-domain samples, makes them difficult to scale and prone to overfitting. This limits their accuracy gains and generalizability to diverse datasets. This paper presents a lightweight method, Norm Voting (NoVo), which harnesses the untapped potential of attention head norms to dramatically enhance factual accuracy in zero-shot multiple-choice questions (MCQs). NoVo begins by automatically selecting truth-correlated head norms with an efficient, inference-only algorithm using only 30 random samples, allowing NoVo to effortlessly scale to diverse datasets. Afterwards, selected head norms are employed in a simple voting algorithm, which yields significant gains in prediction accuracy. On TruthfulQA MC1, NoVo surpasses the current state-of-the-art and all previous methods by an astounding margin---at least 19 accuracy points. NoVo demonstrates exceptional generalization to 20 diverse datasets, with significant gains in over 90\% of them, far exceeding all current representation editing and reading methods. NoVo also reveals promising gains to finetuning strategies and building textual adversarial defence. NoVo's effectiveness with head norms opens new frontiers in LLM interpretability, robustness and reliability. Our code is available at: https://github.com/hozhengyi/novo Zheng Yi Ho, Siyuan Liang 0004, Sen Zhang 0006, Yibing Zhan, Dacheng Tao |
ICLR | 2 |
| 2025 | ICLShield: Exploring and Mitigating In-Context Learning Backdoor AttacksabstractIn-context learning (ICL) has demonstrated remarkable success in large language models (LLMs) due to its adaptability and parameter-free nature. However, it also introduces a critical vulnerability to backdoor attacks, where adversaries can manipulate LLM behaviors by simply poisoning a few ICL demonstrations. In this paper, we propose, for the first time, the dual-learning hypothesis, which posits that LLMs simultaneously learn both the task-relevant latent concepts and backdoor latent concepts within poisoned demonstrations, jointly influencing the probability of model outputs. Through theoretical analysis, we derive an upper bound for ICL backdoor effects, revealing that the vulnerability is dominated by the concept preference ratio between the task and the backdoor. Motivated by these findings, we propose ICLShield, a defense mechanism that dynamically adjusts the concept preference ratio. Our method encourages LLMs to select clean demonstrations during the ICL phase by leveraging confidence and similarity scores, effectively mitigating susceptibility to backdoor attacks. Extensive experiments across multiple LLMs and tasks demonstrate that our method achieves state-of-the-art defense effectiveness, significantly outperforming existing approaches (+26.02% on average). Furthermore, our method exhibits exceptional adaptability and defensive performance even for closed-source models (e.g., GPT-4). Zhiyao Ren, Siyuan Liang 0004, Aishan Liu, Dacheng Tao |
ICML | 2 |
| 2025 | BDefects4NN: A Backdoor Defect Database for Controlled Localization Studies in Neural NetworksabstractPre-trained large deep learning models are now serving as the dominant component for downstream middleware users and have revolutionized the learning paradigm, replacing the traditional approach of training from scratch locally. To reduce development costs, developers often integrate third-party pre-trained deep neural networks (DNNs) into their intelligent software systems. However, utilizing untrusted DNNs presents significant security risks, as these models may contain intentional backdoor defects resulting from the black-box training process. These backdoor defects can be activated by hidden triggers, allowing attackers to maliciously control the model and compromise the overall reliability of the intelligent software. To ensure the safe adoption of DNNs in critical software systems, it is crucial to establish a backdoor defect database for localization studies. This paper addresses this research gap by introducing BDefects4NN, the first backdoor defect database, which provides labeled backdoor-defected DNNs at the neuron granularity and enables controlled localization studies of defect root causes. In BDefects$4 N N$, we define three defect injection rules and employ four representative backdoor attacks across four popular network architectures and three widely adopted datasets, yielding a comprehensive database of$\mathbf{1, 6 5 4}$backdoor-defected DNNs with four defect quantities and varying infected neurons. Based on BDefects4NN, we conduct extensive experiments on evaluating six fault localization criteria and two defect repair techniques, which show limited effectiveness for backdoor defects. Additionally, we investigate backdoor-defected models in practical scenarios, specifically in lane detection for autonomous driving and large language models (LLMs), revealing potential threats and highlighting current limitations in precise defect localization. This paper aims to raise awareness of the threats brought by backdoor defects in our community and inspire future advancements in fault localization methods. Yisong Xiao, Aishan Liu, Xinwei Zhang 0010, Tianyuan Zhang 0004, Tianlin Li, Siyuan Liang 0004, Xianglong Liu 0001, Yang Liu 0088, Dacheng Tao |
ICSE | 6 |
| 2025 | Physical Adversarial Camouflage Through Gradient Calibration and RegularizationabstractThe advancement of deep object detectors has greatly affected safety-critical fields like autonomous driving. However, physical adversarial camouflage poses a significant security risk by altering object textures to deceive detectors. Existing techniques struggle with variable physical environments, facing two main challenges: 1) inconsistent sampling point densities across distances hinder the gradient optimization from ensuring local continuity, and 2) updating texture gradients from multiple angles causes conflicts, reducing optimization stability and attack effectiveness. To address these issues, we propose a novel adversarial camouflage framework based on gradient optimization. First, we introduce a gradient calibration strategy, which ensures consistent gradient updates across distances by propagating gradients from sparsely to unsampled texture points, thereby expanding the attack's effective range. Additionally, we develop a gradient decorrelation method, which prioritizes and orthogonalizes gradients based on loss values, enhancing stability and effectiveness in multi-angle optimization by eliminating redundant or conflicting updates. Extensive experimental results on various detection models, angles, and distances show that our method significantly surpasses the state-of-the-art, with an average attack success rate (ASR) increase of 13.46\% across distances and 11.03\% across angles. Furthermore, experiments in real-world settings confirm the method's threat potential, highlighting the urgent need for more robust autopilot systems less prone to spoofing. Siyuan Liang 0004, Jianjie Huang, Chenxi Si, Xiaochun Cao |
IJCAI | 2 |
| 2025 | MetAdv: A Unified and Interactive Adversarial Testing Platform for Autonomous DrivingabstractEvaluating and ensuring the adversarial robustness of autonomous driving (AD) systems is a critical and unresolved challenge. This paper introduces MetAdv, a novel adversarial testing platform that enables realistic, dynamic, and interactive evaluation by tightly integrating virtual simulation with physical vehicle feedback. At its core, MetAdv establishes a hybrid virtual-physical sandbox, within which we design a three-layer closed-loop testing environment with dynamic adversarial test evolution. This architecture facilitates end-to-end adversarial evaluation, ranging from high-level unified adversarial generation, through mid-level simulation-based interaction, to low-level execution on physical vehicles. Additionally, MetAdv supports a broad spectrum of AD tasks, algorithmic paradigms (e.g., modular deep learning pipelines, end-to-end learning, vision-language models). It supports flexible 3D vehicle modeling and seamless transitions between simulated and physical environments, with built-in compatibility for commercial platforms such as Apollo and Tesla. A key feature of MetAdv is its human-in-the-loop capability: besides flexible environmental configuration for more customized evaluation, it enables real-time capture of physiological signals and behavioral feedback from drivers, offering new insights into human-machine trust under adversarial conditions. We believe MetAdv can offer a scalable and unified framework for adversarial assessment, paving the way for safer AD. Our demo can be found at https://sites.google.com/view/metadv-demo-video. Aishan Liu, Jiakai Wang, Tianyuan Zhang 0004, Hainan Li, Jiangfan Liu 0001, Siyuan Liang 0004, Yilong Ren, Xianglong Liu 0001, Dacheng Tao |
ACM Multimedia | 6 |
| 2025 | Manipulating Multimodal Agents via Cross-Modal Prompt InjectionabstractThe emergence of multimodal large language models has redefined the agent paradigm by integrating language and vision modalities with external data sources, enabling agents to better interpret human instructions and execute increasingly complex tasks. However, in this paper, we identify a critical yet previously overlooked security vulnerability in multimodal agents: cross-modal prompt injection attacks. To exploit this vulnerability, we propose CrossInject, a novel attack framework in which attacker embeds adversarial perturbations across multiple modalities to align with target malicious content, allowing external instructions to hijack the agents' decision-making process and execute unauthorized tasks. Our approach incorporates two key coordinated components. First, we introduce Visual Latent Alignment, where we optimize adversarial features to the malicious instructions in the visual embedding space based on a text-to-image generative model, ensuring that adversarial images subtly encode cues for malicious task execution. Subsequently, we present Textual Guidance Enhancement, where a large language model is leveraged to construct the black-box defensive system prompt through adversarial meta-prompting and generate a malicious textual command based on it that steers the agents' output toward better compliance with attacker's requests. Extensive experiments demonstrate that our method outperforms state-of-the-art attacks, achieving at least a +30.1% increase in attack success rates across diverse tasks. Furthermore, we validate our attack's effectiveness in real-world multimodal autonomous agents, highlighting its potential implications for safety-critical applications. Code can be found in https://github.com/Larry0454/CrossInject. Le Wang 0014, Zonghao Ying, Tianyuan Zhang 0004, Siyuan Liang 0004, Shengshan Hu, Mingchuan Zhang, Aishan Liu, Xianglong Liu 0001 |
ACM Multimedia | 4 |
| 2025 | T2V-OptJail: Discrete Prompt Optimization for Text-to-Video Jailbreak AttacksabstractIn recent years, fueled by the rapid advancement of diffusion models, text-to-video (T2V) generation models have achieved remarkable progress, with notable examples including Pika, Luma, Kling, and Open-Sora. Although these models exhibit impressive generative capabilities, they also expose significant security risks due to their vulnerability to jailbreak attacks, where the models are manipulated to produce unsafe content such as pornography, violence, or discrimination. Existing works such as T2VSafetyBench provide preliminary benchmarks for safety evaluation, but lack systematic methods for thoroughly exploring model vulnerabilities.
To address this gap, we are the first to formalize the T2V jailbreak attack as a discrete optimization problem and propose a joint objective-based optimization framework, called \emph{T2V-OptJail}. This framework consists of two key optimization goals: bypassing the built-in safety filtering mechanisms to increase the attack success rate, preserving semantic consistency between the adversarial prompt and the unsafe input prompt, as well as between the generated video and the unsafe input prompt, to enhance content controllability. In addition, we introduce an iterative optimization strategy guided by prompt variants, where multiple semantically equivalent candidates are generated in each round, and their scores are aggregated to robustly guide the search toward optimal adversarial prompts.
We conduct large-scale experiments on several T2V models, covering both open-source models (\textit{e.g.}, Open-Sora) and real commercial closed-source models (\textit{e.g.}, Pika, Luma, Kling). The experimental results show that the proposed method improves 11.4\% and 10.0\% over the existing state-of-the-art method (SoTA) in terms of attack success rate assessed by GPT-4, attack success rate assessed by human accessors, respectively, verifying the significant advantages of the method in terms of attack effectiveness and content control. This study reveals the potential abuse risk of the semantic alignment mechanism in the current T2V model and provides a basis for the design of subsequent jailbreak defense methods. Siyuan Liang 0004, Shiqian Zhao, Rongcheng Tu, Wenbo Zhou 0004, Aishan Liu, Dacheng Tao, Siew-Kei Lam |
NeurIPS | 2 |
| 2025 | Lie Detector: Unified Backdoor Detection via Cross-Examination FrameworkabstractInstitutions with limited data and computing resources often outsource model training to third-party providers in a semi-honest setting, assuming adherence to prescribed training protocols with pre-defined learning paradigm (e.g., supervised or semi-supervised learning). However, this practice can introduce severe security risks, as adversaries may poison the training data to embed backdoors into the resulting model. Existing detection approaches predominantly rely on statistical analyses, which often fail to maintain universally accurate detection accuracy across different learning paradigms. To address this challenge, we propose a unified backdoor detection framework in the semi-honest setting that exploits cross-examination of model inconsistencies between two independent service providers. Specifically, we integrate central kernel alignment to enable robust feature similarity measurements across different model architectures and learning paradigms, thereby facilitating precise recovery and identification of backdoor triggers. We further introduce backdoor fine-tuning sensitivity analysis to distinguish backdoor triggers from adversarial perturbations, substantially reducing false positives. Extensive experiments demonstrate that our method achieves superior detection performance, improving accuracy by 4.4%, 1.7%, and 10.6% over SoTA baselines across supervised, self-supervised, and autoregressive learning tasks, respectively. Notably, it is the first to effectively detect backdoors in multimodal large language models, further highlighting its broad applicability and advancing secure deep learning. Xuan Wang 0029, Siyuan Liang 0004, Dongping Liao, Aishan Liu, Xiaochun Cao, Yuliang Lu, Ee-Chien Chang |
NeurIPS | 2 |
| 2025 | Detoxifying Large Language Models via Autoregressive Reward Guided Representation EditingabstractLarge Language Models (LLMs) have demonstrated impressive performance across various tasks, yet they remain vulnerable to generating toxic content, necessitating detoxification strategies to ensure safe and responsible deployment. Test-time detoxification methods, which typically introduce static or dynamic interventions into LLM representations, offer a promising solution due to their flexibility and minimal invasiveness. However, current approaches often suffer from imprecise interventions, primarily due to their insufficient exploration of the transition space between toxic and non-toxic outputs. To address this challenge, we propose \textsc{A}utoregressive \textsc{R}eward \textsc{G}uided \textsc{R}epresentation \textsc{E}diting (ARGRE), a novel test-time detoxification framework that explicitly models toxicity transitions within the latent representation space, enabling stable and precise reward-guided editing. ARGRE identifies non-toxic semantic directions and interpolates between toxic and non-toxic representations to reveal fine-grained transition trajectories. These trajectories transform sparse toxicity annotations into dense training signals, enabling the construction of an autoregressive reward model that delivers stable and precise editing guidance. At inference, the reward model guides an adaptive two-step editing process to obtain detoxified representations: it first performs directional steering based on expected reward gaps to shift representations toward non-toxic regions, followed by lightweight gradient-based refinements. Extensive experiments across 8 widely used LLMs show that ARGRE significantly outperforms leading baselines in effectiveness (-62.21\% toxicity) and efficiency (-47.58\% inference time), while preserving the core capabilities of the original model with minimal degradation. Our code is available at the \href{https://anonymous.4open.science/r/ARGRE-6291}{anonymous website}. Yisong Xiao, Aishan Liu, Siyuan Liang 0004, Zonghao Ying, Xianglong Liu 0001, Dacheng Tao |
NeurIPS | 3 |
| 2025 | Bridging the Task Gap: Multi-task Adversarial Transferability in CLIP and Its Derivatives
Kuanrong Liu, Siyuan Liang 0004, Xiaochun Cao |
PRCV (4) | 2 |
| 2025 | VL-Trojan: Multimodal Instruction Backdoor Attacks against Autoregressive Visual Language Models
Siyuan Liang 0004, Aishan Liu, Xiaochun Cao |
Int. J. Comput. Vis. | 2 |
| 2025 | Pre-trained Trojan Attacks for Visual Recognition
Aishan Liu, Xianglong Liu 0001, Xinwei Zhang 0010, Yisong Xiao, Yuguang Zhou, Siyuan Liang 0004, Jiakai Wang, Xiaochun Cao, Dacheng Tao |
Int. J. Comput. Vis. | 6 |
| 2025 | GenderBias-VL: Benchmarking Gender Bias in Vision Language Models via Counterfactual Probing
Yisong Xiao, Xianglong Liu 0001, QianJia Cheng, Zhenfei Yin, Siyuan Liang 0004, Aishan Liu, Dacheng Tao |
Int. J. Comput. Vis. | 5 |
| 2025 | An Intelligent Badminton Handle With Multinode MEMS Sensors for Explainable Motion RecognitionabstractIntelligent sensing technologies are transforming sports training by enabling precise motion analysis, critical for skill development and performance optimization. This study introduces a badminton racket handle embedded with a lightweight, multi-node MEMS-based sensing system designed for real-time motion recognition. To capture distributed grip forces, swing trajectories, and impact mechanics at the player-equipment interface, the system employs an ergonomic design ensuring natural gameplay. A hybrid feature extraction approach, integrating time-and frequency-domain features with a 1D-CNN, achieves a classification accuracy of 97.89% across ten badminton actions. To enhance interpretability and provide actionable insights, explainable AI using SMDL-attribution identifies key motion features, revealing biomechanical inefficiencies in grip strength, swing consistency, and wrist motion. Seamlessly integrated with Virtual Reality (VR) platforms, the system delivers immersive, real-time feedback, transforming training into an interactive and data-driven experience. By combining advanced sensing, machine learning, and explainable AI, this system establishes a new benchmark for intelligent sports monitoring, with broad applications in sports training, rehabilitation, and human-computer interaction. Jian Li 0063, Yibo Fan, Ruoyu Chen 0001, Siyuan Liang 0004, Yuliang Zhao |
IEEE Internet Things J. | 4 |
| 2025 | FOADA: Toward Robust Open-World Mobile App FingerprintingabstractSmartphone users are susceptible to a privacy leakage attack called App Fingerprinting (AF), where traffic analysis is used to infer the apps in use. Despite packet encryption, AF attacks leverage packet size and timing information to identify apps, posing a privacy threat. However, existing attacks fail when a few apps are used concurrently, causing unsegmented traffic with app multiplexing and overlapping. The key reason is that they cannot accurately identify active time boundaries for the apps. This paper presents a novel AF attack, FOADA, the first to accurately predict both the location and label of a target app in traffic. FOADA approaches AF as an object detection problem, training a deep learning model to estimate boundary positions and classify traffic segments. Accurate boundary predictions help the model focus on the most relevant traffic segment, enhancing its classification performance. FOADA excels in handling noisy app traffic. With app multiplexing, it achieves an F1-score of 0.96 for predicting only app labels and an F1-score of 0.92 for predicting both app labels and their locations. FOADA surpasses the state-of-the-art attack PacketPrint, which achieves F1-scores of 0.80 and 0.48 in these two scenarios, respectively. The inference time of FOADA is 2,000 times faster than PacketPrint. Jiajun Gong, Guotao Meng, Siyuan Liang 0004, Tao Wang 0012, Ee-Chien Chang |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | Hard-Label Black-Box Adversarial Attacks for Implicit Scene Interactions
Muxue Liang, Chuan Wang 0002, Siyuan Liang 0004, Aishan Liu, Yanan Cao 0006, Qingyong Li, Zeming Liu, Liang Yang 0002, Xiaochun Cao |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Compromising LLM Driven Embodied Agents With Contextual Backdoor Attacks
Aishan Liu, Yuguang Zhou, Xianglong Liu 0001, Tianyuan Zhang 0004, Siyuan Liang 0004, Jiakai Wang, Yanjun Pu, Tianlin Li, Wenbo Zhou 0004, Qing Guo 0005, Dacheng Tao |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | CleanerCLIP: Fine-Grained Counterfactual Semantic Augmentation for Backdoor Defense in Contrastive LearningabstractWith the rise of the open-source community, multimodal pre-trained models such as CLIP have become increasingly vulnerable to backdoor attacks. Backdoor triggers can manipulate model outputs during inference, posing a significant threat to downstream users. While post-training defenses based on fine-tuning have made some progress, their effectiveness remains limited due to two key challenges: (1) They fail to weaken the connection between the backdoor trigger and the target text space, making it possible for the model to still rely on the backdoor pattern for prediction. (2) Although batch-level fine-tuning expands the data distribution, it lacks precise guidance for vision-language alignment. To address these limitations, we propose a fine-grained counterfactual text-driven sample-level fine-tuning defense. By generating counterfactual sub-texts, we explicitly guide the text space to shift towards a more discriminative and robust representation, thereby indirectly weakening the association between backdoor triggers and target semantics. Furthermore, we introduce intra-sample contrastive learning with hard negative sub-texts, which enforces a more precise gradient direction to enhance vision-language fine-grained alignment. We evaluate our approach against six different backdoor attack methods and conduct a comprehensive zero-shot classification study on ImageNet-1K. Experimental results demonstrate that our method surpasses SoTA CleanCLIP in defending against BadCLIP attacks, reducing the attack success rate (ASR) in Top-1 classification by 52.02% and Top-10 classification by 63.88%. We aim to enhance the robustness of multimodal models against backdoor threats, fostering safer deployment in real-world applications. Yuan Xun, Siyuan Liang 0004, Xiaojun Jia, Jun Chen 0035, Xiaochun Cao |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Jailbreak Vision Language Models via Bi-Modal Adversarial PromptabstractIn the realm of large vision language models (LVLMs), jailbreak attacks serve as a red-teaming approach to bypass guardrails and uncover safety implications. Existing jailbreaks predominantly focus on the visual modality, perturbing solely visual inputs in the prompt for attacks. However, they fall short when confronted with aligned models that fuse visual and textual features simultaneously for generation. To address this limitation, this paper introduces the Bi-Modal Adversarial Prompt Attack (BAP), which executes jailbreaks by optimizing textual and visual prompts cohesively. Initially, we adversarially embed universally adversarial perturbations in an image, guided by a few-shot query-agnostic corpus (e.g., affirmative prefixes and negative inhibitions). This process ensures that the adversarial image prompt LVLMs to respond positively to harmful queries. Subsequently, leveraging the image, we optimize textual prompts with specific harmful intent. In particular, we utilize a large language model to analyze jailbreak failures and employ chain-of-thought reasoning to refine textual prompts through a feedback-iteration manner. To validate the efficacy of our approach, we conducted extensive evaluations on various datasets and LVLMs, demonstrating that our BAP significantly outperforms other methods by large margins (+29.03% in attack success rate on average). Additionally, we showcase the potential of our attacks on black-box commercial LVLMs, such as GPT-4o and Gemini. Zonghao Ying, Aishan Liu, Tianyuan Zhang 0004, Zhengmin Yu, Siyuan Liang 0004, Xianglong Liu 0001, Dacheng Tao |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | Learning to Optimize Permutation Flow Shop Scheduling via Graph-Based Imitation LearningabstractThe permutation flow shop scheduling (PFSS), aiming at finding the optimal permutation of jobs, is widely used in manufacturing systems. When solving large-scale PFSS problems, traditional optimization algorithms such as heuristics could hardly meet the demands of both solution accuracy and computational efficiency, thus learning-based methods have recently garnered more attention. Some work attempts to solve the problems by reinforcement learning methods, which suffer from slow convergence issues during training and are still not accurate enough regarding the solutions. To that end, we propose to train the model via expert-driven imitation learning, which accelerates convergence more stably and accurately. Moreover, in order to extract better feature representations of input jobs, we incorporate the graph structure as the encoder. The extensive experiments reveal that our proposed model obtains significant promotion and presents excellent generalizability in large-scale problems with up to 1000 jobs. Compared to the state-of-the-art reinforcement learning method, our model's network parameters are reduced to only 37% of theirs, and the solution gap of our model towards the expert solutions decreases from 6.8% to 1.3% on average. The code is available at: https://github.com/longkangli/PFSS-IL. Longkang Li, Siyuan Liang 0004, Zihao Zhu 0001, Chris Ding, Hongyuan Zha, Baoyuan Wu |
AAAI | 2 |
| 2024 | Does Few-Shot Learning Suffer from Backdoor Attacks?abstractThe field of few-shot learning (FSL) has shown promising results in scenarios where training data is limited, but its vulnerability to backdoor attacks remains largely unexplored. We first explore this topic by first evaluating the performance of the existing backdoor attack methods on few-shot learning scenarios. Unlike in standard supervised learning, existing backdoor attack methods failed to perform an effective attack in FSL due to two main issues. Firstly, the model tends to overfit to either benign features or trigger features, causing a tough trade-off between attack success rate and benign accuracy. Secondly, due to the small number of training samples, the dirty label or visible trigger in the support set can be easily detected by victims, which reduces the stealthiness of attacks. It seemed that FSL could survive from backdoor attacks. However, in this paper, we propose the Few-shot Learning Backdoor Attack (FLBA) to show that FSL can still be vulnerable to backdoor attacks. Specifically, we first generate a trigger to maximize the gap between poisoned and benign features. It enables the model to learn both benign and trigger features, which solves the problem of overfitting. To make it more stealthy, we hide the trigger by optimizing two types of imperceptible perturbation, namely attractive and repulsive perturbation, instead of attaching the trigger directly. Once we obtain the perturbations, we can poison all samples in the benign support set into a hidden poisoned support set and fine-tune the model on it. Our method demonstrates a high Attack Success Rate (ASR) in FSL tasks with different few-shot learning paradigms while preserving clean accuracy and maintaining stealthiness. This study reveals that few-shot learning still suffers from backdoor attacks, and its security should be given attention. Xiaojun Jia, Jindong Gu, Yuan Xun, Siyuan Liang 0004, Xiaochun Cao |
AAAI | 5 |
| 2024 | BadCLIP: Dual-Embedding Guided Backdoor Attack on Multimodal Contrastive LearningabstractWhile existing backdoor attacks have successfully infected multimodal contrastive learning models such as CLIP, they can be easily countered by specialized backdoor defenses for MCL models. This paper reveals the threats in this practical scenario and introduces the BadCLIP attack, which is resistant to backdoor detection and model fine-tuning defenses. To achieve this, we draw motivations from the perspective of the Bayesian rule and propose a dual-embedding guided framework for backdoor attacks. Specifically, we ensure that visual trigger patterns approximate the textual target semantics in the embedding space, making it challenging to detect the subtle parameter variations induced by backdoor learning on such natural trigger patterns. Additionally, we optimize the visual trigger patterns to align the poisoned samples with target vision features in order to hinder backdoor unlearning through clean fine-tuning. Our experiments show a significant improvement in attack success rate (+45.3% ASR) over current leading methods, even against state-of-the-art backdoor defenses, highlighting our attack's effectiveness in various scenarios, including downstream tasks. Our codes can be found at https://github.com/LiangSiyuan21/BadCLIP. Siyuan Liang 0004, Mingli Zhu, Aishan Liu, Baoyuan Wu, Xiaochun Cao, Ee-Chien Chang |
CVPR | 1 |
| 2024 | Hide in Thicket: Generating Imperceptible and Rational Adversarial Perturbations on 3D Point CloudsabstractAdversarial attack methods based on point manipulation for 3D point cloud classification have revealed the fragility of 3D models, yet the adversarial examples they produce are easily perceived or defended against. The tradeoff between the imperceptibility and adversarial strength leads most point attack methods to inevitably introduce easily detectable outlier points upon a successful attack. An-other promising strategy, shape-based attack, can effectively eliminate outliers, but existing methods often suffer significant reductions in imperceptibility due to irrational deformations. We find that concealing deformation perturbations in areas insensitive to human eyes can achieve a better tradeoff between imperceptibility and adversarial strength, specifically in parts of the object surface that are complex and exhibit drastic curvature changes. Therefore, we propose a novel shape-based adversarial attack method, HiT-ADV, which initially conducts a two-stage search for attack regions based on saliency and imperceptibility scores, and then adds deformation perturbations in each attack region using Gaussian kernel functions. Additionally, HiT-ADV is extendable to physical attack. We propose that by employing benign resampling and benign rigid transformations, we can further enhance physical adversarial strength with little sacrifice to imperceptibility. Extensive experiments have validated the superiority of our method in terms of adversarial and imperceptible properties in both digital and physical spaces. Our code is avaliable at: https://github.com/TRLou/HiT-ADV. Tianrui Lou, Xiaojun Jia, Jindong Gu, Li Liu 0002, Siyuan Liang 0004, Bangyan He, Xiaochun Cao |
CVPR | 5 |
| 2024 | Less is More: Fewer Interpretable Region via Submodular Subset SelectionabstractImage attribution algorithms aim to identify important regions that are highly relevant to model decisions. Although existing attribution solutions can effectively assign importance to target elements, they still face the following challenges: 1) existing attribution methods generate inaccurate small regions thus misleading the direction of correct attribution, and 2) the model cannot produce good attribution results for samples with wrong predictions. To address the above challenges, this paper re-models the above image attribution problem as a submodular subset selection problem, aiming to enhance model interpretability using fewer regions. To address the lack of attention to local regions, we construct a novel submodular function to discover more accurate small interpretation regions. To enhance the attribution effect for all samples, we also impose four different constraints on the selection of sub-regions, i.e., confidence, effectiveness, consistency, and collaboration scores, to assess the importance of various subsets. Moreover, our theoretical analysis substantiates that the proposed function is in fact submodular. Extensive experiments show that the proposed method outperforms SOTA methods on two face datasets (Celeb-A and VGG-Face2) and one fine-grained dataset (CUB-200-2011). For correctly predicted samples, the proposed method improves the Deletion and Insertion scores with an average of 4.9\% and 2.5\% gain relative to HSIC-Attribution. For incorrectly predicted samples, our method achieves gains of 81.0\% and 18.4\% compared to the HSIC-Attribution algorithm in the average highest confidence and Insertion score respectively. The code is released at https://github.com/RuoyuChen10/SMDL-Attribution. Ruoyu Chen 0001, Hua Zhang 0008, Siyuan Liang 0004, Jingzhi Li 0002, Xiaochun Cao |
ICLR | 3 |
| 2024 | Poisoned Forgery Face: Towards Backdoor Attacks on Face Forgery DetectionabstractThe proliferation of face forgery techniques has raised significant concerns within society, thereby motivating the development of face forgery detection methods. These methods aim to distinguish forged faces from genuine ones and have proven effective in practical applications. However, this paper introduces a novel and previously unrecognized threat in face forgery detection scenarios caused by backdoor attack. By embedding backdoors into models and incorporating specific trigger patterns into the input, attackers can deceive detectors into producing erroneous predictions for forged faces. To achieve this goal, this paper proposes \emph{Poisoned Forgery Face} framework, which enables clean-label backdoor attacks on face forgery detectors. Our approach involves constructing a scalable trigger generator and utilizing a novel convolving process to generate translation-sensitive trigger patterns. Moreover, we employ a relative embedding method based on landmark-based regions to enhance the stealthiness of the poisoned samples. Consequently, detectors trained on our poisoned samples are embedded with backdoors. Notably, our approach surpasses SoTA backdoor baselines with a significant improvement in attack success rate (+16.39\% BD-AUC) and reduction in visibility (-12.65\% $L_\infty$). Furthermore, our attack exhibits promising performance against backdoor defenses. We anticipate that this paper will draw greater attention to the potential threats posed by backdoor attacks in face forgery detection scenarios. Our codes will be made available at \url{https://github.com/JWLiang007/PFF}. Siyuan Liang 0004, Aishan Liu, Xiaojun Jia, Junhao Kuang, Xiaochun Cao |
ICLR | 2 |
| 2024 | Towards Robust Object Detection: Identifying and Removing Backdoors via Module Inconsistency Analysis
Siyuan Liang 0004, Chengyang Li 0001 |
ICPR (24) | 2 |
| 2024 | LanEvil: Benchmarking the Robustness of Lane Detection to Environmental IllusionsabstractLane detection (LD) is an essential component of autonomous driving systems, providing fundamental functionalities like adaptive cruise control and automated lane centering. Existing LD benchmarks primarily focus on evaluating common cases, neglecting the robustness of LD models against environmental illusions such as shadows and tire marks on the road. This research gap poses significant safety challenges since these illusions exist naturally in real-world traffic situations. For the first time, this paper studies the potential threats caused by these environmental illusions to LD and establishes the first comprehensive benchmark LanEvil for evaluating the robustness of LD against this natural corruption. We systematically design 14 prevalent yet critical types of environmental illusions (e.g., shadow, reflection) that cover a wide spectrum of real-world influencing factors in LD tasks. Based on real-world environments, we create 94 realistic and customizable 3D cases using the widely used CARLA simulator, resulting in a dataset comprising 90,292 sampled images. Through extensive experiments, we benchmark the robustness of popular LD methods using LanEvil, revealing substantial performance degradation (-5.37% Accuracy and -10.70% F1-Score on average), with shadow effects posing the greatest risk (-7.39% Accuracy). Additionally, we assess the performance of commercial auto-driving systems OpenPilot and Apollo through collaborative simulations, demonstrating that proposed environmental illusions can lead to incorrect decisions and potential traffic accidents. To defend against environmental illusions, we propose the Attention Area Mixing (AAM) approach using hard examples, which witness significant robustness improvement (+3.76%) under illumination effects. We hope our paper can contribute to advancing more robust auto-driving systems in the future. Part of our dataset and demos can be found at the https://lanevil.github.io/. Tianyuan Zhang 0004, Hainan Li, Yisong Xiao, Siyuan Liang 0004, Aishan Liu, Xianglong Liu 0001, Dacheng Tao |
ACM Multimedia | 5 |
| 2024 | Multimodal Unlearnable Examples: Protecting Data against Multimodal Contrastive LearningabstractMultimodal contrastive learning (MCL) has shown remarkable advances in zero-shot classification by learning from millions of image-caption pairs crawled from the Internet. However, this reliance poses privacy risks, as hackers may unauthorizedly exploit image-text data for model training, potentially including personal and privacy-sensitive information. Recent works propose generating unlearnable examples by adding imperceptible perturbations to training images to build shortcuts for protection. However, they are designed for unimodal classification, which remains largely unexplored in MCL. We first explore this context by evaluating the performance of existing methods on image-caption pairs, and they do not generalize effectively to multimodal data and exhibit limited impact to build shortcuts due to the lack of labels and the dispersion of pairs in MCL. In this paper, we propose Multi-step Error Minimization (MEM), a novel optimization process for generating multimodal unlearnable examples. It extends the Error-Minimization (EM) framework to optimize both image noise and an additional text trigger, thereby enlarging the optimized space and effectively misleading the model to learn the shortcut between the noise features and the text trigger. Specifically, we adopt projected gradient descent to solve the noise minimization problem and use HotFlip to approximate the gradient and replace words to find the optimal text trigger. Extensive experiments demonstrate the effectiveness of MEM, with post-protection retrieval results nearly half of random guessing, and its high transferability across different models. Our code is available on the https://github.com/thinwayliu/Multimodal-Unlearnable-Examples Xiaojun Jia, Yuan Xun, Siyuan Liang 0004, Xiaochun Cao |
ACM Multimedia | 4 |
| 2024 | Towards Robust Physical-world Backdoor Attacks on Lane DetectionabstractDeep learning-based lane detection (LD) plays a critical role in autonomous driving systems, such as adaptive cruise control. However, it is vulnerable to backdoor attacks. Existing backdoor attack methods on LD exhibit limited effectiveness in dynamic real-world scenarios, primarily because they fail to consider dynamic scene factors, including changes in driving perspectives (e.g., viewpoint transformations) and environmental conditions (e.g., weather or lighting changes). To tackle this issue, this paper introduces BadLANE, a dynamic scene adaptation backdoor attack for LD designed to withstand changes in real-world dynamic scene factors. To address the challenges posed by changing driving perspectives, we propose an amorphous trigger pattern composed of shapeless pixels. This trigger design allows the backdoor to be activated by various forms or shapes of mud spots or pollution on the road or lens, enabling adaptation to changes in vehicle observation viewpoints during driving. To mitigate the effects of environmental changes, we design a meta-learning framework to train meta-generators tailored to different environmental conditions. These generators produce meta-triggers that incorporate diverse environmental information, such as weather or lighting conditions, as the initialization of the trigger patterns for backdoor implantation, thus enabling adaptation to dynamic environments. Extensive experiments on various commonly used LD models in both digital and physical domains validate the effectiveness of our attacks, outperforming other baselines significantly (+25.15% on average in Attack Success Rate). Our codes can be found in https://github.com/Veee9/BadLANE. Xinwei Zhang 0010, Aishan Liu, Tianyuan Zhang 0004, Siyuan Liang 0004, Xianglong Liu 0001 |
ACM Multimedia | 4 |
| 2024 | Breaking the False Sense of Security in Backdoor Defense through Re-Activation AttackabstractDeep neural networks face persistent challenges in defending against backdoor attacks, leading to an ongoing battle between attacks and defenses. While existing backdoor defense strategies have shown promising performance on reducing attack success rates, can we confidently claim that the backdoor threat has truly been eliminated from the model? To address it, we re-investigate the characteristics of the backdoored models after defense (denoted as defense models). Surprisingly, we find that the original backdoors still exist in defense models derived from existing post-training defense strategies, and the backdoor existence is measured by a novel metric called backdoor existence coefficient. It implies that the backdoors just lie dormant rather than being eliminated. To further verify this finding, we empirically show that these dormant backdoors can be easily re-activated during inference stage, by manipulating the original trigger with well-designed tiny perturbation using universal adversarial attack. More practically, we extend our backdoor re-activation to black-box scenario, where the defense model can only be queried by the adversary during inference stage, and develop two effective methods, i.e., query-based and transfer-based backdoor re-activation attacks. The effectiveness of the proposed methods are verified on both image classification and multimodal contrastive learning (i.e., CLIP) tasks. In conclusion, this work uncovers a critical vulnerability that has never been explored in existing defense strategies, emphasizing the urgency of designing more robust and advanced backdoor defense mechanisms in the future. Mingli Zhu, Siyuan Liang 0004, Baoyuan Wu |
NeurIPS | 2 |
| 2023 | Generating Transferable 3D Adversarial Point Cloud via Random Perturbation FactorizationabstractRecent studies have demonstrated that existing deep neural networks (DNNs) on 3D point clouds are vulnerable to adversarial examples, especially under the white-box settings where the adversaries have access to model parameters. However, adversarial 3D point clouds generated by existing white-box methods have limited transferability across different DNN architectures. They have only minor threats in real-world scenarios under the black-box settings where the adversaries can only query the deployed victim model. In this paper, we revisit the transferability of adversarial 3D point clouds. We observe that an adversarial perturbation can be randomly factorized into two sub-perturbations, which are also likely to be adversarial perturbations. It motivates us to consider the effects of the perturbation and its sub-perturbations simultaneously to increase the transferability for sub-perturbations also contain helpful information. In this paper, we propose a simple yet effective attack method to generate more transferable adversarial 3D point clouds. Specifically, rather than simply optimizing the loss of perturbation alone, we combine it with its random factorization. We conduct experiments on benchmark dataset, verifying our method's effectiveness in increasing transferability while preserving high efficiency. Bangyan He, Jian Liu 0012, Yiming Li 0004, Siyuan Liang 0004, Jingzhi Li 0002, Xiaojun Jia, Xiaochun Cao |
AAAI | 4 |
| 2023 | Improving Robust Fariness via Balance Adversarial TrainingabstractAdversarial training (AT) methods are effective against adversarial attacks, yet they introduce severe disparity of accuracy and robustness between different classes, known as the robust fairness problem. Previously proposed Fair Robust Learning (FRL) adaptively reweights different classes to improve fairness. However, the performance of the better-performed classes decreases, leading to a strong performance drop. In this paper, we observed two unfair phenomena during adversarial training: different difficulties in generating adversarial examples from each class (source-class fairness) and disparate target class tendencies when generating adversarial examples (target-class fairness). From the observations, we propose Balance Adversarial Training (BAT) to address the robust fairness problem. Regarding source-class fairness, we adjust the attack strength and difficulties of each class to generate samples near the decision boundary for easier and fairer model learning; considering target-class fairness, by introducing a uniform distribution constraint, we encourage the adversarial example generation process for each class with a fair tendency. Extensive experiments conducted on multiple datasets (CIFAR-10, CIFAR-100, and ImageNette) demonstrate that our BAT can significantly outperform other baselines in mitigating the robust fairness problem (+5-10\% on the worst class accuracy)(Our codes can be found at https://github.com/silvercherry/Improving-Robust-Fairness-via-Balance-Adversarial-Training). Chunyu Sun, Chenye Xu, Chengyuan Yao, Siyuan Liang 0004, Yichao Wu, Ding Liang, Xianglong Liu 0001, Aishan Liu |
AAAI | 4 |
| 2023 | Exploring the Relationship Between Architectural Design and Adversarially Robust GeneralizationabstractAdversarial training has been demonstrated to be one of the most effective remedies for defending adversarial examples, yet it often suffers from the huge robustness generalization gap on unseen testing adversaries, deemed as the adversarially robust generalization problem. Despite the preliminary understandings devoted to adversarially robust generalization, little is known from the architectural perspective. To bridge the gap, this paper for the first time systematically investigated the relationship between adversarially robust generalization and architectural design. In particular, we comprehensively evaluated 20 most representative adversarially trained architectures on ImageNette and CIFAR-10 datasets towards multiple$\ell_{p}$-norm adversarial attacks. Based on the extensive experiments, we found that, under aligned settings, Vision Transformers (e.g., PVT, CoAtNet) often yield better adversarially robust generalization while CNNs tend to overfit on specific attacks and fail to generalize on multiple adversaries. To better understand the nature behind it, we conduct theoretical analysis via the lens of Rademacher complexity. We revealed the fact that the higher weight sparsity contributes significantly towards the better adversarially robust generalization of Transformers, which can be often achieved by the specially-designed attention blocks. We hope our paper could help to better understand the mechanism for designing robust DNNs. Our model weights can be found at http://robust.art. Aishan Liu, Shiyu Tang, Siyuan Liang 0004, Ruihao Gong, Xianglong Liu 0001, Dacheng Tao |
CVPR | 3 |
| 2023 | Face Encryption via Frequency-Restricted Identity-Agnostic AttacksabstractBillions of people are sharing their daily live images on social media everyday. However, malicious collectors use deep face recognition systems to easily steal their biometric information (e.g., faces) from these images. Some studies are being conducted to generate encrypted face photos using adversarial attacks by introducing imperceptible perturbations to reduce face information leakage. However, existing studies need stronger black-box scenario feasibility and more natural visual appearances, which challenge the feasibility of privacy protection. To address these problems, we propose a frequency-restricted identity-agnostic (FRIA) framework to encrypt face images from unauthorized face recognition without access to personal information. As for the weak black-box scenario feasibility, we obverse that representations of the average feature in multiple face recognition models are similar, thus we propose to utilize the average feature via the crawled dataset from the Internet as the target to guide the generation, which is also agnostic to identities of unknown face recognition systems; in nature, the low-frequency perturbations are more visually perceptible by the human vision system. Inspired by this, we restrict the perturbation in the low-frequency facial regions by discrete cosine transform to achieve the visual naturalness guarantee. Extensive experiments on several face recognition models demonstrate that our FRIA outperforms other state-of-the-art methods in generating more natural encrypted faces while attaining high black-box attack success rates of 96%. In addition, we validate the efficacy of FRIA using real-world black-box commercial API, which reveals the potential of FRIA in practice. Our codes can be found in https://github.com/XinDong10/FRIA. Xin Dong 0015, Rui Wang 0032, Siyuan Liang 0004, Aishan Liu, Lihua Jing |
ACM Multimedia | 3 |
| 2023 | Isolation and Induction: Training Robust Deep Neural Networks against Model Stealing AttacksabstractDespite the broad application of Machine Learning models as a Service (MLaaS), they are vulnerable to model stealing attacks. These attacks can replicate the model functionality by using the black-box query process without any prior knowledge of the target victim model. Existing stealing defenses add deceptive perturbations to the victim's posterior probabilities to mislead the attackers. However, these defenses are now suffering problems of high inference computational overheads and unfavorable trade-offs between benign accuracy and stealing robustness, which challenges the feasibility of deployed models in practice. To address the problems, this paper proposes Isolation and Induction (InI), a novel and effective training framework for model stealing defenses. Instead of deploying auxiliary defense modules that introduce redundant inference time, InI directly trains a defensive model by isolating the adversary's training gradient from the expected gradient, which can effectively reduce the inference computational cost. In contrast to adding perturbations over model predictions that harm the benign accuracy, we train models to produce uninformative outputs against stealing queries, which can induce the adversary to extract little useful knowledge from victim models with minimal impact on the benign performance. Extensive experiments on several visual classification datasets (e.g., MNIST and CIFAR10) demonstrate the superior robustness (up to 48% reduction on stealing accuracy) and speed (up to 25.4× faster) of our InI over other state-of-the-art methods. Our codes can be found in https://github.com/DIG-Beihang/InI-Model-Stealing-Defense. Jun Guo 0009, Xingyu Zheng, Aishan Liu, Siyuan Liang 0004, Yisong Xiao, Yichao Wu, Xianglong Liu 0001 |
ACM Multimedia | 4 |
| 2023 | Exploring Inconsistent Knowledge Distillation for Object Detection with Data AugmentationabstractKnowledge Distillation (KD) for object detection aims to train a compact detector by transferring knowledge from a teacher model. Since the teacher model perceives data in a way different from humans, existing KD methods only distill knowledge that is consistent with labels annotated by human expert while neglecting knowledge that is not consistent with human perception, which results in insufficient distillation and sub-optimal performance. In this paper, we propose inconsistent knowledge distillation (IKD), which aims to distill knowledge inherent in the teacher model's counter-intuitive perceptions. We start by considering the teacher model's counter-intuitive perceptions of frequency and non-robust features. Unlike previous works that exploit fine-grained features or introduce additional regularizations, we extract inconsistent knowledge by providing diverse input using data augmentation. Specifically, we propose a sample-specific data augmentation to transfer the teacher model's ability in capturing distinct frequency components and suggest an adversarial feature augmentation to extract the teacher model's perceptions of non-robust features in the data. Extensive experiments demonstrate the effectiveness of our method which outperforms state-of-the-art KD baselines on one-stage, two-stage and anchor-free object detectors (at most +1.0 mAP). Our codes will be made available at https://github.com/JWLiang007/IKD.git. Siyuan Liang 0004, Aishan Liu, Ke Ma 0001, Jingzhi Li 0002, Xiaochun Cao |
ACM Multimedia | 2 |
| 2023 | X-Adv: Physical Adversarial Object Attacks against X-ray Prohibited Item Detection
Aishan Liu, Jun Guo 0009, Jiakai Wang, Siyuan Liang 0004, Renshuai Tao, Wenbo Zhou 0004, Cong Liu 0006, Xianglong Liu 0001, Dacheng Tao |
USENIX Security Symposium | 4 |
| 2023 | Privacy-Enhancing Face Obfuscation Guided by Semantic-Aware Attribution MapsabstractFace recognition technology is increasingly being integrated into our daily life, e.g. Face ID. With the advancement of machine learning algorithms, the personal information such as age, gender, and race can be easily deduced from the recorded face images in these applications. This poses a serious privacy threat to individuals who do not want to be profiled, as face images are collected for biometric purposes. Existing methods mostly focus on adding the invisible adversarial perturbations into the images to make automatic inference infeasible. However, the application scenarios of these methods are limited due to the perturbations depending on the specific model. In this paper, we introduce a novel face privacy-enhancing framework by obfuscating the stored faces, which could maintain the data utility (face identity) while protecting the privacy of users (facial attributes). Specifically, we first develop a feature attribution module to discover the identity-related facial parts. Within this module, we introduce a pixel importance estimation model based on Shapley value to obtain a pixel-level attribution map, and then each pixel on the attribution map is aggregated into semantic facial parts, which are used to quantify the importance of different facial parts. Next, we design a privacy-enhancing module to generate the high-quality obfuscated images, which can modify the privacy semantic content and preserve the identity-related information. Using the proposed method, users can choose the single or multiple attributes to be obfuscated without affecting identity matching. Extensive experiments conducted on CelebA-HQ and VGGFace2-HQ benchmarks demonstrate the effectiveness and generalization ability of our method. Jingzhi Li 0002, Hua Zhang 0008, Siyuan Liang 0004, Pengwen Dai, Xiaochun Cao |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2022 | A Large-Scale Multiple-objective Method for Black-box Attack Against Object Detection
Siyuan Liang 0004, Longkang Li, Yanbo Fan, Xiaojun Jia, Jingzhi Li 0002, Baoyuan Wu, Xiaochun Cao |
ECCV (4) | 1 |
| 2022 | Imitated Detectors: Stealing Knowledge of Black-box Object DetectorsabstractDeep neural networks have shown great potential in many practical applications, yet their knowledge is at the risk of being stolen via exposed services (\eg APIs). In contrast to the commonly-studied classification model extraction, there exist no studies on the more challenging object detection task due to the sufficiency and efficiency of problem domain data collection. In this paper, we for the first time reveal that black-box victim object detectors can be easily replicated without knowing the model structure and training data. In particular, we treat it as black-box knowledge distillation and propose a teacher-student framework named Imitated Detector to transfer the knowledge of the victim model to the imitated model. To accelerate the problem domain data construction, we extend the problem domain dataset by generating synthetic images, where we apply the text-image generation process and provide short text inputs consisting of object categories and natural scenes; to promote the feedback information, we aim to fully mine the latent knowledge of the victim model by introducing an iterative adversarial attack strategy, where we feed victim models with transferable adversarial examples making victim provide diversified predictions with more information. Extensive experiments on multiple datasets in different settings demonstrate that our approach achieves the highest model extraction accuracy and outperforms other model stealing methods by large margins in the problem domain dataset. Our codes can be found at \urlhttps://github.com/LiangSiyuan21/Imitated-Detectors. Siyuan Liang 0004, Aishan Liu, Longkang Li, Yang Bai 0011, Xiaochun Cao |
ACM Multimedia | 1 |
| 2021 | Parallel Rectangle Flip Attack: A Query-based Black-box Attack against Object DetectionabstractObject detection has been widely used in many safety- critical tasks, such as autonomous driving. However, its vulnerability to adversarial examples has not been sufficiently studied, especially under the practical scenario of black-box attacks, where the attacker can only access the query feedback of predicted bounding-boxes and top- 1 scores returned by the attacked model. Compared with black-box attack to image classification, there are two main challenges in black-box attack to detection. Firstly, even if one bounding-box is successfully attacked, another sub- optimal bounding-box may be detected near the attacked bounding-box. Secondly, there are multiple bounding- boxes, leading to very high attack cost. To address these challenges, we propose a Parallel Rectangle Flip Attack (PRFA) via random search. We explain the difference between our method with other attacks in Fig. 1. Specifically, we generate perturbations in each rectangle patch to avoid sub-optimal detection near the attacked region. Besides, utilizing the observation that adversarial perturbations mainly locate around objects’ contours and critical points under white-box attacks, the search space of attacked rectangles is reduced to improve the attack efficiency. Moreover, we develop a parallel mechanism of attacking multiple rectangles simultaneously to further accelerate the attack process. Extensive experiments demonstrate that our method can effectively and efficiently attack various popular object detectors, including anchor-based and anchor- free, and generate transferable adversarial examples. Siyuan Liang 0004, Baoyuan Wu, Yanbo Fan, Xingxing Wei 0001, Xiaochun Cao |
ICCV | 1 |
| 2020 | Efficient Adversarial Attacks for Visual Object Tracking
Siyuan Liang 0004, Xingxing Wei 0001, Siyuan Yao, Xiaochun Cao |
ECCV (26) | 1 |
| 2019 | Transferable Adversarial Attacks for Image and Video Object DetectionabstractIdentifying adversarial examples is beneficial for understanding deep networks and developing robust models. However, existing attacking methods for image object detection have two limitations: weak transferability---the generated adversarial examples often have a low success rate to attack other kinds of detection methods, and high computation cost---they need much time to deal with video data, where many frames need polluting. To address these issues, we present a generative method to obtain adversarial images and videos, thereby significantly reducing the processing time. To enhance transferability, we manipulate the feature maps extracted by a feature network, which usually constitutes the basis of object detectors. Our method is based on the Generative Adversarial Network (GAN) framework, where we combine a high-level class loss and a low-level feature loss to jointly train the adversarial example generator. Experimental results on PASCAL VOC and ImageNet VID datasets show that our method efficiently generates image and video adversarial examples, and more importantly, these adversarial examples have better transferability, therefore being able to simultaneously attack two kinds of representative object detection models: proposal based models like Faster-RCNN and regression based models like SSD. Xingxing Wei 0001, Siyuan Liang 0004, Ning Chen 0002, Xiaochun Cao |
IJCAI | 2 |