EDBT 2026 Demo / reviewers in the wild / expert
Wenbo Jiang 0001
dblp:34/10703-1
· DBLP profile ↗
50ranked-venue papers
12as first author
45since 2021 · last 2026
0000-0002-4592-8094ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 18 · 4 first-author · 15 since 2021Artificial intelligence and machine learning · 13 · 1 first-author · 13 since 2021Security and privacy · 11 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MPMA: Preference Manipulation Attack Against Model Context ProtocolabstractModel Context Protocol (MCP) standardizes interface mapping for large language models (LLMs) to access external data and tools, which revolutionizes the paradigm of tool selection and facilitates the rapid expansion of the LLM agent tool ecosystem. However, as the MCP is increasingly adopted, third-party customized versions of the MCP server expose potential security vulnerabilities. In this paper, we first introduce a novel security threat, which we term the MCP Preference Manipulation Attack (MPMA). An attacker deploys a customized MCP server to manipulate LLMs, causing them to prioritize it over other competing MCP servers. This can result in economic benefits for attackers, such as revenue from paid MCP services or advertising income generated from free servers. To achieve MPMA, we first design a Direct Preference Manipulation Attack (DPMA) that achieves significant effectiveness by inserting the manipulative word and phrases into the tool name and description. However, such a direct modification is obvious to users and lacks stealthiness. To address these limitations, we further propose Genetic-based Advertising Preference Manipulation Attack (GAPMA). GAPMA employs four commonly used strategies to initialize descriptions and integrates a Genetic Algorithm (GA) to enhance stealthiness. The experiment results demonstrate that GAPMA balances high effectiveness and stealthiness. Our study reveals a critical vulnerability of the MCP in open ecosystems, highlighting an urgent need for robust defense mechanisms to ensure the fairness of the MCP ecosystem. Rui Zhang 0090, Wenshu Fan, Wenbo Jiang 0001, Qingchuan Zhao, Hongwei Li 0001, Guowen Xu |
AAAI | 5 |
| 2026 | ConfGuard: A Simple and Effective Backdoor Detection for Large Language ModelsabstractBackdoor attacks pose a significant threat to Large Language Models (LLMs), where adversaries can embed hidden triggers to manipulate LLM's outputs. Most existing defense methods, primarily designed for classification tasks, are ineffective against the autoregressive nature and vast output space of LLMs, thereby suffering from poor performance and high latency. To address these limitations, we investigate the behavioral discrepancies between benign and backdoored LLMs in output space. We identify a critical phenomenon which we term sequence lock: a backdoored model generates the target sequence with abnormally high and consistent confidence compared to benign generation. Building on this insight, we propose ConfGuard, a lightweight and effective detection method that monitors a sliding window of token confidences to identify sequence lock. Extensive experiments demonstrate ConfGuard achieves a near 100% true positive rate (TPR) and a negligible false positive rate (FPR) in the vast majority of cases. Crucially, the ConfGuard enables real-time detection almost without additional latency, making it a practical backdoor defense for real-world LLM deployments. Rui Zhang 0086, Hongwei Li 0001, Wenshu Fan, Wenbo Jiang 0001, Qingchuan Zhao, Guowen Xu |
AAAI | 5 |
| 2026 | Efficient Privacy-Preserving Genetic Analysis via Distributed Function Secret Sharing
Shenghao Wu, Pengzhi Xing, Meng Hao 0001, Hanxiao Chen 0001, Wenbo Jiang 0001, Hongwei Li 0001 |
ICC | 5 |
| 2026 | Guided by Principles of Composition: A Domain-Specific Priors Based Detector for Recognizing Ritual Implements in ThangkaabstractABSTRACT Detecting ritual implements in Thangka paintings—such as swords and scriptures—remains challenging due to their intricate visual composition and symbolic complexity. Existing object detection models, typically trained on natural scenes, tend to perform poorly in this domain. To address this limitation, we summarize the principles of composition in Thangka and identify key spatial and co‐occurrence priors specific to ritual implements. Based on these insights, we propose GPCDet: a guided by principles of composition detector that integrates domain‐specific priors into the detection process. Specifically, we introduce a spatial coordinate attention module to emphasize critical spatial regions where implements frequently appear. In addition, we design a graph convolution network‐auxiliary detection module to model inter‐category co‐occurrence, thereby enhancing feature representation and improving classification performance. Experiments on the newly curated ritual implements in Thangka (RITK) dataset show that GPCDet achieves substantial improvements over existing methods, establishing a new state‐of‐the‐art baseline for this challenging task. Jiachen Li 0002, Hongyun Wang, Xiaolong Peng, Jinyu Xu 0001, Qing Xie 0002, Yanchun Ma, Wenbo Jiang 0001, Mengzi Tang |
IET Image Process. | 7 |
| 2026 | TrojanEdit: Multimodal backdoor attack against image editing model
Ji Guo, Runjia Zhang, Wenbo Jiang 0001, Yiting Zhu, Jiachen Li 0002, Jiaming He, Hongwei Li 0001 |
Neurocomputing | 3 |
| 2026 | Backdoor Complications: A Comprehensive Analysis and Mitigation of the Unforeseen Consequences of Backdoor AttacksabstractPre-trained language models (PTLMs) have become integral to modern natural language processing (NLP), yet their reuse exposes them to supply chain risks such as backdoor attacks. Existing studies assume that attackers target specific downstream tasks, overlooking how a backdoored PTLM behaves when fine-tuned for unrelated applications. In practice, such unintended adaptation can trigger anomalous and inconsistent predictions, revealing the backdoor and compromising its stealthiness. We define this phenomenon asbackdoor complications, i.e., unintended behavioral side effects emerging on non-target tasks. This work presents the first systematic quantification and mitigation of backdoor complications. Through extensive experiments on 3 widely used PTLMs and 15 benchmark datasets, we show that complications are pervasive across both single- and multi-task attack settings, causing triggered outputs to collapse into arbitrary classes. To address this issue, we propose theComplication-Suppressed Backdoor Attack(CSBA), a task-agnostic, multi-objective framework that leverages auxiliary non-target datasets to suppress backdoor complications. CSBA effectively suppresses complications on unseen downstream tasks while maintaining near-perfect attack success rates. Our work reveals a critical side effect in backdoored PTLMs and provides a new perspective on the stealthiness and robustness of model supply chain security. Rui Zhang 0086, Hongwei Li 0001, Wenbo Jiang 0001, Hanxiao Chen 0001, Yuan Zhang 0006, Guowen Xu, Yang Zhang 0016 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2026 | Hidden Tail: Adversarial Attack for Stealthy Resource Consumption Against Vision-Language ModelsabstractVision-Language Models (VLMs) are increasingly deployed in real-world applications, but their high inference cost makes them vulnerable to resource consumption attacks. Prior attacks attempt to extend VLM output sequences by optimizing adversarial images, thereby increasing inference costs. However, these extended outputs often introduce irrelevant abnormal content, compromising attack stealthiness. This trade-off between effectiveness and stealthiness poses a major limitation for existing attacks. To address this challenge, we proposeHidden Tail, a stealthy resource consumption attack that crafts prompt-agnostic adversarial images, inducing VLMs to generate maximum-length outputs by appending special tokens invisible to users. Our method employs a composite loss function that balances semantic preservation, repetitive special token induction, and suppression of the end-of-sequence (EOS) token, optimized via a dynamic weighting strategy. Extensive experiments show thatHidden Tailoutperforms existing attacks, increasing output length by up to 19.2× and reaching the maximum token limit, while preserving attack stealthiness. These results highlight the urgent need to improve the robustness of VLMs against efficiency-oriented adversarial threats. Our code is available athttps://github.com/zhangrui4041/Hidden_Tail. Rui Zhang 0086, Tianli Yang, Wenbo Jiang 0001, Rui Zhang 0090, Qingchuan Zhao, Hongwei Li 0001, Yang Liu 0003, Guowen Xu |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2026 | Conan: Secure and Reliable Machine Learning Inference Against Malicious Service ProvidersabstractIn the Machine Learning as a Service paradigm, a service provider (e.g., a server) hosting a model offers inference APIs to clients, who can send their queries and receive the inference results. While most recent secure inference works focus on addressing privacy issues, they overlook the importance of checking the service quality and reliability. A malicious server may deviate from the protocol specification to deliberately provide incorrect services such as using low-quality models. Thus, it is necessary to design new solutions to empower clients to verify the server’s model accuracy and inference integrity while protecting both parties’ privacy. We present Conan, a new secure and reliable inference framework against malicious servers to achieve accuracy verification, inference integrity, and privacy simultaneously. In Conan, the server first commits to the model and proves in zero-knowledge that the committed model achieves the claimed accuracy. Then both parties perform secure inference on the committed model against the malicious server. To instantiate the above framework, we design generic maliciously secure two-party computation (2PC) protocols with a fixed corrupted party, which may be of independent interest. Our protocols achieve high efficiency by utilizing the advantage that the semi-honest party can check the behavior of the corrupted party. Furthermore, they support both arithmetic and Boolean circuit evaluation, a crucial attribute for secure inference on complicated machine learning models. We implement the fixed-corruption 2PC protocols for our secure and reliable inference. The experimental results show 1 ~ 2 orders of magnitude improvements over conventional maliciously secure protocols in terms of communication and computation costs. Hanxiao Chen 0001, Hongwei Li 0001, Meng Hao 0001, Pengzhi Xing, Jia Hu 0004, Wenbo Jiang 0001, Tianwei Zhang 0004, Guowen Xu |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2025 | Watch Out for Your Guidance on Generation! Exploring Conditional Backdoor Attacks against Large Language ModelsabstractMainstream backdoor attacks on large language models (LLMs) typically set a fixed trigger in the input instance and specific responses for triggered queries. However, the fixed trigger setting (e.g., unusual words) may be easily detected by human detection, limiting the effectiveness and practicality in real-world scenarios. To enhance the stealthiness of backdoor activation, we present a new poisoning paradigm against LLMs triggered by specifying generation conditions, which are commonly adopted strategies by users during model inference. The poisoned model performs normally for output under normal/other generation conditions, while becomes harmful for output under target generation conditions. To achieve this objective, we introduce BrieFool, an efficient attack framework. It leverages the characteristics of generation conditions by efficient instruction sampling and poisoning data generation, thereby influencing the behavior of LLMs under target conditions. Our attack can be generally divided into two types with different targets: Safety unalignment attack and Ability degradation attack. Our extensive experiments demonstrate that BrieFool is effective across safety domains and ability domains, achieving higher success rates than baseline methods, with 94.3% on GPT-3.5-turbo. Jiaming He, Wenbo Jiang 0001, Guanyu Hou, Wenshu Fan, Rui Zhang 0086, Hongwei Li 0001 |
AAAI | 2 |
| 2025 | DivTrackee versus DynTracker: Promoting Diversity in Anti-Facial Recognition against Dynamic FR StrategyabstractThe widespread adoption of facial recognition (FR) models raises serious concerns about their potential misuse, motivating the development of anti-facial recognition (AFR) to protect user facial privacy. In this paper, we argue that the static FR strategy, predominantly adopted in prior literature for evaluating AFR efficacy, cannot faithfully characterize the actual capabilities of determined trackers who aim to track a specific target identity. In particular, we introduce DynTracker, a dynamic FR strategy where the model's gallery database is iteratively updated with newly recognized target identity images. Surprisingly, such a simple approach renders all the existing AFR protections ineffective. To mitigate the privacy threats posed by DynTracker, we advocate for explicitly promoting diversity in the AFR-protected images. We hypothesize that the lack of diversity is the primary cause of the failure of existing AFR methods. Specifically, we develop DivTrackee, a novel method for crafting diverse AFR protections that builds upon a text-guided image generation framework and diversity-promoting adversarial losses. Through comprehensive experiments on various image benchmarks and feature extractors, we demonstrate DynTracker's strength in breaking existing AFR methods and the superiority of DivTrackee in preventing user facial images from being identified by dynamic FR strategies. We believe our work can act as an important initial step towards developing more effective AFR methods for protecting user facial privacy against determined trackers. Wenshu Fan, Minxing Zhang, Hongwei Li 0001, Wenbo Jiang 0001, Hanxiao Chen 0001, Xiangyu Yue 0001, Michael Backes 0001, Xiao Zhang 0016 |
CCS | 4 |
| 2025 | Evaluating Robustness of Large Audio Language Models to Audio Injection: An Empirical StudyabstractLarge Audio-Language Models (LALMs) are increasingly deployed in real-world applications, yet their robustness against malicious audio injection remains underexplored.To address this gap, this study systematically evaluates five leading LALMs across four attack scenarios: Audio Interference Attack, Instruction Following Attack, Context Injection Attack, and Judgment Hijacking Attack.We quantitatively assess their vulnerabilities and resilience using metrics: the Defense Success Rate, Context Robustness Score, and Judgment Robustness Index.The experiments reveal significant performance disparities, with no single model demonstrating consistent robustness across all attack types.Attack effectiveness is significantly influenced by the position of the malicious content, particularly when injected at the beginning of a sequence.Furthermore, our analysis uncovers a negative correlation between a model's instruction-following capability and its robustness: models that strictly adhere to instructions tend to be more susceptible, whereas safety-aligned models exhibit greater resistance.To facilitate future research, this work introduces a comprehensive benchmark framework.Our findings underscore the critical need for integrating robustness into training pipelines and developing multi-modal defenses, ultimately facilitating the secure deployment of LALMs.The dataset used in this work is available on Hugging Face. Guanyu Hou, Jiaming He, Yinhang Zhou, Ji Guo, Yitong Qiao, Rui Zhang 0086, Wenbo Jiang 0001 |
EMNLP | 7 |
| 2025 | CLBA: A Cross-Lingual Backdoor Attack against Text-to-Image Diffusion ModelsabstractDiffusion-based Text-to-Image (T2I) synthesis has emerged as a transformative multimodal generation technology. However, its reliance on pre-trained models introduces severe backdoor risks, including output manipulation and privacy violations. Existing attacks using specific triggers demonstrate effectiveness in monolingual settings but tend to degrade in crosslingual scenarios due to semantic inconsistency. To bridge this gap, we propose a cross-lingual transfer backdoor attack that maintains robust attack efficacy across languages through single-language poisoning. Specifically, we introduce a linguistically informed trigger selection method that identifies semantically invariant words across languages. The trigger is deliberately and naturally embedded into prompts to minimize semantic disruption, thereby helping to evade potential defense mechanisms. Our approach leverages cross-lingual semantic alignment between high-resource and low-resource languages to enable stealthy and effective backdoor activation. We assume the adversary finetunes with limited data, without knowledge of the model internals or the original training data. Our attack demonstrates the potential to bypass existing defenses in T2I models and exposes critical security vulnerabilities in multilingual T2I systems. These findings highlight the urgent need for targeted security measures to mitigate backdoor threats and prevent malicious exploitation. Hongwei Li 0001, Rui Zhang 0086, Jiaming He, Wenbo Jiang 0001 |
GLOBECOM | 5 |
| 2025 | PromptNeedling: Jailbreaking Text-to-Video Generative ModelsabstractRecent advances in text-to-video (T2V) generation have enabled high-fidelity video synthesis from natural language descriptions. While these models offer promising applications, their capacity to generate not safe for work (NSFW) content poses significant security concerns. In this work, we first explore the vulnerability of T2V models to jailbreak attacks, building upon methods developed for text-to-image (T2I) models. We conduct a systematic evaluation of the transferability of T2I jailbreak techniques to T2V models, revealing that naively adapted attacks yield limited effectiveness due to the unique temporal and semantic challenges in video generation. To address these limitations, we propose PromptNeedling, a jailbreak attack method tailored for T2V models. Specifically, PromptNeedling optimizes the jailbreak prompts to simplify adversarial inputs and injects high-salience NSFW keywords in a controlled manner. We conduct extensive experiments on three open-source T2V models and evaluate two categories of NSFW content (nudity and gore & violence), showing that PromptNeedling achieves higher attack success rates than prior methods. These findings highlight the urgent need for developing effective defenses against jailbreak attacks to ensure the safety of T2V models. Yiyang Mu, Hongwei Li 0001, Rui Zhang 0086, Wenbo Jiang 0001, Wenshu Fan |
GLOBECOM | 4 |
| 2025 | BadComp: Backdoor Attack against Object Detection using Image Compression OperationabstractCurrently, object detection models have achieved widespread success in real-world applications, yet remain vulnerable to backdoor attacks. Existing backdoor methods often suffer from poor stealthiness or can be easily mitigated by standard image processing techniques. In this paper, we propose a novel stealthy backdoor attack to dynamically compress target object-oriented data for backdoor embedding. By modifying the model, a high-frequency feature extraction module is added, so the model learns frequency-domain feature representations of compressed samples and achieve three attack objectives. Extensive experimental results demonstrate the effectiveness and robustness of the proposed method, achieving attack success rates exceeding 90% across three detectors. Xi Nie, Hongwei Li 0001, Wenbo Jiang 0001, Shuai Yuan 0009, Wenshu Fan, Jian Xiong 0007 |
GLOBECOM | 3 |
| 2025 | Exploiting Unknown Samples under Limited Budgets in Open-set Active LearningabstractActive Learning (AL) aims to improve model performance while reducing annotation costs by selectively labeling the most informative samples from an unlabeled dataset. In openset scenarios where the unlabeled pool may include instances from unknown classes, most existing open-set AL methods focus exclusively on known-class samples and neglect the valuable information that unknown-class samples can provide. To address this limitation, we propose ALSO (Active Learning with few-Shot in Open set), a novel framework that incorporates both known and unknown class samples during the selection process. ALSO combines uncertainty-based sampling for known classes with a similarity-driven strategy to identify and utilize representative unknown-class samples. This few-shot approach enables the model to generalize effectively with limited supervision from unknown classes, enhancing overall performance even with a small number of annotated samples. Extensive experiments on CIFAR-10, CIFAR-100, and Tiny-ImageNet demonstrate that ALSO consistently outperforms state-of-the-art methods in accuracy and precision. Zhijing Wang, Ji Guo, Wenbo Jiang 0001 |
GLOBECOM | 4 |
| 2025 | SecInfer: Secure and Efficient Model Inference on Vertically Partitioned DataabstractDeep learning models have achieved unprecedented success in various domains, such as healthcare and finance. However, deploying model inference in real-world applications, where data is distributed among multiple entities, poses significant privacy concerns. Existing secure model inference work has limitations in computational overhead and scalability, especially when dealing with complex models and multiple parties with vertically partitioned data. In this work, we design and implement an efficient and scalable secure inference framework for vertically partitioned data, supporting execution with a large number of parties. Our work considers a semi-honest setting with all-but-one corruptions. The core of our framework is a series of secure and efficient protocols for complex non-linear functions of the model inference, such as ReLU and Maxpool. These protocols are designed based on secure multi-party computation preliminaries, significantly enhancing efficiency while maintaining rigorous security guarantees. We conduct comprehensive experiments to evaluate the performance of our framework. Experimental results show that SecInfer substantially improves the communication and computation performance of secure naive inference works by up to 3.71 × and 3.42 ×, respectively. Robert H. Deng, Hongwei Li 0001, Hanxiao Chen 0001, Meng Hao 0001, Pengzhi Xing, Jia Hu 0004, Rui Zhang 0086, Wenbo Jiang 0001 |
ICC | 8 |
| 2025 | Making Audio Data UnlearnableabstractIn recent years, deep neural networks (DNNs) have driven rapid advancements in various fields. As DNNs continue to grow in size, the amount of training data required is also increasing. Many researchers crawl publicly available data from the internet for training, which raises issues of unauthorized exploitation and potential privacy leakage. Recent work against unauthorized exploitation primarily focuses on the image domain, while the audio domain remains underexplored. In this paper, we propose an effective method to generate audio unlearnable examples, which injects imperceptible perturbations into training samples, making them unlearnable for models to train. Specifically, we employ an error minimization optimization algorithm to iteratively optimize the generated perturbations. To ensure these perturbations remain imperceptible, we leverage both the Short-Time Objective Intelligibility (STOI) score and the$L_{2}$norm as measures to constrain the perceptibility of the unlearnable examples. To further improve the transferability of the unlearnable effects, we use an ensemble model and threshold constraint to effectively enhance generated unlearnable examples. Experimental results demonstrate that our method effectively generates samples that prevent models from accurately learning and making predictions, while remaining indistinguishable from clean samples to human observers. Wenshu Fan, Hongwei Li 0001, Wenbo Jiang 0001 |
ICC | 3 |
| 2025 | Stealthy Physical Backdoor Attacks Against Traffic Sign Recognition SystemsabstractRecent advancements in deep learning have led to remarkable progress in autonomous driving technology, with deep neural network (DNN)-based traffic sign recognition systems (TSRS) playing a crucial role. However, recent studies indicate that TSRS are vulnerable to backdoor attacks, where the backdoor TSRS behaves normally on clean traffic signs but consistently misclassifies backdoor-triggered traffic signs into a designated target class. Notably, while backdoor attacks in the digital domain are effective, their effectiveness may diminish in the physical world due to quality degradation during image transmission. Existing physical backdoor attacks typically rely on specific stickers or transformations as backdoor triggers, which are not stealthy and natural enough in the physical world. To address these limitations, we propose two stealthy physical backdoor attacks against DNN-based TSRS from two different perspectives. On the one hand, we utilize the natural phenomenon of chipped paints on traffic signs as the backdoor trigger. Specifically, we develop an automatic traffic sign segmentation algorithm to identify the edges of the target sign and simulate chipped paint to create poisoned samples. On the other hand, instead of manipulating the target traffic sign, we use the specific filter lens (attached to the in-vehicle camera) as the backdoor trigger, where the parameters of the filter lens are optimized by the Genetic Algorithm (GA). Extensive experiments conducted on the GTSRB and TSRD datasets demonstrate the effectiveness of our proposed backdoor attacks in both digital and physical environments. Wenbo Jiang 0001, Hongwei Li 0001, Shuai Yuan 0009, Rui Zhang 0086, Qiyang Song |
ICC | 1 |
| 2025 | Adversarial Attack with Controllable TransferabilityabstractMachine Learning as a Service (MLaaS) providers often promote the robustness of their models as a selling point for their API services. Existing methods commonly evaluate the robustness of Deep Neural Networks (DNNs) by generating highly transferable adversarial examples However, dishonest MLaaS providers may deceive consumers by falsely exaggerating the robustness of their models. In this paper, contrary to enhancing the transferability of adversarial examples, we attempt to craft adversarial examples that fail under a shielded model. In other words, ensuring that the adversarial examples fail to attack a shielded model but successfully attack other models, which we refer to as controllable transferability. To achieve this goal, we propose the Controllable Transferability Method (CTM), a framework that generates adversarial examples with controllable transferability. CTM involves generating transferable adversarial examples and refining their transferability using gradient antagonism. Experimental results demonstrate that CTM achieves high transferability across models, with controlled adversarial effects on selected models. Jian Xiong 0007, Hongwei Li 0001, Wenbo Jiang 0001, Wenshu Fan, Shuai Yuan 0009 |
ICC | 3 |
| 2025 | Weaponizing Tokens: Backdooring Text-to-Image Generation via Token RemappingabstractText-to-image generative models have garnered immense attention for their ability to produce high-fidelity images from text prompts and enjoyed great popularity among the community. Unfortunately, previous studies have demonstrated that text-to-image models suffer from backdoor attacks, which enforce the text-guided generative models to generate images that align the backdoor target via embedding the textual triggers. However, the currently proposed backdoor attacks rely on numerous training data and complex computing resources for poisoning the core components in generative models, limiting the effectiveness and practicality in real-world scenarios. In this work, we first investigate the backdoor attack against Text-to-image generation by manipulating text tokenizer. Our backdoor attack exploits the semantic conditioning role of text tokenizer in the text-to-image generation. We propose an Automatized Remapping Framework with Optimized Tokens (AROT) for finding the best target tokens to remap the trigger token in the mapping space, according to different tasks. We conduct extensive experiments on Stable Diffusion and two defined tasks to demonstrate the effectiveness, stealthiness and robustness of our attack. Jiaming He, Wenbo Jiang 0001, Guanyu Hou, Qiyang Song, Ji Guo, Hongwei Li 0001 |
ICME | 2 |
| 2025 | Omni-Angle Assault: An Invisible and Powerful Physical Adversarial Attack on Face RecognitionabstractDeep learning models employed in face recognition (FR) systems have been shown to be vulnerable to physical adversarial attacks through various modalities, including patches, projections, and infrared radiation. However, existing adversarial examples targeting FR systems often suffer from issues such as conspicuousness, limited effectiveness, and insufficient robustness. To address these challenges, we propose a novel approach for adversarial face generation, UVHat, which utilizes ultraviolet (UV) emitters mounted on a hat to enable invisible and potent attacks in black-box settings. Specifically, UVHat simulates UV light sources via video interpolation and models the positions of these light sources on a curved surface, specifically the human head in our study. To optimize attack performance, UVHat integrates a reinforcement learning-based optimization strategy, which explores a vast parameter search space, encompassing factors such as shooting distance, power, and wavelength. Extensive experimental evaluations validate that UVHat substantially improves the attack success rate in black-box settings, enabling adversarial attacks from multiple angles with enhanced robustness. Shuai Yuan 0009, Hongwei Li 0001, Rui Zhang 0090, Hangcheng Cao, Wenbo Jiang 0001, Tao Ni 0003, Wenshu Fan, Qingchuan Zhao, Guowen Xu |
ICML | 5 |
| 2025 | The Ripple Effect: On Unforeseen Complications of Backdoor AttacksabstractRecent research highlights concerns about the trustworthiness of third-party Pre-Trained Language Models (PTLMs) due to potential backdoor attacks.
These backdoored PTLMs, however, are effective only for specific pre-defined downstream tasks.
In reality, these PTLMs can be adapted to many other unrelated downstream tasks.
Such adaptation may lead to unforeseen consequences in downstream model outputs, consequently raising user suspicion and compromising attack stealthiness.
We refer to this phenomenon as backdoor complications.
In this paper, we undertake the first comprehensive quantification of backdoor complications.
Through extensive experiments using 4 prominent PTLMs and 16 text classification benchmark datasets, we demonstrate the widespread presence of backdoor complications in downstream models fine-tuned from backdoored PTLMs.
The output distribution of triggered samples significantly deviates from that of clean samples.
Consequently, we propose a backdoor complication reduction method leveraging multi-task learning to mitigate complications without prior knowledge of downstream tasks.
The experimental results demonstrate that our proposed method can effectively reduce complications while maintaining the efficacy and consistency of backdoor attacks. Rui Zhang 0086, Hongwei Li 0001, Wenbo Jiang 0001, Hanxiao Chen 0001, Yuan Zhang 0006, Guowen Xu, Yang Zhang 0016 |
ICML | 4 |
| 2025 | CtrlMark: Controllable Watermarking for ControlNet Against Downstream Fine-TuningabstractText-to-image diffusion models have advanced controllable image generation, with ControlNet plugins enabling precise structural guidance and domain-specific adaptations. As these plugins become widely shared and personalized, protecting their ownership and preventing misuse becomes crucial. Existing watermarking methods address robustness against fine-tuning and personalization at the model level, but fail to address ControlNet-like plugins or modules specifically. To address this gap, we propose CtrlMark, the first watermarking framework designed specifically for ControlNet plugin modules. CtrlMark embeds a robust, triggerable watermark as a benign backdoor, activated by a composite trigger combining text and structural inputs. Furthermore, CtrlMark achieves few misactivations and strong robustness against downstream fine-tuning, by watermark penalization and leveraging a fixed latent residual embedding localized to a spatial region. Extensive experiments demonstrate CtrlMark maintains high watermark activation rates and visual fidelity, providing effective and practical protection for modular ControlNet components in diverse generation scenarios. Rui Zhang 0090, Wenbo Jiang 0001, Hongwei Li 0001, Guowen Xu |
ICPADS | 3 |
| 2025 | Stealthy Backdoor Attack against Object DetectionabstractRecent research has revealed that object detectors are highly susceptible to backdoor attacks, which can introduce detection errors during inference, such as detecting non-existent objects or failing to detect existing objects. Though several backdoor attacks targeting object detection have been proposed to achieve high attack success rates, these methods often involve visible triggers, which can be detected by human inspection or backdoor defenses. To enhance the attack stealthiness, we introduce a stealthy backdoor attack for object detection. Specifically, it employs a uniform shift on each pixel within images as the trigger. The particle swarm optimization is utilized to effectively find the optimal uniform shift to accomplish different attack targets in object detection, including object disappearance, object generation, and object misclassification. To achieve these targets and preserve stealthy, we design corresponding objective functions to maintain a balance between attack stealthiness and attack effectiveness. We have conducted comprehensive experiments to demonstrate the effectiveness of our proposed attack across the three attack targets in object detection, as well as its robustness against existing defense methods. Xiaoyang Ning, Qing Xie 0002, Jinyu Xu 0001, Wenbo Jiang 0001, Xiaoyuan Liu 0002, Jiachen Li 0002, Yanchun Ma |
IJCNN | 4 |
| 2025 | DiffWR: Diffusion for Watermark RemovalabstractDigital watermarking has become a critical technology for copyright protection of digital images. However, the removal of such watermarks to restore original image content presents significant challenges. Current deep learning methods, which rely on GAN or multi-stage encoder-decoder networks, struggle to restore the clean and apparent texture of background images accurately, leading to visual artifacts and degradation in color and texture. Additionally, these methods require precise prediction of watermark positions, limiting their flexibility in handling inaccurate mask predictions.In this paper, we introduce a novel framework, Diffusion Watermark Removal (DiffWR), which employs a diffusion-based approach to address the limitations of existing techniques. DiffWR utilizes Denoising Diffusion Probabilistic Models (DDPM) to generate high-quality, artifact-free restored images. It incorporates an Initial Prediction Module (IPM) to provide preliminary image predictions and guide the diffusion process within accurate boundaries using a binary mask. This approach allows for adaptive inpainting beyond the initially predicted regions, enhancing the flexibility of the restoration process.Our model demonstrates superior performance in terms of visual quality and structural similarity without introducing artifacts. It outperforms state-of-the-art methods in perceptual metrics such as FID and LPIPS, as well as in traditional metrics like SSIM and PSNR. The results showcase the capability of DiffWR to effectively remove watermarks while preserving the original image’s visual fidelity and texture. Lingli Tang, Yanchun Ma, Qing Xie 0002, Anshu Hu, Wenbo Jiang 0001 |
IJCNN | 6 |
| 2025 | A Hidden Backdoor Attack via Formal Text Style Transfer in Language ModelsabstractNatural language processing (NLP) systems have been demonstrated to be vulnerable to backdoor attacks. Specifically, attackers embed the backdoor into the model by poisoning training data, producing the desired results when the input contains pre-defined triggers. Typical textual backdoor attacks adopt static triggers such as words or phrases, which make them detectable by existing defense methods. To enhance stealthiness, this paper introduces a hidden backdoor attack method utilizing formal text style transfer (FTST). Specifically, we adopt a formal text style transfer model to convert part of the benign training samples into formal samples, which serve as the backdoor samples. Compared to static textual triggers, FTST-based triggers can maintain original semantics while evading common defenses and human detections. We conduct extensive experiments on typical NLP tasks, including topic and sentiment classification tasks utilizing three prominent pre-trained language models and four datasets. The results show that our approach achieves the desired attack performance while preserving the normal-functionality of the model. Furthermore, compared to common word-level triggers and sentence-level triggers, our approach has been demonstrated to be more stealthy under GPT-2-based perplexity detection and more robust under backdoor defense methods. Hongwei Li 0001, Wenbo Jiang 0001, Rui Zhang 0086, Jiaming He, Hanxiao Chen 0001, Guowen Xu |
IJCNN | 3 |
| 2025 | When Hallucinated Concepts Cross Modals: Unveiling Backdoor Vulnerability in Multi-modal In-context LearningabstractDue to the remarkable performance of multi-modal large language models (MLLMs) in multi-modal capabilities, multi-modal in-context learning (M-ICL) has garnered widespread attention for fast adapting MLLMs to downstream tasks. However, the vulnerability of M-ICL to attacks remains largely unexplored. In this work, we take the first step to explore the backdoor vulnerability of M-ICL, which allows the adversary only to manipulate the multi-modal demonstration examples to mislead the victim model. We propose a multi-modal backdoor strategy on M-ICL via cross-modal concept mis-matching under black-box attack setting. Extensive experimental results demonstrate that our attacks exhibit high attack effectiveness while preserving the normal functionality of the victim model. Moreover, we further conduct experiments to prove our attacks are robust against backdoor defenses and still remain effective in various real-world conditions. Guanyu Hou, Jiaming He, Yitong Qiao, Jiachen Li 0002, Qiyang Song, Ji Guo, Wenbo Jiang 0001 |
MMAsia | 8 |
| 2025 | You Are Out of My Focus: A Defocus-Blur Backdoor Attack against Deep Learning ModelsabstractWith the widespread adoption of deep learning in image recognition, backdoor attacks have emerged as a significant security threat, drawing increasing attention from the research community. Traditional backdoor attacks are often limited to the digital domain, while few existing physical-world attacks suffer from a lack of stealthiness. In this paper, inspired by the natural defocus blur commonly caused by camera optics in real-world environments, we propose a physically-aware backdoor attack method called DBBA based on the defocus blur phenomenon. By leveraging Gaussian blur to simulate this natural phenomenon, the proposed method enhances both the stealthiness and plausibility of the trigger. To further optimize the attack effectiveness while maintaining stealthiness, we introduce a Particle Swarm Optimization (PSO) algorithm to automatically search for the optimal Gaussian blur parameters that best simulate the defocus phenomenon. We conduct extensive experiments on multiple mainstream image classification datasets and across various model architectures. Experimental results demonstrate that the proposed defocus-blur based trigger achieves a high attack effectiveness with minimal degradation in the classification accuracy of the model. In addition, evaluations against representative defense techniques reveal that the proposed method exhibits strong stealthiness and robustness. Hongwei Li 0001, Wenbo Jiang 0001, Jiaming He, Rui Zhang 0090, Ji Guo, Jiachen Li 0002 |
MMAsia | 3 |
| 2025 | The Fluorescent Veil: A Stealthy and Effective Physical Adversarial Patch Against Traffic Sign RecognitionabstractRecently, traffic sign recognition (TSR) systems have become a prominent target for physical adversarial attacks. These attacks typically rely on conspicuous stickers and projections, or using invisible light and acoustic signals that can be easily blocked. In this paper, we introduce a novel attack medium, i.e., fluorescent ink, to design a stealthy and effective physical adversarial patch, namely FIPatch, to advance the state-of-the-art. Specifically, we first model the fluorescence effect in the digital domain to identify the optimal attack settings, which guide the real-world fluorescence parameters. By applying a carefully designed fluorescence perturbation to the target sign, the attacker can later trigger a fluorescent effect using invisible ultraviolet light, causing the TSR system to misclassify the sign and potentially leading to traffic accidents. We conducted a comprehensive evaluation to investigate the effectiveness of FIPatch, which shows a success rate of 98.31% in low-light conditions. Furthermore, our attack successfully bypasses five popular defenses and achieves a success rate of 96.72%. Shuai Yuan 0009, Xingshuo Han, Hongwei Li 0001, Guowen Xu, Wenbo Jiang 0001, Tao Ni 0003, Qingchuan Zhao, Yuguang Fang |
NeurIPS | 5 |
| 2025 | Backdoor attacks against Hybrid Classical-Quantum Neural Networks
Ji Guo, Wenbo Jiang 0001, Rui Zhang 0090, Wenshu Fan, Jiachen Li 0002, Guoming Lu, Hongwei Li 0001 |
Neural Networks | 2 |
| 2025 | I2I Backdoor: Backdoor Attacks Against Image-to-Image TasksabstractWith the rapid development of deep learning technology, deep learning-based Image-to-Image (I2I) networks have become the predominant choice for I2I tasks like image super-resolution and denoising. Despite their remarkable performance, the security of I2I networks has not been thoroughly investigated. While some studies have probed their susceptibility to adversarial attacks, none have explored the backdoor attack against I2I networks, which is a more stealthy and severe threat. In this work, for the first time, we comprehensively investigate the vulnerability of I2I networks to backdoor attacks. We propose a backdoor attack against I2I tasks, where the backdoored I2I network behaves normally on clean input images, yet outputs a specific inappropriate image when the backdoor trigger appears on the input image. To achieve such an I2I backdoor attack, we design a universal adversarial perturbation (UAP) generation algorithm for I2I networks, where the generated UAP is used as the trigger for the I2I backdoor. Besides, multi-task learning (MTL) with dynamic weighting methods is employed in the backdoor training process to gain better results. Expanding our focus beyond I2I tasks, we extend our I2I backdoor to attack downstream tasks, including image classification and object detection. Specifically, the backdoor-triggered image processed by the backdoored image denoising network can fool the downstream image classifiers and object detectors. Extensive experiments demonstrate the effectiveness of the I2I backdoor on state-of-the-art I2I network architectures as well as the robustness against different backdoor defenses. Wenbo Jiang 0001, Hongwei Li 0001, Jiaming He, Rui Zhang 0090, Guowen Xu, Tianwei Zhang 0004, Rongxing Lu |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2025 | Rethinking the Design of Backdoor Triggers and Adversarial Perturbations: A Color Space PerspectiveabstractDeep neural networks (DNNs) are known to be susceptible to various malicious attacks, such as adversarial and backdoor attacks. However, most of these attacks utilize additive adversarial perturbations (or backdoor triggers) within an$L_{p}$-norm constraint. They can be easily defeated by image preprocessing strategies, such as image compression and image super-resolution. To address this limitation, instead of using additive adversarial perturbations (or backdoor triggers) in the pixel space, this work revisits the design of adversarial perturbations (or backdoor triggers) from the perspective of color space and conducts a comprehensive analysis. Specifically, we propose a color space backdoor attack and a color space adversarial attack where the color space shift is used as the trigger and perturbation. To find the optimal trigger or perturbation in the black-box scenario, we perform an iterative optimization process with the Particle Swarm Optimization algorithm. Experimental results confirm the robustness of the proposed color space attacks against image preprocessing defenses as well as other mainstream defense methods. In addition, we also design adaptive defense strategies and evaluate their effectiveness against color space attacks. Our work emphasizes the importance of the color space when developing malicious attacks against DNN and urges more research in this area. Wenbo Jiang 0001, Hongwei Li 0001, Guowen Xu, Hao Ren 0001, Haomiao Yang, Tianwei Zhang 0004, Shui Yu 0001 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2025 | The Gradient Puppeteer: Adversarial Domination in Gradient Leakage Attacks Through Model Poisoning
Kunlan Xiang, Haomiao Yang, Meng Hao 0001, Shaofeng Li 0001, Haoxin Wang 0004, Zikang Ding, Wenbo Jiang 0001, Tianwei Zhang 0004 |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2024 | Adversarial Robustness Poisoning: Increasing Adversarial Vulnerability of the Model via Data PoisoningabstractDeep neural networks (DNNs) have become prevalent across various domains. However, recent research has revealed their vulnerability to data poisoning attacks, where adversaries inject poisoned data to compromise the usability of the target model. Traditional data poisoning attacks focus on reducing the test accuracy of the model, but they can be detected by model performance evaluation or mitigated by data cleaning. In contrast, we propose an Adversarial Robustness Poisoning Scheme (ARPS) that aims to decrease the adversarial robustness while preserving the normal-functionality of the target model. To achieve ARPS, we first separate the features of data into robust and non-robust features, where the non-robust features are human-imperceptible and more sensitive to adversarial perturbations. After that, we construct a dataset containing only non-robust features, which serves as the poisoning data. For the malicious dataset provider, the poisoned dataset can be constructed by adding poisoning data to the original dataset. For the malicious model provider, we employ the uncertainty-weighted multi-task learning technique to train the poisoned model, facilitating a better balance between functionality-preserving (good accuracy) and attack effectiveness (bad robustness). Extensive experiments are carried out to illustrate the effectiveness of ARPS in weakening adversarial training and amplifying adversarial attacks, as well as the stealthiness of ARPS in escaping the defense of data cleaning and model fine-tuning. Additionally, we propose some potential countermeasures against ARPS, including regularization and data augmentations. Wenbo Jiang 0001, Hongwei Li 0001, Wenshu Fan, Rui Zhang 0086 |
GLOBECOM | 1 |
| 2024 | Backdoor Attack Against Vision Transformers via Attention Gradient-Based Image ErosionabstractVision Transformers (ViTs) have outperformed traditional Convolutional Neural Networks (CNN) across various computer vision tasks. However, akin to CNN, ViTs are vulnerable to backdoor attacks, where the adversary embeds the backdoor into the victim model, causing it to make wrong predictions about testing samples containing a specific trigger. Existing backdoor attacks against ViTs have the limitation of failing to strike an optimal balance between attack stealthiness and attack effectiveness.In this work, we propose an Attention Gradient-based Erosion Backdoor (AGEB) targeted at ViTs. Considering the attention mechanism of ViTs, AGEB selectively erodes pixels in areas of maximal attention gradient, embedding a covert backdoor trigger. Unlike previous backdoor attacks against ViTs, AGEB achieves an optimal balance between attack stealthiness and attack effectiveness, ensuring the trigger remains invisible to human detection while preserving the model’s accuracy on clean samples. Extensive experimental evaluations across various ViT architectures and datasets confirm the effectiveness of AGEB, achieving a remarkable Attack Success Rate (ASR) without diminishing Clean Data Accuracy (CDA). Furthermore, the stealthiness of AGEB is rigorously validated, demonstrating minimal visual discrepancies between the clean and the triggered images. Ji Guo, Hongwei Li 0001, Wenbo Jiang 0001, Guoming Lu |
GLOBECOM | 3 |
| 2024 | BadTTS: Identifying Vulnerabilities in Neural Text-to-Speech ModelsabstractWith the widespread use of deep learning systems in many applications, adversaries have strong incentives to perform attacks against these systems for their adversarial purposes. Reports have indicated that backdoor attacks on deep neural networks represent a novel form of threat. In this attack, the adversary will inject backdoors into the benign model and then mislead the model to classify the input containing backdoor triggers as a target label specified by the adversary. Existing research mainly focuses on backdoor attacks in image and text models, little attention has been paid to the backdoor attacks on text-to-speech (TTS) models. We conduct a systematic investigation of backdoor attacks on text-to-speech models and propose BadTTS, the first backdoor attack against TTS models, which is a general backdoor attack framework that tampers with input texts in three semantic levels to generate malicious output speech, including Char-Backdoor, Word-Backdoor, and Sentence-Backdoor. Our method not only efficiently injects backdoors into a TTS model but is also stealthy and has little impact on the synthesized speech quality. We implement the backdoor attack in a black-box fine-tuning setting, where the adversary has no knowledge of model architectures except for a small amount of training data. We perform empirical experiments on three representative and widely studied TTS models, indicating that backdoors can be injected into TTS models within a few fine-tuning steps. Additionally, we conduct experiments to explore the impact of different types of triggers, as well as the intermediate outputs of models, which provide insights for potential defenses against backdoor attacks. Rui Zhang 0090, Hongwei Li 0001, Wenbo Jiang 0001, Jiaming He |
GLOBECOM | 3 |
| 2024 | QPFFL: Advancing Federated Learning with Quantum-Resistance, Privacy, and FairnessabstractFederated Learning (FL) has gained prominence for collaborative training across multiple devices without data sharing. However, traditional FL overlooks two crucial aspects: collaborative fairness and privacy protection. Typically, all participants receive the same models, regardless of their contribution, and plaintext transmission of model gradients risks privacy. Existing fairness-enhancing approaches often increase privacy risks, while security-focused methods suffer from efficiency limitations, failing to provide a comprehensive solution against multiple threats simultaneously. To address these challenges, we propose QPFFL, a novel fair and secure FL framework. Firstly, we propose Privacy-Preserving Reputation Mechanism (PPRM) that assigns global models to users based on their performance during training, promoting fairness of FL. We employ Functional Encryption (FE) to enable efficient and quantum-resistant aggregation, securing user model parameters. Furthermore, a reputation threshold helps identify malicious behaviors. Theoretical analysis and experiments demonstrate QPFFL’s effectiveness in thwarting various attacks without compromising privacy and efficiency, thereby providing a comprehensive solution for secure and fair FL. Hongwei Li 0001, Xinyuan Qian 0002, Xiaoyuan Liu 0002, Wenbo Jiang 0001 |
GLOBECOM | 5 |
| 2024 | Mtisa: Multi-Target Image-Scaling AttackabstractImage scaling is one of the most common operations in image processing. For instance, it is often conducted before image transferring to preserve resources, image classifiers also require images to be input at a specified size. However, potential threats may come out with the image scaling operation. A recent work called image-scaling attack can change the semantic information of the input image when it is scaled to a specific size. For example, a manipulated image of a sheep may become an image of a wolf when it scales to a specific size. Many works have already demonstrated the effectiveness of this attack and the security risks it poses. However, existing image-scaling attacks only focus on single target with single specific size, and are not applicable to multi-target image-scaling attack. In this paper, we present a multi-target image-scaling attack (MTISA). MTISA can be trained with a single image performs diverse and semantically distinct outputs to fool both human vision and image classifiers. Specifically, to fool human vision, we employ SinGAN to generate semantically different but background-similar samples to serve as the attack target samples. To mislead image classifiers, we employ adversarial attacks to construct adversarial examples to serve as the attack target samples. Finally, we evaluate MTISA on chest X-rays dataset and ImageNet dataset, respectively. The experimental results demonstrate that MTISA achieves high attack success rate against both human vision and image classifiers. Jiaming He, Hongwei Li 0001, Wenbo Jiang 0001, Yuan Zhang 0006 |
ICC | 3 |
| 2024 | An Efficient and Secure Privacy-Preserving Federated Learning Via Lattice-Based Functional EncryptionabstractIn recent times, federated learning (FL) aggregation techniques based on functional encryption (FE) have garnered increased attention. The growing interest stems from the distinct advantages of FE compared to traditional aggregation methods. Especially in terms of computational efficiency, communication costs and functionality, FE markedly surpasses its counterparts. However, privacy-preserving federated learning (PPFL) schemes utilizing FE still grapple with significant privacy and security challenges. For instance, current implementations fail to safe-guard aggregated intermediate outcomes and remain susceptible to quantum attacks, among other concerns. To address these problems, we first propose PIM-MCFE, a new FE scheme based on Learning with Errors (LWE) assumption, which can hide the intermediate aggregated results and is computationally efficient. We extend the scheme to the aggregation task of PPFL and propose an optimization technique, plaintext packaging to accelerate the training process. We provide the security analysis of the proposed PPFL scheme through theoretical analysis and demonstrate its efficiency and practicality through extensive experiments. The results show the encryption efficiency of our scheme improves by 20× and 50× compared to HybridAlpha and CryptoFE, and the decryption operation achieves a 3-orders-of-maanitude efficiency improvement. Hongwei Li 0001, Xinyuan Qian 0002, Wenbo Jiang 0001 |
ICC | 4 |
| 2024 | Instruction Backdoor Attacks Against Customized LLMs
Rui Zhang 0086, Hongwei Li 0001, Rui Wen 0002, Wenbo Jiang 0001, Yuan Zhang 0006, Michael Backes 0001, Yang Zhang 0016 |
USENIX Security Symposium | 4 |
| 2024 | A Comprehensive Defense Framework Against Model Extraction AttacksabstractAs a promising service, Machine Learning as a Service (MLaaS) provides personalized inference functions for clients through paid APIs. Nevertheless, it is vulnerable to model extraction attacks, in which an attacker can extract a functionally-equivalent model by repeatedly querying the APIs with crafted samples. While numerous works have been proposed to defend against model extraction attacks, existing efforts are accompanied by limitations and low comprehensiveness. In this article, we propose AMAO, a comprehensive defense framework against model extraction attacks. Specifically, AMAO consists of four interlinked successive phases: adversarial training is first exploited to weaken the effectiveness of model extraction attacks. Then, malicious query detection is used to detect malicious queries and mark malicious users. After that, we develop a label-flipping poisoning attack to instruct the adaptive query responses to malicious users. Besides, the image pHash algorithm is employed to ensure the indistinguishability of the query responses. Finally, the perturbed results are served as a backdoor to verify the ownership of any suspicious model. Extensive experiments demonstrate that AMAO outperforms existing defenses in defending against model extraction attacks and is also robust against the adaptive adversary who is aware of the defense. Wenbo Jiang 0001, Hongwei Li 0001, Guowen Xu, Tianwei Zhang 0004, Rongxing Lu |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2024 | Incremental Learning, Incremental Backdoor ThreatsabstractClass incremental learning from a pre-trained DNN model is gaining lots of popularity. Unfortunately, the pre-trained model also introduces a new attack vector, which enables an adversary to inject a backdoor into it and further compromise the downstream models learned from it. Prior works proposed backdoor attacks against the pre-trained models in the transfer learning scenario. However, they become less effective when the adversary does not have the knowledge of the downstream tasks or new data, which is more practical and considered in this paper. To this end, we design the first latent backdoor attacks against incremental learning. We propose two novel techniques, which can effectively and stealthily embed a backdoor into the pre-trained model. Such backdoor can only be activated when the pre-trained model is extended to a downstream model with incremental learning. It has a very high attack success rate, and is able to bypass existing backdoor detection approaches. Extensive experiments confirm the effectiveness of our attacks over different datasets and incremental learning methods, as well as strong robustness against state-of-the-art backdoor defense mechanisms includingNeural Cleanse,Fine-PruningandSTRIP. Wenbo Jiang 0001, Tianwei Zhang 0004, Han Qiu 0001, Hongwei Li 0001, Guowen Xu |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2024 | Stealthy Targeted Backdoor Attacks Against Image CaptioningabstractIn recent years, there is an explosive growth in multimodal learning. Image captioning, a classical multimodal task, has demonstrated promising applications and attracted extensive research attention. However, recent studies have shown that image caption models are vulnerable to some security threats such as backdoor attacks. Existing backdoor attacks against image captioning typically pair a trigger either with a predefined sentence or a single word as the targeted output, yet they are unrelated to the image content, making them easily noticeable as anomalies by humans. In this paper, we present a novel method to craft targeted backdoor attacks against image caption models, which are designed to be stealthier than prior attacks. Specifically, our method first learns a special trigger by leveraging universal perturbation techniques for object detection, then places the learned trigger in the center of some specific source object and modifies the corresponding object name in the output caption to a predefined target name. During the prediction phase, the caption produced by the backdoored model for input images with the trigger can accurately convey the semantic information of the rest of the whole image, while incorrectly recognizing the source object as the predefined target. Extensive experiments demonstrate that our approach can achieve a high attack success rate while having a negligible impact on model clean performance. In addition, we show our method is stealthy in that the produced backdoor samples are indistinguishable from clean samples in both image and text domains, which can successfully bypass existing backdoor defenses, highlighting the need for better defensive mechanisms against such stealthy backdoor attacks. Wenshu Fan, Hongwei Li 0001, Wenbo Jiang 0001, Meng Hao 0001, Shui Yu 0001, Xiao Zhang 0016 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2023 | Color Backdoor: A Robust Poisoning Attack in Color SpaceabstractBackdoor attacks against neural networks have been intensively investigated, where the adversary compromises the integrity of the victim model, causing it to make wrong predictions for inference samples containing a specific trigger. To make the trigger more imperceptible and human-unnoticeable, a variety of stealthy backdoor attacks have been proposed, some works employ imperceptible perturbations as the backdoor triggers, which restrict the pixel differences of the triggered image and clean image. Some works use special image styles (e.g., reflection, Instagram filter) as the backdoor triggers. However, these attacks sacrifice the robustness, and can be easily defeated by common preprocessing-based defenses. This paper presents a novel color backdoor attack, which can exhibit robustness and stealthiness at the same time. The key insight of our attack is to apply a uniform color space shift for all pixels as the trigger. This global feature is robust to image transformation operations and the triggered samples maintain natural-looking. To find the optimal trigger, we first define naturalness restrictions through the metrics of PSNR, SSIM and LPIPS. Then we employ the Particle Swarm Optimization (PSO) algorithm to searchfor the optimal trigger that can achieve high attack effectiveness and robustness while satisfying the restrictions. Extensive experiments demonstrate the superiority of PSO and the robustness of color backdoor against different main-stream backdoor defenses. Wenbo Jiang 0001, Hongwei Li 0001, Guowen Xu, Tianwei Zhang 0004 |
CVPR | 1 |
| 2023 | Physical Black-Box Adversarial Attacks Through TransformationsabstractDeep learning has shown impressive performance in numerous applications. However, recent studies have found that deep learning models are vulnerable to adversarial attacks, where the attacker adds imperceptible perturbations into benign samples to induce misclassifications. Adversarial attacks in the digital domain focus on constructing imperceptible perturbations. However, they are always less effective in the physical world because the perturbations may be destroyed when captured by the camera. Most physical adversarial attacks require adding invisible adversarial features (e.g., a sticker or a laser) to the target object, which may be noticed by human eyes. In this work, we propose to employ image transformation to generate more natural adversarial samples in the physical world. Concretely, we propose two attack algorithms to satisfy different attack goals:Efficient-AATRemploys a greedy strategy to generate adversarial samples with fewer queries;Effective-AATRemploys an adaptive particle swarm optimization algorithm to search for the most effective adversarial samples within the given the number of queries. Extensive experiments demonstrate the superiority of our attacks compared with state-of-the-art adversarial attacks under mainstream defenses. Wenbo Jiang 0001, Hongwei Li 0001, Guowen Xu, Tianwei Zhang 0004, Rongxing Lu |
IEEE Trans. Big Data | 1 |
| 2020 | A Practical Black-Box Attack Against Autonomous Speech Recognition ModelabstractWith the wild applications of machine learning (ML) technology, automatic speech recognition (ASR) has made great progress in recent years. Despite its great potential, there are various evasion attacks of ML-based ASR, which could affect the security of applications built upon ASR. Up to now, most studies focus on white-box attacks in ASR, and there is almost no attention paid to black-box attacks where attackers can only query the target model to get output labels rather than probability vectors in audio domain. In this paper, we propose an evasion attack against ASR in the above-mentioned situation, which is more feasible in realistic scenarios. Specifically, we first train a substitute model by using data augmentation, which ensures that we have enough samples to train with a small number of times to query the target model. Then, based on the substitute model, we apply Differential Evolution (DE) algorithm to craft adversarial examples and implement black-box attack against ASR models from the Speech Commands dataset. Extensive experiments are conducted, and the results illustrate that our approach achieves untargeted attacks with over 70% success rate while still maintaining the authenticity of the original data well. Wenshu Fan, Hongwei Li 0001, Wenbo Jiang 0001, Guowen Xu, Rongxing Lu |
GLOBECOM | 3 |
| 2020 | Accelerating Poisoning Attack Through Momentum and Adam AlgorithmsabstractMachine learning has demonstrated promising application prospects in the field of vehicular technology during the past decade, for instance, it effectively propelled the development of autonomous vehicles and intelligent transportation systems. However, machine learning is still vulnerable to numerous malicious attacks. Amongst them, poisoning attack is one of the most severe security threats to the training process of machine learning, where the attacker injects some poisoned samples to the training dataset to make the learned model unavailable. As the crucial part of poisoning attack is generating poisoned samples, most proposals for poisoning attack have employed traditional gradient-based optimization algorithms to optimize the poisoned samples. Nevertheless, conventional gradient-based optimization algorithms are liable to get trapped in local optimums or saddle points and have a slow rate of convergence. As a result, these problems may lead to a reduction of the poisoned samples' effect. To address these issues, we propose two improved gradient-based poisoning attack algorithms. Specifically, in order to accelerate the convergence speed, we propose the first poisoning attack algorithm by employing momentum algorithm. Also, we propose the second poisoning attack algorithm by utilizing adam algorithm, which can get rid of some local optimums and has a faster convergence speed simultaneously. After that, support vector machines (SVM), linear regression and logistics regression are chosen as exemplary algorithms to conduct our attack algorithms and the effectiveness and computational overhead of the two attack algorithms are evaluated. Finally, we propose a countermeasure algorithm, which can detect suspicious samples using mahalanobis distance. Wenbo Jiang 0001, Hongwei Li 0001, Haomiao Yang, Rongxing Lu |
VTC Fall | 1 |
| 2019 | A Flexible Poisoning Attack Against Machine LearningabstractRecent years have witnessed tremendous academic efforts and industry growth in machine learning. The security of machine learning has become increasingly prominent. Poisoning attack is one of the most relevant security threats to machine learning which focuses on polluting the training data that machine learning needs during the training process. Specifically, the attacker blends crafted poisoning samples into training data in order to make the learned model beneficial to him. To the best of our knowledge, existing researches about poisoning attack focused on either integrity attack or availability attack, which did not unify these two attacks together. Aside from that, from the attacker's perspective, attacker's strategy is not flexible enough. Finally, existing proposals only concentrated on increasing the test error of the learned model but ignored the importance of the concealment of attack. To overcome these issues, we firstly present a thorough adversarial model for poisoning attack in which attacker's strategy is defined from two aspects, i.e., the effect of attack and the concealment of attack. Then we unify integrity attack and availability attack together in similar formulations. Furthermore, in order to enhance flexibility, a tradeoff parameter is inserted into attacker's objective function which means the attacker can balance the attraction of effect against the requirement of concealment. Finally, as examples, extensive experiments are conducted on linear regression and logistic regression to demonstrate the effectiveness of attack. Wenbo Jiang 0001, Hongwei Li 0001, Sen Liu 0007, Yanzhi Ren |
ICC | 1 |
| 2019 | PTAS: Privacy-preserving Thin-client Authentication Scheme in blockchain-based PKI
Wenbo Jiang 0001, Hongwei Li 0001, Guowen Xu, Mi Wen, Guishan Dong, Xiaodong Lin 0001 |
Future Gener. Comput. Syst. | 1 |
| 2018 | A Privacy-Preserving Thin-Client Scheme in Blockchain-Based PKIabstractTraditional centralized PKIs are vulnerable due to the single point of failure. A feasible solution is to build a decentralized PKI without certificate authority (CA). Web of Trust is the first step toward realizing a decentralized PKI, but it still has some limitations such as missing incentive and leaking user's privacy. Blockchain's numerous desirable properties, such as cryptographical security, decentralized nature and unalterable transaction record, make it a suitable tool to implement a decentralized PKI. However, the latest research findings about blockchain-based PKI are still incompatible with the thin-clients which have limited storage ability to download the entire blockchain. To combat that, we firstly present a Privacy-preserving Thin-client Scheme (PTS) utilizing the idea of k-anonymity, which enables thin-clients to run normally as full node users and protect user's privacy simultaneously. After that, in order to reduce cost, we further propose an Efficient Privacy preserving Thin-client Scheme (EPTS) employing the method of PIR (private information retrieval). Then security analysis and functional comparison are performed to demonstrate the high security and comprehensive functionality of EPTS compared with existing schemes. Finally, extensive experiments are undertaken to confirm that EPTS can reduce computational cost and communication cost impressively. Wenbo Jiang 0001, Hongwei Li 0001, Guowen Xu, Mi Wen, Guishan Dong, Xiaodong Lin 0001 |
GLOBECOM | 1 |