VLDB 2026 Research / reviewers in the wild / expert
Wenshu Fan
dblp:285/2460
· DBLP profile ↗
15ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0003-1335-7183ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 8 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Security and privacy · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MPMA: Preference Manipulation Attack Against Model Context ProtocolabstractModel Context Protocol (MCP) standardizes interface mapping for large language models (LLMs) to access external data and tools, which revolutionizes the paradigm of tool selection and facilitates the rapid expansion of the LLM agent tool ecosystem. However, as the MCP is increasingly adopted, third-party customized versions of the MCP server expose potential security vulnerabilities. In this paper, we first introduce a novel security threat, which we term the MCP Preference Manipulation Attack (MPMA). An attacker deploys a customized MCP server to manipulate LLMs, causing them to prioritize it over other competing MCP servers. This can result in economic benefits for attackers, such as revenue from paid MCP services or advertising income generated from free servers. To achieve MPMA, we first design a Direct Preference Manipulation Attack (DPMA) that achieves significant effectiveness by inserting the manipulative word and phrases into the tool name and description. However, such a direct modification is obvious to users and lacks stealthiness. To address these limitations, we further propose Genetic-based Advertising Preference Manipulation Attack (GAPMA). GAPMA employs four commonly used strategies to initialize descriptions and integrates a Genetic Algorithm (GA) to enhance stealthiness. The experiment results demonstrate that GAPMA balances high effectiveness and stealthiness. Our study reveals a critical vulnerability of the MCP in open ecosystems, highlighting an urgent need for robust defense mechanisms to ensure the fairness of the MCP ecosystem. Rui Zhang 0090, Wenshu Fan, Wenbo Jiang 0001, Qingchuan Zhao, Hongwei Li 0001, Guowen Xu |
AAAI | 4 |
| 2026 | ConfGuard: A Simple and Effective Backdoor Detection for Large Language ModelsabstractBackdoor attacks pose a significant threat to Large Language Models (LLMs), where adversaries can embed hidden triggers to manipulate LLM's outputs. Most existing defense methods, primarily designed for classification tasks, are ineffective against the autoregressive nature and vast output space of LLMs, thereby suffering from poor performance and high latency. To address these limitations, we investigate the behavioral discrepancies between benign and backdoored LLMs in output space. We identify a critical phenomenon which we term sequence lock: a backdoored model generates the target sequence with abnormally high and consistent confidence compared to benign generation. Building on this insight, we propose ConfGuard, a lightweight and effective detection method that monitors a sliding window of token confidences to identify sequence lock. Extensive experiments demonstrate ConfGuard achieves a near 100% true positive rate (TPR) and a negligible false positive rate (FPR) in the vast majority of cases. Crucially, the ConfGuard enables real-time detection almost without additional latency, making it a practical backdoor defense for real-world LLM deployments. Rui Zhang 0086, Hongwei Li 0001, Wenshu Fan, Wenbo Jiang 0001, Qingchuan Zhao, Guowen Xu |
AAAI | 4 |
| 2025 | Watch Out for Your Guidance on Generation! Exploring Conditional Backdoor Attacks against Large Language ModelsabstractMainstream backdoor attacks on large language models (LLMs) typically set a fixed trigger in the input instance and specific responses for triggered queries. However, the fixed trigger setting (e.g., unusual words) may be easily detected by human detection, limiting the effectiveness and practicality in real-world scenarios. To enhance the stealthiness of backdoor activation, we present a new poisoning paradigm against LLMs triggered by specifying generation conditions, which are commonly adopted strategies by users during model inference. The poisoned model performs normally for output under normal/other generation conditions, while becomes harmful for output under target generation conditions. To achieve this objective, we introduce BrieFool, an efficient attack framework. It leverages the characteristics of generation conditions by efficient instruction sampling and poisoning data generation, thereby influencing the behavior of LLMs under target conditions. Our attack can be generally divided into two types with different targets: Safety unalignment attack and Ability degradation attack. Our extensive experiments demonstrate that BrieFool is effective across safety domains and ability domains, achieving higher success rates than baseline methods, with 94.3% on GPT-3.5-turbo. Jiaming He, Wenbo Jiang 0001, Guanyu Hou, Wenshu Fan, Rui Zhang 0086, Hongwei Li 0001 |
AAAI | 4 |
| 2025 | DivTrackee versus DynTracker: Promoting Diversity in Anti-Facial Recognition against Dynamic FR StrategyabstractThe widespread adoption of facial recognition (FR) models raises serious concerns about their potential misuse, motivating the development of anti-facial recognition (AFR) to protect user facial privacy. In this paper, we argue that the static FR strategy, predominantly adopted in prior literature for evaluating AFR efficacy, cannot faithfully characterize the actual capabilities of determined trackers who aim to track a specific target identity. In particular, we introduce DynTracker, a dynamic FR strategy where the model's gallery database is iteratively updated with newly recognized target identity images. Surprisingly, such a simple approach renders all the existing AFR protections ineffective. To mitigate the privacy threats posed by DynTracker, we advocate for explicitly promoting diversity in the AFR-protected images. We hypothesize that the lack of diversity is the primary cause of the failure of existing AFR methods. Specifically, we develop DivTrackee, a novel method for crafting diverse AFR protections that builds upon a text-guided image generation framework and diversity-promoting adversarial losses. Through comprehensive experiments on various image benchmarks and feature extractors, we demonstrate DynTracker's strength in breaking existing AFR methods and the superiority of DivTrackee in preventing user facial images from being identified by dynamic FR strategies. We believe our work can act as an important initial step towards developing more effective AFR methods for protecting user facial privacy against determined trackers. Wenshu Fan, Minxing Zhang, Hongwei Li 0001, Wenbo Jiang 0001, Hanxiao Chen 0001, Xiangyu Yue 0001, Michael Backes 0001, Xiao Zhang 0016 |
CCS | 1 |
| 2025 | PromptNeedling: Jailbreaking Text-to-Video Generative ModelsabstractRecent advances in text-to-video (T2V) generation have enabled high-fidelity video synthesis from natural language descriptions. While these models offer promising applications, their capacity to generate not safe for work (NSFW) content poses significant security concerns. In this work, we first explore the vulnerability of T2V models to jailbreak attacks, building upon methods developed for text-to-image (T2I) models. We conduct a systematic evaluation of the transferability of T2I jailbreak techniques to T2V models, revealing that naively adapted attacks yield limited effectiveness due to the unique temporal and semantic challenges in video generation. To address these limitations, we propose PromptNeedling, a jailbreak attack method tailored for T2V models. Specifically, PromptNeedling optimizes the jailbreak prompts to simplify adversarial inputs and injects high-salience NSFW keywords in a controlled manner. We conduct extensive experiments on three open-source T2V models and evaluate two categories of NSFW content (nudity and gore & violence), showing that PromptNeedling achieves higher attack success rates than prior methods. These findings highlight the urgent need for developing effective defenses against jailbreak attacks to ensure the safety of T2V models. Yiyang Mu, Hongwei Li 0001, Rui Zhang 0086, Wenbo Jiang 0001, Wenshu Fan |
GLOBECOM | 5 |
| 2025 | BadComp: Backdoor Attack against Object Detection using Image Compression OperationabstractCurrently, object detection models have achieved widespread success in real-world applications, yet remain vulnerable to backdoor attacks. Existing backdoor methods often suffer from poor stealthiness or can be easily mitigated by standard image processing techniques. In this paper, we propose a novel stealthy backdoor attack to dynamically compress target object-oriented data for backdoor embedding. By modifying the model, a high-frequency feature extraction module is added, so the model learns frequency-domain feature representations of compressed samples and achieve three attack objectives. Extensive experimental results demonstrate the effectiveness and robustness of the proposed method, achieving attack success rates exceeding 90% across three detectors. Xi Nie, Hongwei Li 0001, Wenbo Jiang 0001, Shuai Yuan 0009, Wenshu Fan, Jian Xiong 0007 |
GLOBECOM | 5 |
| 2025 | Making Audio Data UnlearnableabstractIn recent years, deep neural networks (DNNs) have driven rapid advancements in various fields. As DNNs continue to grow in size, the amount of training data required is also increasing. Many researchers crawl publicly available data from the internet for training, which raises issues of unauthorized exploitation and potential privacy leakage. Recent work against unauthorized exploitation primarily focuses on the image domain, while the audio domain remains underexplored. In this paper, we propose an effective method to generate audio unlearnable examples, which injects imperceptible perturbations into training samples, making them unlearnable for models to train. Specifically, we employ an error minimization optimization algorithm to iteratively optimize the generated perturbations. To ensure these perturbations remain imperceptible, we leverage both the Short-Time Objective Intelligibility (STOI) score and the$L_{2}$norm as measures to constrain the perceptibility of the unlearnable examples. To further improve the transferability of the unlearnable effects, we use an ensemble model and threshold constraint to effectively enhance generated unlearnable examples. Experimental results demonstrate that our method effectively generates samples that prevent models from accurately learning and making predictions, while remaining indistinguishable from clean samples to human observers. Wenshu Fan, Hongwei Li 0001, Wenbo Jiang 0001 |
ICC | 1 |
| 2025 | Adversarial Attack with Controllable TransferabilityabstractMachine Learning as a Service (MLaaS) providers often promote the robustness of their models as a selling point for their API services. Existing methods commonly evaluate the robustness of Deep Neural Networks (DNNs) by generating highly transferable adversarial examples However, dishonest MLaaS providers may deceive consumers by falsely exaggerating the robustness of their models. In this paper, contrary to enhancing the transferability of adversarial examples, we attempt to craft adversarial examples that fail under a shielded model. In other words, ensuring that the adversarial examples fail to attack a shielded model but successfully attack other models, which we refer to as controllable transferability. To achieve this goal, we propose the Controllable Transferability Method (CTM), a framework that generates adversarial examples with controllable transferability. CTM involves generating transferable adversarial examples and refining their transferability using gradient antagonism. Experimental results demonstrate that CTM achieves high transferability across models, with controlled adversarial effects on selected models. Jian Xiong 0007, Hongwei Li 0001, Wenbo Jiang 0001, Wenshu Fan, Shuai Yuan 0009 |
ICC | 4 |
| 2025 | Omni-Angle Assault: An Invisible and Powerful Physical Adversarial Attack on Face RecognitionabstractDeep learning models employed in face recognition (FR) systems have been shown to be vulnerable to physical adversarial attacks through various modalities, including patches, projections, and infrared radiation. However, existing adversarial examples targeting FR systems often suffer from issues such as conspicuousness, limited effectiveness, and insufficient robustness. To address these challenges, we propose a novel approach for adversarial face generation, UVHat, which utilizes ultraviolet (UV) emitters mounted on a hat to enable invisible and potent attacks in black-box settings. Specifically, UVHat simulates UV light sources via video interpolation and models the positions of these light sources on a curved surface, specifically the human head in our study. To optimize attack performance, UVHat integrates a reinforcement learning-based optimization strategy, which explores a vast parameter search space, encompassing factors such as shooting distance, power, and wavelength. Extensive experimental evaluations validate that UVHat substantially improves the attack success rate in black-box settings, enabling adversarial attacks from multiple angles with enhanced robustness. Shuai Yuan 0009, Hongwei Li 0001, Rui Zhang 0090, Hangcheng Cao, Wenbo Jiang 0001, Tao Ni 0003, Wenshu Fan, Qingchuan Zhao, Guowen Xu |
ICML | 7 |
| 2025 | Backdoor attacks against Hybrid Classical-Quantum Neural Networks
Ji Guo, Wenbo Jiang 0001, Rui Zhang 0090, Wenshu Fan, Jiachen Li 0002, Guoming Lu, Hongwei Li 0001 |
Neural Networks | 4 |
| 2024 | Adversarial Robustness Poisoning: Increasing Adversarial Vulnerability of the Model via Data PoisoningabstractDeep neural networks (DNNs) have become prevalent across various domains. However, recent research has revealed their vulnerability to data poisoning attacks, where adversaries inject poisoned data to compromise the usability of the target model. Traditional data poisoning attacks focus on reducing the test accuracy of the model, but they can be detected by model performance evaluation or mitigated by data cleaning. In contrast, we propose an Adversarial Robustness Poisoning Scheme (ARPS) that aims to decrease the adversarial robustness while preserving the normal-functionality of the target model. To achieve ARPS, we first separate the features of data into robust and non-robust features, where the non-robust features are human-imperceptible and more sensitive to adversarial perturbations. After that, we construct a dataset containing only non-robust features, which serves as the poisoning data. For the malicious dataset provider, the poisoned dataset can be constructed by adding poisoning data to the original dataset. For the malicious model provider, we employ the uncertainty-weighted multi-task learning technique to train the poisoned model, facilitating a better balance between functionality-preserving (good accuracy) and attack effectiveness (bad robustness). Extensive experiments are carried out to illustrate the effectiveness of ARPS in weakening adversarial training and amplifying adversarial attacks, as well as the stealthiness of ARPS in escaping the defense of data cleaning and model fine-tuning. Additionally, we propose some potential countermeasures against ARPS, including regularization and data augmentations. Wenbo Jiang 0001, Hongwei Li 0001, Wenshu Fan, Rui Zhang 0086 |
GLOBECOM | 4 |
| 2024 | Efficient Byzantine-Robust and Privacy-Preserving Federated Learning on Compressive DomainabstractData privacy and resistance against poisoning attack (Byzantine-robustness) are two critical concerns of federated learning (FL). Addressing the two issues simultaneously is challenging, since the privacy-preserving mechanism tends to make the data be indistinguishable, whereas Byzantine-robustness methods require access for the data to make a comprehensive analysis. To solve this problem, in this article, we propose a novel defender for privacy-ensured Byzantine-robust FL on a compressive domain. Unlike existing works that mainly using computation-intensive techniques, our method leverages compressive sensing (CS) as a lightweight encryption to protect the data privacy, while maintaining the possibility of Byzantine-robustness analysis on the encrypted (compressive) model update (i.e., gradient). Our key insight is that the cosine similarity can be approximately measured on the compressive measurements of any two normalized vectors, thus makes it be feasible to identify the malicious gradients on the CS compressive domain. We theoretically prove the correctness of our method. Notably, due to the dimensionality reduction of CS, the computation and communication overhead of our system can be significantly reduced. This makes our scheme be fit for applying in the applications with resource-constrained devices, such as Internet of Things (IoT). Experimental results demonstrate the effectiveness and efficiency of our method. Guiqiang Hu, Hongwei Li 0001, Wenshu Fan, Yushu Zhang 0001 |
IEEE Internet Things J. | 3 |
| 2024 | Stealthy Targeted Backdoor Attacks Against Image CaptioningabstractIn recent years, there is an explosive growth in multimodal learning. Image captioning, a classical multimodal task, has demonstrated promising applications and attracted extensive research attention. However, recent studies have shown that image caption models are vulnerable to some security threats such as backdoor attacks. Existing backdoor attacks against image captioning typically pair a trigger either with a predefined sentence or a single word as the targeted output, yet they are unrelated to the image content, making them easily noticeable as anomalies by humans. In this paper, we present a novel method to craft targeted backdoor attacks against image caption models, which are designed to be stealthier than prior attacks. Specifically, our method first learns a special trigger by leveraging universal perturbation techniques for object detection, then places the learned trigger in the center of some specific source object and modifies the corresponding object name in the output caption to a predefined target name. During the prediction phase, the caption produced by the backdoored model for input images with the trigger can accurately convey the semantic information of the rest of the whole image, while incorrectly recognizing the source object as the predefined target. Extensive experiments demonstrate that our approach can achieve a high attack success rate while having a negligible impact on model clean performance. In addition, we show our method is stealthy in that the produced backdoor samples are indistinguishable from clean samples in both image and text domains, which can successfully bypass existing backdoor defenses, highlighting the need for better defensive mechanisms against such stealthy backdoor attacks. Wenshu Fan, Hongwei Li 0001, Wenbo Jiang 0001, Meng Hao 0001, Shui Yu 0001, Xiao Zhang 0016 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2023 | Membership Inference Attacks Against the Graph ClassificationabstractRecently, there has been increasing interest in extending deep learning approaches to graph data. Graph representation learning has become an important way to fully utilize the information contained in graph data. Graph Neural Networks (GNNs) have demonstrated significant efficacy in various fields. Previous studies have shown that traditional machine learning models may lead to disclosure of private data, but the privacy risks of GNNs have not received enough attention. In this paper, we propose two attack methods based on the ground-truth label. Our attack approach covers two mainstream attack patterns, including training attack models based on neural networks and setting thresholds. To improve the effectiveness of attacks, we consider incorporating label information of the samples. Since the samples' distributions of different classes are different, the possibility of privacy leakage cannot be treated equally. In neural network-based attacks, we concatenate the label information into the input vector of the attack model. In threshold-based attacks, we set separate thresholds for each label category. We systematically evaluate the performance of membership inference attacks against graph-level classification. Our evaluation on three GNN structures and four benchmark datasets shows that GNNs for graph classification are more vulnerable to the improved attacks. On the DD dataset, our attack achieved an accuracy of 79%. Furthermore, we proposed two defense mechanisms to mitigate the privacy leakage caused by membership inference attacks. Junze Yang, Hongwei Li 0001, Wenshu Fan, Meng Hao 0001 |
GLOBECOM | 3 |
| 2020 | A Practical Black-Box Attack Against Autonomous Speech Recognition ModelabstractWith the wild applications of machine learning (ML) technology, automatic speech recognition (ASR) has made great progress in recent years. Despite its great potential, there are various evasion attacks of ML-based ASR, which could affect the security of applications built upon ASR. Up to now, most studies focus on white-box attacks in ASR, and there is almost no attention paid to black-box attacks where attackers can only query the target model to get output labels rather than probability vectors in audio domain. In this paper, we propose an evasion attack against ASR in the above-mentioned situation, which is more feasible in realistic scenarios. Specifically, we first train a substitute model by using data augmentation, which ensures that we have enough samples to train with a small number of times to query the target model. Then, based on the substitute model, we apply Differential Evolution (DE) algorithm to craft adversarial examples and implement black-box attack against ASR models from the Speech Commands dataset. Extensive experiments are conducted, and the results illustrate that our approach achieves untargeted attacks with over 70% success rate while still maintaining the authenticity of the original data well. Wenshu Fan, Hongwei Li 0001, Wenbo Jiang 0001, Guowen Xu, Rongxing Lu |
GLOBECOM | 1 |