VLDB 2026 Research / reviewers in the wild / expert
Rui Zhang 0086
dblp:60/2536-86
· DBLP profile ↗
17ranked-venue papers
6as first author
17since 2021 · last 2026
0000-0001-6412-9338ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 8 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Security and privacy · 3 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ConfGuard: A Simple and Effective Backdoor Detection for Large Language ModelsabstractBackdoor attacks pose a significant threat to Large Language Models (LLMs), where adversaries can embed hidden triggers to manipulate LLM's outputs. Most existing defense methods, primarily designed for classification tasks, are ineffective against the autoregressive nature and vast output space of LLMs, thereby suffering from poor performance and high latency. To address these limitations, we investigate the behavioral discrepancies between benign and backdoored LLMs in output space. We identify a critical phenomenon which we term sequence lock: a backdoored model generates the target sequence with abnormally high and consistent confidence compared to benign generation. Building on this insight, we propose ConfGuard, a lightweight and effective detection method that monitors a sliding window of token confidences to identify sequence lock. Extensive experiments demonstrate ConfGuard achieves a near 100% true positive rate (TPR) and a negligible false positive rate (FPR) in the vast majority of cases. Crucially, the ConfGuard enables real-time detection almost without additional latency, making it a practical backdoor defense for real-world LLM deployments. Rui Zhang 0086, Hongwei Li 0001, Wenshu Fan, Wenbo Jiang 0001, Qingchuan Zhao, Guowen Xu |
AAAI | 2 |
| 2026 | Efficient and Verifiable Data Statistical Analysis via Zero-knowledge Proofs
Hanxiao Chen 0001, Rui Zhang 0086, Pengzhi Xing, Meng Hao 0001, Hongwei Li 0001 |
ICC | 2 |
| 2026 | Backdoor Complications: A Comprehensive Analysis and Mitigation of the Unforeseen Consequences of Backdoor AttacksabstractPre-trained language models (PTLMs) have become integral to modern natural language processing (NLP), yet their reuse exposes them to supply chain risks such as backdoor attacks. Existing studies assume that attackers target specific downstream tasks, overlooking how a backdoored PTLM behaves when fine-tuned for unrelated applications. In practice, such unintended adaptation can trigger anomalous and inconsistent predictions, revealing the backdoor and compromising its stealthiness. We define this phenomenon asbackdoor complications, i.e., unintended behavioral side effects emerging on non-target tasks. This work presents the first systematic quantification and mitigation of backdoor complications. Through extensive experiments on 3 widely used PTLMs and 15 benchmark datasets, we show that complications are pervasive across both single- and multi-task attack settings, causing triggered outputs to collapse into arbitrary classes. To address this issue, we propose theComplication-Suppressed Backdoor Attack(CSBA), a task-agnostic, multi-objective framework that leverages auxiliary non-target datasets to suppress backdoor complications. CSBA effectively suppresses complications on unseen downstream tasks while maintaining near-perfect attack success rates. Our work reveals a critical side effect in backdoored PTLMs and provides a new perspective on the stealthiness and robustness of model supply chain security. Rui Zhang 0086, Hongwei Li 0001, Wenbo Jiang 0001, Hanxiao Chen 0001, Yuan Zhang 0006, Guowen Xu, Yang Zhang 0016 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2026 | Hidden Tail: Adversarial Attack for Stealthy Resource Consumption Against Vision-Language ModelsabstractVision-Language Models (VLMs) are increasingly deployed in real-world applications, but their high inference cost makes them vulnerable to resource consumption attacks. Prior attacks attempt to extend VLM output sequences by optimizing adversarial images, thereby increasing inference costs. However, these extended outputs often introduce irrelevant abnormal content, compromising attack stealthiness. This trade-off between effectiveness and stealthiness poses a major limitation for existing attacks. To address this challenge, we proposeHidden Tail, a stealthy resource consumption attack that crafts prompt-agnostic adversarial images, inducing VLMs to generate maximum-length outputs by appending special tokens invisible to users. Our method employs a composite loss function that balances semantic preservation, repetitive special token induction, and suppression of the end-of-sequence (EOS) token, optimized via a dynamic weighting strategy. Extensive experiments show thatHidden Tailoutperforms existing attacks, increasing output length by up to 19.2× and reaching the maximum token limit, while preserving attack stealthiness. These results highlight the urgent need to improve the robustness of VLMs against efficiency-oriented adversarial threats. Our code is available athttps://github.com/zhangrui4041/Hidden_Tail. Rui Zhang 0086, Tianli Yang, Wenbo Jiang 0001, Rui Zhang 0090, Qingchuan Zhao, Hongwei Li 0001, Yang Liu 0003, Guowen Xu |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2025 | Watch Out for Your Guidance on Generation! Exploring Conditional Backdoor Attacks against Large Language ModelsabstractMainstream backdoor attacks on large language models (LLMs) typically set a fixed trigger in the input instance and specific responses for triggered queries. However, the fixed trigger setting (e.g., unusual words) may be easily detected by human detection, limiting the effectiveness and practicality in real-world scenarios. To enhance the stealthiness of backdoor activation, we present a new poisoning paradigm against LLMs triggered by specifying generation conditions, which are commonly adopted strategies by users during model inference. The poisoned model performs normally for output under normal/other generation conditions, while becomes harmful for output under target generation conditions. To achieve this objective, we introduce BrieFool, an efficient attack framework. It leverages the characteristics of generation conditions by efficient instruction sampling and poisoning data generation, thereby influencing the behavior of LLMs under target conditions. Our attack can be generally divided into two types with different targets: Safety unalignment attack and Ability degradation attack. Our extensive experiments demonstrate that BrieFool is effective across safety domains and ability domains, achieving higher success rates than baseline methods, with 94.3% on GPT-3.5-turbo. Jiaming He, Wenbo Jiang 0001, Guanyu Hou, Wenshu Fan, Rui Zhang 0086, Hongwei Li 0001 |
AAAI | 5 |
| 2025 | Evaluating Robustness of Large Audio Language Models to Audio Injection: An Empirical StudyabstractLarge Audio-Language Models (LALMs) are increasingly deployed in real-world applications, yet their robustness against malicious audio injection remains underexplored.To address this gap, this study systematically evaluates five leading LALMs across four attack scenarios: Audio Interference Attack, Instruction Following Attack, Context Injection Attack, and Judgment Hijacking Attack.We quantitatively assess their vulnerabilities and resilience using metrics: the Defense Success Rate, Context Robustness Score, and Judgment Robustness Index.The experiments reveal significant performance disparities, with no single model demonstrating consistent robustness across all attack types.Attack effectiveness is significantly influenced by the position of the malicious content, particularly when injected at the beginning of a sequence.Furthermore, our analysis uncovers a negative correlation between a model's instruction-following capability and its robustness: models that strictly adhere to instructions tend to be more susceptible, whereas safety-aligned models exhibit greater resistance.To facilitate future research, this work introduces a comprehensive benchmark framework.Our findings underscore the critical need for integrating robustness into training pipelines and developing multi-modal defenses, ultimately facilitating the secure deployment of LALMs.The dataset used in this work is available on Hugging Face. Guanyu Hou, Jiaming He, Yinhang Zhou, Ji Guo, Yitong Qiao, Rui Zhang 0086, Wenbo Jiang 0001 |
EMNLP | 6 |
| 2025 | CLBA: A Cross-Lingual Backdoor Attack against Text-to-Image Diffusion ModelsabstractDiffusion-based Text-to-Image (T2I) synthesis has emerged as a transformative multimodal generation technology. However, its reliance on pre-trained models introduces severe backdoor risks, including output manipulation and privacy violations. Existing attacks using specific triggers demonstrate effectiveness in monolingual settings but tend to degrade in crosslingual scenarios due to semantic inconsistency. To bridge this gap, we propose a cross-lingual transfer backdoor attack that maintains robust attack efficacy across languages through single-language poisoning. Specifically, we introduce a linguistically informed trigger selection method that identifies semantically invariant words across languages. The trigger is deliberately and naturally embedded into prompts to minimize semantic disruption, thereby helping to evade potential defense mechanisms. Our approach leverages cross-lingual semantic alignment between high-resource and low-resource languages to enable stealthy and effective backdoor activation. We assume the adversary finetunes with limited data, without knowledge of the model internals or the original training data. Our attack demonstrates the potential to bypass existing defenses in T2I models and exposes critical security vulnerabilities in multilingual T2I systems. These findings highlight the urgent need for targeted security measures to mitigate backdoor threats and prevent malicious exploitation. Hongwei Li 0001, Rui Zhang 0086, Jiaming He, Wenbo Jiang 0001 |
GLOBECOM | 3 |
| 2025 | PromptNeedling: Jailbreaking Text-to-Video Generative ModelsabstractRecent advances in text-to-video (T2V) generation have enabled high-fidelity video synthesis from natural language descriptions. While these models offer promising applications, their capacity to generate not safe for work (NSFW) content poses significant security concerns. In this work, we first explore the vulnerability of T2V models to jailbreak attacks, building upon methods developed for text-to-image (T2I) models. We conduct a systematic evaluation of the transferability of T2I jailbreak techniques to T2V models, revealing that naively adapted attacks yield limited effectiveness due to the unique temporal and semantic challenges in video generation. To address these limitations, we propose PromptNeedling, a jailbreak attack method tailored for T2V models. Specifically, PromptNeedling optimizes the jailbreak prompts to simplify adversarial inputs and injects high-salience NSFW keywords in a controlled manner. We conduct extensive experiments on three open-source T2V models and evaluate two categories of NSFW content (nudity and gore & violence), showing that PromptNeedling achieves higher attack success rates than prior methods. These findings highlight the urgent need for developing effective defenses against jailbreak attacks to ensure the safety of T2V models. Yiyang Mu, Hongwei Li 0001, Rui Zhang 0086, Wenbo Jiang 0001, Wenshu Fan |
GLOBECOM | 3 |
| 2025 | SecInfer: Secure and Efficient Model Inference on Vertically Partitioned DataabstractDeep learning models have achieved unprecedented success in various domains, such as healthcare and finance. However, deploying model inference in real-world applications, where data is distributed among multiple entities, poses significant privacy concerns. Existing secure model inference work has limitations in computational overhead and scalability, especially when dealing with complex models and multiple parties with vertically partitioned data. In this work, we design and implement an efficient and scalable secure inference framework for vertically partitioned data, supporting execution with a large number of parties. Our work considers a semi-honest setting with all-but-one corruptions. The core of our framework is a series of secure and efficient protocols for complex non-linear functions of the model inference, such as ReLU and Maxpool. These protocols are designed based on secure multi-party computation preliminaries, significantly enhancing efficiency while maintaining rigorous security guarantees. We conduct comprehensive experiments to evaluate the performance of our framework. Experimental results show that SecInfer substantially improves the communication and computation performance of secure naive inference works by up to 3.71 × and 3.42 ×, respectively. Robert H. Deng, Hongwei Li 0001, Hanxiao Chen 0001, Meng Hao 0001, Pengzhi Xing, Jia Hu 0004, Rui Zhang 0086, Wenbo Jiang 0001 |
ICC | 7 |
| 2025 | Stealthy Physical Backdoor Attacks Against Traffic Sign Recognition SystemsabstractRecent advancements in deep learning have led to remarkable progress in autonomous driving technology, with deep neural network (DNN)-based traffic sign recognition systems (TSRS) playing a crucial role. However, recent studies indicate that TSRS are vulnerable to backdoor attacks, where the backdoor TSRS behaves normally on clean traffic signs but consistently misclassifies backdoor-triggered traffic signs into a designated target class. Notably, while backdoor attacks in the digital domain are effective, their effectiveness may diminish in the physical world due to quality degradation during image transmission. Existing physical backdoor attacks typically rely on specific stickers or transformations as backdoor triggers, which are not stealthy and natural enough in the physical world. To address these limitations, we propose two stealthy physical backdoor attacks against DNN-based TSRS from two different perspectives. On the one hand, we utilize the natural phenomenon of chipped paints on traffic signs as the backdoor trigger. Specifically, we develop an automatic traffic sign segmentation algorithm to identify the edges of the target sign and simulate chipped paint to create poisoned samples. On the other hand, instead of manipulating the target traffic sign, we use the specific filter lens (attached to the in-vehicle camera) as the backdoor trigger, where the parameters of the filter lens are optimized by the Genetic Algorithm (GA). Extensive experiments conducted on the GTSRB and TSRD datasets demonstrate the effectiveness of our proposed backdoor attacks in both digital and physical environments. Wenbo Jiang 0001, Hongwei Li 0001, Shuai Yuan 0009, Rui Zhang 0086, Qiyang Song |
ICC | 5 |
| 2025 | The Ripple Effect: On Unforeseen Complications of Backdoor AttacksabstractRecent research highlights concerns about the trustworthiness of third-party Pre-Trained Language Models (PTLMs) due to potential backdoor attacks.
These backdoored PTLMs, however, are effective only for specific pre-defined downstream tasks.
In reality, these PTLMs can be adapted to many other unrelated downstream tasks.
Such adaptation may lead to unforeseen consequences in downstream model outputs, consequently raising user suspicion and compromising attack stealthiness.
We refer to this phenomenon as backdoor complications.
In this paper, we undertake the first comprehensive quantification of backdoor complications.
Through extensive experiments using 4 prominent PTLMs and 16 text classification benchmark datasets, we demonstrate the widespread presence of backdoor complications in downstream models fine-tuned from backdoored PTLMs.
The output distribution of triggered samples significantly deviates from that of clean samples.
Consequently, we propose a backdoor complication reduction method leveraging multi-task learning to mitigate complications without prior knowledge of downstream tasks.
The experimental results demonstrate that our proposed method can effectively reduce complications while maintaining the efficacy and consistency of backdoor attacks. Rui Zhang 0086, Hongwei Li 0001, Wenbo Jiang 0001, Hanxiao Chen 0001, Yuan Zhang 0006, Guowen Xu, Yang Zhang 0016 |
ICML | 1 |
| 2025 | A Hidden Backdoor Attack via Formal Text Style Transfer in Language ModelsabstractNatural language processing (NLP) systems have been demonstrated to be vulnerable to backdoor attacks. Specifically, attackers embed the backdoor into the model by poisoning training data, producing the desired results when the input contains pre-defined triggers. Typical textual backdoor attacks adopt static triggers such as words or phrases, which make them detectable by existing defense methods. To enhance stealthiness, this paper introduces a hidden backdoor attack method utilizing formal text style transfer (FTST). Specifically, we adopt a formal text style transfer model to convert part of the benign training samples into formal samples, which serve as the backdoor samples. Compared to static textual triggers, FTST-based triggers can maintain original semantics while evading common defenses and human detections. We conduct extensive experiments on typical NLP tasks, including topic and sentiment classification tasks utilizing three prominent pre-trained language models and four datasets. The results show that our approach achieves the desired attack performance while preserving the normal-functionality of the model. Furthermore, compared to common word-level triggers and sentence-level triggers, our approach has been demonstrated to be more stealthy under GPT-2-based perplexity detection and more robust under backdoor defense methods. Hongwei Li 0001, Wenbo Jiang 0001, Rui Zhang 0086, Jiaming He, Hanxiao Chen 0001, Guowen Xu |
IJCNN | 4 |
| 2024 | Adversarial Robustness Poisoning: Increasing Adversarial Vulnerability of the Model via Data PoisoningabstractDeep neural networks (DNNs) have become prevalent across various domains. However, recent research has revealed their vulnerability to data poisoning attacks, where adversaries inject poisoned data to compromise the usability of the target model. Traditional data poisoning attacks focus on reducing the test accuracy of the model, but they can be detected by model performance evaluation or mitigated by data cleaning. In contrast, we propose an Adversarial Robustness Poisoning Scheme (ARPS) that aims to decrease the adversarial robustness while preserving the normal-functionality of the target model. To achieve ARPS, we first separate the features of data into robust and non-robust features, where the non-robust features are human-imperceptible and more sensitive to adversarial perturbations. After that, we construct a dataset containing only non-robust features, which serves as the poisoning data. For the malicious dataset provider, the poisoned dataset can be constructed by adding poisoning data to the original dataset. For the malicious model provider, we employ the uncertainty-weighted multi-task learning technique to train the poisoned model, facilitating a better balance between functionality-preserving (good accuracy) and attack effectiveness (bad robustness). Extensive experiments are carried out to illustrate the effectiveness of ARPS in weakening adversarial training and amplifying adversarial attacks, as well as the stealthiness of ARPS in escaping the defense of data cleaning and model fine-tuning. Additionally, we propose some potential countermeasures against ARPS, including regularization and data augmentations. Wenbo Jiang 0001, Hongwei Li 0001, Wenshu Fan, Rui Zhang 0086 |
GLOBECOM | 5 |
| 2024 | Instruction Backdoor Attacks Against Customized LLMs
Rui Zhang 0086, Hongwei Li 0001, Rui Wen 0002, Wenbo Jiang 0001, Yuan Zhang 0006, Michael Backes 0001, Yang Zhang 0016 |
USENIX Security Symposium | 1 |
| 2024 | Vertical Federated Learning Across Heterogeneous Regions for Industry 4.0abstractThis work investigates fine-grained data distribution in real-world federated learning (FL) applications, wherein training samples are distributed across multiple regions, and different clients within each region possess distinct features of local training samples. Furthermore, the datasets and models in these regions often exhibit heterogeneity, characterized by varying label distributions and model architectures, posing challenges to the model construction process. In this article, we propose a vertical federated learning (VFL) framework, named HeteroVFL, to address the data distribution complexities and overcome the hurdles posed by heterogeneous regions. Besides, we enhance the privacy of HeteroVFL by adopting differential privacy, a privacy-preserving technology by injecting measured noise into data based on a stochastic framework. We compare our HeteroVFL with existing solutions on three real-world datasets in simulations. The results demonstrate that HeteroVFL can achieve over 96% accuracy on MNIST, surpassing the accuracy of 90% in the state-of-the-art VFL benchmarks. Rui Zhang 0086, Hongwei Li 0001, Luoding Tian, Meng Hao 0001, Yuan Zhang 0006 |
IEEE Trans. Ind. Informatics | 1 |
| 2022 | Secure Feature Selection for Vertical Federated Learning in eHealth SystemsabstractPrivacy-preserving vertical federated learning (VFL) has been widely applied in electronic health (eHealth) systems. However, existing VFL schemes rarely consider the data pre-processing step including feature selection, which will lead to poor convergence rate and even damaging the model utility. In this paper, we propose an efficient and privacy-preserving feature selection scheme for VFL. Specifically, we first propose a general Gini-impurity based feature selection framework, which is compatible with most existing machine learning models in VFL. With the framework, we present two concrete protocols (dubbed πSS−FSand πH−FS, respectively) customized for different eHealth scenarios. πSS−FSexploits a lightweight additive secret sharing technique, such that it can be executed in comparable time as the evaluation of the plaintext scheme. πH−FSis a hybrid feature selection protocol that additionally utilizes a linear homomorphic encryption technique, to reduce the communication overhead at the cost of a moderate runtime. Moreover, extensive evaluations conducted on real-world medical datasets demonstrate that our scheme realizes up to 27% accuracy gains. Rui Zhang 0086, Hongwei Li 0001, Meng Hao 0001, Hanxiao Chen 0001, Yuan Zhang 0006 |
ICC | 1 |
| 2021 | Towards Lightweight and Efficient Distributed Intrusion Detection FrameworkabstractFederated learning (FL), as a promising distributed learning paradigm, has put many efforts into distributed intrusion detection systems (IDS), for defending against various malicious attacks, such as SQL injection and DDoS attacks. Compared with traditional IDS based on centralized deep learning (DL), FL-based solutions require not to share users' raw data while yielding better detection performance. However, state-of-the-art FL-based methods still suffer from two key limitations: 1) insufficient detection performance on non-independent and identically distributed (non-IID) data, and 2) high communication and computational overheads due to the utilization of large-scale neural network models. In this paper, we propose a lightweight collaborative intrusion detection framework, called CoLGBM, the first of its kind in the regime of decentralized IDS, where decision tree and light gradient boosting machine (LGBM) are combined for constructing the detection scheme. The main insight is that through combining user-trained decision trees (each user's decision tree is derived from its own data with unique distribution), our framework can perform effectively on non-IID data while working efficiently for handling enormous samples. Compared with the current FL-based methods, our CoLGBM achieves higher accuracy and lower overhead on both IID and non-IID data. Extensive experiment results demonstrate our scheme with high-level performance. Shuai Yuan 0009, Hongwei Li 0001, Rui Zhang 0086, Meng Hao 0001, Rongxing Lu |
GLOBECOM | 3 |