EDBT 2026 Demo / reviewers in the wild / expert
Hongwei Yao
dblp:139/4887
· DBLP profile ↗
13ranked-venue papers
4as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Eguard: Defending LLM Embeddings Against Inversion Attacks via Text Mutual Information OptimizationabstractWhile text embeddings enable efficient semantic processing in LLMs, they remain vulnerable to inversion attacks that reconstruct sensitive original text. However, current defense methods typically treat text embeddings from the feature level independently, ignoring the exploitation of the mutual relation among the embedding construction pipeline. To address this limitation, we propose Eguard, a framework that effectively disrupts chains of relationships between the original semantic space and defended functional space. Our improvements manifest at two levels, i.e., the global-level and local-level mutual information. At the global level, we propose to minimize the statistical dependency between protected embeddings and their original inputs, effectively decoupling sensitive content from the semantic space accessible to adversaries. At the local level, we apply keyword-antonym contrastive learning to enforce semantic discriminability within the space of downstream utility. This synergy of global privacy control and local semantic alignment allows Eguard to achieve a superior privacy-utility trade-off than traditional defenses. Our approach significantly reduces privacy risks, protecting over 95 percent of tokens from inversion while maintaining high performance across downstream tasks consistent with original embeddings. Tiantian Liu 0002, Hongwei Yao, Feng Lin 0004, Zhan Qin, Kui Ren 0001 |
AAAI | 2 |
| 2026 | PromptCOS: Towards Content-Only System Prompt Copyright Auditing for LLMs
Yiming Li 0004, Hongwei Yao, Enhao Huang, Shuo Shao 0002, Yuyi Wang 0001, Zhibo Wang 0001, Dacheng Tao, Zhan Qin |
SP | 3 |
| 2026 | Reading Between the Lines: Towards Reliable Black-box LLM Fingerprinting via Zeroth-order Gradient EstimationabstractThe substantial investment required to develop Large Language Models (LLMs) makes them valuable intellectual property, raising significant concerns about copyright protection. LLM fingerprinting has emerged as a key technique to address this, which aims to verify a model's origin by extracting an intrinsic, unique signature (a ''fingerprint'') and comparing it to that of a source model to identify illicit copies. However, existing black-box fingerprinting methods often fail to generate distinctive LLM fingerprints. This ineffectiveness arises because black-box methods typically rely on model outputs, which lose critical information about the model's unique parameters due to the usage of non-linear functions. To address this, we first leverage Fisher Information Theory to formally demonstrate that the gradient of the model's input is a more informative feature for fingerprinting than the output. Based on this insight, we propose ZeroPrint, a novel method that approximates these information-rich gradients in a black-box setting using zeroth-order estimation. ZeroPrint overcomes the challenge of applying this to discrete text by simulating input perturbations via semantic-preserving word substitutions. This operation allows ZeroPrint to estimate the model's Jacobian matrix as a unique fingerprint. Experiments on the standard benchmark show ZeroPrint achieves a state-of-the-art effectiveness and robustness, significantly outperforming existing black-box methods. Shuo Shao 0002, Yiming Li 0004, Hongwei Yao, Zhan Qin |
WWW | 3 |
| 2026 | Combating Knowledge Corruption in Agent Systems: A Byzantine-Tolerant Secure Collaborative RAG FrameworkabstractWhile retrieval-augmented generation systems partially address the hallucination issues in large language models, it also introduces new vulnerabilities to knowledge corruption attacks. Adversaries exploit these vulnerabilities by poisoning documents provided by RAG system to manipulate LLM outputs. To counter this threat, we propose SecureCollaRAG, a Byzantine-tolerant collaborative RAG framework leveraging Multi-source Knowledge Validation Mechanism. Our approach enables agent system to securely verify document provenance through dynamic GNN-based credibility scoring, effectively preventing stealthy knowledge corruption attacks while preserving essential domain knowledge integrity. Through extensive evaluations and formal analysis, we demonstrate that SecureCollaRAG maintains robustness against attackers under non-IID data distributions. Daqing He, Zijian Zhang 0001, Ye Liu 0012, Jiamou Liu, Zhirui Zeng, Zhan Qin, Xin Li 0033, Hongwei Yao, Jincheng An, Yi Li 0008, Xiulei Liu, Liehuang Zhu |
WWW | 10 |
| 2026 | Model Backdoor Attack on Federated Learning Based on Parameter AnalysisabstractWith the increasingly widespread application of federated learning (FL) in various fields, the issue of backdoor attacks against FL has garnered significant attention from both academia and industry. While there has been some progress in researching backdoor attacks against FL, data backdoor attacks are easily mitigated in federated environments, and model backdoor attacks are susceptible to detection by defense mechanisms. Therefore, we propose a refined backdoor attack method tailored for FL under the image classification task. Our method involves the collaborative operation of three key modules. Firstly, the parameter importance analysis module identifies parameters with minimal impact on model performance, and creates parameter importance masks to provide precise targets for subsequent operations. Subsequently, the activation difference computation module calculates the activation differences between backdoor samples and benign samples to locate trigger-sensitive parameters. Our method implants the backdoor by flipping and zeroing the precisely located layer parameters, while maintaining the model's classification performance on benign samples. Experimental results show that our method is feasible to achieve an average attack success rate more than 99% across the three victim models. This demonstrates the effectiveness of our method in FL environments and its robustness against various FL defense mechanisms. Bo Wang 0024, Maozhen Zhang, Wei Wang 0025, Hongwei Yao |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2026 | FIT-Print: Toward False-Claim-Resistant Model Ownership Verification via Targeted FingerprintabstractModel fingerprinting has emerged as a crucial mechanism for safeguarding the intellectual property of open-source models, offering a non-intrusive approach that requires no modifications to the protected model. However, our analysis reveals that existing fingerprinting techniques are fundamentally vulnerable to false claim attacks, wherein adversaries can fraudulently assert ownership over independent third-party models. We demonstrate that this vulnerability stems from the untargeted nature of current methods, which evaluate model similarity based on arbitrary sample outputs rather than alignment with a specific, predefined reference. To mitigate this vulnerability, we introduce FIT-Print, a targeted fingerprinting paradigm that actively counters false claim attacks. Specifically, FIT-Print leverages optimization to transform the fingerprint into a verifiable, targeted signature. Building upon this foundation, we propose two black-box fingerprinting methods, the bit-wise FIT-ModelDiff and the list-wise FIT-LIME, which utilize output distances and feature attributions as robust model signatures, respectively. Extensive evaluations across benchmark models and datasets show that our framework perfectly neutralizes false claim attacks (100% defense success rate) and eliminates false alarms on independent models (0.0%), all while maintaining a 100% ownership verification rate against diverse model reuse techniques. Shuo Shao 0002, Haozhe Zhu, Yiming Li 0004, Hongwei Yao, Tianwei Zhang 0004, Zhan Qin |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | Explanation as a Watermark: Towards Harmless and Multi-bit Model Ownership Verification via Watermarking Feature Attribution
Shuo Shao 0002, Yiming Li 0004, Hongwei Yao, Yiling He, Zhan Qin, Kui Ren 0001 |
NDSS | 3 |
| 2025 | Artificial intelligence security and privacy: a surveyabstractAbstract Artificial intelligence (AI) is revolutionizing both industries and reshaping the global economy. However, the rapid advancement of AI technologies brings significant security and privacy challenges. Recent incidents highlight vulnerabilities in AI systems, such as data leakage and malicious code injection, leading to severe financial losses and privacy breaches. Although existing studies have discussed specific security threats, they often lack detailed granularity and cover a limited scope. In this survey, we fill this gap by systematically categorizing and analyzing the threats and countermeasures in AI systems, which span both the training and inference stages, encompass centralized and distributed settings, and address both conventional and foundation AI models. By reviewing existing literature, we aim to provide AI researchers and practitioners with a thorough understanding of system vulnerabilities and current countermeasures. We hope to inspire further research into robust solutions, ultimately contributing to the development of resilient AI technologies. Xinlei He 0001, Guowen Xu, Xingshuo Han, Qian Wang 0002, Lingchen Zhao, Chao Shen 0001, Chenhao Lin, Zhengyu Zhao 0001, Qian Li 0024, Le Yang 0007, Shouling Ji, Shaofeng Li 0001, Haojin Zhu, Zhibo Wang 0001, Tianqing Zhu, Qi Li 0002, Chaoxiang He, Hongsheng Hu, Shuo Wang 0012, Shifeng Sun 0001, Hongwei Yao, Qinyu Zhang 0001, Kai Chen 0012, Yue Zhao 0027, Hongwei Li 0001, Xinyi Huang 0001, Dengguo Feng |
Sci. China Inf. Sci. | 23 |
| 2025 | FDINet: Protecting Against DNN Model Extraction Using Feature Distortion IndexabstractMachine Learning as a Service (MLaaS) platforms have gained popularity due to their accessibility, cost-efficiency, scalability, and rapid development capabilities. However, recent research has highlighted the vulnerability of cloud-based models in MLaaS to model extraction attacks. In this paper, we introduce FDINet, a novel defense mechanism that leverages the feature distribution of deep neural network (DNN) models. Concretely, by analyzing the feature distribution from the adversary's queries, we reveal that the feature distribution of these queries deviates from that of the model's problem domain. Based on this key observation, we propose Feature Distortion Index (FDI), a metric designed to quantitatively measure the feature distribution deviation of received queries. The proposed FDINet utilizes FDI to train a binary detector and exploits FDI similarity to identify colluding adversaries from distributed extraction attacks. We conduct extensive experiments to evaluate FDINet against six state-of-the-art extraction attacks on four benchmark datasets and four popular model architectures. Empirical results demonstrate the following findings: (1) FDINet proves to be highly effective in detecting model extraction, achieving a100% detection accuracyon DFME and DaST. (2) FDINet is highly efficient, using just 50 queries to raise an extraction alarm with anaverage confidence of 96.08%for GTSRB. (3) FDINet exhibits the capability to identify colluding adversaries with an accuracyexceeding 91%. Additionally, it demonstrates the ability to detect two types of adaptive attacks. Hongwei Yao, Zheng Li 0023, Haiqin Weng, Zhan Qin, Kui Ren 0001 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2024 | PoisonPrompt: Backdoor Attack on Prompt-Based Large Language ModelsabstractPrompts have significantly improved the performance of pre-trained Large Language Models (LLMs) on various downstream tasks recently, making them increasingly indispensable for a diverse range of LLM application scenarios. However, the backdoor vulnerability, a serious security threat that can maliciously alter the victim model’s normal predictions, has not been sufficiently explored for prompt-based LLMs. In this paper, we present PoisonPrompt, a novel backdoor attack capable of successfully compromising both hard and soft prompt-based LLMs. We evaluate the effectiveness, fidelity, and robustness of PoisonPrompt through extensive experiments on three popular prompt methods, using six datasets and three widely used LLMs. Our findings highlight the potential security threats posed by backdoor attacks on prompt-based LLMs and emphasize the need for further research in this area. Hongwei Yao, Jian Lou 0001, Zhan Qin |
ICASSP | 1 |
| 2024 | PromptCARE: Prompt Copyright Protection by Watermark Injection and VerificationabstractLarge language models (LLMs) have witnessed a meteoric rise in popularity among the general public users over the past few months, facilitating diverse downstream tasks with human-level accuracy and proficiency. Prompts play an essential role in this success, which efficiently adapt pre-trained LLMs to task-specific applications by simply prepending a sequence of tokens to the query texts. However, designing and selecting an optimal prompt can be both expensive and demanding, leading to the emergence of Prompt-as-a-Service providers who profit by providing well-designed prompts for authorized use. With the growing popularity of prompts and their indispensable role in LLM-based services, there is an urgent need to protect the copyright of prompts against unauthorized use.In this paper, we propose PromptCARE, the first framework for prompt copyright protection through watermark injection and verification. Prompt watermarking presents unique challenges that render existing watermarking techniques developed for model and dataset copyright verification ineffective. PromptCARE overcomes these hurdles by proposing watermark injection and verification schemes tailor-made for characteristics pertinent to prompts and the natural language domain. Extensive experiments on six well-known benchmark datasets, using three prevalent pre-trained LLMs (BERT, RoBERTa, and Facebook OPT-1.3b), demonstrate the effectiveness, harmlessness, robustness, and stealthiness of PromptCARE. Hongwei Yao, Jian Lou 0001, Zhan Qin, Kui Ren 0001 |
SP | 1 |
| 2024 | RemovalNet: DNN Fingerprint Removal AttacksabstractWith the performance of deep neural networks (DNNs) remarkably improving, DNNs have been widely used in many areas. Consequently, the DNN model has become a valuable asset, and its intellectual property is safeguarded by ownership verification techniques (e.g., DNN fingerprinting). However, the feasibility of the DNN fingerprint removal attack and its potential influence remains an open problem. In this paper, we perform the first comprehensive investigation of DNN fingerprint removal attacks. Generally, the knowledge contained in a DNN model can be categorized into general semantic and fingerprint-specific knowledge. To this end, we propose a min-max bilevel optimization-based DNN fingerprint removal attack namedRemovalNet, to evade model ownership verification. The lower-level optimization is designed to remove fingerprint-specific knowledge. While in the upper-level optimization, we distill the victim model's general semantic knowledge to maintain the surrogate model's performance. We conduct extensive experiments to evaluate thefidelity,effectiveness, andefficiencyof theRemovalNetagainst four advanced defense methods on six metrics. The empirical results demonstrate that (1) theRemovalNetiseffective. After our DNN fingerprint removal attack, the model distance between the target and surrogate models is ×100 times higher than that of the baseline attacks, (2) theRemovalNetisefficient. It uses only 0.2% (400 samples) of the substitute dataset and 1,000 iterations to conduct our attack. Besides, compared with advanced model stealing attacks, theRemovalNetsaves nearly 85% of computational resources at most, (3) theRemovalNetachieves highfidelitythat the created surrogate model maintains high accuracy after the DNN fingerprint removal process. Hongwei Yao, Zheng Li 0023, Kunzhe Huang, Jian Lou 0001, Zhan Qin, Kui Ren 0001 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2022 | Classifying between computer generated and natural images: An empirical study from RAW to JPEG format
Xiangyang Luo 0001, Hongwei Yao |
J. Vis. Commun. Image Represent. | 3 |