Peizhuo Lv

dblp:289/1355 · DBLP profile ↗
← Back
21ranked-venue papers
7as first author
21since 2021 · last 2026
0000-0002-2671-4314ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 12 · 5 first-author · 12 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PROMPRINT: Prompt Fingerprinting via First-Token Response for LLM App Cloning Detection
abstract
As Large Language Model applications (LLM apps) become widespread, system prompts that determine app behavior are increasingly regarded as intellectual property, raising concerns about leakage.Recent studies show that this threat is no longer theoretical, revealing the prevalence of cloned apps replicating system prompts from others on real-world platforms.These clones pose risks of copyright infringement and malicious misuse, highlighting the need for early and reliable detection.In this paper, we propose PROMPRINT, a novel fingerprinting approach for detecting cloned LLM apps without exposing their system prompts.Motivated by the observation that different system prompts yield distinct first output token distributions for the same query, PROMPRINT optimizes queries that induce the LLM to generate a specific first output token associated with the given system prompt, resulting in distinctive query-first-token pairs.Experiments on four instruction-tuned LLMs show that generated pairs effectively identify the corresponding system prompts, achieving over 74% probability of generating the target token while remaining below 2.2% on average under other prompts.Furthermore, we demonstrate that our fingerprinting remains robust to partial system prompt modifications and effective under the injection of adversarial instructions.
Peizhuo Lv, Yeonjoon Lee
ACL (1)2
2026 ReasMark: A Robust Watermark for Attributing LLM Reasoning Under Knowledge Distillation Attacks
abstract
Peizhuo Lv, Ruihua Zhou, Yunpeng Li, Ruigang Liang, Xingshuo Han, XiaoFeng Wang, Wei Dong, Yuling Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Peizhuo Lv, Ruihua Zhou, Ruigang Liang, Xingshuo Han
ACL (1)1
2026 Privacy-preserving for user-uploaded images and text in Vision-Language Models
Zixiang Liu, Chi Chen 0001, Shuguang Yuan 0003, Weilong Huang, Xiaojie Zhu, Peizhuo Lv
Comput. Secur.6
2026 Electromagnetic interference (EMI) backdoor: An EMI-based backdoor attack against computer vision systems
abstract
Recently, computer vision systems, for example, smart traffic surveillance systems, facial recognition systems, etc., have significantly changed our daily life. Even though the neural networks in such systems are known to suffer from backdoor attacks, causing the backdoored models to behave well on benign samples but maliciously on controlled samples (with triggers applied to activate the backdoor), it is generally believed that most of the triggers, when used in physical attacks, are noticeable to victim users and not robust in various settings, such as different angles, distances, lighting conditions, etc. In this paper, we leverage electromagnetic interference (EMI) to produce a specific pattern distortion in images captured by the camera system and utilize the pattern distortion as the backdoor trigger. To avoid the overhead of manually collecting poisoned images, we introduce a simulation sample generation approach, converting clean images to poisoned ones by simulating the distortion caused by EMI against the camera system. Additionally, we propose a contrast loss function to enhance the generalization of backdoor features, improving triggers’ capability to activate the embedded backdoors. We conduct extensive physical experiments using diverse deep neural networks across various camera systems in different practical environments, achieving a 92.54% average backdoor success rate.
Mengjie Sun, Peizhuo Lv, Shengzhi Zhang, Jianshuo Liu, Kai Chen 0012, Hong Li 0004, Zhi Li 0018, Qinhong Jiang, Limin Sun 0001
J. Comput. Secur.2
2026 FedWM: Data-Free Watermarking for Model Ownership Protection in Federated Learning
abstract
The widespread adoption of federated learning has been driven by growing demands for privacy protection in model training. Federated learning enables multiple clients to collaboratively train a global model coordinated by a central server without sharing their raw data. However, when distributing the global model to clients, the central server faces significant security risks from malicious clients who may steal and misuse the model, thereby compromising its ownership. While existing watermarking techniques typically rely on main task data for ownership protection, their application in federated learning is limited since the server lacks access to this data, which remains with the clients. To address this challenge, we propose a novel data-free watermarking method. We utilize substitute data unrelated to the main task and improve efficiency by filtering out redundant samples. To optimize the watermarking process, we introduce a logits alignment-based optimization strategy that uses the substitute dataset with watermark triggers for effective embedding. Additionally, we propose a dynamic optimization algorithm to balance the trade-off between watermark embedding and main task. We comprehensively evaluate our approach across four datasets, four model architectures, and three mainstream deep learning tasks. Our experimental results demonstrate nearly perfect watermark performance while maintaining minimal impact on the main task. Notably, our watermarking method proves resistant to existing backdoor detection techniques, establishing its effectiveness, robustness and stealthiness.
Congyi Li, Peizhuo Lv, Xuejing Yuan, Shengzhi Zhang, Kai Chen 0012, Yingjiu Li
IEEE Trans. Dependable Secur. Comput.2
2026 ERASE: Bypassing Collaborative Detection of AI Counterfeit via Comprehensive Artifacts Elimination
abstract
The rapid advancement of AI-Generated Images (AIGI) has amplified concerns about increasingly undetectable deepfakes. Recent adversarial techniques further worsen this problem by enhancing the imperceptibility of synthetic forgeries to both human viewers and automated detection systems. To simulate realistic adversaries and expose detection vulnerabilities, AI-Generated Image Stealth (AIGI-S) methods specifically aim to make synthetic images harder to detect. However, existing AIGI-S approaches often lack universality and transferability across diverse detection models—especially in collaborative detection settings—and tend to prioritize machine deception over human perceptual fidelity, resulting in visible artifacts. Inspired by real-world antique painting forgery, we propose ERASE (comprehensivE counteRfeit ArtifactS Elimination), a stealth-oriented optimization framework designed for multi-detector environments. ERASE comprehensively suppresses generative artifacts and incorporates a perceptual optimization objective to improve deception against both detection algorithms and human examiners. Extensive evaluations across eight distinct generative subsets from the GenImage benchmark and fifteen detection models demonstrate that ERASE delivers substantially improved attack performance—improving single-detector evasion by +10.5% and collaborative detection evasion by +17.9%—while preserving high image quality.
Qianyun Yang, Peizhuo Lv, Yingjiu Li, Shengzhi Zhang, Zhiwei Chen 0003, Zixu Li 0001, Yupeng Hu 0003
IEEE Trans. Dependable Secur. Comput.2
2026 HEFLGuard: Backdoor Detection in Homomorphic Encryption-Based Federated Learning
abstract
Homomorphic encryption-based federated learning (HEFL) strengthens privacy by aggregating encrypted model updates, but it also renders existing backdoor defenses that assume plaintext updates inapplicable. We present HEFLGuard, a single-server backdoor detection framework for HEFL in which the server constructs overlapping validation models from encrypted client groups and clients locally compare logits of the global and validation models on benign samples to expose backdoor behavior. HEFLGuard further combines consistency verification across non-IID validation groups with Byzantine fault-tolerant aggregation of client reports, ensuring robustness under heterogeneous data and Byzantine participants. We evaluate HEFLGuard on seven vision/text benchmarks under three backdoor types across IID and non-IID settings. HEFLGuard consistently reduces ASR from near 100% to nearly the nobackdoor level while keeping the drop in clean accuracy within 2.5%. Compared with prior work, HEFLGuard achieves higher robustness and deployability.
Congyi Li, Peizhuo Lv, Jinwen He, Kai Chen 0012
IEEE Trans. Inf. Forensics Secur.2
2025 RepeatLeakage: Leak Prompts from Repeating as Large Language Model Is a Good Repeater
abstract
With the development of large language models (LLMs), numerous online applications based on these models have emerged. As system prompts significantly influence the performance of LLMs, many such applications conceal their system prompts and regard them as intellectual property. Consequently, numerous efforts have been made to steal these system prompts. However, for applications that do not publicly disclose their system prompts, previously stolen prompts have low confidence. This is because previous methods rely on confirmation from application developers, which is unrealistic since developers may be unwilling to acknowledge that their system prompts have been leaked. We observed a phenomenon: when an LLM performs repetitive tasks, it accurately repeats based on the context rather than relying on its internal model parameters. We validated this phenomenon by comparing the results of two different inputs—repetitive tasks and knowledge-based tasks—under conditions of normal execution, contaminated execution, and partially restored execution. By contaminating the input nouns and then partially restoring them using data from the normal execution's intermediate layers, we measured the accuracies of both task types across these three execution processes. Based on this phenomenon, we propose a high-confidence leakage method called RepeatLeakage. By specifying the range that the model needs to repeat and encouraging the model not to change the format, we manage to extract its system prompt and conversation contexts. We validated the repetition phenomenon on multiple open-source models and successfully designed prompts using RepeatLeakage to leak contents from the actual system prompts of GPT-Store and publicly available ChatGPT conversation contexts. Finally, we tested RepeatLeakage in real environments such as ChatGPT web, successfully leaking their system prompts and conversation contexts.
Peizhuo Lv
AAAI3
2025 RAG-WM: An Efficient Black-Box Watermarking Approach for Retrieval-Augmented Generation of Large Language Models
abstract
In recent years, tremendous success has been witnessed in Retrieval-Augmented Generation (RAG), widely used to enhance Large Language Models (LLMs) in domain-specific, knowledge-intensive, and privacy-sensitive tasks. However, attackers may steal those valuable RAGs and deploy or commercialize them, making it essential to detect Intellectual Property (IP) infringement. Most existing ownership protection solutions, such as watermarks, are designed for relational databases and texts. They cannot be directly applied to RAGs because relational database watermarks require white-box access to detect IP infringement, which is unrealistic for the knowledge base in RAGs. Meanwhile, post-processing by the adversary's deployed LLMs typically destructs text watermark information. To address those problems, we propose a novel black-box ''knowledge watermark'' approach, named RAG-WM, to detect IP infringement of RAGs. RAG-WM uses a multi-LLM interaction framework, comprising a Watermark Generator, Shadow LLM & RAG, and Watermark Discriminator, to create watermark texts based on watermark entity-relationship tuples and inject them into the target RAG. We evaluate RAG-WM across three domain-specific and two privacy-sensitive tasks on four benchmark LLMs. Experimental results show that RAG-WM effectively detects the stolen RAGs in various deployed LLMs. Furthermore, RAG-WM is robust against paraphrasing, unrelated content removal, knowledge insertion, and knowledge expansion attacks. Lastly, RAG-WM can also evade watermark detection approaches, highlighting its promising application in detecting IP infringement of RAG systems.
Peizhuo Lv, Mengjie Sun, Hao Wang 0034, XiaoFeng Wang 0001, Shengzhi Zhang, Kai Chen 0012, Limin Sun 0001
CCS1
2025 A Model Stealing Attack Against Multi-Exit Networks
abstract
Compared to traditional neural networks with a single output channel, a multi-exit network has multiple exits that allow for early outputs from the model's intermediate layers, thus significantly improving computational efficiency while maintaining similar main task accuracy. Existing model stealing attacks can only steal the model's utility while failing to capture its output strategy, i.e., a set of thresholds used to determine from which exit to output. This leads to a significant decrease in computational efficiency for the extracted model, thereby losing the advantage of multi-exit networks. In this paper, we propose the first model stealing attack against multi-exit networks to extract both the model utility and the output strategy. We employ Kernel Density Estimation to analyze the target model's output strategy and use performance loss and strategy loss to guide the training of the extracted model. Furthermore, we design a novel output strategy search algorithm to maximize the consistency between the victim model and the extracted model's output behaviors. In experiments across multiple multi-exit networks and benchmark datasets, our method always achieves accuracy and efficiency closest to the victim models.
Peizhuo Lv, Kai Chen 0012, Shengzhi Zhang, Yuling Cai, Fan Xiang
ICASSP2
2025 An Efficient White-box LLM Watermarking for IP Protection on Online Market Platforms
abstract
Online market platforms serve as a central hub for sharing and deploying AI models among researchers, developers, and companies. In this context, watermarking techniques are essential to protect intellectual property (IP), preventing unauthorized use and duplication of large language models (LLMs). Two key challenges arise: (i) These platforms host diverse LLMs, yet current watermarking techniques are only tailored to specific models, such as fine-tuned or quantized LLMs. (ii) Efficient watermarking is critical. However, traditional methods require substantial data and costly hardware, which limits their feasibility. In this paper, we propose an efficient white-box LLM watermarking technique called ELLMark. This method treats LLMs as multi-layered matrices while embedding watermarks only relies on modifying the model's weights. To preserve LLMs' performance, it filters weights by correlations with the activation magnitudes and downstream tasks, then modifies weights as minimal as possible via histogram modulation. Notably, all phases are training-free with low hardware resources, making it efficient for online platforms. We conduct extensive experiments to evaluate the effectiveness of ELLMark on LLaMA-3, OPT, and Phi-3 LLMs. The results demonstrate that it achieves 100% success in watermark detection while preserving model performance. Moreover, the preprocessing, encoding, and decoding processes remain efficient, taking less than 7 minutes, 12 minutes, and 18 seconds, respectively, for models with 80B parameters. Lastly, it exhibits robustness against parameter overwriting, re-watermarking, forging, fine-tuning, and pruning attacks.
Shuguang Yuan 0003, Xingyu Su, Peizhuo Lv, Weiji Xue, Jing Yu 0007, Xiaojie Zhu, Chi Chen 0001
KDD (2)3
2025 Hot-Swap MarkBoard: An Efficient Black-box Watermarking Approach for Large-scale Model Distribution
abstract
Recently, Deep Learning (DL) models have been increasingly deployed on end-user devices as On-Device AI, offering improved efficiency and privacy. However, this deployment trend poses more serious Intellectual Property (IP) risks, as models are distributed on numerous local devices, making them vulnerable to theft and redistribution. Most existing ownership protection solutions (e.g., backdoor-based watermarking) are designed for cloud-based AI-as-a-Service (AIaaS) and are not directly applicable to large-scale distribution scenarios, where each user-specific model instance must carry a unique watermark. These methods typically embed a fixed watermark, and modifying the embedded watermark requires retraining the model. To address these challenges, we propose Hot-Swap MarkBoard, an efficient watermarking method. It encodes user-specific n-bit binary signatures by independently embedding multiple watermarks into a multi-branch Low-Rank Adaptation (LoRA) module, enabling efficient watermark customization without retraining through branch swapping. A parameter obfuscation mechanism further entangles the watermark weights with those of the base model, preventing removal without degrading model performance. The method supports black-box verification and is compatible with various model architectures and DL tasks, including classification, image generation, and text generation. Extensive experiments across three types of tasks and six backbone models demonstrate our method's superior efficiency and adaptability compared to existing approaches, achieving 100% verification accuracy.
Zhicheng Zhang 0002, Peizhuo Lv, Mengke Wan, Jiang Fang, Diandian Guo, Yezeng Chen, Yinlong Liu, Jiyan Sun, Liru Geng
ACM Multimedia2
2024 DataElixir: Purifying Poisoned Dataset to Mitigate Backdoor Attacks via Diffusion Models
abstract
Dataset sanitization is a widely adopted proactive defense against poisoning-based backdoor attacks, aimed at filtering out and removing poisoned samples from training datasets. However, existing methods have shown limited efficacy in countering the ever-evolving trigger functions, and often leading to considerable degradation of benign accuracy. In this paper, we propose DataElixir, a novel sanitization approach tailored to purify poisoned datasets. We leverage diffusion models to eliminate trigger features and restore benign features, thereby turning the poisoned samples into benign ones. Specifically, with multiple iterations of the forward and reverse process, we extract intermediary images and their predicted labels for each sample in the original dataset. Then, we identify anomalous samples in terms of the presence of label transition of the intermediary images, detect the target label by quantifying distribution discrepancy, select their purified images considering pixel and feature distance, and determine their ground-truth labels by training a benign model. Experiments conducted on 9 popular attacks demonstrates that DataElixir effectively mitigates various complex attacks while exerting minimal impact on benign accuracy, surpassing the performance of baseline defense methods.
Jiachen Zhou 0001, Peizhuo Lv, Yibing Lan, Guozhu Meng, Kai Chen 0012, Hualong Ma
AAAI2
2024 SSL-WM: A Black-Box Watermarking Approach for Encoders Pre-trained by Self-Supervised Learning
Peizhuo Lv, Shenchen Zhu, Shengzhi Zhang, Kai Chen 0012, Ruigang Liang, Chang Yue, Fan Xiang, Yuling Cai, Hualong Ma, Guozhu Meng
NDSS1
2024 KGDist: A Prompt-Based Distillation Attack against LMs Augmented with Knowledge Graphs
abstract
With Knowledge Graph (KG) increasingly applied in various fields, the integration of KG has gained significant attention to augment the knowledge-specific task capabilities of language models (LMs). However, constructing and maintaining large KGs, much like LMs, can be expensive and challenging, often requiring extensive domain knowledge and human resources. This makes KG a valuable resource potentially vulnerable to theft threats from attackers. In this paper, we present KGDist, the first prompt-based KG distillation technique for extracting KG knowledge from KG+LM augmented models. Through iterations of prompt-based queries, we can steal a substitute KG containing task domain knowledge from the original KG. First of all, we initialize entities from a small scale task-specific corpus. Then, we construct specific task prompts for querying the victim LMs. According to the model outputs, we iteratively select entities showing strong correlation and reconstruct the relation edges for subsequent prompt crafting. We also propose a multi-granularity prompt construction method for reducing the querying cost. After acquiring the extracted KG, we launch a relation type-based pruning to cut off redundant edges forming cycles decreasing the performance of distilled KGs. We evaluate the effectiveness of KGDist on five benchmark KG+LM models designed for various tasks. Results demonstrate that our attack successfully extracts the distilled KGs with minimal performance degradation (under 2.4%) applied on LMs and less storage space. And also, the mechanism we apply greatly saves API queries compared to brute force method. In addition, further experiments demonstrate that we can split the KG knowledge from the LM noises effectively, and the distilled KGs have similar properties in knowledge distribution and graph structures to the original ones. Our code is available at https://github.com/Haro-M/KGDist.
Hualong Ma, Peizhuo Lv, Kai Chen 0012, Jiachen Zhou 0001
RAID2
2024 MEA-Defender: A Robust Watermark against Model Extraction Attack
abstract
Recently, numerous highly-valuable Deep Neural Networks (DNNs) have been trained using deep learning algorithms. To protect the Intellectual Property (IP) of the original owners over such DNN models, backdoor-based watermarks have been extensively studied. However, most of such watermarks fail upon model extraction attack, which utilizes input samples to query the target model and obtains the corresponding outputs, thus training a substitute model using such input-output pairs. In this paper, we propose a novel watermark to protect IP of DNN models against model extraction, named MEA-Defender. In particular, we obtain the watermark by combining two samples from two source classes in the input domain and design a watermark loss function that makes the output domain of the watermark within that of the main task samples. Since both the input domain and the output domain of our watermark are indispensable parts of those of the main task samples, the watermark will be extracted into the stolen model along with the main task during model extraction. We conduct extensive experiments on four model extraction attacks, using five datasets and six models trained based on supervised learning and self-supervised learning algorithms. The experimental results demonstrate that MEA-Defender is highly robust against different model extraction attacks, and various watermark removal/detection approaches.
Peizhuo Lv, Hualong Ma, Kai Chen 0012, Jiachen Zhou 0001, Shengzhi Zhang, Ruigang Liang, Shenchen Zhu
SP1
2023 Invisible Backdoor Attacks Using Data Poisoning in Frequency Domain
abstract
Backdoor attacks have become a significant threat to deep neural networks (DNNs), whereby poisoned models perform well on benign samples but produce incorrect outputs when given specific inputs with a trigger. These attacks are usually implemented through data poisoning by injecting poisoned samples (samples patched with a trigger and mislabelled to the target label) into the dataset, and the models trained with that dataset will be infected with the backdoor. However, most current backdoor attacks lack stealthiness and robustness because of the fixed trigger patterns and mislabelling, which humans or some backdoor defense approach can easily detect. To address this issue, we propose a frequency-domain-based backdoor attack method that implements backdoor implantation without mislabeling the poisoned samples or accessing the training process. We evaluated our approach on four benchmark datasets and two popular scenarios: no-label self-supervised and clean-label supervised learning. The experimental results demonstrate that our approach achieved a high attack success rate (above 90%) on all tasks without significant performance degradation on main tasks and robust against mainstream defense approaches.
Chang Yue, Peizhuo Lv, Ruigang Liang, Kai Chen 0012
ECAI2
2023 DBIA: Data-Free Backdoor Attack Against Transformer Networks
abstract
Recently, transformer architecture has demonstrated its significance in both Natural Language Processing (NLP) and Computer Vision (CV) tasks. Although other network models are known to be vulnerable to the backdoor attack, which embeds triggers in the models and controls the models’ behavior when the triggers are presented, little is known about how such an attack performs on the transformer models. In this paper, we propose DBIA, a novel Data-free1Backdoor Attack against the CV-oriented transformer networks, leveraging the inherent attention mechanism of transformers to generate triggers and injecting the backdoor using a poisoned substitute dataset. We conducted extensive experiments using three benchmark transformers, i.e., ViT, DeiT, and Swin Transformer, on four mainstream image classification tasks, i.e., ImageNet, CIFAR-10, GTSRB, and Youtube Face. The evaluation results demonstrate that, with fewer resources, our approach can embed backdoors with a high success rate and a low impact on the performance of the victim transformers.
Peizhuo Lv, Hualong Ma, Jiachen Zhou 0001, Ruigang Liang, Kai Chen 0012, Shengzhi Zhang, Yunfei Yang 0001
ICME1
2023 A Data-free Backdoor Injection Approach in Neural Networks
Peizhuo Lv, Chang Yue, Ruigang Liang, Yunfei Yang 0001, Shengzhi Zhang, Hualong Ma, Kai Chen 0012
USENIX Security Symposium1
2023 Aliasing Backdoor Attacks on Pre-trained Models
Cheng'an Wei, Yeonjoon Lee, Kai Chen 0012, Guozhu Meng, Peizhuo Lv
USENIX Security Symposium5
2023 A Robustness-Assured White-Box Watermark in Neural Networks
abstract
Recently, stealing highly-valuable and large-scale deep neural network (DNN) models becomes pervasive. The stolen models may be re-commercialized, e.g., deployed in embedded devices, released in model markets, utilized in competitions, etc, which infringes the Intellectual Property (IP) of the original owner. Detecting IP infringement of the stolen models is quite challenging, even with the white-box access to them in the above scenarios, since they may have experienced fine-tuning, pruning, functionality-equivalent adjustment to destruct any embedded watermark. Furthermore, the adversaries may also attempt to extract the embedded watermark or forge a similar watermark to falsely claim ownership. In this article, we propose a novel DNN watermarking solution, named$HufuNet$, to detect IP infringement of DNN models against the above mentioned attacks. Furthermore, HufuNet is the first one theoretically proved to guarantee robustness against fine-tuning attacks. We evaluate HufuNet rigorously on four benchmark datasets with five popular DNN models, including convolutional neural network (CNN) and recurrent neural network (RNN). The experiments and analysis demonstrate that HufuNet is highly robust against model fine-tuning/pruning, transfer learning, kernels cutoff/supplement, functionality-equivalent attacks and fraudulent ownership claims, thus highly promising to protect large-scale DNN models in the real world.
Peizhuo Lv, Shengzhi Zhang, Kai Chen 0012, Ruigang Liang, Hualong Ma, Yue Zhao 0018, Yingjiu Li
IEEE Trans. Dependable Secur. Comput.1