VLDB 2026 Research / reviewers in the wild / expert
Shaofeng Li 0001
dblp:15/8202-1
· DBLP profile ↗
23ranked-venue papers
4as first author
23since 2021 · last 2026
0000-0002-1491-4319ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 9 · 3 first-author · 9 since 2021Computer networks · 8 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Authority Backdoor: A Certifiable Backdoor Mechanism for Authoring DNNsabstractDeep Neural Networks (DNNs), as valuable intellectual property, face unauthorized use. Existing protections, such as digital watermarking, are largely passive; they provide only post-hoc ownership verification and cannot actively prevent the illicit use of a stolen model. This work proposes a proactive protection scheme, dubbed ``Authority Backdoor," which embeds access constraints directly into the model. In particular, the scheme utilizes a backdoor learning framework to intrinsically lock a model's utility, such that it performs normally only in the presence of a specific trigger (e.g., a hardware fingerprint). But in its absence, the DNN's performance degrades to be useless. To further enhance the security of the proposed authority scheme, the certifiable robustness is integrated to prevent an adaptive attacker from removing the implanted backdoor. The resulting framework establishes a secure authority mechanism for DNNs, combining access control with certifiable robustness against adversarial attacks. Extensive experiments on diverse architectures and datasets validate the effectiveness and certifiable robustness of the proposed framework. Shaofeng Li 0001, Tian Dong 0003, Xiangyu Xu 0001, Guangchi Liu, Zhen Ling 0001 |
AAAI | 2 |
| 2026 | A Needle in a Haystack: Defending Federated Learning Backdoor Attacks via Orthogonal Subnetwork Pruning
Zihan Ma 0008, Guangchi Liu, Xiangyu Xu 0001, Shaofeng Li 0001, Zhen Ling 0001, Junzhou Luo |
INFOCOM | 4 |
| 2026 | Cease at the Ultimate Goodness: Towards Efficient Website Fingerprinting Defense via Iterative Mutual Information Minimization
Zhen Ling 0001, Guangchi Liu, Shaofeng Li 0001, Junzhou Luo, Xinwen Fu |
NDSS | 4 |
| 2025 | FlexEmu: Towards Flexible MCU Peripheral EmulationabstractMicrocontroller units (MCUs) are widely used in embedded devices due to their low power consumption and cost-effectiveness. MCU firmware controls these devices and is vital to the security of embedded systems. However, performing dynamic security analyses for MCU firmware has remained challenging due to the lack of usable execution environments -- existing dynamic analyses cannot run on physical devices (e.g., insufficient computational resources), while building emulators is costly due to the massive amount of heterogeneous hardware, especially peripherals. Recent advances in automated peripheral emulation have made MCU emulation more scalable. However, these efforts only support limited peripherals and are hard to extend because they require ad-hoc adaptations. Chongqing Lei, Zhen Ling 0001, Xiangyu Xu 0001, Shaofeng Li 0001, Guangchi Liu, Kai Dong 0001, Junzhou Luo |
CCS | 4 |
| 2025 | The Philosopher's Stone: Trojaning Plugins of Large Language Models
Tian Dong 0003, Minhui Xue 0001, Guoxing Chen, Rayne Holland, Yan Meng 0001, Shaofeng Li 0001, Zhen Liu 0008, Haojin Zhu |
NDSS | 6 |
| 2025 | Depth Gives a False Sense of Privacy: LLM Internal States Inversion
Tian Dong 0003, Yan Meng 0001, Shaofeng Li 0001, Guoxing Chen, Zhen Liu 0008, Haojin Zhu |
USENIX Security Symposium | 3 |
| 2025 | Artificial intelligence security and privacy: a surveyabstractAbstract Artificial intelligence (AI) is revolutionizing both industries and reshaping the global economy. However, the rapid advancement of AI technologies brings significant security and privacy challenges. Recent incidents highlight vulnerabilities in AI systems, such as data leakage and malicious code injection, leading to severe financial losses and privacy breaches. Although existing studies have discussed specific security threats, they often lack detailed granularity and cover a limited scope. In this survey, we fill this gap by systematically categorizing and analyzing the threats and countermeasures in AI systems, which span both the training and inference stages, encompass centralized and distributed settings, and address both conventional and foundation AI models. By reviewing existing literature, we aim to provide AI researchers and practitioners with a thorough understanding of system vulnerabilities and current countermeasures. We hope to inspire further research into robust solutions, ultimately contributing to the development of resilient AI technologies. Xinlei He 0001, Guowen Xu, Xingshuo Han, Qian Wang 0002, Lingchen Zhao, Chao Shen 0001, Chenhao Lin, Zhengyu Zhao 0001, Qian Li 0024, Le Yang 0007, Shouling Ji, Shaofeng Li 0001, Haojin Zhu, Zhibo Wang 0001, Tianqing Zhu, Qi Li 0002, Chaoxiang He, Hongsheng Hu, Shuo Wang 0012, Shifeng Sun 0001, Hongwei Yao, Qinyu Zhang 0001, Kai Chen 0012, Yue Zhao 0027, Hongwei Li 0001, Xinyi Huang 0001, Dengguo Feng |
Sci. China Inf. Sci. | 12 |
| 2025 | DLET-Classifier: A Dynamic and Lightweight Method for Encrypted Traffic ClassificationabstractIn recent years, encrypted traffic has become a critical means of ensuring user information security. However, the widespread adoption of encrypted traffic also introduces new challenges, such as enabling attackers to conceal malicious activities within encrypted channels. Consequently, accurate encrypted traffic classification is crucial for strengthening network security defenses. However, encrypted traffic classification methods often employing complex model structures and feature extraction techniques, while neglecting efficiency and latency, which makes them difficult to apply in low-resource scenarios with slow CPU computation speed, limited memory, and a scarce number of training samples. To address these issues, we propose the Dynamic and Lightweight Encrypted Traffic Classifier (DLET-Classifier), which uses the depthwise separable convolutional neural network and the channel attention mechanism to extract features from encrypted traffic. It efficiently captures byte-level features and the relationships between packets for effective classification. To enable the model to update rapidly and adapt to the ever-changing real-world network environment, we propose the Multi2One algorithm. This algorithm first updates the base model, an ensemble of multiple binary classifiers. Then, we use the knowledge distillation technique to transfer knowledge from the base model to a lightweight model. This process allows for model updates and extensions. The results of the multi-class classification comparison experiment show that among all the compared methods, the DLET-Classifier is the model with the smallest number of parameters and the highest throughput, while also achieving excellent classification accuracy. Incremental expansion experiments demonstrate that the Multi2One algorithm enables fast knowledge updates and extensions for the lightweight model (LWG) while maintaining its classification accuracy above 96%, making our method adapt to complex network environments. Jiayong Wu, Weina Niu, Fushan Wei, Shaofeng Li 0001, Shiping Huang, Jiacheng Gong, Xiaosong Zhang 0001 |
IEEE Internet Things J. | 4 |
| 2025 | The Gradient Puppeteer: Adversarial Domination in Gradient Leakage Attacks Through Model Poisoning
Kunlan Xiang, Haomiao Yang, Meng Hao 0001, Shaofeng Li 0001, Haoxin Wang 0004, Zikang Ding, Wenbo Jiang 0001, Tianwei Zhang 0004 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | Digital Twin-Assisted Adaptive Preloading for Short Video StreamingabstractWe propose a digital twin-assisted adaptive preloading scheme to reduce bandwidth waste as well as enhance user quality of experience (QoE) for short video streaming. Though preloading video content can reduce rebuffering and improve user QoE, non-sequential playback of short videos induced by user swipe can result in substantial bandwidth wastage in mobile networks. To tackle this problem, we first model the short video streaming system and carry out preloading threshold analysis. We then construct a digital twin-assisted adaptive preloading framework for short video streaming. By collecting and analyzing the user's historical throughput and tracking swipe timing information, a throughput prediction model and a probabilistic model can be constructed to accurately predict future throughput and user swipe behavior, respectively. Utilizing the predicted information and real-time running status data from a short video application, we design a preloading strategy to enhance bandwidth efficiency while achieving high user QoE. Simulation results demonstrate the effectiveness of our proposed scheme compared with the state-of-the-art schemes. Shengbo Liu, Wen Wu 0003, Shaofeng Li 0001, Tom H. Luan, Ning Zhang 0007 |
ICC | 3 |
| 2024 | Inferring Activities and Profiles of Users Based on Trajectory Leakage in Mobile Ad NetworkabstractWith the widespread use of smartphones and the development of ad networks, mobile in-app targeted ads have become more and more prevalent, leveraging users' geolocation for targeting purposes. This service involves a large amount of user location data, which may not only expose sensitive locations closely associated with the users, but also reveal the users' activities and profiles. Previous studies have utilized various machine learning methods to infer users' activities or predict their future activities based on the location data from location-based social networks (LBSNs). These approaches, however, often require large datasets for training and are also resource-intensive. Unlike active behaviors, such as checking in, where users intentionally record their location, location data are passively recorded by mobile apps in the background, making inferring activities more challenging. Considering the rapid progress in the reasoning abilities of the large language models (LLMs) in recent years, we aim to evaluate user's activity and profile leakage through LLMs with the assistance of map APIs. We conduct the experiment on the location dataset, which is generated according to specified profiles. The results of the experiment show that the LLM can infer users' activities with an accuracy rate scoring up to 96.1 %, and there is also a high probability of predicting the users' profiles, such as the occupation. Le Yu 0002, Tian Dong 0003, Yan Meng 0001, Shaofeng Li 0001, Guoxing Chen, Haojin Zhu |
MSN | 5 |
| 2024 | Unbridled Icarus: A Survey of the Potential Perils of Image Inputs in Multimodal Large Language Model SecurityabstractMultimodal Large Language Models (MLLMs) demonstrate remarkable capabilities that increasingly influence various aspects of our daily lives, constantly defining the new boundary of Artificial General Intelligence (AGI). Image modalities, enriched with profound semantic information and a more continuous mathematical nature compared to other modalities, greatly enhance the functionalities of MLLMs when integrated. However, this integration serves as a double-edged sword, providing attackers with expansive vulnerabilities to exploit for highly covert and harmful attacks. The pursuit of reliable AI systems like powerful MLLMs has emerged as a pivotal area of contemporary research. In this paper, we endeavor to demostrate the multifaceted risks associated with the incorporation of image modalities into MLLMs. Initially, we delineate the foundational components and training processes of MLLMs. Subsequently, we construct a threat model, outlining the security vulnerabilities intrinsic to MLLMs. Moreover, we analyze and summarize existing scholarly discourses on MLLMs' attack and defense mechanisms, culminating in suggestions for the future research on MLLM security. Through this comprehensive analysis, we aim to deepen the academic understanding of MLLM security challenges and propel forward the development of trustworthy MLLM systems. Yihe Fan, Ziyao Liu, Shaofeng Li 0001 |
SMC | 5 |
| 2024 | Yes, One-Bit-Flip Matters! Universal DNN Model Inference Depletion with Runtime Code Fault Injection
Shaofeng Li 0001, Xinyu Wang 0004, Minhui Xue 0001, Haojin Zhu, Zhi Zhang 0001, Yansong Gao 0001, Wen Wu 0003, Xuemin Shen |
USENIX Security Symposium | 1 |
| 2023 | Privacy Computing with Right to Be Forgotten in Trusted Execution EnvironmentabstractSharing private data is at risk of potential data breaches, including the violation of the “right to be forgot-ten” principle, undermining people's willingness to share their data. A common solution is to involve the Trusted Execution Environment (TEE), which allows the data provider to verify the computation process without trusting others. However, previous works have either encountered incomplete computations or lacked scalability. In this paper, we propose TEERASE,a secure data-sharing framework that addresses these issues. TEERASEprotects every phase of the data lifecycle and enables individuals to share personal data with a predefined privacy budget. In particular, TEERASEapplies comprehensive privacy budgeting mechanisms to efficiently manage privacy budgets and employs an asynchronized execution approach that decouples budget consumption from data computation. TEERASErecords the predefined privacy budgets, verifies privacy consumption requests, updates the remaining budgets, and deletes data that have exhausted their budgets by preventing any attempts to access them. We implement a prototype of TEERASEand evaluate its effectiveness with a realistic case study on Genome-Wide Association Study. Hongzhi Luo, Shaofeng Li 0001, Tian Dong 0003, Guoxing Chen, Yan Meng 0001, Haojin Zhu |
GLOBECOM | 3 |
| 2023 | Data Poisoning Attack Against Anomaly Detectors in Digital Twin-Based NetworksabstractIn this paper, we study the abnormal behaviors detection and the corresponding data poisoning attacks in digital twin (DT)-based networks. We first analyze the abnormal behaviors existing in the DT-based networks, including environment anomalies, hardware and software faults, and network attacks. Specially, we design a machine learning (ML)-based anomaly detector to identify network attacks. Furthermore, due to the strong dependency of ML models on training data, in which the outputs of the trained ML models can be affected by the poisoned samples. We design a data poisoning attack scheme against the proposed ML-based anomaly detector, in which attackers can effectively compromise the output of anomaly detectors. Extensive experimental results adopting three commonly used ML-based models demonstrate that the attack can compromise these detectors with over 80% probability. Shaofeng Li 0001, Wen Wu 0003, Yan Meng 0001, Jiachun Li 0001, Haojin Zhu, Xuemin Shen |
ICC | 1 |
| 2023 | Split Federated Learning: Speed up Model Training in Resource-Limited Wireless NetworksabstractIn this paper, we propose a novel distributed learning scheme, named group-based split federated learning (GSFL), to speed up artificial intelligence (AI) model training. Specifically, the GSFL operates in a split-then-federated manner, which consists of three steps: 1) Model distribution, in which the access point (AP) splits the AI models and distributes the client-side models to clients; 2) Model training, in which each client executes forward propagation and transmit the smashed data to the edge server. The edge server executes forward and backward propagation and then returns the gradient to the clients for updating local client-side models; and 3) Model aggregation, in which edge servers aggregate the server-side and client-side models. Simulation results show that the GSFL outperforms vanilla split learning and federated learning schemes in terms of overall training latency while achieving satisfactory accuracy. Songge Zhang, Wen Wu 0003, Penghui Hu, Shaofeng Li 0001, Ning Zhang 0007 |
ICDCS | 4 |
| 2023 | RAI2: Responsible Identity Audit Governing the Artificial Intelligence
Tian Dong 0003, Shaofeng Li 0001, Guoxing Chen, Minhui Xue 0001, Haojin Zhu, Zhen Liu 0008 |
NDSS | 2 |
| 2023 | Mate! Are You Really Aware? An Explainability-Guided Testing Framework for Robustness of Malware DetectorsabstractNumerous open-source and commercial malware detectors are available. However, their efficacy is threatened by new adversarial attacks, whereby malware attempts to evade detection, e.g., by performing feature-space manipulation. In this work, we propose an explainability-guided and model-agnostic testing framework for robustness of malware detectors when confronted with adversarial attacks. The framework introduces the concept of Accrued Malicious Magnitude (AMM) to identify which malware features could be manipulated to maximize the likelihood of evading detection. We then use this framework to test several state-of-the-art malware detectors' ability to detect manipulated malware. We find that (i) commercial antivirus engines are vulnerable to AMM-guided test cases; (ii) the ability of a manipulated malware generated using one detector to evade detection by another detector (i.e., transferability) depends on the overlap of features with large AMM values between the different detectors; and (iii) AMM values effectively measure the fragility of features (i.e., capability of feature-space manipulation to flip the prediction results) and explain the robustness of malware detectors facing evasion attacks. Our findings shed light on the limitations of current malware detectors, as well as how they can be improved. Ruoxi Sun 0001, Minhui Xue 0001, Gareth Tyson, Tian Dong 0003, Shaofeng Li 0001, Shuo Wang 0012, Haojin Zhu, Seyit Ahmet Çamtepe, Surya Nepal |
ESEC/SIGSOFT FSE | 5 |
| 2022 | Fingerprinting Deep Neural Networks Globally via Universal Adversarial PerturbationsabstractIn this paper, we propose a novel and practical mechanism to enable the service provider to verify whether a suspect model is stolen from the victim model via model extraction attacks. Our key insight is that the profile of a DNN model's decision boundary can be uniquely characterized by its Universal Adversarial Perturbations (UAPs). UAPs belong to a low-dimensional subspace and piracy models' subspaces are more consistent with victim model's subspace compared with non-piracy model. Based on this, we propose a UAP fingerprinting method for DNN models and train an encoder via contrastive learning that takes fingerprints as inputs, outputs a similarity score. Extensive studies show that our framework can detect model Intellectual Property (IP) breaches with confidence > 99.99 % within only 20 fingerprints of the suspect model. It also has good generalizability across different model architectures and is robust against post-modifications on stolen models. Zirui Peng, Shaofeng Li 0001, Guoxing Chen, Cheng Zhang 0014, Haojin Zhu, Minhui Xue 0001 |
CVPR | 2 |
| 2021 | Hidden Backdoors in Human-Centric Language ModelsabstractNatural language processing (NLP) systems have been proven to be vulnerable to backdoor attacks, whereby hidden features (backdoors) are trained into a language model and may only be activated by specific inputs (called triggers), to trick the model into producing unexpected behaviors. In this paper, we create covert and natural triggers for textual backdoor attacks, hidden backdoors, where triggers can fool both modern language models and human inspection. We deploy our hidden backdoors through two state-of-the-art trigger embedding methods. The first approach via homograph replacement, embeds the trigger into deep neural networks through the visual spoofing of lookalike characters replacement. The second approach uses subtle differences between text generated by language models and real natural text to produce trigger sentences with correct grammar and high fluency. We demonstrate that the proposed hidden backdoors can be effective across three downstream security-critical NLP tasks, representative of modern human-centric NLP systems, including toxic comment detection, neural machine translation (NMT), and question answering (QA). Our two hidden backdoor attacks can achieve an Attack Success Rate (ASR) of at least 97% with an injection rate of only 3% in toxic comment detection, 95.1% ASR in NMT with less than 0.5% injected data, and finally 91.12% ASR against QA updated with only 27 poisoning data samples on a model previously trained with 92,024 samples (0.029%). We are able to demonstrate the adversary's high success rate of attacks, while maintaining functionality for regular users, with triggers inconspicuous by the human administrators. Shaofeng Li 0001, Tian Dong 0003, Benjamin Zi Hao Zhao, Minhui Xue 0001, Haojin Zhu |
CCS | 1 |
| 2021 | BatFL: Backdoor Detection on Federated Learning in e-HealthabstractFederated Learning (FL) has received significant interest both from the research field and industry perspective. One of the most promising cross-silo applications on FL is electronic health records mining which trains a model on siloed data. In this application, clients can be different hospitals or health centers that are located in geo-distributed data centers. A central orchestration server (superior health center) organizes the training, while never seeing patients’ raw data. In this paper, we demonstrate that any local hospital in such a collaborative training framework can introduce hidden backdoor functionality into the joint global model. The backdoored joint global model will produce an adversary-expected output when a predefined trigger is attached to its input but it will behave normally for clean inputs. This vulnerability is exacerbated by the distributed nature of FL, making detecting backdoor attacks on FL a challenging work. Based on the coalitional game and Shapley value, we propose an effective and real-time backdoor detection system on FL. Extensive experiments over two machine learning tasks show that our techniques achieve high accuracy and are robust against multi-attackers settings. Binhan Xi, Shaofeng Li 0001, Jiachun Li 0001, Haojin Zhu |
IWQoS | 2 |
| 2021 | Automatic Permission Optimization Framework for Privacy Enhancement of Mobile ApplicationsabstractMobile applications play a crucial role in the IoT system, which is experiencing unprecedented growth. However, users possessing little knowledge of permission configurations often accept app permission requests without reading them, which opens a backdoor for the potential adversaries to launch the future attacks. Proposing an automatic permission management scheme is an attractive solution to solve this issue, but since users have varying attitudes toward privacy, such a scheme would be neither straightforward nor user friendly. In this study, an automatic permission optimization framework, Permizer, is proposed to recommend different app permission configurations to users with different privacy preferences. Permizer estimates the permission risks and builds the permission-functionality mapping to each app, then regulates the relationship between permission and app functionality. Permizer is the first module to achieve a balance between privacy protection and app functionality under the personal privacy preference condition. Finally, we develop Permizer as a one-button service on the real-world Android OS with 58 apps. Case studies conducted on TikTok and Amazon Alexa also demonstrate its practicability and effectiveness. Yiting Qu, Suguo Du, Shaofeng Li 0001, Yan Meng 0001, Haojin Zhu |
IEEE Internet Things J. | 3 |
| 2021 | Invisible Backdoor Attacks on Deep Neural Networks Via Steganography and RegularizationabstractDeep neural networks (DNNs) have been proven vulnerable to backdoor attacks, where hidden features (patterns) trained to a normal model, which is only activated by some specific input (called triggers), trick the model into producing unexpected behavior. In this article, we create covert and scattered triggers for backdoor attacks, invisible backdoors, where triggers can fool both DNN models and human inspection. We apply our invisible backdoors through two state-of-the-art methods of embedding triggers for backdoor attacks. The first approach on Badnets embeds the trigger into DNNs through steganography. The second approach of a trojan attack uses two types of additional regularization terms to generate the triggers with irregular shape and size. We use the Attack Success Rate and Functionality to measure the performance of our attacks. We introduce two novel definitions of invisibility for human perception; one is conceptualized by the Perceptual Adversarial Similarity Score (PASS) and the other is Learned Perceptual Image Patch Similarity (LPIPS). We show that the proposed invisible backdoors can be fairly effective across various DNN models as well as four datasets MNIST, CIFAR-10, CIFAR-100, and GTSRB, by measuring their attack success rates for the adversary, functionality for the normal users, and invisibility scores for the administrators. We finally argue that the proposed invisible backdoor attacks can effectively thwart the state-of-the-art trojan backdoor detection approaches. Shaofeng Li 0001, Minhui Xue 0001, Benjamin Zi Hao Zhao, Haojin Zhu, Xinpeng Zhang 0001 |
IEEE Trans. Dependable Secur. Comput. | 1 |