EDBT 2026 Demo / reviewers in the wild / expert
Yifeng Cai
dblp:122/4894
· DBLP profile ↗
15ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0002-0049-6670ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Computer networks · 3 · 2 first-authorSoftware engineering, systems software and programming languages · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VulSCA: A Community-Level SCA Approach for Accurate C/C++ Supply Chain Vulnerability Analysis
Yueming Wu 0001, Yifeng Cai, Deqing Zou |
NDSS | 4 |
| 2025 | Membership and Memorization in LLM Knowledge DistillationabstractRecent advances in Knowledge Distillation (KD) aim to mitigate the high computational demands of Large Language Models (LLMs) by transferring knowledge from a large "teacher" to a smaller "student" model.However, students may inherit the teacher's privacy when the teacher is trained on private data.In this work, we systematically characterize and investigate membership and memorization privacy risks inherent in six LLM KD techniques.Using instruction-tuning settings that span seven NLP tasks, together with three teacher model families (GPT-2, LLAMA-2, and OPT), and various size student models, we demonstrate that all existing LLM KD approaches carry membership and memorization privacy risks from the teacher to its students.However, the extent of privacy risks varies across different KD techniques.We systematically analyse how key LLM KD components (KD objective functions, student training data and NLP tasks) impact such privacy risks.We also demonstrate a significant disagreement between memorization and membership privacy risks of LLM KD techniques.Finally, we characterize per-block privacy risk and demonstrate that the privacy risk varies across different blocks by a large margin. Ziqi Zhang 0017, Ali Shahin Shamsabadi, Hanxiao Lu, Yifeng Cai, Hamed Haddadi 0001 |
EMNLP | 4 |
| 2025 | I Can Tell Your Secrets: Inferring Privacy Attributes from Mini-app Interaction History in Super-apps
Yifeng Cai, Mengyu Yao, Xiaoke Zhao, Zhe Liu 0001, Xiangqun Chen, Yao Guo 0001, Ding Li 0001 |
USENIX Security Symposium | 1 |
| 2025 | Game of Arrows: On the (In-)Security of Weight Obfuscation for On-Device TEE-Shielded LLM Partition Algorithms
Pengli Wang, Bingyou Dong, Yifeng Cai, Huanran Xue |
USENIX Security Symposium | 3 |
| 2025 | TEESlice: Protecting Sensitive Neural Network Models in Trusted Execution Environments when Attackers Have Pre-Trained ModelsabstractTrusted Execution Environments (TEEs) are used to safeguard on-device models. However, directly employing TEEs to secure the entire DNN model is challenging due to the limited computational speed. Utilizing GPU can accelerate DNN’s computation speed but widely available commercial GPUs usually lack security protection. To this end, scholars introduce TEE-Shielded DNN Partition (TSDP), a method that protects privacy-sensitive weights within TEEs and offloads insensitive weights to GPUs. Nevertheless, current methods do not consider the presence of a knowledgeable adversary who can access abundant publicly available pre-trained models and datasets. This article investigates the security of the existing methods against such a knowledgeable adversary and reveals their inability to fulfill their security promises. Consequently, we introduce a novel partition before training strategy, which effectively separates privacy-sensitive weights from other components of the model. Our evaluation demonstrates that our approach can offer full model protection with a computational cost reduced by a factor of 10. In addition to traditional CNN models, we also demonstrate the scalability to large language models. Our approach can compress the private functionalities of the large language model to lightweight slices and achieve the same level of protection as the shielding-whole-model baseline. Ding Li 0001, Ziqi Zhang 0017, Mengyu Yao, Yifeng Cai, Yao Guo 0001, Xiangqun Chen |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2024 | No Privacy Left Outside: On the (In-)Security of TEE-Shielded DNN Partition for On-Device MLabstractOn-device ML introduces new security challenges: DNN models become white-box accessible to device users. Based on white-box information, adversaries can conduct effective model stealing (MS) against model weights and membership inference attack (MIA) against training data privacy. Using Trusted Execution Environments (TEEs) to shield on-device DNN models aims to downgrade (easy) white-box attacks to (harder) black-box attacks. However, one major shortcoming of TEEs is the sharply increased latency (up to 50×). To accelerate TEE-shield DNN computation with GPUs, researchers proposed several model partition techniques. These solutions, referred to as TEE-Shielded DNN Partition (TSDP), partition a DNN model into two parts, offloading1the privacy-insensitive part to the GPU while shielding the privacy-sensitive part within the TEE. However, the community lacks an in-depth understanding of the seemingly encouraging privacy guarantees offered by existing TSDP solutions during DNN inference. This paper benchmarks existing TSDP solutions using both MS and MIA across a variety of DNN models, datasets, and metrics. We show important findings that existing TSDP solutions are vulnerable to privacy-stealing attacks and are not as safe as commonly believed. We also unveil the inherent difficulty in deciding the optimal DNN partition configurations, which vary across datasets and models. Based on lessons harvested from the experiments, we present TEESlice, a novel TSDP method that defends against MS and MIA during DNN inference. Unlike existing approaches, TEESlice follows a partition-before-training strategy, which allows for accurate separation between privacy-related weights from public weights. TEESlice delivers the same security protection as shielding the entire DNN model inside TEE (the "upper-bound" security guarantees) with over 10×less overhead (in both experimental and real-world environments) than prior TSDP solutions and no accuracy loss. We make the code and artifacts publicly available on the Internet. Ziqi Zhang 0017, Yifeng Cai, Yuanyuan Yuan 0001, Ding Li 0001, Yao Guo 0001, Xiangqun Chen |
SP | 3 |
| 2024 | FAMOS: Robust Privacy-Preserving Authentication on Payment Apps via Federated Multi-Modal Contrastive Learning
Yifeng Cai, Jiaping Gui, Xiaoke Zhao, Ding Li 0001 |
USENIX Security Symposium | 1 |
| 2023 | FedSlice: Protecting Federated Learning Models from Malicious Participants with Model SlicingabstractCrowdsourcing Federated learning (CFL) is a new crowdsourcing development paradigm for the Deep Neural Network (DNN) models, also called “software 2.0”. In practice, the privacy of CFL can be compromised by many attacks, such as free-rider attacks, adversarial attacks, gradient leakage attacks, and inference attacks. Conventional defensive techniques have low efficiency because they deploy heavy encryption techniques or rely on Trusted Execution Environments (TEEs). To improve the efficiency of protecting CFL from these attacks, this paper proposes FedSlice to prevent malicious participants from getting the whole server-side model while keeping the performance goal of CFL. FedSlice breaks the server-side model into several slices and delivers one slice to each participant. Thus, a malicious participant can only get a subset of the server-side model, preventing them from effectively conducting effective attacks. We evaluate FedSlice against these attacks, and results show that FedSlice provides effective defense: the server-side model leakage is reduced from 100% to 43.45%, the success rate of adversarial attacks is reduced from 100% to 11.66%, the average accuracy of membership inference is reduced from 71.91% to 51.58%, and the data leakage from shared gradients is reduced to the level of random guesses. Besides, FedSlice only introduces less than 2% accuracy loss and about 14% computation overhead. To the best of our knowledge, this is the first paper to discuss defense methods against these attacks to the CFL framework. Ziqi Zhang 0017, Yuanchun Li 0003, Yifeng Cai, Ding Li 0001, Yao Guo 0001, Xiangqun Chen |
ICSE | 4 |
| 2023 | Beyond Fine-Tuning: Efficient and Effective Fed-Tuning for Mobile/Web UsersabstractFine-tuning is a typical mechanism to achieve model adaptation for mobile/web users, where a model trained by the cloud is further retrained to fit the target user task. While traditional fine-tuning has been proved effective, it only utilizes local data to achieve adaptation, failing to take advantage of the valuable knowledge from other mobile/web users. In this paper, we attempt to extend the local-user fine-tuning to multi-user fed-tuning with the help of Federated Learning (FL). Following the new paradigm, we propose EEFT, a framework aiming to achieve Efficient and Effective Fed-Tuning for mobile/web users. The key idea is to introduce lightweight but effective adaptation modules to the pre-trained model, such that we can freeze the pre-trained model and just focus on optimizing the modules to achieve cost reduction and selective task cooperation. Extensive experiments on our constructed benchmark demonstrate the effectiveness and efficiency of the proposed framework. Yifeng Cai, Hongzhe Bi, Ziqi Zhang 0015, Ding Li 0001, Yao Guo 0001, Xiangqun Chen |
WWW | 2 |
| 2021 | TransTailor: Pruning the Pre-trained Model for Improved Transfer LearningabstractThe increasing of pre-trained models has significantly facilitated the performance on limited data tasks with transfer learning. However, progress on transfer learning mainly focuses on optimizing the weights of pre-trained models, which ignores the structure mismatch between the model and the target task. This paper aims to improve the transfer performance from another angle - in addition to tuning the weights, we tune the structure of pre-trained models, in order to better match the target task. To this end, we propose TransTailor, targeting at pruning the pre-trained model for improved transfer learning. Different from traditional pruning pipelines, we prune and fine-tune the pre-trained model according to the target-aware weight importance, generating an optimal sub-model tailored for a specific target task. In this way, we transfer a more suitable sub-structure that can be applied during fine-tuning to benefit the final performance. Extensive experiments on multiple pre-trained models and datasets demonstrate that TransTailor outperforms the traditional pruning methods and achieves competitive or even better performance than other state-of-the-art transfer learning methods while using a smaller model. Notably, on the Stanford Dogs dataset, TransTailor can achieve 2.7% accuracy improvement over other transfer methods with 20% fewer FLOPs. Yifeng Cai, Yao Guo 0001, Xiangqun Chen |
AAAI | 2 |
| 2016 | Subgraph matching route navigation by UAV and ground robot cooperationabstractIn this paper, we propose a novel and robust cognitive sharing method based on edge-label subgraph matching for ROI sharing and object recognition. This method enhances a vision based cooperation for UAV and ground robot system under an unknown environment where GPS is not available. The proposed subgraph matching method is an extension of Nema in that only edge-label graph is assumed and hence it can be applied to the situation where the node-labels are hard to distinguish. This method is applied to the navigation system where a UAV navigates a ground robot to the target position via ROI sharing. We test the method through a series of simulations and confirm a stable and good performance in different situations. Yifeng Cai, Kousuke Sekiyama |
CEC | 1 |
| 2015 | System design of a time-controlled broadband piezoelectric energy harvesting interface circuitabstractThis paper presents a novel system design of an adaptive time-controlled energy harvesting interface circuit for broadening the bandwidth of piezoelectric transducers. The proposed system monitors the excitation frequency and delays the switching activity accordingly to achieve optimal impedance matching. The system and the theory for calculating the time delay values are described in detail. Simulation result shows that the proposed system is able to generate nearly as much power as the theoretical limitation over a large range of excitation frequency, and the bandwidth increases from 3.7 Hz to more than 70 Hz. Yifeng Cai, Yiannos Manoli |
ISCAS | 1 |
| 2013 | Improving WLAN throughput via reactive jamming in the presence of hidden terminalsabstractIn the area of performance analysis of wireless networks, one critical issue is the hidden terminal problem, which is considered as one of the severest reasons for the degradation of network performance. In this paper, we incorporate reactive jamming scheme with distributed coordination function (DCF) in IEEE 802.11 based wireless local area networks (WLANs) to improve network throughput in the presence of hidden terminals. In the proposed protocol, we schedule access point (AP) to broadcast jamming signal reactively to hinder the simultaneous transmission of hidden terminals. Both analytical and numerical results show that our reactive jamming based protocol can constantly improve WLAN throughput for a wide range of conditions, compared with the traditional RTS/CTS. Yifeng Cai, Kunjie Xu, Yijun Mo, Bang Wang 0001, Mu Zhou |
WCNC | 1 |
| 2013 | Joint reactive jammer detection and localization in an enterprise WiFi network
Yifeng Cai, Konstantinos Pelechrinis, Prashant Krishnamurthy, Yijun Mo |
Comput. Networks | 1 |
| 2012 | A novel indoor localization method based on virtual AP estimationabstractIndoor localization is important to many location based applications and services. Many indoor localization methods have been proposed and they can be roughly categorized into two groups: One is based on the distance estimation between a target point and the Access Points (APs); and the other is based on the Received Signal Strength (RSS) fingerprint map. However, the two approaches all assume that the locations of Access Points (APs) are known beforehand. In this paper, we consider a scenario that the locations of APs are not known, and we use measured RSS (Received Signal Strength) at some location-known reference points and their geographical information to estimate virtual APs' parameters. Then those parameters are used to locate new user by optimization method. Experiments show that the position accuracy of the proposed method approaches that of RADAR with less effort. Yijun Mo, Yifeng Cai, Bang Wang 0001 |
ICC | 2 |