VLDB 2026 Research / reviewers in the wild / expert
Yuke Hu
dblp:330/4599
· DBLP profile ↗
11ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0001-5780-6898ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 7 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Shadow in the Cache: Unveiling and Mitigating Privacy Risks of KV-cache in LLM Inference
Zhifan Luo, Shuo Shao 0002, Lijing Zhou, Yuke Hu, Zhan Qin |
NDSS | 5 |
| 2026 | LFS: A Locally Private Framework for Degree Statistic Estimation With Laplace MechanismabstractAs a fundamental task in graph data analysis, degree statistic estimation serves as the foundation for many complex tasks. Local differential privacy (LDP) preserves the privacy inherent in raw degrees without a trusted third party. Existing methods struggle to balance different types of degree statistics. They either introduce excessive noise when estimating degree distribution due to not fully leveraging the properties of edge LDP, or are limited to polynomial statistic estimation only. We design a locally private framework for degree statistic estimation (LFS), using Laplace mechanism to provide appropriate privacy protection under edge LDP. LFS can estimate three types of degree statistics: polynomial, distribution and single-point. According to degrees with Laplace noise, we transform degree distribution estimation into a linear regression problem, then post-process the estimated distribution to mitigate the excessive smoothing introduced by the regularization term. We also achieve single-point statistic estimation considering the degree distribution and properties of Laplace noise. Systematic experiments on five datasets demonstrate that LFS consistently outperforms existing methods in four utility metrics. Yuke Hu, Shiqi Zhou, Fenghua Li 0001, Ben Niu 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2026 | Differentially Private Zeroth-Order Methods for Scalable Large Language Model Fine-TuningabstractFine-tuning large language models (LLMs) on down-stream tasks has become a standard approach to adapt their capabilities. However, the process raises privacy concerns when using sensitive datasets, prompting increasing interest in differentially private (DP) fine-tuning methods. While existing approaches build upon the seminal work of DP-SGD, they are constrained by the inherent inefficiency bottlenecks. In this paper, we investigate the potential of DP zeroth-order methods for LLM fine-tuning, which avoids the scalability bottleneck of SGD by approximating gradients with more efficient zeroth-order gradients. We propose the stagewise DP zeroth-order method (DP-ZOSO) that dynamically schedules key hyperparameters to leverage the synergy between DP random perturbation and the gradient approximation error. To further enhance the scalability, we propose DP zeroth-order stagewise pruning method (DP-ZOPO) which reduces the trainable parameters by a data-free pruning technique requiring no extra privacy budget. We provide theoretical analysis for both proposed methods and conduct extensive empirical analysis on both encoder-only masked and decoder-only autoregressive language model, achieving impressive results in terms of scalability and utility across diverse tasks. Jian Lou 0001, Wenjie Bao, Yuke Hu, Zhan Qin, Kui Ren 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | Membership Inference Attacks Against Vision-Language Models
Yuke Hu, Zheng Li 0023, Yang Zhang 0016, Zhan Qin, Kui Ren 0001, Chun Chen 0001 |
USENIX Security Symposium | 1 |
| 2025 | Privacy Risks of Federated Knowledge Graph Embedding: New Membership Inference Attacks and Personalized Differential Privacy DefenseabstractKnowledge Graph Embedding (KGE) has been widely studied as an important semantic enhancement technique that extracts expressive representation from Knowledge Graph (KG) to facilitate various downstream applications, such as knowledge reasoning, Semantic Web, and question answering. The advent of Federated KGE (FKGE) allows for collaborative training across distributed KGs without revealing clients’ private raw KGs that contain sensitive knowledge graph triples. Despite this, FKGE remains susceptible to privacy threats as demonstrated in previously studied federated learning models. However, utilizing and addressing these vulnerabilities remain uninvestigated for FKGE which exhibits unique characteristics distinct from other models. In this work, we conduct the first comprehensive study of the privacy issues in FKGE from both attack and defense perspectives. On the attack side, we introduce five new inference attacks, highlighting the privacy vulnerabilities by successfully deducing the presence of KG triples from the targeted dataset. On the defense side, we present PDP-Flames, a novel differentially private FKGE scheme that leverages the sparse gradient nature of FKGE for better privacy-utility trade-off by integrating advanced private selection techniques. We further introduce a dynamic defense policy based on the observation that the privacy risk fluctuates throughout the training procedure. Additionally, we incorporate a personalized procedure to provide a customized model tailored to the unique data distributions of individual clients. Joint differential privacy is introduced to guarantee the privacy of the personalized models. Comprehensive experiments demonstrate that PDP-Flames effectively mitigates privacy concerns, notably diminishing the attack success rate while maintaining decent model utility. Yuke Hu, Jian Lou 0001, Weiqiang Wang 0002, Jinfei Liu, Zhan Qin |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2024 | ERASER: Machine Unlearning in MLaaS via an Inference Serving-Aware ApproachabstractOver the past years, Machine Learning-as-a-Service (MLaaS) has received a surging demand for supporting Machine Learning-driven services to offer revolutionized user experience across diverse application areas. MLaaS provides inference service with low inference latency based on an ML model trained using a dataset collected from numerous individual data owners. Recently, for the sake of data owners' privacy and to comply with the "right to be forgotten (RTBF)" as enacted by data protection legislation, many machine unlearning methods have been proposed to remove data owners' data from trained models upon their unlearning requests. However, despite their promising efficiency, almost all existing machine unlearning methods handle unlearning requests independently from inference requests, which unfortunately introduces a new security issue of inference service obsolescence and a privacy vulnerability of undesirable exposure for machine unlearning in MLaaS. Yuke Hu, Jian Lou 0001, Jiaqi Liu 0003, Wangze Ni, Feng Lin 0004, Zhan Qin, Kui Ren 0001 |
CCS | 1 |
| 2024 | SWAT: A System-Wide Approach to Tunable Leakage Mitigation in Encrypted Data StoresabstractNumerous studies have underscored the significant privacy risks associated with various leakage patterns in encrypted data stores. While many solutions have been proposed to mitigate these leakages, they either (1) incur substantial overheads, (2) focus on specific subsets of leakage patterns, or (3) apply the same security notion across various workloads, thereby impeding the attainment of fine-tuned privacy-efficiency trade-offs. In light of various detrimental leakage patterns, this paper starts with an investigation into which specific leakage patterns require our focus in the contexts of key-value, range-query, and dynamic workloads, respectively. Subsequently, we introduce new security notions tailored to the specific privacy requirements of these workloads. Accordingly, we propose and instantiate Swat, an efficient construction that progressively enables these workloads, while provably mitigating system-wide leakage via a suite of algorithms with tunable privacy-efficiency trade-offs. We conducted extensive experiments and compiled a detailed result analysis, showing the efficiency of our solution. Swat is about an order of magnitude slower than an encryption-only data store that reveals various leakage patterns and is two orders of magnitude faster than a trivial zero-leakage solution. Meanwhile, the performance of Swat remains highly competitive compared to other designs that mitigate specific types of leakage. Leqian Zheng, Lei Xu 0019, Cong Wang 0001, Sheng Wang 0011, Yuke Hu, Zhan Qin, Feifei Li 0001, Kui Ren 0001 |
Proc. VLDB Endow. | 5 |
| 2024 | Privacy Enhancement Via Dummy Points in the Shuffle ModelabstractThe shuffle model is recently proposed to address the issue of severe utility loss in Local Differential Privacy (LDP) due to distributed data randomization. In the shuffle model, a shuffler is utilized to break the link between the user identity and the message uploaded to the data analyst. Since less noise needs to be introduced to achieve the same privacy guarantee, following this paradigm, the utility of privacy-preserving data collection is improved. We propose DUMP (DUMmy-Point-based), a framework for privacy-preserving histogram estimation in the shuffle model. The core of DUMP is a new concept ofdummy blanket, which enables enhancing privacy by just introducing dummy points on the user side and further improving the utility of the shuffle model. We instantiate DUMP by proposing two protocols: pureDUMP and mixDUMP, and conduct a comprehensive experimental evaluation to compare them with existing protocols. The experimental results show that, under the same privacy guarantee, (1) the proposed protocols have significant improvements in communication efficiency over all existing multi-message protocols, by at least 3 orders of magnitude; (2) they achieve competitive utility, while the only known protocol (Ghaziet al., PMLR 2020) having better utility than ours employs hard-to-exactly-sample distributions which are vulnerable to floating-point attacks (CCS 2012). Hanwen Feng 0001, Kunzhe Huang, Yuke Hu, Jinfei Liu, Kui Ren 0001, Zhan Qin |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2024 | Location Privacy-Aware Task Offloading in Mobile Edge ComputingabstractIn mobile edge computing (MEC), users can offload tasks to nearby MEC servers to reduce computation cost. Considering that the size of offloaded tasks could disclose user location information, several location privacy-preserving task offloading mechanisms have been proposed under the single-server scenario. However, to the best of our knowledge, none of them could provide a strict privacy protection guarantee or be applicable to the multi-server scenario where the user's location can be inferred more accurately if servers collude with each other. In this paper, we propose a novel location privacy-aware task offloading framework (LPA-Offload) for both single-server and multi-server scenarios, which provides strict and provable location privacy protection while achieving efficient task offloading. Specifically, we propose a location perturbation mechanism that allows each user to perturb its real location within a rational perturbation region and provides a differential privacy guarantee. To make a satisfactory offloading strategy, we propose a perturbation region determination mechanism and an offloading strategy generation mechanism that adaptively select a proper perturbation region according to the customized privacy factor, and then generate an optimal offloading strategy based on the perturbed location within the decided region. The determination of the perturbation region could achieve personalized privacy requirements while reducing computation cost. LPA-Offload is proved to satisfy$(\epsilon,\delta)$-differential privacy, and the experiments demonstrate the effectiveness of our framework. Zhibo Wang 0001, Yunan Sun, Defang Liu, Jiahui Hu 0001, Xiaoyi Pang, Yuke Hu, Kui Ren 0001 |
IEEE Trans. Mob. Comput. | 6 |
| 2023 | Quantifying and Defending against Privacy Threats on Federated Knowledge Graph EmbeddingabstractKnowledge Graph Embedding (KGE) is a fundamental technique that extracts expressive representation from knowledge graph (KG) to facilitate diverse downstream tasks. The emerging federated KGE (FKGE) collaboratively trains from distributed KGs held among clients while avoiding exchanging clients’ sensitive raw KGs, which can still suffer from privacy threats as evidenced in other federated model trainings (e.g., neural networks). However, quantifying and defending against such privacy threats remain unexplored for FKGE which possesses unique properties not shared by previously studied models. In this paper, we conduct the first holistic study of the privacy threat on FKGE from both attack and defense perspectives. For the attack, we quantify the privacy threat by proposing three new inference attacks, which reveal substantial privacy risk by successfully inferring the existence of the KG triple from victim clients. For the defense, we propose DP-Flames, a novel differentially private FKGE with private selection, which offers a better privacy-utility tradeoff by exploiting the entity-binding sparse gradient property of FKGE and comes with a tight privacy accountant by incorporating the state-of-the-art private selection technique. We further propose an adaptive privacy budget allocation policy to dynamically adjust defense magnitude across the training procedure. Comprehensive evaluations demonstrate that the proposed defense can successfully mitigate the privacy threat by effectively reducing the success rate of inference attacks from to on average with only a modest utility decrease. Yuke Hu, Weiqiang Wang 0002, Jinfei Liu, Zhan Qin |
WWW | 1 |
| 2022 | OpBoost: A Vertical Federated Tree Boosting Framework Based on Order-Preserving DesensitizationabstractVertical Federated Learning (FL) is a new paradigm that enables users with non-overlapping attributes of the same data samples to jointly train a model without directly sharing the raw data. Nevertheless, recent works show that it's still not sufficient to prevent privacy leakage from the training process or the trained model. This paper focuses on studying the privacy-preserving tree boosting algorithms under the vertical FL. The existing solutions based on cryptography involve heavy computation and communication overhead and are vulnerable to inference attacks. Although the solution based on Local Differential Privacy (LDP) addresses the above problems, it leads to the low accuracy of the trained model. This paper explores to improve the accuracy of the widely deployed tree boosting algorithms satisfying differential privacy under vertical FL. Specifically, we introduce a framework called OpBoost. Three order-preserving desensitization algorithms satisfying a variant of LDP called distance-based LDP (dLDP) are designed to desensitize the training data. In particular, we optimize the dLDP definition and study efficient sampling distributions to further improve the accuracy and efficiency of the proposed algorithms. The proposed algorithms provide a trade-off between the privacy of pairs with large distance and the utility of desensitized values. Comprehensive evaluations show that OpBoost has a better performance on prediction accuracy of trained models compared with existing LDP approaches on reasonable settings. Our code is open source. Yuke Hu, Hanwen Feng 0001, Yuan Hong 0001, Kui Ren 0001, Zhan Qin |
Proc. VLDB Endow. | 2 |