VLDB 2026 Research / reviewers in the wild / expert
Xiaoyi Pang
dblp:240/8960
· DBLP profile ↗
22ranked-venue papers
6as first author
20since 2021 · last 2026
0000-0002-2763-2695ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 14 · 2 first-author · 12 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Security and privacy · 3 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Two-Dimensional Stackelberg Game-Based Incentive Mechanism for Differential Private Federated Learning With Non-IID DataabstractIncentive mechanisms are essential for boosting client engagement in differential private federated learning (DP-FL). However, existing Stakelberg games-based incentive mechanisms typically assume that client decisions are one-dimension and that data is independent and identically distributed (IID) across clients. In reality, data distributions are often non-IID and clients have two-dimensional resources decisions, including data quantity and privacy. Therefore, in this paper, we present a novel two-dimensional Stackelberg game-based incentive mechanism for DP-FL with non-IID data, aiming to maximize the total utility of clients and server by seeking a balance between the clients' two-dimensional decisions and the server's payment. Specifically, we first formulate the utility functions of both server and clients under two-dimensional decisions and then model the interactions between server and clients as a single-leader-multiple-followers Stackelberg game. To derive the optimal decisions that maximize their utilities, we theoretically prove the existence of a Stackelberg equilibrium between server and clients. Due to the difficulty to directly calculate the Stackelberg equilibrium, we propose a bi-level multi-agent reinforment learning algorithm to learn the optimal decisions for both server and clients by trial and error. Extensive simulation results demonstrate that our proposed method outperforms the baselines in terms of total utility. Dan Wang 0031, Xiaoyi Pang, Jiahui Hu 0001, Sheng Yue 0001, Ju Ren 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2025 | Is Your Image a Good Storyteller?abstractQuantifying image complexity at the entity level is straightforward, but the assessment of semantic complexity has been largely overlooked. In fact, there are differences in semantic complexity across images. Images with richer semantics can tell vivid and engaging stories and offer a wide range of application scenarios. For example, the Cookie Theft picture is such a kind of image and is widely used to assess human language and cognitive abilities due to its higher semantic complexity. Additionally, semantically rich images can benefit the development of vision models, as images with limited semantics are becoming less challenging for them. However, such images are scarce, highlighting the need for a greater number of them. For instance, there is a need for more images like Cookie Theft to cater to people from different cultural backgrounds and eras. Assessing semantic complexity requires human experts and empirical evidence. Automatic evaluation of how semantically rich an image will be the first step of mining or generating more images with rich semantics, and benefit human cognitive assessment, Artificial Intelligence, and various other applications. In response, we propose the Image Semantic Assessment (ISA) task to address this problem. We introduce the first ISA dataset and a novel method that leverages language to solve this vision problem. Experiments on our dataset demonstrate the effectiveness of our approach. Xiujie Song, Xiaoyi Pang, Haifeng Tang, Mengyue Wu, Kenny Q. Zhu |
AAAI | 2 |
| 2025 | Textual Unlearning Gives a False Sense of UnlearningabstractLanguage Models (LMs) are prone to ''memorizing'' training data, including substantial sensitive user information. To mitigate privacy risks and safeguard the right to be forgotten, machine unlearning has emerged as a promising approach for enabling LMs to efficiently ''forget'' specific texts. However, despite the good intentions, is textual unlearning really as effective and reliable as expected? To address the concern, we first propose Unlearning Likelihood Ratio Attack+ (U-LiRA+), a rigorous textual unlearning auditing method, and find that unlearned texts can still be detected with very high confidence after unlearning. Further, we conduct an in-depth investigation on the privacy risks of textual unlearning mechanisms in deployment and present the Textual Unlearning Leakage Attack (TULA), along with its variants in both black- and white-box scenarios. We show that textual unlearning mechanisms could instead reveal more about the unlearned texts, exposing them to significant membership inference and data reconstruction risks. Our findings highlight that existing textual unlearning actually gives a false sense of unlearning, underscoring the need for more robust and secure unlearning mechanisms. Jiacheng Du, Zhibo Wang 0001, Jie Zhang 0081, Xiaoyi Pang, Jiahui Hu 0001, Kui Ren 0001 |
ICML | 4 |
| 2025 | ICLScan: Detecting Backdoors in Black-Box Large Language Models via Targeted In-context IlluminationabstractThe widespread deployment of large language models (LLMs) allows users to access their capabilities via black-box APIs, but backdoor attacks pose serious security risks for API users by hijacking the model behavior. This highlights the importance of backdoor detection technologies to help users audit LLMs before use. However, most existing LLM backdoor defenses require white-box access or costly reverse engineering, limiting their practicality for resource-constrained users. Moreover, they mainly target classification tasks, leaving broader generative scenarios underexplored. To solve the problem, this paper introduces ICLScan, a lightweight framework that exploits targeted in-context learning (ICL) as illumination for backdoor detection in black-box LLMs, which effectively supports generative tasks without additional training or model modifications. ICLScan is based on our finding of backdoor susceptibility amplification: LLMs with pre-embedded backdoors are highly susceptible to new trigger implantation via ICL. Including only a small ratio of backdoor examples (containing ICL-triggered input and target output) in the ICL prompt can induce ICL trigger-specific malicious behavior in backdoored LLMs. ICLScan leverages this phenomenon to detect backdoored LLMs by statistically analyzing whether the success rate of new trigger injection via targeted ICL exceeds a threshold. It requires only multiple queries to estimate the backdoor success rate, overcoming black-box access and computational resource limitations. Extensive experiments across diverse LLMs and backdoor attacks demonstrate ICLScan's effectiveness and efficiency, achieving near-perfect detection performance (precision/recall/F1-score/ROC-AUC all approaching 1) with minimal additional overhead across all settings. Xiaoyi Pang, Xuanyi Hao, Song Guo 0001, Zhibo Wang 0001 |
NeurIPS | 1 |
| 2025 | PoiSAFL: Scalable Poisoning Attack Framework to Byzantine-resilient Semi-asynchronous Federated Learning
Xiaoyi Pang, Zhibo Wang 0001, Jiahui Hu 0001, Yinggui Wang, Lei Wang 0251, Tao Wei 0002, Kui Ren 0001, Chun Chen 0001 |
USENIX Security Symposium | 1 |
| 2025 | U-DPAP: Utility-aware Efficient Range Counting on Privacy-preserving Spatial Data FederationabstractRange counting is a fundamental operation in spatial data applications. There is a growing demand to facilitate this operation over a data federation, where spatial data are separately held by multiple data providers (a.k.a., data silos). Most existing data federation schemes employ Secure Multiparty Computation (SMC) to protect privacy, but this approach is computationally expensive and leads to high latency. Consequently, private data federations are often impractical for typical database workloads.This challenge highlights the need for a private data federation scheme capable of providing fast and accurate query responses while maintaining strong privacy. To address this issue, we propose U-DPAP, a utility-aware efficient privacy-preserving method. It is the first scheme to exclusively use differential privacy for privacy protection in spatial data federation, without employing SMC. Moreover, it combines approximate query processing to further enhance efficiency. Our experimental results indicate that a straightforward combination of the two techniques results in unacceptable impacts on data utility. Thus, we design two novel algorithms: one to make differential privacy practical by optimizing the privacy-utility trade-off, and another to address the efficiency-utility trade-off in approximate query processing. The grouping-based perturbation algorithm reduces noise by grouping similar data and applying noise to the groups. The representative data silos selection algorithm minimizes approximate error by selecting representative silos using the similarity between data silos. We rigorously prove the privacy guarantees of U-DPAP. Moreover, experimental results demonstrate that U-DPAP enhances data utility by an order of magnitude while maintaining high communication efficiency. Yahong Chen, Xiaoyi Pang, Ben Niu 0001, Shengnan Hu |
Proc. ACM Manag. Data | 2 |
| 2025 | Poisoning Attacks to Knowledge Distillation-Based Federated Learning Under Robust Aggregation RulesabstractFederated learning (FL) is susceptible to poisoning attacks. To defend against such threats, robust aggregation rules (AGRs) are typically deployed on the server to identify or filter clients’ potentially malicious submissions based on statistical similarity. Recently, knowledge distillation (KD) has been widely used in FL to facilitate collaborative learning among clients that have heterogeneous model architectures by aggregating and distilling architecture-independent model outputs (i.e., logits). However, the KD process introduces a novel poisoning attack surface, where adversaries can manipulate local model output logits to ruin the global model performance. To fully reveal and explore such a new security vulnerability and effectively poison the global model in the existence of robust AGRs, in this paper, we propose the first untargeted poisoning attack scheme to KD-based FL under robust AGRs, named ManipulatingKD. It manipulates compromised clients to send well-designed malicious logits during the KD process. To ensure attack effectiveness and stealthiness, ManipulatingKD models attacks as constrained optimization problems. This allows for crafting satisfactory malicious logits that are statistically similar to benign logits but can generate poisoned aggregated logits to provide deviated supervision and mislead the global model. Extensive experiments demonstrate the effectiveness of ManipulatingKD under both non-robust and robust AGRs. Particularly, under robust AGRs, the global model accuracy degradation caused by our attacks can exceed 2× that of state-of-the-art attacks. Xiaoyi Pang, Zhibo Wang 0001, Defang Liu, Jiahui Hu 0001, Peng Sun 0003, Meng Luo 0002, Kui Ren 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2025 | An Incentive Framework for Task Offloading in Edge Computing Marketplaces Under Price CompetitionabstractTo efficiently execute tasks, computation resource requesters (CRRs) with limited resources can offload their tasks to nearby computation resource providers (CRPs) with spare computing capacity. These CRPs require appropriate incentives to compensate for their incurred costs when helping process the offloaded tasks. Although several mechanisms have been designed to incentivize CRPs, none of them have investigated the incentive mechanism considering price-setting and price-taking CRPs simultaneously. In this work, we propose an incentive framework for task offloading in the edge computing marketplace that includes both price-setting and price-taking CRPs. We model the CRR's interactions with both types of CRPs as a three-stage Stackelberg game to maximize the profit for both the CRR and CRPs. We prove the existence of a unique subgame perfect equilibrium (SPE) of the formulated game and further develop iterative algorithms for the CRR and price-setting CRPs to achieve the equilibrium. Through the designed algorithms, each CRP does not require complete information about the CRR and other CRPs. Extensive simulations demonstrate that offloading tasks to both price-setting and price-taking CRPs achieves higher profits for the CRR and price-setting CRPs compared to offloading tasks solely to price-setting CRPs. Additionally, the obtained SPE can achieve near-optimal social welfare. Liantao Wu, Peng Sun 0003, Zhibo Wang 0001, Xiaoyi Pang, Jiahui Hu 0001, Honglong Chen, Yang Yang 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | Breaking Secure Aggregation: Label Leakage from Aggregated Gradients in Federated LearningabstractFederated Learning (FL) exhibits privacy vulnerabilities under gradient inversion attacks (GIAs), which can extract private information from individual gradients. To enhance privacy, FL incorporates Secure Aggregation (SA) to prevent the server from obtaining individual gradients, thus effectively resisting GIAs. In this paper, we propose a stealthy label inference attack to bypass SA and recover individual clients’ private labels. Specifically, we conduct a theoretical analysis of label inference from the aggregated gradients that are exclusively obtained after implementing SA. The analysis results reveal that the inputs (embeddings) and outputs (logits) of the final fully connected layer (FCL) contribute to gradient disaggregation and label restoration. To preset the embeddings and logits of FCL, we craft a fishing model by solely modifying the parameters of a single batch normalization (BN) layer in the original model. Distributing client-specific fishing models, the server can derive the individual gradients regarding the bias of FCL by resolving a linear system with expected embeddings and the aggregated gradients as coefficients. Then the labels of each client can be precisely computed based on preset logits and gradients of FCL’s bias. Extensive experiments show that our attack achieves large-scale label recovery with 100% accuracy on various datasets and model architectures. Zhibo Wang 0001, Zhiwei Chang, Jiahui Hu 0001, Xiaoyi Pang, Jiacheng Du, Yongle Chen, Kui Ren 0001 |
INFOCOM | 4 |
| 2024 | Towards Efficient Asynchronous Federated Learning in Heterogeneous Edge EnvironmentsabstractFederated learning (FL) is widely used in edge environments as a privacy-preserving collaborative learning paradigm. However, edge devices often have heterogeneous computation capabilities and data distributions, hampering the efficiency of co-training. Existing works develop staleness-aware semi-asynchronous FL that reduces the contribution of slow devices to the global model to mitigate their negative impacts. But this makes data on slow devices unable to be fully leveraged in global model updating, exacerbating the effects of data heterogeneity. In this paper, to cope with both system and data heterogeneity, we propose a clustering and two-stage aggregation-based Efficient Asynchronous Federated Learning (EAFL) framework, which can achieve better learning performance with higher efficiency in heterogeneous edge environments. In EAFL, we first propose a gradient similarity-based dynamic clustering mechanism to cluster devices with similar system and data characteristics together dynamically during the training process. Then, we develop a novel two-stage aggregation strategy consisting of staleness-aware semi-asynchronous intra-cluster aggregation and data size-aware synchronous inter-cluster aggregation to efficiently and comprehensively aggregate training updates across heterogeneous clusters. With that, the negative impacts of slow devices and Non-IID data can be simultaneously alleviated, thus achieving efficient collaborative learning. Extensive experiments demonstrate that EAFL is superior to state-of-the-art methods. Xiaoyi Pang, Zhibo Wang 0001, Jiahui Hu 0001, Peng Sun 0003, Kui Ren 0001 |
INFOCOM | 2 |
| 2024 | Location and Bid Privacy Preserving-Based Quality-Aware Worker Recruitment Scheme in MCSabstractMobile Crowd Sensing (MCS) has become a prevalent large-scale and low-cost data collection paradigm by employing workers, and the location and bid privacy of both task and workers should not be leaked to the third party to prevent the adversary from attacking. Existing privacy preserving worker recruitment schemes have taken the location and quality into consideration, but ignore the bid privacy. To tackle this issue, a two-stage Location and Bid Privacy Preserving based Quality-aware Worker Recruitment (LBPP-QWR) scheme is proposed in this paper. In the first stage, to select those workers who satisfy the specified location and bid range of the task in the encrypted state, we propose a hybrid encryption scheme of matrix encryption and asymmetric encryption technique in the MCS platform. For the second stage, after obtaining the preliminary worker set via the platform, we propose a Knapsack Worker Selection (KWS) algorithm to recruit those high-quality and low bid workers under the budget constraint in the Data Requester (DR). Considering that there are quality-unknown workers, we further propose an improved.-KWS algorithm based on.-greedy algorithm by combining the exploration and exploitation mechanism to learn the quality of worker. Extensive experiments conducted on real-world datasets demonstrate that our proposed scheme can improve the average total quality by 17.96%-83.34%, and the cost efficiency by 27.99%-67.90% for the DR compared with other benchmark methods. Weifan Shi, Qingyong Deng, Zhetao Li, Saiqin Long, Haolin Liu 0001, Xiaoyi Pang |
IEEE Internet Things J. | 6 |
| 2024 | Does Differential Privacy Really Protect Federated Learning From Gradient Leakage Attacks?abstractFederated Learning (FL) is susceptible to the gradient leakage attack (GLA), which can recover local private training data from the shared gradients or model updates. To ensure privacy, differential privacy is applied in FL by clipping and adding noise to local gradients (i.e., Local Differential Privacy (LDP)) or the global model update (i.e., Central Differential Privacy (CDP)). However, the effectiveness of DP in defending GLAs needs to be thoroughly investigated since some works briefly verify that DP can guard FL against GLAs while others question its defense capability. In this paper, we empirically evaluate CDP and LDP on the resistance of GLAs, and pay close attention to the trade-offs between privacy and utility in FL. Our findings reveal that: 1) existing GLAs can be defended by CDP using a per-layer clipping strategy and LDP with a reasonable privacy guarantee and 2) both CDP and LDP ensure the trade-off between privacy and utility in training shallow model, but cannot guarantee this trade-off in deeper model training (e.g., ResNets). Triggered by the crucial role of clipping operation for DP, we propose an improved attack that incorporates the clipping operation into existing GLAs without requiring additional information. The experimental results show our attack can destruct the protection of CDP and weaken the effectiveness of LDP. Overall, our work validates the effectiveness as well as reveals the vulnerability of DP under GLAs. We hope this work can provide guidance on utilizing DP for defending against GLA in FL and inspire the design of future privacy-preserving FL. Jiahui Hu 0001, Jiacheng Du, Zhibo Wang 0001, Xiaoyi Pang, Peng Sun 0003, Kui Ren 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | Location Privacy-Aware Task Offloading in Mobile Edge ComputingabstractIn mobile edge computing (MEC), users can offload tasks to nearby MEC servers to reduce computation cost. Considering that the size of offloaded tasks could disclose user location information, several location privacy-preserving task offloading mechanisms have been proposed under the single-server scenario. However, to the best of our knowledge, none of them could provide a strict privacy protection guarantee or be applicable to the multi-server scenario where the user's location can be inferred more accurately if servers collude with each other. In this paper, we propose a novel location privacy-aware task offloading framework (LPA-Offload) for both single-server and multi-server scenarios, which provides strict and provable location privacy protection while achieving efficient task offloading. Specifically, we propose a location perturbation mechanism that allows each user to perturb its real location within a rational perturbation region and provides a differential privacy guarantee. To make a satisfactory offloading strategy, we propose a perturbation region determination mechanism and an offloading strategy generation mechanism that adaptively select a proper perturbation region according to the customized privacy factor, and then generate an optimal offloading strategy based on the perturbed location within the decided region. The determination of the perturbation region could achieve personalized privacy requirements while reducing computation cost. LPA-Offload is proved to satisfy$(\epsilon,\delta)$-differential privacy, and the experiments demonstrate the effectiveness of our framework. Zhibo Wang 0001, Yunan Sun, Defang Liu, Jiahui Hu 0001, Xiaoyi Pang, Yuke Hu, Kui Ren 0001 |
IEEE Trans. Mob. Comput. | 5 |
| 2024 | Shield Against Gradient Leakage Attacks: Adaptive Privacy-Preserving Federated LearningabstractFederated learning (FL) requires frequent uploading and updating of model parameters, which is naturally vulnerable to gradient leakage attacks (GLAs) that reconstruct private training data through gradients. Although some works incorporate differential privacy (DP) into FL to mitigate such privacy issues, their performance is not satisfactory since they did not notice that GLA incurs heterogeneous risks of privacy leakage (RoPL) with respect to gradients from different communication rounds and clients. In this paper, we propose an Adaptive Privacy-Preserving Federated Learning (Adp-PPFL) framework to achieve satisfactory privacy protection against GLA, while ensuring good performance in terms of model accuracy and convergence speed. Specifically, a leakage risk-aware privacy decomposition mechanism is proposed to provide adaptive privacy protection to different communication rounds and clients by dynamically allocating the privacy budget according to the quantified RoPL. In particular, we exploratively design a round-level and a client-level RoPL quantification method to measure the possible risks of GLA breaking privacy from gradients in different communication rounds and clients respectively, which only employ the limited information in general FL settings. Furthermore, to improve the FL model training performance (i.e., convergence speed and global model accuracy), we propose an adaptive privacy-preserving local training mechanism that dynamically clips the gradients and decays the noises added to the clipped gradients during the local training process. Extensive experiments show that our framework outperforms the existing differentially private FL schemes on model accuracy, convergence, and attack resistance. Jiahui Hu 0001, Zhibo Wang 0001, Yongsheng Shen, Bohan Lin, Peng Sun 0003, Xiaoyi Pang, Jian Liu 0012, Kui Ren 0001 |
IEEE/ACM Trans. Netw. | 6 |
| 2023 | Towards Fairness-aware Adversarial Network PruningabstractNetwork pruning aims to compress models while minimizing loss in accuracy. With the increasing focus on bias in AI systems, the bias inheriting or even magnification nature of traditional network pruning methods has raised a new perspective towards fairness-aware network pruning. Straightforward pruning plus debias methods and recent designs for monitoring disparities of demographic attributes during pruning have endeavored to enhance fairness in pruning. However, neither simple assembling of two tasks nor specifically designed pruning strategies could achieve the optimal trade-off among pruning ratio, accuracy, and fairness. This paper proposes an end-to-end learnable framework for fairness-aware network pruning, which optimizes both pruning and debias tasks jointly by adversarial training against those final evaluation metrics like accuracy for pruning, and disparate impact (DI) and equalized odds (DEO) for fairness. In other words, our fairness-aware adversarial pruning method would learn to prune without any handcraft rules. Therefore, our approach could flexibly adapt to variate network structures. Exhaustive experimentation demonstrates the generalization capacity of our approach, as well as superior performance on pruning and debias simultaneously. To highlight, the proposed method could preserve the SOTA pruning performance while significantly improving fairness by around 50% as compared to traditional pruning methods. Lei Zhang 0006, Zhibo Wang 0001, Xiaowei Dong, Yunhe Feng, Xiaoyi Pang, Kui Ren 0001 |
ICCV | 5 |
| 2023 | Towards Class-Balanced Privacy Preserving Heterogeneous Model AggregationabstractHeterogeneous model aggregation (HMA) is an effective paradigm that integrates on-device trained models heterogeneous in architecture and target task into a comprehensive model. Recent works adopt knowledge distillation to amalgamate the knowledge of learned features and predictions from heterogeneous on-device models to realize HMA. However, most of them ignore that the disclosure of learned features exposes on-device models to privacy attacks. Moreover, the aggregated model may suffer from the imbalanced supervision caused by the uneven distribution of amalgamated knowledge about each class and show class bias. In this article, to address these issues, we propose a response-based class-balanced heterogeneous model aggregation mechanism, called CBHMA. It can effectively achieve HMA in a privacy-preserving manner and alleviate class bias in the aggregated model. Specifically, CBHMA aggregates on-device models by using only their response information to reduce their privacy leakage risk. To mitigate the impact of imbalanced supervision, CBHMA quantitatively measures the imbalanced supervision level for each class. Based on that, CBHMA customizes fine-grained misclassification costs for each class and utilizes such costs to adjust the importance of each class (more importance to classes with weaker supervision) in the response-based HMA algorithm. Extensive experiments on two real-world datasets demonstrate the effectiveness of CBHMA. Xiaoyi Pang, Zhibo Wang 0001, Zeqing He, Peng Sun 0003, Meng Luo 0010, Ju Ren 0001, Kui Ren 0001 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2022 | Towards Online Privacy-preserving Computation Offloading in Mobile Edge ComputingabstractMobile Edge Computing (MEC) is a new paradigm where mobile users can offload computation tasks to the nearby MEC server to reduce their resource consumption. Some works have pointed out that the true amount of offloaded tasks may reveal the sensitive information (e.g., device usage pattern and location information) of users, and proposed several privacy-preserving offloading mechanisms. However, to the best of our knowledge, none of them can provide strict and provable privacy guarantee. In this paper, we focus on the privacy leakage issue in computation offloading in MEC with a honest-but-curious server, and propose a novel online privacy-preserving computation offloading mechanism, called OffloadingGuard, to generate efficient offloading strategies for users in real time, which provide strict user privacy guarantee while minimizing the total cost of task computation. To this end, we design a deep reinforcement learning-based offloading model which allows each user to adaptively determine the satisfactory perturbed offloading ratio according to the time-varying channel state at each time slot to achieve trade-off between user privacy and computation cost. In particular, to strictly protect the true amount of offloaded tasks and prevent the untrusted MEC server from revealing mobile users’ privacy, a range-constrained Laplace distribution is designed to obfuscate the original offloading ratio of each user and restrict the perturbed offloading ratio in a rational range. OffloadingGuard is proved to satisfy ϵ-differential privacy, and extensive experiments demonstrate its effectiveness. Xiaoyi Pang, Zhibo Wang 0001, Jingxin Li, Ruiting Zhou, Ju Ren 0001, Zhetao Li |
INFOCOM | 1 |
| 2022 | Towards Personalized Privacy-Preserving Incentive for Truth Discovery in Mobile Crowdsensing SystemsabstractIncentive mechanisms are essential for stimulating adequate worker participation to achieve good truth discovery performance in mobile crowdsensing (MCS) systems. However, most of existing incentive mechanisms only consider compensating workers’ sensing cost, while the cost incurred by potential privacy leakage has been largely neglected. Moreover, none of existing privacy-preserving incentive mechanisms has incorporated workers’ different privacy preferences to provide personalized payments for them. In this paper, we propose a contract-based personalized privacy-preserving incentive mechanism for truth discovery in MCS systems, named Paris-TD, which provides personalized payments for workers as a compensation for privacy cost while achieving accurate truth discovery. The basic idea is that the platform offers a set of different contracts to workers with different privacy preferences, and each worker chooses to sign a contract which specifies a privacy-preserving degree (PPD) and the corresponding payment the worker will receive if she submits perturbed data with that PPD. Specifically, we respectively design a set of optimal contracts analytically under both full and incomplete information models, which maximize the truth discovery accuracy under a given budget, while satisfying the individual rationality and incentive compatibility properties. The feasibility and effectiveness of Paris-TD are validated through experiments on both synthetic and real-world datasets. Peng Sun 0003, Zhibo Wang 0001, Liantao Wu, Yunhe Feng, Xiaoyi Pang, Hairong Qi 0001, Zhi Wang 0003 |
IEEE Trans. Mob. Comput. | 5 |
| 2022 | Privacy-Preserving Streaming Truth Discovery in Crowdsourcing With Differential PrivacyabstractDifferential privacy (DP) has gained popularity in truth discovery recently due to its strong privacy guarantee. However, existing DP mechanisms for streaming data publication are not suitable for truth discovery as they fail to consider the different reliabilities of individuals, while the DP-based approaches for truth discovery are not suitable for streaming data because they ignore the correlations between truths over time. Directly applying these existing methods to streaming crowdsourced data would lead to low accuracy of the discovered truth. To solve this problem, in this paper, we propose an edge computing based privacy-preserving truth discovery mechanism, named PrivSTD, for streaming crowdsourced data to realize high accuracy of discovered truth while protecting the privacy of workers. Specifically, edge servers are introduced between the untrusted cloud server and workers to securely calculate the local truths and workers’ reliabilities. A truth-dependent budget recycle mechanism is proposed for each edge server to adaptively determine the perturbed timestamp and allocate the privacy budget according to the changing pattern of local truths. Besides, a reliability-based perturbation mechanism is proposed to reduce the perturbation magnitude on the basis of worker's reliability. We theoretical analyze the data utility and computation cost of PrivSTD, and prove that PrivSTD can satisfy$w$-event ($\epsilon,\delta$)-differential privacy. Extensive experimental results on synthetic and real-world datasets demonstrate that PrivSTD achieves better utility than the state-of-the-art approaches. Dan Wang 0031, Ju Ren 0001, Zhibo Wang 0001, Xiaoyi Pang, Yaoxue Zhang, Xuemin Shen |
IEEE Trans. Mob. Comput. | 4 |
| 2022 | Towards Personalized Privacy-Preserving Truth Discovery Over Crowdsourced Data StreamsabstractTruth discovery is an effective paradigm which could reveal the truth from crowdsouced data with conflicts, enabling data-driven decision-making systems to make quick and smart decisions. The increasing privacy concern promotes users to perturb or encrypt their private data before outsourcing, which poses significant challenges for truth discovery. Although several privacy-preserving truth discovery mechanisms have been proposed, none of them take personal privacy expectation into consideration. In this work, we propose a novel personalized privacy-preserving truth discovery (PPPTD) framework over crowdsourced data streams to achieve timely and accurate truth discovery while guaranteeing the protection of individual privacy. The key challenges of PPPTD lie in improving the accuracy of truth estimation from the perturbed streaming data with personalized protection level. To address these challenges, we first develop a personalized budget initialization mechanism to quantify each user’s privacy protection requirement, and allocate personalized privacy budgets to users according to their privacy requirements. Then we propose a deviation-aware weighted aggregation method to improve the accuracy of truth discovery from streaming data with varying degrees of perturbation. In order to achieve privacy-utility tradeoff, we further propose an influence-aware adaptive budget adjustment mechanism that adaptively re-allocates privacy budgets to users based on the evolution of their influence in the weighted aggregation. We prove that PPPTD can achieve$\epsilon $-differential privacy over the whole data generated by users and satisfy individual personalized privacy requirements. Extensive experiments on two real-world datasets demonstrate the effectiveness of PPPTD. Xiaoyi Pang, Zhibo Wang 0001, Defang Liu, John C. S. Lui, Qian Wang 0002, Ju Ren 0001 |
IEEE/ACM Trans. Netw. | 1 |
| 2020 | Towards Pattern-aware Privacy-preserving Real-time Data CollectionabstractAlthough time-series data collected from users can be utilized to provide services for various applications, they could reveal sensitive information about users. Recently, local differential privacy (LDP) has emerged as the state-of-art approach to protect data privacy by perturbing data locally before outsourcing. However, existing works based on LDP perturb each data point separately without considering the correlations between consecutive data points in time-series. Thus, the important patterns of each time-series might be distorted by existing LDP-based approaches, leading to severe degradation of data utility. In this paper, we focus on real-time data collection under a honest-but-curious server, and propose a novel pattern-aware privacy-preserving approach, called PatternLDP, to protect data privacy while the pattern of time-series can still be preserved. To this end, instead of providing the same level of privacy protection at each data point, each user only samples remarkable points in time-series and adaptively perturbs them according to their impacts on local patterns. In particular, we propose a pattern-aware sampling method based on Piecewise Linear Approximation (PLA) to determine whether to sample and perturb current data point. To reduce the utility loss caused by pattern change after perturbation, we propose an importance-aware randomization mechanism to adaptively perturb sampled data locally while achieving better trade-off between privacy and utility. A novel metric-based w-event privacy is introduced to measure the privacy protection degree for pattern-rich time-series. We prove that PatternLDP can provide the above privacy guarantee, and extensive experiments on real-world datasets demonstrate that PatternLDP outperforms existing mechanisms and can effectively preserve the important patterns. Zhibo Wang 0001, Xiaoyi Pang, Ju Ren 0001, Zhe Liu 0001, Yongle Chen |
INFOCOM | 3 |
| 2019 | Privacy-Preserving Crowd-Sourced Statistical Data Publishing with An Untrusted ServerabstractThe continuous publication of aggregate statistics over crowd-sourced data to the public has enabled many data mining applications (e.g., real-time traffic analysis). Existing systems usually rely on a trusted server to aggregate the spatio-temporal crowd-sourced data and then apply differential privacy mechanism to perturb the aggregate statistics before publishing to provide strong privacy guarantee. However, the privacy of users will be exposed once the server is hacked or cannot be trusted. In this paper, we study the problem of real-time crowd-sourced statistical data publishing with strong privacy protection under an untrusted server. We propose a novel distributed agent-based privacy-preserving framework, called DADP, that introduces a new level of multiple agents between the users and the untrusted server. Instead of directly uploading the check-in information to the untrusted server, a user can randomly select one agent and upload the check-in information to it with the anonymous connection technology. Each agent aggregates the received crowd-sourced data and perturbs the aggregated statistics locally with Laplace mechanism. The perturbed statistics from all the agents are further combined together to form the entire perturbed statistics for publication. In particular, we propose a distributed budget allocation mechanism and an agent-based dynamic grouping mechanism to realize global w-event ε-differential privacy in a distributed way. We prove that DADP can provide w-event ε-differential privacy for real-time crowd-sourced statistical data publishing under the untrusted server. Extensive experiments on real-world datasets demonstrate the effectiveness of DADP.. Zhibo Wang 0001, Xiaoyi Pang, Yahong Chen, Huajie Shao, Qian Wang 0002, Honglong Chen, Hairong Qi 0001 |
IEEE Trans. Mob. Comput. | 2 |