Weiqi Wang 0003

dblp:51/5775-3 · DBLP profile ↗
← Back
31ranked-venue papers
13as first author
29since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 15 · 7 first-author · 15 since 2021Computer networks · 7 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Pixel-Depth Prototypical Knowledge Consolidation Network for Deepfake Detection
Ahmed Asiri, Weiqi Wang 0003, Luoyu Chen, Shui Yu 0001
ACISP (2)2
2026 Alleviating Budgetary Challenges in Machine Unlearning Services Through Insurance
Mingjian Tang 0002, Weiqi Wang 0003, Shui Yu 0001
ICC2
2026 Forget Me, Not My Friends! Object Unlearning Based on Scene Graphs
abstract
Machine unlearning offers a practical technical means for fulfilling users' requests to remove personally identifiable information (PII) under ''right to be forgotten'' regulations such as GDPR and COPPA. Traditionally, unlearning is performed with the removal of entire data samples (sample unlearning) or whole features across the dataset (feature unlearning). However, when the removal request targets only certain parts of the PII, such as specific objects within a sample, these traditional unlearning approaches fall short of meeting such finer-grained unlearning requirements. To address this gap, we propose a scene graph-based object unlearning framework. This framework utilizes scene graphs, rich in semantic representation, transparently translate unlearning requests into actionable steps. The result, is the preservation of the overall semantic integrity of the generated image, bar the unlearned object. Furthermore, we develop three distinct approaches for object unlearning, grounded in the mainstream unlearning techniques of fine-tuning and model redaction. For validation, we evaluate the unlearned object's fidelity in outputs under the tasks of image reconstruction and image synthesis. Our proposed framework demonstrates improved object unlearning outcomes, with the preservation of unrequested samples in contrast to sample and feature learning methods. This work addresses critical privacy issues by increasing the granularity of targeted machine unlearning through forgetting specific object-level details without sacrificing the utility of the whole data sample or dataset feature.
Chenhan Zhang, Benjamin Zi Hao Zhao, Hassan Jameel Asghar, Weiqi Wang 0003, An Liu 0002, Mohamed Ali Kâafar
WSDM4
2026 BlindU: Blind Machine Unlearning Without Revealing Erasing Data
abstract
Machine unlearning enables data holders to remove the contribution of their specified samples from trained models to protect their privacy. However, it is paradoxical that most unlearning methods require the unlearning requesters to first upload their data to the server as a prerequisite for unlearning. These methods are infeasible in many privacy-preserving scenarios where servers are prohibited from accessing users' data, such as federated learning (FL). In this paper, we explore how to implement unlearning under the condition of not uncovering the erasing data to the server. We propose Blind Unlearning (BlindU), which carries out unlearning using compressed representations instead of original inputs. BlindU only involves the server and the unlearning user: the user locally generates privacy-preserving representations, and the server performs unlearning solely on these representations and their labels. For the FL model training, we employ the information bottleneck (IB) mechanism. The encoder of the IB-based FL model learns representations that distort maximum task-irrelevant information from inputs, allowing FL users to generate compressed representations locally. For effective unlearning using compressed representation, BlindU integrates two dedicated unlearning modules tailored explicitly for IB-based models and uses a multiple gradient descent algorithm to balance forgetting and utility retaining. While IB compression already provides protection for task-irrelevant information of inputs, to further enhance the privacy protection, we introduce a noise-free differential privacy (DP) masking method to deal with the raw erasing data before compressing. Theoretical analysis and extensive experimental results illustrate the superiority of BlindU in privacy protection and unlearning effectiveness compared with the best existing privacy-preserving unlearning benchmarks.
Weiqi Wang 0003, Zhiyi Tian, Chenhan Zhang, Shui Yu 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 SMS: Self-Supervised Model Seeding for Verification of Machine Unlearning
abstract
Many machine unlearning methods have been proposed recently to uphold users' right to be forgotten. However, offering users verification of their data removal post-unlearning is an important yet under-explored problem. Current verifications typically rely on backdooring, i.e., adding backdoored samples to influence model performance. Nevertheless, the backdoor methods can merely establish a connection between backdoored samples and models but fail to connect the backdoor with genuine samples. Thus, the backdoor removal can only confirm the unlearning of backdoored samples, not users' genuine samples, as genuine samples are independent of backdoored ones. In this paper, we propose a Self-supervised Model Seeding (SMS) scheme to provide unlearning verification for genuine samples. Unlike backdooring, SMS links user-specific seeds (such as users' unique indices), original samples, and models, thereby facilitating the verification of unlearning genuine samples. However, implementing SMS for unlearning verification presents two significant challenges. First, embedding the seeds into the service model while keeping them secret from the server requires a sophisticated approach. We address this by employing a self-supervised model seeding task, which learns the entire sample, including the seeds, into the model's latent space. Second, maintaining the utility of the original service model while ensuring the seeding effect requires a delicate balance. We design a joint-training structure that optimizes both the self-supervised model seeding task and the primary service task simultaneously on the model, thereby maintaining model utility while achieving effective model seeding. The effectiveness of the proposed SMS scheme is evaluated through extensive experiments on three representative datasets, utilizing various model architectures and exact and approximate unlearning benchmarks. The results demonstrate that SMS provides effective verification for genuine sample unlearning, effectively addressing the limitations of existing solutions.
Weiqi Wang 0003, Chenhan Zhang, Zhiyi Tian, Shui Yu 0001
IEEE Trans. Dependable Secur. Comput.1
2026 Ellipsoid Control: A White-List Jailbreak Defense via Benign Latent Modeling
Luoyu Chen, Weiqi Wang 0003, Zhiyi Tian, Ahmed Asiri, Shui Yu 0001
IEEE Trans. Inf. Forensics Secur.2
2025 Fine-Grained Privacy-Preserving Semantic Communication against Eavesdropping Attacks
abstract
Semantic communication has emerged as a pivotal technology for future wireless communication, where privacy preservation plays a critical role. However, existing privacy-preserving semantic communication methods have a negative impact on communication performance. In this paper, we focus on fine-grained privacy-preserving semantic communication to achieve a better performance-privacy trade-off. Specifically, we formalize it as an optimization problem and propose a Dual Information Bottleneck (Dual-IB) method to solve it. Firstly, Dual-IB extracts task-relevant information from input and compresses task-irrelevant information. Secondly, it minimizes sensitive attributes within the task-relevant representation by information bottleneck. This mitigates the risks of privacy exposure while maintaining the task performance. We validate the effectiveness of our method using an eavesdropping attack that attempts to extract private information in the transmission. The experimental results demonstrate that our method not only achieves a task performance of approximately 97%, but also reduces eavesdropping accuracy by 68%, effectively addressing a critical limitation of existing approaches.
Zhiyi Tian, Weiqi Wang 0003, Shui Yu 0001
GLOBECOM3
2025 IDIR: Interpolated Diffusion Image Reconstruction for Generalizable Detection of Synthetic Images
Weiqi Wang 0003, Zhiyi Tian, Shui Yu 0001
PAKDD (4)2
2025 Can Self Supervision Rejuvenate Similarity-Based Link Prediction?
Chenhan Zhang, Weiqi Wang 0003, Zhiyi Tian, James Jian Qiao Yu, Mohamed Ali Kâafar, An Liu 0002, Shui Yu 0001
PAKDD (7)2
2025 Incentive-Compatible Pricing for Truthful Data Sharing
abstract
With the rapid development of Artificial Intelligence, fake data are being generated and circulated on an unprecedented scale. While such data may be entertaining in some contexts, data collectors are generally reluctant to accept them due to concerns about service quality and security risks. A central challenge, therefore, is how to incentivise contributors to provide authentic data. Existing research has mainly focused on regulating the use of fake data, with limited attention to designing mechanisms that motivate contributors to share real data. In this paper, we propose a pricing strategy with an embedded incentive mechanism that ensures contributors obtain higher financial benefits from sharing real data rather than fake data. We formulate the problem as a Stackelberg game, where the data collector acts as the leader and contributors are the followers. To enhance reliability, we integrate a peer-review method to help verify data authenticity, which in turn informs the incentive-compatible pricing design. We establish the existence and uniqueness of the equilibrium solution, and further conduct numerical simulations to demonstrate the equilibrium-seeking process and the influence of key parameters.
Mingjian Tang 0002, Weiqi Wang 0003, Shui Yu 0001
TrustCom2
2025 TAPE: Tailored Posterior Difference for Auditing of Machine Unlearning
abstract
With the increasing prevalence of Web-based platforms handling vast amounts of user data, machine unlearning has emerged as a crucial mechanism to uphold users' right to be forgotten, enabling individuals to request the removal of their specified data from trained models. However, the auditing of machine unlearning processes remains significantly underexplored. Although some existing methods offer unlearning auditing by leveraging backdoors, these backdoor-based approaches are inefficient and impractical, as they necessitate involvement in the initial model training process to embed the backdoors. In this paper, we propose a TAilored Posterior diffErence (TAPE) method to provide unlearning auditing independently of original model training. We observe that the process of machine unlearning inherently introduces changes in the model, which contains information related to the erased data. TAPE leverages unlearning model differences to assess how much information has been removed through the unlearning operation. Firstly, TAPE mimics the unlearned posterior differences by quickly building unlearned shadow models based on first-order influence estimation. Secondly, we train a Reconstructor model to extract and evaluate the private information of the unlearned posterior differences to audit unlearning. Existing privacy reconstructing methods based on posterior differences are only feasible for model updates of a single sample. To enable the reconstruction effective for multi-sample unlearning requests, we propose two strategies, unlearned data perturbation and unlearned influence-based division, to augment the posterior difference. Extensive experimental results indicate the significant superiority of TAPE over the state-of-the-art unlearning verification methods, at least 4.5x efficiency speedup and supporting the auditing for broader unlearning scenarios.
Weiqi Wang 0003, Zhiyi Tian, An Liu 0002, Shui Yu 0001
WWW1
2025 Backdoored Sample Cleansing for Unlabeled Datasets via Bootstrapped Dual Set Purification
abstract
Self-Supervised Learning (SSL) excels in utilizing unlabeled data for feature representation learning. However, recent studies have revealed that SSL is vulnerable to data poisoning-based backdoor attacks. To remove backdoored samples from the SSL training dataset, model optimization methods often fine-tune a trained model by contrasting the training dataset with a reserved clean dataset. This contrastive training effectively marginalizes backdoored samples from the distribution of benign ones,if and only ifboth the reserved clean dataset and the training dataset are from the same data distribution. However, presuming identical distributions between the web-scraped data and reserved data is impractical. To address this impractical assumption, our proposed Bootstrapped Dual SetPurification (AUTO) method distinguishes backdoored from benign samples by contrasting a mined ‘positive set’ and a mined ‘negative set’ within the training dataset itself. We exploit the resistance of backdoored samples in data mixing to mine a highly poisoned ‘positive set’ and a minimally poisoned ‘negative set’. Besides, AUTO mitigates unstable detection performance within different optimization steps by continuously refining the dual sets by the optimized model, enhancing the model's poison distinguishability from consistently improving supervision signals. Our extensive experiments on Cifar10, Cifar100, and Imagenet100 against existing data poisoning SSL backdoor attacks demonstrate AUTO's superiority in detection performance over all existing defenses.
Luoyu Chen, Weiqi Wang 0003, Zhiyi Tian, Chenhan Zhang, Shui Yu 0001
IEEE Trans. Dependable Secur. Comput.2
2025 CRFU: Compressive Representation Forgetting Against Privacy Leakage on Machine Unlearning
abstract
Machine unlearning allows data owners to erase the impact of their specified data from trained models. Unfortunately, recent studies have shown that adversaries can recover the erased data, posing serious threats to user privacy. An effective unlearning method removes the information of the specified data from the trained model, resulting in different outputs for the same input before and after unlearning. Adversaries can exploit these output differences to conduct privacy leakage attacks, such as reconstruction and membership inference attacks. However, directly applying traditional defenses to unlearning leads to significant model utility degradation. In this article, we introduce a Compressive Representation Forgetting Unlearning scheme (CRFU), designed to safeguard against privacy leakage on unlearning. CRFU achieves data erasure by minimizing the mutual information between the trained compressive representation (learned through information bottleneck theory) and the erased data, thereby maximizing the distortion of data. This ensures that the model's output contains less information that adversaries can exploit. Furthermore, we introduce a remembering constraint and an unlearning rate to balance the forgetting of erased data with the preservation of previously learned knowledge, thereby reducing accuracy degradation. Theoretical analysis demonstrates that CRFU can effectively defend against privacy leakage attacks. Our experimental results show that CRFU significantly increases the reconstruction mean square error (MSE), achieving a defense effect improvement of approximately 200% against privacy reconstruction attacks with only 1.5% accuracy degradation on MNIST.
Weiqi Wang 0003, Chenhan Zhang, Zhiyi Tian, Shushu Liu, Shui Yu 0001
IEEE Trans. Dependable Secur. Comput.1
2025 FedU: Federated Unlearning via User-Side Influence Approximation Forgetting
abstract
Machine unlearning has become a significant research topic on a global scale due to the increasing importance of privacy protection, particularly in light of the right to be forgotten legislation. Although many solutions are proposed, the current mainstream centralized machine unlearning studies are not feasible in federated learning (FL), where the server has no access to any users’ unlearning samples. In this paper, we aim to tackle thefederated unlearningproblem by proposing a Federated Unlearning (FedU) scheme via a user-side influence approximation forgetting method, thereby eliminating the need to share raw data with the server. In FedU, only users who have unlearning needs execute the influence approximation forgetting, while other users and the server just conduct the same operations as they did in FL. The proposed influence approximation forgetting method achieves unlearning by estimating the influence of the erased samples relying on only the user's local data and eliminating this influence from the model. However, the model utility is still negatively influenced by directly removing the influence estimation. To mitigate the side effects of unlearning, we propose a utility preservation method that simultaneously trains the unlearned model based on the unlearning requesters’ remaining local dataset. We design an adaptive optimization method to balance the forgetting and utility preservation effectiveness optimally during the unlearning process. Extensive evaluations on three representative public datasets demonstrate that our proposed method significantly outperforms state-of-the-art methods in both effectiveness and efficiency, avoiding more than 3% accuracy degradation when the number of unlearning requesters is large.
Weiqi Wang 0003, Chenhan Zhang, Zhiyi Tian, Shui Yu 0001
IEEE Trans. Dependable Secur. Comput.1
2025 SCU: An Efficient Machine Unlearning Scheme for Deep Learning Enabled Semantic Communications
abstract
Deep learning (DL) enabled semantic communications leverage DL to train encoders and decoders (codecs) to extract and recover semantic information. However, most semantic training datasets contain personal private information. Such concerns call for enormous requirements for specified data erasure from semantic codecs when previous users hope to move their data from the semantic system. Existing machine unlearning solutions remove data contribution from trained models, yet usually in supervised sole model scenarios. These methods are infeasible in semantic communications that often need to jointly train unsupervised encoders and decoders. In this paper, we investigate the unlearning problem in DL-enabled semantic communications and propose a semantic communication unlearning (SCU) scheme to tackle the problem. SCU includes two key components. Firstly, we customize the joint unlearning method for semantic codecs, including the encoder and decoder, by minimizing mutual information between the learned semantic representation and the erased samples. Secondly, to compensate for semantic model utility degradation caused by unlearning, we propose a contrastive compensation method, which considers the erased data as the negative samples and the remaining data as the positive samples to retrain the unlearned semantic models contrastively. Theoretical analysis and extensive experimental results on three representative datasets demonstrate the effectiveness and efficiency of our proposed methods.
Weiqi Wang 0003, Zhiyi Tian, Chenhan Zhang, Shui Yu 0001
IEEE Trans. Inf. Forensics Secur.1
2025 Evaluation of Machine Unlearning Through Model Difference
abstract
Increasing attention is being paid to machine unlearning, which supports individuals’ “right to be forgotten.” While most studies focus on the efficiency and effectiveness of unlearning algorithms, the evaluation of machine unlearning effectiveness remains underexplored. Offering robust evaluation services for unlearning is critical, not only to uphold privacy legislation but also to assess and improve existing unlearning methods. Lots of existing methods employ backdoor methods to evaluate unlearning effectiveness, which can only verify the unlearning effect of backdoored samples and negatively impact the model utility as they need to embed backdoors into the model first. In this paper, we propose an evaluating machine unlearning (EMU) method, which aims to evaluate the effectiveness of unlearning and verify data removal without the aforementioned adverse effects. Machine unlearning inherently creates a difference on the model before and after unlearning. The model difference contains information about the unlearned samples, which can be extracted through reconstruction models for unlearning effectiveness evaluation. To efficiently generate the model differences as input for evaluation, we simulate the model changes based on the influence function theory. Additionally, we design a multi-task information bottleneck structure to enhance the scalability of EMU and simplify the analysis of different learning tasks. We provide a theoretical analysis of how the similarity between erased and remaining samples, as well as task types, affects the extent of unlearning—factors that have been largely overlooked. Extensive experiments on various model architectures and representative datasets confirm our analysis, demonstrating the effective evaluation for unlearning without any degradation in the service model utility.
Weiqi Wang 0003, Chenhan Zhang, Zhiyi Tian, Shui Yu 0001, Zhou Su 0001
IEEE Trans. Inf. Forensics Secur.1
2024 UFL: Unlinkable Federated Learning Through Shuffle and Shamir's Secret Sharing
Jingxue Chen, Zhiwei Si, Jingcheng Song, Manoranjan Mohanty, Weiqi Wang 0003, Hu Xiong
ADMA (2)5
2024 OPMUS: A Win-Win Pricing Strategy for Machine Unlearning Service
Mingjian Tang 0002, Weiqi Wang 0003, Shui Yu 0001
ADMA (1)2
2024 A Mutation-Based Method for Backdoored Sample Detection Without Clean Data
abstract
Backdoor attacks significantly threaten machine learning-based vision systems. Existing detection methods typically require clean data from a similar distribution as the dataset under inspection, limiting practical deployment. This work proposes a Mutation-Based Method (MBM) for detecting and filtering backdoored samples in image training dataset, without referencing any external clean data. MBM aims at distinguishing backdoored and benign samples distribution via their distinct stability in feature space under certain data augmentations. Firstly, MBM applies multiple data augmentation techniques, generating mutated versions of each sample to ‘deactivate’ potential triggers while maintaining natural semantics not heavily distorted. Secondly, MBM measures how sample features diverge after mutating from its origin as poison score, which we call ‘Feature Stability’. Thirdly, by analyzing extreme scores within each class, MBM effectively identifies the backdoored class, and isolates samples not from backdoored class as clean data. Finally, a benign distribution is fit to benchmark against backdoored samples from backdoored class. We validated MBM on the CIFAR-10 dataset, achieving a true positive rate above 95% and a false positive rate below 0.2% for all defense settings. Our results confirm MBM’s efficacy without reliance on external clean data.
Luoyu Chen, Tao Zhang 0165, Ahmed Asiri, Weiqi Wang 0003, Shui Yu 0001
GLOBECOM5
2024 A Low-cost Black-box Jailbreak Based on Custom Mapping Dictionary with Multi-round Induction
abstract
Note:This paper contains many malicious contents generated by LLMs. Jailbreak can cause large language models (LLMs) to violate moral and ethical guidelines, generating harmful text and images. However, the existing jailbreaks are high-cost (e.g., fine-tuning LLMs) and underperform when facing some specific jailbreak tasks (e.g., generating pornographic and bloody texts), which has too many limitations for reflecting the threat of jailbreak on LLMs’ security. To fill this gap, in this paper, we propose a black-box jailbreak with a higher attack success rate and lower cost, called Dictionary Jailbreak with Multi-round Induction (DJMI). First, in DJMI, we create a simple custom language dictionary based on specific jailbreak tasks to be performed and send it to the LLM. Then, we use the custom language from the dictionary to construct malicious prompts, instructing the LLM to respond in the custom language as much as possible. Generally, in the first round, the LLM will provide a neutral response (such as translating the malicious prompts) or refuse to answer. To counteract this, we need to emphasize that the prompts comply with ethical guidelines and repeat the prompts using the custom language. The process can be conducted with multiple rounds to guide LLM in outputting harmful content. During the attack process, the only attack cost is the transmission cost of the prompts through the official interactive interface, without any other costs of LLM fine-tuning, algorithm design and prompt generation. Thus, DJMI is a low-cost jailbreak method. We verified DJMI on multiple mainstream LLMs across various jailbreak tasks. Experiments show that compared to existing black-box jailbreaks, DJMI achieves a higher attack success rate (> 80% on average) across specific jailbreak tasks with various risk levels. Additionally, DJMI enables the LLM to provide highly detailed execution steps for harmful behaviors within a session (such as specifying proportions of ingredients, synthetic chemical formulas, experimental conditions, and other detailed information in illicit drug production), further highlighting the severity of jailbreak attacks.
Weiqi Wang 0003, Youyang Qu, Shui Yu 0001
TrustCom2
2024 Machine Unlearning via Representation Forgetting With Parameter Self-Sharing
abstract
Machine unlearning enables data owners to remove the contribution of their specified samples from trained models. However, existing methods fail to strike an optimal balance between erasure effectiveness and model utility preservation. Previous studies focused on removing the impact of user-specified data from the model as much as possible to implement unlearning. These methods usually result in significant model utility degradation, commonly called catastrophic unlearning. To address the issue, we systematically consider machine unlearning and formulate it as a two-objective optimization problem that involves forgetting the erased data and retaining the previously learned knowledge, highlighting accuracy preservation during the unlearning process. We propose an unlearning method called representation-forgetting unlearning with parameter self-sharing (RFU-SS) to achieve the two-objective unlearning goal. Firstly, we design a representation-forgetting unlearning (RFU) method that aims to remove the contribution of specified samples from a trained representation by minimizing the mutual information between the representation and the erased data. The representation is learned using the information bottleneck (IB) method. RFU is tailored to the IB structure models for ease of introduction. Secondly, we customize a parameter self-sharing structural optimization method for RFU (i.e., RFU-SS) to simultaneously optimize the forgetting and retention objectives to find the optimal balance. Extensive experimental results demonstrate a significant effectiveness improvement of RFU-SS over the state-of-the-art methods. RFU-SS almost eliminates catastrophic unlearning, reducing model accuracy degradation from over 6% to less than 0.2% on the MNIST dataset with an even better removal effect. The source code is available athttps://github.com/wwq5-code/RFU-SS.git.
Weiqi Wang 0003, Chenhan Zhang, Zhiyi Tian, Shui Yu 0001
IEEE Trans. Inf. Forensics Secur.1
2024 Forgetting and Remembering Are Both You Need: Balanced Graph Structure Unlearning
abstract
In light of the growing emphasis on the right to be forgotten of graph data, machine unlearning has been extended to unlearn the graph structures’ knowledge from graph neural networks (GNNs), namely, structure unlearning. Whereas the complex dependencies in graph data, structure unlearning is intrinsically prone to imbalanced performance between the objectives of knowledge forgetting and model utility maintenance. Nevertheless, most existing methods fall short in addressing the two objectives in tandem and developing balanced solutions. In this paper, we propose imbalanced Structure Unlearning Mitigation using MultI-objective OpTimization (SUMMIT), which aims to develop balanced solutions regarding both knowledge forgetting and model utility maintenance effects. Corresponding to the two aspects, we first construct two tailored objectives that specifically address the challenges inherent in structure unlearning. Specifically, for the forgetting objective, we introduce a higher-order forgetting enhancement strategy aimed at mitigating the adverse effects of GNN oversmoothing on node decoupling. For the remembering objective, we adhere to the principle of ideal unlearning and propose to minimize the distributional distance between the node embeddings developed by unlearned and well-trained GNNs. Considering the potential competitive relationship between the two objectives during the optimization process, we present an adaptive two-objective balancer based on multi-objective optimization to reconcile the two objectives and strike a balance between them. We conduct comprehensive experiments to evaluate the efficacy of SUMMIT on three representative GNNs and four datasets, and compare the performance of SUMMIT with its ablation variants and a cadre of baselines. We demonstrate the superiority of SUMMIT in its ability to yield optimal and balanced solutions, addressing both the facets of knowledge forgetting and model utility maintenance.
Chenhan Zhang, Weiqi Wang 0003, Zhiyi Tian, Shui Yu 0001
IEEE Trans. Inf. Forensics Secur.2
2023 BFU: Bayesian Federated Unlearning with Parameter Self-Sharing
abstract
As the right to be forgotten has been legislated worldwide, many studies attempt to design machine unlearning mechanisms to enable data erasure from a trained model. Existing machine unlearning studies focus on centralized learning, where the server can access all users’ data. However, in a popular scenario, federated learning (FL), the server cannot access users’ training data. In this paper, we investigate the problem of machine unlearning in FL. We formalize a federated unlearning problem and propose a bayesian federated unlearning (BFU) approach to implement unlearning for a trained FL model without sharing raw data with the server. Specifically, we first introduce an unlearning rate in BFU to balance the trade-off between forgetting the erased data and remembering the original global model, making it adaptive to different unlearning tasks. Then, to mitigate accuracy degradation caused by unlearning, we propose BFU with parameter self-sharing (BFU-SS). BFU-SS considers data erasure and maintaining learning accuracy as two tasks and optimizes them together during unlearning. Extensive comparisons between our methods and the state-of-art federated unlearning method demonstrate the superiority of our proposed realizations.
Weiqi Wang 0003, Zhiyi Tian, Chenhan Zhang, An Liu 0002, Shui Yu 0001
AsiaCCS1
2023 Extracting Privacy-Preserving Subgraphs in Federated Graph Learning using Information Bottleneck
abstract
As graphs are getting larger and larger, federated graph learning (FGL) is increasingly adopted, which can train graph neural networks (GNNs) on distributed graph data. However, the privacy of graph data in FGL systems is an inevitable concern due to multi-party participation. Recent studies indicated that the gradient leakage of trained GNN can be used to infer private graph data information utilizing model inversion attacks (MIA). Moreover, the central server can legitimately access the local GNN gradients, which makes MIA difficult to counter if the attacker is at the central server. In this paper, we first identify a realistic crowdsourcing-based FGL scenario where MIA from the central server towards clients’ subgraph structures is a nonnegligible threat. Then, we propose a defense scheme, Subgraph-Out-of-Subgraph (SOS), to mitigate such MIA and meanwhile, maintain the prediction accuracy. We leverage the information bottleneck (IB) principle to extract task-relevant subgraphs out of the clients’ original subgraphs. The extracted IB-subgraphs are used for local GNN training and the local model updates will have less information about the original subgraphs, which renders the MIA harder to infer the original subgraph structure. Particularly, we devise a novel neural network-powered approach to overcome the intractability of graph data’s mutual information estimation in IB optimization. Additionally, we design a subgraph generation algorithm for finally yielding reasonable IB-subgraphs from the optimization results. Extensive experiments demonstrate the efficacy of the proposed scheme, the FGL system trained on IB-subgraphs is more robust against MIA attacks with minuscule accuracy loss.
Chenhan Zhang, Weiqi Wang 0003, James Jian Qiao Yu, Shui Yu 0001
AsiaCCS2
2023 CP-FL: Practical Gradient Leakage Defense in Federated Learning with Compressive Privacy
abstract
Federated learning (FL) requires clients to train constituted models based on their local datasets. Clients usually directly train local models using their entire datasets without distinguishing which information of data is task-relevant or irrelevant. Task-irrelevant information does not contribute to the learning task but exposes additional privacy information to adversaries. Studies have shown that unintended information leakage from gradients during FL iterations threatens clients' privacy. Researchers applied differential privacy (DP) to protect clients' gradients, but it does not help to reduce task-irrelevant information from the gradients. In this paper, we propose a compressive privacy federated learning (CP-FL) scheme to protect the task-irrelevant information from gradient leakage attacks. In CP-FL, clients train a local compressive model according to the global task. The local compressive model constructs a new representation, which extracts task-relevant and removes task-irrelevant information from clients' data. Since the global model is updated based on the compressed representation that eliminates the task-irrelevant information, it can effectively prevent adversaries from inferring those property values from the uploaded gradients. Moreover, with the help of a powerful local compressive model that sanitizes the challenging data into a low-dimension space representation, CP-FL can use a small global model instead of a sizeable one, significantly reducing communication. Both theoretical analysis and extensive experimental results demonstrate that CP-FL can effectively defend against gradient leakage attacks while maintaining practical utility.
Weiqi Wang 0003, Shushu Liu, Chenhan Zhang, Mingjian Tang 0002, Shui Yu 0001
GLOBECOM1
2023 Conditional Matching GAN Guided Reconstruction Attack in Machine Unlearning
abstract
Machine unlearning allows data owners to erase certain data and its impact from learning models for the right to be forgotten. However, privacy risks during the unlearning process have been identified. Earlier studies have used differences in model outputs before and after unlearning to conduct membership inference attacks. Nevertheless, the current attacks on machine unlearning are limited to inference and cannot reconstruct data without access to the victim's dataset. In this paper, we propose a reconstruction attack towards machine unlearning (RAU), which can reconstruct the unlearned data by exploiting the privacy leakage from the two models. To improve reconstruction quality, we propose a Conditional Matching Generative Adversarial Network (CMGAN), a novel variant of generative adversarial networks which introduces a reconstructive loss. Our work demonstrates the possible privacy leakage of current machine unlearning scenarios. Experimental results on MNIST and Fashion-MNIST show that the proposed attack achieves high label recovery accuracy and good data recovery performance.
Weiqi Wang 0003, Zipei Fan, Xuan Song 0001, Shui Yu 0001
GLOBECOM2
2023 FedMC: Federated Learning with Mode Connectivity Against Distributed Backdoor Attacks
abstract
Federated learning (FL) has become a hot research domain due to its privacy protection for model collaboratively training in edge computing systems. However, recent studies indicated that most FL algorithms have desperately suffered from backdoor attacks. Although many backdoor defence FL algorithms were proposed, their effects were highly related to the ratio of malicious clients (RMC) of all participated edge nodes. To be more specific, most of them only set RMC around 10% to 30% in their experiments, and their results also showed that the rate of successful backdoor defence seriously drops when RMC increases. In the paper, we propose a novel federated learning scheme with mode connectivity (FedMC) to defend against backdoor attacks, mitigating the sharp defence effect degradation as RMC increases. Conventional mode connectivity mainly focuses on training a connecting curve between two end models, which is inapplicable in distributed multiple clients FL situations. We extend the two-ends mode connectivity to multi-ends by introducing a scalable regularization term consisting of the edge clients' models to involve their knowledge in the connective model training. In each communication round, the FL-Server aggregates and absorbs the contribution of clients by training a connective model based on a small set of clean samples, which builds a pathway to accurately connect all edge clients' models and mitigates the backdoor triggers of models. Extensive experiments and results demonstrate that FedMC can effectively defend against backdoor attacks while maintaining the accuracy on untampered test data.
Weiqi Wang 0003, Chenhan Zhang, Shushu Liu, Mingjian Tang 0002, An Liu 0002, Shui Yu 0001
ICC1
2023 RUE: Realising Unlearning from the Perspective of Economics
abstract
Machine unlearning has quickly emerged as a technique to withdraw users’ data from the trained model to protect their privacy. Yet the cost of completely unlearning by retraining is high and the unlearning service is thus hard to proceed in the market. We work on the SISA method that greatly lowers the cost of unlearning as an example in this manuscript. We model the problem with a game theory model that balances customers’ benefit and the service provider’s profit, and the optimal price is acquired that both parties could accept. More specifically, in the game model, we calculate customers’ average waiting time with bulk service queueing model, and linearly estimate the customers’ benefit in terms of privacy from withdrawing their data. Our method RUE shows that both parties get more benefit or profit than others, which could make the unlearning service run smoothly when the concerns about the high price are removed. We also analyse the influence of the waiting time on the number of unlearning requests and on the price.
Mingjian Tang 0002, Weiqi Wang 0003, Chenhan Zhang, Shui Yu 0001
TrustCom2
2022 Locally Random Sampling for Practical Privacy Protection in Federated Learning
abstract
Federated learning (FL) is an emerging solution for machine learning model training in edge/fog computing systems. Unlike traditional systems that collect and train models on clouds, FL allows multiple edge/fog nodes to train a global model collaboratively without revealing their local data to clouds. Compared with traditional systems, it is inherited with better privacy protection ability. Although the basic privacy protection is inherited in FL, the privacy leakage from shard models is still unsolved. Existing solutions attempt to enhance the privacy of shared model parameters by adding differential privacy (DP) noise. However, these solutions all suffer from accuracy loss and convergence problems owing to the injected noise. In this paper, we propose a novel federated learning protocol to solve the above problem. The model trained on a carefully selected sampling subset can achieve the same level privacy protection as DP while preserving the model accuracy. Experimentally, we proved that our protocol achieves better model accuracy in the same privacy guarantee compared with noise injecting DP methods.
Weiqi Wang 0003, Shushu Liu, An Liu 0002, Christy Jie Liang, Shui Yu 0001
GLOBECOM1
2019 Protecting multi-party privacy in location-aware social point-of-interest recommendation
Weiqi Wang 0003, An Liu 0002, Zhixu Li, Xiangliang Zhang 0001, Qing Li 0001, Xiaofang Zhou 0001
World Wide Web1
2018 Efficient task assignment in spatial crowdsourcing with worker and task privacy protection
An Liu 0002, Weiqi Wang 0003, Shuo Shang, Qing Li 0001, Xiangliang Zhang 0001
GeoInformatica2