EDBT 2026 Demo / reviewers in the wild / expert
Zhiyi Tian
dblp:195/7695
· DBLP profile ↗
30ranked-venue papers
2as first author
30since 2021 · last 2026
0000-0001-8905-0941ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 15 · 1 first-author · 15 since 2021Computer networks · 6 · 6 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dual-Verbalizer with Label Correlation Modeling for Few-Shot Multi-Label Text Classification
Zhiyi Tian, Guangzhong Sun, Jingwei Sun 0001 |
DASFAA (4) | 1 |
| 2026 | ASWmark: A copyright protection approach for audio classification datasets
Xuefeng Fan, Zhiyi Tian, Fan Xing, Jixin Ma 0001, Xiaoyi Zhou |
Expert Syst. Appl. | 3 |
| 2026 | BlindU: Blind Machine Unlearning Without Revealing Erasing DataabstractMachine unlearning enables data holders to remove the contribution of their specified samples from trained models to protect their privacy. However, it is paradoxical that most unlearning methods require the unlearning requesters to first upload their data to the server as a prerequisite for unlearning. These methods are infeasible in many privacy-preserving scenarios where servers are prohibited from accessing users' data, such as federated learning (FL). In this paper, we explore how to implement unlearning under the condition of not uncovering the erasing data to the server. We propose Blind Unlearning (BlindU), which carries out unlearning using compressed representations instead of original inputs. BlindU only involves the server and the unlearning user: the user locally generates privacy-preserving representations, and the server performs unlearning solely on these representations and their labels. For the FL model training, we employ the information bottleneck (IB) mechanism. The encoder of the IB-based FL model learns representations that distort maximum task-irrelevant information from inputs, allowing FL users to generate compressed representations locally. For effective unlearning using compressed representation, BlindU integrates two dedicated unlearning modules tailored explicitly for IB-based models and uses a multiple gradient descent algorithm to balance forgetting and utility retaining. While IB compression already provides protection for task-irrelevant information of inputs, to further enhance the privacy protection, we introduce a noise-free differential privacy (DP) masking method to deal with the raw erasing data before compressing. Theoretical analysis and extensive experimental results illustrate the superiority of BlindU in privacy protection and unlearning effectiveness compared with the best existing privacy-preserving unlearning benchmarks. Weiqi Wang 0003, Zhiyi Tian, Chenhan Zhang, Shui Yu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | SMS: Self-Supervised Model Seeding for Verification of Machine UnlearningabstractMany machine unlearning methods have been proposed recently to uphold users' right to be forgotten. However, offering users verification of their data removal post-unlearning is an important yet under-explored problem. Current verifications typically rely on backdooring, i.e., adding backdoored samples to influence model performance. Nevertheless, the backdoor methods can merely establish a connection between backdoored samples and models but fail to connect the backdoor with genuine samples. Thus, the backdoor removal can only confirm the unlearning of backdoored samples, not users' genuine samples, as genuine samples are independent of backdoored ones. In this paper, we propose a Self-supervised Model Seeding (SMS) scheme to provide unlearning verification for genuine samples. Unlike backdooring, SMS links user-specific seeds (such as users' unique indices), original samples, and models, thereby facilitating the verification of unlearning genuine samples. However, implementing SMS for unlearning verification presents two significant challenges. First, embedding the seeds into the service model while keeping them secret from the server requires a sophisticated approach. We address this by employing a self-supervised model seeding task, which learns the entire sample, including the seeds, into the model's latent space. Second, maintaining the utility of the original service model while ensuring the seeding effect requires a delicate balance. We design a joint-training structure that optimizes both the self-supervised model seeding task and the primary service task simultaneously on the model, thereby maintaining model utility while achieving effective model seeding. The effectiveness of the proposed SMS scheme is evaluated through extensive experiments on three representative datasets, utilizing various model architectures and exact and approximate unlearning benchmarks. The results demonstrate that SMS provides effective verification for genuine sample unlearning, effectively addressing the limitations of existing solutions. Weiqi Wang 0003, Chenhan Zhang, Zhiyi Tian, Shui Yu 0001 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2026 | Ellipsoid Control: A White-List Jailbreak Defense via Benign Latent Modeling
Luoyu Chen, Weiqi Wang 0003, Zhiyi Tian, Ahmed Asiri, Shui Yu 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | MG-Det: Deepfake Detection with Multi-granularity
Ahmed Asiri, Luoyu Chen, Zhiyi Tian, Shui Yu 0001 |
ACISP (3) | 3 |
| 2025 | Introducing Graph Context into Language Models through Parameter-Efficient Fine-Tuning for Lexical Relation MiningabstractLexical relation refers to the way words are related within a language. Prior work has demonstrated that pretrained language models (PLMs) can effectively mine lexical relations between word pairs. However, they overlook the potential of graph structures composed of lexical relations, which can be integrated with the semantic knowledge of PLMs. In this work, we propose a parameter-efficient fine-tuning method through graph context, which integrates graph features and semantic representations for lexical relation classification (LRC) and lexical entailment (LE) tasks. Our experiments show that graph features can help PLMs better understand more complex lexical relations, establishing a new state-of-the-art for LRC and LE. Finally, we perform an error analysis, identifying the bottlenecks of language models in lexical relation mining tasks and providing insights for future improvements. Zhiyi Tian, Jingwei Sun 0001, Guangzhong Sun |
ACL (1) | 2 |
| 2025 | Fine-Grained Privacy-Preserving Semantic Communication against Eavesdropping AttacksabstractSemantic communication has emerged as a pivotal technology for future wireless communication, where privacy preservation plays a critical role. However, existing privacy-preserving semantic communication methods have a negative impact on communication performance. In this paper, we focus on fine-grained privacy-preserving semantic communication to achieve a better performance-privacy trade-off. Specifically, we formalize it as an optimization problem and propose a Dual Information Bottleneck (Dual-IB) method to solve it. Firstly, Dual-IB extracts task-relevant information from input and compresses task-irrelevant information. Secondly, it minimizes sensitive attributes within the task-relevant representation by information bottleneck. This mitigates the risks of privacy exposure while maintaining the task performance. We validate the effectiveness of our method using an eavesdropping attack that attempts to extract private information in the transmission. The experimental results demonstrate that our method not only achieves a task performance of approximately 97%, but also reduces eavesdropping accuracy by 68%, effectively addressing a critical limitation of existing approaches. Zhiyi Tian, Weiqi Wang 0003, Shui Yu 0001 |
GLOBECOM | 2 |
| 2025 | IDIR: Interpolated Diffusion Image Reconstruction for Generalizable Detection of Synthetic Images
Weiqi Wang 0003, Zhiyi Tian, Shui Yu 0001 |
PAKDD (4) | 3 |
| 2025 | Inversion Triplet - A Contrastive Backdoor Mitigation Method for Self-Supervised Vision Encoders
Hiep Vo, Zhiyi Tian, Chenhan Zhang, James Xi Zheng, Shui Yu 0001 |
PAKDD (6) | 2 |
| 2025 | Can Self Supervision Rejuvenate Similarity-Based Link Prediction?
Chenhan Zhang, Weiqi Wang 0003, Zhiyi Tian, James Jian Qiao Yu, Mohamed Ali Kâafar, An Liu 0002, Shui Yu 0001 |
PAKDD (7) | 3 |
| 2025 | MOUSSE: A Multimodality-Oriented Unified Semantic Communication System by Contrastive LearningabstractThe sixth generation (6G) communication posed higher requirements for the communication system regarding accurate semantic transmission. The existing studies about semantic communication principally concentrate on tackling task-oriented problems, rather than directly design modality-oriented system which is more generic and adaptive to different tasks. To improve the flexibility and robustness of communication system, we propose a Multimodality-Oriented Unified Semantic Communication SystEm (MOUSSE) based on contrastive learning. MOUSSE is designed to firstly orient modality then matches different modalities combination up to various tasks. Existing task-oriented philosophy primarily considers tasks whilst restricting modal versatility. MOUSSE could also intake and output various tasks with multiple modalities while transmit them in a concise and unified representation. Specifically, the system consists of structure-symmetric twin encoder-decoder for modality unification, which cascades joint source-channel coding (JSCC) module and contrastive learning based alignment module. Finally, the experiments verify the validity of proposed MOUSSE by quantitative results from different modalities with their respective tasks. The reliability and robustness are also improved from the point of entire communication system view. Tao Zhang 0165, Zhiyi Tian, Chenhan Zhang, Shui Yu 0001 |
WCNC | 3 |
| 2025 | TAPE: Tailored Posterior Difference for Auditing of Machine UnlearningabstractWith the increasing prevalence of Web-based platforms handling vast amounts of user data, machine unlearning has emerged as a crucial mechanism to uphold users' right to be forgotten, enabling individuals to request the removal of their specified data from trained models. However, the auditing of machine unlearning processes remains significantly underexplored. Although some existing methods offer unlearning auditing by leveraging backdoors, these backdoor-based approaches are inefficient and impractical, as they necessitate involvement in the initial model training process to embed the backdoors. In this paper, we propose a TAilored Posterior diffErence (TAPE) method to provide unlearning auditing independently of original model training. We observe that the process of machine unlearning inherently introduces changes in the model, which contains information related to the erased data. TAPE leverages unlearning model differences to assess how much information has been removed through the unlearning operation. Firstly, TAPE mimics the unlearned posterior differences by quickly building unlearned shadow models based on first-order influence estimation. Secondly, we train a Reconstructor model to extract and evaluate the private information of the unlearned posterior differences to audit unlearning. Existing privacy reconstructing methods based on posterior differences are only feasible for model updates of a single sample. To enable the reconstruction effective for multi-sample unlearning requests, we propose two strategies, unlearned data perturbation and unlearned influence-based division, to augment the posterior difference. Extensive experimental results indicate the significant superiority of TAPE over the state-of-the-art unlearning verification methods, at least 4.5x efficiency speedup and supporting the auditing for broader unlearning scenarios. Weiqi Wang 0003, Zhiyi Tian, An Liu 0002, Shui Yu 0001 |
WWW | 2 |
| 2025 | CRCGAN: Toward robust feature extraction in finger vein recognition
Zhongxia Zhang, Zhengchun Zhou, Zhiyi Tian |
Pattern Recognit. | 3 |
| 2025 | Backdoored Sample Cleansing for Unlabeled Datasets via Bootstrapped Dual Set PurificationabstractSelf-Supervised Learning (SSL) excels in utilizing unlabeled data for feature representation learning. However, recent studies have revealed that SSL is vulnerable to data poisoning-based backdoor attacks. To remove backdoored samples from the SSL training dataset, model optimization methods often fine-tune a trained model by contrasting the training dataset with a reserved clean dataset. This contrastive training effectively marginalizes backdoored samples from the distribution of benign ones,if and only ifboth the reserved clean dataset and the training dataset are from the same data distribution. However, presuming identical distributions between the web-scraped data and reserved data is impractical. To address this impractical assumption, our proposed Bootstrapped Dual SetPurification (AUTO) method distinguishes backdoored from benign samples by contrasting a mined ‘positive set’ and a mined ‘negative set’ within the training dataset itself. We exploit the resistance of backdoored samples in data mixing to mine a highly poisoned ‘positive set’ and a minimally poisoned ‘negative set’. Besides, AUTO mitigates unstable detection performance within different optimization steps by continuously refining the dual sets by the optimized model, enhancing the model's poison distinguishability from consistently improving supervision signals. Our extensive experiments on Cifar10, Cifar100, and Imagenet100 against existing data poisoning SSL backdoor attacks demonstrate AUTO's superiority in detection performance over all existing defenses. Luoyu Chen, Weiqi Wang 0003, Zhiyi Tian, Chenhan Zhang, Shui Yu 0001 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2025 | CRFU: Compressive Representation Forgetting Against Privacy Leakage on Machine UnlearningabstractMachine unlearning allows data owners to erase the impact of their specified data from trained models. Unfortunately, recent studies have shown that adversaries can recover the erased data, posing serious threats to user privacy. An effective unlearning method removes the information of the specified data from the trained model, resulting in different outputs for the same input before and after unlearning. Adversaries can exploit these output differences to conduct privacy leakage attacks, such as reconstruction and membership inference attacks. However, directly applying traditional defenses to unlearning leads to significant model utility degradation. In this article, we introduce a Compressive Representation Forgetting Unlearning scheme (CRFU), designed to safeguard against privacy leakage on unlearning. CRFU achieves data erasure by minimizing the mutual information between the trained compressive representation (learned through information bottleneck theory) and the erased data, thereby maximizing the distortion of data. This ensures that the model's output contains less information that adversaries can exploit. Furthermore, we introduce a remembering constraint and an unlearning rate to balance the forgetting of erased data with the preservation of previously learned knowledge, thereby reducing accuracy degradation. Theoretical analysis demonstrates that CRFU can effectively defend against privacy leakage attacks. Our experimental results show that CRFU significantly increases the reconstruction mean square error (MSE), achieving a defense effect improvement of approximately 200% against privacy reconstruction attacks with only 1.5% accuracy degradation on MNIST. Weiqi Wang 0003, Chenhan Zhang, Zhiyi Tian, Shushu Liu, Shui Yu 0001 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2025 | FedU: Federated Unlearning via User-Side Influence Approximation ForgettingabstractMachine unlearning has become a significant research topic on a global scale due to the increasing importance of privacy protection, particularly in light of the right to be forgotten legislation. Although many solutions are proposed, the current mainstream centralized machine unlearning studies are not feasible in federated learning (FL), where the server has no access to any users’ unlearning samples. In this paper, we aim to tackle thefederated unlearningproblem by proposing a Federated Unlearning (FedU) scheme via a user-side influence approximation forgetting method, thereby eliminating the need to share raw data with the server. In FedU, only users who have unlearning needs execute the influence approximation forgetting, while other users and the server just conduct the same operations as they did in FL. The proposed influence approximation forgetting method achieves unlearning by estimating the influence of the erased samples relying on only the user's local data and eliminating this influence from the model. However, the model utility is still negatively influenced by directly removing the influence estimation. To mitigate the side effects of unlearning, we propose a utility preservation method that simultaneously trains the unlearned model based on the unlearning requesters’ remaining local dataset. We design an adaptive optimization method to balance the forgetting and utility preservation effectiveness optimally during the unlearning process. Extensive evaluations on three representative public datasets demonstrate that our proposed method significantly outperforms state-of-the-art methods in both effectiveness and efficiency, avoiding more than 3% accuracy degradation when the number of unlearning requesters is large. Weiqi Wang 0003, Chenhan Zhang, Zhiyi Tian, Shui Yu 0001 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2025 | SCU: An Efficient Machine Unlearning Scheme for Deep Learning Enabled Semantic CommunicationsabstractDeep learning (DL) enabled semantic communications leverage DL to train encoders and decoders (codecs) to extract and recover semantic information. However, most semantic training datasets contain personal private information. Such concerns call for enormous requirements for specified data erasure from semantic codecs when previous users hope to move their data from the semantic system. Existing machine unlearning solutions remove data contribution from trained models, yet usually in supervised sole model scenarios. These methods are infeasible in semantic communications that often need to jointly train unsupervised encoders and decoders. In this paper, we investigate the unlearning problem in DL-enabled semantic communications and propose a semantic communication unlearning (SCU) scheme to tackle the problem. SCU includes two key components. Firstly, we customize the joint unlearning method for semantic codecs, including the encoder and decoder, by minimizing mutual information between the learned semantic representation and the erased samples. Secondly, to compensate for semantic model utility degradation caused by unlearning, we propose a contrastive compensation method, which considers the erased data as the negative samples and the remaining data as the positive samples to retrain the unlearned semantic models contrastively. Theoretical analysis and extensive experimental results on three representative datasets demonstrate the effectiveness and efficiency of our proposed methods. Weiqi Wang 0003, Zhiyi Tian, Chenhan Zhang, Shui Yu 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Evaluation of Machine Unlearning Through Model DifferenceabstractIncreasing attention is being paid to machine unlearning, which supports individuals’ “right to be forgotten.” While most studies focus on the efficiency and effectiveness of unlearning algorithms, the evaluation of machine unlearning effectiveness remains underexplored. Offering robust evaluation services for unlearning is critical, not only to uphold privacy legislation but also to assess and improve existing unlearning methods. Lots of existing methods employ backdoor methods to evaluate unlearning effectiveness, which can only verify the unlearning effect of backdoored samples and negatively impact the model utility as they need to embed backdoors into the model first. In this paper, we propose an evaluating machine unlearning (EMU) method, which aims to evaluate the effectiveness of unlearning and verify data removal without the aforementioned adverse effects. Machine unlearning inherently creates a difference on the model before and after unlearning. The model difference contains information about the unlearned samples, which can be extracted through reconstruction models for unlearning effectiveness evaluation. To efficiently generate the model differences as input for evaluation, we simulate the model changes based on the influence function theory. Additionally, we design a multi-task information bottleneck structure to enhance the scalability of EMU and simplify the analysis of different learning tasks. We provide a theoretical analysis of how the similarity between erased and remaining samples, as well as task types, affects the extent of unlearning—factors that have been largely overlooked. Extensive experiments on various model architectures and representative datasets confirm our analysis, demonstrating the effective evaluation for unlearning without any degradation in the service model utility. Weiqi Wang 0003, Chenhan Zhang, Zhiyi Tian, Shui Yu 0001, Zhou Su 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | Unveiling the Unseen: Video Recognition Attacks on Social Software
Hangyu Zhao, Hua Wu 0004, Xuqiong Bian, Guang Cheng 0001, Xiaoyan Hu 0007, Zhiyi Tian |
ACISP (2) | 7 |
| 2024 | From the Perspective of AI Safety: Analyzing the Impact of XAI Performance on Adversarial AttackabstractThe outstanding performance of machine learning models across various fields enables them to be used in sensitive and high-risk activities such as healthcare, automated driving, and security services. However, the explainability of their outputs and the safety of their operation are major concerns. Although Explainable AI (XAI) can enhance the interpretability of AI models, further research is necessary to evaluate its effectiveness in explaining adversarial attacks. The use of XAI techniques and the rise of adversarial attacks are important issues related to the explainability and security of deep learning, respectively. Furthermore, the relation between the explainability and safety of deep learning, such as the vulnerability of the XAI explanation to an adversarial attack, is key to unlocking these concerns. In this paper, we use the Saliency Map as an XAI technique to explain the behavior of the Fast Gradient Sign Method (FGSM) adversarial attack on the ResNet model, and show the vulnerability related to such an explanation with respect to the attack. Extensive experiments show that as the severity of FGSM attack on the ResNet model increases, the Saliency Map gradually fails, exposing its potential vulnerability. Ahmed Asiri, Zhiyi Tian, Shui Yu 0001 |
GLOBECOM | 3 |
| 2024 | Modeling and Analyzing the Spatial-Temporal Propagation of Malware in Mobile Wearable IoT NetworksabstractWearable Internet of Things (IoT) devices are easily compromised by malware due to their security vulnerabilities. The bots infected by malware may continue to infect healthy neighbor devices through wireless communication technology in mobile wearable IoT networks (WIoT). All bots form a botnet which eventually leads to a series of malicious attacks. Therefore, it is necessary to predict the dynamic malware propagation path between wearable devices, which can help provide target immunization measures on devices to prevent the formation of botnets. In this article, we capture the local interaction and spatial–temporal propagation behavior of malware utilizing the individual-based cellular automata (CA) model. First, taking into account the mobility of walking users carrying wearable devices in the actual WIoT, we present a human mobility model called Gauss–Markov truncated Levy walk (GM-TLW) to describe the movement patterns of mobile users. Second, based on the moving coordinates of all wearable devices obtained from the GM-TLW mobility model, we leverage the improved CA propagation model to study the time evolution of the number of bots and the spreading spatial distribution of malware. We compare our propagation model with the differential equation model and traditional CA model, and analyze the impact of various parameters on the dynamics of botnet formation using numerical simulations. Finally, detailed simulation results show that the GM-TLW model is more suitable for realistic human mobility scenarios. In addition, the proposed CA-based model is more precise than the differential equation model to modeling the malware propagation and provides a basis for defenders to adopt the optimal malware control strategies. Jie Dou, Gang Xie 0001, Zhiyi Tian, Lei Cui 0006, Shui Yu 0001 |
IEEE Internet Things J. | 3 |
| 2024 | The Role of Class Information in Model Inversion Attacks Against Image Deep Learning ClassifiersabstractModel inversion attacks can reconstruct the training samples of victim deep learning models. The existing efforts heavily rely on auxiliary information of the target samples (prior target information) to achieve their adversarial goals. However, prior target information is hard to obtain in practice. In this paper, we explore the effect of class information in model inversion attacks to reduce the reliance of prior target information. Our contributions on class information exploitation are two-fold. Firstly, we propose a supervised inversion model, Supervised Model Inversion (SMI). The proposed inversion model learns pixel-level features and data-to-class features from the rounded-outputs of the victim model and labeled auxiliary dataset. Secondly, we leverage victim model's rounded-outputs to guide the optimization of reconstructing inversion samples after trained inversion model. Our experimental results show that inversion samples reconstructed by SMI are more visually plausible with more details, comparing to the three representative model inversion attacks. We further perform an extensive study on various auxiliary dataset settings. It is found that the class combination in the auxiliary dataset rather than the number of classes that determines the quality of inversion samples. The ground-truth labels can improve the qualities of inversion samples but not essential to inversion attacks. Zhiyi Tian, Lei Cui 0006, Chenhan Zhang, Shuaishuai Tan, Shui Yu 0001, Yonghong Tian 0001 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2024 | Machine Unlearning via Representation Forgetting With Parameter Self-SharingabstractMachine unlearning enables data owners to remove the contribution of their specified samples from trained models. However, existing methods fail to strike an optimal balance between erasure effectiveness and model utility preservation. Previous studies focused on removing the impact of user-specified data from the model as much as possible to implement unlearning. These methods usually result in significant model utility degradation, commonly called catastrophic unlearning. To address the issue, we systematically consider machine unlearning and formulate it as a two-objective optimization problem that involves forgetting the erased data and retaining the previously learned knowledge, highlighting accuracy preservation during the unlearning process. We propose an unlearning method called representation-forgetting unlearning with parameter self-sharing (RFU-SS) to achieve the two-objective unlearning goal. Firstly, we design a representation-forgetting unlearning (RFU) method that aims to remove the contribution of specified samples from a trained representation by minimizing the mutual information between the representation and the erased data. The representation is learned using the information bottleneck (IB) method. RFU is tailored to the IB structure models for ease of introduction. Secondly, we customize a parameter self-sharing structural optimization method for RFU (i.e., RFU-SS) to simultaneously optimize the forgetting and retention objectives to find the optimal balance. Extensive experimental results demonstrate a significant effectiveness improvement of RFU-SS over the state-of-the-art methods. RFU-SS almost eliminates catastrophic unlearning, reducing model accuracy degradation from over 6% to less than 0.2% on the MNIST dataset with an even better removal effect. The source code is available athttps://github.com/wwq5-code/RFU-SS.git. Weiqi Wang 0003, Chenhan Zhang, Zhiyi Tian, Shui Yu 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | Forgetting and Remembering Are Both You Need: Balanced Graph Structure UnlearningabstractIn light of the growing emphasis on the right to be forgotten of graph data, machine unlearning has been extended to unlearn the graph structures’ knowledge from graph neural networks (GNNs), namely, structure unlearning. Whereas the complex dependencies in graph data, structure unlearning is intrinsically prone to imbalanced performance between the objectives of knowledge forgetting and model utility maintenance. Nevertheless, most existing methods fall short in addressing the two objectives in tandem and developing balanced solutions. In this paper, we propose imbalanced Structure Unlearning Mitigation using MultI-objective OpTimization (SUMMIT), which aims to develop balanced solutions regarding both knowledge forgetting and model utility maintenance effects. Corresponding to the two aspects, we first construct two tailored objectives that specifically address the challenges inherent in structure unlearning. Specifically, for the forgetting objective, we introduce a higher-order forgetting enhancement strategy aimed at mitigating the adverse effects of GNN oversmoothing on node decoupling. For the remembering objective, we adhere to the principle of ideal unlearning and propose to minimize the distributional distance between the node embeddings developed by unlearned and well-trained GNNs. Considering the potential competitive relationship between the two objectives during the optimization process, we present an adaptive two-objective balancer based on multi-objective optimization to reconcile the two objectives and strike a balance between them. We conduct comprehensive experiments to evaluate the efficacy of SUMMIT on three representative GNNs and four datasets, and compare the performance of SUMMIT with its ablation variants and a cadre of baselines. We demonstrate the superiority of SUMMIT in its ability to yield optimal and balanced solutions, addressing both the facets of knowledge forgetting and model utility maintenance. Chenhan Zhang, Weiqi Wang 0003, Zhiyi Tian, Shui Yu 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2023 | BFU: Bayesian Federated Unlearning with Parameter Self-SharingabstractAs the right to be forgotten has been legislated worldwide, many studies attempt to design machine unlearning mechanisms to enable data erasure from a trained model. Existing machine unlearning studies focus on centralized learning, where the server can access all users’ data. However, in a popular scenario, federated learning (FL), the server cannot access users’ training data. In this paper, we investigate the problem of machine unlearning in FL. We formalize a federated unlearning problem and propose a bayesian federated unlearning (BFU) approach to implement unlearning for a trained FL model without sharing raw data with the server. Specifically, we first introduce an unlearning rate in BFU to balance the trade-off between forgetting the erased data and remembering the original global model, making it adaptive to different unlearning tasks. Then, to mitigate accuracy degradation caused by unlearning, we propose BFU with parameter self-sharing (BFU-SS). BFU-SS considers data erasure and maintaining learning accuracy as two tasks and optimizes them together during unlearning. Extensive comparisons between our methods and the state-of-art federated unlearning method demonstrate the superiority of our proposed realizations. Weiqi Wang 0003, Zhiyi Tian, Chenhan Zhang, An Liu 0002, Shui Yu 0001 |
AsiaCCS | 2 |
| 2023 | Construct New Graphs Using Information Bottleneck Against Property Inference AttacksabstractGraphs provide a unique representation of real- world data. However, recent studies found that inference attacks can extract private property information of graph data from trained graph neural networks (GNNs), which arouses privacy concerns about graph data, especially in collaborative learning systems where model information is more accessible. While there has been a few research efforts on the property inference attacks against GNNs, how to defend against such attacks has seldom been studied. In this paper, we propose to leverage the information bottleneck (IB) principle to defend against the property inference attacks. Particularly, we involve a threat model, where the attacker can extract graph property from the graph embedding developed by GNNs. To defend against the attacks, we use IB to construct new graph structures from the original graphs. The change in graph structures enables the new graphs to contain less information related to the property information of the original graphs, making it harder for attackers to infer property information of the original graphs from the graph embeddings. Meantime, the IB principle enables task-relevant information to be sufficiently contained in the new graph, enabling GNNs to develop accurate predictions. The experimental results demonstrate the efficacy of the proposed approach in resisting property inference attacks and developing accurate predictions. Chenhan Zhang, Zhiyi Tian, James Jian Qiao Yu, Shui Yu 0001 |
ICC | 2 |
| 2022 | Utility-Aware Privacy-Preserving Federated Learning through Information BottleneckabstractFederated learning (FL) as a privacy-preserving machine learning (ML) algorithm provides an efficient distributed training paradigm. Existing FL frameworks still suffer from privacy leakage hazards such as membership inference attacks. The current popular defense approaches are mainly based on differential privacy (DP) strategy. However, privacy preservation is undertaken with an inevitable loss of model utility in DP. As a result, it performs miserably in practical deployments. To solve this problem, we modify the FL framework through the information bottleneck (IB) method to attain a trade-off between privacy protection and model utility. Firstly, we adapt the training process on client side by applying IB in the local training. It is intended to squeeze out privacy through the bottleneck. Secondly, we further modify the training process on server side. A validation process is used to evaluate whether the IB-based local training is squeezing out privacy. Clients that extrude the right information will occupy an important place in aggregation phase. Extensive experiments on classic datasets demonstrate the superiority of the proposed scheme in terms of privacy preservation and model utility. Shaolong Guo, Zhou Su 0001, Zhiyi Tian, Shui Yu 0001 |
TrustCom | 3 |
| 2022 | Sneaking Through Security: Mutating Live Network Traffic to Evade Learning-Based NIDSabstractMachine learning based network intrusion system (NIDS) is known to be vulnerable to evasions. Attackers conceal intrusion activities to make them undetected. Researching evasion techniques contributes to evaluating and increasing the robustness of NIDS. Previous evasion approaches modify feature values or packets of an offline network trace as a whole. However, in real scenarios, attackers are constrained to manipulate only outbound packets on the fly. To bridge this assumption gap, we present the first evasion solution for live network traffic against learning based NIDSs. The solution consists of three components: a devised Kalman filter based algorithm to predicate the feature values of live flows, a set of formally constructed atomic packet mutation operators, and a proposed Strength Enhanced Deep Q-learning (SE-DQN) to determine effective mutation operators on outbound packets according to the predicted features. A defense scheme based on adaptive decision threshold adjustment is also provided. Experimental evaluation is presented on various NIDS classifiers and cyber attacks. Results show that SE-DQN achieves an evasion rate of at least 64.2% on most classifiers and even more than 90% on certain ones, and it is three times faster than DQN on learning mutation policy. The defense scheme shows an improvement of at least 76.4% on recall measurement. Shuaishuai Tan, Xiaoxiong Zhong, Zhiyi Tian, Qingkuan Dong |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2021 | A Novel Android Malware Detection Method Based on Visible User InterfaceabstractMachine learning has been increasingly adopted to detect Android malwares. Most existing studies depend on features in code space such as information flows and API calls. Malware variants would engage these models in a never-ending war. Inspired by the observation that some variants share similar or even identical user interfaces (UIs), this paper explores employing visible UI screenshot as the indicator to build a novel Android malware detection method. To achieve this vision, we built the first Android Application Screenshot Dataset (AnASD) consisting of more than twenty thousand UI screenshots produced by both benign applications and malwares. A thorough analysis was conducted to characterize the dataset, especially the UI difference between benign applications and malwares. Then a set of state of the art deep learning classifiers on AnASD were trained and evaluated. The results of both sim-ilarity measurement and classification performance proved the feasibility to detect Android malwares based on user interfaces. To facilitate the research community, the dataset is free available at https://doi.org/10.6084/m9.figshare.14445768. Shuaishuai Tan, Zhiyi Tian, Xiaoxiong Zhong, Shui Yu 0001, Weizhe Zhang, Guozhong Dong |
TrustCom | 2 |