VLDB 2026 Research / reviewers in the wild / expert
Gaoyang Liu
dblp:133/3350
· DBLP profile ↗
34ranked-venue papers
12as first author
26since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 10 · 6 first-author · 10 since 2021Computer networks · 8 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 4 since 2021Systems, architecture and hardware · 4 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Breaking the Boundary Barrier: Robust Model Fingerprinting via Unlearnable Examples in Model-Parameter SpaceabstractDeep learning models represent valuable intellectual property due to their high development costs. To protect model ownership, existing fingerprinting techniques have been proposed to use adversarial examples to fingerprint a model's decision boundaries. However, these fingerprints are inherently fragile, as model decision boundaries are highly sensitive to common model modifications such as fine-tuning, pruning, and adversarial training. In this paper, we propose MFUE (Model Fingerprinting via Unlearnable Examples), a novel fingerprinting methodology that leverages the stable unlearnability of unlearnable examples to fingerprint arbitrary modified models in parameter space, fundamentally circumventing the inherent vulnerability of decision boundaries. To achieve robust model fingerprinting in parameter space, we are the first to identify that unlearnable examples, owing to their persistent training resistance, can serve as stable fingerprints beyond the model's decision boundaries. To endow unlearnable examples with robustness against arbitrary model modifications, we introduce adversarial training that simulates the randomness of model modifications by jointly optimizing the unlearnable examples over models at different training stages. We evaluate the performance of MFUE against six different attack types, including both model and input tampering. Through extensive experiments, we demonstrate that MFUE outperforms four existing methods in terms of robustness and uniqueness. Tianlong Xu, Zixiong Wang, Gaoyang Liu, Jian Chen 0046, Ahmed M. Abdelmoniem, Chen Wang 0011 |
KDD (1) | 3 |
| 2025 | Prototype Surgery: Tailoring Neural Prototypes via Soft Labels for Efficient Machine UnlearningabstractThe rapid advancements and widespread application of deep neural networks (DNNs), coupled with their reliance on sensitive and private data, have sparked growing concerns regarding data privacy and the ''right to be forgotten''. To address these concerns, machine unlearning has been proposed to efficiently eliminate the influence of specific training data from trained DNNs. However, existing machine unlearning methods struggle with the large number of parameters in trained DNNs, which lead to slow execution and high memory consumption, making them impractical for large-scale models. In this paper, we shift our focus to the small set of weights in the final classification layer of DNNs, which are defined as as ''prototypes'' for different classes. Our key observation is that the prototype associated with the unlearned training data undergoes a significant shift, whereas prototypes of unrelated classes exhibit only minor changes when comparing the prototypes of original and retrained models. Based on this observation, we propose a novel machine unlearning approach that efficiently achieves machine unlearning by directly adjusting the prototypes of DNNs. We first introduce Naive Prototype Surgery (Naive PS), a fast and simplified method that uses a closed-form solution to approximate unlearning effect by directly adjusting the prototype associated with the unlearned data. Next, we propose Prototype Surgery (PS), which incorporates soft label information to fine-tune the prototypes of all classes, to achieve a more effective unlearning. Both methods achieve data unlearning by only modifying the prototypes in the DNNs, thus avoiding the challenges posed by the large number of model parameters. Extensive experiments on four datasets demonstrate that our methods significantly accelerate the unlearning process while achieving comparable results to five existing methods in terms of both unlearning performance and privacy guarantee. Gaoyang Liu, Xijie Wang, Zixiong Wang, Chen Wang 0011, Ahmed M. Abdelmoniem, Desheng Wang 0001 |
CCS | 1 |
| 2025 | Generating Is Believing: Membership Inference Attacks against Retrieval-Augmented GenerationabstractRetrieval-Augmented Generation (RAG) is a state-of-the-art technique that mitigates issues such as hallucinations and knowledge staleness in Large Language Models (LLMs) by retrieving relevant knowledge from an external database to assist in content generation. Existing research has demonstrated potential privacy risks associated with the LLMs of RAG. However, the privacy risks posed by the integration of an external database, which often contains sensitive data such as medical records or personal identities, have remained largely unexplored. In this paper, we aim to bridge this gap by focusing on membership privacy of RAG’s external database, with the aim of determining whether a given sample is part of the RAG’s database. Our basic idea is that if a sample is in the external database, it will exhibit a high degree of semantic similarity to the text generated by the RAG system. We present S2MIA, a Membership Inference Attack that utilizes the Semantic Similarity between a given sample and the content generated by the RAG system. With our proposed S2MIA, we demonstrate the potential to breach the membership privacy of the RAG database. Extensive experimental results demonstrate that S2MIA outperforms five existing MIAs, even when the system is protected by three representative defenses. Gaoyang Liu, Chen Wang 0011, Yang Yang 0060 |
ICASSP | 2 |
| 2025 | From Expansion to Retraction: Long-tailed Machine Unlearning via Boundary ManipulationabstractMachine unlearning aims to remove the information of specific data from a trained machine learning model while retaining its utility for the remaining data, so as to meet the requirements of privacy regulations. Existing unlearning methods often assume a balanced data distribution, but neglect the real-world, long-tailed scenarios, where the decision boundaries of tail classes are frequently distorted due to insufficient sample representation, thereby reducing the unlearning efficacy. In this paper, we propose the first Long-Tailed Machine Unlearning (LTMU) framework from a unified decision-boundary perspective. Our framework begins with a directional boundary repair scheme designed to enrich the distorted decision boundary of the tail class, and then develop a novel boundary retraction approach tailored for long-tailed unlearning, dispersing both the augmented and original features throughout the feature space. This bidirectional manipulation not only offers a unified interpretation of the relationship between long-tailed learning and unlearning, but also enables flexible control over both repair and unlearning processes through the generation of augmented features, thereby effectively accomplishing the long-tailed unlearning task. Extensive experiments across multiple datasets and neural network architectures demonstrate the effectiveness of our framework in achieving complete unlearning of tail classes in long-tailed distributions. Weizhuo Gao, Chen Wang 0011, Gaoyang Liu, Ahmed M. Abdelmoniem, Kai Peng 0001 |
KDD (2) | 4 |
| 2025 | Freehand Sketch-Based 3D Reconstruction with Contour Constraints via Elastic MetricsabstractSketch-based 3D reconstruction enables intuitive content creation through freehand drawings, yet generating high-fidelity 3D models from geometrically ambiguous, structurally simplified, and sparse sketches remain challenging. To overcome existing methods’ limitations in sketch-style generalization, contour accuracy, and suboptimal texture effects, we propose an end-to-end framework that generates textured 3D models directly from a single freehand sketch and semantic labels. To address the scarcity of paired freehand sketch training data, we introduce a 3D model-based automated sketch generation method for extracting mesh contours via a 3D mesh-to-sketch pipeline and synthesizing freehand-style sketches employing a Transformer-based stroke generator to construct a paired dataset of hand-drawn sketches and 3D models. Meanwhile, we design a contour constraint mechanism that jointly optimizes projection-space Chamfer distances and elastic metrics, significantly enhancing the reconstruction accuracy of complex geometries. Furthermore, we integrate a semantic-guided texture generation module using Text2Tex with depth-aware diffusion models and dynamic view-optimization strategies, achieving a complete geometry-appearance integrated modeling pipeline. Finally, extensive experimental results demonstrate that our method outperforms existing structural reconstruction and texture synthesis approaches, exhibiting strong generalization capabilities and practical applicability. Gaoyang Liu, Chunyang Huo, Zhentong Xu, Junli Zhao, Yishan Dong, Baodong Wang, Xinbin Sun |
VRST | 1 |
| 2025 | A real-time welding defect detection framework based on RT-DETR deep neural networkabstract• A real-time welding defect detection framework based on the RT-DETR deep neural network is proposed. • By exploring three different data augmentation strategies, the most suitable method for weld defect detection is selected. • The proposed method outperforms other models in terms of both mAP and FPS. • The proposed method is validated in the WAAM weld defect detection. The quality of welds is critical to the safety and reliability of steel structure connections, underscoring the importance of accurate inspection during the welding process. To enhance inspection effectiveness, deep learning methods have gained popularity in weld defect detection for their ability to automatically learn and refine image features. However, the complex multi-stage training and inference process of these methods often fails to meet the requirements of real-time performance and accuracy. To address this problem, a framework based on the Real-Time DEtection TRansformer (RT-DETR) for deep learning-based welding defect detection is proposed. This framework improves the Transformer backbone by eliminating the most time-consuming non-maximum suppression (NMS) step, achieving real-time detection without sacrificing accuracy. A diverse welding dataset with 1,134 images from real-world manufacturing and construction environments was developed for model training and validation. In addition, three data enhancement algorithms were explored to enhance the model’s generalization ability. The model achieved detection accuracy scores of [email protected] at 0.996 and [email protected]:0.95 at 0.801, with a detection speed of 67 frames per second (FPS). Compared to the previous Faster R-CNN, SSD, YOLOv5, YOLOv11 and DETR models, the proposed RT-DETR model demonstrates superior efficiency and accuracy. The proposed framework was further validated in the on-site inspections of metal additive manufacturing, and the results confirmed that the RT-DETR-based model meets the stringent requirements for real-time inspection in metal additive manufacturing. Gaoyang Liu, Duanrui Yang, Hongjia Lu |
Adv. Eng. Informatics | 1 |
| 2025 | Poisoning as a Post-Protection: Mitigating Membership Privacy Leakage From Gradient and Prediction of Federated ModelsabstractFederated learning (FL) is a distributed learning paradigm that enables multiple clients to train a unified model without sharing their private data. However, recent works demonstrate that FL models are vulnerable to membership inference attacks (MIAs), which can infer whether a data sample was used to train a given FL model. Existing countermeasures either require far-reaching modifications of FL training process or enforce extra processing in prediction phase, yielding them unlikely to be applied well in practice. In this paper, we design a post-protection mechanism, dubbedP$^{2}$-Protection, which degrades the inference performance of MIAs by simultaneously poisoning the prediction and gradient of the target FL model to reduce the privacy leakage of training data while keeping the model prediction accuracy.P$^{2}$-Protectiononly involves one additional training round to embed the poisoned prediction and gradient into the target FL model, without requiring model retraining or training process modification. We evaluateP$^{2}$-Protectionand compare it with two state-of-the-art defenses against three MIAs on five realistic datasets. Experimental results show thatP$^{2}$-Protectionoutperforms the existing defenses by offering limited implement overhead and improved utility-privacy trade-off. Gaoyang Liu, Tianlong Xu, Yang Yang 0060, Ahmed M. Abdelmoniem, Chen Wang 0011, Jiangchuan Liu |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2025 | LLMGraph: Label-Free Detection Against APTs in Edge Networks via LLM and GCNabstractIn the growing trend of remote working, millions of edge networks (e.g., homes or branch offices) are increasingly threatened by Advanced Persistent Threats (APTs), because of the weakened segmentation between business and non-business devices in remote working environment. Despite the fact that numerous APT detection mechanisms have been proposed, all of them are struggling to handle thecomplex structure, themassive scaleand thediverse topologyof edge networks.Can recent machine-learning advances tackle these APT detection pain points in edge networks?The GNNs (Graph Neural Networks) seems to be suited to capture thecomplex structure, but its adjacency matrix fails to capture key network flow context. Additionally, GNNs require extensive manual labeling, which is not scalable. LLMs (Large Language Models) have the potential to provide automatic labeling for the GNNs, but they lack the supplementary security context needed for effective labeling. To address these gaps, we presentLLMGraph, which incorporates extended GCNs (Graph Convolutional Networks) and domain-specific RAG (Retrieval-Augmented Generation) pipeline to achieve label-free detection against APTs in edge networks.LLMGraph's extended GCNs model can capture network flow context and direction.LLMGraph's domain-specific RAG pipeline can supplement key security contexts, including device vulnerability and network flow, for effective labeling. Additionally,LLMGraphprovides an LLM aggregator to augment and merge thediverse topologyof the edge networks. Compared to the state-of-the-art mechanisms,LLMGraphproveseffectiveandscalable, improving the F1-score by at least 46.9%, and the training time for 1 million edge networks is within 1000 s. Tianlong Yu, Gaoyang Liu, Chen Wang 0011, Yang Yang 0060 |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2025 | Unlearning Attacks for Regression LearningabstractRecently, the machine unlearning has emerged as a popular method for efficiently erasing the impact of personal data in machine learning (ML) models upon the data owner's removal request. However, few studies take into consideration the security concerns that may exist in the unlearning process. In this article, we propose the first unlearning attack dubbed unlearning attack for regression learning (UnAR) to deliberately influence the predictive behavior of the target sample against regression learning models. The central concept of UnAR revolves around misleading the regression model into erasing the information associated with the influential samples for the target sample. Observing that the influential samples for target data are generally located far away from the regression plane, we thus propose two novel methods, known as influential sample selection (ISS) and influential sample unlearning (ISU), to identify and subsequently eliminate the lineage of the influential samples. By doing so, we can substantially introduce bias into the prediction pertaining to the target sample, yielding the deliberate manipulation for the user adversely. We extensively evaluate UnAR on five public datasets, and the experimental results indicate our attacks can achieve prediction deviations over 35% by unlearning only 0.5% data as the influential samples. Jian Chen 0046, Wenlong Shi, Wanyu Lin, Chen Wang 0011, Wei Liu 0004, Hailong Sun 0001, Gaoyang Liu |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2024 | Few-Shot Camouflaged Object SegmentationabstractIn the domain of computer vision, Camouflaged Object Segmentation (COS) is a crucial task aimed at identifying objects that blend into their surroundings, with applications spanning diverse sectors such as military, medical, and beyond. Traditional COS techniques, which primarily depend on supervised learning, necessitate large-scale labeled datasets. However, the acquisition of sufficient camouflage images for such purposes is often constrained due to their scarcity and the high cost of manual annotation. Additionally, these conventional methods frequently struggle to generalize to novel, unseen classes. In response to these challenges, this paper proposes the Camouflage Few-Shot (CAMFS) framework, an innovative approach integrating few-shot learning into COS. The CAMFS framework comprises two main components: the Camouflaged-Meta module, which converts the semantic information of camouflaged objects into compact feature vectors to facilitate knowledge transfer from support to query images; and the Camouflaged-Base module, focused on refining edge detection and enhancing the contrast between foreground and background elements. To overcome the limitations of existing COS datasets, which are primarily designed for supervised learning, we have developed the COS-FSS dataset, the first public few-shot COS dataset. It is based on the COD10K dataset and supplemented with approximately 3200 additional camouflage images. We conducted extensive evaluations of our CAMFS framework on the COS-FSS dataset. Compared to existing COS models, CAMFS demonstrates an average improvement of 5.7% in Sαand 14.02% in $F_\beta ^w$, while against few-shot segmentation models, it achieves a 5.97% increase in m-IoU. The dataset and additional resources are available at https://github.com/CAM-FSS/FSS-COD. Ziqiu Wang, Yang Yang 0060, Gaoyang Liu |
IJCNN | 5 |
| 2024 | A Continuous Verification Mechanism for Clients in Federated Unlearning to Defend the Right to be ForgottenabstractIn Federated Learning (FL), the regulatory need for the "right to be forgotten" requires efficient Federated Unlearning (FU) methods, which enable FL models to unlearn appointed training data. Associating with the emergence FU, verifying the performance of FU plays a critical role in evaluating the consistency FU methods, in case of the unexpected degradation of the FL model. Though well developed, none of the existing verification methods in FU stands for the clients who opt out of the FL process, which is a universal demand in FL. More specifically, after the clients quit the FL cooperation, they can no longer verify whether the FL model unlearns their data after the FL keeps training for several rounds. To this end, we introduce a continuous verification mechanism for FL clients, called Backdoor Attack-based Forgetting Verification (BAFV). Inspired by backdoor attack, BAFV embedded a persistent mark for the client that proposes to leave, with the intention that the client still has the right to verify of FU after leaving the FL cooperation for a relatively long period. Extensive experiments across diverse FU environments and datasets demonstrated that our method maintains the accuracy of model and provides clients with a continuous verification mechanism to defend their rights. Our code of BAFV is publicly available at: https://github.com/paper-liu/BAFV-master.git. Yang Yang 0060, Gaoyang Liu, Chen Wang 0011 |
ISPA | 6 |
| 2024 | A Contribution Assessment Method Based on Model Performance Gains in Federated LearningabstractFederated Learning is a distributed model training system that uses the computing resources and private data of participants to collaboratively train Machine Learning models. In order to incentivize data holders to actively participate in Federated Learning, a crucial issue is how to fairly assess each participant’s contribution. One class of the mainstream methods estimates client contributions by calculating the correlation between local and global model parameters, substantially reducing the reliance on large computing resources and test datasets. However, these methods overlook an implicit key factor, leading to inaccuracy in assessing the contribution of the participants. We propose FedCA, which provides a fairness assessment of the client’s contribution based on the performance gain of the global model in each communication round. The experimental results demonstrate that FedCA accurately identifies the contributions of clients with varying data qualities and effectively approximates the optimal fairness method of Shapley values. The source code for this paper is available at https://github.com/paper-liu/FedCA-master.git. Yuyin Li, Gaoyang Liu, Yang Yang 0060 |
ISPA | 4 |
| 2024 | United We Stand, Divided We Fall: Fingerprinting Deep Neural Networks via Adversarial TrajectoriesabstractIn recent years, deep neural networks (DNNs) have witnessed extensive applications, and protecting their intellectual property (IP) is thus crucial. As a non-invasive way for model IP protection, model fingerprinting has become popular. However, existing single-point based fingerprinting methods are highly sensitive to the changes in the decision boundary, and may suffer from the misjudgment of the resemblance of sparse fingerprinting, yielding high false positives of innocent models. In this paper, we propose ADV-TRA, a more robust fingerprinting scheme that utilizes adversarial trajectories to verify the ownership of DNN models. Benefited from the intrinsic progressively adversarial level, the trajectory is capable of tolerating greater degree of alteration in decision boundaries. We further design novel schemes to generate a surface trajectory that involves a series of fixed-length trajectories with dynamically adjusted step sizes. Such a design enables a more unique and reliable fingerprinting with relatively low querying costs. Experiments on three datasets against four types of removal attacks show that ADV-TRA exhibits superior performance in distinguishing between infringing and innocent models, outperforming the state-of-the-art comparisons. Tianlong Xu, Chen Wang 0011, Gaoyang Liu, Yang Yang 0060, Kai Peng 0001, Wei Liu 0004 |
NeurIPS | 3 |
| 2024 | Manipulating Pre-Trained Encoder for Targeted Poisoning Attacks in Contrastive LearningabstractIn recent years, contrastive learning has become very powerful for representation learning using large-scale unlabeled data, by involving pre-trained encoders to fine-tune downstream classifiers. However, the latest research indicates that contrastive learning can potentially suffer from the risks of data poisoning attacks, where the attacker injects maliciously crafted poisoned samples into the unlabeled pre-training data. To step forward, in this paper, we present a more stealthy poisoning attack dubbed PA-CL to directly poison the pre-trained encoder, such that the downstream classifier’s behavior on a single target instance to the attacker-desired class can be manipulated without affecting the overall downstream classification performance. We observe that a high similarity exists between the feature representation generated by the poisoned pre-trained encoder for the target sample and samples from the attacker-desired class. This leads to the downstream classifier misclassifying the target sample with the attacker-desired class. Therefore, we formulate our attack as an optimization problem, and design two novel loss functions, namely, the target effectiveness loss to effectively poison the pre-trained encoder, and the model utility loss to maintain the downstream classification performance. Experimental results on four real-world datasets demonstrate that the attack success rate of the proposed attack is 40% higher on average than that of the three baseline attacks, and the fluctuation of the downstream classifier’s prediction accuracy is within 5%. Jian Chen 0046, Gaoyang Liu, Ahmed M. Abdelmoniem, Chen Wang 0011 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | Gradient-Leaks: Enabling Black-Box Membership Inference Attacks Against Machine Learning ModelsabstractMachine Learning (ML) techniques have been applied to many real-world applications to perform a wide range of tasks. In practice, ML models are typically deployed as the black-box APIs to protect the model owner’s benefits and/or defend against various privacy attacks. In this paper, we present Gradient-Leaks as the first evidence showcasing the possibility of performing membership inference attacks (MIAs), with mere black-box access, which aim to determine whether a data record was utilized to train a given target ML model or not. The key idea of Gradient-Leaks is to construct a local ML model around the given record which locally approximates the target model’s prediction behavior. By extracting the membership information of the given record from the gradient of the substituted local model using an intentionally modified autoencoder, Gradient-Leaks can thus breach the membership privacy of the target model’s training data in an unsupervised manner, without any priori knowledge about the target model’s internals or its training data. Extensive experiments on different types of ML models with real-world datasets have shown that Gradient-Leaks can achieve a better performance compared with state-of-the-art attacks. Gaoyang Liu, Tianlong Xu, Rui Zhang 0066, Zixiong Wang, Chen Wang 0011, Ling Liu 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2024 | EgoMUIL: Enhancing Spatio-Temporal User Identity Linkage in Location-Based Social Networks With Ego-Mo HypergraphabstractUsers tend to own multiple accounts on different location-based social network (LBSN) platforms, and they typically engage with diverse social circles on each platform within the same locations. Consequently, linking these accounts across separate networks becomes essential, playing a critical role in information fusion. Previous works accomplishing user identity linkage (UIL) utilize individual mobility records, which are significantly affected by the issue of data scarcity. In this paper, we propose EgoMUIL, a heterogeneous graph embedding approach specifically devised for information propagation, aiming to alleviate the scarcity problem to some extent. Considering that follow relations of respective networks also hold great significance for the UIL task, we are inspired to enrich individual limited mobility records through follow relations. Our preliminary research reveals that direct common follow relations are quite insufficient. Since the followers with the same spatio-temporal mode tend to have social connections, we first mine closely-related users for each user through topology and locality similarity, generating respective cross-domain ego-networks. Subsequently, we construct a heterogeneous ego-mo hypergraph consisting of mobility and ego-networks. We propose a novel graph convolutional network (GCN)-based approach to learn user representations, which enables the aggregation of information from surrounding nodes, incorporating topological similarities, stay locality similarities, and co-occurrence frequencies. The resulting embeddings provide comprehensive representations of users and locations, capturing their characteristics and relationships across platforms, which further facilitates the UIL task. Our experimental results on real-world check-in datasets from Foursquare and Twitter demonstrate that EgoMUIL outperforms the state-of-the-art methods on the UIL task. Notably, EgoMUIL exhibits superior performance in scenarios involving limited check-in records and follow relations. Haojun Huang, Fengxiang Ding, Gaoyang Liu, Chen Wang 0011, Dapeng Oliver Wu |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | Accurate Prediction of Network Distance via Federated Deep Reinforcement LearningabstractA large number of distributed applications necessitate accurate network distance, for example, in the form of delay or latency, to ensure the Quality of Service (QoS). Due to high network measurement overhead and severe traffic congestion, network distance prediction has been introduced, instead of direct network measurements, to infer the unknown network distance with the partial measurements. However, most existing efforts neglect to fully capitalize on the potential latent factors, such as spatial correlations, long-existing temporal results and multi-rule exploration fusion, to achieve better accuracies with quicker convergence. To fill this gap, in this paper, we propose an Accurate Prediction of Network Distance (APND) solution via Federated Deep Reinforcement Learning (FDRL), which has four novel features distinguishing from the previous work. Firstly, a local feature-based matrix with low rank is established in each network cluster, referring to a set of neighbor nodes, to represent the potential spatial correlations among reachable node-pairs. Secondly, the parallel FDRL-based matrix factorization with multi-rule exploration fusion is introduced into APND and executed in all local clusters to minimize prediction errors and accelerate learning convergence. Thirdly, the long-existing learning experience is designed for local model training via Deep Reinforcement Learning (DRL) with rapid convergence. Fourthly, following the real-world routing paths, the cross-domain network nodes are simultaneously classified into adjacent clusters, built on the spatial correlations among them, and their coordinates will be further refined with error-based and average-based policies. Extensive experiments built on available real-world datasets illustrate that APND can accurately predict network distance compared with state-of-the-art approaches at the moderate computing cost. Haojun Huang, Yiming Cai, Geyong Min, Haozhe Wang 0001, Gaoyang Liu, Dapeng Oliver Wu |
IEEE/ACM Trans. Netw. | 5 |
| 2023 | Boundary Unlearning: Rapid Forgetting of Deep Networks via Shifting the Decision BoundaryabstractThe practical needs of the “right to be forgotten” and poisoned data removal call for efficient machine unlearning techniques, which enable machine learning models to unlearn, or to forget a fraction of training data and its lineage. Recent studies on machine unlearning for deep neural networks (DNNs) attempt to destroy the influence of the forgetting data by scrubbing the model parameters. However, it is prohibitively expensive due to the large dimension of the parameter space. In this paper, we refocus our attention from the parameter space to the decision space of the DNN model, and propose Boundary Unlearning, a rapid yet effective way to unlearn an entire class from a trained DNN model. The key idea is to shift the decision boundary of the original DNN model to imitate the decision behavior of the model retrained from scratch. We develop two novel boundary shift methods, namely Boundary Shrink and Boundary Expanding, both of which can rapidly achieve the utility and privacy guarantees. We extensively evaluate Boundary Unlearning on CIFAR-10 and Vggface2 datasets, and the results show that Boundary Unlearning can effectively forget the forgetting class on image classification and face recognition tasks, with an expected speed-up of 17x and 19x, respectively, compared with retraining from the scratch. Weizhuo Gao, Gaoyang Liu, Kai Peng 0001, Chen Wang 0011 |
CVPR | 3 |
| 2023 | AIoT-Empowered Smart Grid Energy Management with Distributed Control and Non-Intrusive Load MonitoringabstractToday's electrical grid is experiencing a fast transition toward a smart infrastructure. Modern smart grid is expected to integrate Artificial Intelligence of Things (AIoT)-empowered energy management systems (EMS) to sense, analyze, and optimize the power consumption and QoS of diverse end users. Non-Intrusive Load Monitoring (NILM) plays a key role in this transition, particularly considering that many legacy devices/appliances may not have built-in sensors. Yet most of the NILM solutions rely on large (often impractical) datasets for training. In this paper, we address this challenge through a meta learning-inspired approach, which implements a hierarchical architecture with a “meta-learner” to supervise the training of each appliance. Current EMS also relies on a central controller to access long-term information across all participants, which mismatches their distributed nature, and so often with slow responses. To this end, we develop a deep reinforcement learning based controller to make dynamic decisions for each component in the system. The experiment results based on real-world data sets and simulation data show that applying the meta learning approach can greatly improve the performance of NILM and the QoS of the whole system. Linfeng Shen, Feng Wang 0001, Miao Zhang 0003, Jiangchuan Liu, Gaoyang Liu, Xiaoyi Fan 0001 |
IWQoS | 5 |
| 2023 | Membership Inference Attacks Against Machine Learning Models via Prediction SensitivityabstractMachine learning (ML) has achieved huge success in recent years, but is also vulnerable to various attacks. In this article, we concentrate on membership inference attacks and propose Aster, which merely requires the target model's black-box API and a data sample to determine whether this sample was used to train the given ML model or not. The key idea of Aster is that the training data of a fully trained ML model usually has lower prediction sensitivities compared with that of the non-training data (i.e., testing data). Less sensitivity means that when perturbing a training sample's feature value in the corresponding feature space, the prediction of the perturbed sample obtained from the target model tends to be consistent with the original prediction. In this article, we quantify the prediction sensitivity with the Jacobian matrix which could reflect the relationship between each feature's perturbation and the corresponding prediction's change. Then we regard the samples with a lower as training data. Aster can breach the membership privacy of the target model's training data with no prior knowledge about the target model or its training data. The experiment results on four datasets show that our method outperforms three state-of-the-art inference attacks. Yi Wang 0150, Gaoyang Liu, Kai Peng 0001, Chen Wang 0011 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2023 | TEAR: Exploring Temporal Evolution of Adversarial Robustness for Membership Inference Attacks Against Federated LearningabstractFederated learning (FL) is a privacy-preserving machine learning paradigm that enables multiple clients to train a unified model without disclosing their private data. However, susceptibility to membership inference attacks (MIAs) arises due to the natural inclination of FL models to overfit on the training data during the training process, thereby enabling MIAs to exploit the subtle differences in the FL model’s parameters, activations, or predictions between the training and testing data to infer membership information. It is worth noting that most if not all existing MIAs against FL require access to the model’s internal information or modification of the training process, yielding them unlikely to be performed in practice. In this paper, we present with TEAR the first evidence that it is possible for an honest-but-curious federated client to perform MIA against an FL system, by exploring the Temporal Evolution of the Adversarial Robustness between the training and non-training data. We design a novel adversarial example generation method to quantify the target sample’s adversarial robustness, which can be utilized to obtain the membership features to train the inference model in a supervised manner. Extensive experiment results on five realistic datasets demonstrate that TEAR can achieve a strong inference performance compared with two existing MIAs, and is able to escape from the protection of two representative defenses. Gaoyang Liu, Zehao Tian, Jian Chen 0046, Chen Wang 0011, Jiangchuan Liu |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2023 | Knowledge Representation of Training Data With Adversarial Examples Supporting Decision BoundaryabstractDeep learning (DL) has achieved tremendous success in recent years in many fields. The success of DL typically relies on a considerable amount of training data and the expensive model optimization process. Therefore, a trained DL model and its corresponding training data have become valuable assets whose intellectual property (IP) needs to be protected. Once a DL model or its training dataset is released, there is currently no mechanism for the entity that owns one part to establish a clear relationship with the other. In this paper, we aim to reveal the integrated relationship between a given DL model and the corresponding training dataset, by framing the problem of knowledge representation of a dataset with respect to DL models trained on it:how to effectively represent the knowledge transferred from a training dataset to a DL model?Our basic idea is that the knowledge transferred from a training dataset to a DL model can be uniquely represented by the model’s decision boundary. Therefore, we design a novel generation method that utilizes geometric consistency to find the samples supporting the decision boundary, which can serve as the proxy for the knowledge representation. We evaluate our method in three different cases: IP audit of training data, IP audit of DL models, and adversarial knowledge distillation. The experimental results show that our method can improve the performance of existing works in all cases, which confirm that our method can effectively represent the knowledge transferred from a training dataset to a DL model. Zehao Tian, Zixiong Wang, Ahmed M. Abdelmoniem, Gaoyang Liu, Chen Wang 0011 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2022 | Your Model Trains on My Data? Protecting Intellectual Property of Training Data via Membership Fingerprint AuthenticationabstractIn recent years, data has become the new oil that fuels various machine learning (ML) applications. Just as the oil refining, providing data to an ML model is a product of massive costs and expertise efforts. However, how to protect the intellectual property (IP) of the training data in ML remains largely open. In this paper, we present MeFA, a novel framework for detecting training data IP embezzlement via Membership Fingerprint Authentication, which is able to determine whether a suspect ML model is trained on the to be protected target data or not. The key observation is that a part of data has a similar influence on the prediction behavior of different ML models. On this basis, MeFA leverages membership inference techniques to extract these data as the fingerprints of the target data and constructs an authentication model to verify the data’s ownership by identifying the obtained membership fingerprints. MeFA has several salient features. It does not assume any knowledge of the suspect model except for its black-box prediction API, through which we can merely get the prediction output of a given input, and also does not require any modification to the dataset or the training process, since it takes advantage of the inherent membership property of the data. As a by-product, MeFA can also serve as a post-protection to verify the ownership of ML models, without modifying the training process of the model. Extensive experiments on three realistic datasets and seven types of ML models validate the effectiveness of MeFA, and demonstrate that it is also robust to scenarios when the training data is partially used or preprocessed with representative membership inference defenses. Gaoyang Liu, Tianlong Xu, Xiaoqiang Ma, Chen Wang 0011 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2021 | FedEraser: Enabling Efficient Client-Level Data Removal from Federated Learning ModelsabstractFederated learning (FL) has recently emerged as a promising distributed machine learning (ML) paradigm. Practical needs of the "right to be forgotten" and countering data poisoning attacks call for efficient techniques that can remove, or unlearn, specific training data from the trained FL model. Existing unlearning techniques in the context of ML, however, are no longer in effect for FL, mainly due to the inherent distinction in the way how FL and ML learn from data. Therefore, how to enable efficient data removal from FL models remains largely under-explored. In this paper, we take the first step to fill this gap by presenting FedEraser, the first federated unlearning method-ology that can eliminate the influence of a federated client’s data on the global FL model while significantly reducing the time used for constructing the unlearned FL model. The basic idea of FedEraser is to trade the central server’s storage for unlearned model’s construction time, where FedEraser reconstructs the unlearned model by leveraging the historical parameter updates of federated clients that have been retained at the central server during the training process of FL. A novel calibration method is further developed to calibrate the retained updates, which are further used to promptly construct the unlearned model, yielding a significant speed-up to the reconstruction of the unlearned model while maintaining the model efficacy. Experiments on four realistic datasets demonstrate the effectiveness of FedEraser, with an expected speed-up of 4× compared with retraining from the scratch. We envision our work as an early step in FL towards compliance with legal and ethical criteria in a fair and transparent manner. Gaoyang Liu, Xiaoqiang Ma, Yang Yang 0060, Chen Wang 0011, Jiangchuan Liu |
IWQoS | 1 |
| 2021 | ML-Stealer: Stealing Prediction Functionality of Machine Learning Models with Mere Black-Box AccessabstractMachine Learning (ML) models are progressively deployed in many real-world applications to perform a wide range of tasks, but are exposed to the security and privacy threats which aim to infer the details and even steal the functionality of the ML models. Despite extensive attacking efforts which rely on white-box or gray-box access, how to perform attacks with black-box access continues to be elusive. Aspiring to fill this gap, we move one step further and present ML-Stealer that can steal the functionality of any type of ML models with mere black-box access. With two algorithm designs, namely, synthetic data generation and replica model construction, ML-Stealer can construct a deep neural network (DNN)-based replica model which has the similar prediction functionality to the victim ML model. ML-Stealer does not require any knowledge about the victim model, nor does it enforce the access to statistical information or samples of the victim's training data. Experiment results demonstrate that ML-Stealer can achieve the consistent prediction results with the victim model of an averaged testing accuracy of 85.6%, and up to 93.6% at best. Gaoyang Liu, Shijie Wang 0007, Borui Wan, Chen Wang 0011 |
TrustCom | 1 |
| 2021 | GPS spoofed or not? Exploiting RSSI and TSS in crowdsourced air traffic control data
Gaoyang Liu, Rui Zhang 0066, Yang Yang 0060, Chen Wang 0011, Ling Liu 0001 |
Distributed Parallel Databases | 1 |
| 2020 | AFA: Adversarial fingerprinting authentication for deep neural networks
Qingyue Hu, Gaoyang Liu, Xiaoqiang Ma, Fei Chen 0014, Mohammad Mehedi Hassan |
Comput. Commun. | 3 |
| 2020 | A survey of local differential privacy for securing internet of vehicles
Ping Zhao 0001, Guanglin Zhang, Shaohua Wan 0001, Gaoyang Liu, Tariq Umer |
J. Supercomput. | 4 |
| 2020 | MIASec: Enabling Data Indistinguishability Against Membership Inference Attacks in MLaaSabstractThe emerging of machine learning has massively promoted the abilities of computational sustainability in natural resource management and allocation. Many Internet giants such as Google, Amazon, and Microsoft now provide Machine Learning as a Service (MLaaS) to meet the increasing demand for machine learning services. However, the prediction results of training data and testing data with the same machine learning model in MLaaS have remarkable differences, and thus the attackers can leverage machine learning techniques to launch the so-called membership inference attacks, i.e., to infer whether a record is in the training data or not. In this paper, we propose MIASec that can guarantee the data indistinguishability of the training data and thereby has the ability to defend against membership inference attacks in MLaaS. The key idea of MIASec is to narrow the dynamic ranges of vital features in the training data, such that the training data, the testing data, and even the synthetic data have almost semblable prediction results by the same machine learning model. With elaborated design on modifying the values of vital features in the training data, MIASec can thus reduce the differences between the model's outcomes of training data and testing data, thereby protecting the training data in effect while keeping the model's accuracy stable. We empirically evaluate MIASec on machine learning models trained by off-line neural networks and on-line MLaaS. Using realistic data and classification tasks, our experiment results show that MIASec can defend the membership inference attacks effectively. In particular, MIASec can reduce the precision and recall of attacks respectively by 11.7 and 15.4 percent in average, and by 18.6 and 21.8 percent at best. Chen Wang 0011, Gaoyang Liu, Haojun Huang, Weijie Feng, Kai Peng 0001, Lizhe Wang 0001 |
IEEE Trans. Sustain. Comput. | 2 |
| 2020 | Recent advances in consensus protocols for blockchain: a survey
Shaohua Wan 0001, Meijun Li, Gaoyang Liu, Chen Wang 0011 |
Wirel. Networks | 3 |
| 2019 | Synchronization-Free GPS Spoofing Detection with Crowdsourced Air Traffic Control DataabstractGPS-dependent localization, navigation and air traffic control (ATC) applications have had a significant impact on the modern aviation industry. However, the lack of encryption and authentication makes GPS vulnerable to spoofing attacks with the purpose of hijacking aerial vehicles or threatening air safety. In this paper, we propose GPS-Probe, a GPS spoofing detection algorithm that leverages the ATC messages that are periodically broadcasted by aerial vehicles. By continuously analyzing the received signal strength indicator (RSSI) and the timestamps at server (TSS) of the ATC messages, which are monitored by multiple ground sensors, GPS-Probe constructs a machine learning enabled framework to estimate the real position of the target aerial vehicle and to detect whether or not the position data is compromised by GPS spoofing attacks. Unlike existing techniques, GPS-Probe neither requires any updates of the GPS infrastructure nor updates of the GPS receivers. More importantly, it releases the requirement on time synchronization of the ground sensors distributed around the world. Using the real-world ATC data crowdsourced by the OpenSky Network, our experiment results show that GPS-Probe can achieve the detection accuracy and precision, of 81.7% and 85.3% respectively on average, and up to 89.7% and 91.5% respectively at the best. Gaoyang Liu, Rui Zhang 0066, Chen Wang 0011, Ling Liu 0001 |
MDM | 1 |
| 2019 | Classifying transportation mode and speed from trajectory data via deep multi-scale learning
Rui Zhang 0066, Chen Wang 0011, Gaoyang Liu, Shaohua Wan 0001 |
Comput. Networks | 4 |
| 2019 | On the Performance of $k$ -Anonymity Against Inference Attacks With Background InformationabstractInternet of Things (IoT) applications bring in a great convenience for human’s life, but users’ data privacy concern is the major barrier toward the development of IoT.${k}$-anonymity is a method to protect users’ data privacy, but it is presently known to suffer from inference attacks. Thus far, existing work only relies on a number of experimental examples to validate${k}$-anonymity’s performance against inference attacks, and thereby lacks of a theoretical guarantee. To tackle this issue, in this paper we propose the first theoretical foundation that gives a nonasymptotic bound on the performance of${k}$-anonymity against inference attacks, taking into consideration of adversaries’ background information. The main idea is to first quantify adversaries’ background information, and from the point of the view of adversaries, classify users’ data into four kinds: 1) independent with unknown data values; 2) local dependent with unknown data values; 3) independent with certain known data values; and 4) local dependent with certain known data values. We then move one step further, theoretically proving the bound on the performance of${k}$-anonymity corresponding to each of the four kinds of users’ data through cooperating with the noiseless privacy. We argue that such a theoretical foundation links${k}$-anonymity with noiseless privacy, theoretically proving${k}$-anonymity provides noiseless privacy. Additionally, this paper theoretically explains why${k}$-anonymity is vulnerable to inference attacks using the modified Stein method. Simulations on real check-in dataset from the location-based social network have validated our results. We believe that this paper can bridge the gap between design and evaluation, enabling a designer to construct a more practical${k}$-anonymity technique in real-life scenarios to resist inference attacks. Ping Zhao 0001, Hongbo Jiang 0001, Chen Wang 0011, Haojun Huang, Gaoyang Liu, Yang Yang 0060 |
IEEE Internet Things J. | 5 |
| 2019 | SocInf: Membership Inference Attacks on Social Media Health Data With Machine LearningabstractSocial media networks have shown rapid growth in the past, and massive social data are generated which can reveal behavior or emotion propensities of users. Numerous social researchers leverage machine learning technology to build social media analytic models which can detect the abnormal behaviors or mental illnesses from the social media data effectively. Although the researchers only public the prediction interfaces of the machine learning models, in general, these interfaces may leak information about the individual data records on which the models were trained. Knowing a certain user's social media record was used to train a model can breach user privacy. In this paper, we present SocInf and focus on the fundamental problem known as membership inference. The key idea of SocInf is to construct a mimic model which has a similar prediction behavior with the public model, and then we can disclose the prediction differences between the training and testing data set by abusing the mimic model. With elaborated analytics on the predictions of the mimic model, SocInf can thus infer whether a given record is in the victim model's training set or not. We empirically evaluate the attack performance of SocInf on machine learning models trained by Xgboost, logistics, and online cloud platform. Using the realistic data, the experiment results show that SocInf can achieve an inference accuracy and precision of 73% and 84%, respectively, in average, and of 83% and 91% at best. Gaoyang Liu, Chen Wang 0011, Kai Peng 0001, Haojun Huang, Wenqing Cheng |
IEEE Trans. Comput. Soc. Syst. | 1 |