EDBT 2026 Demo / reviewers in the wild / expert
Di Wu 0050
dblp:52/328-50
· DBLP profile ↗
40ranked-venue papers
6as first author
32since 2021 · last 2026
0000-0002-4753-8161ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 5 first-author · 17 since 2021Security and privacy · 7 · 1 first-author · 4 since 2021Computer networks · 5 · 5 since 2021Systems, architecture and hardware · 4 · 2 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PAGPL: Privacy-Aware Graph Prompt Learning Scheme via Adaptive Perturbation-Estimated Topology RecoveryabstractGraph prompt learning (GPL) serves as a crucial framework for mitigating the knowledge transfer by reconciling the substantial mismatch between pre-training models and downstream tasks. However, prevalent GPL paradigm fail to accommodate graph data affected by privacy-induced noise. Specifically, 1) GPL typically relies on the stability of original graph structures for the design of effective prompt templates; 2) the construction of prompts lacks explicit guidance to suppress noise introduced by privacy perturbations; 3) prompt optimization on single disturbed graphs can easily lead to overfitting to noise patterns. To address these issues, we propose a novel privacy-aware graph prompt learning (PAGPL) scheme, which alleviates spurious clues caused by privacy noise injection. Initially, an adaptive structure-wise Bayesian estimation is applied to reconstruct the privacy-perturbed graphs. Subsequently, to suppress the impact of residual perturbation, a noise-resilient prompt generation is employed to filter unreliable structural and signals. Ultimately, we incorporate a multi-view-based progressive privacy consistency to promote the robustness of prompts against the semantic misalignment while improving the task-specific consistency. The experimental results reveal that our scheme outperforms state-of-the-art (SOTA) GPL approaches with a 10%–60% improvement in accuracy under various real-world privacy-perturbed scenarios. Ju Jia, Jiansen Song, Jingxuan Yu, Jiabao Guo, Xiaoshuang Jia, Di Wu 0050, Yali Yuan, Guang Cheng 0001 |
AAAI | 6 |
| 2026 | Cross-Modal Unlearning via Influential Neuron Path Editing in Multimodal Large Language ModelsabstractMultimodal Large Language Models (MLLMs) extend foundation models to real-world applications by integrating inputs such as text and vision. However, their broad knowledge capacity raises growing concerns about privacy leakage, toxicity mitigation, and intellectual property violations. Machine Unlearning (MU) offers a practical solution by selectively forgetting targeted knowledge while preserving overall model utility. When applied to MLLMs, existing neuron-editing-based MU approaches face two fundamental challenges: (i) forgetting becomes inconsistent across modalities because existing point-wise attribution methods fail to capture the structured, layer-by-layer information flow that connects different modalities; and (ii) general knowledge performance declines when sensitive neurons that also support important reasoning paths are pruned, as this disrupts the model’s ability to generalize. To alleviate these limitations, we propose a multimodal influential neuron path editor (MIP-Editor) for MU. Our approach introduces modality-specific attribution scores to identify influential neuron paths responsible for encoding forget-set knowledge and applies influential-path-aware neuron-editing via representation misdirection. This strategy also enables effective and coordinated forgetting across modalities while preserving the model's general capabilities. Experimental results demonstrate that MIP-Editor achieves a superior unlearning performance on multimodal tasks, with a maximum forgetting rate of 87.75% and up to 54.26% improvement in general knowledge retention. On textual tasks, MIP-Editor achieves up to 80.65% forgetting and preserves 77.90% of general performance. Kunhao Li, Di Wu 0050, Ju Jia, Minhui Xue 0001 |
AAAI | 3 |
| 2026 | LegiCode: A blockchain-legal LLM framework for real-time compliance in smart contract generation
Linkai Zhu, Di Wu 0050, Longxiang Gao |
Empir. Softw. Eng. | 4 |
| 2026 | GResMark: A swin transformer-based watermarking framework with geometric attack resilience
Weitong Chen 0002, Jiale Zhang 0001, Chunpeng Ge 0001, Di Wu 0050, Willy Susilo, Palaiahnakote Shivakumara |
Expert Syst. Appl. | 5 |
| 2026 | Dynamic Generator: A Stealth-Preserving Generator-Based Data Poisoning Attack in Federated Learning With Defensive Countermeasures in IoT SystemsabstractFederated Learning (FL) is a distributed learning framework that enables collaborative training of a global model across multiple IoT devices while preserving the privacy of local data. However, this collaborative paradigm also introduces security challenges to the global model. Compromised IoT devices can carry out data poisoning attacks on their local data, uploading malicious model gradients to the cloud server and ultimately degrading the global model performance. Most existing data poisoning methods focus primarily on maximizing attack effectiveness, causing a significant drop in global model accuracy, while often neglecting the stealthiness of the attack. This leads to malicious updates being easily detected and filtered, reducing their threat in IoT deployments. To fill this gap, we propose a stealthy data poisoning method based on a dynamic generator. The approach dynamically generates perturbations with gradient-level stealthiness, achieving a trade-off between poisoning effectiveness and stealthiness in realistic IoT settings. We conducted comprehensive experiments on 3 public IoT datasets: with 30% malicious clients, our attack reduces global accuracy by 10.78% on DeepHealth, 6.07% on ChestMNIST and 7.81% on SVHN, achieving degradation comparable to baselines while preserving stealthiness at the gradient level. To counter this threat, We propose PCATrace, a trajectory-based defense that consistently detects and removes malicious clients across different attack ratios, effectively preventing poisoned updates from degrading the global model. Xiuheng Liao, Ziang Wu, Buzhen He, Shuai Shang, Di Wu 0050, Chunhua Su |
IEEE Internet Things J. | 6 |
| 2026 | A lightweight privacy-preserving fingerprint authentication system for IoT devices via pruned and secured minutia cylinder codeabstractFingerprint authentication is extensively adopted due to its ease of capture, low cost sensors and high recognition accuracy. The Minutia Cylinder Code (MCC) is a high-quality feature representation widely used in fingerprint authentication. However, there are two main limitations in the direct use of MCC: redundancy in the feature representation due to overlap between minutiae vicinities, which can lead to inefficient resource utilization; and vulnerability to template inversion attacks, which may expose the original fingerprint data and threaten user privacy. In this paper, we propose a lightweight privacy-preserving fingerprint authentication system that overcomes these limitations through two novel algorithms. The first algorithm, P-MCC, uses the Pearson correlation coefficient to prune MCC features to effectively reduce redundancy and improve resource utilisation, yielding a lightweight design. The second algorithm, S-MCC, applies a secure Boolean function which transforms the pruned MCC features non-invertibly to ensure privacy, thus preventing the reconstruction of original fingerprint data. Together, P-MCC and S-MCC provide a lightweight privacy-preserving fingerprint authentication system, which is well suited to resource-constrained environments, such as the Internet of Things (IoT). Experimental results demonstrate the effectiveness of the proposed system and its practicality in IoT applications. Wencheng Yang, Song Wang 0003, Yan Li 0002, Di Wu 0050, Ji Zhang 0001, Xu Yang 0002 |
J. Inf. Secur. Appl. | 4 |
| 2026 | FedMLAC: Mutual learning driven heterogeneous federated audio classificationabstractFederated Learning (FL) offers a privacy-preserving framework for training audio classification (AC) models across decentralized clients without sharing raw data. However, Federated Audio Classification faces three major challenges: data heterogeneity , model heterogeneity , and data corruption , which degrade performance in real-world settings. While existing methods often address these issues separately, a unified solution remains underexplored. We propose FedMLAC, a mutual learning-based FL framework that tackles all three challenges simultaneously. Each client maintains a personalized local AC model and a lightweight, globally shared Plug-in model. These models interact via bidirectional knowledge distillation, enabling global knowledge sharing while adapting to local data distributions, thus addressing both data and model heterogeneity. To counter data corruption, we introduce a Layer-wise Pruning Aggregation (LPA) strategy that filters anomalous Plug-in updates based on parameter deviations during aggregation. Extensive experiments on four diverse AC benchmarks, including both speech and non-speech tasks, show that FedMLAC consistently outperforms state-of-the-art baselines in classification accuracy and robustness to noisy data. Rajib Rana, Di Wu 0050, Youyang Qu, Xiaohui Tao 0001, Ji Zhang 0001, Carlos Busso, Palaiahnakote Shivakumara |
Pattern Recognit. | 3 |
| 2026 | Privacy-preserving federated SAR image target recognition with adaptive resource management in space-air-ground integrated networks
Yuchao Hou, Zhiqin Yang, Wei Xiang 0001, Di Wu 0050, Minghui LiWang, Xiaoyu Xia 0001, Zijian Li 0007, Youliang Tian, Yuzhou Sun |
Pattern Recognit. | 6 |
| 2026 | Federated continual learning with domain-routed historical expert matching for multimodal perception
Xiuheng Liao, Zhiwei Si, Di Wu 0050, Buzhen He, Chunhua Su |
Pattern Recognit. | 4 |
| 2026 | Explanation-guided backdoor defense for ID and OOD attacks in graph neural networks
Hao Sui 0003, Bing Chen 0002, Jiale Zhang 0001, Di Wu 0050, Palaiahnakote Shivakumara |
Pattern Recognit. | 5 |
| 2026 | FiDD: Secure Fine-Grained Deduplication and Dynamic Auditing Scheme for Cloud StorageabstractWith the rapid development of cloud computing, more and more users tend to store their data remotely to the cloud. Taking into account data security and resource utilization comprehensively, in addition to providing users with basic remote data integrity verification, cloud servers also need to conduct redundancy checks. However, current deduplication schemes primarily focus on static file-level data and auditing processes, rendering them inadequate for managing resources with dynamic attributes. In this paper, we propose a fine-grained deduplication and dynamic auditing model (FiDD) for cloud storage to address these challenges. FiDD utilizes homomorphic verifier-based data tags to seamlessly integrate deduplication and auditing processes, allowing both block-level and file-level deduplications. Additionally, FiDD employs doubly linked lists and multi-set hash functions to enhance the efficiency of data updates. The security of FiDD is validated through rigorous security proofs, while its efficiency is demonstrated through comprehensive experimental analysis. The experiments demonstrate an average improvement of at least 35% in audit efficiency and at least 50% in dynamics efficiency. Consequently, FiDD enhanes the security of data management in cloud computing while improving its overall efficiency. Longxia Huang, Lei Zhou 0026, Di Wu 0050, Longxiang Gao, Tom H. Luan |
IEEE Trans. Netw. | 4 |
| 2025 | Enhancing Privacy in Face Recognition With Dual-Path Feature Compression and Homomorphic EncryptionabstractFace recognition offers seamless human-machine interaction and efficiency. However, its widespread adoption has heightened security and privacy concerns due to the risks associated with compromised biometric data, such as spoofing and unauthorized tracking. To mitigate these concerns, this paper introduces a novel privacy-preserving face recognition framework that integrates an enhanced dual-path feature compression approach with homomorphic encryption (HE) for secure and efficient authentication. We leverage the robust deep neural network model FaceNet to extract discriminative 512-dimensional feature vectors and propose two significantly improved complementary feature compression methods tailored specifically for encrypted biometric systems: (1) Partitioned Principal Component Analysis (P-PCA), which employs a novel segment-wise PCA transformation, preserving localized discriminative information and supporting revocable biometric templates; and (2) Segment-wise Locality-Sensitive Hashing (S-LSH), introducing segment-specific hashing optimized for efficient binary representation and privacy-preserving encrypted-domain computations. Both compressed real-valued and binary features are securely encrypted using HE, enabling direct encrypted-domain similarity computations without exposing sensitive biometric data. Extensive experiments demonstrate that our method achieves competitive authentication performance while maintaining computational efficiency and practical feasibility. Wencheng Yang, Song Wang 0003, Di Wu 0050, Xu Yang 0002, Hui Cui 0001, Michael N. Johnstone, Yan Li 0002 |
IJCB | 3 |
| 2025 | A Unified Solution to Diverse Heterogeneities in One-Shot Federated LearningabstractOne-Shot Federated Learning (OSFL) restricts communication between the server and clients to a single round, significantly reducing communication costs and minimizing privacy leakage risks compared to traditional Federated Learning (FL), which requires multiple rounds of communication. However, existing OSFL frameworks remain vulnerable to distributional heterogeneity, as they primarily focus on model heterogeneity while neglecting data heterogeneity. To bridge this gap, we propose FedHydra, a unified, data-free, OSFL framework designed to effectively address both model and data heterogeneity. Unlike existing OSFL approaches, FedHydra introduces a novel two-stage learning mechanism. Specifically, it incorporates model stratification and heterogeneity-aware stratified aggregation to mitigate the challenges posed by both model and data heterogeneity. By this design, the data and model heterogeneity issues are simultaneously monitored from different aspects during learning. Consequently, FedHydra can effectively mitigate both issues by minimizing their inherent conflicts. We compared FedHydra with five SOTA baselines on four benchmark datasets. Experimental results show that our method outperforms the previous OSFL methods in both homogeneous and heterogeneous settings. The code is available at https://github.com/Jun-B0518/FedHydra. Yiliao Song, Di Wu 0050, Atul Sajjanhar, Yong Xiang 0001, Wei Zhou 0044, Xiaohui Tao 0001, Yan Li 0002, Yue Li 0017 |
KDD (2) | 3 |
| 2025 | Prompt as a Double-Edged Sword: A Dynamic Equilibrium Gradient-Assigned Attack against Graph Prompt LearningabstractGraph prompt learning (GPL) is designed to bridge the gap between graph pretraining models and downstream graph tasks, providing advantages in terms of graph knowledge transfer. However, GPL is vulnerable to poisoned graph attacks that induce abnormal training via adversarial malicious perturbations. We observe that the prevalent meta-gradient attacks, which heavily rely on the training of surrogate graph neural networks (GNNs), fail to account for the impact of perturbations on GPL where the pretrained GNN remains frozen and graph prompt tokens are tuned. Moreover, their gradient-assigned strategies tend to corrupt the topological semantics on a few influential labeled graphs, which in turn diminishes the trustworthiness of the surrogate training. To address this issue, we propose a dynamic equilibrium gradient-assigned attack against GPL, named MetaGpro. To guarantee the transferability of MetaGpro, the surrogate GPL is utilized in our simulation across various downstream tasks. To dynamically equilibrate the relationships between the reliability of surrogate models and instable structures, the over-robust contrastive learning is integrated into the surrogate training. In this way, the gradient bias caused by excessive perturbations of labeled nodes can be effectively mitigated. Subsequently, the topology perturbation generation is exploited to assign more gradient weights to nodes that are closer to the misclassification area. The experimental results reveal that the surrogate GPL outperforms the surrogate GNN in 96% of downstream evaluations, and our MetaGpro reduces the accuracy of GPL by 2%∼20% compared to the state-of-the-art (SOTA) works mostly. The code for our MetaGpro is available here. Ju Jia, Jingxuan Yu, Di Wu 0050, Cong Wu 0003, Hengjie Zhu, Lina Wang 0001 |
KDD (2) | 3 |
| 2025 | Beyond Dataset Watermarking: Model-Level Copyright Protection for Code Summarization Models
Jiale Zhang 0001, Di Wu 0050, Xiaobing Sun 0001, Qinghua Lu 0001, Guodong Long |
WWW | 3 |
| 2025 | EPAD: Ethereum phishing scam detection via graph contrastive learning
Hao Sui 0003, Jiale Zhang 0001, Bing Chen 0002, Di Wu 0050, Xiaobing Sun 0001, Palaiahnakote Shivakumara |
Expert Syst. Appl. | 4 |
| 2025 | SFFL: Self-aware fairness federated learning framework for heterogeneous data distributions
Jiale Zhang 0001, Ye Li 0041, Di Wu 0050, Yanchao Zhao, Palaiahnakote Shivakumara |
Expert Syst. Appl. | 3 |
| 2025 | Advancing DoA assessment through federated learning: A one-shot pseudo data approachabstractAccurately measuring the Depth of Anaesthesia (DoA) during surgical procedures is crucial for patient safety. A significant challenge in developing effective machine learning models for DoA assessment is the lack of data from single organisations and preserving data privacy between institutions. Federated learning offers a solution by enabling multiple parties to collaboratively train models without exchanging data. However, traditional federated learning algorithms perform poorly in data heterogeneous, non-identically distributed data distribution scenarios. To address these challenges, we propose a one-shot federated learning framework, DoAFedP-NN, which facilitates federated learning with heterogeneous model development. The framework is tested in a range of model and data heterogeneity environments. This method enables the training of a global DoA prediction model across different medical facilities without sharing local data. The DoAFedP-NN model, utilising neural network design with entropy and spectral feature extraction, is compared to benchmark federated learning architectures, demonstrating its advantage in handling heterogeneous medical data. Experimental results show that DoAFedP-NN achieves robust DoA estimation when compared to the Bispectral (BIS) index, with high correlation coefficients of 0.8472 and 0.8542 across independent databases. The proposed model outperforms locally developed models, showing significant improvements when validated against external datasets from different medical facilities. This paper makes the key contributions: (1) introduces a one-shot pseudo-data method for federated learning; (2) demonstrates the effectiveness of this approach for EEG-based DoA using real-world databases; (3) showcases the model’s ability to achieve high correlation with the BIS index while preserving patient privacy in a range of client distribution scenarios and under cross-validation. • This study introduces a novel federated learning framework, DoAFedP-NN, which utilises neural network architecture to improve the accuracy of EEG-based Depth of Anaesthesia (DoA) monitoring. By integrating data across multiple databases without sharing local patient data, the framework respects patient privacy while enhancing model performance. • The DoAFedP-NN model, employing entropy and spectral analysis, demonstrated robust DoA estimation capabilities, achieving high correlation coefficients across independent databases. • The research utilises a novel pseudo-data federated learning aggregation method to enable heterogeneous model development through one-shot a novel federated learning framework for EEG-based DoA analysis. The DoAFedP-NN model’s superior performance compared to locally trained models and its comparability to a traditional full aggregation model demonstrate the value of federated learning in achieving high analytical precision without compromising patient privacy. Thomas Schmierer, Tianning Li, Di Wu 0050, Yan Li 0002 |
Neurocomputing | 3 |
| 2025 | FedMLC: White-Box Model Watermarking for Copyright Protection in Federated Learning for IoT EnvironmentabstractWith the widespread application of the Internet of Things (IoT), data processing has gradually migrated to edge devices that are closer to the data source. This shift has significantly improved the ability of real-time data analysis while effectively reducing bandwidth requirements and latency. Furthermore, Federated Learning (FL) has been introduced as a decentralized training method to achieve collaborative training of multiple devices while ensuring local data privacy. However, malicious clients in FL may theft trained models for unauthorized use, which causes model misuse or copyright challenges. To address these issues, this paper proposes FedMLC (Malicious client detection, Leakage tracing, and Copyright verification), a server-side white-box watermarking scheme. FedMLC utilizes the embedded watermark at different stages to achieve both traceability and copyright verification, simplifying the watermarking process. Additionally, the watermarking can also detect malicious clients in FL. Specifically, FedMLC uses the regularization term to guide the parameter signs of the normalization layer to be consistent with the watermark sign, thereby achieving watermark embedding. Experimental results show that our FL model watermarking scheme excels in malicious client detection, leakage tracing, and copyright verification, with minimal impact on model performance, able to resist various attacks such as fine-tuning, pruning, and quantization. Weitong Chen 0002, Wei Zhang 0098, Di Wu 0050, Anja Keskinarkaus, Tapio Seppänen, Jiale Zhang 0001, Longxiang Gao, Tom H. Luan |
IEEE Internet Things J. | 3 |
| 2025 | Federated Average Clustering Learning Based on Time-Asynchronous SimilarityabstractTo address the challenge of non-independent and identically distributed data in federated learning, clustering federated learning extracts local model features from clients and groups clients with high local model similarity into the same cluster to optimize the global model for the corresponding data distribution. However, current federated clustering algorithms bring additional communication cost for the accuracy of cluster division, which increases the communication burden of some performance-limited edge devices. To address the communication efficiency issue of clustering federated learning, we proposed an optimized communication efficiency federated average clustering learning that uses time weights to select partial clients to participate in training randomly. At the same time, a time-asynchronous similarity calculation method is proposed to improve the accuracy of local model similarity calculation for randomly selected clients. Extensive experimental evaluations show that our federated average clustering learning can achieve or even surpass the model accuracy of existing federated clustering algorithms. In the communication evaluation experiment, we can achieve the specified model accuracy using 5% to 30% of the communication of existing federated clustering algorithms. Xiaoyao Zheng, Di Wu 0050, Tianran Bu, Ji Zhang 0001 |
IEEE Trans. Big Data | 3 |
| 2025 | HgtJIT: Just-in-Time Vulnerability Detection Based on Heterogeneous Graph TransformerabstractVulnerability detection plays a crucial role in the software development lifecycle. Commit-level vulnerability detection aims to detect whether the changed code contributed to potential vulnerabilities by the developer when submitting the code, which is also referred to as Just-In-Time (JIT) vulnerability detection. Previous JIT vulnerability detection approaches relied on code metrics and textual features, which were unable to effectively characterize vulnerability-contributing commits (VCCs). Recently, CodeJIT (a code-centric learning-based approach) has been proposed to detect vulnerability at the commit-level. However, CodeJIT still has its limitations: imprecise feature representation, static code embedding, and underutilized heterogeneous information. In this paper, we propose HgtJIT, a JIT vulnerability detection approach based on a Heterogeneous Graph Transformer (HGT) in order to address several limitations of the state-of-the-art CodeJIT approach. We propose diffPDG to represent code changes and use the CCT5 model (the latest feature encoder pre-trained on a large-scale code change corpus) to embed graph nodes to generate the most meaningful vector representations. In addition, we employ HGT to adequately utilize heterogeneous information of the graph to learn vulnerability features. Extensive experiments have shown that HgtJIT is the best-performing model, with F1 and AUC improvement of 14.6%-37.5% and 12.2%-53.7% compared to the baseline model Xiaobing Sun 0001, Mingxuan Zhou, Sicong Cao, Xiaoxue Wu 0001, Lili Bo, Di Wu 0050, Bin Li 0006, Yang Xiang 0001 |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2025 | Machine Unlearning for Source-Free Unsupervised Partial-Domain Adaptation in Remote SensingabstractSource-Free Unsupervised Domain Adaptation (SFUDA) enables model adaptation to unlabeled target domains without accessing source data. However, when the source domain contains classes absent in the target domain, existing methods suffer from negative transfer: knowledge of irrelevant source-only classes interferes with target class recognition, significantly degrading classification accuracy. We propose Machine Unlearning-based SFUDA (MUSFUDA), which addresses this problem by selectively unlearning source-only class knowledge from the pre-trained model rather than adding compensatory mechanisms. This machine unlearning approach allows the model to focus on shared classes, fundamentally eliminating negative transfer. Remote sensing images with large intra-class variations and high inter-class similarity cause over-unlearning of target classes when forgetting source-only classes, thus we design the Model-Disruption Based Dual-Teacher Unlearning Strategy (MDUS), which uses dual teachers to manage target class preservation and source-only class erasure through knowledge distillation. MDUS is lightweight and easily integrated into existing SFUDA frameworks. Experiments on remote sensing datasets demonstrate that combining MDUS with representative baselines consistently reduces negative transfer and improves classification performance, maintaining high efficiency, validating the effectiveness and generalizability of our approach. Jielong Yang, Xialun Yun, Xionghu Zhong, Di Wu 0050 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | FedInverse: Evaluating Privacy Leakage in Federated LearningabstractFederated Learning (FL) is a distributed machine learning technique where multiple devices (such as smartphones or IoT devices) train a shared global model by using their local data. FL claims that the data privacy of local participants is preserved well because local data will not be shared with either the server-side or other training participants. However, this paper discovers a pioneering finding that a model inversion (MI) attacker, who acts as a benign participant, can invert the shared global model and obtain the data belonging to other participants. This will lead to severe data-leakage risk in FL because it is difficult to identify attackers from benign participants.
In addition, we found even the most advanced defense approaches could not effectively address this issue. Therefore, it is important to evaluate such data-leakage risks of an FL system before using it. To alleviate this issue, we propose FedInverse to evaluate whether the FL global model can be inverted by MI attackers. In particular, FedInverse can be optimized by leveraging the Hilbert-Schmidt independence criterion (HSIC) as a regularizer to adjust the diversity of the MI attack generator. We test FedInverse with three typical MI attackers, GMI, KED-MI, and VMI, and the experiments show our FedInverse method can successfully obtain the data belonging to other participants. The code of this work is available at https://github.com/Jun-B0518/FedInverse Di Wu 0050, Yiliao Song, Wei Zhou 0044, Yong Xiang 0001, Atul Sajjanhar |
ICLR | 1 |
| 2024 | BADFSS: Backdoor Attacks on Federated Self-Supervised Learning
Jiale Zhang 0001, Di Wu 0050, Xiaobing Sun 0001, Jianming Yong, Guodong Long |
IJCAI | 3 |
| 2024 | EXVul: Toward Effective and Explainable Vulnerability Detection for IoT DevicesabstractAs with anything connected to the internet, Internet of Things (IoT) devices are also subject to severe cybersecurity threats because an adversary could exploit vulnerabilities in their internal software to perform malicious attacks. Despite the promising results of Deep Learning (DL)-based approaches, the lack of well-labeled IoT vulnerability samples available for training and explainability pose a critical challenge to deploy them in practice. In this paper, we propose, a novel DL-based approach for Effective and eXplainable IoT VULnerability detection. Specifically, inspired by recent advances of self-supervised learning in label-expensive tasks, we propose a new combinatorial contrastive loss to combine the strengths of large-scale unlabeled code corpus and limited IoT vulnerability samples. Then, given a binary detection result, provides a set of faithful and stable code statements positively contributing to the model’s predictions as understandable explanations. Experimental results indicate that outperforms state-of-the-art baselines by 33.44%-72.91% and 19.52%-98.78% with respect to the accuracy and F1 score metrics, respectively. For vulnerability explanation, improves over the best-performing baseline explainer PGExplainer by 22.97% in MSP, 49.55% in MSR, and 48.40% in MIoU, demonstrating that the explanations provided by can correctly point out the vulnerable statements relevant to the detected vulnerabilities. Sicong Cao, Xiaobing Sun 0001, Wei Liu 0010, Di Wu 0050, Jiale Zhang 0001, Yan Li 0002, Tom H. Luan, Longxiang Gao |
IEEE Internet Things J. | 4 |
| 2024 | From Wide to Deep: Dimension Lifting Network for Parameter-Efficient Knowledge Graph EmbeddingabstractKnowledge graph embedding (KGE) that maps entities and relations into vector representations is essential for downstream applications. Conventional KGE methods require high-dimensional representations to learn the complex structure of knowledge graph, but lead to oversized model parameters. Recent advances reduce parameters by low-dimensional entity representations, while developing techniques (e.g., knowledge distillation or reinvented representation forms) to compensate for reduced dimension. However, such operations introduce complicated computations and model designs that may not benefit large knowledge graphs. To seek a simple strategy to improve the parameter efficiency of conventional KGE models, we take inspiration from that deeper neural networks require exponentially fewer parameters to achieve expressiveness comparable to wider networks for compositional structures. We view all entity representations as a single-layer embedding network, and conventional KGE methods that adopt high-dimensional entity representations equal widening the embedding network to gain expressiveness. To achieve parameter efficiency, we instead propose a deeper embedding network for entity representations, i.e., a narrow entity embedding layer plus a multi-layer dimension lifting network (LiftNet). Experiments on three public datasets show that by integrating LiftNet, four conventional KGE methods with 16-dimensional representations achieve comparable link prediction accuracy as original models that adopt 512-dimensional representations, saving 68.4% to 96.9% parameters. Borui Cai, Yong Xiang 0001, Longxiang Gao, Di Wu 0050, He Zhang 0034, Jiong Jin, Tom H. Luan |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2023 | Hybrid KD-NFT: A multi-layered NFT assisted robust Knowledge Distillation framework for Internet of Things
Nai Wang, Di Wu 0050, Wencheng Yang, Yong Xiang 0001, Atul Sajjanhar |
J. Inf. Secur. Appl. | 3 |
| 2022 | Campus Network Intrusion Detection based on Federated LearningabstractTo solve the problem of data scarcity and data silos in campus network intrusion detection, an intrusion detection method based on federated learning is proposed. This method allows multiple participants to collaboratively train a global detection model without sharing their training data with third parties, protecting data privacy. Federated learning is connected to transfer learning, as federated learning allows participants' knowledge transfer via its training mechanism. The resampling method is used in the federated learning training process to improve the global detection model's performance on rare class data. Besides, a contribution evaluation method is proposed, which evaluates participants' contribution in federated learning from two aspects of data quality and quantity. Experimental results show that the proposed method can achieve intrusion detection performance similar to traditional centralized collaborative learning under the premise of protecting participant data privacy. Zhongnan Fu, Qun Shang, Di Wu 0050 |
IJCNN | 6 |
| 2022 | A Blockchain-based Multi-layer Decentralized Framework for Robust Federated LearningabstractWith the expansion of the Internet of Things (IoT) development and application, federated learning has gained higher popularity in industrial researching fields. However, the security issues in federated learning have become hot-spots in the research area, such as privacy-preserving and poisoning attacks. This paper proposes a robust blockchained multi-layer decentralized federated learning (RBML-DFL) framework to ensure the federated learning's robustness. Firstly, by adopting the three-layered framework, the blockchain connects the federated learning components to secure the privacy and data safety of federated learning. Secondly, the proposed framework provides resilience on poisoning attacks to the central model compared to typical federated learning frameworks. Lastly, the decentralized structure associated with the blockchain tracing back mechanism can prevent the central server failure or mal-function compared to centralized federated learning. We evaluate and compare the proposed framework with other state-of-the-art federated learning frameworks on the accuracy, latency, and system robustness under poisoning attacks. The results show that the proposed RBML-DFL framework outperforms state-of-the-art baseline frameworks on all three metrics: accuracy, latency, and the robustness of the federated learning. Di Wu 0050, Nai Wang, Jiale Zhang 0001, Yuan Zhang 0007, Yong Xiang 0001, Longxiang Gao |
IJCNN | 1 |
| 2022 | From distributed machine learning to federated learning: In the view of data privacy and securityabstractSummary Federated learning is an improved version of distributed machine learning that further offloads operations which would usually be performed by a central server. The server becomes more like an assistant coordinating clients to work together rather than micromanaging the workforce as in traditional DML. One of the greatest advantages of federated learning is the additional privacy and security guarantees it affords. Federated learning architecture relies on smart devices, such as smartphones and IoT sensors, that collect and process their own data, so sensitive information never has to leave the client device. Rather, clients train a submodel locally and send an encrypted update to the central server for aggregation into the global model. These strong privacy guarantees make federated learning an attractive choice in a world where data breaches and information theft are common and serious threats. This survey outlines the landscape and latest developments in data privacy and security for federated learning. We identify the different mechanisms used to provide privacy and security, such as differential privacy, secure multiparty computation and secure aggregation. We also survey the current attack models, identifying the areas of vulnerability and the strategies adversaries use to penetrate federated systems. The survey concludes with a discussion on the open challenges and potential directions of future work in this increasingly popular learning paradigm. Sheng Shen 0005, Tianqing Zhu, Di Wu 0050, Wanlei Zhou 0001 |
Concurr. Comput. Pract. Exp. | 3 |
| 2022 | Detecting and mitigating poisoning attacks in federated learning using generative adversarial networksabstractSummary In the age of the Internet of Things (IoT), large numbers of sensors and edge devices are deployed in various application scenarios; Therefore, collaborative learning is widely used in IoT to implement crowd intelligence by inviting multiple participants to complete a training task. As a collaborative learning framework, federated learning is designed to preserve user data privacy, where participants jointly train a global model without uploading their private training data to a third party server. Nevertheless, federated learning is under the threat of poisoning attacks, where adversaries can upload malicious model updates to contaminate the global model. To detect and mitigate poisoning attacks in federated learning, we propose a poisoning defense mechanism, which uses generative adversarial networks to generate auditing data in the training procedure and removes adversaries by auditing their model accuracy. Experiments conducted on two well‐known datasets, MNIST and Fashion‐MNIST, suggest that federated learning is vulnerable to the poisoning attack, and the proposed defense method can detect and mitigate the poisoning attack. Ying Zhao 0011, Jiale Zhang 0001, Di Wu 0050, Michael Blumenstein, Shui Yu 0001 |
Concurr. Comput. Pract. Exp. | 4 |
| 2021 | Defending against Membership Inference Attacks in Federated learning via Adversarial ExampleabstractFederated learning has attracted attention in recent years due to its native privacy-preserving features. However, it is still vulnerable to various membership inference attacks, such as backdoor, poisoning, and adversarial attacks. Membership Inference attack aims to discover the data used to train the model, which leads to privacy leaking ramifications on participants who use their local data to train the shared model. Recent research on countermeasure methods mainly focuses on protecting the parameters and has limitations in guaranteeing privacy while restraining the loss of the model. This paper proposes Fedefend, which applies adversarial examples to defend against membership inference attacks in federated learning. The proposed approach adds well-designed noise to the attack features of the target model of each iteration becomes an adversarial example. In addition, we also consider the utility loss of the model and use an adversarial method to generate noise to constrain the loss to a certain extent, which efficiently achieves a trade-off between privacy security and loss of the federated learning model. We evaluate the proposed Fedefend on two benchmark datasets, and the experimental results demonstrate that Fedefend has a good performance. Yuanyuan Xie, Bing Chen 0002, Jiale Zhang 0001, Di Wu 0050 |
MSN | 4 |
| 2019 | A Privacy-Preserving Access Control Scheme with Verifiable and Outsourcing Capabilities in Fog-Cloud Computing
Jiale Zhang 0001, Hongyan Qian, Mingrong Xiang, Di Wu 0050 |
ICA3PP (1) | 5 |
| 2019 | PDGAN: A Novel Poisoning Defense Method in Federated Learning Using Generative Adversarial Network
Ying Zhao 0011, Jiale Zhang 0001, Di Wu 0050, Jian Teng, Shui Yu 0001 |
ICA3PP (1) | 4 |
| 2019 | Adversarial Action Data Augmentation for Similar Gesture Action RecognitionabstractHuman gestures are unique for recognizing and describing human actions, and video-based human action recognition techniques are effective solutions to varies real-world applications, such as surveillance, video indexing, and human-computer interaction. Most existing video human action recognition approaches either using handcraft features from the frames or deep learning models such as convolutional neural networks (CNN) and recurrent neural networks (RNN); however, they have mostly overlooked the similar gestures between different actions when processing the frames into the models. The classifiers suffer from similar features extracted from similar gestures, which are unable to classify the actions in the video streams. In this paper, we propose a novel framework with generative adversarial networks (GAN) to generate the data augmentation for similar gesture action recognition. The contribution of our work is tri-fold: 1) we proposed a novel action data augmentation framework (ADAF) to enlarge the differences between the actions with very similar gestures; 2) the framework can boost the classification performance either on similar gesture action pairs or the whole dataset; 3) experiments conducted on both KTH and UCF101 datasets show that our data augmentation framework boost the performance on both similar gestures actions as well as the whole dataset compared with baseline methods such as 2DCNN and 3DCNN. Di Wu 0050, Nabin Sharma, Shirui Pan, Guodong Long, Michael Blumenstein |
IJCNN | 1 |
| 2019 | Feature-Dependent Graph Convolutional Autoencoders with Adversarial Training MethodsabstractGraphs are ubiquitous for describing and modeling complicated data structures, and graph embedding is an effective solution to learn a mapping from a graph to a low-dimensional vector space while preserving relevant graph characteristics. Most existing graph embedding approaches either embed the topological information and node features separately or learn one regularized embedding with both sources of information, however, they mostly overlook the interdependency between structural characteristics and node features when processing the graph data into the models. Moreover, existing methods only reconstruct the structural characteristics, which are unable to fully leverage the interaction between the topology and the features associated with its nodes during the encoding-decoding procedure. To address the problem, we propose a framework using autoencoder for graph embedding (GED) and its variational version (VEGD). The contribution of our work is two-fold: 1) the proposed frameworks exploit a feature-dependent graph matrix (FGM) to naturally merge the structural characteristics and node features according to their interdependency; and 2) the Graph Convolutional Network (GCN) decoder of the proposed framework reconstructs both structural characteristics and node features, which naturally possesses the interaction between these two sources of information while learning the embedding. We conducted the experiments on three real-world graph datasets such as Cora, Citeseer and PubMed to evaluate our framework and algorithms, and the results outperform baseline methods on both link prediction and graph clustering tasks. Di Wu 0050, Ruiqi Hu, Yu Zheng 0013, Jing Jiang 0002, Nabin Sharma, Michael Blumenstein |
IJCNN | 1 |
| 2017 | Recent advances in video-based human action recognition using deep learning: A reviewabstractVideo-based human action recognition has become one of the most popular research areas in the field of computer vision and pattern recognition in recent years. It has a wide variety of applications such as surveillance, robotics, health care, video searching and human-computer interaction. There are many challenges involved in human action recognition in videos, such as cluttered backgrounds, occlusions, viewpoint variation, execution rate, and camera motion. A large number of techniques have been proposed to address the challenges over the decades. Three different types of datasets namely, single viewpoint, multiple viewpoint and RGB-depth videos, are used for research. This paper presents a review of various state-of-the-art deep learning-based techniques proposed for human action recognition on the three types of datasets. In light of the growing popularity and the recent developments in video-based human action recognition, this review imparts details of current trends and potential directions for future work to assist researchers. Di Wu 0050, Nabin Sharma, Michael Blumenstein |
IJCNN | 1 |
| 2015 | Detecting stepping stones by abnormal causality probabilityabstractAbstract Locating the real source of the Internet attacks has long been an important but difficult problem to be addressed. In the real world, attackers can easily hide their identities and evade punishment by relaying their attacks through a series of compromised systems or devices called stepping stones. Currently, researchers mainly use similar features from the network traffic, such as packet timestamps and frequencies, to detect stepping stones. However, these features can be easily destroyed by attackers using evasive techniques. In addition, it is also difficult to implement an appropriate threshold of similarity that can help justify the stepping stones. In order to counter these problems, in this paper, we introduce the consistent causality probability to detect the stepping stones. We formulate the ranges of abnormal causality probabilities according to the different network conditions, and on the basis of it, we further implement to self‐adaptive methods to capture stepping stones. To evaluate our proposed detection methods, we adopt theoretic analysis and empirical studies, which demonstrate accuracy of the abnormal causality probability. Moreover, we compare our proposed methods with previous works. The result shows that our methods in this paper significantly outperform previous works in the accuracy of detection malicious stepping stones, even when evasive techniques are adopted by attackers. Copyright © 2014 John Wiley & Sons, Ltd. Sheng Wen, Di Wu 0050, Ping Li 0019, Yang Xiang 0001, Wanlei Zhou 0001, Guiyi Wei |
Secur. Commun. Networks | 2 |
| 2014 | On Addressing the Imbalance Problem: A Correlated KNN Approach for Network Traffic Classification
Di Wu 0050, Xiao Chen 0002, Chao Chen 0015, Jun Zhang 0010, Yang Xiang 0001, Wanlei Zhou 0001 |
NSS | 1 |
| 2011 | A Survey on Latest Botnet Attack and DefenseabstractA botnet is a group of compromised computers, which are remotely controlled by hackers to launch various network attacks, such as DDoS attack and information phishing. Botnet has become a popular and productive tool behind many cyber attacks. Recently, the owners of some botnets, such as storm worm, torpig and conflicker, are employing fluxing techniques to evade detection. Therefore, the understanding of their fluxing tricks is critical to the success of defending from botnet attacks. Motivated by this, we survey the latest botnet attacks and defenses in this paper. We begin with introducing the principles of fast fluxing (FF) and domain fluxing (DF), and explain how these techniques were employed by botnet owners to fly under the radar. Furthermore, we investigate the state-of-art research on fluxing detection. We also compare and evaluate those fluxing detection methods by multiple criteria. Finally, we discuss future directions on fighting against botnet based attacks. Shui Yu 0001, Di Wu 0050, Paul A. Watters |
TrustCom | 3 |