Gaolei Li

dblp:176/5848 · DBLP profile ↗
← Back
95ranked-venue papers
6as first author
84since 2021 · last 2026
0000-0003-3913-5001ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 38 · 3 first-author · 30 since 2021Security and privacy · 21 · 21 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 9 · 9 since 2021Systems, architecture and hardware · 9 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 9 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Splats in Splats: Robust and Effective 3D Steganography Towards Gaussian Splatting
abstract
3D Gaussian splatting (3DGS) has demonstrated impressive 3D reconstruction performance with explicit scene representations. Given the widespread application of 3DGS in 3D reconstruction and generation tasks, there is an urgent need to protect the copyright of 3DGS assets. However, existing copyright protection techniques for 3DGS overlook the usability of 3D assets, posing challenges for practical deployment. Here we describe splats in splats, the first 3DGS steganography framework that embeds 3D content in 3DGS itself without modifying any attributes. To achieve this, we take a deep insight into spherical harmonics (SH) and devise an importance-graded SH coefficient encryption strategy to embed the hidden SH coefficients. Furthermore, we employ a convolutional autoencoder to establish a mapping between the original Gaussian primitives' opacity and the hidden Gaussian primitives' opacity. Extensive experiments indicate that our method significantly outperforms existing 3D steganography techniques, with 5.31% higher scene fidelity and 3x faster rendering speed, while ensuring security, robustness, and user experience.
Yijia Guo, Wenkai Huang 0003, Gaolei Li, Hang Zhang 0010, Liwen Hu 0002, Jianhua Li 0001, Tiejun Huang 0001, Lei Ma 0008
AAAI4
2026 Can Protective Watermarking Safeguard the Copyright of 3D Gaussian Splatting?
abstract
3D Gaussian Splatting (3DGS) has emerged as a powerful representation for 3D scenes, widely adopted due to its exceptional efficiency and high-fidelity visual quality. Given the significant value of 3DGS assets, recent works have introduced specialized watermarking schemes to ensure copyright protection and ownership verification. However, can existing 3D Gaussian watermarking approaches genuinely guarantee robust protection of the 3D assets? In this paper, for the first time, we systematically explore and validate possible vulnerabilities of 3DGS watermarking frameworks. We demonstrate that conventional watermark removal techniques designed for 2D images do not effectively generalize to the 3DGS scenario due to the specialized rendering pipeline and unique attributes of each gaussian primitives. Motivated by this insight, we propose GSPure, the first watermark purification framework specifically for 3DGS watermarking representations. By analyzing view-dependent rendering contributions and exploiting geometrically accurate feature clustering, GSPure precisely isolates and effectively removes watermark-related Gaussian primitives while preserving scene integrity. Extensive experiments demonstrate that our GSPure achieves the best watermark purification performance, reducing watermark PSNR by up to 16.34dB while minimizing degradation to original scene fidelity with less than 1dB PSNR loss. Moreover, it consistently outperforms existing methods in both effectiveness and generalization.
Wenkai Huang 0003, Yijia Guo, Gaolei Li, Lei Ma 0008, Hang Zhang 0010, Liwen Hu 0002, Jiazheng Wang 0001, Jianhua Li 0001, Tiejun Huang 0001
AAAI3
2026 Ownership-Protected Semantic Communication via Signal Processing-Driven Robust Watermark
Xiao Yang 0016, Gaolei Li, Zhaohui Yang 0001, Yuchen Liu 0001, Jianhua Li 0001
ICC3
2026 LSFL: A Lightweight and Secure Federated Learning scheme for Internet of Vehicles
Dan Peng, Lei Shi 0001, Gaolei Li, Huijuan Lian
Inf. Process. Manag.5
2026 ReSLC: Defending backdoor attacks on intelligent vulnerability detection via redundant semantic LLM compression
Gaolei Li, Jin Pang, Jianhua Li 0001
J. Inf. Secur. Appl.4
2026 Persistent Clean-Label Backdoor Attacks on Semisupervised Social Graph Node Classification
abstract
Semisupervised social graph node classification (SSGNC) attempts to deduce node-related information of social graph with limited labeled training samples. It is primarily deployed in large-scale graph processing, e.g., malicious client detection, knowledge graph, and recommender system. However, in this article, we identify that the SSGNC model is also extremely sensitive to backdoor attacks. We present a novel persistent clean-label backdoor attack (PerCBA) on SSGNC, which selectively poisons unmarked training nodes before learning to compel the trained model to misclassify trigger-embedded inputs into malicious class. Specifically, PerCBA employs a style-agnostic trigger generator with adjustable perturbation strategy to produce perturbed triggers. These triggers are pasted onto a small subset of unmarked nodes ($< \, 4\%$), enabling the adversary to covertly poison the training graph and implant backdoors into the model without modifying labels. Additionally, to ensure SSGNC robustness when confronted with homogenous threats, we present a testing sample filtering-based defense strategy for PerCBA. It employs feature distribution to identify poisoned nodes and applies Gaussian blur and thresholding to remove the trigger fraction, thereby restoring suspicious data to clean states. Extensive experiments on SOTA SSGNC models and datasets indicate that PerCBA performs high attack success rates (maxima 96.25%) while remaining evasive, and the defense method can effectively mitigate attacks and purify backdoored models.
Xiao Yang 0016, Gaolei Li, Xinzheng Feng, Xiaoyu Yi 0003, Jianhua Li 0001
IEEE Trans. Comput. Soc. Syst.2
2026 AgentChain: Blockchain-Empowered Multi-Agent Coordination for Trustworthy LLM Question-Answering Systems
abstract
Multi-agent architectures leveraging Large Language Models (LLMs) have significantly advanced the precision of Question Answering (QA) systems across diverse domains. However, existing frameworks remain vulnerable to adversarial manipulations, including poisoning, backdoor, and jailbreak at tacks, primarily due to their reliance on centralized orchestration. To mitigate these risks, we propose AgentChain, a framework that substitutes centralized control with a distributed semantic consensus process. By modeling the blockchain as an ideal functionality, AgentChain establishes a secure distributed layer to coordinate role allocation, answer proposal, evaluation and voting through a decentralized council. Specifically, we design Proof-of-Content-Quality (PoCQ) mechanism to ensure that the f inal answers reflect a robust semantic agreement among the majority of honest agents. Furthermore, we propose an incentive mechanism based on stake reassignment that penalizes malicious agents by reducing their rewards, ultimately phasing them out of the network. Comprehensive evaluations across eight datasets demonstrate that AgentChain achieves superior performance and resilience. AgentChain minimizes the impact of poisoning attacks on precision to less than 3% and reduces the success rate of backdoor and jailbreak attacks to less than 4%. These findings highlight the effectiveness and trustworthiness of AgentChain in mitigating security threats while maintaining high QA accuracy.
Bei Chen 0004, Gaolei Li, Jun Wu 0001, Jianhua Li 0001, Mingzhe Chen, Jiacheng Wang 0001
IEEE Trans. Dependable Secur. Comput.2
2026 HGAFA: Heterogeneous Graph Attention-Based Featureless Aggregation for IoC Joint Identification
abstract
Malicious cyber activities can potentially be detected through indicators of compromise (IoCs). As attacks become more complex, IoCs can be increasingly interconnected; thus, motivating the use of graph-based modeling. However, current approaches face three key challenges: limited and small-scale benchmarks that hinder industrial applicability, reliance on expert-designed meta-paths that restricts generalization in heterogeneous graphs, and insufficient interpretability, which increases the cost of verifying false positives. To address these challenges, we propose a web-scale IoC heterogeneous graph (IoCHG) that models domains, files, IPs, and URLs with seven interaction types, constructed through malware sandbox execution and open-source threat intelligence. Building on IoCHG, we develop Heterogeneous Graph Attention-based Featureless Aggregation (HGAFA) to support joint IoC identification. HGAFA leverages node and edge attention to capture IoC subgraph structures without meta-paths or hand-crafted features, thereby reducing reliance on expert knowledge. Our approach further improves interpretability through edge masking. To our knowledge, this is the first approach to model large-scale IoCs and identify malicious IoCs without expert-designed features. Experiments on millions of nodes from an industrial dataset show that HGAFA outperforms five competing approaches by an average of 8% in precision, while its interpretable subgraphs assist security experts in analyzing attack scenarios.
Hongjie Gu, Daojing He, Jialun Cao, Gaolei Li, Kim-Kwang Raymond Choo
IEEE Trans. Dependable Secur. Comput.4
2026 Adversarial Robustness of Link Sign Prediction in Signed Graphs
abstract
Signed graphs serve as fundamental data structures for representing positive and negative relationships in social networks, with signed graph neural networks (SGNNs) emerging as the primary tool for their analysis. Our investigation reveals that balance theory, while essential for modeling signed relationships in SGNNs, inadvertently introduces exploitable vulnerabilities to black-box attacks. To showcase this, we propose balance-attack, a novel adversarial strategy specifically designed to compromise graph balance degree, and develop an efficient heuristic algorithm to solve the associated NP-hard optimization problem. While existing approaches attempt to restore attacked graphs through balance learning techniques, they face a critical challenge we term “Irreversibility of Balance-related Information,” as restored edges fail to align with original attack targets. To address this limitation, we introduce Balance Augmented-Signed Graph Contrastive Learning (BA-SGCL), an innovative framework that combines contrastive learning with balance augmentation techniques to achieve robust graph representations. By maintaining high balance degree in the latent space, BA-SGCL not only effectively circumvents the irreversibility challenge but also significantly enhances model resilience. Extensive experiments across multiple SGNN architectures and real-world datasets demonstrate both the effectiveness of our proposed balance-attack and the superior robustness of BA-SGCL, advancing the security and reliability of signed graph analysis in social networks. Datasets and codes of the proposed framework are at the github repositoryhttps://github.com/JialongZhou666/BA-SGCL.git.
Jialong Zhou, Xing Ai, Yuni Lai, Tomasz P. Michalak, Gaolei Li, Jianhua Li 0001, Mengpei Yang, Kai Zhou 0001
IEEE Trans. Dependable Secur. Comput.5
2026 Revisiting Adversarial Robustness of GNNs Against Structural Attacks: A Simple and Fast Approach
abstract
To defend against adversarial structural attacks on graphs, we analyze attacks through the lens of mutual information and discover the “pairwise effect". This effect reveals that structural attacks effectively degrade the performance of victim GNNs when these GNNs receive the modified structure paired with the given node attributes as training input. Therefore, we propose a novel defense strategy that renders structural attacks ineffective by disrupting the pairing of modified structures and node attributes during the training of victim GNNs, which we call “disrupting the pairwise effect". To implement this idea, we propose two simple yet effective training strategies: Structural Fine-Tuning (SF) and Progressive Structural Training (PST), which disrupt the pairwise effect through node attributes pre-training followed by structure fine-tuning and progressive structure training, respectively. Compared to existing robust GNNs, our strategies avoid time-consuming techniques, thereby improving the robustness of GNNs while enhancing training speed. Additionally, these strategies can be easily applied to a wide range of commonly used GNNs, including robust GNN variants, making them highly adaptable to different models and applications. We provide theoretical analysis of the proposed training strategies and conduct extensive experiments on various datasets to demonstrate their effectiveness. Datasets and codes of this paper are available at https://github.com/Xing-Ai1003/Revisiting-Adversarial-Robustness-of-GNNs.
Xing Ai, Yulin Zhu 0001, Yu Zheng 0021, Gaolei Li, Jianhua Li 0001, Kai Zhou 0001
IEEE Trans. Inf. Forensics Secur.4
2026 PromptFishing: Active Hallucination Inducement to Distinguish LLMs From Humans
Bei Chen 0004, Gaolei Li, Jun Wu 0001, Jianhua Li 0001, He Fang
IEEE Trans. Inf. Forensics Secur.2
2026 BPF-DAG: Byte-Packet-Flow Features Fusion via Dynamic Attributed Graph for Reliable Encrypted Traffic Classification
abstract
Reliable encrypted traffic classification is crucial for fine-grained and efficient network security management, enabling accurate user behavior recognition and cybercrime forensics. While AI-based methods can automatically extract subtle features from traffic data, existing approaches often fail to effectively capture and integrate features across different levels of traffic granularity, namely the byte, packet and flow levels. Current graph-based methods heavily rely on manual feature engineering to construct global IP-based graphs, overlooking critical packet-level temporal features and byte-level raw information. Focusing on only one or two levels of traffic granularity is unreliable and insufficient, ultimately compromising model accuracy and robustness. To address these limitations, we propose BPF-DAG, a byte-packet-flow feature fusion framework based on dynamic attributed graphs, for reliable encrypted traffic classification. To the best of our knowledge, this is the first method that integrates temporal packet relations into flow interaction patterns while directly leveraging raw byte-level data. Specifically, we introduce a multi-granularity feature fusion strategy that dynamically updates an IP-based graph by iteratively assigning edge attributes derived from evolving flow representations. During the joint training of the Transformer and the graph neural network, temporal representations are learned from raw packet sequences and reflected in edge attributes dynamically for further message aggregation. Experiments on the ISCX VPN-nonVPN, Tor-nonTor, MIRAGE-2019 and MIRAGE-2024 datasets show that BPF-DAG outperforms recent state-of-the-art methods in terms of classification performance.
Yunxiao Shi, Gaolei Li, Jun Wu 0001, Jianhua Li 0001, He Fang
IEEE Trans. Inf. Forensics Secur.2
2026 Toward Polymorphic Backdoor Against Semantic Communication via Intensity-Based Poisoning
Xiao Yang 0016, Yuni Lai, Gaolei Li, Jun Wu 0001, Kai Zhou 0001, Jianhua Li 0001, Mingzhe Chen
IEEE Trans. Inf. Forensics Secur.3
2026 Beamforming Feedback-Driven Wireless Positioning: A Transferable Vision Transformer Approach
abstract
WiFi-based indoor positioning plays a crucial role in a variety of location-based services due to its widespread avail ability and cost-effectiveness. However, most existing indoor positioning systems predominantly utilize channel state information (CSI) to learn channel characteristics and apply fingerprinting for position estimation. Unfortunately, CSI can only be extracted from a limited set of commercial WiFi devices, hindering its widespread application in practice. In this work, we introduce BFMLoc, a novel indoor positioning framework that exploits the beamforming feedback matrix (BFM), which is readily available on commercial WiFi devices. Although BFM provides broader sensing coverage, it sacrifices detailed channel information due to the data compression applied to reduce feedback overhead. To address this limitation, we explore the feasibility of using BFM derivatives for indoor positioning and propose a U-net model to reconstruct the angle-delay profiles (ADP) from the compressed BFM data, thereby enhancing positioning accuracy. A Vision Transformer (ViT) model is then developed to extract spatial features from the predicted ADP maps to perform localization. Additionally, we design a model adaptation module based on transfer learning, integrated into the overall framework. This allows the positioning model to be easily deployed and adapted to various indoor environments with minimal retraining overhead. Extensive evaluations and validation on a digital twin testbed demonstrate that our framework achieves high positioning ac curacy and enhanced robustness compared to state-of-the-art methods.
Zhizhen Li, Xuanhao Luo, Mingzhe Chen, Gaolei Li, Yuchen Liu 0001
IEEE Trans. Mob. Comput.4
2026 Agile, Reliable and Communication-Efficient Metaverse 3D Reconstruction Via Gaussian Semantic Splatting
abstract
3D reconstruction is a cornerstone for creating immersive digital experiences in metaverse. Owing to explicit scene representation and efficient rendering, Gaussian splatting (GS) has become a prominent research focus in 3D reconstruction. However, the input images for GS are often imperfect, as those collected via highly-interfered wireless environment (HIWE) tend to be distorted, thereby undermining the accuracy of 3D reconstruction and limiting scalability. This paper proposes a novel Gaussian semantic splatting (GSS) scheme, designed for agile, reliable, and communication-efficient 3D reconstruction in the metaverse. Specifically, the semantic communication encoder/decoder (SCED) within GSS performs sequential semantic encoding and channel encoding using the proposed reliable and efficient semantic communication (RESC) algorithm, enabling the receiver to recover images with near-perfect accuracy. These images are then processed by the memory-efficient Gaussian renderer (MEGR), which employs an agile Gaussian splatting rendering (AGSR) algorithm to complete the 3D reconstruction and render a series of new viewpoint images. Additionally, a semantic control unit (SCU) is designed to oversee the components, enhancing the overall efficiency of the 3D reconstruction process. Experimental results demonstrate that GSS achieves competitive 3D reconstruction quality in HIWE, delivering real-time rendering speeds of 145 FPS at an$800\times 800$resolution while reducing storage memory overhead by more than 12 times compared to the state-of-the-art (SOTA) scheme.
Gaolei Li, Changze Li, Jianhua Li 0001, Yuchen Liu 0001, Mingzhe Chen
IEEE Trans. Mob. Comput.2
2026 SemanAegis: Toward Credential-Aware Semantic Communication Against Knowledge Leakage Threats
abstract
Semantic Communication (SC) achieves meaning transmission instead of bitstreams by deep semantic encoding decoding. Since the encoder-decoder contains sensitive and proprietary knowledge, its illicit leakage infringes commercial benefits and copyright, which warrants corresponding protection. However, current SC security paradigms narrowly emphasize transmission data protection while neglecting encoding knowl edge safeguarding. To bridge this gap, we present SemanAegis, the first SC knowledge protection framework. SemanAegisinte grates a built-in-system access control mechanism that remains effective even if the system is stolen, ensuring that unauthorized access attempts yield unacceptable low-fidelity outputs, while credential-embedded inputs from authorized entities are met with accurate responses. Specifically, we establish access control through backdoor implantation, whereby only inputs embedded with credentials activate the backdoor and access system, while source inputs are constrained to generate erroneous results. Moreover, we adopt a synthesizer to generate imperceptible credentials, thus guaranteeing their confidentiality. Additionally, a dedicated contrastive learning strategy is implemented to accelerate the convergence of backdoor implanting. Empirical evaluations across SC systems and benchmark datasets demonstrate SemanAegis precisely rejects unauthorized inputs, effectively mitigates knowledge extractions, and consistently preserves SC regular functionality.
Xiao Yang 0016, Yuni Lai, Gaolei Li, Jun Wu 0001, Kai Zhou 0001, Mingzhe Chen
IEEE Trans. Mob. Comput.3
2026 Neural Optimization for Image Registration via Joint Modeling of Global Affine and Local Deformation Transformations
abstract
Conventional registration approaches frequently underperform when applied to sparse feature alignment (e.g., retinal vessels and filamentous collagen fibers in second-harmonic generation (SHG) and bright-field (BF) images), as these tasks demand simultaneous handling of global affine registration and local deformation correction. End-to-end learning-based approaches struggle with minimal effective gradients from loss back-propagation of these sparse features, while descriptor matching methods, though helpful, lack fidelity loss and fail to adapt to local deformation. To address these issues, we propose Neural Affine Optimization (NeOn), which implicitly approximates discrete optimization using a few neural network layers, combined with a sampling-regression layer to handle affine transformations. NeOn allows iterative refinement with fidelity loss and provides a flexible transition between a purely affine configuration and a linear weighted blend of affine and deformation fields. NeOn's performance was validated on four public datasets. In multi-modal SHG-BF microscopy registration, NeOn achieved top rankings on the validation leaderboard for Task 3 of the Learn2Reg Challenge 2024. For retinal image registration, NeOn outperformed existing methods on both mono-modal and multi-modal datasets, reducing target registration error from 6.3 to 2.1 pixels in mono-modal and from 2.6 to 1.8 pixels in multi-modal registration. Furthermore, NeOn demonstrates strong generalization and can be effectively extended to 3D multi-modality image registration scenarios.
Xiang Chen 0008, Renjiu Hu, Jiacheng Wang 0001, Min Liu 0008, Yaonan Wang 0001, Jiazheng Wang 0001, Rongguang Wang, Gaolei Li, Hang Zhang 0010
IEEE Trans. Medical Imaging8
2025 WatCOM: Unconscious Watermarking for Semantic Communication Intellectual Property Protection
abstract
Semantic Communication (SC) enhances communication efficacy by abstracting and decoding semantic information via shared knowledge instead of bitstream, while considerably reducing redundancy and reinforcing efficiency in downstream tasks including image recognition, language processing, internet of things, etc. Due to the extensive data collection, processing, and training, the SC shared knowledge is invaluable Intellectual Property (IP), and despite the owners' desire to prevent misuse, the knowledge still remains vulnerable to theft while related IP protection has yet to be explored. To bridge this gap, we propose WatCom, the first SC IP protection methodology via watermarking. Specifically, we implant a stealthy backdoor into the semantic shared knowledge to verify model ownership, which can solely be activated by the owner-exclusive implicit trigger to validate ownership. The backdoor is implanted by poisoning-training strategy, facilitating SC system to respond normally to regular inputs while producing verification outputs (i.e., backdoor activation) for trigger-infected samples. To ensure imperceptibility, we leverage one generator to synthesize infected data that is nearly indistinguishable from regular data, which thereby obfuscates the verification information presence and enhances security against adversarial detection. Experiments based on multiple datasets and systems demonstrate WatCom can effectively verify system ownership (IP Verification Rate$\sim 100 \%$) while maintaining transmission efficacy (Peak Signal-to-Noise Ratio drop$< 2 ~\text{dB}$).
Xiao Yang 0016, Yuanhang He, Gaolei Li, Jianhua Li 0001
ICC3
2025 3D-MGW: A Memory-Efficient Grouped Watermark for Multi-Object 3D Gaussian Splatting
abstract
Multi-object 3D Gaussian Splatting (3DGS) technology aims to efficiently synthesize complex 3D scenes from images while allowing users to manipulate objects through textual prompts. Training multi-object 3DGS models requires substantial computational resources, making it necessary to protect the generated 3D objects from unauthorized reproduction, modification, and distribution. Existing watermarking solutions suffer from high memory consumption and require additional time overhead. Moreover, they cannot precisely localize watermarks to specific objects, making it difficult to trace individual contributions when multiple creators collaborate on a same multi-object scene. To address these challenges, we propose 3D-MGW, a novel memory-efficient grouped watermark for multi-object 3DGS. Within 3D-MGW, background scenes are reconstructed from images, while diffusion models are incorporated to guide highquality 3DGS synthesis from prompts. To eliminate additional training overhead, watermark embedding is integrated within the 3DGS training process rather than implementing it separately. Subsequently, a grouped Gaussian strategy is introduced to enable granular, high-capacity multi-object watermark. Additionally, a Gaussian compression module is proposed to eliminate redundant primitives, reducing the storage footprint of Gaussian models. Through comprehensive experiments, our 3D-MGW demonstrates exceptional performance, achieving 95% watermark extraction accuracy under 64-bit watermarks while reducing storage utilization by up to 61%, highlighting its substantial potential for multi-object 3DGS applications.
Hui Su, Gaolei Li, Wenkai Huang 0003, Xiaoyu Yi 0003, Jianhua Li 0001
ICPADS2
2025 GraphProt: Certified Black-Box Shielding Against Backdoored Graph Models
abstract
Graph learning models have been empirically proven to be vulnerable to backdoor threats, wherein adversaries submit trigger-embedded inputs to manipulate the model predictions. Current graph backdoor defenses manifest several limitations: 1) dependence on model-related details, 2) necessitation of additional fine-tuning, and 3) reliance on extra explainability tools, all of which are infeasible under stringent privacy policies. To address those limitations, we propose GraphProt, a certified black-box defense method to suppress backdoor attacks on GNN-based graph classifiers. Our GraphProt operates in a model-agnostic manner and solely leverages graph input. Specifically, GraphProt first introduces designed topology-feature-filtration to mitigate graph anomalies. Subsequently, subgraphs are sampled via a formulated strategy integrating topology and features, followed by a robust model inference through a majority vote-based subgraph prediction ensemble. Our results across benchmark attacks and datasets show GraphProt effectively reduces attack success rates while preserving regular graph classification accuracy.
Xiao Yang 0016, Yuni Lai, Kai Zhou 0001, Gaolei Li, Jianhua Li 0001, Hang Zhang 0010
IJCAI4
2025 Gaussian Primitive Optimized Deformable Retinal Image Registration
Jiazheng Wang 0001, Xiang Chen 0008, Renjiu Hu, Gaolei Li, Min Liu 0008, Hang Zhang 0010
MICCAI (4)6
2025 VoxelOpt: Voxel-Adaptive Message Passing for Discrete Optimization in Deformable Abdominal CT Registration
Hang Zhang 0010, Jiazheng Wang 0001, Xiang Chen 0008, Renjiu Hu, Gaolei Li, Min Liu 0008
MICCAI (4)7
2025 Semantic-Graph-Indistinguishability: A Novel Approach to Location Privacy Protection Under Road Networks
abstract
The core challenge in location privacy protection for location-based services (LBS) remains balancing location privacy and data utility. Differential privacy, backed by mathematical proofs, offers an effective framework for location protection. However, existing extended differential privacy methods for this purpose have some limitations. On one hand, most such methods focus on Euclidean spaces, making them ill-suited for road network-based LBS. They fail to align with road network contexts, potentially disrupting path planning, and often introduce excessive noise that inflates distance loss and degrades service quality. On the other hand, location semantics, a critical component of data utility, are frequently overlooked. Their degradation directly undermines utility. To address these issues, this paper introduces Semantic-Graph-Indistinguishability (SEM-G-IND) within the differential privacy paradigm, aiming to enhance location protection under road network and semantic constraints. First, a POI-based location semantic hash is designed to quantify location semantics. Then, integrating semantic distance and shortest path distance, a novel graph-based metric, Semantic-Graph-Distance (SGD), is proposed to measure inter-location distances. Finally, based on SGD, we propose a semantic penalty-based differential privacy (SPDP) location protection mechanism under road networks that satisfies ϵ-SEM-G-IND. We validated on real-world datasets that the SPDP mechanism effectively reduces the semantic loss of locations while ensuring minimal path distance loss, and it is feasible in terms of time overhead.
Huijuan Lian, Lei Shi 0001, Gaolei Li
TrustCom4
2025 HyBiGraph: Toward Multi-Order Malicious Encrypted Traffic Classification via Hyper-Bipartite Graph Fusion
abstract
Malicious attacks frequently exploit encrypted traffic as a covert channel for intrusion, rendering the accurate identification of malicious encrypted traffic essential for early threat detection. Existing encrypted traffic classification methods primarily focus on low-order IP topological graph and single flow features. However, malicious IPs usually send encrypted flows mixed attack flows with benign traffic to hide themselves, while attack flows have high-order relations, leaving complex interactions between IP nodes and traffic flows. To address these limitations, we propose a novel Hyper-Bipartite Graph Fusion (HyBiGraph) framework for malicious encrypted traffic classification that integrates a bipartite graph for modeling low-order relationships and a hypergraph for propagating higher-order structural information. HyBiGraph constructs an IP-Flow bipartite graph with trainable IP embeddings updated via flow features, which enables the model to efficiently capture the contextual relationships between source and target IP. It further employs hypergraph attention with learnable hyperedges to power precise modeling of higher-order interactions among flows. Finally, residual fusion of hypergraph and bipartite graph offers a robust and efficient mechanism for integrating structural representations, enhancing classification performance. HyBiGraph was evaluated on benchmark encrypted malicious traffic datasets—USTC-TFC2016, CICIoT2023, and CICAndMal2017—attaining accuracy improvements of 2.51%, 23.93%, and 51.97% and requiring only approximately 10% training cost of baselines. Also, ablation studies validate that integrating hypergraph and bipartite graph promotes accuracy gains between 1.27% and 53.98%.
Yibin Zhou, Yunxiao Shi, Xiao Yang 0016, Gaolei Li, Jianhua Li 0001
TrustCom5
2025 Non-Orthogonal Multiple Access Based Multi-Objective Optimization in Emergency Task Offloading
abstract
The escalating frequency of geological disasters, local conflicts, and public health crises necessitates robust emergency communication systems within Internet of Things (IoT) frameworks to ensure public safety. Employing Power-Domain NOMA (PD-NOMA) and Successive Interference Cancellation (SIC) enables effective multi-device communication that is crucial for emergency responses. However, the practical implementation of NOMA technologies in emergency scenarios is hindered by existing resource constraints and the requirement of prioritizing critical emergency information. To address these limitations, we propose a multi-objective optimization-based NOMA framework to maximize network throughput while ensuring quality of service (QoS) and energy efficiency. This framework incorporates a novel communication grouping algorithm based on queue penalty optimization and an Actor-Critic-based dynamic updating algorithm for multiple resource allocation. Our study proposes a practical approach to enhance emergency communication in disaster areas, ensuring efficient and reliable network performance under various constraints.
Caijuan Chen, Gaolei Li, Yuchen Liu 0001
WCNC3
2025 Enhanced Resource Orchestration in Near-Field and Far-Field NOMA Communications for Emergency IoT Applications
Caijuan Chen, Gaolei Li
IEEE Internet Things J.4
2025 Anti-traceable backdoor: Blaming malicious poisoning on innocents in non-IID federated learning
Bei Chen 0004, Gaolei Li, Haochen Mei, Jianhua Li 0001, Mingzhe Chen, Mérouane Debbah
J. Inf. Secur. Appl.2
2025 Large language model-enhanced probabilistic modeling for effective static analysis alarms
abstract
Static analysis presents significant challenges in alarm handling, where probabilistic models and alarm prioritization are essential methods for addressing these issues. These models prioritize alarms based on user feedback, thereby alleviating the burden on users to manually inspect alarms. However, they often encounter limitations related to efficiency and issues such as false generalization. While learning-based approaches have demonstrated promise, they typically incur high training costs and are constrained by the predefined structures of existing models. Moreover, the integration of large language models (LLMs) in static analysis has yet to reach its full potential, often resulting in lower accuracy rates in vulnerability identification. To tackle these challenges, we introduce BinLLM, a novel framework that harnesses the generalization capabilities of LLMs to enhance alarm probability models through rule learning. Our approach integrates LLM-derived abstract rules into the probabilistic model, using alarm paths and critical statements from static analysis. This integration enhances the model’s reasoning capabilities, improving its effectiveness in prioritizing genuine bugs while mitigating false generalizations. We evaluated BinLLM on a suite of C programs and observed 40.1% and 9.4% reduction in the number of checks required for alarm verification compared to two state-of-the-art baselines, Bingo and BayeSmith, respectively, underscoring the potential of combining LLMs with static analysis to improve alarm management.
Xinlong Pan, Jianhua Li 0001, Zhi Hong Zhou, Gaolei Li, Xiuzhen Chen, Jun Wu 0001, Quanhai Zhang
Frontiers Inf. Technol. Electron. Eng.4
2025 Silent Penetrator: Breaching Cross-Domain Federated Fine-Tuning via Feature Shift-Induced Backdoor
abstract
To improve communication efficiency and handle data heterogeneity challenges in federated learning (FL), fine-tuning the pre-trained large models rather than training neural networks from scratch has received increasing attention in recent years, especially under cross-domain settings. However, such a cross-domain federated fine-tuning scenario opens up a broader attack surface for new threats, especially backdoors, posing significant security risks. Existing backdoor attacks mainly focus on label shift scenarios and use explicit triggers, which lack transferability and effectiveness in cross-domain settings, thereby exhibiting significant weaknesses. In this paper, we propose Silent Penetrator, an innovative penetration scheme tailored for cross-domain federated fine-tuning, which exploits a feature shift-induced backdoor to elicit specific symptoms in the trusted private data of targeted victims. In Silent Penetrator, the attacker can obtain a high-quality poisoned dataset by leveraging the available domain information as the text prompts for Stable Diffusion, and inject a domain-sensitive backdoor that can be unconsciously triggered by unmodified private data of the victims. To achieve stronger and more persistent penetration, we thoroughly explore the adversary’s configurable space and enhance our backdoor injection utilizing contrastive-enhanced boundary deviation and cross-domain predictive confrontation. Extensive experiments on three cross-domain datasets and four state-of-the-art federated fine-tuning frameworks validate the effectiveness of Silent Penetrator in successfully compromising target clients. Furthermore, our backdoor enhancement strategy improves the penetration accuracy by over 10% in most scenarios and significantly enhances the durability of the penetration compared to four state-of-the-art backdoor enhancement techniques.
Wenkai Huang 0003, Gaolei Li, Mingzhe Chen, Jianhua Li 0001, Haojin Zhu
IEEE Trans. Inf. Forensics Secur.2
2025 FeCoGraph: Label-Aware Federated Graph Contrastive Learning for Few-Shot Network Intrusion Detection
abstract
With increasing cyber attacks over the Internet, network intrusion detection systems (NIDS) have been an indispensable barrier to protecting network security. Taking advantage of automatically capturing topology connections, recent deep graph learning approaches have achieved remarkable performance in distinguishing different types of malicious flows. However, there remain some critical challenges. 1) previous supervised learning methods rely heavily on abundant and high-quality annotated samples, while label annotation requires abundant time and expert knowledge. 2) Centralized methods require all data to be uploaded to a server for learning behavior patterns, which results in high detection latency and critical privacy leakage. 3) Diverse attack scenarios exhibit highly imbalanced distribution, making it hard to characterize abnormal behaviors. To address these issues, we proposed FeCoGraph, a label-aware federated graph contrastive learning framework for intrusion detection in few-shot scenarios. The line graph is introduced to directly process flow embeddings, which are compatible with diverse GNNs. Furthermore, We formulate a graph contrastive learning task to effectively leverage label information, allowing intra-class embeddings more compact than inter-class embeddings. To improve the scalability of NIDS, we utilize federated learning to cover more attack scenarios while protecting data privacy. Experiment results show that FeCoGraph surpass E-graphSAGE with an average 8.36% accuracy on binary classification and 6.77% accuracy on multiclass classification, demonstrating the efficiency of our approach.
Qinghua Mao, Xi Lin 0003, Wenchao Xu 0001, Yuxin Qi 0001, Xiu Su, Gaolei Li, Jianhua Li 0001
IEEE Trans. Inf. Forensics Secur.6
2025 Spatially Covariant Image Registration With Text Prompts
abstract
Medical images are often characterized by their structured anatomical representations and spatially inhomogeneous contrasts. Leveraging anatomical priors in neural networks can greatly enhance their utility in resource-constrained clinical settings. Prior research has harnessed such information for image segmentation, yet progress in deformable image registration has been modest. Our work introduces textSCF, a novel method that integrates spatially covariant filters and textual anatomical prompts encoded by visual-language models, to fill this gap. This approach optimizes an implicit function that correlates text embeddings of anatomical regions to filter weights. textSCF not only boosts computational efficiency but can also retain or improve registration accuracy. By capturing the contextual interplay between anatomical regions, it offers impressive interregional transferability and the ability to preserve structural discontinuities during registration. textSCF's performance has been rigorously tested on intersubject brain magnetic resonance imaging (MRI) and abdominal computerized tomography (CT) registration tasks, outperforming existing state-of-the-art models in the MICCAI Learn2Reg 2021 challenge and leading the leaderboard. In abdominal registrations, textSCF's larger model variant improved the Dice score by 11.3% over the second-best model, while its smaller variant maintained similar accuracy but with an 89.13% reduction in network parameters and a 98.34% decrease in computational operations.
Xiang Chen 0008, Min Liu 0008, Rongguang Wang, Renjiu Hu, Gaolei Li, Yaonan Wang 0001, Hang Zhang 0010
IEEE Trans. Neural Networks Learn. Syst.6
2025 HFL-RD: Heterogeneous Federated Learning-Empowered Ransomware Detection via APIs and Traffic Features
abstract
Ransomware has evolved into a more organized attack threat with stronger anti detection and analysis capabilities, resulting in significant global losses. However, traditional methods separate the external and internal behaviors of ransomware infiltration into attack targets, making it difficult to discover the complex and covert evolution and iteration characteristics of advanced ransomware. The main contribution of this study lies in three aspects: a) The integration of Command-and-control (C&C) traffic behavior analysis and local API call operation analysis can effectively discern and capture the concealed characteristics of ransomware; b) The non-IID problem in aggregating ransomware features using federated learning can be resolved using dynamic regularization methods and penalty terms; c) By preprocessing the original data of ransomware traffic through one-dimensional convolution, the structural characteristics of network traffic in the process of attack operation can be retained to the greatest extent. Comprehensive experiments are conducted to validate the effectiveness of this model, specifically, the heterogeneous federated learning-empowered ransomware detection (HFL-RD) scheme outperformed existing methods, the experimental dataset gathered runnable ransomware from three public websites, including 300 ransomware samples from 30 families and 200 benign software samples from 7 categories. HFL-RD obtained a high accuracy over 95%. In terms of detecting unknown ransomware variants, it has demonstrated superior detection capabilities in terms of detection time and number of file corruption.
Lan Kun, Gaolei Li, Wenkai Huang 0003, Jianhua Li 0001
IEEE Trans. Netw. Serv. Manag.2
2025 Toward Covert and Reliable Communication for Anti-Eavesdropping Transmission in V2X Networks
abstract
The integration of covert communication in vehicle-to-everything (V2X) network has recently shown great potential to improve efficiency and reliability of data transmission under adversarial eavesdropping scenarios. In this paper, we propose a covert and reliable communication (CRC) framework for V2X networks, where the legitimate transmitter (Alice) attempts to communicate with a mobile receiver (Bob) in the presence of the location uncertainties of the eavesdropper (Willie). Specifically, the Bob adjusts the artificial noise power and position dynamically to communicate with Alice aided by full duplex antenna. In this context, we derive two key performance indicators of covert communication, namely the detection error probability and the effect covert throughput (ECT). Subsequently, we consider the worst case of CRC in the presence of single uncertain Willie, and derive the approximate maximum ECT expression by two-stage robust optimization. Building on this foundation, for more complex CRC scenario with multi uncertain Willies exist, we propose a deep reinforcement learning-empowered adaptation (DRLA) algorithm to maximize accumulated ECT. Extensive experiments compared to benchmarks (including stochastic selection, TD3 and DDPG) demonstrate the superiority of CRC. Specially, the designated DRLA algorithm not only can achieve a higher accumulated ECT but also can converge quickly compared with the benchmark schemes.
Gaolei Li, Jun Wu 0001, Jianhua Li 0001, Yue Zhao 0010, Yuchen Liu 0001, Mingzhe Chen
IEEE Trans. Wirel. Commun.2
2024 What Makes Good Collaborative Views? Contrastive Mutual Information Maximization for Multi-Agent Perception
abstract
Multi-agent perception (MAP) allows autonomous systems to understand complex environments by interpreting data from multiple sources. This paper investigates intermediate collaboration for MAP with a specific focus on exploring "good" properties of collaborative view (i.e., post-collaboration feature) and its underlying relationship to individual views (i.e., pre-collaboration features), which were treated as an opaque procedure by most existing works. We propose a novel framework named CMiMC (Contrastive Mutual Information Maximization for Collaborative Perception) for intermediate collaboration. The core philosophy of CMiMC is to preserve discriminative information of individual views in the collaborative view by maximizing mutual information between pre- and post-collaboration features while enhancing the efficacy of collaborative views by minimizing the loss function of downstream tasks. In particular, we define multi-view mutual information (MVMI) for intermediate collaboration that evaluates correlations between collaborative views and individual views on both global and local scales. We establish CMiMNet based on multi-view contrastive learning to realize estimation and maximization of MVMI, which assists the training of a collaborative encoder for voxel-level feature fusion. We evaluate CMiMC on V2X-Sim 1.0, and it improves the SOTA average precision by 3.08% and 4.44% at 0.5 and 0.7 IoU (Intersection-over-Union) thresholds, respectively. In addition, CMiMC can reduce communication volume to 1/32 while achieving performance comparable to SOTA. Code and Appendix are released at https://github.com/77SWF/CMiMC.
Wanfang Su, Lixing Chen, Yang Bai 0010, Xi Lin 0003, Gaolei Li, Pan Zhou 0001
AAAI5
2024 Backdoor NLP Models via AI-Generated Text
abstract
Backdoor attacks pose a critical security threat to natural language processing (NLP) models by establishing covert associations between trigger patterns and target labels without affecting normal accuracy. Existing attacks usually disregard fluency and semantic fidelity of poisoned text, rendering the malicious data easily detectable. However, text generation models can produce coherent and content-relevant text given prompts. Moreover, potential differences between human-written and AI-generated text may be captured by NLP models while being imperceptible to humans. More insidious threats could arise if attackers leverage latent features of AI-generated text as trigger patterns. We comprehensively investigate backdoor attacks on NLP models using AI-generated poisoned text obtained via continued writing or paraphrasing, exploring three attack scenarios: data, model and pre-training. For data poisoning, we fine-tune generators with attribute control to enhance the attack performance. For model poisoning, we leverage downstream tasks to derive specialized generators. For pre-training poisoning, we train multiple attribute-based generators and align their generated text with pre-defined vectors, enabling task-agnostic migration attacks. Experiments demonstrate that our method achieves effective attacks while maintaining fluency and semantic similarity across all scenarios. We hope this work can raise awareness of the security risks hidden in AI-generated text.
Tianjie Ju, Gaolei Li, Gongshen Liu
LREC/COLING4
2024 Trading Trust for Privacy: Socially-Motivated Personalized Privacy-Preserving Collaborative Learning in IoT
abstract
Nowadays, collaborative federated learning (CFL) is developing rapidly in the Internet of Things (IoT), which allows clients to jointly train models without compromising private data. The existing research has studied alone either trust enhancement or privacy preservation issues in CFL. Due to the highly coupled nature of trust and privacy in a collaborative environment, it is worth investigating how to balance appropriate trust and privacy tradeoffs for realizing high-quality CFL. In this paper, we come up with the idea of "trading Trust for Privacy", and propose a novel socially-motivated personalized privacy-preserving federated learning (SP-PFL) framework, which aims to realize social trust-grained privacy protection. First, we design a social trust evaluation method among CFL clients, which is based on topological relation and attribute similarity. Based on the obtained trust value, we then propose a trust-grained privacy budget allocation strategy for SP-PFL, which could further adaptively adjust the differential privacy (DP) noise perturbation. Besides, we provide an analysis of privacy and convergence for our SP-PFL. Finally, we experiment with different models and parameter settings on different datasets. Extensive experimental results show that our method maintains personalized privacy and effectively improves the accuracy by 6.11% on the CNN model and MNIST dataset.
Yuliang Chen, Xi Lin 0003, Gaolei Li, Lixing Chen, Siyi Liao, Jianhua Li 0001
CSCWD3
2024 Diffusion-based Reinforcement Learning for Dynamic UAV-assisted Vehicle Twins Migration in Vehicular Metaverses
abstract
Air-ground integrated networks can relieve communication pressure on ground transportation networks and provide 6G-enabled vehicular Metaverses services offloading in remote areas with sparse RoadSide Units (RSUs) coverage and downtown areas where users have a high demand for vehicular services. Vehicle Twins (VTs) are the digital twins of physical vehicles to enable more immersive and realistic vehicular services, which can be offloaded and updated on RSU, to manage and provide vehicular Metaverses services to passengers and drivers. The high mobility of vehicles and the limited coverage of RSU signals necessitate VT migration to ensure service continuity when vehicles leave the signal coverage of RSUs. However, uneven VT task migration might overload some RSUs, which might result in increased service latency, and thus impactive immersive experiences for users. In this paper, we propose a dynamic Unmanned Aerial Vehicle (UAV)-assisted VT migration framework in air-ground integrated networks, where UAVs act as aerial edge servers to assist ground RSUs during VT task offloading. In this framework, we propose a diffusion-based Reinforcement Learning (RL) algorithm, which can efficiently make immersive VT migration decisions in UAV-assisted vehicular networks. To balance the workload of RSUs and improve VT migration quality, we design a novel dynamic path planning algorithm based on a heuristic search strategy for UAVs. Simulation results show that the diffusion-based RL algorithm with UAV-assisted performs better than other baseline schemes.
Yongju Tong, Jiawen Kang 0001, Minrui Xu, Gaolei Li, Weiting Zhang, Xincheng Yan
GLOBECOM5
2024 Mixture Gaussian Distribution-Based Collaborative Reinforcement Learning for 3D UAV Localization Optimization Against Jamming Attacks
abstract
In this paper, the optimization of unmanned aerial vehicle (UAV) localization under jamming attacks is studied. In the considered network, a base station (BS) collaborates with an active UAV to localize a target UAV. During this positioning process, a jamming UAV transmits discontinuous signals to passive UAVs to interfere the distance information measurement. To localize the target UAV under jamming attacks, the BS jointly use two localization methods: 1) generative adversarial network (GAN)-based positioning method and 2) time difference of arrival (TDOA)-based positioning method. Since GAN-based positioning method cannot defense in a strong jamming signal while TDOA-based positioning method may consume more energy and sacrifice localization accuracy, the BS must select an appropriate positioning method (GAN-based or TDOA-based methods) and four distance measurement information of passive UAVs to estimate the position of the target UAV. This problem is formulated as an optimization problem whose goal is to minimize the positioning error between the estimated and the ground truth positions of the target UAV while considering jamming attacks and the trajectory of passive UAVs. To solve this problem, we propose a mixture Gaussian distribution model-based collaborative reinforcement learning (RL) method which enables the active UAV to determine its transmit power and trajectory, and enables the BS to select the most appropriate subsets of distance measurement information and the optimal positioning method according to the movement of passive UAVs and the unknown jamming attack pattern of the jamming UAV. Simulation results show the proposed method can reduce the positioning error of the target UAV by up to 36.5% compared to the method that does not consider the GAN-based positioning method.
Yujiao Zhu, Mingzhe Chen, Sihua Wang, Yuchen Liu 0001, Gaolei Li, Changchuan Yin, Tony Q. S. Quek
GLOBECOM5
2024 MKPL: Multi-dimensional Knowledge-embedded Prompt Learning for Few-shot Malware Family Recognition
abstract
Large language models (LLMs) bring great potential for next-generation malware family recognition with their capacity to understand complex code semantics by integrating multi-dimensional data features. However, existing fine-tuning methods still rely on well-labelled datasets and powerful computation resources, which is particularly challenging when the variety and amount of malware grow in real-time. To more effectively recognize unknown malware varieties based on LLMs, a novel multi-dimensional knowledge-embedded prompt learning (MKPL) framework is proposed, in which prompts are generated through two main steps: 1) cross-linguistic prompt paraphrasing (CPP) for embedding multi-dimensional knowledge into templates, and 2) prompt scoring for selecting the most effective prompt templates. Moreover, to reduce feature loss during prompt tuning, a sampling-infer-concatenation pipeline is designated to process these long API malware sequences. Specifically, a single-sentence template can be upgraded to a multi-sentence template by integrating statistic features into CPP, which is essential to improve the robustness of recognition results. Comprehensive experiments across eight malware families in few-shot scenarios demonstrate the proposed method’s superior performance in all metrics.
Shuilin Li, Gaolei Li, Xiaoyu Yi 0003, Jianhua Li 0001, Mianxiong Dong, Kaoru Ota
HPCC3
2024 Learning-Based DApp Task Scheduling for Elastic Hybrid Computing in Edge Web 3.0
abstract
Web 3.0 and Edge computing are inherently compatible, making them an ideal combination for building a secure and efficient distributed service platform to support decentralized applications (DApps). This paper investigates an elastic hybrid computing architecture in Edge Web 3.0, allowing DApp tasks to be executed in a hybrid manner by integrating on-chain and off-chain execution. The principle is to transfer a portion of DApp to an off-chain execution environment, along with an appropriate result verification process, to enhance computing efficiency and reduce blockchain overhead. We formulate a DApp task scheduling problem that jointly optimizes the execution pattern and offloading decision of user tasks. A learning-based DApp task scheduling scheme is designed based on Proximal Policy Optimization (PPO) to minimize the gas cost and service delay of DApps. Particularly, we tailor PPO to handle the hard constraints of service delay, gas consumption, and computing capacity in Edge Web 3.0 by adding regularization terms in the learning objective function. We establish an Edge Web 3.0 testbed based on Goerli, ZkSync, and Ethereum to evaluate the proposed method. The experimental results show that our method outperforms state-of-the-art benchmarks.
Xichun Cai, Lixing Chen, Yang Bai 0010, Xi Lin 0003, Gaolei Li, Jianhua Li 0001
ICC5
2024 Scale Wisely, Secure Wholly: P2P Swarm Learning Over Consortium Blockchain in Edge Networks
abstract
Swarm Learning (SL) provides a secure distributed learning environment to edge computing (EC) networks by leveraging blockchain technology for certified participation, encrypted information transmission, and immutable data storage. However, vanilla SL faces scalability limitations due to system-wide model aggregation, which bottlenecks its communication and blockchain efficiency. This paper presents a novel framework called Peer-to-peer Swarm Learning Over consOrtium blOckchain$(\mathbf{PSLO}_{3})$to enhance the scalability of SL over EC networks. PSLO3proposes a peer-to-peer swarm learning (P2P-SL) mechanism that only requires local communications for P2P model aggregation, thereby reducing the communication overhead of vanilla SL for system-wide model aggregation. Furthermore, PSLO3delivers P2P-SL over consortium blockchain and strategically organizes the edge servers into subchains to minimize the overhead of P2P-SL over consortium blockchain. A subchain formation scheme is designed based on graph partitioning by jointly analyzing the topological property of EC networks, message-passing patterns of P2P-SL, and the overhead of cross-/intra-chain interactions. An evaluation environment is built based on the Wecross platform to evaluate the performance of PSLO3. Experimental results demonstrate that PSLO3provides a reduction of 81.1% in communication overhead and a reduction of 26.03% in blockchain cost compared to vanilla swarm learning while demonstrating comparable learning performances.
Lixing Chen, Quanhai Zhang, Gaolei Li, Xi Lin 0003, Yang Bai 0010, Jianhua Li 0001
ICC4
2024 On-demand Quantization for Green Federated Generative Diffusion in Mobile Edge Networks
abstract
Generative Artificial Intelligence (GAI) shows remarkable productivity and creativity in Mobile Edge Networks, such as the metaverse and the Industrial Internet of Things. Federated learning is a promising technique for effectively training GAI models in mobile edge networks due to its data distribution. However, there is a notable issue with communication consumption when training large GAI models like generative diffusion models in mobile edge networks. Additionally, the substantial energy consumption associated with training diffusion-based models, along with the limited resources of edge devices and complexities of network environments, pose challenges for improving the training efficiency of GAI models. To address this challenge, we propose an on-demand quantized energy-efficient federated diffusion approach for mobile edge networks. Specifically, we first design a dynamic quantized federated diffusion training scheme considering various demands from the edge devices. Then, we study an energy efficiency problem based on specific quantization requirements. Numerical results show that our proposed method significantly reduces system energy consumption and transmitted model size compared to both baseline federated diffusion and fixed quantized federated diffusion methods while effectively maintaining reasonable quality and diversity of generated data.
Bingkun Lai, Jiawen Kang 0001, Gaolei Li, Minrui Xu, Tao Zhang 0063, Shengli Xie 0001
ICC4
2024 Graph Anomaly Detection at Group Level: A Topology Pattern Enhanced Unsupervised Approach
abstract
Graph anomaly detection (GAD) has achieved success and has been widely applied in various domains, such as fraud detection, cybersecurity, finance security, and biochemistry. However, existing graph anomaly detection algorithms focus on distinguishing individual entities (nodes or graphs) and overlook the possibility of anomalous groups within the graph. To address this limitation, this paper introduces a novel unsupervised framework for a new task called Group-level Graph Anomaly Detection (Gr-GAD). The proposed framework first employs a variant of Graph AutoEncoder (GAE) to locate anchor nodes that belong to potential anomaly groups by capturing long-range inconsistencies. Subsequently, group sampling is employed to sample candidate groups, which are then fed into the proposed Topology Pattern-based Graph Contrastive Learning (TPGCL) method. TPGCL utilizes the topology patterns of groups as clues to generate embeddings for each candidate group and thus distinct anomaly groups. The experimental results on both real-world and synthetic datasets demonstrate that the proposed framework shows superior performance in identifying and localizing anomaly groups, highlighting it as a promising solution for Gr-GAD. Datasets and codes of the proposed framework are at the github repository https://github.com/STiL-Team/Topology-Pattern-Enhanced-Unsupervised-Group-level-Graph-Anomaly-Detection.git.
Xing Ai, Jialong Zhou, Yulin Zhu 0001, Gaolei Li, Tomasz P. Michalak, Xiapu Luo, Kai Zhou 0001
ICDE4
2024 SegDaemon: Actively Protecting Semantic Segmentation Models Against Intellectual Property Infringement
Gaolei Li, Xiaoyang Jiang, Shiyong Qiu, Liangjie Liu, Shuilin Li, Shenghong Li 0001
ICONIP (8)2
2024 LateBA: Latent Backdoor Attack on Deep Bug Search via Infrequent Execution Codes
abstract
Backdoor attacks can mislead deep bug search models by exploring model-sensitive assembly code, which can change alerts to benign results and cause buggy binaries to enter production environments. But assembly instructions have strict constraints and dependencies, and these additional model-sensitive assembly codes destroy semantics and syntax and are easily detected by dynamic analysis or context-based detection. To escape from the dynamic analysis-based detection, we propose a novel latent backdoor attack (LateBA) scheme based on the locality principle of program execution, which only poisons a few of infrequent execution codes, minimizing the effects on the original code logic. In LateBA, a progressive seed mutating strategy is designated to change the American Fuzzy Lop (AFL)-based path search tool to pay more attention to infrequent execution codes. With this strategy, the optimal range to positions in the whole program is determined. Subsequently, triggers are target model-sensitive assembly instructions, and try to minimize the variables that have been called in the context instructions in the trigger. Finally, we employ code semantic feature comparisons to select precise trigger injection positions within these ranges. The selection criteria of the trigger injection position is whether the corresponding code segments in this position have a data dependency relationship with other code segments. We evaluate the performance of LateBA over 7 deep bug search tasks. The results demonstrate the attack success rate of the proposed LateBA is considerable and competitive against the baselines.
Xiaoyu Yi 0003, Gaolei Li, Wenkai Huang 0003, Xi Lin 0003, Jianhua Li 0001, Yuchen Liu 0001
Internetware2
2024 InviINS: Invisible Instruction Backdoor Attacks on Peer-to-Peer Semantic Networks
abstract
Recently, Peer-to-Peer Semantic Network (P2PSN) has significantly boosted transmission efficiency among humans, machine agents, and smart devices. Despite these enhancements, the intelligent components within P2PSN pose vulnerabilities to backdoor attacks, where adversaries introduce specific pattern triggers to poison the training set, which prompts the well-trained P2PSN system to generate targeted malicious predictions when inputted with trigger-embedded data. Current backdoor methodologies exhibit several deficiencies: 1) pattern-based trigger lacking physical meaning and explainability; 2) visible trigger design that can be easily detected by defenders; 3) unstable attack performance resulting from communication interference. To overcome these shortcomings, we propose a novel invisible instruction backdoor attack scheme on Peer-to-Peer Semantic Networks: InviINS. The proposed method embeds text instructions on partial training samples as invisible triggers instead of pattern triggers, thereby poisoning the training set of P2PSN model before learning without visually discernible changes in data, and subsequently backdooring the model via training. In InviINS, adversaries can directly set instructions based on practical scenarios to launch attacks. Meanwhile, to accelerate backdoor convergence, a contrastive backdoor training methodology is presented to enhance the model’s sensitivity to instruction triggers and bolster its prediction performance on normal samples. Experiments with different poisoning-rates, signal-to-noise ratios, channel usages, and trigger types demonstrate that the InviINS can achieve a high attack success rate (~ 100%) while preserving the model performance on main tasks (accuracy drop < 3%).
Xiao Yang 0016, Gaolei Li, Mianxiong Dong, Kaoru Ota, Jun Wu 0001, Jianhua Li 0001
ISPA2
2024 ActIPP: Active Intellectual Property Protection of Edge-Level Graph Learning for Distributed Vehicular Networks
abstract
Edge-Level Graph Learning System (EGLS) exhibits diverse applicability in management of distributed vehicular networks, e.g., flow prediction, route planning, and accident forecasting. For the EGLS training, expensive hardware resource consumption, traffic data collection, and dedicated training procedures make the learning algorithms become valuable intellectual property (IP) for the EGLS owner (e.g., Uber and Lyft), and they cannot tolerate the infringement act of their models’ intellectual property. To enhance its IP protection, we present ActIPP, the first active IP protection methodology for EGLS, which incorporates a built-in access control function in the model to safeguard against unauthorized queries. Specifically, it is achieved via a creative edge backdoor mechanism, wherein the edge training samples are poisoned via user-specific access tokens to induce legal outputs from a well-trained EGLS model for authorized users. Moreover, related token regulating strategies were proposed to dynamically realize the addition and revocation of user tokens by model retraining to guarantee access control in EGLS. Additionally, a Graph Mutual Information-based adaptive token generation method is presented to augment the access control embedding. Based on experiments with various real-world datasets, ActIPP demonstrates high success rates of IP protection (accuracy drop < 4%) under various scenarios and efficiently prevents unauthorized access (unauthorized access accuracy < 6%).
Xiao Yang 0016, Gaolei Li, Mianxiong Dong, Kaoru Ota, Xiting Peng, Jianhua Li 0001
ISPA2
2024 MemWarp: Discontinuity-Preserving Cardiac Registration with Memorized Anatomical Filters
Hang Zhang 0010, Xiang Chen 0008, Renjiu Hu, Gaolei Li, Rongguang Wang
MICCAI (3)5
2024 OSNeRF: On-demand Semantic Neural Radiance Fields for Fast and Robust 3D Object Reconstruction
abstract
By leveraging multi-view inputs to synthesize novel-view images, Neural Radiance Fields (NeRF) have emerged as a prominent technique in the realm of 3D object reconstruction. However, existing methods primarily focus on global scene reconstruction using large datasets, which necessitate substantial computational resources and impose high-quality requirements on input images. Nevertheless, in practical applications, users prioritize the 3D reconstruction results of on-demand specific object (OSO) based on their individual demands . Furthermore, the collected images transmitted through high-interference wireless environment (HIWE) leads to negatively impact the accuracy of NeRF reconstruction, thereby limiting its scalability. In this paper, we propose a novel on-demand Semantic Neural Radiance Fields (OSNeRF) scheme, which offers fast and robust 3D object reconstruction for diverse tasks. Within OSNeRF, semantic encoder is employed to extract core semantic features of OSOs from the collected scene images, semantic decoder is utilized to facilitate robust image recovery under HIWE conditions, lightweight renderer is employed for fast and efficient object reconstruction. Moreover, a semantic control unit (SCU) is introduced to guide above components, thereby enhancing the efficiency of reconstruction. Demonstrative experiments demonstrate that the proposed OSNeRF enables fast and robust object reconstruction in HIWE, surpassing the performance of state-of-the-art (SOTA) methods in terms of reconstruction quality.
Gaolei Li, Changze Li, Zhaohui Yang 0001, Yuchen Liu 0001, Mingzhe Chen
ACM Multimedia2
2024 ZeroTKS: Zero-trust Knowledge Synchronization via Federated Fine-tuning for Secure Semantic Communications
abstract
Semantic communication has experienced considerable growth and advancement due to its potential to support future intelligent applications (e.g., augmented reality). The realization of the above potential superiority depends on the construction and synchronization of semantic knowledge base among multiple ends. However, existing methods for constructing semantic knowledge base fail to adhere to the zero-trust architecture, where all communication ends can act as knowledge contributors without rigorous authentication. Motivated by this insight, we propose a novel zero-trust knowledge synchronization (ZeroTKS) scheme for secure semantic communication based on federated fine-tuning. In the proposed scheme, we firstly explore to introduce homomorphic encryption into federated fine-tuning of large models to securely synchronize the semantic knowledge base against privacy leakage risks. And also, to prevent from malicious model tampering attacks, a novel Age-of-Update-based access control mechanism is designated, in which only nodes with high AoU values can be authorized to participate in updating the semantic knowledge base by combining with hash-based message authentication code. Extensive experiments based on four different datasets in the GLUE benchmark show that our proposed scheme can securely synchronize distributed semantic knowledge base with maintaining acceptable performance.
Gaolei Li, Shuilin Li, Jianhua Li 0001
MobiHoc2
2024 ActiveDaemon: Unconscious DNN Dormancy and Waking Up via User-specific Invisible Token
Gaolei Li, Shenghong Li 0001, Kui Ren 0001
NDSS2
2024 HyperBC: Hypergraph-Based Approach for Behavior Cluster of Suspicious APT Attacks
Wenhui Du, Shuilin Li, Gaolei Li, Jianhua Li 0001
SecureComm (2)3
2024 Leveraging Neural Radiance Field and Semantic Communication for Robust 3D Reconstruction
abstract
By leveraging multi-view inputs to synthesize novel-view images, Neural Radiance Fields (NeRF) have emerged as a prominent technique in the realm of 3D object reconstruction. However, the input images of NeRF transmitted through high-interference wireless environment (HIWE) leads to negatively impact the accuracy of 3D reconstruction, thereby limiting its scalability. Fortunately, semantic communication has been proved a effective method to solve the above problem. In this paper, we propose a novel NeRF based 3D semantic communication (NeRF-3DSC) system, which offers robust 3D reconstruction in HIWE. Within NeRF-3DSC, semantic encoder and decoder are employed to extract and recover core semantic features of task-specific specific object (TSO) from the collected images, channel encoder and decoder ensure robust transmission of compressed semantic information in HIWE, lightweight renderer based on NeRF is employed for fast and efficient 3D reconstruction. Moreover, a semantic control unit (SCU) is introduced to guide above components, thereby enhancing the efficiency of reconstruction. Demonstrative experiments demonstrate that the proposed NeRF-3DSC enables robust object reconstruction in HIWE, surpassing the performance of state-of-the-art (SOTA) methods in terms of reconstruction quality.
Gaolei Li, Xi Lin 0003, Yuchen Liu 0001, Mingzhe Chen, Jianhua Li 0001
VTC Fall2
2024 Covert and Reliable Semantic Communication Against Cross-Layer Privacy Inference over Wireless Edge Networks
abstract
Semantic communication has emerged as a revolutionary paradigm within wireless edge networks, showcasing remarkable communication efficiency. In contrast to traditional bit-level communication systems, semantic communication systems exhibit superior effectiveness and precision, particularly in scenarios characterized by low signal-to-noise ratios (SNR). Nonetheless, the privacy of semantic communication poses a critical challenge that demands attention. Once the attacker intercepts the semantic information through continuous eaves-dropping, the private data would be leaked under adversarial environment. Moreover, in low SNR scenario, joint optimization of anti -eavesdropping and privacy reconstruction has not yet been studied, coupled with the intricate nature of designing a cross-layer semantic protection strategy. To address this concern, this paper presents a covert and reliable semantic communication (CRSC) framework via full-duplex receiver to counter continuous eavesdropper by concealing the entire transmission process. Furthermore, a newly-defined metric, namely covert semantic throughput (CST), is introduced to quantify the system's performance. Furthermore, we formulate the maximization of average CST during the semantic transmission period as a multi-constraint optimization problem. Subsequently, we propose a reinforcement learning (RL)-empowered adaptation algorithm to address the formulated problem. Through simulation results, the effectiveness and feasibility of proposed CRSC framework are demonstrated, with an observed maximum average CST improvement of up to 42% compared to conventional communication systems in the low SNR scenario.
Gaolei Li, Zhaohui Yang 0001, Mingzhe Chen, Yuchen Liu 0001, Jianhua Li 0001
WCNC2
2024 Local Differential Private Spatio- Temporal Dynamic Graph Learning for Wireless Social Networks
abstract
Differential-Private Graph Neural Networks (DP-GNNs) have generated remarkable research results, enabling them to effectively tackle the privacy leakage problem in graph learning. However, most DP-GNNs do not consider the temporal-dimensional scenarios. In tasks involving spatiotemporal graph training, such as wireless social networks analysis, the sensitive interactive information in each time graph should be protected. Therefore, we propose spatio-temporal dynamic graph learning with enhanced local differential privacy (LDP-STG). First, we design a weighted graph perturbation encoder based on the Bernoulli distribution and Laplace mechanism, which protects the structures of weighted graphs in time series under edge-level local differential privacy. Second, we employ an attention mechanism to learn dynamic graph node embedding from spatial and temporal dimensions. The theoretical analysis proves that our LDP-STG realizes differential privacy guarantees. We conduct experiments mainly on two communication datasets (i.e., Enron and UCI), which shows that our LDP-STG can achieve better privacy utility tradeoffs compared with traditional mechanisms in terms of snatiotemnoral dimensions.
Jiani Zhu, Xi Lin 0003, Yuxin Qi 0001, Gaolei Li, Jianhua Li 0001
WCNC4
2024 BenchMFC: A benchmark dataset for trustworthy malware family classification under concept drift
Yongkang Jiang, Gaolei Li, Shenghong Li 0001, Ying Guo 0004
Comput. Secur.2
2024 Joint Top-K Sparsification and Shuffle Model for Communication-Privacy-Accuracy Tradeoffs in Federated-Learning-Based IoV
abstract
The Internet of Vehicles (IoV) connects a massive amount of smart vehicles for inter/intra-vehicle information sharing. Data privacy issues, such as privacy leakage and privacy cost are the key challenges that hinder vehicle operators from sharing their data safely. Traditional privacy-preserving techniques, including Federated Learning (FL) and Differential Privacy (DP) techniques, can protect data privacy and security, but the high privacy cost severely limits learning performance. In addition, the IoV services place high demands on low communication latency, which can be obtained by reducing the communication bits, but it also limits the learning performance. Thus, how to solve the communication-privacy-accuracy tradeoffs to achieve low latency, high privacy preservation and model performance has been a complicated issue in IoV. In this paper, a privacy-enhancement differentially private federated learning framework (FedSDP) is proposed based on the shuffle model to ensure secure and efficient data sharing under the constraint of low latency in IoV. In our proposed framework, four privacy enhancement methods are proposed, including data subsampling, vehicle sampling, shuffle model and dummy points, to amplify the privacy and obtain higher learning performance. Then, a Top-K sparsification mechanism of the vehicle training process is proposed to reduce communication bits. Finally, the experimental results indicate that our approach can reduce the communication latency by 31.66%, enhance the privacy ϵc by 30.77% and improve the test accuracy by 48.56%, compared with the traditional SDP mechanism.
Hansong Xu, Kun Hua, Xi Lin 0003, Gaolei Li, Tigang Jiang, Jianhua Li 0001
IEEE Internet Things J.5
2024 HSESR: Hierarchical Software Execution State Representation for Ultralow-Latency Threat Alerting Over Internet of Things
abstract
To reduce attack risks in Internet of Things (IoT), many security vendors conduct software security analysis on IoT devices all the time. However, how to build an ultralow-latency threat alerting strategy using software vulnerability information still faces challenges. First, existing terminal threat detection methods for IoT systems relying on Indicators of Compromise (IoC) threat intelligence can only cover limited software vulnerabilities so the alert validity rate is still very low. Second, most users lack security knowledge and cannot proactively distinguish high-risk vulnerabilities, resulting in untimely reporting. In this article, a novel hierarchical software execution state representation (HSESR) scheme is proposed for ultralow latency threat alerting over IoT systems based on Beyond 5G. In HSESR, function call graphs are recorded and delivered to edge servers for swiftly identifying suspicious threat behaviors based on deep graph representation, while corresponding instruction sequences are delivered to the cloud data center for further matching the vulnerability information via recurrent semantic representation. To improve the effectiveness of HSESR, the graph representation is also actively encapsulated into the corresponding semantic representation, together acting as an implicit threat behavior signature, which is essential to associate with a security patch. Moreover, to accelerate the detection of suspicious behaviors, we also propose a deep reinforcement learning-based graph searching (DRL-GS) strategy to crop the huge function call graph of the entire software to timely report high-risk threat behaviors with minimized resource consumption. By instancing 1-day attacks on a simulated beyond 5G IoT system, the performance of HSESR is trustfully competitive against existing baselines, and the efficiency of threat detection was increased by 21.63%.
Xiaoyu Yi 0003, Gaolei Li, Bei Chen 0004, Xi Lin 0003, Yuchen Liu 0001, Jianhua Li 0001
IEEE Internet Things J.2
2024 Securing Distributed Network Digital Twin Systems Against Model Poisoning Attacks
abstract
In the era of 5G and beyond, the increasing complexity of wireless networks necessitates innovative frameworks for efficient management and deployment. Digital twins (DTs), embodying real-time monitoring, predictive configurations, and enhanced decision-making capabilities, stand out as a promising solution in this context. Within a time-series data-driven framework that effectively maps wireless networks into digital counterparts, encapsulated by integrated vertical and horizontal twinning phases, this study investigates the security challenges in distributed network DT (NDT) systems, which potentially undermine the reliability of subsequent network applications, such as wireless traffic forecasting. Specifically, we consider a minimal-knowledge scenario for all attackers, in that they do not have access to network data and other specialized knowledge, yet can interact with previous iterations of server-level models. In this context, we spotlight a novel fake traffic injection attack designed to compromise a distributed NDT system for wireless traffic prediction. In response, we then propose a defense mechanism, termed global-local inconsistency detection (GLID), to counteract various model poisoning threats. GLID strategically removes abnormal model parameters that deviate beyond a particular percentile range, thereby fortifying the security of network twinning process. Through extensive experiments on real-world wireless traffic data sets, our experimental evaluations show that both our attack and defense strategies significantly outperform existing baselines, highlighting the importance of security measures in the design and implementation of DTs for 5G and beyond network systems.
Minghong Fang, Mingzhe Chen, Gaolei Li, Xi Lin 0003, Yuchen Liu 0001
IEEE Internet Things J.4
2024 Protecting Intellectual Property With Reliable Availability of Learning Models in AI-Based Cybersecurity Services
abstract
Artificial intelligence (AI)-based cybersecurity services offer significant promise in many scenarios, including malware detection, content supervision, and so on. Meanwhile, many commercial and government applications have raised the need for intellectual property protection of using deep neural network (DNN). Existing studies (e.g., watermarking techniques) on intellectual property protection only aim at inserting secret information into DNNs, allowing producers to detect whether the given DNN infringes on their own copyrights. However, since the availability protection of learning models is rarely considered, the piracy model can still work with high accuracy. In this paper, a novel model locking (M-LOCK) scheme for the DNN is proposed to enhance its availability protection, where the DNN produces poor accuracy if a specific token is absent, while it maps only the tokenized inputs into correct predictions. The proposed scheme performs the verification process during the DNN inference operation, actively protecting models' intellectual property copyright at each query. Specifically, to train the token-sensitive decision-making boundaries of DNNs, a data poisoning-based model manipulation (DPMM) method is also proposed, which minimizes the correlation between the dummy outputs and correct predictions. Extensive experiments demonstrate the proposed scheme could achieve high reliability and effectiveness across various benchmark datasets as well as typical model protection methods.
Jun Wu 0001, Gaolei Li, Shenghong Li 0001, Mohsen Guizani
IEEE Trans. Dependable Secur. Comput.3
2024 FocusedCleaner: Sanitizing Poisoned Graphs for Robust GNN-Based Node Classification
abstract
Graph Neural Networks (GNNs) are vulnerable to data poisoning attacks, which will generate a poisoned graph as the input to the GNN models. We present FocusedCleaner as a poisoned graph sanitizer to effectively identify the poison injected by attackers. Specifically, FocusedCleaner provides a sanitation framework consisting of two modules: bi-level structural learning and victim node detection. In particular, the structural learning module will reverse the attack process to steadily sanitize the graph while the detection module provides the “focus” – a narrowed and more accurate search region – to structural learning. These two modules will operate in iterations and reinforce each other to sanitize a poisoned graph step by step. As an important application, we show that the adversarial robustness of GNNs trained over the sanitized graph for the node classification task is significantly improved. Extensive experiments demonstrate that FocusedCleaner outperforms the state-of-the-art baselines both on poisoned graph sanitation and improving robustness.
Yulin Zhu 0001, Liang Tong, Gaolei Li, Xiapu Luo, Kai Zhou 0001
IEEE Trans. Knowl. Data Eng.3
2024 Map-Driven mmWave Link Quality Prediction With Spatial-Temporal Mobility Awareness
abstract
The susceptibility of millimeter-wave (mmWave) links to blockages poses challenges for maintaining consistent high-rate performance. By predicting link quality in advance at specific locations or times of interest, proactive resource allocation techniques, such as link-quality-aware scheduling, can be employed to optimize the utilization of network resources. In this paper, we introduce a map-driven link quality prediction framework that divides the problem into long-term and short-term link quality predictions to cater to the needs of mobile computing. The first stage aims to predict a long-term radio map considering static network characteristics. We propose to separate LoS and NLoS scenarios, and build an analytical model and a regression-based approach to construct a complete link quality map in the spatial domain. Next, short-term link quality prediction is explored to anticipate future variations in link quality through a spatial-temporal attention-based prediction framework. The essence of this approach lies in capturing the spatial correlation and temporal dependency of mmWave wireless characteristics, followed by an attention mechanism to complement the dynamic link quality prediction task. On top of that, we also design a regional training mechanism with a weighted loss function to address the classical data imbalance problem of map-driven prediction. Extensive experimental and simulation results show that our integrated framework effectively captures comprehensive spatial-temporal knowledge and achieves significantly higher accuracy than other baseline prediction methods, making it a promising solution for a wide range proactive configuration tasks in mobile mmWave networks.
Zhizhen Li, Mingzhe Chen, Gaolei Li, Xi Lin 0003, Yuchen Liu 0001
IEEE Trans. Mob. Comput.3
2024 Crowdsourcing Malware Family Annotation: Joint Class-Determined Tag Extraction and Weakly-Tagged Sample Inference
abstract
Anti-malware engines report malware labels to detail malice, typically including tags of family, behavior, and platform classes. This capability has been heavily used by the security community to annotate malware families and build reference datasets, which is referred to as crowdsourcing malware family annotation. However, how to associate tags with their corresponding classes in chaotic malware labels (extract class-determined tags) and how to infer ground truth for weakly-tagged samples that hold controversial tags remain open problems. In this paper, we present a novel annotation pipeline to advance further, which includes an incremental parsing scheme and a maximum likelihood estimation scheme. The incremental parsing scheme treats behavior and platform tags as locators and achieves incremental parsing by introducing and iterating the following two algorithms: location first search, which hits family tags using locators, and co-occurrence first search, which finds new locators by family tags. The maximum likelihood estimating scheme models an engine’s ability to identify different families as a confusion matrix and introduces an expectation-maximization algorithm to estimate the matrix, as well as the unknown truth of samples. Experiments across four benchmark datasets indicate that our pipeline outperforms existing work, improving label-level parsing accuracy by an average of 29%, and improving inferring accuracy on weakly-tagged samples by an average of 9%. Our pipeline decouples parsing and inferring, which would pave the way for research on crowdsourcing malware family annotation.
Yongkang Jiang, Gaolei Li, Shenghong Li 0001, Ying Guo 0004, Kai Zhou 0001
IEEE Trans. Netw. Serv. Manag.2
2024 Exploring Semantic Redundancy using Backdoor Triggers: A Complementary Insight into the Challenges Facing DNN-based Software Vulnerability Detection
abstract
To detect software vulnerabilities with better performance, deep neural networks (DNNs) have received extensive attention recently. However, these vulnerability detection DNN models trained with code representations are vulnerable to specific perturbations on code representations. This motivates us to rethink the bane of software vulnerability detection and find function-agnostic features during code representation which we name as semantic redundant features. This paper first identifies a tight correlation between function-agnostic triggers and semantic redundant feature space (where the redundant features reside) in these DNN models. For correlation identification, we propose a novel Backdoor-based Semantic Redundancy Exploration (BSemRE) framework. In BSemRE, the sensitivity of the trained models to function-agnostic triggers is observed to verify the existence of semantic redundancy in various code representations. Specifically, acting as the typical manifestations of semantic redundancy, naming conventions, ternary operators and identically-true conditions are exploited to generate function-agnostic triggers. Extensive comparative experiments on 1,613,823 samples of eight representative vulnerability datasets and state-of-the-art code representation techniques and vulnerability detection models demonstrate that the existence of semantic redundancy determines the upper trustworthiness limit of DNN-based software vulnerability detection. To the best of our knowledge, this is the first work exploring the bane of software vulnerability detection using backdoor triggers.
Changjie Shao, Gaolei Li, Jun Wu 0001, James Xi Zheng
ACM Trans. Softw. Eng. Methodol.2
2023 TagClass: A Tool for Extracting Class-Determined Tags from Massive Malware Labels via Incremental Parsing
abstract
VirusTotal is widely used for malware annotation by providing malware labels from a large set of anti-malware engines. A long-standing challenge in using these inconsistent labels is extracting class-determined tags. In this paper, we present Tagclass,a tool based on incremental parsing to associate tags with their corresponding family, behavior, and platform classes. Tagclasstreats behavior and platform tags as locators and achieves incremental parsing by introducing and iterating the following two algorithms: 1) location first search, which hits family tags using locators, and 2) co-occurrence first search, which finds new locators by family tags. Experiments across two benchmark datasets indicate Tagclassoutperforms existing methods, improving the parsing accuracy by 21% and 28%, respectively. To the best of our knowledge, Tagclassis the first tag class-determined malware label parsing tool, which would pave the way for research on crowdsourcing malware annotation. Tagclasshas been released to the community11https://github.com/crowdma/tagclass.
Yongkang Jiang, Gaolei Li, Shenghong Li 0001
DSN2
2023 Spatial-Temporal Attention-Based mmWave Link Quality Prediction Under Dynamic Blockages
abstract
Millimeter-wave (mmWave) communication is a promising technology that has become a key component of next-generation wireless networks due to its large available band-width. However, the susceptibility of mmWave link to dynamic blockages makes it challenging to maintain consistently high rate performance. Hence, it is imperative to have the knowledge of link quality in advance at the location of interest to proactively optimize the use of network resources. In this work, we propose a Spatial-Temporal Attention-based Prediction (STAP) framework to predict the link quality at arbitrary locations in the presence of dynamic blockages. Specifically, our STAP model is built to capture the spatial correlation and temporal dependency of mmWave wireless characteristics in an integrated module, followed by an attention mechanism to complement the link quality prediction task. On top of that, we also design a regional training approach with a weighted loss function to address the data imbalance problem of map-based prediction. Extensive evaluation results show that our framework effectively captures comprehensive spatial-temporal knowledge and achieves significantly higher accuracy than other baseline prediction methods.
Zhizhen Li, Mingzhe Chen, Gaolei Li, Yuchen Liu 0001
GLOBECOM3
2023 Dynamic Path Planning Based on Traffic Flow Prediction and Traffic Light Status
Bingyi Liu, Weizhen Han, Gaolei Li
ICA3PP (1)4
2023 Black-Box Graph Backdoor Defense
Xiao Yang 0016, Gaolei Li, Xiaoyi Tao, Jianhua Li 0001
ICA3PP (5)2
2023 Persistent Clean-Label Backdoor on Graph-Based Semi-supervised Cybercrime Detection
Xiao Yang 0016, Gaolei Li
ICDF2C (1)2
2023 SemSBA: Semantic-perturbed Stealthy Backdoor Attack on Federated Semi-supervised Learning
abstract
Federated semi-supervised learning (FSSL) has been perceived as a promising approach that leverages semi-supervised learning and federated learning (FL) to provide powerful privacy preservation while reducing the burden on human supervision. However, due to the lack of strict participant identification and the significant proportion of unlabeled samples, FSSL is more susceptible to covert backdoor attacks than traditional machine learning. To validate this speculation, a novel semantic-perturbed stealthy backdoor attack (SemSBA) scheme is proposed for FSSL-based systems. In SemSBA, we select original natural semantic features in the unlabeled training samples as backdoor triggers and then generate poisoned samples by adding adversarial perturbations that move them across the model decision boundary. With SemSBA, the adversary can trigger the hidden backdoor in the victim model during the inference stage without any deliberate modifications on testing samples. To further improve the strength and robustness of the attack, a pseudo label steering enhancement strategy is also designed to perturb the weakly-augmented version of unlabeled samples to induce target pseudo label allocations. Additionally, to improve the attack success rate, we amplify the weight of the local backdoored model during FSSL’s model aggregation process to manipulate the game between benign clients and malicious clients. Extensive experiments based on two benchmark datasets demonstrate that the proposed SemSBA scheme can achieve comparable stealthiness against existing attacks.
Yingrui Tong, Jun Feng 0007, Gaolei Li, Xi Lin 0003, Chengcheng Zhao, Xiaoyu Yi 0003, Jianhua Li 0001
ICPADS3
2023 Privacy Inference-Empowered Stealthy Backdoor Attack on Federated Learning under Non-IID Scenarios
abstract
Federated learning (FL) naturally faces the problem of data heterogeneity in real-world scenarios, but this is often overlooked by studies on FL security and privacy. On the one hand, the effectiveness of backdoor attacks on FL may drop significantly under non-IID scenarios. On the other hand, malicious clients may steal private data through privacy inference attacks. Therefore, it is necessary to have a comprehensive perspective of data heterogeneity, backdoor, and privacy inference. In this paper, we propose a novel privacy inference-empowered stealthy backdoor attack (PI-SBA) scheme for FL under non-IID scenarios. Firstly, a diverse data reconstruction mechanism based on generative adversarial networks (GANs) is proposed to produce a supplementary dataset, which can improve the attacker's local data distribution and support more sophisticated strategies for backdoor attacks. Based on this, we design a source-specified backdoor learning (SSBL) strategy as a demonstration, allowing the adversary to arbitrarily specify which classes are susceptible to the backdoor trigger. Since the PI-SBA has an independent poisoned data synthesis process, it can be integrated into existing backdoor attacks to improve their effectiveness and stealthiness in non-IID scenarios. Extensive experiments based on MNIST, CIFAR10 and Youtube Aligned Face datasets demonstrate that the proposed PI-SBA scheme is effective in non-IID FL and stealthy against state-of-the-art defense methods.
Haochen Mei, Gaolei Li, Jun Wu 0001, Longfei Zheng
IJCNN2
2023 DPG-DT: Differentially Private Generative Digital Twin for Imbalanced Learning in Industrial IoT
abstract
The existing Artificial Intelligence (AI)-based industrial defect detection methods have received extensive attention in the industrial Internet of Things (IoT). However, due to the limited defect samples, it is difficult for discriminative models to achieve better performance in imbalanced learning. In addition, the privacy concerns surrounding sensitive information hinder the sharing of synthetic industrial data. In this paper, we propose a novel framework called the Differentially Private Generative AI-empowered Digital Twin (DPG-DT) framework, aiming to synthesize realistic samples while satisfying differential privacy and empowering the construction of digital space and its connection with physical space. Specifically, the core of the DPG-DT framework is the proposed Private Synthetic Industry Energy-guided model (PSIE), in which we privatize the energybased model-empowered Langevin Markov Chain Monte Carlo (MCMC) sampling method with Gaussian noise and random response. Our method could replace the conventional generator while guaranteeing privacy. Extensive experiments on real-world industrial datasets NEU-CLS and DeepPCB demonstrate that the proposed framework is capable of generating synthetic industrial images with both high fidelity and differential privacy. Moreover, the achieved downstream accuracy outperforms baselines by 23.9 % in industrial scenarios.
Siyuan Li 0005, Xi Lin 0003, Gaolei Li, Lixing Chen, Siyi Liao, Jianhua Li 0001
MSN3
2023 CATFL: Certificateless Authentication-based Trustworthy Federated Learning for 6G Semantic Communications
abstract
Federated learning (FL) provides an emerging approach for collaboratively training semantic encoder/decoder models of semantic communication systems, without private user data leaving the devices. Most existing studies on trustworthy FL aim to eliminate data poisoning threats that are produced by malicious clients, but in many cases, eliminating model poisoning attacks brought by fake servers is also an important objective. In this paper, a certificateless authentication-based trustworthy federated learning (CATFL) framework is proposed, which mutually authenticates the identity of clients and server. In CATFL, each client verifies the server’s signature information before accepting the delivered global model to ensure that the global model is not delivered by false servers. On the contrary, the server also verifies the server’s signature information before accepting the delivered model updates to ensure that they are submitted by authorized clients. Compared to PKI-based methods, the CATFL can avoid too high certificate management overheads. Meanwhile, the anonymity of clients shields data poisoning attacks, while real-name registration may suffer from user-specific privacy leakage risks. Therefore, a pseudonym generation strategy is also presented in CATFL to achieve a trade-off between identity traceability and user anonymity, which is essential to conditionally prevent from user-specific privacy leakage. Theoretical security analysis and evaluation results validate the superiority of CATFL.
Gaolei Li
WCNC1
2023 Multitentacle Federated Learning Over Software-Defined Industrial Internet of Things Against Adaptive Poisoning Attacks
abstract
Software-defined industrial Internet of things (SD-IIoT) exploits federated learning to process the sensitive data at edges, while adaptive poisoning attacks threat the security of SD-IIoT. To address this problem, this article proposes a multi-tentacle federated learning (MTFL) framework, which is essential to guarantee the trustness of training data in SD-IIoT. In MTFL, participants with similar learning tasks are assigned to the same tentacle group. To identify adaptive poisoning attacks, a tentacle distribution-based efficient poisoning attack detection (TD-EPAD) algorithm is presented. And also, to minimize the impact of adaptive poisoning data, a stochastic tentacle data exchanging (STDE) protocol is also proposed. Simultaneously, to protect the tentacle’s privacy in STDE, all exchanged data will be processed by differential privacy technology. A MTFL prototype system is implemented, which provides extensive ablation experiments and comparison experiments, demonstrating that the accuracy of the global model under attack scenario can be improved with 40%.
Gaolei Li, Jun Wu 0001, Shenghong Li 0001, Wu Yang 0001, Changlian Li
IEEE Trans. Ind. Informatics1
2023 Friend-as-Learner: Socially-Driven Trustworthy and Efficient Wireless Federated Edge Learning
abstract
Recently, wireless edge networks have realized intelligent operation and management with edge artificial intelligence (AI) techniques (i.e., federated edge learning). However, the trustworthiness and effective incentive mechanisms of federated edge learning (FEL) have not been fully studied. Thus, the current FEL framework will still suffer untrustworthy or low-quality learning parameters from malicious or inactive learners, which undermines the viability and stability of FEL. To address these challenges, the potential social attributes among edge devices and their users can be exploited, while not included in previous works. In this paper, we propose a novelSocialFederatedEdgeLearning framework (SFEL) over wireless networks, which recruits trustworthy social friends as learning partners. First, we build a social graph model to find like-minded friends, comprehensively considering the mutual trust and learning task similarity. Besides, we propose a social effect based incentive mechanism for better personal federated learning behaviors with both complete and incomplete information. Finally, we conduct extensive simulations with the Erdos-Renyi random network, the Facebook network, and the classic MNIST/CIFAR-10 datasets. Simulation results demonstrate our framework could realize trustworthy and efficient federated learning over wireless edge networks, and it is superior to the existing FEL incentive mechanisms that ignore social effects.
Xi Lin 0003, Jun Wu 0001, Jianhua Li 0001, James Xi Zheng, Gaolei Li
IEEE Trans. Mob. Comput.5
2022 Digital Transformation (DX) for Skill Learners: The Design Methodology and Implementation of Educational Chatbot using Knowledge Connection and Emotional Expression
abstract
With the continuous evolution of online education, online educational chatbots have become an indispensable help between teachers and learners. Due to the unified answers, it is hard for the educational chatbots to replace part of the teacher’s work, which should be more efficient than textbooks when retrieving knowledge and establishing emotional connections. In this paper, we propose an AI-based educational chatbot paradigm that introduces context recognition, learning experience, and emotion management to improve learners’ emotional confidence through the dialogue template while strengthening the chatbots’ ability to guide materials to inspire learners. We utilize long-term attention mechanisms to establish emotional connections with learners and strengthen skill learners’ self-efficacy. We investigate two types of up-to-date AI chatbots and conduct a mathematical analysis of the learners’ dialogue trust and self-efficacy by DNN (deep learning neural network). Referring to the analysis results, the introduction of mechanism rationale knowledge sharing and emotional connection has a significant effect on dialogue trust when chatting, and the learners’ self-efficacy is satisfied for future skill learning.
Gaolei Li, Hiroshi Hashimoto, Zejun Zhang 0009
EDUCON2
2022 Propagable Backdoors over Blockchain-based Federated Learning via Sample-Specific Eclipse
abstract
Blockchain-based federated learning, also being named as swarm learning, is perceived to have great potential to support decentralized and privacy-enhancing big data processing. However, numerous serious vulnerabilities found on blockchain and federated learning enforce us to concern about the security of swarm learning. Some seemingly-unrelated combinations of known vulnerabilities may derive highly-converted and unknown threats to swarm learning. In this paper, we first investigate the security threats of the swarm learning framework. And then, leveraging backdoor attacks and eclipse attacks, a novel hybrid vulnerability that can furtively propagate backdoors among swarm learning nodes is identified. To speed up the backdoor propagation and reduce attack costs, a sample-specific eclipse (SSE) strategy that can select the swarm network node with a high data contribution rate as the attack object is also proposed. Finally, by adjusting the trigger size, the data distribution rate, and the poisoning ratio, we conduct various comparison experiments to validate the feasibility of the proposed methods. To the best of our knowledge, this is the first article to study the epidemicity of backdoors in swarm learning.
Zheng Yang 0002, Gaolei Li, Jun Wu 0001, Wu Yang 0001
GLOBECOM2
2022 Explainable Intelligence-Driven Defense Mechanism Against Advanced Persistent Threats: A Joint Edge Game and AI Approach
abstract
Advanced persistent threats (APT) have novel features such as long-term latency, precision strikes and uncertain strategies. APT poses severe threats to the resource-limited edge devices in advanced networks. Cyber threat intelligence (CTI) conducts data analysis on attack strategies by artificial intelligence (AI) and generates threat intelligence to optimize the detection model and guide defense strategies. However, AI lacks explanations for the decisions and thus reduces the transparency and performance of the detection model. Besides, the tradeoff between the detection accuracy and the computational resource limitation of edge devices needs an optimal and rapid dynamic resource allocation method, which edge game and AI can help. In this paper, we propose an explainable intelligence-driven APT edge defense mechanism. The proposed mechanism provides guidelines and explanations for designing the defense strategy and resource allocation scheme of the edge defender to detect APT. The edge defense strategy model is based on edge Bayesian Stackelberg game and CTI. Meanwhile, we implement a DRL-based resource allocation scheme to meet rapid response requirements at the edges. We demonstrate that the proposed mechanism can improve the protection level of edges and defense capability against APT through extensive experiments.
Jun Wu 0001, Hansong Xu, Gaolei Li, Mohsen Guizani
IEEE Trans. Dependable Secur. Comput.4
2022 FLAG: Few-Shot Latent Dirichlet Generative Learning for Semantic-Aware Traffic Detection
abstract
The number of malware attempts that try to bypass the existing Network Intrusion Detection System (NIDS) is increasing. To detect illegal access to servers, deep analysis of the server-side network traffic has become increasingly important. However, the existing approaches have serious performance limitations in terms of real-time and accurate traffic detection. These limitations are mainly because of i) the rigid feature extraction and rule matching techniques of NIDS, which are insensitive to incremental network traffic, and ii) the strong correlation and coupling of malicious traffic to large normal traffic. To address these limitations, we propose a Few-shot Latent Dirichlet Generative Learning (FLAG) scheme for semantic-aware traffic detection in this paper. In FLAG, a Latent Dirichlet Allocation (LDA)-based pseudo samples generation algorithm is designated to augment the few-shot training data, which is essential to improve traffic classification accuracy. Furthermore, we propose a Fuzziness Recycle Method (FRM) to further improve the long short-term memory (LSTM)-based classifier’s robustness. Experimental results in real scenarios demonstrate that malicious traffic can be efficiently detected when only few-shot samples are learned. The results also reveal that the proposed scheme outperforms the state-of-the-art methods in detection accuracy.
Tianpeng Ye, Gaolei Li, Ijaz Ahmad 0001, Jianhua Li 0001
IEEE Trans. Netw. Serv. Manag.2
2022 Progressive Slice Recovery With Guaranteed Slice Connectivity After Massive Failures
abstract
In presence of multiple failures affecting their network infrastructure, operators are faced with the Progressive Network Recovery (PNR) problem, i.e., deciding the best sequence of repairs during recovery. With incoming deployments of 5G networks, PNR must evolve to incorporate new recovery opportunities offered by network slicing. In this study, we introduce the new problem of Progressive Slice Recovery (PSR), which is addressed with eight different strategies, i.e., allowing or not to change slice embedding during the recovery, and/or by enforcing different versions of slice connectivity (i.e., network vs. content connectivity). We propose a comprehensive PSR scheme, which can be applied to all recovery strategies and achieves fast recovery of slices. We first prove the PSR’s NP-hardness and design an integer linear programming (ILP) model, which can obtain the best recovery sequence and is extensible for all the recovery strategies. Then, to address scalability issues of the ILP model, we devise an efficient two-phases progressive slice recovery (2-phase PSR) meta-heuristic algorithm, small optimality gap, consisting of two main steps: i) determination of recovery sequence, achieved through a linear-programming relaxation that works in polynomial time; and ii) slice-embedding recovery, for which we design an auxiliary-graph-based column generation to re-embed failed slice nodes/links to working substrate elements within a given number of actions. Numerical results compare the different strategies and validate that amount of recovered slices can be improved up to 50% if operators decide to reconfigure only few slice nodes and guarantee content connectivity.
Qiaolun Zhang, Omran Ayoub, Jun Wu 0001, Francesco Musumeci 0001, Gaolei Li, Massimo Tornatore
IEEE/ACM Trans. Netw.5
2021 DeHiB: Deep Hidden Backdoor Attack on Semi-supervised Learning via Adversarial Perturbation
abstract
The threat of data-poisoning backdoor attacks on learning algorithms typically comes from the labeled data. However, in deep semi-supervised learning (SSL), unknown threats mainly stem from the unlabeled data. In this paper, we propose a novel deep hidden backdoor (DeHiB) attack scheme for SSL-based systems. In contrast to the conventional attacking methods, the DeHiB can inject malicious unlabeled training data to the semi-supervised learner so as to enable the SSL model to output premeditated results. In particular, a robust adversarial perturbation generator regularized by a unified objective function is proposed to generate poisoned data. To alleviate the negative impact of the trigger patterns on model accuracy and improve the attack success rate, a novel contrastive data poisoning strategy is designed. Using the proposed data poisoning scheme, one can implant the backdoor into the SSL model using the raw data without hand-crafted labels. Extensive experiments based on CIFAR10 and CIFAR100 datasets demonstrated the effectiveness and crypticity of the proposed scheme.
Zhicong Yan, Gaolei Li, Yuan Tian 0017, Jun Wu 0001, Shenghong Li 0001, Mingzhe Chen, H. Vincent Poor
AAAI2
2021 PFCC: Predictive Fast Consensus Convergence for Mobile Blockchain over 5G Slicing-enabled IoT
abstract
As the security requirements increases, 5G slicing-enabled Internet of things needs to adopt end-to-end standalone networking, which limits the consensus convergence of mobile blockchain. Although the rise of FIBRE (Fast Internet Bitcoin Relay Engine) gives a huge promotion to block propagation, existing approaches cannot consider the influence of link outages among 5G slices on mobile blockchain. In this paper, we focus on decreasing the block propagation time among blockchain peers under the scenario where link outages among 5G slices exist. A predictive fast consensus convergence (PFCC) scheme is proposed for mobile blockchain over 5G slicing-enabled internet of things. In PFCC, federated semi-supervised learning is used to learn the features of withdraw packets, reroutes the packets of blockchain peers, and ultimately reduces the scale of link outages quickly. With PFCC, different blockchain peers located in standalone 5G slices of IoT can transact local sensing data more efficiently. To the best of our knowledge, this is the first work to improve the consensus convergence speed of mobile blockchain by optimizing communications between 5G slices. Experiments shows the feasibility of proposed scheme.
Gaolei Li, Jun Wu 0001, Jianhua Li 0001
GLOBECOM2
2021 GradMFL: Gradient Memory-Based Federated Learning for Hierarchical Knowledge Transferring Over Non-IID Data
Guanghui Tong, Gaolei Li, Jun Wu 0001, Jianhua Li 0001
ICA3PP (1)2
2021 Deep Neural Backdoor in Semi-Supervised Learning: Threats and Countermeasures
abstract
Semi-Supervised Learning (SSL) is a powerful derivative for humans to discover the hidden knowledge, and will be a great substitute for data taggers. Although the availability of unlabeled data rises up a huge passion to SSL, the untrustness of unlabeled data leads to many unknown security risks. In this paper, we first identify an insidious backdoor threat of SSL where unlabeled training data are poisoned by backdoor methods migrated from supervised settings. Then, to further exploit this threat, a Deep Neural Backdoor (DeNeB) scheme is proposed, which requires less data poisoning budgets and produces stronger backdoor effectiveness. By poisoning a fraction of unlabeled training data, the DeNeB achieves the illegal manipulation on the trained model without modifying the training process. Finally, an efficient detection-and-purification defense (DePuD) framework is proposed to thwart the proposed scheme. In DePuD, we construct a deep detector to locate trigger patterns in the unlabeled training data, and perform secured SSL training with purified unlabeled data where the detected trigger patterns are obfuscated. Extensive experiments based on benchmark datasets are performed to demonstrate the huge threatening of DeNeB and the effectiveness of DePuD. To our best knowledge, this is the first work to achieve the backdoor and its defense in semi-supervised learning.
Zhicong Yan, Jun Wu 0001, Gaolei Li, Shenghong Li 0001, Mohsen Guizani
IEEE Trans. Inf. Forensics Secur.3
2020 DeSVig: Decentralized Swift Vigilance Against Adversarial Attacks in Industrial Artificial Intelligence Systems
abstract
Individually reinforcing the robustness of a single deep learning model only gives limited security guarantees especially when facing adversarial examples. In this article, we propose DeSVig, a decentralized swift vigilance framework to identify adversarial attacks in an industrial artificial intelligence systems (IAISs), which enables IAISs to correct the mistake in a few seconds. The DeSVig is highly decentralized, which improves the effectiveness of recognizing abnormal inputs. We try to overcome the challenges on ultralow latency caused by dynamics in industries using peculiarly designated mobile edge computing and generative adversarial networks. The most important advantage of our work is that it can significantly reduce the failure risks of being deceived by adversarial examples, which is critical for safety-prioritized and delay-sensitive environments. In our experiments, adversarial examples of industrial electronic components are generated by several classical attacking models. Experimental results demonstrate that the DeSVig is more robust, efficient, and scalable than some state-of-art defenses.
Gaolei Li, Kaoru Ota, Mianxiong Dong, Jun Wu 0001, Jianhua Li 0001
IEEE Trans. Ind. Informatics1
2020 Processing capability and QoE driven optimized computation offloading scheme in vehicular fog based F-RAN
Tianpeng Ye, Jun Wu 0001, Gaolei Li, Jianhua Li 0001
World Wide Web4
2018 Sema-ICN: Toward Semantic Information-Centric Networking Supporting Smart Anomalous Access Detection
abstract
As a next-generation networking architecture, information-centric networks (ICN) has strengthened the focus on physical location-independent content sharing, which introduces abundant semantic features and novel access approach. However, the semantic modeling of ICN is an unresolved problem, thus current ICN lacks the capabilities of smart content analysis and understanding to support the knowledge decision for optimized user experience. To address this issue, we propose a semantic ICN model, Sema-ICN, that can provide logically related information depending on content-relevance-based relationships extraction and name-based weight setting. Moreover, besides the great benefits brought into ICN by semantic features, Sema- ICN will also contribute to security protection against anomalous access, which usually the basis of further threats. In this paper, we additionally design a smart anomalous access detection scheme supported by Sema- ICN, in which semantic communities are partitioned utilizing spectral clustering according to content name with semantic attributes. And a forecast model is introduced to predict access situation based on triple exponential smoothing algorithm using historical request data, access traffic that is beyond the forecast results will be considered as anomalous. The simulation results demonstrate the efficiency of the proposed scheme. To the best of our knowledge, this work is the first to propose a novel semantic model for ICN.
Jun Wu 0001, Jianhua Li 0001, Gaolei Li
GLOBECOM4
2018 Resource-Efficient Secure Data Sharing for Information Centric E-Health System Using Fog Computing
abstract
Recently, an accelerating number of studies are dedicated to deploying various IoT applications in the information centric network paradigm which has lower system complexity than traditional network architectures. However, such a paradigm poses a number of security challenges especially when it is applied in real-time e-health applications. Firstly, it is difficult to ensure security of sensitive data in such a distributed data caching environment because after the data is published in the form of a packet to the information centric network (ICN), it is no longer controlled by the data publisher. Secondly, in some real-time e-health applications, terminal medical sensors are usually resource-constrained, limiting the direct adoption of expensive cryptographic primitives. In order to address these challenges, a resource-efficient secure data sharing scheme in information centric e-health system is proposed, one that utilizes ciphertext-policy attribute based encryption (CP-ABE) and adapts it to the above-mentioned system with respect to necessary security requirements. It also exploits computation resources of fog nodes and employs outsourcing cryptography to improve system efficiency.The evaluation demonstrates that the scheme can significantly reduce the computation overheads of the resource-constrained terminal medical devices, and can better support the real-time e-health applications.
Lintao Dang, Mianxiong Dong, Kaoru Ota, Jun Wu 0001, Jianhua Li 0001, Gaolei Li
ICC6
2018 MapReduce Enabling Content Analysis Architecture for Information-Centric Networks Using CNN
abstract
Information Centric Network (ICN) is one of the promising architectures in the next generation networks. The content-based routing in ICN can satisfy the content distribution of large-scale data. For prompt content obtainment, it is important to realize the content analysis before the content reaches application layer. The novel characteristics of data naming in ICN make it possible to search and analyse content during the transmission of content, which can directly get the critical content without the process of the application layer. In this paper, we propose a MapReduce enabling content analysis architecture for ICN. MapReduce framework can realize the parallelization of content collection and analysis during the routing process. For more efficient content collection, we put forward an optimal selection for mapper nodes. Moreover, Convolutional Neural Network (CNN) is deployed in the MapReduce architecture providing further analysis for ICN content. The simulation result shows the advantages of the proposed architecture.
Chengcheng Zhao, Mianxiong Dong, Kaoru Ota, Jun Wu 0001, Jianhua Li 0001, Gaolei Li
ICC6
2018 Service Popularity-Based Smart Resources Partitioning for Fog Computing-Enabled Industrial Internet of Things
abstract
Recently, fog computing has gained increasing attention in processing the computing tasks of the industrial Internet of things (IIoT) with different service popularity. In task-diversified fog computing-enabled IIoT (F-IIoT), the mismatch between expected computing efficiency and partitioned resources on fog nodes (FNs) may pose serious traffic congestion even large-scale industrial service interruptions. The existing works mainly studied offloading which type of computing tasks into FNs, but few studies enabled smart resource partitioning of FNs. In this paper, a service popularity-based smart resources partitioning (SPSRP) scheme is proposed for fog computing-enabled IIoT. We first exploit Zipf's law to model the relationship between popularity ranks and computing costs of IIoT services. Moreover, we propose an implementation architecture of the SPSRP scheme for F-IIoT, which decouples the computing control layer from data processing layer of IIoT through a specified SPSRP controller. Besides, a mobility and heterogeneity-aware partitioning algorithm is presented for extending SPSRP scheme to seamlessly support cross-domain resources partitioning. The simulations demonstrate that the SPSRP scheme can bring notable performance improvements on delay time, successful response rate and fault tolerance for fog computing to deal with the large-scale IIoT services.
Gaolei Li, Jun Wu 0001, Jianhua Li 0001, Kuan Wang 0001, Tianpeng Ye
IEEE Trans. Ind. Informatics1
2017 SD-OPTS: Software-Defined On-Path Time Synchronization for Information-Centric Smart Grid
abstract
Information-centric networking (ICN) and software defined networking (SDN) has been perceived as a promising paradigm for integrating distributed generation (DG) into smart grid networks for flexibility and dynamic features. However, flexible, controllable and reliable time synchronization still remains an open issue for supervisory control and data sensing in smart grid. Firstly, for information-centric smart grid, smart grid entities may obtain available data from caching routers, and data delivery between caching routers with edge devices results in time synchronization requirements of on-path caching routers. Without on-path time synchronization of caching routers, it may lead to maliciously fluctuations if energy data for monitoring the energy supplying and consumption are collected and traversed at an inappropriate time. Secondly, as scale expanding and network environment of smart grid becomes complex and changeable, on-path time synchronization needs a unified and dynamic management and control. To address these issues, in this paper we propose a novel scheme of software- defined on-path time synchronization (SD-OPTS) scheme for information-centric smart grid. In proposed scheme, all on-path caching routers share the time stamps from master clock and synchronize the time of local clock during one-time synchronization process. Besides, SDN controller estimates on-path caching routers' sync error, before choosing the nearest nodes to implement accurate time synchronization. The simulation results demonstrate the efficiency of software-defined on-path time synchronization scheme. The SD-OPTS scheme supports the flexibility, controllability and reliability of time synchronization.
Weiyi Han, Mianxiong Dong, Kaoru Ota, Jun Wu 0001, Jianhua Li 0001, Gaolei Li
GLOBECOM6
2017 Software-Defined Efficient Service Reconstruction in Fog Using Content Awareness and Weighted Graph
abstract
Fog computing, shifting intelligence and resources from remote cloud to edge networks, has the potential of providing low-latency for the end-to-end communication from data sources to users. However, it's hard to enhance resource-efficiency in the existing relatively static and proprietary framework of fog nodes due to the diversity of service requirements. With the growing deployment of fog computing, the overall resource consumption in fog will be huge without considering the efficient service provisioning in each fog node. On one hand, different fog users require diverse local services policies which are carried out within the fog nodes. Moreover, for one user, the requirements on services are time-varying. On the other hand, the processing strategies on different types of content (e.g. video, audio, etc.) are also distinct. These dynamic features impose the need for user-driven and content-based service reconstruction in order to achieve the high recycling utilization of resources of fog system. To this end, we propose a software-defined efficient service reconstruction (SDSR) scheme in fog using content awareness and weighted graph. Service reconstruction mechanism is devised to dynamically recycle modularized resources after mapping different contents to relevant operations. Weighted graph is introduced to schedule and optimize the services reconstruction in terms of resource saving during content-driven controlling. User-defined interfaces are designed to enable fog users to reconfigure the recyclable resource modules. Simulation results demonstrate that the service cost of each fog nodes is reduced significantly, thus promote efficient service provisioning for the whole fog system.
Mianxiong Dong, Kaoru Ota, Jun Wu 0001, Jianhua Li 0001, Gaolei Li
GLOBECOM6
2017 Towards QoE named content-centric wireless multimedia sensor networks with mobile sinks
abstract
To enforce surrounding surveillance efficiently and reduce the heavy cost to deploy various infrastructures, mobile sinks are perceived to have potentials for utilization by wireless multimedia sensor networks (WMSNs). However, since high-mobility usually causes communication disconnections and the high re-transmission rate will consume more network resources, quality of experience (QoE) monitoring and control is a must that WMSNs with mobile sinks (MS-WMSNs) should provide satisfactory services with constrained resources. In this paper, we propose a novel QoE-named content-centric network paradigm for MS-WMSNs, which supports location independent networking and low redundancy data aggregation. Each network node constructs a hierarchical content naming tree (HCNT) negotiated by QoE parameters. The MS prioritizes the sensing data and caches them differentially by identifying these QoE parameters based content names. Simultaneously, to verify the feasibility, we design a stochastic network calculus model to analyse the performances of our proposed network paradigm at worst-case situation. Simulation results show that the proposed paradigm reduces end-to-end communication delay.
Gaolei Li, Mianxiong Dong, Kaoru Ota, Jun Wu 0001, Jianhua Li 0001, Tianpeng Ye
ICC1
2016 Deep Packet Inspection Based Application-Aware Traffic Control for Software Defined Networks
abstract
Software defined networks (SDN) is perceived to have specific capabilities for utilization by network infrastructures automatically. The success of OpenFlow protocol is to decouple control plane from data plane completely. However, current SDN still regards the network as a group of devices rather than a holistic resource, and traffic monitoring and control only relies on network states but not including traffic behaviours. Although speed of packet forwarding is improved significantly, QoS demands can not be satisfied when network congested, unavailability of SDN in some resource constrained scenes does not present well. To address this, we propose an application-aware traffic control scheme, in which both network states and traffic behaviours are exploited cooperatively. Deep Packet Inspection (DPI) is introduced into SDN controller. Meanwhile, a mechanism for packet classification and behaviour matching is designed. To perform information exchange between components, a publish/subscribe based middle ware is designed. Besides, mathematical models for analysing network throughput and latency are established. Simulation results show that proposed scheme can facilitate the improvement of throughput and reduce latency time of end-to-end communication.
Gaolei Li, Mianxiong Dong, Kaoru Ota, Jun Wu 0001, Jianhua Li 0001, Tianpeng Ye
GLOBECOM1
2015 Chance Discovery Based Security Service Selection for Social P2P Based Sensor Networks
abstract
Social Peer-to-Peer (P2P) is a novel model to organize sensor networks, which can establish social relationships in an autonomous way with the benefits of extending the network boundaries and enhancing the network scalability. However, the complexity and time dependence characteristics introduced by social P2P model raise difficulties for assessing and selecting security services accurately and effectively in sensor networks. To address this, we propose a chance discovery based security service selection scheme for social P2P based sensor networks. We firstly establish the security assessment model for services in social P2P based sensor networks, regarding the security factors of exploitability, credibility, severity, confidentiality, integrality, availability and importance weight. More importantly, time dependence characteristics introduced by social P2P are considered during security assessment. Next, a security service selection scheme is proposed based on KeyGraph construction as well as the computation of its connection and tightness values. Finally, the service request forwarding model is established. The simulation results show the effectiveness and accuracy of the proposed security service selection scheme, which improves the feasibility and security of integrated sensor network and social networks.
Jun Wu 0001, Mianxiong Dong, Kaoru Ota, Jianhua Li 0001, Longhua Guo, Gaolei Li
GLOBECOM6