Longxiang Gao

dblp:44/7500 · DBLP profile ↗
← Back
150ranked-venue papers
4as first author
115since 2021 · last 2026
0000-0002-3026-7537ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 47 · 3 first-author · 29 since 2021Artificial intelligence and machine learning · 32 · 32 since 2021Databases, data management, data science and information retrieval · 15 · 11 since 2021Systems, architecture and hardware · 14 · 10 since 2021Security and privacy · 14 · 1 first-author · 9 since 2021Software engineering, systems software and programming languages · 13 · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 13 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 since 2021Human-computer interaction and ubiquitous computing · 1Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TransHER2: Prediction of HER2 Expression Status in Breast Ultrasound Videos Based on Transformer Spatiotemporal Interactive Feature Fusion
Xuejing Li, Longxiang Gao, Lei Cui 0006, Kexue Fu 0001
ICIC (3)3
2026 LegiCode: A blockchain-legal LLM framework for real-time compliance in smart contract generation
Linkai Zhu, Di Wu 0050, Longxiang Gao
Empir. Softw. Eng.5
2026 APSM: Adaptive privacy budget control in differentially private matching in electric vehicles
abstract
The rapid growth of Electric Vehicles (EVs) has brought significant challenges in ensuring the privacy of sensitive data generated, particularly in Vehicle-to-Vehicle (V2V) energy trading systems. This study examines methods to balance data privacy preservation with the utility required for EV-related services. Existing privacy-preserving techniques often struggle to strike a balance between privacy and utility, particularly in dynamic environments where data sensitivity and usage patterns are constantly changing. In this paper, we propose an Adaptive Private Stable Matching (APSM) algorithm that incorporates a dynamic privacy budget algorithm for Differential Privacy (DP). APSM provides stable, privacy-preserving matches for EVs participating in V2V energy trading. The dynamic privacy budget mechanism adjusts allocation according to the number of EVs, offering enhanced privacy protection when necessary and increased utility when feasible. The proposed approach optimizes the utilization of the privacy budget, meeting both strict privacy requirements and ensuring efficient service delivery. Experimental results show that the technique outperforms static approaches in terms of privacy budget management, thereby enhancing privacy protection while maintaining high data utility. This combination renders APSM highly suitable for practical V2V energy trading scenarios, delivering robust privacy safeguards without compromising system performance.
Saad Masood, Muneeb Ul Hassan 0001, Pei-Wei Tsai, Kai Zhang 0074, Longxiang Gao, Mianxiong Dong, Jinjun Chen
Expert Syst. Appl.5
2026 Fragment-energy audio watermarking resilient to de-synchronization attacks
Juan Zhao 0007, Tianrui Zong, Iynkaran Natgunanathan, Yong Xiang 0001, Guang Hua 0001, Longxiang Gao, Wanlei Zhou 0001
Expert Syst. Appl.7
2026 Dual-channel time-aware graph attention network for session-based recommendation
abstract
Session-based recommender systems face significant challenges in accurately predicting user preferences due to the limited availability of long-term historical interactions. While recent advances in deep learning and graph-based approaches have improved recommendation performance, the temporal aspects of user interactions remain underutilized. This paper identifies three critical temporal challenges in session-based recommendations: interest shifts indicated by long intervals between interactions, interaction noise from brief engagements, and system popularity effects during high-traffic periods. To address these challenges, we propose a novel Dual-channel Time-aware Graph Attention Network (DT-GAT) to incorporate temporal signal, i.e., time intervals between interactions and time differences between sessions, into session representations from both item and session perspectives. The item-wise learning channel employs a temporal graph attention network to capture interest shifts and filter interaction noise, while the session-wise learning channel utilizes a temporal graph attention network to handle inconsistent popularity trends. Additionally, we introduce a multi-temporal window processing mechanism to construct robust session representations that effectively capture short-term interests while filtering noise. Extensive experiments conducted on three real-world datasets demonstrate that DT-GAT consistently outperforms state-of-the-art baseline models. Our code is available at: https://github.com/downw/DT-GAT • We propose DT-GAT to integrate item- and session-level temporal signals. • Dual temporal GATs capture dependencies via temporal intra- and inter-session graphs. • Contrastive learning aligns dual channels to enhance session representations. • Experiments on three datasets validate the effectiveness of DT-GAT.
Linjiang Guo, Shiqing Wu 0001, Dan Lu 0004, Longxiang Gao, Guandong Xu
Inf. Sci.4
2026 Invisible watermarking framework for unlearned diffusion model in online service
Tianqing Zhu, Longxiang Gao, Wanlei Zhou 0001
Neural Networks3
2026 Structure prediction and opportunity-cost scheduler for LLM inference
Weiyan Huang, Guomao Xin, Bruce Gu, Youyang Qu, Longxiang Gao
Pattern Recognit.5
2026 Fast Convergent Federated Learning via Decaying SGD Updates
Md Palash Uddin, Yong Xiang 0001, Mahmudul Hasan 0018, Yao Zhao 0006, Youyang Qu, Longxiang Gao
IEEE Trans. Big Data6
2026 Safe and Reliable Diffusion Models via Subspace Projection
abstract
Large-scale text-to-image (T2I) diffusion models have revolutionized image generation, enabling the synthesis of highly detailed visuals from textual descriptions. However, these models may inadvertently generate inappropriate content, such as copyrighted works or offensive images. While existing methods attempt to eliminate specific unwanted concepts, they often fail to ensure robust removal-allowing the concept to reappear in subtle forms. For instance, a model may successfully avoid generating images in Van Gogh's style when explicitly prompted with “Van Gogh”, yet still reproduce his signature artwork when given the prompt “Starry Night”. In this paper, we propose SAFER, a novel and efficient approach for thoroughly removing target concepts from diffusion models. At a high level, SAFER is inspired by the observed low-dimensional structure of the text embedding space. The method first identifies a concept-specific subspace$\mathcal {S}_{c}$associated with the target concept$c$. It then projects the prompt embeddings onto the complementary subspace of$\mathcal {S}_{c}$, effectively erasing the concept from the generated images. Since concepts can be abstract and difficult to fully capture using natural language alone, we employ textual inversion to learn an optimized embedding of the target concept from a reference image. This enables more precise subspace estimation and enhances removal performance. Furthermore, we introduce a subspace expansion strategy to ensure comprehensive and robust concept erasure. Extensive experiments demonstrate that SAFER consistently and effectively erases unwanted concepts from diffusion models while preserving generation quality.
Huiqiang Chen, Tianqing Zhu, Xin Yu 0002, Longxiang Gao, Wanlei Zhou 0001
IEEE Trans. Dependable Secur. Comput.5
2026 Collusion-Resistant and Time-Aware Co-Verification for Edge Data Integrity
abstract
MobileEdgeComputing (MEC) has incentivized App vendors to outsource various services and applications to distributed edge nodes for low access latency. However, the data cached on these nodes is vulnerable to both intentional and accidental corruption, necessitating periodic audits ofEdgeDataIntegrity (EDI). Existing solutions either rely on a “fully trustworthy”ThirdPartyAuditor (TPA) or leverage blockchain to enhance trust. However, they overlook the security risks brought by the use of blockchain, particularly collusion attacks. Furthermore, while they employ achallenge-responsemechanism to enhance efficiency by batch verification, they fail to account for the heterogeneity of edge nodes. To address these challenges, we propose$\mathtt {CTCV}$, aCollusion-resistant andTime-awareCollaborativeVerification framework.$\mathtt {CTCV}$aims to accommodate edge node heterogeneity while enabling public audits and batch verification without introducing additional security risks. Specifically, it incorporates blockchain to allow edge nodes to collaboratively verify EDI without trust dependencies, while mitigating collusion attacks through a carefully designed proof generation and verification approach. Considering the resource and state heterogeneity of edge nodes,$\mathtt {CTCV}$employs atime-constrained challenge-responsemechanism that sets a time threshold$\mathcal {T}$between the verification request issuance and the integrity proof inspection to avoid excessive delays. The selection guideline of$\mathcal {T}$, along with the correctness, efficiency, and collusion resistance of$\mathtt {CTCV}$, are rigorously analyzed. Extensive experiments validate that$\mathtt {CTCV}$is computationally and communicationally efficient compared to three baselines: EdgeWatch, EDI-S, and EDI-V. On average, given 10 edge nodes,$\mathtt {CTCV}$outperforms EdgeWatch, EDI-S, and EDI-V with computation efficiency improvements of 7.9, 9.0, and 5.0 times, and communication efficiency improvement of 2063.0, 4.8, and 2.6 times, respectively.
Yao Zhao 0006, Youyang Qu, Bo Li 0103, Lu Zhao 0001, Feifei Chen 0001, Yong Xiang 0001, Longxiang Gao
IEEE Trans. Dependable Secur. Comput.7
2026 FiDD: Secure Fine-Grained Deduplication and Dynamic Auditing Scheme for Cloud Storage
abstract
With the rapid development of cloud computing, more and more users tend to store their data remotely to the cloud. Taking into account data security and resource utilization comprehensively, in addition to providing users with basic remote data integrity verification, cloud servers also need to conduct redundancy checks. However, current deduplication schemes primarily focus on static file-level data and auditing processes, rendering them inadequate for managing resources with dynamic attributes. In this paper, we propose a fine-grained deduplication and dynamic auditing model (FiDD) for cloud storage to address these challenges. FiDD utilizes homomorphic verifier-based data tags to seamlessly integrate deduplication and auditing processes, allowing both block-level and file-level deduplications. Additionally, FiDD employs doubly linked lists and multi-set hash functions to enhance the efficiency of data updates. The security of FiDD is validated through rigorous security proofs, while its efficiency is demonstrated through comprehensive experimental analysis. The experiments demonstrate an average improvement of at least 35% in audit efficiency and at least 50% in dynamics efficiency. Consequently, FiDD enhanes the security of data management in cloud computing while improving its overall efficiency.
Longxia Huang, Lei Zhou 0026, Di Wu 0050, Longxiang Gao, Tom H. Luan
IEEE Trans. Netw.5
2025 Focus on Local: Finding Reliable Discriminative Regions for Visual Place Recognition
abstract
Visual Place Recognition (VPR) is aimed at predicting the location of a query image by referencing a database of geotagged images. For VPR task, often fewer discriminative local regions in an image produce important effects while mundane background regions do not contribute or even cause perceptual aliasing because of easy overlap. However, existing methods lack precisely modeling and full exploitation of these discriminative regions. In addition, the lack of pixel-level correspondence supervision in the VPR dataset hinders further improvement of the local feature matching capability in the re-ranking stage. In this paper, we propose the Focus on Local (FoL) approach to stimulate the performance of image retrieval and re-ranking in VPR simultaneously by mining and exploiting reliable discriminative local regions in images and introducing pseudo-correlation supervision. First, we design two losses, Extraction-Aggregation Spatial Alignment Loss (SAL) and Foreground-Background Contrast Enhancement Loss (CEL), to explicitly model reliable discriminative local regions and use them to guide the generation of global representations and efficient re-ranking. Second, we introduce a weakly-supervised local feature training strategy based on pseudo-correspondences obtained from aggregating global features to alleviate the lack of local correspondences ground truth for the VPR task. Third, we suggest an efficient re-ranking pipeline that is efficiently and precisely based on discriminative region guidance. Finally, experimental results show that our FoL achieves the state-of-the-art on multiple VPR benchmarks in both image retrieval and re-ranking stages and also significantly outperforms existing two-stage VPR methods in terms of computational efficiency.
Changwei Wang 0001, Shunpeng Chen, Rongtao Xu, Jiguang Zhang, Haoran Yang 0003, Yu Zhang 0133, Kexue Fu 0001, Shide Du, Zhiwei Xu 0005, Longxiang Gao, Li Guo 0004, Shibiao Xu
AAAI12
2025 Spatial-Temporal Graph Diffusion Policy with Kinematic Modeling for Bimanual Robotic Manipulation
abstract
Despite the significant success of imitation learning in robotic manipulation, its application to bimanual tasks remains highly challenging. Existing approaches mainly learn a policy to predict a distant next-best end-effector pose (NBP) and then compute the corresponding joint rotation angles for motion using inverse kinematics. However, they suffer from two important issues: (1) rarely considering the physical robotic structure, which may cause self-collisions or interferences, and (2) overlooking the kinematics constraint, which may result in the predicted poses not conforming to the actual limitations of the robot joints. In this paper, we propose Kinematics enhanced Spatial-TemporAl gRaph Diffuser (KStar Diffuser). Specifically, (1) to incorporate the physical robot structure information into action prediction, KStar Diffuser maintains a dynamic spatial-temporal graph according to the physical bimanual joint motions at continuous timesteps. This dynamic graph serves as the robot-structure condition for denoising the actions; (2) to make the NBP learning objective consistent with kinematics, we introduce the differentiable kinematics to provide the reference for optimizing KStar Diffuser. This module regularizes the policy to predict more reliable and kinematics-aware next end-effector poses. Experimental results show that our method effectively leverages the physical structural information and generates kinematics-aware actions in both simulation and real-world.
Qi Lv 0001, Xiang Deng 0002, Rui Shao 0001, Yinchuan Li, Jianye Hao, Longxiang Gao, Michael Yu Wang, Liqiang Nie
CVPR7
2025 Dual Focus-Attention Transformer for Robust Point Cloud Registration
abstract
Recently, coarse-to-fine methods for point cloud registration have achieved great success, but few works deeply explore the impact of feature interaction at both coarse and fine scales. By visualizing attention scores and correspondences, we find that existing methods fail to achieve effective feature aggregation at the two scales during the feature interaction. To tackle this issue, we propose a Dual Focus-Attention Transformer framework, which only focuses on points relevant to the current point for feature interaction, avoiding interactions with irrelevant points. For the coarse scale, we design a superpoint focus-attention transformer guided by sparse keypoints, which are selected from the neighborhood of superpoints. For the fine scale, we only perform feature interaction between the point sets that belong to the same superpoint. Experiments show that our method achieve the state-of-the-art performance on three standard benchmarks. The code and pre-trained models are available at https://github.com/fukexue/DFAT.git.
Kexue Fu 0001, Mingzhi Yuan, Changwei Wang 0001, Weiguang Pang, Jing Chi, Manning Wang, Longxiang Gao
CVPR7
2025 AASD: Accelerate Inference by Aligning Speculative Decoding in Multimodal Large Language Models
abstract
Multimodal Large Language Models (MLLMs) have achieved notable success in visual instruction tuning, yet their inference is time-consuming due to the auto-regressive decoding of Large Language Model (LLM) backbone. Traditional methods for accelerating inference, including model compression and migration from language model acceleration, often compromise output quality or face challenges in effectively integrating multimodal features. To address these issues, we propose AASD, a novel framework for Accelerating inference with refined KV Cache and Aligning speculative decoding in MLLMs. Our approach leverages the target model’s cached KeyValue (KV) pairs to extract vital information for generating draft tokens, enabling efficient speculative decoding. To reduce the computational burden associated with long multimodal token sequences, we introduce a KV Projector to compress the KV Cache while maintaining representational fidelity. Additionally, we design a Target-Draft Attention mechanism that optimizes the alignment between the draft model and the target model, achieving the benefits of real inference scenarios with minimal computational overhead. Extensive experiments on mainstream MLLMs demonstrate that our method achieves up to a $2 \times$ inference speedup without sacrificing accuracy. This study not only provides an effective and lightweight solution for accelerating MLLM inference but also introduces a novel alignment strategy for speculative decoding in multimodal contexts, laying a strong foundation for future research in efficient MLLMs. Code is availiable at https://github.com/transcend-0/ASD
Muyang Zhang, Weiguang Pang, Yuzhi Chen, Rongtao Xu, Kexue Fu 0001, Changwei Wang 0001, Longxiang Gao
DAC9
2025 ContxE: Attention-based Context Aggregation for Temporal Knowledge Graph Completion
abstract
Knowledge graph completion (KGC) methods aim to predict missing links by learning from existing facts in a knowledge graph. Different from KGC, Temporal Knowledge Graph Completion (TKGC) further incorporates the time validity of facts (tagged timestamps) during the learning and inference to improve the completion accuracy. Many TKGC methods achieve this by projecting the static entity representations (time-invariant) of KGC embedding methods to time-dependent representations, which vary across timestamps. However, when measuring a fact, these TKGC methods only consider its subject/object entity representations corresponding to the tagged timestamp, but ignore their historical contexts that normally carry essential supportive information. With this observation, we propose a novel context aggregation (ContxE) method to include historical contexts of subject/object entities for TKGC. To achieve that, we propose a linear-rotary time embedding to obtain time-dependent entity representations that can preserve temporal relationships, and a relation-based attention to aggregate historical context for the score measurement. Comprehensive experiments on three temporal knowledge graph datasets show that the proposed ContxE achieves improved knowledge graph completion results compared to strong counterpart methods.
Borui Cai, Yong Xiang 0001, Longxiang Gao, Jiong Jin, Junfeng Wu 0010, Tom H. Luan
IJCNN3
2025 Resisting Catastrophic Recall: Persistent Unlearning via Knowledge Distillation with Feature Suppression
Zonghao Ji, Youyang Qu, Longxiang Gao, Taihao Zhang
KSEM (3)3
2025 Multi-scale Masked Transformer for Robust Point Cloud Registration
Taihao Zhang, Longxiang Gao, Youyang Qu, Zonghao Ji
KSEM (4)2
2025 Collaboration Wins More: Dual-Modal Collaborative Attention Reinforcement for Mitigating Large Vision Language Models Hallucination
abstract
Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in visual-language understanding for downstream multimodal tasks. However, these models often generate descriptions containing objects or details not present in the input image, a phenomenon commonly referred to as ''hallucination''. Existing methods focus solely on single-side hallucination mitigation: Intra-modal-only reinforcement (e.g. visual attention enhancement) ignores prompt-based guidance; Inter-modal-only correlation correction may introduce low-information visual tokens to mislead reasoning. To tackle this challenge, we propose Dual-Modal Collaborative Attention Reinforcement (DuCAR). Specifically, DuCAR is equipped with intra-visual CLS-driven sampling and cross-modal dynamic sampling, extracting important visual tokens guided by intra- and inter-modal joint information. During the multimodal fusion stage, DuCAR adaptively enhances the attention weights of these visual tokens. Our sampling and enhancement strategies in DuCAR simultaneously reinforces informative visual tokens, and suppresses attention dispersion towards question-irrelevant visual information. We conduct extensive experiments on the POPE and CHAIR hallucination benchmarks, demonstrating that our method outperforms existing state-of-the-art mitigation baselines and effectively reduces hallucinations in text generated by LVLMs. The code is available in the https://github.com/xjy2020/DuCAR.
Jiye Xie, Liangliang You, Zhiqiang Kou, Kexue Fu 0001, Youyang Qu, Wenjie Yang 0005, Jianwei Guo 0003, Weiliang Meng, Longxiang Gao, Haoran Yang 0003, Changwei Wang 0001, Yu Zhang 0133
ACM Multimedia12
2025 Wisdom is Knowing What not to Say: Hallucination-Free LLMs Unlearning via Attention Shifting
abstract
The increase in computing power and the necessity of AI-assisted decision-making boost the growing application of large language models (LLMs). Along with this, the potential retention of sensitive data of LLMs has spurred increasing research into machine unlearning. However, existing unlearning approaches face a critical dilemma: Aggressive unlearning compromises model utility, while conservative strategies preserve utility but risk hallucinated responses. This significantly limits LLMs' reliability in knowledge-intensive applications. To address this, we introduce a novel Attention-Shifting (AS) framework for selective unlearning. AS is driven by two design objectives: (1) context-preserving suppression that attenuates attention to fact-bearing tokens without disrupting LLMs' linguistic structure; and (2) hallucination-resistant response shaping that discourages fabricated completions when queried about unlearning content. AS realizes these objectives through two attention-level interventions, which are importance-aware suppression applied to the unlearning set to reduce reliance on memorized knowledge and attention-guided retention enhancement that reinforces attention toward semantically essential tokens in the retained dataset to mitigate unintended degradation. These two components are jointly optimized via a dual-loss objective, which forms a soft boundary that localizes unlearning while preserving unrelated knowledge under representation superposition. Experimental results show that AS improves performance preservation over the state-of-the-art unlearning methods, achieving up to 15\% higher accuracy on the ToFU benchmark and 10\% on the TDEC benchmark, while maintaining competitive hallucination-free unlearning effectiveness. Compared to existing methods, AS demonstrates a superior balance between unlearning effectiveness, generalization, and response reliability.
Chenchen Tan, Youyang Qu, Xinghao Li, Shujie Cui, Cunjian Chen, Longxiang Gao
NeurIPS7
2025 LAT: Luminance Information Assisted Collaborative Attention Transformer for Single Image Deraining
Weiyan Huang, Bruce Gu, Youyang Qu, Lei Cui 0006, Longxiang Gao
PRCV (8)6
2025 PointMM: A Hybrid Mamba-Transformer Framework for Point Cloud Analysis with Morton Reordering Strategy
Changwei Wang 0001, Shujun Gu, Chuanfu Wu, Longxiang Gao, Kexue Fu 0001, Youyang Qu
PRCV (10)6
2025 A Mamba-KAN Joint UNet Framework for Medical Image Segmentation
Haoyu Zhou, Changwei Wang 0001, Weiguang Pang, Lei Cui 0006, Shujun Gu, Longxiang Gao, Kexue Fu 0001, Youyang Qu
PRCV (3)6
2025 FedMLC: White-Box Model Watermarking for Copyright Protection in Federated Learning for IoT Environment
abstract
With the widespread application of the Internet of Things (IoT), data processing has gradually migrated to edge devices that are closer to the data source. This shift has significantly improved the ability of real-time data analysis while effectively reducing bandwidth requirements and latency. Furthermore, Federated Learning (FL) has been introduced as a decentralized training method to achieve collaborative training of multiple devices while ensuring local data privacy. However, malicious clients in FL may theft trained models for unauthorized use, which causes model misuse or copyright challenges. To address these issues, this paper proposes FedMLC (Malicious client detection, Leakage tracing, and Copyright verification), a server-side white-box watermarking scheme. FedMLC utilizes the embedded watermark at different stages to achieve both traceability and copyright verification, simplifying the watermarking process. Additionally, the watermarking can also detect malicious clients in FL. Specifically, FedMLC uses the regularization term to guide the parameter signs of the normalization layer to be consistent with the watermark sign, thereby achieving watermark embedding. Experimental results show that our FL model watermarking scheme excels in malicious client detection, leakage tracing, and copyright verification, with minimal impact on model performance, able to resist various attacks such as fine-tuning, pruning, and quantization.
Weitong Chen 0002, Wei Zhang 0098, Di Wu 0050, Anja Keskinarkaus, Tapio Seppänen, Jiale Zhang 0001, Longxiang Gao, Tom H. Luan
IEEE Internet Things J.7
2025 An Optimized Privacy-Utility Tradeoff Framework for Differentially Private Data Sharing in Blockchain-Based Internet of Things
abstract
Differential private (DP) query and response mechanisms have been widely adopted in various applications based on Internet of Things (IoT) to leverage variety of benefits through data analysis. The protection of sensitive information is achieved through the addition of noise into the query response which hides the individual records in a dataset. However, the noise addition negatively impacts the accuracy which gives rise to privacy-utility tradeoff. Moreover, the DP budget or cost$\epsilon $is often fixed and it accumulates due to the sequential composition which limits the number of queries. Therefore, in this article, we propose a framework known as optimized privacy-utility tradeoff framework for data sharing in IoT (OPU-TF-IoT). First, OPU-TF-IoT uses an adaptive approach to utilize the DP budget$\epsilon $by considering a new metric of population or dataset size along with the query. Second, our proposed heuristic search algorithm (HSA) reduces the DP budget accordingly whereas satisfying both data owner and data user. Third, to make the utilization of DP budget transparent to the data owners, a blockchain-based verification mechanism is also proposed. Finally, the proposed framework is evaluated using real-world datasets and compared with the traditional DP model and other related state-of-the-art works. The results demonstrate that our proposed framework not only utilizes the DP budget$\epsilon $efficiently, but also optimizes the number of queries by 49% and 54% on average compared to state-of-the-art and standard DP models, respectively. Furthermore, the data owners can effectively make sure that their data are shared accordingly through our blockchain-based verification mechanism which encourages them to share their data into the IoT system.
Muhammad Islam 0001, Mubashir Husain Rehmani, Longxiang Gao, Jinjun Chen
IEEE Internet Things J.3
2025 A Hybrid-Feature-Based Autoencoder Model for Predicting Traffic Flow to Accelerate Ambulance Medical Response Time in Internet of Vehicles
abstract
Traffic congestion during ambulance travel can delay medical response times. With the growing availability of traffic data, computational methods for traffic flow prediction are attracting significant attention. However, current computational prediction models have the following shortcomings. Fully connected networks require extensive feature engineering yet struggle to capture local traffic flow features. Their high parameter count increases training time and overfitting risk, while varying traffic conditions hinder model generalization. Therefore, in this work, we propose a hybrid-feature-based autoencoder (HFAE) model to predict traffic flow and accelerate ambulance medical response time in Internet of Vehicles. Specifically, the HFAE model simultaneously considers local features from the convolutional neural network and global features from the fully connected neural network. Additionally, the HFAE model uses an autoencoder to reconstruct the original input, capturing the intrinsic structure of the traffic flow data. Specifically, the root mean-squared error and mean absolute error scores of the HFAE model are better than those of the three mainstream models.
Yongmeng Li, Longxiang Gao, Lumin Xing
IEEE Internet Things J.4
2025 Morality-Driven Mechanism Design: Application in Hierarchical Carbon Trading Markets
Ruhan Liu, Yao Zhang 0005, Youyang Qu, Longxiang Gao, Yong Xiang 0001, Shang Gao 0003, Tom H. Luan
IEEE Internet Things J.4
2025 Federated Unlearning With Reinforcement Learning: Adaptive Privacy Preservation for Clients
abstract
With growing attention to data privacy in federated learning, federated unlearning has become an important solution to meet increasing demands for privacy compliance. However, unlearning may bring in new security concerns, such as dangers of adversarial manipulation, where the adversary may launch malicious updates or inputs to hurt the model performance or prediction, privacy-attacks, as the sensitive data can be possibly deduced from the process of unlearning, and performance degradation, because the unlearning process may break the consistency or performance of the model. In this paper, to address such issues and acquire a good and adaptive unlearning policy without causing much negative effect to the federated system, we present a reinforcement learning based method to facilitate the data unlearning method in federated learning. Our approach iteratively disposes of clients through partial unlearning, complete unlearning, or no unlearning using a DQN combined with clients’ properties like contribution, privacy cost, and computational overhead. We show that by utilizing the reinforcement learning technique, the performance decay can be defended effectively, and adversarial behaviors are indeed a common concern for the federated unlearning scenario. Our analysis can inform the development of federated unlearning frameworks that defend against performance and security threats.
Kun Gao 0006, Tianqing Zhu, Dayong Ye, Longxiang Gao, Wanlei Zhou 0001
J. Inf. Secur. Appl.4
2025 AirDIV: Over-the-Air Cloud-Fog Data Integrity Verification Scheme for Industrial Cyber-Physical Systems
abstract
Industrial Cyber-Physical Systems (ICPSs) have been motivating various Industry 4.0 endeavours, particularly with the integration of fog computing. Cloud-fog data caching paradigms, as supportive elements of ICPSs, have been adopted to cache user data, catering to diverse ICPS requirements such as data sensitivity and reduced access latency. In this hierarchical caching context, ensuring Cloud-Fog Data Integrity (CFDI) is crucial for maintaining the consistent functionality of ICPSs. Existing solutions primarily focus on examining the integrity of data cached solely on either cloud or fog nodes. However, cloud-cached data and fog-cached data are tightly coupled and should be considered simultaneously when checking data integrity. In this work, we introduce an over-the-air CFDI verification scheme, namely AirDIV, with a high accuracy and security guarantee. Instead of aggregating integrity proofs after proof transmission, AirDIV completes proof aggregation and transmission over the air for efficiency improvement. To enhance practicability, we derive adjustable parameters and formulate an optimization problem to minimize over-the-air aggregation errors. Furthermore, with an effective proof generation method, AirDIV can defend against two common attacks, i.e., replay and forge attacks. We provide a theoretical analysis of AirDIV’s correctness, accuracy and security, while conducting extensive experiments on both simulated and real platforms to validate its efficiency.
Yao Zhao 0006, Yong Xiang 0001, Md Palash Uddin, Yushu Zhang 0001, Lu Liu 0001, Longxiang Gao
IEEE J. Sel. Areas Commun.7
2025 Ddog: optimizing multi-hop inference via dual-driven retrieval and reasoning path
Bruce Gu, Longxiang Gao, Kexue Fu 0001, Youyang Qu, Lei Cui 0006
Mach. Learn.3
2025 Towards Efficient Consistency Auditing of Dynamic Data in Cross-Chain Interaction
abstract
As blockchain technology matures and its adoption grows across multiple industries, there is a growing need to make the data stored on blockchains adaptable to prevent misuse. However, such modifications can lead to inconsistency when interacting with other blockchains, necessitating the preservation of dynamic data consistency during these cross-chain interactions. While ensuring consistency for static data is relatively straightforward, doing so for dynamic data with efficiency remains a significant challenge. In response, we propose an efficient dynamic cross-chain data consistency auditing model (EDCA), which artfully integrates an advanced gamma multi-signature approach (AGMS) with the designed dynamic Merkle hash tree (D-MHT) to facilitate effective auditing of the consistency for dynamic data in cross-chain interaction. EDCA can produce relatively small auditor states while maintaining the storage proof. Moreover, EDCA is proven to satisfy strong security and privacy guarantees, tag unforgeability, and proof unforgeability. Experimental evaluations confirm that EDCA has high computational and communicational efficiency and can retain a small and relatively constant auditing overhead.
Yushu Zhang 0001, Longxiang Gao, Liehuang Zhu, Zhihong Tian 0001
IEEE Trans. Dependable Secur. Comput.4
2025 Trustworthy and Fair Federated Learning via Reputation-Based Consensus and Adaptive Incentives
abstract
Federated Learning (FL) allows collaborative training of a Machine Learning (ML) model while preserving data privacy across participating clients. Most existing studies consider FL clients to be proactive and completely honest in their participation. However, in reality, clients might lack the motivation to participate, and malicious behavior among some clients could negatively impact the interests of others. For these reasons, ensuring trust and fairness among FL clients is paramount but remains challenging due to limitations in FL consensus mechanisms and incentive strategies. To address these challenges, we introduce a Trustworthy and Fair FL (TFFL) framework that develops a reputation-based consensus mechanism called Dynamic Reputation Consensus (DRC), where clients’ reputations are dynamically assessed based on subjective opinions by evaluating real-time client behavior. We also incorporate time decay and temporal discounting of TFFL interactions along with the weighted measures of clients’ data quality, performance, and reliability to accurately reflect the evolving nature of client behavior over time. By adaptively adjusting clients’ incentives based on reputations and a cooperative game theory, DRC incentivizes honest participation and discourages malicious intent. In addition, we utilize blockchain and smart contracts to provide decentralized, regularized, and secure reputation management that is resistant to tampering and non-repudiation. Theoretical analysis and empirical results on widely used datasets (MNIST, CIFAR-10, and CIFAR-100) demonstrate the effectiveness of DRC in enhancing trust and fairness, improving performance, and providing robust security in FL settings. Results further exhibit that DRC offers superior performance in local model validation, consensus decision, and convergence time compared to related research approaches across various experimental settings.
Yong Xiang 0001, Md Palash Uddin, Jine Tang, Keshav Sood, Longxiang Gao
IEEE Trans. Inf. Forensics Secur.6
2025 DPNM: A Differential Private Notary Mechanism for Privacy Preservation in Cross-Chain Transactions
abstract
Notary cross-chain transaction technologies have obtained broad affirmation from industry and academia as they can avoid data islands and enhance chain interoperability. However, the increased privacy concern in data sharing makes the participants hesitate to upload sensitive information without the trust foundation of the external network. To address this issue, this paper proposes a differential private notary mechanism (DPNM) to preserve privacy in blockchain interoperations. It establishes a fully trusted notary organization to conduct data perturbation before replying query to the external blockchain network. In addition, the DPNM contains two built-in privacy budget allocation schemes: Efficiency priority scheme (EPS) and Privacy priority scheme (PPS). These schemes unify the privacy preferences among different nodes based on multi-node consensus in the decentralized environment. The EPS can generate noise linearly and work efficiently, and the PPS reflects better on nodes’ preferences. This paper utilizes several metrics including mechanism errors, elapsed time, latency, and gas consumption to evaluate the performance of DPNM compared to the traditional mechanisms. The experiment results indicate that the proposed mechanism can meet privacy preferences among different nodes and provide better utility with little extra cost.
Kai Zhang 0074, Pei-Wei Tsai, Jiao Tian, Ke Yu 0006, Hongwang Xiao, Xinyi Cai, Longxiang Gao, Jinjun Chen
IEEE Trans. Inf. Forensics Secur.8
2025 Privacy and Fairness Analysis in the Post-Processed Differential Privacy Framework
abstract
The post-processed Differential Privacy (DP) framework has been routinely adopted to preserve privacy while maintaining important invariant characteristics of datasets in data-release applications such as census data. Typical invariant characteristics include non-negative counts and total population. Subspace DP has been proposed to preserve total population while guaranteeing DP for sub-populations. Non-negativity post-processing has been identified to inherently incur fairness issues. In this work, we study privacy and unfairness (i.e., accuracy disparity) concerns in the post-processed DP framework. On one hand, we propose the post-processed DP framework with both non-negativity and accurate total population as constraints would inadvertently violate privacy guarantee desired by it. Instead, we propose thepost-processed subspace DP frameworkto accurately define privacy guarantees against adversaries. On the other hand, we identify unfairness level is dependent on privacy budget, count sizes as well as their imbalance level via empirical analysis. Particularly concerning is severe unfairness in the setting of strict privacy budgets. We further trace unfairness back touniform privacy budget setting over different population subgroups. To address this, we propose avarying privacy budget settingmethod and develop optimization approaches using ternary search and golden ratio search to identify optimal privacy budget ranges that minimize unfairness while maintaining privacy guarantees. Our extensive theoretical and empirical analysis demonstrates the effectiveness of our approaches in addressing severe unfairness issues across different privacy settings and several canonical privacy mechanisms. Using datasets of Australian Census data, Adult dataset, and delinquent children by county and household head education level, we validate both our privacy analysis framework and fairness optimization methods, showing significant reduction in accuracy disparities while maintaining strong privacy guarantees.
Ying Zhao 0012, Kai Zhang 0074, Longxiang Gao, Jinjun Chen
IEEE Trans. Inf. Forensics Secur.3
2025 Multiple Edge Data Integrity Verification With Multi-Vendors and Multi-Servers in Mobile Edge Computing
abstract
Ensuring Edge Data Integrity (EDI) is imperative in providing reliable and low-latency services in mobile edge computing. Existing EDI schemes typically address single-vendor (App Vendor, AV) single-server (Edge Server, ES), single-vendor multi-server, and multi-vendor multi-server scenarios, which consider a single data replica cached by an ES from the AVs. However, the most practical scenario of Multi-Vendors and Multi-Servers with Multiple Data (MVMS-MD) cached by an ES from different AVs remains unexplored. Current solutions struggle when applied to this scenario due to increased computation and communication costs in the verification process across all ESs using the classicalchallenge-response per-data multi-roundstrategy. To tackle this issue, we propose a Multiple EDI-Verification (MEDI-V) approach in this paper. In particular, our MEDI-V utilizes an adaptive Merkle Hash Tree (ad-MHT) to efficiently generate a tree of multiple data replicas within each AV. Next, the dynamic mechanism computes minimal verification information using ad-MHT to create achallengefor individual ESs to produce EDI proofs. The ES then leverages its ad-MHT and the ES's proof to send the reconstructed ad-MHT root to the AV for verification. Theoretical insights into MEDI-V's correctness, efficiency, security, and comprehensive evaluations demonstrate its superiority in addressing MEDI issues in the MVMS-MD scenario.
Yong Xiang 0001, Md Palash Uddin, Yao Zhao 0006, Jonathan Kua, Longxiang Gao
IEEE Trans. Mob. Comput.6
2025 SRIF: Data-Free Knowledge Distillation via Stable Regulation and Input Filtering
abstract
Data-free knowledge distillation (DFKD) enables knowledge transfer from a pre-trained teacher to a student network without accessing the real dataset. However, generator-based DFKD methods struggle to ensure that the synthetic images accurately reflect the real dataset distribution. The update of the generator network relies heavily on teacher category guidance, but varying teacher prediction accuracy across categories leads to inconsistent synthetic image quality. Such variations introduce a distribution shift between synthetic and real datasets, negatively impacting student network performance during knowledge distillation. To address this challenge, we propose the SRIF, comprising two components: Student-Driven Flexible Filtering (SDFF) and Re-weighting for Independent Regularization (RIR). SDFF filters out synthetic images affected by the category distribution shift during data generation, producing a more reliable dataset. RIR, applied during distillation, encourages the student to learn stable causal relationships through sample reweighting. Both components flexibly integrate into existing DFKD frameworks, improving performance while reducing training costs.
Rongtao Xu, Changwei Wang 0001, Shibiao Xu, Jie Zhou 0001, Longxiang Gao, Wenbo Xu 0003, Li Guo 0004
IEEE Trans. Multim.7
2025 Data Re-Outsourcing Detection With Latency-Constraint for Edge Storage
abstract
Edge storage has become a widely used solution for providing low-latency data access services, which motivates data owners to outsource data on geographically distributed edge nodes to deliver a positive user experience. Nevertheless, various security concerns raise in terms of data availability. Among them, edge data geo-location verification becomes a prominent concern when the data is out of owners' control, since outsourced data may be re-outsourced to other economical yet unknown third-party devices by dishonest edge nodes for saving storage space and pocketing the difference. Existing geo-localization approaches for cloud architectures can not be practically applied to identify such re-outsourcing behaviors due to the uniqueness of edge storage. To close this gap, we make the first attempt to investigate theedgedatare-outsourcingdetection (EDRD) problem, enabling the data owner to inspect if outsourced data is consistently cached on the rented edge nodes with agreed geo-location. We leverage timedChallenge-Responsemechanisms for data possession proof while measuring verification latency to detect re-outsourcing behaviors by comparing with re-outsourcing detection threshold$\mathbb {C}$. We prove that the edge node whose verification latency exceeds$\mathbb {C}$is dishonest. To obtain the optimal$\mathbb {C}$, we formulate thethresholddetermination (TD) problem and transform it to an easy-to-handle form for problem complexity reduction. Then, apreference-based approach named TD-P is developed to efficiently address the transformed TD problem. On top of that, we propose a$\mathbb {C}$-aware edge data re-outsourcing detection scheme entitled EDRD-$\mathbb {C}$to tackle the EDRD problem effectively. The efficiency and effectiveness of TD-P and EDRD-$\mathbb {C}$are verified by extensive theoretical analysis and experimental evaluations on both simulated and real platforms. Notably, EDRD-$\mathbb {C}$achieves 100% detection accuracy by sacrificing a reasonable amount of computing resources and 92.97% in the worst case.
Yao Zhao 0006, Youyang Qu, Yong Xiang 0001, Feifei Chen 0001, Longxiang Gao
IEEE Trans. Serv. Comput.5
2025 MM-SCS: Leveraging Multimodal Features to Enhance Smart Contract Code Search
abstract
Semantic code search technology allows searching for existing code snippets through natural language, which can greatly improve programming efficiency. Smart contracts, programs that run on the blockchain, have a code reuse rate of more than 79%, which means developers have a great demand for semantic code search tools. However, the existing code search models still have a semantic gap between code and query and perform poorly on specialized queries of smart contracts. In this paper, we propose a Multi-Modal Smart contract Code Search (MM-SCS) model. Specifically, we construct a Contract Elements Dependency Graph (CEDG) for MM-SCS as an additional modality to capture the data flow and control flow information of the code. To make the model more focused on the key contextual information, we use a multi-head attention network to generate embeddings for code features. In addition, we use a fine-tuned pretrained model to ensure the model's effectiveness when the training data is small. We compared MM-SCS with four state-of-the-art models on a dataset with 470K (code, docstring) pairs collected from Github and Etherscan. Experimental results show that MM-SCS achieves an MRR (Mean Reciprocal Rank) of 0.572, outperforming four state-of-the-art models UNIF, DeepCS, CARLCS-CNN, and TAB-CS by 34.2%, 59.3%, 36.8%, and 14.1%, respectively. Additionally, the search speed of MM-SCS is second only to UNIF, reaching 0.34s/query.
Chaochen Shi, Yong Xiang 0001, Jiangshan Yu, Longxiang Gao
IEEE Trans. Software Eng.4
2024 Control Flow Divergence Optimization by Exploiting Tensor Cores
abstract
Kernels are scheduled on Graphics Processing Units (GPUs) in the granularity of GPU warp, which is a bunch of threads that must be scheduled together. When executing kernels with conditional branches, the threads within a warp may execute different branches sequentially, resulting in a considerable utilization loss and unpredictable execution time. This problem is known as the control flow divergence. In this work, we propose a novel method to predict threads' execution path before the launch of the kernel by deploying a branch prediction network on the GPU's tensor cores, which can efficiently parallel run with the kernels on CUDA cores, so that the divergence problem can be eased in a large extent with the lowest overhead. Combined with a well-designed thread data reorganization algorithm, this solution can better mitigate GPUs' control flow divergence problem.
Weiguang Pang, Xu Jiang 0004, Songran Liu, Lei Qiao 0002, Kexue Fu 0001, Longxiang Gao, Wang Yi 0001
DAC6
2024 From Data Integrity to Global Model Integrity for Federated Learning: An MHT-based Approach
abstract
Federated Learning (FL) is a distributed machine learning (ML) approach that enables multiple edge nodes to collaboratively train ML models by sharing model parameters, thus addressing privacy concerns. However, in highly distributed, dynamic, and volatile FL environments, the global model is vulnerable to various corruptions. For instance, edge nodes might falsely claim that the received global model is incomplete, or the channel that transmits the global model is untrustworthy. Effectively verifying the integrity of the global model poses a critical challenge. To tackle this issue, we introduce a method for verifying model integrity called Federated learning global Model Integrity Verification (FMIV). It leverages Merkle Hash Tree (MHT) to generate integrity proofs of the global model during verification. To improve security, we integrate random security codes during proof generation. FMIV is capable of verifying the global model updated by the central server and shared with untrusted edge nodes, while efficiently identifying the edge node caching the incomplete global model. Furthermore, we conduct theoretical analysis and extensive experiments to validate the performance of FMIV. Compared to the two state-of-the-art approaches, FMIV consistently exhibits a notable improvement in verification efficiency and effectiveness in detecting model corruption.
Yao Zhao 0006, Y. Neil Qu, Bruce Gu, Keshav Sood, Longxiang Gao, Shui Yu 0001
GLOBECOM6
2024 Federated Meta Continual Learning for Efficient and Autonomous Edge Inference
Bingze Li, Stella Ho, Youyang Qu, Chenhao Xu 0003, Tom H. Luan, Longxiang Gao
ICA3PP (5)6
2024 Federated Learning and Parallel Prompt Scheduling Strategies for Large Language Models
Guangtong Lv, Bruce Gu, Xiaocong Jia, Longxiang Gao, Youyang Qu, Lei Cui 0006
ICA3PP (2)4
2024 DT-UPD: User Privacy Data Protection Through Distribution Transformation in Unlearning Cloud Service
Shouyue Sun, Lei Cui 0006, Longxiang Gao, Shui Yu 0001
ICA3PP (5)5
2024 Modal-Centric Insights Into Multimodal Federated Learning for Smart Healthcare: A Survey
Di Wang 0051, Longxiang Gao, Y. Neil Qu, Jihong Shi
ICA3PP (5)3
2024 Mitigating Over-Unlearning in Machine Unlearning with Synthetic Data Augmentation
Baohai Wang, Youyang Qu, Longxiang Gao, Conggai Li, Lin Li 0066, David B. Smith 0001
ICA3PP (4)3
2024 Incremental Graph Computation: Anchored Vertex Tracking in Dynamic Social Networks (Extended Abstract)
abstract
User engagement has recently received significant attention in understanding the decay and expansion of communities in many online social networking platforms. Many user engagement studies have been conducted to find a set of critical (anchored) users in the static social network. However, social networks are highly dynamic and their structures are continuously evolving. In this paper, we target a new research problem called Anchored Vertex Tracking (AVT), aiming to track the anchored users at each timestamp of evolving networks. To address the AVT problem, we develop a greedy algorithm inspired by the previous anchored k-core study in the static networks. Furthermore, we design an incremental algorithm to efficiently solve the AVT problem by utilizing the smoothness of the network structure's evolution. The extensive experiments demonstrate the performance of our proposed algorithms.
Taotao Cai, Shuiqiao Yang, Jianxin Li 0001, Quan Z. Sheng, Jian Yang 0001, Xin Wang 0030, Wei Zhang 0098, Longxiang Gao
ICDE8
2024 SGD-YOLOv5: A Small Object Detection Model for Complex Industrial Environments
abstract
Due to the complexity of industrial environments, such as construction sites and production workshops, the objects to be detected are easily occluded and perceived as small objects, which poses certain challenges for object detection. To ensure a safe industrial environment, this study adopts YOLOv5 as the basic framework and integrates the depth-to-space convolution module to improve the model’s ability to extract feature information of small targets. Second, the global attention mechanism is incorporated into the network to enhance the global interaction information, reducing the feature information loss, and improving the model performance. Finally, to alleviate the contradiction between classification and regression tasks in object detection, the YOLOv5 head is replaced with a decoupled head to achieve better classification and accelerate model convergence. To improve data diversity and enhance model robustness, we augmented the open-source safety helmet wearing dataset (SHWD) and smoking behavior detection dataset (SBDD). We test the performance of the proposed model (SGD-YOLOv5) on the large and small object detection dataset (SODA-D) and VisDrone2021-DET datasets. Furthermore, its ability to detect small objects was also evaluated. Experiments show that our model outperforms all baseline models on SHWD and SBDD. Compared to the TPH-YOLOv5 on the SODA-D dataset, AP and Recall achieve improvements of 18.3% to 19.2% and 11.9% to 13.3% respectively. On the Visdrone 2021-DET dataset, [email protected] achieved an improvement from 35.45% to 35.70% compared to the state-of-the-art model YOLO-Drone.
Jiabin Pei, Xiangzhi Liu, Longxiang Gao, Shui Yu 0001, James Xi Zheng
IJCNN4
2024 Grouped Federated Meta-Learning for Privacy-Preserving Rare Disease Diagnosis
abstract
Federated learning (FL) has been widely applied in medical field, which allows clients to collaboratively train global models without sharing local data. Nevertheless, the diversity and scarcity of samples from rare diseases may result in a decline in the performance of local models on client-side due to using a singular global model. Moreover, direct transmission of local models or parameters will likely lead to user privacy violations. To solve these problems, we propose a Grouped Federated Meta-Learning (GrFML) method to improve the performance of local personalization models while protecting data privacy. Specifically, we first utilize a self-attention mechanism to extract partial features from the client’s local data, which are uploaded to the server (medical data is susceptible to perturbation and data integrity, thus this process does not expose the private data). The server groups clients with similar features based on these extracted features. Then, multiple meta-models are trained on these groups and distributed back to the clients to enhance the performance of the client’s local models. Furthermore, during the FL process, we introduce dynamic perturbation to the uploaded gradients based on the model’s test accuracy to protect their privacy. Typically, the perturbation magnitude is directly proportional to the model’s test accuracy. Extensive experiments shown that the GrFML model significantly improves client personalization model accuracy and achieves a good privacy-utility trade-off.
Xinru Song, Zongchao Xie, Longxiang Gao, Lei Cui 0006, Youyang Qu, Shujun Gu
IJCNN4
2024 From Data Integrity to Global ModeI Integrity for Decentralized Federated Learning: A Blockchain-based Approach
abstract
Decentralized Federated Learning (DFL) is extensively applied in various areas, e.g., healthcare, finance, and Internet of Things (loT), offering practical solutions for distributed intelligent applications and data collaboration. In DFL systems, participants, e.g., edge devices, organizations, or nodes, collaborate in the training of a shared global model by aggregating local models from various participants. During this process, participants need to communicate frequently with a central authority/node/server to share model parameters. Such communication is vulnerable to malicious attacks or tampering, posing a significant threat to the integrity of model training. The integrity verification method can provide an integrity guarantee for the global model of DFL. However, most of the existing integrity verification schemes are centralized and not suitable for resource-constrained DFL scenarios. Therefore, how to verify the integrity of the global model becomes an important issue in DFL. To address it, we devise a global model integrity verification method for DFL. Specifically, we generate a digital signature for each global model parameter as proof of integrity, while improving the efficiency of integrity verification by electing delegates to conduct the verification process. A series of experiments is conducted to validate the performance of the proposed method. The experimental results demonstrate that our approach not only effectively ensures the integrity of the global model but also functions well under limited resources.
Yao Zhao 0006, Youyang Qu, Lei Cui 0006, Longxiang Gao
IJCNN6
2024 PDLG: Prevent Deep Leakage from Gradients based on Dataset Condensation in Federated Meta-Learning
abstract
Federated Meta-Learning (FML) has achieved privacy protection and intelligent improvement in many fields, such as healthcare and finance. However, the small dataset characteristic of FML may satisfy the strong assumptions of Deep Leakage from Gradients (DLG) attacks, making it susceptible to such attacks. Previous research has considered the strong assumptions of DLG to be impractical in real-world scenarios, resulting in a gap in the defense against DLG attacks in FML scenarios. In this paper, we propose a method called Prevent Deep Leakage Gradients (PDLG), PDLG utilizes condensation of training dataset to prevent DLG in FML, addressing the strong assumption problem of DLG in real-world scenarios. Specifically, PDLG first condenses the local raw training data through the client. Second use condensation data instead of the raw training data to train the model, and evaluate the changes in model classification performance in FML. Third DLG is used for the condensed dataset. As the Mean Square Error (MSE) between the real and the dummy gradients decreased, DLG could not restore the original images. The experimental results demonstrated that the classification performance of the models trained on condensation data remained within 10%. PDLG achieving a balance between model performance and privacy.
Wenhang Bian, Lei Cui 0006, Shouyue Sun, Longxiang Gao
MSN6
2024 FAST: A Dual-tier Few-Shot Learning Paradigm for Whole Slide Image Classification
abstract
The expensive fine-grained annotation and data scarcity have become the primary obstacles for the widespread adoption of deep learning-based Whole Slide Images (WSI) classification algorithms in clinical practice. Unlike few-shot learning methods in natural images that can leverage the labels of each image, existing few-shot WSI classification methods only utilize a small number of fine-grained labels or weakly supervised slide labels for training in order to avoid expensive fine-grained annotation. They lack sufficient mining of available WSIs, severely limiting WSI classification performance. To address the above issues, we propose a novel and efficient dual-tier few-shot learning paradigm for WSI classification, named FAST. FAST consists of a dual-level annotation strategy and a dual-branch classification framework. Firstly, to avoid expensive fine-grained annotation, we collect a very small number of WSIs at the slide level, and annotate an extremely small number of patches. Then, to fully mining the available WSIs, we use all the patches and available patch labels to build a cache branch, which utilizes the labeled patches to learn the labels of unlabeled patches and through knowledge retrieval for patch classification. In addition to the cache branch, we also construct a prior branch that includes learnable prompt vectors, using the text encoder of visual-language models for patch classification. Finally, we integrate the results from both branches to achieve WSI classification. Extensive experiments on binary and multi-class datasets demonstrate that our proposed method significantly surpasses existing few-shot classification methods and approaches the accuracy of fully supervised methods with only 0.22% annotation costs. All codes and models will be publicly available on https://github.com/fukexue/FAST.
Kexue Fu 0001, Xiaoyuan Luo, Linhao Qu, Shuo Wang 0011, Ilias Maglogiannis, Longxiang Gao, Manning Wang
NeurIPS7
2024 A Data Synchronization Incentive Scheme in Vehicular Digital Twin Network with Stackelberg Game
abstract
The evolving digital twin technology translates physical entities into the digital realm, allowing the exploration of abundant digital resources to optimize the task execution of these physical entities. Real-time data synchronization between physical entities and their digital twins is essential for the effective functioning of digital twin systems. In this paper, we investigate the challenge of data synchronization in vehicular digital twin networks operating in open street scenarios, where multiple vehicles rely on cellular networks for continuous data synchronization with their digital twins. Given the contention for cellular bandwidth among vehicles, a coordination scheme is required to manage resource allocation. As vehicles are fully distributed driven by self-interests only, a game-theoretic approach is proposed that leverages a cloud center controller to guide the sharing of cellular resources among digital twins. An optimal incentive mechanism is introduced to encourage digital twins to adhere to the center's guidance, promoting global social welfare. Through extensive simulations, we demonstrate that the proposed scheme successfully motivates vehicles to follow the center's guidance, leading to efficient data synchronization and mutual benefit maximization.
Jingru Tan, Jinkai Zheng, Tom H. Luan, Longxiang Gao, Zhou Su 0001
VTC Spring5
2024 POS-BERT: Point cloud one-stage BERT pre-training
Kexue Fu 0001, Peng Gao 0007, Shaolei Liu, Linhao Qu, Longxiang Gao, Manning Wang
Expert Syst. Appl.5
2024 EXVul: Toward Effective and Explainable Vulnerability Detection for IoT Devices
abstract
As with anything connected to the internet, Internet of Things (IoT) devices are also subject to severe cybersecurity threats because an adversary could exploit vulnerabilities in their internal software to perform malicious attacks. Despite the promising results of Deep Learning (DL)-based approaches, the lack of well-labeled IoT vulnerability samples available for training and explainability pose a critical challenge to deploy them in practice. In this paper, we propose, a novel DL-based approach for Effective and eXplainable IoT VULnerability detection. Specifically, inspired by recent advances of self-supervised learning in label-expensive tasks, we propose a new combinatorial contrastive loss to combine the strengths of large-scale unlabeled code corpus and limited IoT vulnerability samples. Then, given a binary detection result, provides a set of faithful and stable code statements positively contributing to the model’s predictions as understandable explanations. Experimental results indicate that outperforms state-of-the-art baselines by 33.44%-72.91% and 19.52%-98.78% with respect to the accuracy and F1 score metrics, respectively. For vulnerability explanation, improves over the best-performing baseline explainer PGExplainer by 22.97% in MSP, 49.55% in MSR, and 48.40% in MIoU, demonstrating that the explanations provided by can correctly point out the vulnerable statements relevant to the detected vulnerabilities.
Sicong Cao, Xiaobing Sun 0001, Wei Liu 0010, Di Wu 0050, Jiale Zhang 0001, Yan Li 0002, Tom H. Luan, Longxiang Gao
IEEE Internet Things J.8
2024 An Optimized Privacy-Protected Blockchain System for Supply Chain on Internet of Things
abstract
The consortium blockchain is being utilized in supply chains on the Internet of Things (IoT) for tracking and protecting supply chain data, such as manufacture, storage, and shipment. However, the supply chain data in a consortium blockchain is publicly accessible for all parties, which attracts widespread concerns about supply chain data privacy. Several existing attribute-based encryption (ABE)-based blockchain systems targeting to address the supply chain data privacy problem either bring about additional security problems or lack the feasibility analysis on IoTs. To address the aforementioned issues, in this article, a novel multiauthority ABE (MA-ABE)-based blockchain system is proposed to protect the data privacy for the supply chain on IoTs. Specifically, a four-way tradeoff optimization framework is designed so that the system decentralization, scalability, and storage consumption are not significantly affected by the improved privacy. The optimal attribute setting policies for different scale blockchain networks are dynamically generated by the nondominated sorting genetic algorithm II (NSGA-II). Extensive experiment results show that the proposed scheme remarkably improves data privacy protection for the supply chain without downgrading the other three key factors.
Chenhao Xu 0003, Youyang Qu, Yong Xiang 0001, Tom H. Luan, Longxiang Gao
IEEE Internet Things J.5
2024 A Learning-Based Hierarchical Edge Data Corruption Detection Framework in Edge Intelligence
abstract
Edge intelligence, an emerging distributed paradigm, is driven by the increasing number of Internet of Things devices and the development of edge computing and artificial intelligence. This paradigm revolutionizes the way of data caching by encouraging latency-sensitive data to be distributed across multiple edge nodes. In such data caching scenarios, ensuring the integrity of data stored at the edge nodes is critical for business continuity guarantee. Existing Edge Data Integrity (EDI) verification solutions rely on the interactive Challenge-Response mechanism. However, this mechanism imposes significant communication overhead on participants, leading to low verification efficiency. To address this challenge, we propose a Learning-based Hierarchical Edge Data Corruption Detection framework (LH-EDCD), aiming to enhance verification efficiency from a round perspective by reducing communication interaction between edge nodes and the data owner. LH-EDCD involves two layers of verification: internal and external. In the internal verification layer, each edge node self-inspects the cached data replica by running a corruption detection model distributedly trained by blockchain-based Federated Learning (FL). With such filtration, potential corruption can be efficiently identified without complex interaction. Considering the false positive existence in the model, in the external verification layer, LH-EDCD adopts a smart contract in blockchain to verify identified potentially corrupted data replicas for corruption confirmation, mitigating the trust concerns among edge nodes while reducing communication overhead on backbone networks. With the combination of these two layers, the overall EDI verification efficiency can be improved by reducing interaction verification time. Additionally, we make the first attempt to investigate the optimal verification time to improve the applicability and practicality of LH-EDCD. Extensive experimental results substantiate the advantages of employing FL in the first layer of LH-EDCD and demonstrate that LH-EDCD outperforms two state-of-the-art EDI approaches, i.e., EDI-S and EDI-V. Specifically, LH-EDCD achieves better model accuracy and convergence speed compared to centralized training, while exhibiting superior efficiency over EDI-S and EDI-V with 3.5 and 2.8 times performance improvements, respectively.
Yao Zhao 0006, Chenhao Xu 0003, Youyang Qu, Yong Xiang 0001, Feifei Chen 0001, Longxiang Gao
IEEE Internet Things J.6
2024 Enhancing Robustness of Speech Watermarking Using a Transformer-Based Framework Exploiting Acoustic Features
abstract
Digital watermarking serves as an effective approach for safeguarding speech signal copyrights, achieved by the incorporation of ownership information into the original signal and its subsequent extraction from the watermarked signal. While traditional watermarking methods can embed and extract watermarks successfully when the watermarked signals are not exposed to severe alterations, these methods cannot withstand attacks such as de-synchronization. In this work, we introduce a novel transformer-based framework designed to enhance the imperceptibility and robustness of speech watermarking. This framework incorporates encoders and decoders built on multi-scale transformer blocks to effectively capture local and long-range features from inputs, such as acoustic features extracted by Short-Time Fourier Transformation (STFT). Further, a deep neural networks (DNNs) based generator, notably the Transformer architecture, is employed to adaptively embed imperceptible watermarks. These perturbations serve as a step for simulating noise, thereby bolstering the watermark robustness during the training phase. Experimental results show the superiority of our proposed framework in terms of watermark imperceptibility and robustness against various watermark attacks. When compared to the currently available related techniques, the framework exhibits an eightfold increase in embedding rate. Further, it also presents superior practicality with scalability and reduced inference time of DNN models.
Chuxuan Tong, Iynkaran Natgunanathan, Yong Xiang 0001, Jianhua Li 0002, Tianrui Zong, James Xi Zheng, Longxiang Gao
IEEE ACM Trans. Audio Speech Lang. Process.7
2024 Context-Aware Consensus Algorithm for Blockchain-Empowered Federated Learning
abstract
Supported by cloud computing,FederatedLearning (FL) has experienced rapid advancement, as a promising technique to motivate clients to collaboratively train models without sharing local data. To improve the security and fairness of FL implementation, numerousBlockchain-empoweredFederatedLearning (BFL) frameworks have emerged accordingly. Among them, consensus algorithms play a pivotal role in determining the scalability, security, and consistency of BFL systems. Existing consensus solutions to block producer selection and reward allocation either focus on well-resourced scenarios or accommodate BFL based on clients' contributions to model training. However, these approaches limit consensus efficiency and undermine reward fairness, due to involving intricate consensus processes, disregarding clients' contributions during blockchain consensus, and failing to address lazy client problems (malicious clients plagiarizing local model updates from others to reap rewards). Given the aforementioned challenges, we make the first attempt to design a joint solution for efficient consensus and fair reward allocation in heterogeneous BFL systems with lazy clients. Specifically, we introduce a generalizable BFL workflow that can address lazy client problems well. Based on it, the global contribution of BFL clients is decoupled into five dominant metrics, and the block producer selection problem is formulated as a reward-constraint contribution maximization problem. By addressing this problem, the optimal block producer that maximizes global contribution can be identified to orchestrate consensus processes, and rewards are distributed to clients in proportion to their respective global contributions. To achieve it, we develop aContext-awareProof-of-Contribution consensus algorithm named CPoC to reach consensus and incentive simultaneously, followed by theoretical analysis of lazy client problems and privacy issues. Empirical results on widely-used datasets demonstrate the effectiveness of our design in improving consensus efficiency and maximizing global contribution.
Yao Zhao 0006, Youyang Qu, Yong Xiang 0001, Feifei Chen 0001, Longxiang Gao
IEEE Trans. Cloud Comput.5
2024 Design and Robust Evaluation of Next Generation Node Authentication Approach
abstract
The flexibility of 5G-NGNs makes them an ideal infrastructure for supporting mission-critical IoT applications that require low latency and high bandwidth. However, due to the rapid proliferation and the integration of IoTs with 5 G, the threat surface has considerably expanded. Hence the security of IoT devices is a big concern. Unfortunately, IoT devices have limited resources, and the traditional security approaches (authentication and intrusion detection approaches) of cryptography do not work effectively on 5G-IoT ecosystems. Motivated from this, we leverage the distinctive RF (Radio Frequency) fingerprinting signatures of IoT devices and used them to train a Deep learning model, Mahalanobis Distance theory in addition to the Chi-square distribution theory, to authenticate the IoT nodes. Under robust scenarios we have tested the approach shows detection accuracy (99.35%) as well as significant amount of reduction in model's training time as these two metrics are one of the primary key performance indicators (KPIs). In order to evaluate the effectiveness of the proposed method in real-time scenarios, we tested the proposed solution with a real RF dataset and the OSM-MANO 5 G platform. The model underwent formal verification using the Tamarin Prover tool, and the proposal was also compared with recent research works.
Dinh Duc Nha Nguyen, Keshav Sood, Yong Xiang 0001, Longxiang Gao, Lianhua Chi, Shui Yu 0001
IEEE Trans. Dependable Secur. Comput.4
2024 BASS: A Blockchain-Based Asynchronous SignSGD Architecture for Efficient and Secure Federated Learning
abstract
Federated learning (FL) is a distributed framework for machine learning that enables collaborative training of a shared model across data silos while preserving data privacy. However, the FL aggregation server faces a challenge in waiting for a large volume of model parameters from selected nodes before generating a global model, which leads to inefficient communication and aggregation. Although transmitting only the signs of stochastic gradient descent (SignSGD) reduces the transmission load, it decreases model accuracy, and the time waiting for local model collection remains substantial. Moreover, the security of FL is severely compromised by prevalent poisoning, backdoor, and DDoS attacks, causing ineffective and inaccurate model training. To overcome these challenges, this paper proposes aBlockchain-basedAsynchronousSignSGD (BASS) architecture for efficient and secure federated learning. By integrating a blockchain-based semi-asynchronous aggregation scheme with sign-based gradient compression, BASS considerably improves communication and aggregation efficiency, while providing resistance against attacks. Besides, a novel node-summarized sign aggregation algorithm is developed for the blockchain leaders to ensure the convergence and accuracy of the global model. An open-source prototype is developed, on top of which extensive experiments are conducted. The results validate the superiority of BASS in terms of efficiency, model accuracy, and security.
Chenhao Xu 0003, Jiaqi Ge, Longxiang Gao, Mengshi Zhang, Wanlei Zhou 0001, James Xi Zheng
IEEE Trans. Dependable Secur. Comput.4
2024 From Wide to Deep: Dimension Lifting Network for Parameter-Efficient Knowledge Graph Embedding
abstract
Knowledge graph embedding (KGE) that maps entities and relations into vector representations is essential for downstream applications. Conventional KGE methods require high-dimensional representations to learn the complex structure of knowledge graph, but lead to oversized model parameters. Recent advances reduce parameters by low-dimensional entity representations, while developing techniques (e.g., knowledge distillation or reinvented representation forms) to compensate for reduced dimension. However, such operations introduce complicated computations and model designs that may not benefit large knowledge graphs. To seek a simple strategy to improve the parameter efficiency of conventional KGE models, we take inspiration from that deeper neural networks require exponentially fewer parameters to achieve expressiveness comparable to wider networks for compositional structures. We view all entity representations as a single-layer embedding network, and conventional KGE methods that adopt high-dimensional entity representations equal widening the embedding network to gain expressiveness. To achieve parameter efficiency, we instead propose a deeper embedding network for entity representations, i.e., a narrow entity embedding layer plus a multi-layer dimension lifting network (LiftNet). Experiments on three public datasets show that by integrating LiftNet, four conventional KGE methods with 16-dimensional representations achieve comparable link prediction accuracy as original models that adopt 512-dimensional representations, saving 68.4% to 96.9% parameters.
Borui Cai, Yong Xiang 0001, Longxiang Gao, Di Wu 0050, He Zhang 0034, Jiong Jin, Tom H. Luan
IEEE Trans. Knowl. Data Eng.3
2024 ARFL: Adaptive and Robust Federated Learning
abstract
Federated Learning (FL) is a machine learning technique that enables multiple local clients holding individual datasets to collaboratively train a model, without exchanging the clients' datasets. Conventional FL approaches often assign a fixed workload (local epoch) and step size (learning rate) to the clients during the client-side local model training and utilize all collaborating trained models' parameters evenly during the server-side global model aggregation. Consequently, they frequently experience problems with data heterogeneity and high communication costs. In this paper, we propose a novel FL approach to mitigate the above problems. On the client side, we propose an adaptive model update approach that optimally allocates a needful number of local epochs and dynamically adjusts the learning rate to train the local model and regularizes the conventional objective function by adding a proximal term to it. On the server side, we propose a robust model aggregation strategy that potentially supplants the local outlier updates (models' weights) prior to the aggregation. We provide the theoretical convergence results and perform extensive experiments on different data setups over the MNIST, CIFAR-10, and Shakespeare datasets, which manifest that our FL scheme surpasses the baselines in terms of communication speedup, test-set performance, and global convergence.
Md Palash Uddin, Yong Xiang 0001, Borui Cai, Xuequan Lu, John Yearwood, Longxiang Gao
IEEE Trans. Mob. Comput.6
2024 SCEI: A Smart-Contract Driven Edge Intelligence Framework for IoT Systems
abstract
Federated learning (FL) enables collaborative training of a shared model on edge devices while maintaining data privacy. FL is effective when dealing with independent and identically distributed (iid) datasets, but struggles with non-iid datasets. Various personalized approaches have been proposed, but such approaches fail to handle underlying shifts in data distribution, such as data distribution skew commonly observed in real-world scenarios (e.g., driver behavior in smart transportation systems changing across time and location). Additionally, trust concerns among unacquainted devices and security concerns with the centralized aggregator pose additional challenges. To address these challenges, this paper presents a dynamically optimized personal deep learning scheme based on blockchain and federated learning. Specifically, the innovative smart contract implemented in the blockchain allows distributed edge devices to reach a consensus on the optimal weights of personalized models. Experimental evaluations using multiple models and real-world datasets demonstrate that the proposed scheme achieves higher accuracy and faster convergence compared to traditional federated and personalized learning approaches.
Chenhao Xu 0003, Jiaqi Ge, Longxiang Gao, Mengshi Zhang, Yong Xiang 0001, James Xi Zheng
IEEE Trans. Mob. Comput.5
2024 Data Integrity Verification in Mobile Edge Computing With Multi-Vendor and Multi-Server
abstract
The emergingMobileEdgeComputing (MEC) paradigm reforms the way of data caching by motivating App vendors to store latency-sensitive data on distributed edge servers. In volatile MEC environments, ensuringEdgeDataIntegrity (EDI) is a major concern for App vendors. Existing EDI solutions only consider the scenario with a single App vendor and multiple edge servers, neglecting more complex multi-vendor and multi-server cases. If multiple App vendors check their data replicas cached on the same edge server simultaneously, integrity verification efficiency will drop exponentially. To mitigate this challenge, we make the first attempt to develop aSmartInspectionAlgorithm (SIA) to pre-select unreliable data replicas for different App vendors in each verification round by jointly considering cache services' QoS (Quality-of-Service) and data replicas' unverified time. By implementing this approach, edge servers can merely verify the selected data replicas, greatly reducing computation and communication overheads in EDI verification. Theoretically, SIA can achieve$\mathcal {O}(n)$expected time complexity. Supported by SIA, we expand the EDI problem in multi-vendor and multi-server MEC environments (referred to as the MVMS-EDI problem) and propose a smart contract-based approach entitled MVMS-SC to tackle the problem efficiently and impartially. We provide a rigorous theoretical analysis of the correctness, security, and efficiency of MVMS-SC. Both large-scale and small-scale experiments with real-world datasets are correspondingly performed on a single machine and a real platform to validate the superiority of MVMS-SC in terms of computation and communication efficiencies.
Yao Zhao 0006, Youyang Qu, Feifei Chen 0001, Yong Xiang 0001, Longxiang Gao
IEEE Trans. Mob. Comput.5
2024 Long-Term Over One-Off: Heterogeneity-Oriented Dynamic Verification Assignment for Edge Data Integrity
abstract
EdgeIntelligence (EI), a burgeoning research area, motivates App vendors to cache data replicas on geographically distributed edge servers to deliver better services. On the downside, this benefit also incurs more data integrity audit overhead on App vendors, which calls for more efficientEdgeDataIntegrity (EDI) verification approaches. However, existing EDI solutions totally rely on an implicitresource homogeneity assumption-edge servers have identical resource availability throughout EDI inspection execution in each round-but it rarely holds in reality. The edge servers with insufficient computation and/or communication capacity greatly limit overall EDI verification efficiency from a round perspective. Thus, in this work, we release the identified impractical assumption and accordingly study the EDIDynamicVerificationAssignment (DVA) problem for the first time. The problem aims to maximize the number of data replicas being verified in the long term under the constraints of verification delay in resource-limited environments. In this way, App vendors merely need to check the integrity of selected data replicas in each round for efficiency improvement. Specifically, we first formalize the DVA problem as a delay-constrained long-term stochastic optimization problem and further prove its$\mathcal {NP}$-hardness. To resolve the problem efficiently, we decompose it to an easy-to-handle form and then develop a polynomial-timePriority-based approach named DVA-P with a theoretical analysis of its time complexity and performance bound. Finally, experimental evaluations validate that DVA-P can be seamlessly incorporated into existing EDI solutions to enhance overall verification efficiency while guaranteeing verification performance.
Yao Zhao 0006, Youyang Qu, Yong Xiang 0001, Chaochen Shi, Feifei Chen 0001, Longxiang Gao
IEEE Trans. Mob. Comput.6
2024 Prototype-Guided Memory Replay for Continual Learning
abstract
Continual learning (CL) is a machine learning paradigm that accumulates knowledge while learning sequentially. The main challenge in CL is catastrophic forgetting of previously seen tasks, which occurs due to shifts in the probability distribution. To retain knowledge, existing CL models often save some past examples and revisit them while learning new tasks. As a result, the size of saved samples dramatically increases as more samples are seen. To address this issue, we introduce an efficient CL method by storing only a few samples to achieve good performance. Specifically, we propose a dynamic prototype-guided memory replay (PMR) module, where synthetic prototypes serve as knowledge representations and guide the sample selection for memory replay. This module is integrated into an online meta-learning (OML) model for efficient knowledge transfer. We conduct extensive experiments on the CL benchmark text classification datasets and examine the effect of training set order on the performance of CL models. The experimental results demonstrate the superiority our approach in terms of accuracy and efficiency.
Stella Ho, Ming Liu 0028, Lan Du 0002, Longxiang Gao, Yong Xiang 0001
IEEE Trans. Neural Networks Learn. Syst.4
2024 Blockchained Dual-Asynchronous Federated Learning Services for Digital Twin Empowered Edge-Cloud Continuum
abstract
The booming of learning-based Artificial Intelligence (AI) enables the integration of Big Data and emerging computing architectures, which facilitate the Edge-AI-as-a-Service (EAaaS) in the edge-cloud continuum. To meet the emerging demands, such as privacy preservation and autonomy, blockchain-enabled federated learning (B-FL) is proposed, which further provides decentralized processing, data falsification avoidance, and learning model reliability. However, synchronous global aggregation, which is deployed in most existing B-FL paradigms, is dragging down the performances due to the data and computing resources heterogeneity of diverse edge devices. In addition, the restricted resources of edge devices pose further challenges in executing learning tasks and blockchain-based consensus simultaneously. To solve these issues, we propose a blockchained dual-asynchronous federated learning (BAFL-DT) service model for EAaaS in the digital twin empowered edge-cloud continuum. In BAFL-DT, federated learning services are run on local edge devices, while the global aggregation is achieved by the consensus process of digital twins implemented in the cloud. Besides, dual-asynchronous FL allows both local training and global aggregation to be performed in an asynchronous manner, which is uniquely enabled by the proposed paradigm. Extensive evaluations of real-world datasets testify to the superior performances of EAaaS by improving accuracy and efficiency.
Youyang Qu, Shui Yu 0001, Longxiang Gao, Keshav Sood, Yong Xiang 0001
IEEE Trans. Serv. Comput.3
2024 Adaptive Regularization and Resilient Estimation in Federated Learning
abstract
Federated Learning (FL) is an emerging research area that produces a globally trained model using numerous local users' data and maintains their privacy. Heterogeneous or non-Independent and Identically Distributed ( non-IID) data affect the global model's convergence and, therefore, cause high communication costs. These are because traditional FL approaches often disregard an adaptive regularized objective for the user-side training and utilize conventional arithmetic mean on the locally trained models for the server-side aggregation. To alleviate these issues, we propose a novel FL scheme in this paper. In particular, we propose an adaptive regularization approach to add to the classical objective function of the users' local models during training and a resilient estimation approach to the locally trained models during aggregation. The adaptive regularization approach is derived using the users' local and global performance diversification while the resilient estimation scheme uses a modified geometric mean aggregation over the local models' parameters. We provide consolidated theoretical results and perform extensive experiments on the IID and non-IID settings of MNIST, CIFAR-10, and Shakespeare datasets with various deep networks. The results manifest that our FL scheme outperforms the state-of-the-art approaches in terms of communication speedup, test-set performance, training convergence stability, and resiliency against attacks.
Md Palash Uddin, Yong Xiang 0001, Yao Zhao 0006, Mumtaz Ali 0003, Yushu Zhang 0001, Longxiang Gao
IEEE Trans. Serv. Comput.6
2024 Long-Term Proof-of-Contribution: An Incentivized Consensus Algorithm for Blockchain-Enabled Federated Learning
abstract
The surge in data collected by local devices has given rise to a distributed machine learning architecture namedFederatedLearning (FL) for privacy-preserving model training. However, the security of centralized aggregation of local models becomes a primary concern, which can be mitigated byBlockchain-enabledFederatedLearning (BFL) to facilitate decentralized model aggregation. In BFL, consensus and incentive are two of the key components that impact the scalability, security, and consistency of the system. Existing joint solutions focus on selecting a block producer based on client contributions to model training but overlook contributions to blockchain consensus and lack consideration for correlations across communication rounds, inevitably affecting incentive performance. Motivated by these, we make the first attempt to achieve blockchain consensus with along-term incentive guaranteefor BFL systems. Following a generalizable BFL workflow, we decouple the global contribution of BFL clients into four rigorously modeled metrics, and formulate the block producer selection problem as a long-term total contribution maximization problem with reward constraints. ALong-termProof-of-Contribution algorithm named LPoC is developed to handle this problem efficiently. In each communication round, LPoC identifies an optimal block producer that can maximize total contributions from a long-term perspective while allocating rewards to continuously motivate clients to contribute to BFL. We provide a detailed analysis of time complexity and performance bounds, followed by extensive experimental evaluations. The results demonstrate the effectiveness of LPoC in maximizing long-term total contribution, improving consensus efficiency, and upgrading training performance.
Yao Zhao 0006, Youyang Qu, Yong Xiang 0001, Feifei Chen 0001, Longxiang Gao
IEEE Trans. Serv. Comput.5
2024 Winning at the Starting Line: Unreliable Data Replica Selection for Edge Data Integrity Verification
abstract
MobileEdgeComputing (MEC) is an emerging technology, where App vendors are allowed to cache multiple data replicas on geographically distributed edge servers to serve adjacent mobile subscribers. However, this benefit introduces an extra workload for edge servers and App vendors, as they must audit the integrity of multiple data replicas periodically considering various threats caused by distributed and dynamic MEC environments. The large-scale growth of data replicas certainly is a challenge to design more efficientEdgeDataIntegrity (EDI) verification approaches. Existing solutions are mostly limited to improving efficiency by optimizing proof generation and verification methods, while the improvement is still far from satisfactory due to adopting indiscriminate inspection philosophy (checking all data replicas without discrimination). In this paper, we make the first attempt to abstract a pre-processing phase and correspondingly study theUnreliable dataReplicaSelection (URS) problem. It can be seamlessly integrated into existing EDI solutions by solving the URS problem at the start of each verification round. Such pre-selection can significantly enhance overall EDI verification efficiency by incorporating the cache serviceQualityofService (QoS) and verification success rate, especially in scenarios with a large number of data replicas. Specifically, we first formalize the URS problem as a constrained optimization problem, and further prove its$\mathcal {NP}$-hardness. To address the problem efficiently, we transform it into an easy-to-handle form and develop aPriority-based approach named URS-P. Both theoretical analysis and experimental evaluation validate the effectiveness and efficiency of our proposed solution.
Yao Zhao 0006, Youyang Qu, Yong Xiang 0001, Feifei Chen 0001, Md Palash Uddin, Longxiang Gao
IEEE Trans. Serv. Comput.6
2023 Temporal Knowledge Graph Completion: A Survey
abstract
Knowledge graph completion (KGC) predicts missing links and is crucial for real-life knowledge graphs, which widely suffer from incompleteness. KGC methods assume a knowledge graph is static, but that may lead to inaccurate prediction results because many facts in the knowledge graphs change over time. Emerging methods have recently shown improved prediction results by further incorporating the temporal validity of facts; namely, temporal knowledge graph completion (TKGC). With this temporal information, TKGC methods explicitly learn the dynamic evolution of the knowledge graph that KGC methods fail to capture. In this paper, for the first time, we comprehensively summarize the recent advances in TKGC research. First, we detail the background of TKGC, including the preliminary knowledge, benchmark datasets, and evaluation metrics. Then, we summarize existing TKGC methods based on how the temporal validity of facts is used to capture the temporal dynamics. Finally, we conclude the paper and present future research directions of TKGC.
Borui Cai, Yong Xiang 0001, Longxiang Gao, He Zhang 0034, Jianxin Li 0001
IJCAI3
2023 Bio-Inspired Dual-Network Model to Tackle Statistical Heterogeneity in Federated Learning
abstract
The problem of statistical heterogeneity in Federated Learning has been a major challenge, with existing solutions making unrealistic assumptions about the availability of shared datasets and high bandwidth between clients and the server. Solving this problem is crucial for the success of Federated Learning in real-world scenarios. In this work, we propose a biologically inspired dual-network model FedDual, which mimics how the human brain learns and memorizes the information. The model consists of a neocortical and a hippocampal network similar to those in the human brain. The hippocampal network is comprised by an image classification model, while the neocortical network is a variational auto-encoder responsible for long-term and re-callable memory. In this manner, FedDual uses the neocortical network to generate pseudo-patterns (synthetic data) on the server (global model). This allows for the hippocampal network to be trained with these pseudo-patterns. The dual-network architecture allows devices to share information via the weight updates of the neocortical network to the server without sending the actual data. We compare FedDual against alternatives elsewhere in the literature when applied to widely available datasets. FedDual not only achieves a margin of accuracy improvement over the alternatives, but also converges faster, requiring less communication rounds.
Adnan Ahmad, Vinh Loi Chau, Antonio Robles-Kelly, Shang Gao 0003, Longxiang Gao, Lianhua Chi, Wei Luo 0001
IJCNN5
2023 Machine translation-based fine-grained comments generation for solidity smart contracts
Chaochen Shi, Yong Xiang 0001, Jiangshan Yu, Keshav Sood, Longxiang Gao
Inf. Softw. Technol.5
2023 Trustworthy Sensor Fusion Against Inaudible Command Attacks in Advanced Driver-Assistance Systems
abstract
There are increasing concerns about malicious attacks on autonomous vehicles. In particular, inaudible voice command attacks pose a significant threat as voice commands become available in autonomous driving systems. How to empirically defend against these inaudible attacks remains an open question. Previous research investigates utilizing deep learning-based multimodal fusion for defense, without considering the model uncertainty in trustworthiness. As deep learning has been applied to increasingly sensitive tasks, uncertainty measurement is crucial in helping improve model robustness, especially in mission-critical scenarios. In this article, we propose the multimodal fusion framework (MFF) as an intelligent security system to defend against inaudible voice command attacks. MFF fuses heterogeneous audio–vision modalities using VGG family neural networks and achieves the detection accuracy of 92.25% in the comparative fusion method empirical study. Additionally, extensive experiments on audio–vision tasks reveal the model’s uncertainty. Using expected calibration errors, we measure calibration errors and Monte Carlo Dropout to estimate the predictive distribution for the proposed models. Our findings show empirically to train robust multimodal models, improve standard accuracy and provide a further step toward interpretability. Finally, we discuss the pros and cons of our approach and its applicability for advanced driver assistance systems.
Jiwei Guan, Lei Pan 0002, Chen Wang 0008, Shui Yu 0001, Longxiang Gao, James Xi Zheng
IEEE Internet Things J.5
2023 Toward IoT Node Authentication Mechanism in Next Generation Networks
abstract
Although the next generation networks (5G-NGNs) provide a flexible infrastructure to support latency-sensitive and bandwidth-hungry mission-critical Internet of Things (IoT) applications, however, the 5G-IoT integration in NGNs has increased the threat surface. Unfortunately, IoT devices are resource constrained, and the traditional intrusion detection systems (IDS) approaches based on cryptography are not effective on 5G-IoT ecosystems. In this article, we propose an effective 5G-IoT node authentication approach that leverages unique radio frequency (RF) fingerprinting data to train the Deep learning model to detect legitimate and nonlegitimate IoT nodes. Our approach is based on Mahalanobis Distance theory and Chi-square distribution theories. The proposed approach achieves a higher detection accuracy (99.35%) as well as lower training time compared to other existing approaches which is a key benefit of our approach in NGNs. The experiments are conducted using ETSI-open source NFV management and orchestration (OSM-MANO) platform on Amazon Web Services (AWSs) cloud platform to verify how the proposed approach would fit in real-life scenarios. The method can be used as a standalone security system or as a part of multifactor authentication.
Dinh Duc Nha Nguyen, Keshav Sood, Yong Xiang 0001, Longxiang Gao, Lianhua Chi, Shui Yu 0001
IEEE Internet Things J.4
2023 Designing a Secure Blockchain-Based Supply Chain Management Framework
abstract
Supply chain management (SCM) faces a critical security issue because of the asymmetry of information delivered to various parties in the ecosystem and the lack of corresponding supervision. In response, we propose the use of blockchain technology to address the SCM security issues and put forward a blockchain-based SCM framework. We apply design science paradigm to guide the blockchain-based SCM framework development and implementation of a proof-of-concept prototype. We use Hyperledger Fabric and Composer to develop the prototype artifact. Performance evaluation results issued from Hyperledger Caliper prove the superiority and robustness of the proposed blockchain-based framework in terms of security and efficiency requirements, and performance metrics including throughput and latency. Also, the evaluation results show that the IT artifact is stable, and the high stability can reduce the risks of system vulnerabilities and breakdown.
Jiongbin Liu, William Yeoh 0002, Longxiang Gao, Shang Gao 0003, Ojelanki K. Ngwenyama
J. Comput. Inf. Syst.3
2023 Hybrid variational autoencoder for time series forecasting
abstract
Variational autoencoders (VAE) are powerful generative models that learn the latent representations of input data as random variables. Recent studies show that VAE can flexibly learn the complex temporal dynamics of time series and achieve more promising forecasting results than deterministic models. However, a major limitation of existing works is that they fail to jointly learn the local patterns (e.g., seasonality and trend) and temporal dynamics of time series for forecasting. Accordingly, we propose a novel hybrid variational autoencoder (HyVAE) to integrate the learning of local patterns and temporal dynamics by variational inference for time series forecasting. Experimental results on four real-world datasets show that the proposed HyVAE achieves better forecasting results than various counterpart methods, as well as two HyVAE variants that only learn the local patterns or temporal dynamics of time series, respectively.
Borui Cai, Shuiqiao Yang, Longxiang Gao, Yong Xiang 0001
Knowl. Based Syst.3
2023 Query-Efficient Black-Box Adversarial Attacks on Automatic Speech Recognition
abstract
The susceptibility of Deep Neural Networks (DNNs) to adversarial attacks has raised concerns regarding their practical applications in real-world scenarios. Although the vulnerability of DNNs to adversarial attacks has been extensively studied in the image domain, research in the audio domain, particularly in the black-box setting with Automatic Speech Recognition (ASR) models, remains limited. While various black-box attacks have been proposed for ASR models, such as transfer attacks, hardware attacks, and query-based attacks, this study concentrates on query-based black-box attacks. The article introduces a new gradient estimation technique, Temporal Natural Evolution Strategies (T-NES), to generate adversarial audio samples more efficiently than existing attacks. T-NES leverages the temporal correlation present in audio to speed up gradient estimation based on the probability scores returned by the target model. The empirical results on benchmark datasets, LibriSpeech and TEDLIUM, and two state-of-the-art ASR models, DeepSpeech2 and Wav2Letter, demonstrate that T-NES can generate successful attacks with up to 30% fewer queries than existing attacks within 500 queries. T-NES could provide a robust baseline for evaluating the black-box adversarial vulnerability of ASR systems.
Chuxuan Tong, James Xi Zheng, Jianhua Li 0002, Xingjun Ma, Longxiang Gao, Yong Xiang 0001
IEEE ACM Trans. Audio Speech Lang. Process.5
2023 SSVS-SSVD Based Desynchronization Attacks Resilient Watermarking Method for Stereo Signals
abstract
Most of the audio signals in real-world applications are stereo signals. However, the previous desynchronization attacks resilient watermarking methods cannot preserve perceptual quality or achieve robustness when constrained by high embedding rates and stereo host. In this paper, based on two novel features segmental singular values summation (SSVS) and segmental singular values difference (SSVD) that are generated using discrete cosine transform (DCT) and singular value decomposition (SVD), we present a robust watermarking method for stereo signals that not only is robust to desynchronization attacks and common signal processing attacks but also has a larger embedding rate compared with the previous methods. In the proposed method, we first apply DCT and SVD on each segment of the host signal to extract the SSVS feature and the SSVD feature. Then we generate the adaptive embedding parameters and embed watermark bits via optimized embedding strategies based on these features. Due to the use of the adaptive embedding parameters and the optimized embedding strategies, the proposed method significantly increases the embedding rate without compromising the robustness and perceptual quality. Analysis results show our proposed method outperforms the state-of-the-art methods by a large margin, where the perceptual quality improvement is over 14%, and the robustness against desynchronization attacks is improved by more than 49% when the embedding rate is 70 bps.
Juan Zhao 0007, Tianrui Zong, Yong Xiang 0001, Longxiang Gao, Guang Hua 0001, Keshav Sood, Yushu Zhang 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2023 Frequency Spectrum Modification Process-Based Anti-Collusion Mechanism for Audio Signals
abstract
The collusion attack combines multiple multimedia files into one new file to erase the user identity information. The traditional anti-collusion methods (which aim to trace the traitors) can defend the collusion attack, but they cannot well defend some hybrid collusion attacks (e.g., a collusion attack combined with desynchronization attacks). To address this issue, we propose a frequency spectrum modification process (FSMP) to defend the collusion attack by significantly downgrading the perceptual quality of the colluded file. The severe perceptual quality degradation can demotivate the attackers from launching the collusion attack. Because FSMP is orthogonal to the existing traitor-trace-based methods, it can be combined with the existing methods to provide a double-layer protection against different attacks. In FSMP, after several signal processing procedures (e.g., uneven framing and smoothing), multiple signals (called FSMP signals) can be generated from the host signal. Launching collusion attack using the generated FSMP signals would lead to the energy disturbance and attenuation effect (EDAE) over the colluded signals. Due to the EDAE, FSMP can significantly degrade the perceptual quality of the colluded audio file, thereby thwarting the collusion attack. In addition, FSMP can well defend different hybrid collusion attacks. Theoretical analysis and experimental results confirm the validity of the proposed method.
Juan Zhao 0007, Tianrui Zong, Yong Xiang 0001, Guang Hua 0001, Longxiang Gao, Gleb Beliakov
IEEE Trans. Cybern.6
2023 Incremental Graph Computation: Anchored Vertex Tracking in Dynamic Social Networks
abstract
User engagement has recently received significant attention in understanding the decay and expansion of communities in many online social networking platforms. When a user chooses to leave a social networking platform, it may cause a cascading dropping out among her friends. In many scenarios, it would be a good idea to persuade critical users to stay active in the network and prevent such a cascade because critical users can have significant influence on user engagement of the whole network. Many user engagement studies have been conducted to find a set of critical(anchored)users in the static social network. However, social networks are highly dynamic and their structures are continuously evolving. In order to fully utilize the power of anchored users in evolving networks, existing studies have to mine multiple sets of anchored users at different times, which incurs an expensive computational cost. To better understand user engagement in evolving network, we target a new research problem calledAnchored Vertex Tracking(AVT) in this paper, aiming to track the anchored users at each timestamp of evolving networks. Nonetheless, it is nontrivial to handle the AVT problem which we have proved to be NP-hard. To address the challenge, we develop a greedy algorithm inspired by the previous anchored$k$-core study in the static networks. Furthermore, we design an incremental algorithm to efficiently solve the AVT problem by utilizing the smoothness of the network structure's evolution. The extensive experiments conducted on real and synthetic datasets demonstrate the performance of our proposed algorithms and the effectiveness in solving the AVT problem.
Taotao Cai, Shuiqiao Yang, Jianxin Li 0001, Quan Z. Sheng, Jian Yang 0001, Xin Wang 0030, Wei Zhang 0098, Longxiang Gao
IEEE Trans. Knowl. Data Eng.8
2023 A Comprehensive Survey on Multi-View Clustering
abstract
The development of information gathering and extraction technology has led to the popularity of multi-view data, which enables samples to be seen from numerous perspectives. Multi-view clustering (MVC), which groups data samples by leveraging complementary and consensual information from several views, is gaining popularity. Despite the rapid evolution of MVC approaches, there has yet to be a study that provides a full MVC roadmap for both stimulating technical improvements and orienting research newbies to MVC. In this article, we review recent MVC techniques with the purpose of exhibiting the concepts of popular methodologies and their advancements. This survey not only serves as a unique MVC comprehensive knowledge for researchers but also has the potential to spark new ideas in MVC research. We summarise a large variety of current MVC approaches based on two technical mechanisms: heuristic-based multi-view clustering (HMVC) and neural network-based multi-view clustering (NNMVC). We end with four technological approaches within the category of HMVC: nonnegative matrix factorisation, graph learning, latent representation learning, and tensor learning. Deep representation learning and deep graph learning are two technical methods that we demonstrate in NNMVC. We also show 15 publicly available multi-view datasets and examine how representative MVC approaches perform on them. In addition, this study identifies the potential research directions that may require further investigation in order to enhance the further development of MVC.
Uno Fang, Jianxin Li 0001, Longxiang Gao, Tao Jia 0001, Yanchun Zhang
IEEE Trans. Knowl. Data Eng.4
2023 Federated Learning via Disentangled Information Bottleneck
abstract
Existing Federated Learning (FL) algorithms generally suffer from high communication costs and data heterogeneity due to the use of conventional loss function for local model update and the equal consideration of each local model for global model aggregation. In this paper, we propose a novel FL approach to address the above issues. For local model update, we propose a disentangled Information Bottleneck (IB) principle-based loss function. For global model aggregation, we suggest a model selection strategy based on Mutual Information (MI). Particularly, we design a Lagrangian-based loss function using the IB principle and “disentanglement” for maximizing MI between the ground truth and model prediction and minimizing MI between the intermediate representations. We calculate MI ratio between the ground truth and model prediction, and between the original input and ground truth to select the effective models for aggregation. We analyze the theoretical optimal cost of the loss function and manifest optimal convergence rate, and quantify the outlier robustness of the aggregation scheme. Experiments demonstrate the superiority of the proposed FL approach, in terms of testing performance and communication speedup (i.e., 3.00-14.88 times for IID MNIST, 2.5-50.75 times for non-IID MNIST, 1.87-18.40 times for IID CIFAR-10, and 1.24-2.10 times for non-IID MIMIC-III).
Md Palash Uddin, Yong Xiang 0001, Xuequan Lu, John Yearwood, Longxiang Gao
IEEE Trans. Serv. Comput.5
2023 A Lightweight Model-Based Evolutionary Consensus Protocol in Blockchain as a Service for IoT
abstract
Internet of Things (IoT) is experiencing fast proliferation with emerging trends in autonomy and local decision-making to avoid the explosive burden on network infrastructure between cloud and edge. Thereby, blockchain as a Service (BaaS) for IoT, as an emerging distributed services computing paradigm, has drawn intense attention due to its decentralization, auditability, and tamper-resistance. However, the primary challenge is to design a tailor-made consensus protocol that is applicable to BaaS for IoT. Existing consensus protocols generally focus on power-intensive environments, which is not feasible for power-constrained BaaS-enabled IoT systems. In this article, to fully exploit BaaS's superiority (e.g., to sharing data securely), we propose a lightweight model-based evolutionary consensus protocol called Proof of Evolutionary Model (PoEM) that can improve the quality of BaaS in IoT environments. Beyond existing rule-based consensus protocols, PoEM iteratively trains a machine learning model to achieve consensus. In this way, PoEM enhances consensus efficiency and enables low-performance IoT devices to be involved. Moreover, considering IoT environments’ dynamics, a novel mechanism is designed to manage nodes joining and exiting dynamically. Extensive analytical and experimental results show PoEM's improved consensus efficiency and applicability in dynamic BaaS-based IoT environments while providing high-level security guarantees.
Yao Zhao 0006, Youyang Qu, Yong Xiang 0001, Yushu Zhang 0001, Longxiang Gao
IEEE Trans. Serv. Comput.5
2023 CoSS: Leveraging Statement Semantics for Code Summarization
abstract
Automated code summarization tools allow generating descriptions for code snippets in natural language, which benefits software development and maintenance. Recent studies demonstrate that the quality of generated summaries can be improved by using additional code representations beyond token sequences. The majority of contemporary approaches mainly focus on extracting code syntactic and structural information from abstract syntax trees (ASTs). However, from the view of macro-structures, it is challenging to identify and capture semantically meaningful features due to fine-grained syntactic nodes involved in ASTs. To fill this gap, we investigate how to learn more code semantics and control flow features from the perspective of code statements. Accordingly, we propose a novel model entitled CoSS for code summarization. CoSS adopts a Transformer-based encoder and a graph attention network-based encoder to capture token-level and statement-level semantics from code token sequence and control flow graph, respectively. Then, after receiving two-level embeddings from encoders, a joint decoder with a multi-head attention mechanism predicts output sequences verbatim. Performance evaluations on Java, Python, and Solidity datasets validate that CoSS outperforms nine state-of-the-art (SOTA) neural code summarization models in effectiveness and is competitive in execution efficiency. Further, the ablation study reveals the contribution of each model component.
Chaochen Shi, Borui Cai, Yao Zhao 0006, Longxiang Gao, Keshav Sood, Yong Xiang 0001
IEEE Trans. Software Eng.4
2023 Data Synchronization in Vehicular Digital Twin Network: A Game Theoretic Approach
abstract
A fundamental issue of the vehicular digital twin (DT) is efficiently synchronizing the data between the DT and the vehicular user (VUE). In this paper, we consider the heterogeneous vehicular networks (HetVNets) in which a VUE can connect to the network through different networks. The HetVNets can improve the efficiency of communication by providing seamless connections. However, the uneven distribution of VUEs and the dynamics of HetVNets make the environment more complex. Therefore, we propose the network selection algorithm for data synchronization between VUEs and DTs in the HetVNets, where the behaviour between the VUEs is considered as a competition for wireless resources. A learning-based prediction model residing in the DT is developed where the DT can predict the waiting time of each relay and transmit the predicted results to the VUE for decision-making. We model the network selection problem as a potential game considering both the transmission time and the waiting time obtained from the prediction model and prove the existence of Nash equilibrium (NE). We analyze the performance of the proposed algorithm, and simulation results show that our approach can effectively find the optimal strategy while achieving a fast convergence speed and high-level performance compared to the baselines.
Jinkai Zheng, Tom H. Luan, Yao Zhang 0005, Rui Li 0047, Yilong Hui, Longxiang Gao, Mianxiong Dong
IEEE Trans. Wirel. Commun.6
2022 Semi-supervised Continual Learning with Meta Self-training
abstract
Continual learning (CL) aims to enhance sequential learning by alleviating the forgetting of previously acquired knowledge. Recent advances in CL lack consideration of the real-world scenarios, where labeled data are scarce and unlabeled data are abundant. To narrow this gap, we focus on semi-supervised continual learning (SSCL). We exploit unlabeled data under limited supervision in the CL setting and demonstrate the feasibility of semi-supervised learning in CL. In this work, we propose a novel method, namely Meta-SSCL, which combines meta-learning with pseudo-labeling and data augmentations to learn a sequence of semi-supervised tasks without catastrophic forgetting. Extensive experiments on CL benchmark text classification datasets show that our method achieves promising results in SSCL.
Stella Ho, Ming Liu 0028, Lan Du 0002, Longxiang Gao, Shang Gao 0003
CIKM5
2022 BASS: Blockchain-Based Asynchronous SignSGD for Robust Collaborative Data Mining
abstract
Federated learning (FL) is a machine learning framework for collaborative data mining in many scenarios (e.g. Internet of Things) due to its privacy-preserving feature. However, various attacks arise security concerns of FL, such as poisoning, backdoor, and DDoS attacks. Several blockchain-based FL schemes strengthen credibility and security without considering the increased communication overhead. Some existing work compresses local updated gradients to sign vectors to lower communication overhead at the expense of model accuracy. To address the above concerns, this paper offers a blockchain-based asynchronous SignSGD (BASS) scheme. A novel asynchronous sign aggregation algorithm is introduced to ensure model accuracy even if the local updated gradients are compressed to sign vectors. Considering the unstable network connection on IoT, a consensus algorithm that elects multiple leader nodes enables reliable global model aggregation. The introduced blockchain improves credibility and security without downgrading efficiency. Empirical studies show that BASS outperforms other schemes in efficiency, model accuracy, and security.
Chenhao Xu 0003, Youyang Qu, Yong Xiang 0001, Longxiang Gao, David B. Smith 0001, Shui Yu 0001
DSAA4
2022 Impersonation Attack Detection in IoT Networks
abstract
The deployment of Internet of Things (IoT) networks is growing at an extraordinary speed from last decade and has expanded the interconnection of billions of nodes, providing a range of flexible communication and computing services, etc. We note that this significant expansion of the IoT surface has expanded the attack surfaces and is a danger to companies of every size from security aspects. The IoT devices are easy to compromise and therefore the attacker can easily act as an impersonator to impersonate other legitimate IoT nodes. This is known as impersonation attacks or spoofing attacks in wireless IoT networks. In this paper, we propose a new methodology to detect an impersonation attack in IoT networks. We use Mahalanobis Distance correlation theory based two-stage attack detection model to resist IoT node spoofing. The approach is evaluated on cloud platforms and is compared with the recent state-of-the-art literature. The proposal is deployed as a pluggable module in cloud networks. The key metrics of our evaluation and comparisons are accuracy with respect to the varying size of the IoT network, classification metrics, attack detection time, and CPU utilization.
Dinh Duc Nha Nguyen, Keshav Sood, Yong Xiang 0001, Longxiang Gao, Lianhua Chi
GLOBECOM4
2022 Personalized Privacy-Preserving Medical Data Sharing for Blockchain-based Smart Healthcare Networks
abstract
With the growing proliferation of intelligent end devices and data analytics techniques, real momentum towards the development of smart healthcare networks (SHN) has already been evident. Multiple parties in SHNs continuously exchange medical data in order to achieve a precise diagnosis and process optimization. Privacy issue emerges since medical data are susceptible, while the combination of a series of medical data may lead to further privacy leakage. Adversaries launch unceasingly launch poisoning attacks, a dominant attack to maliciously manipulate data, severely impact the authenticity of the data transmitting over the SHNs, leading to misdiagnosing or even physical damage. In this paper, we propose a personalized differential privacy model built upon blockchain, in which the community density is exploited to customize the degree of privacy protection and inject corresponding noise data. Besides using blockchain as the underlying network architecture to defeat poisoning attacks. The proposed model can guarantee the authentication of the differentially private data, traceability of data, and single-point failure avoidance in SHN. Evaluation and extensive results using real-world data sets demonstrate the superiority of the proposed model.
Youyang Qu, Shiping Chen 0001, Longxiang Gao, Lei Cui 0006, Keshav Sood, Shui Yu 0001
ICC3
2022 Attention Distraction: Watermark Removal Through Continual Learning with Selective Forgetting
abstract
Fine-tuning attacks are effective in removing the embedded watermarks in deep learning models. However, when the source data is unavailable, it is challenging to just erase the watermark without jeopardizing the model performance. In this context, we introduce Attention Distraction (AD), a novel source data-free watermark removal attack, to make the model selectively forget the embedded watermarks by customizing continual learning. In particular, AD first anchors the model's attention on the main task using some unlabeled data. Then, through continual learning, a small number of lures (randomly selected natural images) that are assigned a new label distract the model's attention away from the watermarks. Experimental results from different datasets and networks corroborate that AD can thoroughly remove the watermark with a small resource budget without compromising the model's performance on the main task, which outperforms the state-of-the-art works.
Leo Yu Zhang, Shengshan Hu, Longxiang Gao, Jun Zhang 0010, Yong Xiang 0001
ICME4
2022 A Blockchain-based Multi-layer Decentralized Framework for Robust Federated Learning
abstract
With the expansion of the Internet of Things (IoT) development and application, federated learning has gained higher popularity in industrial researching fields. However, the security issues in federated learning have become hot-spots in the research area, such as privacy-preserving and poisoning attacks. This paper proposes a robust blockchained multi-layer decentralized federated learning (RBML-DFL) framework to ensure the federated learning's robustness. Firstly, by adopting the three-layered framework, the blockchain connects the federated learning components to secure the privacy and data safety of federated learning. Secondly, the proposed framework provides resilience on poisoning attacks to the central model compared to typical federated learning frameworks. Lastly, the decentralized structure associated with the blockchain tracing back mechanism can prevent the central server failure or mal-function compared to centralized federated learning. We evaluate and compare the proposed framework with other state-of-the-art federated learning frameworks on the accuracy, latency, and system robustness under poisoning attacks. The results show that the proposed RBML-DFL framework outperforms state-of-the-art baseline frameworks on all three metrics: accuracy, latency, and the robustness of the federated learning.
Di Wu 0050, Nai Wang, Jiale Zhang 0001, Yuan Zhang 0007, Yong Xiang 0001, Longxiang Gao
IJCNN6
2022 Towards Accurate Knowledge Transfer between Transformer-based Models for Code Summarization
abstract
Automatic code summarization generates high-level natural language descriptions of code snippets, which can benefit software maintenance and code comprehension.Recently, Transformer-based models achieved state-of-the-art performance on code summarization tasks.However, there are data gaps in neural model training for some programming languages.To fill this gap, we propose a novel transfer learning approach to accurately transfer knowledge between Transformer-based models.We train a discriminator to identify which heads of the multi-head attention module should be transferred.On this basis, we define a transfer strategy of parameter matrices.We evaluated the proposed transfer learning approach on four state-of-the-art Transformer-based code summarization models.Experimental results show that models with transferred knowledge outperform original models up to 10.70% in BLEU, 5.36% in ROUGE-L, and 4.34% in METEOR.
Chaochen Shi, Yong Xiang 0001, Jiangshan Yu, Longxiang Gao
SEKE4
2022 A Bytecode-based Approach for Smart Contract Classification
abstract
With the development of blockchain technologies, the number of smart contracts deployed on blockchain platforms is growing exponentially, which makes it difficult for users to find desired services by manual screening. The automatic classification of smart contracts can provide blockchain users with keyword-based contract searching and helps to manage smart contracts effectively. Current research on smart contract classification focuses on Natural Language Processing (NLP) solutions which are based on contract source code. However, more than 94% of smart contracts are not open-source, so the application scenarios of NLP methods are very limited. Meanwhile, NLP models are vulnerable to adversarial attacks. This paper proposes a classification model based on features from contract bytecode instead of source code to solve these problems. We also use feature selection and ensemble learning to optimize the model. Our experimental studies on over 11K real-world Ethereum smart contracts show that our model can classify smart contracts without source code and has better performance than baseline models. Our model also has good resistance to adversarial attacks compared with NLP-based models. In addition, our analysis reveals that account features used in many smart contract classification models have little effect on classification and can be excluded.
Chaochen Shi, Yong Xiang 0001, Jiangshan Yu, Longxiang Gao, Keshav Sood, Robin Doss
SANER4
2022 Privacy-preserving blockchain-enabled federated learning for B5G-Driven edge computing
Yichen Wan, Youyang Qu, Longxiang Gao, Yong Xiang 0001
Comput. Networks3
2022 LSP: Lightweight Smart-Contract-Based Transaction Prioritization Scheme for Smart Healthcare
abstract
In recent years, several blockchain-based models have emerged to provide a secure way to store and access sensitive electronic medical records (EMRs) across the healthcare sector. These records are of different priorities and business requirements. From our comprehensive literature review, we observe that the existing models have no provision of prioritizing the EMR transactions. This critically affects the quick and streamline sharing of emergency EMRs in a smart healthcare environment. Furthermore, the lack of prioritization significantly restricts the optimal usage of the blockchain network. Motivated by this, we first propose a lightweight and deterministic method to prioritize the flow of emergency healthcare transactions through the smart contract. We also propose logical stateless transaction models for different entities involved in the system with varying levels of trust. Finally, the performance of the model based on the private Ethereum is verified and it outperforms the existing benchmark model in terms of usefulness in the healthcare setting and computation overhead with the use of a simple prioritizing algorithm. The obtained results demonstrate the feasibility of the proposed scheme in the real-time smart healthcare system.
Akanksha Saini, Dimaz Wijaya, Navneesh Kaur, Yong Xiang 0001, Longxiang Gao
IEEE Internet Things J.5
2022 A Lightweight and Attack-Proof Bidirectional Blockchain Paradigm for Internet of Things
abstract
Diverse technologies, such as machine learning and big data, have been driving the prosperity of the Internet of Things (IoT) and the ubiquitous proliferation of IoT devices. Consequently, it is natural that IoT becomes the driving force to meet the increasing demand for frictionless transactions. To secure transactions in IoT, blockchain is widely deployed since it can remove the necessity of a trusted central authority. However, the mainstream blockchain-based IoT payment platforms, dominated by Proof-of-Work (PoW) and Proof-of-Stake (PoS) consensus algorithms, face several major security and scalability challenges that result in system failures and financial loss. Among the three leading attacks in this scenario, double-spend attacks and long-range attacks threaten the tokens of blockchain users, while eclipse attacks target Denial of Service. To defeat these attacks, a novel bidirectional-linked blockchain (BLB) using chameleon hash functions is proposed, where bidirectional pointers are constructed between blocks. Furthermore, a new committee members auction (CMA) consensus algorithm is designed to improve the security and attack resistance of BLB while guaranteeing high scalability. In CMA, distributed blockchain nodes elect committee members through a verifiable random function. The smart contract uses Shamir’s secret-sharing scheme to distribute the trapdoor keys to committee members. To better investigate BLB’s resistance against double-spend attacks, an improved Nakamoto’s attack analysis is presented. In addition, a modified entropy metric is devised to measure eclipse attack resistance across different consensus algorithms. Extensive evaluation results show the superior resistance against attacks and demonstrate high scalability of BLB compared with current leading paradigms based on PoS and PoW.
Chenhao Xu 0003, Youyang Qu, Tom H. Luan, Peter W. Eklund, Yong Xiang 0001, Longxiang Gao
IEEE Internet Things J.6
2022 Robust Federated Averaging via Outlier Pruning
abstract
Federated Averaging (FedAvg) is the baseline Federated Learning (FL) algorithm that applies the stochastic gradient descent for local model training and the arithmetic averaging of the local models’ parameters for global model aggregation. Succeeding FL works commonly utilize the arithmetic averaging scheme of FedAvg for the aggregation. However, such arithmetic averaging is prone to the outlier model-updates, especially when the clients’ data are non-Independent and Identically Distributed (non-IID). As such, the classical aggregation approach suffers from the dominance of the outlier updates and, consequently, causes high communication costs towards producing a decent global model. In this letter, we propose a robust aggregation strategy to alleviate the above issues. In particular, we propose first pruning the node-wise outlier updates (weights) from the local trained models and then performing the aggregation on the selected effective weights-set at each node. We provide the theoretical result of our method and conduct extensive experiments on the MNIST, CIFAR-10, and Shakespeare datasets with IID and non-IID settings, which demonstrate that our aggregation approach outperforms the state-of-the-art methods in terms of communication speedup, test-set performance and training convergence.
Md Palash Uddin, Yong Xiang 0001, John Yearwood, Longxiang Gao
IEEE Signal Process. Lett.4
2022 A Covert Electricity-Theft Cyberattack Against Machine Learning-Based Detection Models
abstract
A Covert Electricity-Theft Cyberattack Against Machine Learning-Based Detection Models
Lei Cui 0006, Lei Guo 0019, Longxiang Gao, Borui Cai, Youyang Qu, Yipeng Zhou, Shui Yu 0001
IEEE Trans. Ind. Informatics3
2022 Blockchain-Based Audio Watermarking Technique for Multimedia Copyright Protection in Distribution Networks
abstract
Copyright protection in multimedia protection distribution is a challenging problem. To protect multimedia data, many watermarking methods have been proposed in the literature. However, most of them cannot be used effectively in a multimedia distribution network (MDN) as they are not designed to support multi-layer watermark embedding. Multi-layer watermarking mechanisms were developed to protect multimedia data across different layers in an MDN. However, in those mechanisms, we need to trust the entities in the MDN, such as regional and country distributors. To overcome this potential drawback, in this article, we propose a novel privacy protection mechanism for MDNs by combining the advantages of both blockchain and watermarking technologies. A specifically designed watermarking algorithm is used to link the copyright information with the audio file, while a novel blockchain-based smart contract mechanism is developed to enforce the proper functioning of each entity in the distribution network. Moreover, the new audio mechanism is computationally efficient. Although audio signals are used to show the effectiveness of the proposed mechanism, the proposed approach can easily be extended to other multimedia objects, such as an image. The validity of the proposed mechanism is demonstrated by our simulation results. The proposed mechanism can benefit multimedia production companies and other entities in the MDN.
Iynkaran Natgunanathan, Purathani Praitheeshan, Longxiang Gao, Yong Xiang 0001, Lei Pan 0002
ACM Trans. Multim. Comput. Commun. Appl.3
2022 Desynchronization-attack-resilient audio watermarking mechanism for stereo signals using the linear correlation between channels
Tianrui Zong, Juan Zhao 0007, Yong Xiang 0001, Iynkaran Natgunanathan, Longxiang Gao, Wanlei Zhou 0001
World Wide Web5
2021 Differentially Privacy-Preserving Federated Learning Using Wasserstein Generative Adversarial Network
abstract
Artificial intelligence (AI) requires a large amount of data to train high-quality machine learning (ML) models. However, due to privacy issues, individuals or organizations are not willing to share data with others, which results in “data islands”. This motivates the emergence of Federated Learning (FL), a novel ML framework allowing clients to exchange model parameters rather than the raw data. Unfortunately, the private data may be reconstructed by malicious participants by exploiting the context of model parameters in FL. This poses further challenges to privacy protection. To address this issue, we propose to integrate Wasserstein Generative Adversarial Network (WGAN) and differential privacy (DP) to protect the model parameters. WGAN is used to generate controllable random noise, which is then injected into model parameters. The new mechanism satisfies DP requirements while the data utility is highly improved. We experimentally demonstrate superior performances from aspects of convergence, accuracy, and data utility.
Yichen Wan, Youyang Qu, Longxiang Gao, Yong Xiang 0001
ISCC3
2021 BAFL: An Efficient Blockchain-Based Asynchronous Federated Learning Framework
abstract
With the widespread of 5G networks, the application of Federated Learning (FL) in Internet of Things (IoT) has become a trend. However, the trust problem caused by the centralized aggregation server, and the inefficiency problem caused by the low-performance devices, are still key challenges. Several studies involving asynchronous FL have been conducted to accelerate the training process, but they usually have a decreased model performance. In this paper, a blockchain-based asynchronous federated learning framework with a dynamic scaling factor is proposed. By adopting the blockchain, the trust problem among devices can be addressed. Meanwhile, the novel dynamic scaling factor is proposed to help improve the FL efficiency and accuracy. Extensive experiments are conducted on heterogeneous devices and the results show that the proposed framework mitigates the impact of low-performance devices while being as efficient as traditional FL with the extra benefit of alleviating the trust problem among IoT devices.
Chenhao Xu 0003, Youyang Qu, Peter W. Eklund, Yong Xiang 0001, Longxiang Gao
ISCC5
2021 A Blockchain-Based Cooperative Perception in Internet of Vehicles
abstract
In the Internet of Vehicles (IoVs), cooperative perception allows vehicles to share the sensor data with each other, so as to increase the perception range of vehicles beyond their field of view. This enables vehicles to obtain more accurate sensing information while driving and improves the safety of vehicles on the road. However, malicious nodes can send false information and poison the cooperative perception process. Therefore, how to guarantee the security of cooperative perception is crucial. In this paper, we propose a blockchain-based scheme for the post hoc electronic forensics of cooperative perception. Applying blockchain in cooperative perception is however challenging. Due to the high mobility of vehicles, the connection of vehicles to the Internet is intermittent and unpredictable. In this scenario, the security and effectiveness of blockchain can not be guaranteed as vehicles are offline.11Offline refers to that the vehicles can not connect to the Internet, but they can use vehicle-to-vehicle communication to share data. and can not update the blockchain on time. On addressing the issue, We develop an offline blockchain scheme that is composed of offline and online phases. In the offline phase, vehicles transact by sending a commitment. In the online phases before and after the offline phase, the deposit and arbitration mechanisms are proposed to defeat potential attacks in the offline phase. Lastly, the consortium blockchain is deployed to store the records of the cooperative perception process so that the records are tamper-proof and can be traced. Using extensive evaluations, we show the effectiveness of the proposed scheme.
Xinghao Li, Chenchen Tan, Minghao Liu 0012, Tom H. Luan, Longxiang Gao, Youyang Qu
VTC Fall5
2021 Digital Twin Based Remote Resource Sharing in Internet of Vehicles using Consortium Blockchain
abstract
With the evolving Internet of Vehicles (IoVs), the onboard resources of vehicles in computing and communication are experiencing fast growth. The sharing of road information and computing results among vehicles in proximity can effectively improve the utility of IoVs. However, remote inter-vehicular resource sharing, e.g., information and computing resource sharing, remains an under-explored issue. Motivated by this, we propose a novel digital twin based fair trading platform built upon consortium blockchain to enable city-wide vehicular resource sharing. Specifically, we first develop a digital twin based vehicular platform to enable vehicular resource sharing in the cloud. To track and secure the resource sharing among digital twins, the consortium blockchain is deployed, which is enforced by the designed smart contracts with an efficient Proof-of-Stake (PoS) consensus algorithm. In addition, an innovative incentive mechanism is devised to motivate the city-wide resource sharing for vehicles, which can maximize the profits of task publishers. Using extensive evaluations, we show the effectiveness of the proposed system.
Chenchen Tan, Xinghao Li, Tom H. Luan, Bruce Gu, Youyang Qu, Longxiang Gao
VTC Fall6
2021 Variational auto-encoder based Bayesian Poisson tensor factorization for sparse and imbalanced count data
Ming Liu 0028, Ruohua Xu, Lan Du 0002, Longxiang Gao, Yong Xiang 0001
Data Min. Knowl. Discov.6
2021 A fast and scalable authentication scheme in IOT for smart living
Jianhua Li 0002, Jiong Jin, Lingjuan Lyu, Dong Yuan 0001, Longxiang Gao, Chao Shen 0001
Future Gener. Comput. Syst.6
2021 Self-supervised cross-iterative clustering for unlabeled plant disease images
Uno Fang, Jianxin Li 0001, Xuequan Lu, Longxiang Gao, Mumtaz Ali 0003, Yong Xiang 0001
Neurocomputing4
2021 A Smart-Contract-Based Access Control Framework for Cloud Smart Healthcare System
abstract
In current healthcare systems, electronic medical records (EMRs) are always located in different hospitals and controlled by a centralized cloud provider. However, it leads to single point of failure as patients being the real owner lose track of their private and sensitive EMRs. Hence, this article aims to build an access control framework based on smart contract, which is built on the top of distributed ledger (blockchain), to secure the sharing of EMRs among different entities involved in the smart healthcare system. For this, we propose four forms of smart contracts for user verification, access authorization, misbehavior detection, and access revocation, respectively. In this framework, considering the block size of ledger and huge amount of patient data, the EMRs are stored in cloud after being encrypted through the cryptographic functions of elliptic curve cryptography (ECC) and Edwards-curve digital signature algorithm (EdDSA), while their corresponding hashes are packed into blockchain. The performance evaluation based on a private Ethereum system is used to verify the efficiency of proposed access control framework in the real-time smart healthcare system.
Akanksha Saini, Qingyi Zhu, Navneet Singh, Yong Xiang 0001, Longxiang Gao, Yushu Zhang 0001
IEEE Internet Things J.5
2021 DP-GAN: Differentially private consecutive data publishing using generative adversarial nets
Stella Ho, Youyang Qu, Bruce Gu, Longxiang Gao, Jianxin Li 0001, Yong Xiang 0001
J. Netw. Comput. Appl.4
2021 Segmental DCT Coefficient Reversal Based Anti-Collusion Audio Fingerprinting Mechanism
abstract
Collusion attacks are challenging to tackle in audio fingerprinting. A new direction to resist collusion attacks is to degrade the perceptual quality of the colluded files so that these files cannot be reused. The existing method in this direction has low embedding capacity and limited anti-collusion performance when the number of colluders is odd. In this letter, we present an anti-collusion mechanism that has a higher embedding capacity and can significantly degrade the perceptual quality of the colluded files regardless of the number of colluders. In the proposed mechanism, we first segment the host audio file into frames and perform the discrete cosine transform (DCT) on each frame. Then multiple fingerprint bits are embedded into each frame by reversing the DCT coefficients of the corresponding frequency band. As a result, when a collusion attack occurs, our proposed embedding mechanism can introduce perceptibly annoying differences between frames in the colluded file, which leads to severe perceptual quality degradation. Theoretical analysis and experimental results validate the superiority of the proposed anti-collusion mechanism.
Juan Zhao 0007, Tianrui Zong, Yong Xiang 0001, Longxiang Gao, Guang Hua 0001
IEEE Signal Process. Lett.4
2021 Desynchronization Attacks Resilient Watermarking Method Based on Frequency Singular Value Coefficient Modification
abstract
Desynchronization is a very challenging type of attack in audio watermarking. The traditional singular value decomposition (SVD) based audio watermarking methods embed the watermark information by modifying the singular value of individual segment, which have little resistance against desynchronization attacks. In this paper, we propose a novel frequency singular value coefficient (FSVC) feature, which reflects the ratio between the singular values of two consecutive segments and is insensitive to desynchronization attacks, to carry the watermark bits. To our best knowledge, it is the first time that the ratio between singular values is employed for audio watermarking. In the proposed method, the discrete cosine transform (DCT) is performed on two consecutive segments of the host audio signal and SVD is applied to the DCT coefficients of the mid frequency band of each segment to extract the FSVC. Then the watermark bits are embedded by adjusting the values of the FSVC. The watermark embedding procedure is optimized to minimize the perceptual quality degradation and an error buffer is created to enhance the robustness. As a result, the proposed method can achieve a much higher embedding capacity than the existing methods tackling desynchronization attacks. The impact of desynchronization and common signal processing attacks on the proposed watermarking method is mathematically modeled, and the effectiveness of the proposed method against these attacks is theoretically and experimentally validated.
Juan Zhao 0007, Tianrui Zong, Yong Xiang 0001, Longxiang Gao, Wanlei Zhou 0001, Gleb Beliakov
IEEE ACM Trans. Audio Speech Lang. Process.4
2021 Non-Linear-Echo Based Anti-Collusion Mechanism for Audio Signals
abstract
Collusion attacks are considered to be challenging attacks in audio copyright protection. The traditional watermarking algorithms cannot identify the traitors when other attacks, such as desynchronization attacks, are applied with a collusion attack. Instead of tracing the traitors, in this paper we aim to tackle collusion attacks by removing the commercial value from the colluded copy, which will demotivate the attackers from launching collusion attacks. Since the commercial value of an audio signal is directly reflected by its perceptual quality, we propose a novel non-linear-echo generation (NLEG) based algorithm to significantly degrade the perceptual quality of the colluded copy by embedding a time delay sequence into the host signal. The proposed NLEG is also designed to be resilient to common signal processing attacks and desynchronization attacks. Furthermore, the proposed NLEG can be combined with other digital watermarking techniques to enhance its performance on protecting the copyright information. Experimental results show the validity of the proposed NLEG.
Tianrui Zong, Yong Xiang 0001, Iynkaran Natgunanathan, Longxiang Gao, Guang Hua 0001, Wanlei Zhou 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2021 A Blockchained Federated Learning Framework for Cognitive Computing in Industry 4.0 Networks
abstract
Cognitive computing, a revolutionary AI concept emulating human brain's reasoning process, is progressively flourishing in the Industry 4.0 automation. With the advancement of various AI and machine learning technologies the evolution toward improved decision making as well as data-driven intelligent manufacturing has already been evident. However, several emerging issues, including the poisoning attacks, performance, and inadequate data resources, etc., have to be resolved. Recent research works studied the problem lightly, which often leads to unreliable performance, inefficiency, and privacy leakage. In this article, we developed a decentralized paradigm for big data-driven cognitive computing (D2C), using federated learning and blockchain jointly. Federated learning can solve the problem of “data island” with privacy protection and efficient processing while blockchain provides incentive mechanism, fully decentralized fashion, and robust against poisoning attacks. Using blockchain-enabled federated learning help quick convergence with advanced verifications and member selections. Extensive evaluation and assessment findings demonstrate D2C's effectiveness relative to existing leading designs and models.
Youyang Qu, Shiva Raj Pokhrel, Sahil Garg, Longxiang Gao, Yong Xiang 0001
IEEE Trans. Ind. Informatics4
2021 Mutual Information Driven Federated Learning
abstract
Federated Learning (FL) is an emerging research field that yields a global trained model from different local clients without violating data privacy. Existing FL techniques often ignore the effective distinction between local models and the aggregated global model when doing the client-side weight update, as well as the distinction of local models for the server-side aggregation. In this article, we propose a novel FL approach with resorting to mutual information (MI). Specifically, in client-side, the weight update is reformulated through minimizing the MI between local and aggregated models and employing Negative Correlation Learning (NCL) strategy. In server-side, we select top effective models for aggregation based on the MI between an individual local model and its previous aggregated model. We also theoretically prove the convergence of our algorithm. Experiments conducted on MNIST, CIFAR-10, ImageNet, and the clinical MIMIC-III datasets manifest that our method outperforms the state-of-the-art techniques in terms of both communication and testing performance.
Md Palash Uddin, Yong Xiang 0001, Xuequan Lu, John Yearwood, Longxiang Gao
IEEE Trans. Parallel Distributed Syst.5
2020 Reliable Customized Privacy-Preserving in Fog Computing
abstract
Fog computing is an emergent computing paradigm that extends the cloud paradigm to the edge. With the explosive growth of smart devices and massive data generated everyday, cloud computing no longer matches the requirements of the Internet of Things (IoT) era, such as low latency, uninterrupted service and location awareness. Thus, fog computing has been introduced as a complement of the current cloud computing model to meet the requirements in IoT. Fog computing is a relatively new networking paradigm and considered as a promising solution to support IoT scenarios. On the one hand, fog computing inherits many features from cloud; on the other hand, fog computing also inherits some challenges and issues from cloud computing: privacy issue is one of them. In this paper, we propose a personalized differential privacy model based on the distance between two fog nodes in a fog network. We also identify the collusion attack in differential privacy framework which compromised the personalized Laplace function. Based on that, we develop a personalized differential privacy model, which not only eliminate this particular attack but also optimize the trade-off between privacy preserving and data utility.
Xiaodong Wang 0017, Bruce Gu, Youyang Qu, Yongli Ren, Yong Xiang 0001, Longxiang Gao
ICC6
2020 Non-Technical Losses Detection in Smart Grids: An Ensemble Data-Driven Approach
abstract
Non technical losses (NTL) detection plays a crucial role in protecting the security of smart grids. Employing massive energy consumption data and advanced artificial intelligence (AI) techniques for NTL detection are helpful. However, there are concerns regarding the effectiveness of existing AI-based detectors against covert attack methods. In particular, the tampered metering data with normal consumption patterns may result in low detection rate. Motivated by this, we propose a hybrid data-driven detection framework. In particular, we introduce a wide & deep convolutional neural networks (CNN) model to capture the global and periodic features of consumption data. We also leverage the maximal information coefficient algorithm to analysis and detect those covert abnormal measurements. Our extensive experiments under different attack scenarios demonstrate the effectiveness of the proposed method.
Yufeng Xing, Lei Guo 0005, Zongchao Xie, Lei Cui 0006, Longxiang Gao, Shui Yu 0001
ICPADS5
2020 A Privacy Preserving Aggregation Scheme for Fog-Based Recommender System
Xiaodong Wang 0017, Bruce Gu, Youyang Qu, Yongli Ren, Yong Xiang 0001, Longxiang Gao
NSS6
2020 Protecting IP of Deep Neural Networks with Watermarking: A New Label Helps
Leo Yu Zhang, Jun Zhang 0010, Longxiang Gao, Yong Xiang 0001
PAKDD (2)4
2020 SummPip: Unsupervised Multi-Document Summarization with Sentence Graph Compression
abstract
Obtaining training data for multi-document Summarization (MDS) is time consuming and resource-intensive, so recent neural models can only be trained for limited domains. In this paper, we propose SummPip: an unsupervised method for multi-document summarization, in which we convert the original documents to a sentence graph, taking both linguistic and deep representation into account, then apply spectral clustering to obtain multiple clusters of sentences, and finally compress each cluster to generate the final summary. Experiments on Multi-News and DUC-2004 datasets show that our method is competitive to previous unsupervised methods and is even comparable to the neural supervised approaches. In addition, human evaluation shows our system produces consistent and complete summaries compared to human written ones.
Jinming Zhao, Ming Liu 0028, Longxiang Gao, Lan Du 0002, He Zhao 0001, He Zhang 0034, Gholamreza Haffari
SIGIR3
2020 Robust Blockchain-Based Cross-Platform Audio Copyright Protection System Using Content-Based Fingerprint
Juan Zhao 0007, Tianrui Zong, Yong Xiang 0001, Longxiang Gao, Gleb Beliakov
WISE (2)4
2020 Channel Correlation Based Robust Audio Watermarking Mechanism for Stereo Signals
Tianrui Zong, Yong Xiang 0001, Iynkaran Natgunanathan, Longxiang Gao, Wanlei Zhou 0001
WISE (2)4
2020 QoS-Aware Personalized Privacy With Multipath TCP for Industrial IoT: Analysis and Design
abstract
With the ensuing surge in data communication volume and the growing need for privacy protection, limiting centralized data collection to the minimum required for specific tasks has been mandatory in industries. This is now guided by the modern privacy legislation, namely, the General Data Protection Regulation and the California Consumer Protection Act. Privacy leakage has become increasingly serious because of massive volume and a variety of data transmission and Quality-of-Service (QoS) requirements in the Industrial Internet-of-Things (IIoT) networks. Although differential privacy is the core privacy protection paradigm, most of its extensions assume all parties share the same level of privacy requirements, which cannot meet varying needs and QoS of IIoT devices in practice. In addition, with multiple paths access to the cloud server (often operated by the trusted third party in IIoT) for higher reliability and performance, satisfying both the privacy and QoS is nontrivial during the data transmission. The usual transmission over both the cellular and WiFi interfaces simultaneously for continuous connectivity among devices, edge networks, and the server is crucial. As a result, we observe that IIoT data privacy is highly vulnerable to collusion attacks. Motivated by this observation, we develop a detailed QoS modeling for multipath TCP over IIoT and propose a QoS-aware personalized privacy protection model. Our model works in two different layers: one at the cloud server and another at the network edges (access points/base station). The aim is not only to balance the load but also to achieve the required QoS and optimize the tradeoff between privacy protection and efficiency. The extensive experimental results based on the real-world data sets illustrate the superiority of the proposed model in terms of privacy protection and efficiency.
Shiva Raj Pokhrel, Youyang Qu, Longxiang Gao
IEEE Internet Things J.3
2020 Decentralized Privacy Using Blockchain-Enabled Federated Learning in Fog Computing
abstract
As the extension of cloud computing and a foundation of IoT, fog computing is experiencing fast prosperity because of its potential to mitigate some troublesome issues, such as network congestion, latency, and local autonomy. However, privacy issues and the subsequent inefficiency are dragging down the performances of fog computing. The majority of existing works hardly consider a reasonable balance between them while suffering from poisoning attacks. To address the aforementioned issues, we propose a novel blockchain-enabled federated learning (FL-Block) scheme to close the gap. FL-Block allows local learning updates of end devices exchanges with a blockchain-based global learning model, which is verified by miners. Built upon this, FL-Block enables the autonomous machine learning without any centralized authority to maintain the global model and coordinates by using a Proof-of-Work consensus mechanism of the blockchain. Furthermore, we analyze the latency performance of FL-Block and further derive the optimal block generation rate by taking communication, consensus delays, and computation cost into consideration. Extensive evaluation results show the superior performances of FL-Block from the aspects of privacy protection, efficiency, and resistance to the poisoning attack.
Youyang Qu, Longxiang Gao, Tom H. Luan, Yong Xiang 0001, Shui Yu 0001, Gavin Zheng
IEEE Internet Things J.2
2020 A Fog-Based Recommender System
abstract
Fog computing is an emergent computing paradigm that extends the cloud paradigm. With the explosive growth of smart devices and mobile users, cloud computing no longer matches the requirements of the Internet of Things (IoT) era. Fog computing is a promising solution to satisfying these new requirements, such as low latency, uninterrupted service, and location awareness. As a typical new computing paradigm and network architecture, fog computing raises new challenges, such as privacy, data management, data analytics, information overload, and participatory sensing. In this article, we present a fog-based hybrid recommender system to address the issue of information overload in fog computing. Our proposed system not only abstracts useful information from the fog environment but can also be considered as an optimization tool due to its ability to provide suggestions to improve system performance. In particular, we demonstrate that the proposed system provides personalized and localized recommendations to users, and also advise the system itself to precache the content to optimize the storage capacity of the fog server.
Xiaodong Wang 0017, Bruce Gu, Yongli Ren, Shui Yu 0001, Yong Xiang 0001, Longxiang Gao
IEEE Internet Things J.7
2020 Detecting false data attacks using machine learning techniques in smart grid: A survey
Lei Cui 0006, Youyang Qu, Longxiang Gao, Gang Xie 0001, Shui Yu 0001
J. Netw. Comput. Appl.3
2019 Generative Adversarial Nets Enhanced Continual Data Release Using Differential Privacy
Stella Ho, Youyang Qu, Longxiang Gao, Jianxin Li 0001, Yong Xiang 0001
ICA3PP (2)3
2019 Context-Aware Privacy Preservation in a Hierarchical Fog Computing System
abstract
Fog computing faces various security and privacy threats. Internet of Things (IoTs) devices have limited computing, storage, and other resources. They are vulnerable to attack by adversaries. Although the existing privacy-preserving solutions in fog computing can be migrated to address some privacy issues, specific privacy challenges still exist because of the unique features of fog computing, such as the decentralized and hierarchical infrastructure, mobility, location and content-aware applications. Unfortunately, privacy-preserving issues and resources in fog computing have not been systematically identified, especially the privacy preservation in multiple fog node communication with end users. In this paper, we propose a dynamic MDP-based privacy-preserving model in zero-sum game to identify the efficiency of the privacy loss and payoff changes to preserve sensitive content in a fog computing environment. First, we develop a new dynamic model with MDP-based comprehensive algorithms. Then, extensive experimental results identify the significance of the proposed model compared with others in more effectively and feasibly solving the discussed issues.
Bruce Gu, Xiaodong Wang 0017, Youyang Qu, Jiong Jin, Yong Xiang 0001, Longxiang Gao
ICC6
2019 GAN-DP: Generative Adversarial Net Driven Differentially Privacy-Preserving Big Data Publishing
abstract
Increasing massive volume of data are generated every single second in this big data era. With big data from multiple sources, adversaries continuously mine private information for potential benefits. Motivated by this, we propose a generative adversarial net (GAN) driven noise generation method under the framework of differential privacy. We add one more perceptron, which is a specifically devised differential privacy identifier. After the generator produces the noise, the discriminator and the proposed identifier game with each other to derive the Nash Equilibrium. Extensive experimental results demonstrate the proposed model meets differential privacy constraints and upgrade data utility simultaneously.
Youyang Qu, Shui Yu 0001, Huynh Thi Thanh Binh, Longxiang Gao, Wanlei Zhou 0001
ICC5
2019 Pre-adjustment Based Anti-collusion Mechanism for Audio Signals
Juan Zhao 0007, Tianrui Zong, Yong Xiang 0001, Longxiang Gao, Gleb Beliakov
NSS4
2019 Efficient and privacy preserving access control scheme for fog-enabled IoT
Kai Fan 0001, Huiyue Xu, Longxiang Gao, Hui Li 0006, Yintang Yang
Future Gener. Comput. Syst.3
2019 Cyber security framework for Internet of Things-based Energy Internet
Abubakar Sadiq Sani, Dong Yuan 0001, Jiong Jin, Longxiang Gao, Shui Yu 0001, Zhao Yang Dong
Future Gener. Comput. Syst.4
2019 Sustainability Analysis for Fog Nodes With Renewable Energy Supplies
abstract
There is a growing interest in the use of renewable energy sources to power fog networks in order to mitigate the detrimental effects of conventional energy production. However, renewable energy sources, such as solar and wind, are by nature unstable in their availability and capacity. The dynamics of energy supply hence impose new challenges for network planning and resource management. In this paper, the sustainable performance of a fog node powered by renewable energy sources is studied. We develop a generic analytical model to study the energy sustainability of fog nodes powered by renewable energy sources, by generalizing the leaky bucket model to shape and police traffic source for rate-based congestion control in high-speed fog networks. Based on the closed-form solutions of energy buffer analysis, i.e., the energy depletion probability and mean energy length, we study the energy sustainability in two special but real-happening scenarios. The experimental results show that with proper design the leaky bucket model effectively reflects the energy sustainability of data traffic in fog networks. Numerical results also reveal that the model performance is sensitive to certain traffic source characteristics in fog networks.
Jiaojiao Jiang 0001, Longxiang Gao, Jiong Jin, Tom H. Luan, Shui Yu 0001, Yong Xiang 0001, Saurabh Kumar Garg 0001
IEEE Internet Things J.2
2019 Progressive Average-Based Smart Meter Privacy Enhancement Using Rechargeable Batteries
abstract
Usage of smart meters (SMs) have significantly increased in the recent days due to the advantages they offer. However, it is possible for an adversary to extract private information about a consumer by observing the SM readings. Therefore, it is important to protect the privacy of consumers using SMs. Among the SM privacy protection mechanisms, rechargeable battery (RB)-based mechanisms are preferred as they do not alter SM readings. The existing mechanisms cannot protect the privacy when the consumer energy usage is either low or high for a longer period. Furthermore, these mechanisms do not perform well in online scenario where the energy management unit (EMU) only knows the current and past consumer energy demands. To solve this problem, in this article, we proposed a novel online privacy protection mechanism to protect the privacy of SMs using a progressive average-based algorithm (PABA). The proposed PABA uses two uniquely designed algorithms to ensure protection of privacy and energy cost reduction during peak and off-peak periods. Moreover, an adaptive output smoothing technique is used to further enhance the privacy. Compared with the privacy protection mechanisms designed for tackling SM privacy, the proposed mechanism achieves higher amount of privacy while reducing the energy cost. The validity of the proposed privacy enhancement mechanism is demonstrated by simulation results.
Iynkaran Natgunanathan, Mohammad Belayet Hossain, Yong Xiang 0001, Longxiang Gao, Dezhong Peng, Jianxin Li 0001
IEEE Internet Things J.4
2019 Location Privacy Protection in Smart Health Care System
abstract
In a smart health system, patients' location information is periodically sent to hospitals and this information helps hospitals to provide improved health care services. The location information together with time stamp alone can reveal a patient's private information, such as person's life style, places frequently visited by the person, and personal interests. Thus, it is important to protect the location privacy of a patient. In the existing privacy protection mechanisms, trusted third party (TTP) and location perturbation techniques are used. However, in TTP-based mechanism, an adversary who illegally gets access to TTP server will have access to the private location information. On the other hand, in location perturbation technique, utility of the location information is significantly compromised. In this paper, we propose a location privacy protection mechanism in which location privacy is protected while maintaining the utility of the location data. In the proposed mechanism, a main processing unit attached to a patient's body generates the perturbed location by considering the distance between the patient's location and the preidentified patient's sensitive locations. This adaptive generation of perturbed location, removes the necessity to trust other parties while preserving the privacy and utility of the location data. The validity of the proposed mechanism is demonstrated by simulation results.
Iynkaran Natgunanathan, Abid Mehmood, Yong Xiang 0001, Longxiang Gao, Shui Yu 0001
IEEE Internet Things J.4
2018 A Trust-Grained Personalized Privacy-Preserving Scheme for Big Social Data
abstract
In the age of big data, the rapid development of social networking applications has become an improtant data source, while the massive collection of personal data leads to significant privacy concerns. Differential privacy emerged as an effective tool to get access to useful information while provide strong privacy guarantees. However, most the current proposed solutions suppose that all individuals across the network require a uniform level of privacy protection, which rules out of individuals' personalized requirements. Aiming at solving this problem, in this paper, we propose a trust-grained personalized differential privacy mechanism, called TGDP, by combining the notion of trust. Specifically, whenever a user wants to get another user's personal information, the proposed mechanism returns a corresponding private response in which the privacy level selected for each individual depend on the trust value between them in the network. Compared with traditional methods, the scheme can provide a fine-grained differential privacy protection method, while guarantee the utility of social networks. Finally, the scheme is evaluated analytically, and demonstrated experimentally on the real- world data, which reflects its effectiveness and utility.
Lei Cui 0006, Youyang Qu, Shui Yu 0001, Longxiang Gao, Gang Xie 0001
ICC4
2018 A Hybrid Privacy Protection Scheme in Cyber-Physical Social Networks
abstract
The rapid proliferation of smart mobile devices has significantly enhanced the popularization of the cyber-physical social network, where users actively publish data with sensitive information. Adversaries can easily obtain these data and launch continuous attacks to breach privacy. However, existing works only focus on either location privacy or identity privacy with a static adversary. This results in privacy leakage and possible further damage. Motivated by this, we propose a hybrid privacy-preserving scheme, which considers both location and identity privacy against a dynamic adversary. We study the privacy protection problem as the tradeoff between the users aiming at maximizing data utility with high-level privacy protection while adversaries possessing the opposite goal. We first establish a game-based Markov decision process model, in which the user and the adversary are regarded as two players in a dynamic multistage zero-sum game. To acquire the best strategy for users, we employ a modified state-action-reward-state-action reinforcement learning algorithm. Iteration times decrease because of cardinality reduction from n to 2, which accelerates the convergence process. Our extensive experiments on real-world data sets demonstrate the efficiency and feasibility of the propose method.
Youyang Qu, Shui Yu 0001, Longxiang Gao, Wanlei Zhou 0001, Sancheng Peng
IEEE Trans. Comput. Soc. Syst.3
2018 Malware Propagations in Wireless Ad Hoc Networks
abstract
Accurate malware propagation modeling in wireless ad hoc networks (WANETs) represents a fundamental and open research issue which shows distinguished challenges due to complicated access competition, severe channel interference, and dynamic connectivity. As an effort towards the issue, in this paper, we investigate the malware propagation under two spread schemes including Unicast and Broadcast, in Spread Mode and Communication Mode, respectively. We highlight our contributions in three-fold in the light of previous literature works. First, a bound of malware infection rate for each scheme is provided by applying the wireless network capacity theories. Second, the impact of mobility on malware propagations has been studied. Third, discussion of the relationship between different schemes and practical applications is provided. Numerical simulations and detailed performance analysis show that the Broadcast Scheme with Spread Mode is most dangerous in the sense of malware propagation speed in WANETs, and mobility will greatly increase the risk further. The results achieved in this paper not only provide insights on the malware propagation characteristics in WANETs, but also serve as fundamental guidelines on designing defense schemes.
Bo Liu 0001, Wanlei Zhou 0001, Longxiang Gao, Tom H. Luan, Sheng Wen
IEEE Trans. Dependable Secur. Comput.3
2017 Big data set privacy preserving through sensitive attribute-based grouping
abstract
There is a growing trend towards attacks on database privacy due to great value of privacy information stored in big data set. Public's privacy are under threats as adversaries are continuously cracking their popular targets such as bank accounts. We find a fact that existing models such as K-anonymity, group records based on quasi-identifiers, which harms the data utility a lot. Motivated by this, we propose a sensitive attribute-based privacy model. Our model is the early work of grouping records based on sensitive attributes instead of quasi-identifiers which is popular in existing models. Random shuffle is used to maximize information entropy inside a group while the marginal distribution maintains the same before and after shuffling, therefore, our method maintains a better data utility than existing models. We have conducted extensive experiments which confirm that our model can achieve a satisfying privacy level without sacrificing data utility while guarantee a higher efficiency.
Youyang Qu, Shui Yu 0001, Longxiang Gao, Jianwei Niu 0002
ICC3
2017 Towards an Analysis of Traffic Shaping and Policing in Fog Networks Using Stochastic Fluid Models
abstract
This paper gives models and analytic techniques for studying shaping and policing data traffic in fog networks. The traffic in these networks is expected to be highly diverse and bursty, and regulation will be required as an integral part of congestion control. We generalize the Leaky Bucket model to shape and police traffic source for rate-based congestion control in high-speed fog networks. In particular, the Markov modulated fluid sources reflect the bursty characteristics of data traffic. To measure the performance of the model in shaping and policing traffic, we derive four performance metrics. The experimental results show that with proper design the Leaky Bucket model effectively controls a 4-way trade-off between throughput, loss probability, delay and burstiness of data traffic. Numerical results also reveal that the model performance is sensitive to certain traffic source characteristics.
Jiaojiao Jiang 0001, Longxiang Gao, Jiong Jin, Tom H. Luan, Shui Yu 0001, Dong Yuan 0001, Yong Xiang 0001, Dongfeng Yuan
MobiQuitous2
2017 FogRoute: DTN-Based Data Dissemination Model in Fog Computing
abstract
Fog computing, known as “cloud closed to ground,” deploys light-weight compute facility, called Fog servers, at the proximity of mobile users. By precatching contents in the Fog servers, an important application of Fog computing is to provide high-quality low-cost data distributions to proximity mobile users, e.g., video/live streaming and ads dissemination, using the single-hop low-latency wireless links. A Fog computing system is of a three tier Mobile–Fog–Cloud structure; mobile user gets service from Fog servers using local wireless connections, and Fog servers update their contents from Cloud using the cellular or wired networks. This, however, may incur high content update cost when the bandwidth between the Fog and Cloud servers is expensive, e.g., using the cellular network, and is therefore inefficient for nonurgent, high volume contents. How to economically utilize the Fog–Cloud bandwidth with guaranteed download performance of users thus represents a fundamental issue in Fog computing. In this paper, we address the issue by proposing a hybrid data dissemination framework which applies software-defined network and delay-tolerable network (DTN) approaches in Fog computing. Specifically, we decompose the Fog computing network with two planes, where the cloud is a control plane to process content update queries and organize data flows, and the geometrically distributed Fog servers form a data plane to disseminate data among Fog servers with a DTN technique. Using extensive simulations, we show that the proposed framework is efficient in terms of data-dissemination success ratio and content convergence time among Fog servers.
Longxiang Gao, Tom H. Luan, Shui Yu 0001, Wanlei Zhou 0001, Bo Liu 0001
IEEE Internet Things J.1
2016 Complex network theoretical analysis on information dissemination over vehicular networks
abstract
How to enhance the communication efficiency and quality on vehicular networks is one critical important issue. While with the larger and larger scale of vehicular networks in dense cities, the real-world datasets show that the vehicular networks essentially belong to the complex network model. Meanwhile, the extensive research on complex networks has shown that the complex network theory can both provide an accurate network illustration model and further make great contributions to the network design, optimization and management. In this paper, we start with analyzing characteristics of a taxi GPS dataset and then establishing the vehicular-to-infrastructure, vehicle-to-vehicle and the hybrid communication model, respectively. Moreover, we propose a clustering algorithm for station selection, a traffic allocation optimization model and an information source selection model based on the communication performances and complex network theory.
Jingjing Wang 0001, Chunxiao Jiang, Longxiang Gao, Shui Yu 0001, Zhu Han 0001, Yong Ren 0001
ICC3
2016 A scalable and automatic mechanism for resource allocation in self-organizing cloud
Xiaotong Wu, Meng Liu 0010, Wan-Chun Dou, Longxiang Gao, Shui Yu 0001
Peer-to-Peer Netw. Appl.4
2013 Multidimensional Routing Protocol in Human-Associated Delay-Tolerant Networks
abstract
Human-associated delay-tolerant networks (HDTNs) are new networks where mobile devices are associated with humans and can be viewed from multiple dimensions including geographic and social aspects. The combination of these different dimensions enables us to comprehend delay-tolerant networks and consequently use this multidimensional information to improve overall network efficiency. Alongside the geographic dimension of the network, which is concerned with geographic topology of routing, social dimensions such as social characters can be used to guide the routing message to improve not only the routing efficiency for individual nodes, but also efficiency for the entire network. We propose a multidimensional routing protocol (M-Dimension) for the human-associated delay-tolerant networks which uses local information derived from multiple dimensions to identify a mobile node more accurately. The importance of each dimension has been measured by the weight function and it is used to calculate the best route. The greedy routing strategy is applied to select an intermediary node to forward message. We compare M-Dimension to the existing benchmark routing protocols via MIT reality Data Set and INFOCOM 2006 Data Set, which are real human-associated mobile network trace files. The results of our simulations show that M-Dimension significantly increases the average success ratio with a competitive end-to-end delay when compared with other multicast DTNs routing protocols.
Longxiang Gao, Ming Li 0010, Alessio Bonti, Wanlei Zhou 0001, Shui Yu 0001
IEEE Trans. Mob. Comput.1
2012 Effects of Social Characters in Viral Propagation Seeding Strategies in Online Social Networks
abstract
Online social networks have not only become a point of aggregation and exchange of information, they have so radically rooted into our everyday behaviors that they have become the target of important network attacks. We have seen an increasing trend in Sybil based activity, such as in personification, fake profiling and attempts to maliciously subvert the community stability in order to illegally create benefits for some individuals, such as online voting, and also from more classic informatics assaults using specifically mutated worms. Not only these attacks, in the latest months, we have seen an increase in spam activities on social networks such as Facebook and RenRen, and most importantly, the first attempts at propagating worms within these communities. What differentiates these attacks from normal network attacks, is that compared to anonymous and stealthy activities, or by commonly untrusted emails, social networks regain the ability to propagate within consentient users, who willingly accept to partake. In this paper, we will demonstrate the effects of influential nodes against non-influential nodes through in simulated scenarios and provide an overview and analysis of the outcomes.
Alessio Bonti, Ming Li 0010, Longxiang Gao
TrustCom3
2012 AMDD: Exploring Entropy Based Anonymous Multi-dimensional Data Detection for Network Optimization in Human Associated DTNs
abstract
Human associated delay-tolerant networks (HDTNs) are new networks where mobile devices are associated with humans and demonstrate social-related communication characteristics. Most of recent works use real social trace file to analyse its social characteristics, however social-related data is sensitive and has concern of privacy issues. In this paper, we propose an anonymous method that anonymize the original data by coding to preserve individual's privacy. The Shannon entropy is applied to the anonymous data to keep rich useful social characteristics for network optimization, e.g. routing optimization. We use an existing MIT reality dataset and Infocom 06 dataset, which are human associated mobile network trace files, to simulate our method. The results of our simulations show that this method can make data anonymously while achieving network optimization.
Longxiang Gao, Ming Li 0010, Tianqing Zhu, Alessio Bonti, Wanlei Zhou 0001, Shui Yu 0001
TrustCom1
2012 A minimum disclosure approach to authentication and privacy in RFID systems
Robin Doss, Wanlei Zhou 0001, Saravanan Sundaresan, Shui Yu 0001, Longxiang Gao
Comput. Networks5
2012 M-Dimension: Multi-characteristics based routing protocol in human associated delay-tolerant networks with improved performance over one dimensional classic models
Longxiang Gao, Ming Li 0010, Alessio Bonti, Wanlei Zhou 0001, Shui Yu 0001
J. Netw. Comput. Appl.1
2011 Multi-level virtual ring: An architecture for content routing in wireless sensor network
abstract
Two main problems prevent the deployment of content delivery in a wireless sensor network: the address, which is widely used in the Internet as the identifier, is meaningless in wireless network, and the routing efficiency is a big concern in wireless sensor network. This paper presents an embedded multi-level ring (MVR) structure to address those two problems. The MVR uses names rather than addresses to identify sensor nodes. The MVR routes packets on the name identifiers without being aware the location. Some sensor nodes are selected as the backbone nodes and are placed on the different levels of the virtual rings. MVR hashes nodes and contents identifiers, and stores them at the backbone nodes. MVR takes the cross-level routing to improve the routing efficiency. Further, MVR is constructed decentralized and runs on the mobile nodes themselves, requiring no central control. Experiments using ns2 simulator for up to 200 nodes show that the storage and bandwidth requirements of MVR grow slowly with the size of the network. Furthermore, MVR has demonstrated as self-administrating, fault-tolerant, and resilient under the different workloads. We also discuss alternative implementation options, and future work.
Ming Li 0010, Longxiang Gao
APCC2
2010 S-Kcore: A Social-aware Kcore Decomposition Algorithm in Pocket Switched Networks
abstract
The key nodes in network play the critical role in system recovery and survival. Many traditional key nodes selection algorithms utilize the characters of the physical topology to find the key nodes. But they can hardly succeed in the mobile ad hoc network due to the mobility nature of the network. In this paper we propose a social-aware Kcore selection algorithm to work in the Pocket Switched Network. The social view of the network suggests the social position of the mobile nodes can help to find the key nodes in the Pocket Switched Network. The S-Kcore selection algorithm is designed to exploit the nodes' social features to improve the performance in data communication. Experiments use the NS2 shows S-Kcore selection algorithm workable in the Pocket Switched Network. Furthermore, with the social behavior information, those key nodes are more suitable to represent and improve the whole network's performance.
Ming Li 0010, Longxiang Gao, Wanlei Zhou 0001
EUC2