VLDB 2026 Research / reviewers in the wild / expert
Yao Zhao 0006
dblp:45/2091-6
· DBLP profile ↗
21ranked-venue papers
13as first author
20since 2021 · last 2026
0000-0002-5870-7370ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 6 · 4 first-author · 6 since 2021Software engineering, systems software and programming languages · 6 · 4 first-author · 6 since 2021Security and privacy · 3 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fast Convergent Federated Learning via Decaying SGD Updates
Md Palash Uddin, Yong Xiang 0001, Mahmudul Hasan 0018, Yao Zhao 0006, Youyang Qu, Longxiang Gao |
IEEE Trans. Big Data | 4 |
| 2026 | Collusion-Resistant and Time-Aware Co-Verification for Edge Data IntegrityabstractMobileEdgeComputing (MEC) has incentivized App vendors to outsource various services and applications to distributed edge nodes for low access latency. However, the data cached on these nodes is vulnerable to both intentional and accidental corruption, necessitating periodic audits ofEdgeDataIntegrity (EDI). Existing solutions either rely on a “fully trustworthy”ThirdPartyAuditor (TPA) or leverage blockchain to enhance trust. However, they overlook the security risks brought by the use of blockchain, particularly collusion attacks. Furthermore, while they employ achallenge-responsemechanism to enhance efficiency by batch verification, they fail to account for the heterogeneity of edge nodes. To address these challenges, we propose$\mathtt {CTCV}$, aCollusion-resistant andTime-awareCollaborativeVerification framework.$\mathtt {CTCV}$aims to accommodate edge node heterogeneity while enabling public audits and batch verification without introducing additional security risks. Specifically, it incorporates blockchain to allow edge nodes to collaboratively verify EDI without trust dependencies, while mitigating collusion attacks through a carefully designed proof generation and verification approach. Considering the resource and state heterogeneity of edge nodes,$\mathtt {CTCV}$employs atime-constrained challenge-responsemechanism that sets a time threshold$\mathcal {T}$between the verification request issuance and the integrity proof inspection to avoid excessive delays. The selection guideline of$\mathcal {T}$, along with the correctness, efficiency, and collusion resistance of$\mathtt {CTCV}$, are rigorously analyzed. Extensive experiments validate that$\mathtt {CTCV}$is computationally and communicationally efficient compared to three baselines: EdgeWatch, EDI-S, and EDI-V. On average, given 10 edge nodes,$\mathtt {CTCV}$outperforms EdgeWatch, EDI-S, and EDI-V with computation efficiency improvements of 7.9, 9.0, and 5.0 times, and communication efficiency improvement of 2063.0, 4.8, and 2.6 times, respectively. Yao Zhao 0006, Youyang Qu, Bo Li 0103, Lu Zhao 0001, Feifei Chen 0001, Yong Xiang 0001, Longxiang Gao |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2026 | Split Learning With Local Epoch Regulation and Time-Aware DetectionabstractFederated learning (FL) has become a popular approach in Edge AI for extracting valuable knowledge within edge computing (EC) systems. To enhance AI application performance, large-scale models have gained increasing attention due to their strong generalization capabilities. However, training and transmitting such models impose substantial computational and communication overhead on resource-constrained clients at the edge, and exchanging complete models may also compromise model privacy. To alleviate these burdens and safeguard privacy, split learning (SL) has been introduced by combining data and model parallelism. Although SL alleviates resource constraints, it still encounters efficiency and security challenges in EC environments, where heterogeneous clients can slow down training without enhancing accuracy, and malicious clients may manipulate model behavior. To address these challenges, we propose a novel SL framework, CoDefend, which integrates local epoch regulation and time-aware detection. Specifically, local epoch regulation dynamically assigns heterogeneous clients with appropriate local epoch numbers to improve training efficiency, while time-aware detection provides an effective detection window to identify clients' malicious manipulation to improve model security. Moreover, CoDefend jointly optimizes these two strategies by leveraging their interdependence to further improve SL performance. Extensive experiments on both simulated and real-world platforms using NVIDIA Jetson edge nodes demonstrate that CoDefend achieves approximately 2× faster training speed than baseline methods, while maintaining comparable model accuracy and effectively identifying malicious manipulations even under collusion. Yao Zhao 0006, Zahir Tari, Nasrin Sohrabi, Qin Wang 0008, Xiaoyu Xia 0001 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2025 | AirDIV: Over-the-Air Cloud-Fog Data Integrity Verification Scheme for Industrial Cyber-Physical SystemsabstractIndustrial Cyber-Physical Systems (ICPSs) have been motivating various Industry 4.0 endeavours, particularly with the integration of fog computing. Cloud-fog data caching paradigms, as supportive elements of ICPSs, have been adopted to cache user data, catering to diverse ICPS requirements such as data sensitivity and reduced access latency. In this hierarchical caching context, ensuring Cloud-Fog Data Integrity (CFDI) is crucial for maintaining the consistent functionality of ICPSs. Existing solutions primarily focus on examining the integrity of data cached solely on either cloud or fog nodes. However, cloud-cached data and fog-cached data are tightly coupled and should be considered simultaneously when checking data integrity. In this work, we introduce an over-the-air CFDI verification scheme, namely AirDIV, with a high accuracy and security guarantee. Instead of aggregating integrity proofs after proof transmission, AirDIV completes proof aggregation and transmission over the air for efficiency improvement. To enhance practicability, we derive adjustable parameters and formulate an optimization problem to minimize over-the-air aggregation errors. Furthermore, with an effective proof generation method, AirDIV can defend against two common attacks, i.e., replay and forge attacks. We provide a theoretical analysis of AirDIV’s correctness, accuracy and security, while conducting extensive experiments on both simulated and real platforms to validate its efficiency. Yao Zhao 0006, Yong Xiang 0001, Md Palash Uddin, Yushu Zhang 0001, Lu Liu 0001, Longxiang Gao |
IEEE J. Sel. Areas Commun. | 1 |
| 2025 | A federated compositional knowledge graph embedding for communication efficiencyabstractKnowledge Graph Embedding (KGE), which automatically capture structural information from Knowledge Graphs (KGs), are essential for enhancing various downstream tasks, such as recommender systems. To further improve the effectiveness of KGE models, Federated Knowledge Graph Embedding (FKGE) has been introduced, enabling the privacy-preserving integration of KGs across multiple organizations. However, existing FKGE frameworks require aggregation of a large global KGE model (embeddings). resulting in significant communication overhead, thereby reducing the efficiency and utility of FKGE in practical scenarios. To address this challenge, we propose Federated Compositional Knowledge Graph Embedding (FedComp), which enhances communication efficiency by leveraging the compositional characteristics of KG entities. In FedComp, we design a lightweight global model that represents shareable latent features of entities. These global latent features are composed into personalized KGE models with local embedding generators on the clients, improving both local adaptability and performance. By this, FedComp can significantly reduce the number of parameters that need to be transmitted Experimental results show that FedComp outperforms state-of-the-art FKGE frameworks on link prediction accuracy, with only around 1.0% communication overhead compared to counterpart frameworks. Borui Cai, Yong Xiang 0001, Yao Zhao 0006, Md Palash Uddin, Keshav Sood |
Knowl. Based Syst. | 4 |
| 2025 | Intelligent Edge Data Integrity Verification With Dynamic Unreliable Data Replica SelectionabstractWith the advancement of Mobile Edge Computing (MEC), App vendors are increasingly motivated to cache multiple data replicas on geographically distributed edge servers to ensure rapid responses for latency-sensitive applications. However, the security of data replicas is a critical concern due to the dynamic nature and resource limitations of MEC environments. To this end, data replicas’ integrity must be regularly verified to maintain the accuracy of data-driven decision-making. Existing Edge Data Integrity (EDI) verification solutions suffer from low efficiency due to relying on indiscriminative verification, where all data replicas are checked at each round without considering their inherent reliability characteristics. This paper designs an Intelligent framework called I-EDI, which enables discriminative EDI verification by integrating a novel Long-term Unreliable data Replica Selection (L-URS) mechanism. This framework aims to reduce verification costs without compromising accuracy, while resisting spoofing, forgery, outsourcing, collusion, alteration-before-verification, delayed-response, and adaptive attacks. Specifically, each data replica is associated with a reliability representation by evaluating its long-term performance. Based on that, the L-URS problem is defined as stochastically minimizing the global reliability representation over time, subject to constraints on the number of data replicas to be verified. To make it easy-to-handle, the L-URS problem is decomposed into a series of online minimization problems. An Online Opportunistic-based Replica Selection approach called O2RS is developed. O2RS allows App vendors to significantly decrease verification costs by targetedly inspecting unreliable data replicas. Moreover, this work provides a thorough theoretical analysis of O2RS’s time complexity and approximation bound, as well as I-EDI’s security. Extensive experiments are conducted to validate the effectiveness and efficiency of O2RS and I-EDI. The results demonstrate that, compared to commonly used alternatives, O2RS achieves an approximate 50% improvement in selection efficiency, while I-EDI reduces verification costs by 1.23 times on average. Yao Zhao 0006, Youyang Qu, Nasrin Sohrabi, Md. Redowan Mahmud, Zahir Tari |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2025 | Multiple Edge Data Integrity Verification With Multi-Vendors and Multi-Servers in Mobile Edge ComputingabstractEnsuring Edge Data Integrity (EDI) is imperative in providing reliable and low-latency services in mobile edge computing. Existing EDI schemes typically address single-vendor (App Vendor, AV) single-server (Edge Server, ES), single-vendor multi-server, and multi-vendor multi-server scenarios, which consider a single data replica cached by an ES from the AVs. However, the most practical scenario of Multi-Vendors and Multi-Servers with Multiple Data (MVMS-MD) cached by an ES from different AVs remains unexplored. Current solutions struggle when applied to this scenario due to increased computation and communication costs in the verification process across all ESs using the classicalchallenge-response per-data multi-roundstrategy. To tackle this issue, we propose a Multiple EDI-Verification (MEDI-V) approach in this paper. In particular, our MEDI-V utilizes an adaptive Merkle Hash Tree (ad-MHT) to efficiently generate a tree of multiple data replicas within each AV. Next, the dynamic mechanism computes minimal verification information using ad-MHT to create achallengefor individual ESs to produce EDI proofs. The ES then leverages its ad-MHT and the ES's proof to send the reconstructed ad-MHT root to the AV for verification. Theoretical insights into MEDI-V's correctness, efficiency, security, and comprehensive evaluations demonstrate its superiority in addressing MEDI issues in the MVMS-MD scenario. Yong Xiang 0001, Md Palash Uddin, Yao Zhao 0006, Jonathan Kua, Longxiang Gao |
IEEE Trans. Mob. Comput. | 4 |
| 2025 | Data Re-Outsourcing Detection With Latency-Constraint for Edge StorageabstractEdge storage has become a widely used solution for providing low-latency data access services, which motivates data owners to outsource data on geographically distributed edge nodes to deliver a positive user experience. Nevertheless, various security concerns raise in terms of data availability. Among them, edge data geo-location verification becomes a prominent concern when the data is out of owners' control, since outsourced data may be re-outsourced to other economical yet unknown third-party devices by dishonest edge nodes for saving storage space and pocketing the difference. Existing geo-localization approaches for cloud architectures can not be practically applied to identify such re-outsourcing behaviors due to the uniqueness of edge storage. To close this gap, we make the first attempt to investigate theedgedatare-outsourcingdetection (EDRD) problem, enabling the data owner to inspect if outsourced data is consistently cached on the rented edge nodes with agreed geo-location. We leverage timedChallenge-Responsemechanisms for data possession proof while measuring verification latency to detect re-outsourcing behaviors by comparing with re-outsourcing detection threshold$\mathbb {C}$. We prove that the edge node whose verification latency exceeds$\mathbb {C}$is dishonest. To obtain the optimal$\mathbb {C}$, we formulate thethresholddetermination (TD) problem and transform it to an easy-to-handle form for problem complexity reduction. Then, apreference-based approach named TD-P is developed to efficiently address the transformed TD problem. On top of that, we propose a$\mathbb {C}$-aware edge data re-outsourcing detection scheme entitled EDRD-$\mathbb {C}$to tackle the EDRD problem effectively. The efficiency and effectiveness of TD-P and EDRD-$\mathbb {C}$are verified by extensive theoretical analysis and experimental evaluations on both simulated and real platforms. Notably, EDRD-$\mathbb {C}$achieves 100% detection accuracy by sacrificing a reasonable amount of computing resources and 92.97% in the worst case. Yao Zhao 0006, Youyang Qu, Yong Xiang 0001, Feifei Chen 0001, Longxiang Gao |
IEEE Trans. Serv. Comput. | 1 |
| 2024 | From Data Integrity to Global Model Integrity for Federated Learning: An MHT-based ApproachabstractFederated Learning (FL) is a distributed machine learning (ML) approach that enables multiple edge nodes to collaboratively train ML models by sharing model parameters, thus addressing privacy concerns. However, in highly distributed, dynamic, and volatile FL environments, the global model is vulnerable to various corruptions. For instance, edge nodes might falsely claim that the received global model is incomplete, or the channel that transmits the global model is untrustworthy. Effectively verifying the integrity of the global model poses a critical challenge. To tackle this issue, we introduce a method for verifying model integrity called Federated learning global Model Integrity Verification (FMIV). It leverages Merkle Hash Tree (MHT) to generate integrity proofs of the global model during verification. To improve security, we integrate random security codes during proof generation. FMIV is capable of verifying the global model updated by the central server and shared with untrusted edge nodes, while efficiently identifying the edge node caching the incomplete global model. Furthermore, we conduct theoretical analysis and extensive experiments to validate the performance of FMIV. Compared to the two state-of-the-art approaches, FMIV consistently exhibits a notable improvement in verification efficiency and effectiveness in detecting model corruption. Yao Zhao 0006, Y. Neil Qu, Bruce Gu, Keshav Sood, Longxiang Gao, Shui Yu 0001 |
GLOBECOM | 2 |
| 2024 | From Data Integrity to Global ModeI Integrity for Decentralized Federated Learning: A Blockchain-based ApproachabstractDecentralized Federated Learning (DFL) is extensively applied in various areas, e.g., healthcare, finance, and Internet of Things (loT), offering practical solutions for distributed intelligent applications and data collaboration. In DFL systems, participants, e.g., edge devices, organizations, or nodes, collaborate in the training of a shared global model by aggregating local models from various participants. During this process, participants need to communicate frequently with a central authority/node/server to share model parameters. Such communication is vulnerable to malicious attacks or tampering, posing a significant threat to the integrity of model training. The integrity verification method can provide an integrity guarantee for the global model of DFL. However, most of the existing integrity verification schemes are centralized and not suitable for resource-constrained DFL scenarios. Therefore, how to verify the integrity of the global model becomes an important issue in DFL. To address it, we devise a global model integrity verification method for DFL. Specifically, we generate a digital signature for each global model parameter as proof of integrity, while improving the efficiency of integrity verification by electing delegates to conduct the verification process. A series of experiments is conducted to validate the performance of the proposed method. The experimental results demonstrate that our approach not only effectively ensures the integrity of the global model but also functions well under limited resources. Yao Zhao 0006, Youyang Qu, Lei Cui 0006, Longxiang Gao |
IJCNN | 2 |
| 2024 | Hierarchical Service Composition via Blockchain-enabled Federated LearningabstractAbstract In recent years, the transformative evolution of cloud computing has reshaped organizational practices by enabling the outsourcing of web service applications. This shift has led to the emergence of the cloud environment, characterized by the involvement of Cloud Service Providers (CSPs) and intelligent applications. Cloud Service Composition (CSC) has become pivotal in this context, playing a crucial role in enhancing efficiency, Quality of Service (QoS), and customer satisfaction through the aggregation of diverse Cloud Services (CSs) to create composite services. However, the vast array of available CSs presents a challenge in efficiently addressing specified QoS requirements, turning CSC into a recognized NP-hard problem. Existing solutions, often involving third-party brokers, struggle with scalability in large-scale systems and overlook crucial security concerns. To address these limitations, we propose the Hierarchical Service Composition (HSC) approach, leveraging blockchain and federated learning to minimize computational complexity. The integration of Blockchain-enabled Federated Learning (BFL) facilitates machine learning model training with decentralized data, ensuring practicality and fairness. HSC comprises an initialization phase and two selection layers. The first selection layer enables each CSP to efficiently select services using a pre-trained model, while the second selection layer employs a blockchain-based QoS-aware mechanism for the final composition result, addressing privacy concerns. HSC introduces a novel framework, collaborative service selection methods, and a smart selection algorithm, demonstrating remarkable composition efficiency in extensive simulations compared to the baseline approach. Lu Zhao 0001, Yao Zhao 0006 |
Data Sci. Eng. | 4 |
| 2024 | A Learning-Based Hierarchical Edge Data Corruption Detection Framework in Edge IntelligenceabstractEdge intelligence, an emerging distributed paradigm, is driven by the increasing number of Internet of Things devices and the development of edge computing and artificial intelligence. This paradigm revolutionizes the way of data caching by encouraging latency-sensitive data to be distributed across multiple edge nodes. In such data caching scenarios, ensuring the integrity of data stored at the edge nodes is critical for business continuity guarantee. Existing Edge Data Integrity (EDI) verification solutions rely on the interactive Challenge-Response mechanism. However, this mechanism imposes significant communication overhead on participants, leading to low verification efficiency. To address this challenge, we propose a Learning-based Hierarchical Edge Data Corruption Detection framework (LH-EDCD), aiming to enhance verification efficiency from a round perspective by reducing communication interaction between edge nodes and the data owner. LH-EDCD involves two layers of verification: internal and external. In the internal verification layer, each edge node self-inspects the cached data replica by running a corruption detection model distributedly trained by blockchain-based Federated Learning (FL). With such filtration, potential corruption can be efficiently identified without complex interaction. Considering the false positive existence in the model, in the external verification layer, LH-EDCD adopts a smart contract in blockchain to verify identified potentially corrupted data replicas for corruption confirmation, mitigating the trust concerns among edge nodes while reducing communication overhead on backbone networks. With the combination of these two layers, the overall EDI verification efficiency can be improved by reducing interaction verification time. Additionally, we make the first attempt to investigate the optimal verification time to improve the applicability and practicality of LH-EDCD. Extensive experimental results substantiate the advantages of employing FL in the first layer of LH-EDCD and demonstrate that LH-EDCD outperforms two state-of-the-art EDI approaches, i.e., EDI-S and EDI-V. Specifically, LH-EDCD achieves better model accuracy and convergence speed compared to centralized training, while exhibiting superior efficiency over EDI-S and EDI-V with 3.5 and 2.8 times performance improvements, respectively. Yao Zhao 0006, Chenhao Xu 0003, Youyang Qu, Yong Xiang 0001, Feifei Chen 0001, Longxiang Gao |
IEEE Internet Things J. | 1 |
| 2024 | Context-Aware Consensus Algorithm for Blockchain-Empowered Federated LearningabstractSupported by cloud computing,FederatedLearning (FL) has experienced rapid advancement, as a promising technique to motivate clients to collaboratively train models without sharing local data. To improve the security and fairness of FL implementation, numerousBlockchain-empoweredFederatedLearning (BFL) frameworks have emerged accordingly. Among them, consensus algorithms play a pivotal role in determining the scalability, security, and consistency of BFL systems. Existing consensus solutions to block producer selection and reward allocation either focus on well-resourced scenarios or accommodate BFL based on clients' contributions to model training. However, these approaches limit consensus efficiency and undermine reward fairness, due to involving intricate consensus processes, disregarding clients' contributions during blockchain consensus, and failing to address lazy client problems (malicious clients plagiarizing local model updates from others to reap rewards). Given the aforementioned challenges, we make the first attempt to design a joint solution for efficient consensus and fair reward allocation in heterogeneous BFL systems with lazy clients. Specifically, we introduce a generalizable BFL workflow that can address lazy client problems well. Based on it, the global contribution of BFL clients is decoupled into five dominant metrics, and the block producer selection problem is formulated as a reward-constraint contribution maximization problem. By addressing this problem, the optimal block producer that maximizes global contribution can be identified to orchestrate consensus processes, and rewards are distributed to clients in proportion to their respective global contributions. To achieve it, we develop aContext-awareProof-of-Contribution consensus algorithm named CPoC to reach consensus and incentive simultaneously, followed by theoretical analysis of lazy client problems and privacy issues. Empirical results on widely-used datasets demonstrate the effectiveness of our design in improving consensus efficiency and maximizing global contribution. Yao Zhao 0006, Youyang Qu, Yong Xiang 0001, Feifei Chen 0001, Longxiang Gao |
IEEE Trans. Cloud Comput. | 1 |
| 2024 | Data Integrity Verification in Mobile Edge Computing With Multi-Vendor and Multi-ServerabstractThe emergingMobileEdgeComputing (MEC) paradigm reforms the way of data caching by motivating App vendors to store latency-sensitive data on distributed edge servers. In volatile MEC environments, ensuringEdgeDataIntegrity (EDI) is a major concern for App vendors. Existing EDI solutions only consider the scenario with a single App vendor and multiple edge servers, neglecting more complex multi-vendor and multi-server cases. If multiple App vendors check their data replicas cached on the same edge server simultaneously, integrity verification efficiency will drop exponentially. To mitigate this challenge, we make the first attempt to develop aSmartInspectionAlgorithm (SIA) to pre-select unreliable data replicas for different App vendors in each verification round by jointly considering cache services' QoS (Quality-of-Service) and data replicas' unverified time. By implementing this approach, edge servers can merely verify the selected data replicas, greatly reducing computation and communication overheads in EDI verification. Theoretically, SIA can achieve$\mathcal {O}(n)$expected time complexity. Supported by SIA, we expand the EDI problem in multi-vendor and multi-server MEC environments (referred to as the MVMS-EDI problem) and propose a smart contract-based approach entitled MVMS-SC to tackle the problem efficiently and impartially. We provide a rigorous theoretical analysis of the correctness, security, and efficiency of MVMS-SC. Both large-scale and small-scale experiments with real-world datasets are correspondingly performed on a single machine and a real platform to validate the superiority of MVMS-SC in terms of computation and communication efficiencies. Yao Zhao 0006, Youyang Qu, Feifei Chen 0001, Yong Xiang 0001, Longxiang Gao |
IEEE Trans. Mob. Comput. | 1 |
| 2024 | Long-Term Over One-Off: Heterogeneity-Oriented Dynamic Verification Assignment for Edge Data IntegrityabstractEdgeIntelligence (EI), a burgeoning research area, motivates App vendors to cache data replicas on geographically distributed edge servers to deliver better services. On the downside, this benefit also incurs more data integrity audit overhead on App vendors, which calls for more efficientEdgeDataIntegrity (EDI) verification approaches. However, existing EDI solutions totally rely on an implicitresource homogeneity assumption-edge servers have identical resource availability throughout EDI inspection execution in each round-but it rarely holds in reality. The edge servers with insufficient computation and/or communication capacity greatly limit overall EDI verification efficiency from a round perspective. Thus, in this work, we release the identified impractical assumption and accordingly study the EDIDynamicVerificationAssignment (DVA) problem for the first time. The problem aims to maximize the number of data replicas being verified in the long term under the constraints of verification delay in resource-limited environments. In this way, App vendors merely need to check the integrity of selected data replicas in each round for efficiency improvement. Specifically, we first formalize the DVA problem as a delay-constrained long-term stochastic optimization problem and further prove its$\mathcal {NP}$-hardness. To resolve the problem efficiently, we decompose it to an easy-to-handle form and then develop a polynomial-timePriority-based approach named DVA-P with a theoretical analysis of its time complexity and performance bound. Finally, experimental evaluations validate that DVA-P can be seamlessly incorporated into existing EDI solutions to enhance overall verification efficiency while guaranteeing verification performance. Yao Zhao 0006, Youyang Qu, Yong Xiang 0001, Chaochen Shi, Feifei Chen 0001, Longxiang Gao |
IEEE Trans. Mob. Comput. | 1 |
| 2024 | Adaptive Regularization and Resilient Estimation in Federated LearningabstractFederated Learning (FL) is an emerging research area that produces a globally trained model using numerous local users' data and maintains their privacy. Heterogeneous or non-Independent and Identically Distributed ( non-IID) data affect the global model's convergence and, therefore, cause high communication costs. These are because traditional FL approaches often disregard an adaptive regularized objective for the user-side training and utilize conventional arithmetic mean on the locally trained models for the server-side aggregation. To alleviate these issues, we propose a novel FL scheme in this paper. In particular, we propose an adaptive regularization approach to add to the classical objective function of the users' local models during training and a resilient estimation approach to the locally trained models during aggregation. The adaptive regularization approach is derived using the users' local and global performance diversification while the resilient estimation scheme uses a modified geometric mean aggregation over the local models' parameters. We provide consolidated theoretical results and perform extensive experiments on the IID and non-IID settings of MNIST, CIFAR-10, and Shakespeare datasets with various deep networks. The results manifest that our FL scheme outperforms the state-of-the-art approaches in terms of communication speedup, test-set performance, training convergence stability, and resiliency against attacks. Md Palash Uddin, Yong Xiang 0001, Yao Zhao 0006, Mumtaz Ali 0003, Yushu Zhang 0001, Longxiang Gao |
IEEE Trans. Serv. Comput. | 3 |
| 2024 | Long-Term Proof-of-Contribution: An Incentivized Consensus Algorithm for Blockchain-Enabled Federated LearningabstractThe surge in data collected by local devices has given rise to a distributed machine learning architecture namedFederatedLearning (FL) for privacy-preserving model training. However, the security of centralized aggregation of local models becomes a primary concern, which can be mitigated byBlockchain-enabledFederatedLearning (BFL) to facilitate decentralized model aggregation. In BFL, consensus and incentive are two of the key components that impact the scalability, security, and consistency of the system. Existing joint solutions focus on selecting a block producer based on client contributions to model training but overlook contributions to blockchain consensus and lack consideration for correlations across communication rounds, inevitably affecting incentive performance. Motivated by these, we make the first attempt to achieve blockchain consensus with along-term incentive guaranteefor BFL systems. Following a generalizable BFL workflow, we decouple the global contribution of BFL clients into four rigorously modeled metrics, and formulate the block producer selection problem as a long-term total contribution maximization problem with reward constraints. ALong-termProof-of-Contribution algorithm named LPoC is developed to handle this problem efficiently. In each communication round, LPoC identifies an optimal block producer that can maximize total contributions from a long-term perspective while allocating rewards to continuously motivate clients to contribute to BFL. We provide a detailed analysis of time complexity and performance bounds, followed by extensive experimental evaluations. The results demonstrate the effectiveness of LPoC in maximizing long-term total contribution, improving consensus efficiency, and upgrading training performance. Yao Zhao 0006, Youyang Qu, Yong Xiang 0001, Feifei Chen 0001, Longxiang Gao |
IEEE Trans. Serv. Comput. | 1 |
| 2024 | Winning at the Starting Line: Unreliable Data Replica Selection for Edge Data Integrity VerificationabstractMobileEdgeComputing (MEC) is an emerging technology, where App vendors are allowed to cache multiple data replicas on geographically distributed edge servers to serve adjacent mobile subscribers. However, this benefit introduces an extra workload for edge servers and App vendors, as they must audit the integrity of multiple data replicas periodically considering various threats caused by distributed and dynamic MEC environments. The large-scale growth of data replicas certainly is a challenge to design more efficientEdgeDataIntegrity (EDI) verification approaches. Existing solutions are mostly limited to improving efficiency by optimizing proof generation and verification methods, while the improvement is still far from satisfactory due to adopting indiscriminate inspection philosophy (checking all data replicas without discrimination). In this paper, we make the first attempt to abstract a pre-processing phase and correspondingly study theUnreliable dataReplicaSelection (URS) problem. It can be seamlessly integrated into existing EDI solutions by solving the URS problem at the start of each verification round. Such pre-selection can significantly enhance overall EDI verification efficiency by incorporating the cache serviceQualityofService (QoS) and verification success rate, especially in scenarios with a large number of data replicas. Specifically, we first formalize the URS problem as a constrained optimization problem, and further prove its$\mathcal {NP}$-hardness. To address the problem efficiently, we transform it into an easy-to-handle form and develop aPriority-based approach named URS-P. Both theoretical analysis and experimental evaluation validate the effectiveness and efficiency of our proposed solution. Yao Zhao 0006, Youyang Qu, Yong Xiang 0001, Feifei Chen 0001, Md Palash Uddin, Longxiang Gao |
IEEE Trans. Serv. Comput. | 1 |
| 2023 | A Lightweight Model-Based Evolutionary Consensus Protocol in Blockchain as a Service for IoTabstractInternet of Things (IoT) is experiencing fast proliferation with emerging trends in autonomy and local decision-making to avoid the explosive burden on network infrastructure between cloud and edge. Thereby, blockchain as a Service (BaaS) for IoT, as an emerging distributed services computing paradigm, has drawn intense attention due to its decentralization, auditability, and tamper-resistance. However, the primary challenge is to design a tailor-made consensus protocol that is applicable to BaaS for IoT. Existing consensus protocols generally focus on power-intensive environments, which is not feasible for power-constrained BaaS-enabled IoT systems. In this article, to fully exploit BaaS's superiority (e.g., to sharing data securely), we propose a lightweight model-based evolutionary consensus protocol called Proof of Evolutionary Model (PoEM) that can improve the quality of BaaS in IoT environments. Beyond existing rule-based consensus protocols, PoEM iteratively trains a machine learning model to achieve consensus. In this way, PoEM enhances consensus efficiency and enables low-performance IoT devices to be involved. Moreover, considering IoT environments’ dynamics, a novel mechanism is designed to manage nodes joining and exiting dynamically. Extensive analytical and experimental results show PoEM's improved consensus efficiency and applicability in dynamic BaaS-based IoT environments while providing high-level security guarantees. Yao Zhao 0006, Youyang Qu, Yong Xiang 0001, Yushu Zhang 0001, Longxiang Gao |
IEEE Trans. Serv. Comput. | 1 |
| 2023 | CoSS: Leveraging Statement Semantics for Code SummarizationabstractAutomated code summarization tools allow generating descriptions for code snippets in natural language, which benefits software development and maintenance. Recent studies demonstrate that the quality of generated summaries can be improved by using additional code representations beyond token sequences. The majority of contemporary approaches mainly focus on extracting code syntactic and structural information from abstract syntax trees (ASTs). However, from the view of macro-structures, it is challenging to identify and capture semantically meaningful features due to fine-grained syntactic nodes involved in ASTs. To fill this gap, we investigate how to learn more code semantics and control flow features from the perspective of code statements. Accordingly, we propose a novel model entitled CoSS for code summarization. CoSS adopts a Transformer-based encoder and a graph attention network-based encoder to capture token-level and statement-level semantics from code token sequence and control flow graph, respectively. Then, after receiving two-level embeddings from encoders, a joint decoder with a multi-head attention mechanism predicts output sequences verbatim. Performance evaluations on Java, Python, and Solidity datasets validate that CoSS outperforms nine state-of-the-art (SOTA) neural code summarization models in effectiveness and is competitive in execution efficiency. Further, the ablation study reveals the contribution of each model component. Chaochen Shi, Borui Cai, Yao Zhao 0006, Longxiang Gao, Keshav Sood, Yong Xiang 0001 |
IEEE Trans. Software Eng. | 3 |
| 2018 | Blockchain-Based UDDI Data Replication and SharingabstractUniversal Description, Discovery and Integration (UDDI), as the main technical support of Web service, plays an increasingly important role in SOA (Service Oriented Architecture). Smart contract technology based on blockchain is used to solve the problem of data replication and sharing between UDDI registries. The UDDI registry only stores the information of simple service description and reference to the blockchain, and provides identity management, authority management, and other functions. The detailed information of the service provider and the services it provides is stored in the blockchain. The cost of the UDDI registry is greatly reduced, and the sharing and storage of data can be easily realized. Meanwhile, the security and auditability of data are guaranteed by the characteristics of the blockchain. This paper innovatively introduces the blockchain into the field of service computing, which provides a new idea for future research. Yao Zhao 0006, Lu Zhao 0001 |
CSCWD | 1 |