EDBT 2026 Demo / reviewers in the wild / expert
Zhiyuan Wang 0002
dblp:32/3351-2
· DBLP profile ↗
26ranked-venue papers
4as first author
26since 2021 · last 2026
0000-0002-5368-1132ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 19 · 3 first-author · 19 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Accelerating Decentralized Federated Learning With Probabilistic Communication in Heterogeneous Edge ComputingabstractDecentralized federated learning (DFL) has gained popularity for training machine learning models on massive data in edge computing, as it avoids the potential bottleneck of conventional parameter server architectures. However, the existing DFL solutions typically use deterministic topologies that struggle with both system heterogeneity and non-IID local data, resulting in high bandwidth costs and slow convergence rates. In this paper, we propose a novel mechanism called Communication-efficient Decentralized Federated Learning (CedFL) to accelerate model training. InCedFL, each worker will communicate with each of its neighbors (i.e., model exchange) according to a certain probability at each epoch, so as to reduce bandwidth consumption. To this end, we then propose an efficient algorithm to adaptively determine the optimal probability for each worker pair according to real-time system situations (e.g., data distribution and bandwidth resource). Our proposed mechanism has been extensively tested on classical models and datasets, and the results demonstrate its high effectiveness.CedFLhas been shown to reduce completion time for model training by approximately 55% and improve test accuracy by 11% under the bandwidth constraint, compared to state-of-the-art solutions. Jianchun Liu, Jiaming Yan, Hongli Xu 0001, Lun Wang 0003, Zhiyuan Wang 0002, Jinyang Huang, Chunming Qiao |
IEEE Trans. Netw. | 5 |
| 2025 | Hier-FUN: Hierarchical Federated Learning and Unlearning in Heterogeneous Edge ComputingabstractFederated learning (FL) has emerged as a pivotal paradigm for distributed model training in edge computing (EC), enabling cooperation among numerous Internet of Things devices while safeguarding their data privacy. Despite its successes in machine learning, concerns regarding data security and model fidelity necessitate the efficient unlearning of target device, i.e., federated unlearning (FUN). However, due to resource constraints, device heterogeneity, and non-independent and identically distributed (Non-IID) data, securely eliminating a device’s impact without retraining the model from scratch presents a complex challenge. In response to these challenges, we propose a hierarchical FUN framework, called Hier-FUN. Hier-FUN organizes edge devices into K clusters, each managed by a head device responsible for aggregating local models within the cluster. To expedite both the learning and unlearning processes of Hier-FUN, we design a heuristic algorithm to determine an appropriate value for K based on devices’ data distributions and available resources. In addition, Hier-FUN denies the communication between the server and cluster heads during training, which can constrain the influence sphere of target device and accelerate the unlearning process. We conduct extensive experiments using real-world datasets, and the experimental results illustrate that Hier-FUN can improve test accuracy by 3.19% during the learning phase and achieve a$6.8\times $speedup during unlearning compared with the baseline methods. Zhen-guo Ma, Huaqing Tu, Pengli Ji, Xiaoran Yan, Hongli Xu 0001, Zhiyuan Wang 0002, Suo Chen |
IEEE Internet Things J. | 7 |
| 2025 | FedSNN: Training Slimmable Neural Network With Federated Learning in Edge ComputingabstractTo provide a flexible tradeoff between inference accuracy and resource requirement at runtime, the slimmable neural network (SNN), a single network executable at different widths with the same deploying and management cost as that of a single model, has been proposed. However, how to effectively train SNN among massive devices in edge computing without revealing their local data remains an open problem. To this end, we leverage a novel distributed machine learning paradigm, i.e., federated learning, to realize effective on-device SNN training. As current FL schemes often train only one model with fixed architecture, and the existing SNN training algorithm is resource-intensive, integrating FL and SNN is non-trivial. Furthermore, two intrinsic features in edge computing, i.e., data and system heterogeneity, exacerbate the difficulty. Motivated by this, we redesign the model distribution, local training, and model aggregation phases in traditional FL, and propose FedSNN, a framework that ensures all widths in SNN can obtain high accuracy with less resource consumption. Specifically, for devices with heterogeneous training capacities and data distributions, the parameter server will distribute each of them with one proper width for adaptive local training guided by their uploaded model features, and their trained models will be weighted-averaged using the proposed multi-width SNN aggregation to improve their statistical utility. Extensive experiments on a distributed testbed show that FedSNN improves the model accuracy by about 2.18%-8.1%, and accelerates training by about$1.31\times $-$6.84\times $, compared with existing solutions. Yang Xu 0020, Yunming Liao, Hongli Xu 0001, Zhiyuan Wang 0002, Lun Wang 0003, Jianchun Liu, Chen Qian 0001 |
IEEE Trans. Netw. | 4 |
| 2025 | Enhancing Federated Learning Through Layer-Wise Aggregation Over Non-IID DataabstractNowadays, federated learning (FL) has been widely adopted to train deep neural networks (DNNs) among massive devices without revealing their local data in edge computing (EC). To relieve the communication bottleneck of the central server in FL, hierarchical federated learning (HFL), which leverages edge servers as intermediaries to perform model aggregation among devices in proximity, comes into being. Nevertheless, the existing HFL systems may not perform training effectively due to bandwidth constraints and non-IID issues on devices. To conquer these challenges, we introduce anHFL system with device-edgeassignment andlayer selection, namely Heal. Specifically, Heal organizes all the devices into a hierarchical structure (i.e., device-edge assignment) and enables each device to forward only a sub-model with several valuable layers for aggregation (i.e., layer selection). This processing procedure is called layer-wise aggregation. To further save communication resource and improve the convergence performance, we then design an iteration-based algorithm to optimize the development of our layer-wise aggregation strategy by considering the data distribution as well as resource constraints among devices. Extensive experiments on both the physical platform and the simulated environment show that Heal accelerates DNN training by about 1.4–12.5×, and reduces the network traffic consumption by about 31.9–64.1%, compared with the existing HFL systems. Yang Xu 0020, Zhiyuan Wang 0002, Hongli Xu 0001, Yunming Liao |
IEEE Trans. Serv. Comput. | 3 |
| 2024 | Clients Help Clients: Alternating Collaboration for Semi-Supervised Federated LearningabstractFederated learning (FL) provides a distributed framework for multiple clients to collaboratively train models without exposing raw data. Most FL research assumes that all clients have fully labeled data, which is impractical for many real-world applications. To this end, we focus on semi-supervised FL (SSFL), where data samples of each client are partially labeled. However, existing SSFL methods ignore two inherent characteristics of FL: limited communication resources and heterogeneous data distribution, which severely hinder convergence stability and efficiency. This paper proposes a novel SSFL mechanism, called FedAC, to address the above two challenges by alternating client-to-client (C2C) collaboration. Specifically, we group all clients using different clustering strategies at two different training stages. During each global round, FedAC first performs similarity clustering based on local data distribution, which gathers the knowledge from similar clients to generate high-quality pseudo-labels for unlabeled data. Then the clients are re-grouped using dissimilarity clustering strategy to approximate the IID setting at the cluster level, thereby alleviating the bias induced by Non-IID data. FedAC adopts a reinforcement learning algorithm to achieve a balance between labeling assistance from similar clients and unbiased optimization from dissimilar clients. Extensive evaluations demonstrate that FedAC can improve model accuracy and save up to 59.65% of communication costs compared with existing benchmarks. Zhida Jiang, Yang Xu 0020, Hongli Xu 0001, Zhiyuan Wang 0002, Chunming Qiao |
ICDE | 4 |
| 2024 | Enhancing Decentralized and Personalized Federated Learning With Topology ConstructionabstractThe emerging Federated Learning (FL) permits all workers (e.g., mobile devices) to cooperatively train a model using their local data at the network edge. In order to avoid the possible bottleneck of conventional parameter server architecture, the decentralized federated learning (DFL) is developed on the peer-to-peer (P2P) communication. Non-IID issue is a key challenge in FL and will significantly degrade the model training performance. To this end, we propose a personalized solution called TOPFL, in which only parts of the local models (not the entire models) are shared and aggregated. Moreover, considering the limited communication bandwidth on workers, we propose a topology construction algorithm to accelerate the training process. To verify the convergence of the decentralized training framework, we theoretically analyze the impact of the data heterogeneity and topology on the convergence upper bound. Extensive simulation results show that TOPFL can achieve 2.2× speedup when reaching convergence and 5.8% higher test accuracy under the same resource consumption, compared with the baseline solutions. Suo Chen, Yang Xu 0020, Hongli Xu 0001, Zhen-guo Ma, Zhiyuan Wang 0002 |
IEEE Trans. Mob. Comput. | 5 |
| 2024 | Semi-Supervised Decentralized Machine Learning With Device-to-Device CooperationabstractThe massive data from mobile and embedded devices have huge potential for training machine learning models. Decentralized machine learning (DML) can avoid the inherent bottleneck of the parameter server (PS) by collaboratively training models in a device-to-device (D2D) fashion. However, the previous DML works often assume that the local data are fully annotated with ground-truth labels, which is unrealistic for many Internet of Things (IoT) applications. This arises a new practical DML scenario, namely semi-supervised DML, where the local data of distributed workers are partially labeled in the D2D network. The existing semi-supervised learning techniques are proposed for standalone or the PS architecture, which ignore the impact of D2D topology on the performance of semi-supervised learning. Thus, they cannot adequately leverage the unlabeled data of decentralized workers, leading to performance degradation. Herein, we propose a novel framework, called SSD, to address the problem of semi-supervised DML by exploiting D2D cooperation. The key insight behind SSD is that neighbor selection has a crucial impact on pseudo-label quality and communication overhead. In SSD, each worker adaptively selects its neighbors with high-quality models and similar data distribution under communication resource constraints, which helps to generate high-confidence pseudo-labels for local unlabeled data and further boosts the DML performance. Extensive empirical evaluations on both testbed and simulated environments show that SSD significantly outperforms other baselines. Zhida Jiang, Yang Xu 0020, Hongli Xu 0001, Zhiyuan Wang 0002, Jianchun Liu, Chunming Qiao |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | Computation and Communication Efficient Federated Learning With Adaptive Model PruningabstractFederated learning (FL) has emerged as a promising distributed learning paradigm that enables a large number of mobile devices to cooperatively train a model without sharing their raw data. The iterative training process of FL incurs considerable computation and communication overhead. The workers participating in FL are usually heterogeneous and the workers with poor capabilities may become the bottleneck of model training. To address the challenges of resource overhead and system heterogeneity, this article proposes an efficient FL framework, called FedMP, that improves both computation and communication efficiency over heterogeneous workers through adaptive model pruning. We theoretically analyze the impact of pruning ratio on training performance, and employ a Multi-Armed Bandit based online learning algorithm to adaptively determine different pruning ratios for heterogeneous workers, even without any prior knowledge of their capabilities. As a result, each worker in FedMP can train and transmit the sub-model that fits its own capabilities, accelerating the training process without hurting model accuracy. To prevent the diverse structures of pruned models from affecting the training convergence, we further present a new parameter synchronization scheme, called Residual Recovery Synchronous Parallel (R2SP). Besides, our proposed framework can be extended to the peer-to-peer (P2P) setting. Extensive experiments on physical devices demonstrate that FedMP is effective for different heterogeneous scenarios and data distributions, and can provide up to 4.1× speedup compared to the existing FL methods. Zhida Jiang, Yang Xu 0020, Hongli Xu 0001, Zhiyuan Wang 0002, Jianchun Liu, Chen Qian 0001, Chunming Qiao |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | Finch: Enhancing Federated Learning With Hierarchical Neural Architecture SearchabstractFederated learning (FL) has been widely adopted to train machine learning models over massive data in edge computing. Most works of FL employ pre-defined model architectures on all participating clients for model training. However, these pre-defined architectures may not be the optimal choice for the FL setting since manually designing a high-performance neural architecture is complicated and burdensome with intense human expertise and effort, which easily makes the model training fall into the local suboptimal solution. To this end, Neural Architecture Search (NAS) has been applied to FL to address this critical issue. Unfortunately, the search space of existing federated NAS approaches is extraordinarily large, resulting in unacceptable completion time on the resource-constrained edge clients, especially under the non-independent and identically distributed (non-IID) setting. In order to remedy this, we propose a novel framework, calledFinch, which adopts hierarchical neural architecture search to enhance federated learning. InFinch, we first divide the clients into several clusters according to the data distribution. Then, some subnets are sampled from a pre-trained supernet and allocated to the specific client clusters for searching the optimal model architecture in parallel, so as to significantly accelerate the process of model searching and training. The extensive experimental results demonstrate the high effectiveness of our proposed framework. Specifically,Finchcan reduce the completion time by about 30.6%, and achieve an average accuracy improvement of around 9.8% compared with the baselines. Jianchun Liu, Jiaming Yan, Hongli Xu 0001, Zhiyuan Wang 0002, Jinyang Huang, Yang Xu 0020 |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | Federated Learning With Client Selection and Gradient Compression in Heterogeneous Edge SystemsabstractFederated learning (FL) has recently gained tremendous attention in edge computing and Internet of Things, due to its capability of enabling distributed clients to cooperatively train models while keeping raw data locally. However, the existing works usually suffer from limited communication resources, dynamic network conditions and heterogeneous client properties, which hinder efficient FL. To simultaneously tackle the above challenges, we propose a heterogeneity-aware FL framework, called FedCG, with adaptive client selection and gradient compression. Specifically, FedCG introduces diversity to client selection and aims to select a representative client subset considering statistical heterogeneity. These selected clients are assigned different compression ratios based on heterogeneous and time-varying capabilities. After local training, they upload sparse model updates matching their capabilities for global aggregation, which can effectively reduce the communication cost and mitigate the straggler effect. More importantly, instead of naively combining client selection and gradient compression, we highlight that their decisions are tightly coupled and indicate the necessity of joint optimization. We theoretically analyze the impact of both client selection and gradient compression on convergence performance. Guided by the convergence rate, we develop an iteration-based algorithm to jointly optimize client selection and compression ratio decision using submodular maximization and linear programming. On this basis, we propose the quantized extension of FedCG, termed Q-FedCG, which further adjusts quantization levels based on gradient innovation. Extensive experiments on both real-world prototypes and simulations show that FedCG and its extension can provide up to 6.4× speedup. Yang Xu 0020, Zhida Jiang, Hongli Xu 0001, Zhiyuan Wang 0002, Chen Qian 0001, Chunming Qiao |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | Enhancing Federated Learning With Server-Side Unlabeled Data by Adaptive Client and Data SelectionabstractFederated learning (FL) has been widely applied to collaboratively train deep learning (DL) models on massive end devices (i.e., clients). Due to the limited storage capacity and high labeling cost, the data on each client may be insufficient for model training. Conversely, in cloud datacenters, there exist large-scale unlabeled data, which are easy to collect from public access (e.g., social media). Herein, we propose theAda-FedSemisystem, which leverages both on-device labeled data and in-cloud unlabeled data to boost the performance of DL models. In each round, local models are aggregated to produce pseudo-labels for the unlabeled data, which are utilized to enhance the global model. Considering that the number of participating clients and the quality of pseudo-labels will have a significant impact on the training performance, we introduce a multi-armed bandit (MAB) based online algorithm to adaptively determine the participating fraction and confidence threshold. Besides, to alleviate the impact of stragglers, we assign local models of different depths for heterogeneous clients. Extensive experiments on benchmark models and datasets show that given the same resource budget, the model trained by Ada-FedSemi achieves 3%$\sim$14.8% higher test accuracy than that of the baseline methods. When achieving the same test accuracy, Ada-FedSemi saves up to 48% training cost, compared with the baselines. Under the scenario with heterogeneous clients, the proposed HeteroAda-FedSemi can further speed up the training process by$1.3\times \sim 1.5\times$. Yang Xu 0020, Lun Wang 0003, Hongli Xu 0001, Jianchun Liu, Zhiyuan Wang 0002, Liusheng Huang |
IEEE Trans. Mob. Comput. | 5 |
| 2024 | Peaches: Personalized Federated Learning With Neural Architecture Search in Edge ComputingabstractIn edge computing (EC), federated learning (FL) enables numerous distributed devices (or workers) to collaboratively train AI models without exposing their local data. Most works of FL adopt a predefined architecture on all participating workers for model training. However, since workers' local data distributions vary heavily in EC, the predefined architecture may not be the optimal choice for every worker. It is also unrealistic to manually design a high-performance architecture for each worker, which requires intense human expertise and effort. In order to tackle this challenge, neural architecture search (NAS) has been applied in FL to automate the architecture design process. Unfortunately, the existing federated NAS frameworks often suffer from the difficulties of system heterogeneity and resource limitation. To remedy this problem, we present a novel framework, termedPeaches, to achieve efficient searching and training in the resource-constrained EC system. Specifically, the local model of each worker is stacked by base cell and personal cell, where the base cell is shared by all workers to capture the common knowledge and the personal cell is customized for each worker to fit the local data. We determine the number of base cells, shared by all workers, according to the bandwidth budget on the parameters server. Besides, to relieve the data and system heterogeneity, we find the optimal number of personal cells for each worker based on its computing capability. In addition, we gradually prune the search space during training to mitigate the resource consumption. We evaluate the performance ofPeachesthrough extensive experiments, and the results show thatPeachescan achieve an average accuracy improvement of about 6.29% and up to 3.97× speed up compared with the baselines. Jiaming Yan, Jianchun Liu, Hongli Xu 0001, Zhiyuan Wang 0002, Chunming Qiao |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | YOGA: Adaptive Layer-Wise Model Aggregation for Decentralized Federated LearningabstractTraditional Federated Learning (FL) is a promising paradigm that enables massive edge clients to collaboratively train deep neural network (DNN) models without exposing raw data to the parameter server (PS). To avoid the bottleneck on the PS, Decentralized Federated Learning (DFL), which utilizes peer-to-peer (P2P) communication without maintaining a global model, has been proposed. Nevertheless, DFL still faces two critical challenges, i.e., limited communication bandwidth and not independent and identically distributed (non-IID) local data, thus hindering efficient model training. Existing works commonly assume full model aggregation at periodic intervals, i.e., clients periodically collect models from peers. To reduce the communication cost, these methods allow clients to collect model(s) from selected peers, but often result in a significant degradation of model accuracy when dealing with non-IID data. Alternatively, the layer-wise aggregation mechanism has been proposed to alleviate communication overhead under the PS architecture, but its potential in DFL remains rarely explored yet. To this end, we propose an efficient DFL framework YOGA that adaptively performs layer-wise model aggregation and training. Specifically, YOGA first generates the ranking of layers in the model according to the learning speed and layer-wise divergence. Combining with the layer ranking and peers’ status information (i.e., data distribution and communication capability), we propose the max-match (MM) algorithm to generate the proper layer-wise model aggregation policy for the clients. Extensive experiments on DNN models and datasets show that YOGA saves communication cost by about 45% without sacrificing the model performance compared with the baselines, and provides 1.53-$3.5\times $speedup on the physical platform. Jun Liu 0083, Jianchun Liu, Hongli Xu 0001, Yunming Liao, Zhiyuan Wang 0002, Qianpiao Ma |
IEEE/ACM Trans. Netw. | 5 |
| 2024 | Adaptive Block-Wise Regularization and Knowledge Distillation for Enhancing Federated LearningabstractFederated Learning (FL) is a distributed model training framework that allows multiple clients to collaborate on training a global model without disclosing their local data in edge computing (EC) environments. However, FL usually faces statistical heterogeneity (e.g., non-IID data) and system heterogeneity (e.g., computing and communication capabilities), resulting in poor model training performance. To deal with the above two challenges, we propose an efficient FL framework, named FedBR, which integrates the idea of block-wise regularization and knowledge distillation (KD) into the pioneering FL algorithm FedAvg, for resource-constrained edge computing. Specifically, we first divide the model into multiple blocks according to the layer order of deep neural network (DNN). The server only sends some consecutive model blocks instead of an entire model to clients for communication efficiency. Then, the clients make use of knowledge distillation to absorb the knowledge of global model blocks to alleviate statistical heterogeneity during local training. We provide a theoretical convergence guarantee for FedBR and show that the convergence bound will decrease as the increasing number of model blocks sent by the server. Besides, since the increasing number of model blocks brings more computing and communication costs, we design a heuristic algorithm (GMBS) to determine the appropriate number of model blocks for clients according to their varied data distributions, computing, and communication capabilities. Extensive experimental results show that FedBR can reduce the bandwidth consumption by about 31%, and achieve an average accuracy improvement of around 5.6% compared with the baselines under heterogeneous settings. Jianchun Liu, Qingmin Zeng, Hongli Xu 0001, Yang Xu 0020, Zhiyuan Wang 0002, He Huang 0001 |
IEEE/ACM Trans. Netw. | 5 |
| 2024 | FAST: Enhancing Federated Learning Through Adaptive Data Sampling and Local TrainingabstractThe emerging paradigm of federated learning (FL) strives to enable devices to cooperatively train models without exposing their raw data. In most cases, the data across devices are non-independently and identically distributed in FL. Thus, the local models trained over different data distributions will inevitably deviate from the global optima, which induces optimization inconsistency and even hurts global convergence. Moreover, the resource-constrained devices with heterogeneous training capacities (e.g., computing and communication) further slow down the convergence rate. To this end, we introduce anFL framework withadaptive datasampling and localtraining, namely FAST. Specifically, even without devices’ private data distributions, FAST enables each device to sample different rates of data points from each of its local classes to rebuild a dataset for training, thus adjusting the convergence direction of the aggregated global model to be closer to the global optima. The theoretical analysis shows that the convergence bound depends on the sampling rates as well as the number of local iterations executed on the sampled data. To achieve resource-effective and convergence-guaranteed FL, we then design an online learning algorithm that jointly optimizes the data sampling and local training strategies so as to encourage the decrease of global loss under the given time budget. Extensive experiments on physical and simulated environments show that, FAST improves the model accuracy by about 1.55%-6.78% given the same time budget, and accelerates training by about 1.39-5.89× with the same target accuracy, compared with the baselines. Zhiyuan Wang 0002, Hongli Xu 0001, Yang Xu 0020, Zhida Jiang, Jianchun Liu, Suo Chen |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2023 | FedCD: A Hybrid Centralized-Decentralized Architecture for Efficient Federated LearningabstractWith billions of IoT devices producing vast data globally, privacy and efficiency challenges arise in AI applications. Federated learning (FL) has been widely adopted to train deep neural networks (DNNs) without privacy leakage. Existing centralized and decentralized FL architectures have limitations, including memory burden, huge bandwidth pressure and non-IID data issues. This paper introduces a novel framework, named FedCD, merging the benefits of both centralized and decentralized FL architectures. FedCD strategically distributes the model based on layer sizes and consensus distances (measuring the deviation between the local models and the global average models), effectively relieving network bandwidth pressures and accelerating training speed even under the non-IID setting. This method significantly mitigates resource constraints and improves model accuracy, offering a promising solution to the challenges in distributed machine learning. Extensive experiment results show the high effectiveness of FedCD. The total completion time of FedCD is reduced by 16.3%-53% and the average accuracy improvement is 1.85% compared to the existing FL systems. Pengcheng Qu, Jianchun Liu, Zhiyuan Wang 0002, Qianpiao Ma, Jinyang Huang |
ICPADS | 3 |
| 2023 | Heterogeneity-Aware Federated Learning with Adaptive Client Selection and Gradient CompressionabstractFederated learning (FL) allows multiple clients cooperatively train models without disclosing local data. However, the existing works fail to address all these practical concerns in FL: limited communication resources, dynamic network conditions and heterogeneous client properties, which slow down the convergence of FL. To tackle the above challenges, we propose a heterogeneity-aware FL framework, called FedCG, with adaptive client selection and gradient compression. Specifically, the parameter server (PS) selects a representative client subset considering statistical heterogeneity and sends the global model to them. After local training, these selected clients upload compressed model updates matching their capabilities to the PS for aggregation, which significantly alleviates the communication load and mitigates the straggler effect. We theoretically analyze the impact of both client selection and gradient compression on convergence performance. Guided by the derived convergence rate, we develop an iteration-based algorithm to jointly optimize client selection and compression ratio decision using submodular maximization and linear programming. Extensive experiments on both real-world prototypes and simulations show that FedCG can provide up to 5.3× speedup compared to other methods. Zhida Jiang, Yang Xu 0020, Hongli Xu 0001, Zhiyuan Wang 0002, Chen Qian 0001 |
INFOCOM | 4 |
| 2023 | Enhanced Federated Learning with Adaptive Block-wise Regularization and Knowledge DistillationabstractFederated Learning (FL) has emerged as an efficient distributed model training framework that enables multiple clients cooperatively to train a global model without exposing their local data in edge computing (EC). However, FL usually faces statistical heterogeneity (e.g., non-IID data) and system heterogeneity (e.g., computing and communication capabilities), resulting in poor model training performance. To deal with the above two challenges, we propose an efficient FL framework, named FedBR, which integrates the idea of block-wise regularization and knowledge distillation (KD) into the pioneer FL algorithm FedAvg, for resource-constrained edge computing. Besides, we design a heuristic algorithm (GMBS) to determine the appropriate number of model blocks for clients according to their varied data distributions, computing, and communication capabilities. Extensive experimental results show that FedBR can reduce the time cost by 19.5% and the communication cost by 27% on average compared with the other three baselines when achieving the target testing accuracy under heterogeneous settings. Qingmin Zeng, Jianchun Liu, Hongli Xu 0001, Zhiyuan Wang 0002, Yang Xu 0020, Yangming Zhao |
IWQoS | 4 |
| 2023 | CoopFL: Accelerating federated learning with DNN partitioning and offloading in heterogeneous edge computing
Zhiyuan Wang 0002, Hongli Xu 0001, Yang Xu 0020, Zhida Jiang, Jianchun Liu |
Comput. Networks | 1 |
| 2023 | Accelerating Federated Learning With Cluster Construction and Hierarchical AggregationabstractFederated learning (FL) has emerged in edge computing to address the limited bandwidth and privacy concerns of traditional cloud-based training. However, the existing FL mechanisms may lead to a long training time and consume massive communication resources. In this paper, we propose an efficient FL mechanism, namely FedCH, to accelerate FL in heterogeneous edge computing. Different from existing works which adopt the pre-defined system architecture and train models in a synchronous or asynchronous manner, FedCH will construct a special cluster topology and perform hierarchical aggregation for training. Specifically, FedCH arranges all clients into multiple clusters based on their heterogeneous training capacities. The clients in one cluster synchronously forward their local updates to the cluster header for aggregation, while all cluster headers take the asynchronous method for global aggregation. Our analysis shows that the convergence bound depends on the number of clusters and the training epochs. We propose efficient algorithms to determine the optimal number of clusters with resource budgets and then construct the cluster topology to address the client heterogeneity. Extensive experiments on both physical platform and simulated environment show that FedCH reduces the completion time by 49.5-79.5% and the network traffic by 57.4-80.8%, compared with the existing FL mechanisms. Zhiyuan Wang 0002, Hongli Xu 0001, Jianchun Liu, Yang Xu 0020, He Huang 0001, Yangming Zhao |
IEEE Trans. Mob. Comput. | 1 |
| 2022 | FedMP: Federated Learning through Adaptive Model Pruning in Heterogeneous Edge ComputingabstractFederated learning (FL) has been widely adopted to train machine learning models over massive distributed data sources in edge computing. However, the existing FL frameworks usually suffer from the difficulties of resource limitation and edge heterogeneity. Herein, we design and implement FedMP, an efficient FL framework through adaptive model pruning. We theoretically analyze the impact of pruning ratio on model training performance, and propose to employ a Multi-Armed Bandit based online learning algorithm to adaptively determine different pruning ratios for heterogeneous edge nodes, even without any prior knowledge of their computation and communication capabilities. With adaptive model pruning, FedMP can not only reduce resource consumption but also achieve promising accuracy. To prevent the diverse structures of pruned models from affecting the training convergence, we further present a new parameter synchronization scheme, called Residual Recovery Synchronous Parallel (R2SP), and provide a theoretical convergence guarantee. Extensive experiments on the classical models and datasets demonstrate that FedMP is effective for different heterogeneous scenarios and data distributions, and can provide up to 4.1× speedup compared to the existing FL methods. Zhida Jiang, Yang Xu 0020, Hongli Xu 0001, Zhiyuan Wang 0002, Chunming Qiao, Yangming Zhao |
ICDE | 4 |
| 2022 | Enhancing Federated Learning with Intelligent Model Migration in Heterogeneous Edge ComputingabstractTo approach the challenges of non-IID data and limited communication resource raised by the emerging federated learning (FL) in mobile edge computing (MEC), we propose an efficient framework, called FedMigr, which integrates a deep reinforcement learning (DRL) based model migration strategy into the pioneer FL algorithm FedAvg. According to the data distribution and resource constraints, our FedMigr will intelligently guide one client to forward its local model to another client after local updating, rather than directly sending the local models to the server for global aggregation as in FedAvg. Intuitively, migrating a local model from one client to another is equivalent to training it over more data from different clients, contributing to alleviating the influence of non-IID issue. We prove that FedMigr can help to reduce the parameter divergences between different local models and the global model from a theoretical perspective, even over local datasets with non-IID settings. Extensive experiments on three popular benchmark datasets demonstrate that FedMigr can achieve an average accuracy improvement of around 13%, and reduce bandwidth consumption for global communication by 42% on average, compared with the baselines. Jianchun Liu, Yang Xu 0020, Hongli Xu 0001, Yunming Liao, Zhiyuan Wang 0002, He Huang 0001 |
ICDE | 5 |
| 2022 | Enhancing Federated Learning with In-Cloud Unlabeled DataabstractFederated learning (FL) has been widely applied to collaboratively train deep learning (DL) models on massive end devices (i.e., clients). Due to the limited storage capacity and high labeling cost, there are always insufficient data stored and annotated on each client. Conversely, in cloud datacenters, there exist large-scale unlabeled data, which are easy to collect from public access (e.g., social media). Herein, upon the federated semi-supervised learning (FSSL) technology, we propose the Ada-FedSemi system, which leverages both on-device labeled data and in-cloud unlabeled data to boost the performance of DL models. Given the limited communication and massive quantity of the clients, in each training round, we decide to select partial clients to participate in FL, and their local models are aggregated by the parameter server (PS) to produce pseudo-labels for the unlabeled data, which are utilized to enhance the global model. Considering that the number of participating clients and the quality of pseudo-labels will have a significant impact on the training performance (e.g., efficiency and accuracy), we introduce a multi-armed bandit (MAB) based online algorithm to adaptively determine the participating fraction and confidence threshold during federated model training. Extensive experiments on benchmark models and datasets show that, given the same resource budget, the model trained by Ada-FedSemi achieves 3%-14.8 % higher test accuracy than that of the baseline methods. Besides, when achieving the same test accuracy, Ada-FedSemi saves up to 48% training cost, compared with the baselines. Lun Wang 0003, Yang Xu 0020, Hongli Xu 0001, Jianchun Liu, Zhiyuan Wang 0002, Liusheng Huang |
ICDE | 5 |
| 2021 | Resource-Efficient Federated Learning with Hierarchical Aggregation in Edge ComputingabstractFederated learning (FL) has emerged in edge computing to address limited bandwidth and privacy concerns of traditional cloud-based centralized training. However, the existing FL mechanisms may lead to long training time and consume a tremendous amount of communication resources. In this paper, we propose an efficient FL mechanism, which divides the edge nodes into K clusters by balanced clustering. The edge nodes in one cluster forward their local updates to cluster header for aggregation by synchronous method, called cluster aggregation, while all cluster headers perform the asynchronous method for global aggregation. This processing procedure is called hierarchical aggregation. Our analysis shows that the convergence bound depends on the number of clusters and the training epochs. We formally define the resource-efficient federated learning with hierarchical aggregation (RFL-HA) problem. We propose an efficient algorithm to determine the optimal cluster structure (i.e., the optimal value of K) with resource constraints and extend it to deal with the dynamic network conditions. Extensive simulation results obtained from our study for different models and datasets show that the proposed algorithms can reduce completion time by 34.8%-70% and the communication resource by 33.8%-56.5% while achieving a similar accuracy, compared with the well-known FL mechanisms. Zhiyuan Wang 0002, Hongli Xu 0001, Jianchun Liu, He Huang 0001, Chunming Qiao, Yangming Zhao |
INFOCOM | 1 |
| 2021 | DNN Inference Acceleration with Partitioning and Early Exiting in Edge Computing
Hongli Xu 0001, Yang Xu 0020, Zhiyuan Wang 0002, Liusheng Huang |
WASA (1) | 4 |
| 2021 | Communication-efficient asynchronous federated learning in resource-constrained edge computing
Jianchun Liu, Hongli Xu 0001, Yang Xu 0020, Zhen-guo Ma, Zhiyuan Wang 0002, Chen Qian 0001, He Huang 0001 |
Comput. Networks | 5 |