VLDB 2026 Research / reviewers in the wild / expert
Lun Wang 0003
dblp:130/1339-3
· DBLP profile ↗
20ranked-venue papers
3as first author
20since 2021 · last 2026
0000-0002-9436-7924ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 16 · 2 first-author · 16 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Accelerating Decentralized Federated Learning With Probabilistic Communication in Heterogeneous Edge ComputingabstractDecentralized federated learning (DFL) has gained popularity for training machine learning models on massive data in edge computing, as it avoids the potential bottleneck of conventional parameter server architectures. However, the existing DFL solutions typically use deterministic topologies that struggle with both system heterogeneity and non-IID local data, resulting in high bandwidth costs and slow convergence rates. In this paper, we propose a novel mechanism called Communication-efficient Decentralized Federated Learning (CedFL) to accelerate model training. InCedFL, each worker will communicate with each of its neighbors (i.e., model exchange) according to a certain probability at each epoch, so as to reduce bandwidth consumption. To this end, we then propose an efficient algorithm to adaptively determine the optimal probability for each worker pair according to real-time system situations (e.g., data distribution and bandwidth resource). Our proposed mechanism has been extensively tested on classical models and datasets, and the results demonstrate its high effectiveness.CedFLhas been shown to reduce completion time for model training by approximately 55% and improve test accuracy by 11% under the bandwidth constraint, compared to state-of-the-art solutions. Jianchun Liu, Jiaming Yan, Hongli Xu 0001, Lun Wang 0003, Zhiyuan Wang 0002, Jinyang Huang, Chunming Qiao |
IEEE Trans. Netw. | 4 |
| 2025 | FedSNN: Training Slimmable Neural Network With Federated Learning in Edge ComputingabstractTo provide a flexible tradeoff between inference accuracy and resource requirement at runtime, the slimmable neural network (SNN), a single network executable at different widths with the same deploying and management cost as that of a single model, has been proposed. However, how to effectively train SNN among massive devices in edge computing without revealing their local data remains an open problem. To this end, we leverage a novel distributed machine learning paradigm, i.e., federated learning, to realize effective on-device SNN training. As current FL schemes often train only one model with fixed architecture, and the existing SNN training algorithm is resource-intensive, integrating FL and SNN is non-trivial. Furthermore, two intrinsic features in edge computing, i.e., data and system heterogeneity, exacerbate the difficulty. Motivated by this, we redesign the model distribution, local training, and model aggregation phases in traditional FL, and propose FedSNN, a framework that ensures all widths in SNN can obtain high accuracy with less resource consumption. Specifically, for devices with heterogeneous training capacities and data distributions, the parameter server will distribute each of them with one proper width for adaptive local training guided by their uploaded model features, and their trained models will be weighted-averaged using the proposed multi-width SNN aggregation to improve their statistical utility. Extensive experiments on a distributed testbed show that FedSNN improves the model accuracy by about 2.18%-8.1%, and accelerates training by about$1.31\times $-$6.84\times $, compared with existing solutions. Yang Xu 0020, Yunming Liao, Hongli Xu 0001, Zhiyuan Wang 0002, Lun Wang 0003, Jianchun Liu, Chen Qian 0001 |
IEEE Trans. Netw. | 5 |
| 2025 | PairingFL: Efficient Federated Learning With Model Splitting and Client PairingabstractFederated learning (FL) has recently gained tremendous attention in edge computing and the Internet of Things, due to its capability of enabling clients to perform model training at the network edge or end devices (i.e., clients). However, these end devices are usually resource-constrained without the ability to train large-scale models. In order to accelerate the training of large-scale models on these devices, we incorporate Split Learning (SL) into Federated Learning (FL), and propose a novel FL framework, termedPairingFL. Specifically, we split a full model into a bottom model and a top model, and arrange participating clients into pairs, each of which collaboratively trains the two partial models as one client does in typical FL. Driven by the advantages of SL and FL, PairingFL is able to relax the computation burden on clients and protect model privacy. However, considering the features of system and statistical heterogeneity in edge networks, it is challenging to pair the clients by carefully developing the strategies of client partitioning and matching for efficient model training. To this end, we first theoretically analyze the convergence property of PairingFL, and obtain a convergence upper bound. Guided by this, we then design a greedy and efficient algorithm, which makes the joint decision of client partitioning and matching, so as to well balance the trade-off between convergence rate and model accuracy. The performance of PairingFL is evaluated through extensive simulation experiments. The experimental results demonstrate that PairingFL can speed up the training process by$4.6\times $compared to baselines when achieving the corresponding convergence accuracy. Ji Qi 0005, Yang Xu 0020, Yunming Liao, Hongli Xu 0001, Lun Wang 0003 |
IEEE Trans. Netw. | 6 |
| 2024 | Federated Semi-Supervised Learning with Local and Global Updating Frequency OptimizationabstractFederated learning (FL) has gained Important attention for training deep learning models across clients in edge computing. The scarcity of labeled data poses critical challenges in professional fields (e.g., medical diagnosis) for FL, motivating the development of federated semi-supervised learning (FSSL) to exploit unlabeled data. The existing works of FSSL always assign fixed values for global updating frequency and/or local updating frequency during training. Though the previous works can make full use of unlabeled data, they may still hinder the improvement of model performance on heterogeneous clients, due to that the two hyper-parameters are coupled and have a joint impact on training process. To this end, we propose a novel FSSL framework, named LOGO, to enhance training efficiency by jointly optimizing local and global updating frequencies. We analyze the coupled impact of these two hyper-parameters on model training. Then, we develop a multi-armed bandit (MAB) based online algorithm to adaptively determine diverse local updating frequencies for heterogeneous clients as well as appropriate global updating frequency, so as to address system heterogeneity and improve training efficiency. The performance of LOGO is evaluated through extensive simulation experiments. The experimental results demonstrate that LOGO achieves the model training speedup by 2.3 × and reduces communication cost by 44%, compared to the baselines. Xin Hang, Yang Xu 0020, Hongli Xu 0001, Yunming Liao, Lun Wang 0003 |
CCGrid | 5 |
| 2024 | MergeSFL: Split Federated Learning with Feature Merging and Batch Size RegulationabstractRecently, federated learning (FL) has emerged as a popular technique for edge AI to mine valuable knowledge in edge computing (EC) systems. To boost the performance of AI applications, large-scale models have received increasing attention due to their excellent generalized abilities. However, training and transmitting large-scale models will incur significant computing and communication burden on the resource-constrained workers, and the exchange of entire models may violate model privacy. To relax the burden of workers and protect model privacy, split federated learning (SFL) has been released by integrating both data and model parallelism. Despite resource limitations, SFL also faces two other critical challenges in EC systems, i.e., statistical heterogeneity and system heterogeneity. In order to address these challenges, we propose a novel SFL framework, termed MergeSFL, by incorporating feature merging and batch size regulation in SFL. Concretely, feature merging aims to merge the features from workers into a mixed feature sequence, which is approximately equivalent to the features derived from IID data and is employed to promote model accuracy. While batch size regulation aims to assign diverse and suitable batch sizes for heterogeneous workers to improve training efficiency. Moreover, MergeSFL explores to jointly optimize these two strategies upon their coupled relationship to better enhance the performance of SFL. Extensive experiments are conducted on a physical platform with 80 NVIDIA Jetson edge devices, and the experimental results show that MergeSFL can improve the final model accuracy by 5.82% to 26.22%, with a speedup by about 1.39x to 4.14x, compared to the baselines. Yunming Liao, Yang Xu 0020, Hongli Xu 0001, Lun Wang 0003, Chunming Qiao |
ICDE | 4 |
| 2024 | Towards Communication-Efficient Federated Graph Learning: An Adaptive Client Selection PerspectiveabstractFederated graph learning (FGL) has been proposed to collaboratively train the increasing graph data with graph neural networks (GNNs) in a recommendation system, aggregating the features of graph nodes and edges among these nodes. Nevertheless, implementing an efficient recommendation system with FGL still faces two primary challenges, i.e., limited communication bandwidth and non-IID local graph data. Existing works typically reduce communication frequency or transmission amount, which may suffer significant performance degradation under non-IID settings. Furthermore, some researchers propose to share the underlying structure information among all clients, which brings massive communication cost. To this end, we propose an efficient FGL framework, named FedACS, which adaptively selects a subset of clients for model training, to alleviate communication overhead and non-IID issues simultaneously. In FedACS, the global GNN model can learn significant hidden edges and the structure of graph data among selected clients, enhancing recommendation efficiency. This capability distinguishes it from the traditional FL client selection methods. To optimize the client selection process, we introduce a multi-armed bandit (MAB) based algorithm to select participating clients according to the resource budgets and the training performance (i.e., RMSE) under different data distributions. Experimental results show that, given the same resource budget, FedACS achieves the RMSE improvement of 5.4% over the baselines. Besides, when achieving the same RMSE performance, FedACS saves up to approximately 70.7% communication cost, compared with the baselines. Xianjun Gao, Jianchun Liu, Hongli Xu 0001, Qianpiao Ma, Lun Wang 0003 |
IWQoS | 5 |
| 2024 | Decentralized Federated Learning With Adaptive Configuration for Heterogeneous ParticipantsabstractData generated at the network edge can be processed locally by leveraging the paradigm of edge computing (EC). Aided by EC, decentralized federated learning (DFL), which overcomes the single-point-of-failure problem in the parameter server based federated learning, is becoming a practical and popular approach for machine learning over distributed data. However, DFL faces two critical challenges,i.e., system heterogeneity and statistical heterogeneity introduced by edge devices. To ensure fast convergence with the existence of slow edge devices, we present an efficient DFL method, termed FedHP, which integrates adaptive control of both local updating frequency and network topology to better support the heterogeneous participants. We establish a theoretical relationship between local updating frequency and network topology regarding model training performance and obtain a convergence upper bound. Upon the convergence bound, we propose an optimization algorithm that adaptively determines local updating frequencies and constructs the network topology, so as to speed up convergence and improve the model accuracy. We evaluate the performance of FedHP through extensive simulation and testbed experiments. Evaluation results show that the proposed FedHP can reduce the completion time by about 51% and improve model accuracy by at least 5% in heterogeneous scenarios, compared with the baselines. Yunming Liao, Yang Xu 0020, Hongli Xu 0001, Lun Wang 0003, Chen Qian 0001, Chunming Qiao |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | Overcoming Noisy Labels and Non-IID Data in Edge Federated LearningabstractFederated learning (FL) enables edge devices to cooperatively train models without exposing their raw data. However, implementing a practical FL system at the network edge mainly faces three challenges: label noise, data non-IIDness, and device heterogeneity, which seriously harm model performance and slow down convergence speed. Unfortunately, none of the existing works tackle all three challenges simultaneously. To this end, we develop a novel FL system, called Aorta, which features adaptive dataset construction and aggregation weightassignment. On each client, Aorta first calibrates potentially noisy labels and then constructs a training dataset with low noise, balanced distribution, and proper size. To fully utilize limited data on clients, we propose a global model guided method to select clean data and progressively correct noisy labels. To achieve balanced class distribution and proper dataset size, we propose a distribution-and-capability-aware data augmentation method to generate local training data. On the server, Aorta assigns aggregation weights based on the quality of local models to ensure that high-quality models have a greater influence on the global model. The model quality is measured through its cosine similarity with a benchmark model, which is trained on a clean and balanced dataset. We conduct extensive experiments on four datasets with various settings, including different noise types/ratios and non-IID types/levels. Compared to the baselines, Aorta improves model accuracy up to 9.8% on the datasets with moderate noise and non-IIDness, while providing a speedup of 4.2× on average when achieving the same target accuracy. Yang Xu 0020, Yunming Liao, Lun Wang 0003, Hongli Xu 0001, Zhida Jiang, Wuyang Zhang |
IEEE Trans. Mob. Comput. | 3 |
| 2024 | Enhancing Federated Learning With Server-Side Unlabeled Data by Adaptive Client and Data SelectionabstractFederated learning (FL) has been widely applied to collaboratively train deep learning (DL) models on massive end devices (i.e., clients). Due to the limited storage capacity and high labeling cost, the data on each client may be insufficient for model training. Conversely, in cloud datacenters, there exist large-scale unlabeled data, which are easy to collect from public access (e.g., social media). Herein, we propose theAda-FedSemisystem, which leverages both on-device labeled data and in-cloud unlabeled data to boost the performance of DL models. In each round, local models are aggregated to produce pseudo-labels for the unlabeled data, which are utilized to enhance the global model. Considering that the number of participating clients and the quality of pseudo-labels will have a significant impact on the training performance, we introduce a multi-armed bandit (MAB) based online algorithm to adaptively determine the participating fraction and confidence threshold. Besides, to alleviate the impact of stragglers, we assign local models of different depths for heterogeneous clients. Extensive experiments on benchmark models and datasets show that given the same resource budget, the model trained by Ada-FedSemi achieves 3%$\sim$14.8% higher test accuracy than that of the baseline methods. When achieving the same test accuracy, Ada-FedSemi saves up to 48% training cost, compared with the baselines. Under the scenario with heterogeneous clients, the proposed HeteroAda-FedSemi can further speed up the training process by$1.3\times \sim 1.5\times$. Yang Xu 0020, Lun Wang 0003, Hongli Xu 0001, Jianchun Liu, Zhiyuan Wang 0002, Liusheng Huang |
IEEE Trans. Mob. Comput. | 2 |
| 2024 | Ferrari: A Personalized Federated Learning Framework for Heterogeneous Edge ClientsabstractFederated semi-supervised learning (FSSL) has been proposed to address the insufficient labeled data problem by training models with pseudo-labeling. In previous FSSL systems, a single global model is always trained without an equivalent generalization ability for the clients under the non-IID setting. Accordingly, model personalization methods have been proposed to overcome this problem. Intuitively, seeking labeling assistance from other clients with similar data distribution,i.e., model migration, can effectively improve the personalization on the clients with scarce labeled data. However, previous works require to migrate a pre-fixed number of models among the clients, causing unnecessary resource waste and accuracy degradation due to resource heterogeneity. Considering that the number of model migrations and the quality of pseudo-labels have a significant impact on the training performance (e.g., efficiency and accuracy), we propose a novel personalized FSSL system, called Ferrari, to boost the efficiency of pseudo-labeling and training accuracy through adaptive model migrations among the clients. Specifically, Ferrari first generates the similarity-based ranking using a Gaussian KD-Tree, considering the varied data distributions among the clients. Combined with the ranking and clients' heterogeneous resource constraints, Ferrari then adaptively determines the proper model migration policy and confidence thresholds for high-quality pseudo-labeling and personalized training for clients. Extensive experiments on a physical platform show that Ferrari provides a 1.2$\sim 5.5\times$speedup without sacrificing model accuracy, compared to existing methods. Jianchun Liu, Hongli Xu 0001, Lun Wang 0003, Chen Qian 0001, Yunming Liao |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | Asynchronous Decentralized Federated Learning for Heterogeneous DevicesabstractData generated at the network edge can be processed locally by leveraging the emerging technology of Federated Learning (FL). However, non-IID local data will lead to degradation of model accuracy and the heterogeneity of edge nodes inevitably slows down model training efficiency. Moreover, to avoid the potential communication bottleneck in the parameter-server-based FL, we concentrate on the Decentralized Federated Learning (DFL) that performs distributed model training in Peer-to-Peer (P2P) manner. To address these challenges, we propose an asynchronous DFL system by incorporating neighbor selection and gradient push, termed AsyDFL. Specifically, we require each edge node to push gradients only to a subset of neighbors for resource efficiency. Herein, we first give a theoretical convergence analysis of AsyDFL under the complicated non-IID and heterogeneous scenario, and further design a priority-based algorithm to dynamically select neighbors for each edge node so as to achieve the trade-off between communication cost and model performance. We evaluate the performance of AsyDFL through extensive experiments on a physical platform with 30 NVIDIA Jetson edge devices. Evaluation results show that AsyDFL can reduce the communication cost by 57% and the completion time by about 35% for achieving the same test accuracy, and improve model accuracy by at least 6% under the non-IID scenario, compared to the baselines. Yunming Liao, Yang Xu 0020, Hongli Xu 0001, Min Chen 0033, Lun Wang 0003, Chunming Qiao |
IEEE/ACM Trans. Netw. | 5 |
| 2024 | Accelerating Federated Learning With Data and Model Parallelism in Edge ComputingabstractRecently, edge AI has been launched to mine and discover valuable knowledge at network edge. Federated Learning, as an emerging technique for edge AI, has been widely deployed to collaboratively train models on many end devices in data-parallel fashion. To alleviate the computation/communication burden on the resource-constrained workers (e.g., end devices) and protect user privacy, Spilt Federated Learning (SFL), which integrates both data parallelism and model parallelism in Edge Computing (EC), is becoming a practical and popular approach for model training over distributed data. However, apart from the resource limitation, SFL still faces two other critical challenges in EC, i.e., system heterogeneity and context dynamics. To overcome these challenges, we present an efficient SFL method, named AdaSFL, which controls both local updating frequency and batch size to better accelerate model training. We theoretically analyze the model convergence rate and obtain a convergence upper bound regarding local updating frequency given a fixed batch size. Upon this, we develop a control algorithm to determine adaptive local updating frequency and diverse batch sizes for heterogeneous workers to enhance the training efficiency. The experimental results show that AdaSFL can reduce the completion time by about 43% and the network traffic consumption by about 31% for achieving the similar test accuracy, compared to the baselines. Yunming Liao, Yang Xu 0020, Hongli Xu 0001, Lun Wang 0003, Chunming Qiao |
IEEE/ACM Trans. Netw. | 5 |
| 2024 | BOSE: Block-Wise Federated Learning in Heterogeneous Edge ComputingabstractAt the network edge, federated learning (FL) has gained attention as a promising approach for training deep learning (DL) models collaboratively across a large number of devices while preserving user privacy. However, FL still faces specific challenges related to the limited, heterogeneous and dynamic resources of devices. In most FL systems, all devices train the same model, while the devices with constrained resources, referred to as stragglers, will significantly slow down overall training process. It is intuitive to alleviate computation and communication load on the stragglers by training and transmitting a part of the model. Inspired by multi-exit models, we divide an original DL model into several non-overlapping blocks, which can be trained separately on the low-capability devices. Furthermore, we propose BOSE, a novel FL system that performs adaptiveblock-wisemodel training under resource constraints. Considering the diverse impacts of different blocks on model convergence and the varying training loads they incur, a naive block assignment strategy, e.g., uniformly random assignment, may not yield optimal model performance and fail to fully utilize available resources. To this end, we introduce two metrics, includinglearning speedanddevice-wise divergence, to measure the potential of blocks in promoting model convergence. Given resource budget, BOSE initially identifies a set of candidate blocks for each device and subsequently selects specific training blocks based on their potential for promoting model convergence. In general, blocks with higher potential are more likely to be chosen for training. Extensive experiments on a physical platform show that BOSE provides a 1.4$\times$$\sim$3.8$\times$speedup without sacrificing model accuracy, compared to the baselines. Lun Wang 0003, Yang Xu 0020, Hongli Xu 0001, Zhida Jiang, Min Chen 0033, Wuyang Zhang, Chen Qian 0001 |
IEEE/ACM Trans. Netw. | 1 |
| 2023 | Adaptive Configuration for Heterogeneous Participants in Decentralized Federated LearningabstractData generated at the network edge can be processed locally by leveraging the paradigm of edge computing (EC). Aided by EC, decentralized federated learning (DFL), which overcomes the single-point-of-failure problem in the parameter server (PS) based federated learning, is becoming a practical and popular approach for machine learning over distributed data. However, DFL faces two critical challenges, i.e., system heterogeneity and statistical heterogeneity introduced by edge devices. To ensure fast convergence with the existence of slow edge devices, we present an efficient DFL method, termed FedHP, which integrates adaptive control of both local updating frequency and network topology to better support the heterogeneous participants. We establish a theoretical relationship between local updating frequency and network topology regarding model training performance and obtain a convergence upper bound. Upon this, we propose an optimization algorithm, that adaptively determines local updating frequencies and constructs the network topology, so as to speed up convergence and improve the model accuracy. Evaluation results show that the proposed FedHP can reduce the completion time by about 51% and improve model accuracy by at least 5% in heterogeneous scenarios, compared with the baselines. Yunming Liao, Yang Xu 0020, Hongli Xu 0001, Lun Wang 0003, Chen Qian 0001 |
INFOCOM | 4 |
| 2023 | Adaptive Asynchronous Federated Learning in Resource-Constrained Edge ComputingabstractFederated learning (FL) has been widely adopted to train machine learning models over massive data in edge computing. However, machine learning faces critical challenges, e.g., data imbalance, edge dynamics, and resource constraints, in edge computing. The existing FL solutions cannot well cope with data imbalance or edge dynamics, and may cause high resource cost. In this paper, we propose an adaptive asynchronous federated learning (AAFL) mechanism. To deal with edge dynamics, a certain fraction$\alpha$of all local updates will be aggregated by their arrival order at the parameter server in each epoch. Moreover, the system can intelligently vary the number of local updated models for global model aggregation in different epochs with network situations. We then propose experience-driven algorithms based on deep reinforcement learning (DRL) to adaptively determine the optimal value of$\alpha$in each epoch for two cases of AAFL, single learning task and multiple learning tasks, so as to achieve less completion time of training under resource constraints. Extensive experiments on the classical models and datasets show high effectiveness of the proposed algorithms. Specifically, AAFL can reduce the completion time by about 70 percent and improve the learning accuracy by about 28 percent under resource constraints, compared with the state-of-the-art solutions. Jianchun Liu, Hongli Xu 0001, Lun Wang 0003, Yang Xu 0020, Chen Qian 0001, Jinyang Huang, He Huang 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2023 | Accelerating Decentralized Federated Learning in Heterogeneous Edge ComputingabstractIn edge computing (EC), federated learning (FL) enables massive devices to collaboratively train AI models without exposing local data. In order to avoid the possible bottleneck of the parameter server (PS) architecture, we concentrate on the decentralized federated learning (DFL), which adopts peer-to-peer (P2P) communication without maintaining a global model. However, due to the intrinsic features of EC, e.g., resource limitation and heterogeneity, network dynamics and non-IID data, DFL with a fixed P2P topology and/or an identical model compression ratio for all workers results in a slow convergence rate. In this paper, we propose an efficient algorithm (termed CoCo) to accelerate DFL by integrating optimization of topology Construction and model Compression. Concretely, we adaptively construct P2P topology and determine specific compression ratios for each worker to conquer the system dynamics and heterogeneity under bandwidth constraints. To reflect how the non-IID data influence the consistency of local models in DFL, we introduce the consensus distance, i.e., the discrepancy between local models, as the quantitative metric to guide the fine-grained operations of the joint optimization. Extensive simulation results show that CoCo achieves 10× speedup, and reduces the communication cost by about 50% on average, compared with the existing DFL baselines. Lun Wang 0003, Yang Xu 0020, Hongli Xu 0001, Min Chen 0033, Liusheng Huang |
IEEE Trans. Mob. Comput. | 1 |
| 2023 | Adaptive Control of Local Updating and Model Compression for Efficient Federated LearningabstractData generated at the network edge can be processed locally by leveraging the paradigm of Edge Computing (EC). Aided by EC, Federated Learning (FL) has been becoming a practical and popular approach for distributed machine learning over locally distributed data. However, FL faces three critical challenges, i.e., resource constraint, system heterogeneity and context dynamics in EC. To address these challenges, we present a training-efficient FL method, termedFedLamp, by optimizing both theLocal updating frequency andmodel compression ratio in the resource-constrained EC systems. We theoretically analyze the model convergence rate and obtain a convergence upper bound related to the local updating frequency and model compression ratio. Upon the convergence bound, we propose a control algorithm, that adaptively determines diverse and appropriate local updating frequencies and model compression ratios for different edge nodes, so as to reduce the waiting time and enhance the training efficiency. We evaluate the performance ofFedLampthrough extensive simulation and testbed experiments. Evaluation results show thatFedLampcan reduce the traffic consumption by 63% and the completion time by about 52% for achieving the similar test accuracy, compared to the baselines. Yang Xu 0020, Yunming Liao, Hongli Xu 0001, Zhen-guo Ma, Lun Wang 0003, Jianchun Liu |
IEEE Trans. Mob. Comput. | 5 |
| 2023 | Joint Model Pruning and Topology Construction for Accelerating Decentralized Machine LearningabstractRecently, mobile and embedded devices worldwide generate a massive amount of data at the network edge. To efficiently exploit the data from distributed devices, we concentrate on decentralized machine learning (DML), where the workers collaboratively train models under the peer-to-peer (P2P) setting. DML avoids the bottleneck of the parameter server (PS) by enabling the workers to exchange local models with their neighbors rather than the PS. However, DML still faces some key challenges, i.e., resource limitation, system heterogeneity, network dynamics and non-IID data. In this article, we design and implement MOTOR, an efficient DML mechanism that simultaneously addresses these challenges by applying model pruning and topology construction, thus accelerating DML. Specifically, MOTOR assigns different pruning ratios to heterogeneous workers. After model pruning, each worker will train and transmit a sub-model that fits its capabilities, reducing both computation and communication overhead. Besides, MOTOR dynamically constructs the network topology considering the time-varying network conditions and non-IID data distributions. We theoretically analyze the impact of pruning ratio and network topology on model training performance. Guided by the theoretical analysis, we develop a joint optimization algorithm for pruning ratio decision and topology construction to achieve the trade-off between resource overhead and training performance. We implement MOTOR on commercial devices and evaluate the performance with different DML tasks. Extensive experiments show that MOTOR achieves up to 4.2× speedup compared to the existing DML approaches. Zhida Jiang, Yang Xu 0020, Hongli Xu 0001, Lun Wang 0003, Chunming Qiao, Liusheng Huang |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2022 | Enhancing Federated Learning with In-Cloud Unlabeled DataabstractFederated learning (FL) has been widely applied to collaboratively train deep learning (DL) models on massive end devices (i.e., clients). Due to the limited storage capacity and high labeling cost, there are always insufficient data stored and annotated on each client. Conversely, in cloud datacenters, there exist large-scale unlabeled data, which are easy to collect from public access (e.g., social media). Herein, upon the federated semi-supervised learning (FSSL) technology, we propose the Ada-FedSemi system, which leverages both on-device labeled data and in-cloud unlabeled data to boost the performance of DL models. Given the limited communication and massive quantity of the clients, in each training round, we decide to select partial clients to participate in FL, and their local models are aggregated by the parameter server (PS) to produce pseudo-labels for the unlabeled data, which are utilized to enhance the global model. Considering that the number of participating clients and the quality of pseudo-labels will have a significant impact on the training performance (e.g., efficiency and accuracy), we introduce a multi-armed bandit (MAB) based online algorithm to adaptively determine the participating fraction and confidence threshold during federated model training. Extensive experiments on benchmark models and datasets show that, given the same resource budget, the model trained by Ada-FedSemi achieves 3%-14.8 % higher test accuracy than that of the baseline methods. Besides, when achieving the same test accuracy, Ada-FedSemi saves up to 48% training cost, compared with the baselines. Lun Wang 0003, Yang Xu 0020, Hongli Xu 0001, Jianchun Liu, Zhiyuan Wang 0002, Liusheng Huang |
ICDE | 1 |
| 2022 | Decentralized Machine Learning Through Experience-Driven Method in Edge NetworksabstractData generated at the network edge can be processed locally by leveraging the paradigm of edge computing. To fully utilize the widely distributed data, we concentrate on a wireless edge computing system that conducts model training using decentralized peer-to-peer (P2P) methods. However, there are two major challenges on the way towards efficient P2P model training: limited resources (e.g., network bandwidth and battery life of mobile devices) and time-varying network connectivity due to device mobility or wireless channel dynamics, which receives less attention in recent years. To address these two challenges, this paper studies the impact of topology construction on the P2P training performance. Specifically, we dynamically construct an efficient P2P topology, where model aggregation occurs at the edge. In a nutshell, we first formulate the topology construction for P2P learning (TCPL) problem with resource constraints as an integer programming problem. Then a learning-driven method is proposed to adaptively construct a topology at each training epoch. We evaluate the performance of our proposed algorithm through extensive simulations and physical platform. Evaluation results show that our method can improve the model training efficiency by about 11% with resource constraints, reduce the communication cost by 30% and the network traffic consumption by about 60% under the same accuracy requirement compared to the benchmarks. Hongli Xu 0001, Min Chen 0033, Zeyu Meng, Yang Xu 0020, Lun Wang 0003, Chunming Qiao |
IEEE J. Sel. Areas Commun. | 5 |