VLDB 2026 Research / reviewers in the wild / expert
Yunming Liao
dblp:325/9681
· DBLP profile ↗
36ranked-venue papers
6as first author
36since 2021 · last 2026
0000-0002-5065-2600ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 28 · 5 first-author · 28 since 2021Systems, architecture and hardware · 5 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CoSine: Enhancing LLM Serving via Collaborative and Decoupled Speculative Inference
Luyao Gao, Jianchun Liu, Xichong Zhang, Guoju Gao, Yunming Liao |
INFOCOM | 5 |
| 2026 | Optimizing Split Federated Learning through Adaptive Pipeline Parallelism
Zuan Xie, Yang Xu 0020, Yunming Liao |
INFOCOM | 3 |
| 2026 | A Novel Hat-Shaped Device-Cloud Collaborative Inference Framework for Large Language Models
Zuan Xie, Yang Xu 0020, Hongli Xu 0001, Yunming Liao |
INFOCOM | 4 |
| 2026 | Federated fine-tuning on heterogeneous data with alternating device-to-device collaboration
Xitong Fu, Yang Xu 0020, Hongli Xu 0001, Yunming Liao |
Comput. Networks | 5 |
| 2026 | Decentralized federated learning with vertical model parallelism and topology construction
Yang Xu 0020, Zuan Xie, Hongli Xu 0001, Yunming Liao |
Comput. Networks | 4 |
| 2026 | DySTop: Dynamic Staleness Control and Topology Construction for Asynchronous Decentralized Federated LearningabstractFederated Learning (FL) has emerged as a potential distributed learning paradigm that enables model training on edge devices (i.e., workers) while preserving data privacy. However, its reliance on a centralized server leads to limited scalability. Decentralized federated learning (DFL) eliminates the dependency on a centralized server by enabling peer-to peer model exchange. Existing DFL mechanisms mainly employ synchronous communication, which may result in training inefficiencies under heterogeneous and dynamic edge environments. Although a few recent asynchronous DFL (ADFL) mechanisms have been proposed to address these issues, they typically yield stale model aggregation and frequent model transmission, leading to degraded training performance on non-IID data and high communication overhead. To overcome these issues, we present DySTop, an innovative mechanism that jointly optimizes dynamic staleness control and topology construction in ADFL. In each round, multiple workers are activated, and a subset of their neighbors is selected to transmit models for aggregation, followed by local training. We provide a rigorous convergence analysis for DySTop, theoretically revealing the quantitative relationships between the convergence bound and key factors such as maximum staleness, activating frequency, and data distribution among workers. From the insights of the analysis, we propose a worker activation algorithm (WAA) for staleness control and a phase-aware topology construction algorithm (PTCA) to reduce communication overhead and handle data non-IID. Extensive evaluations through both large-scale simulations and real-world testbed experiments demonstrate that our DySTop reduces completion time by 46.7% and the communication resource consumption by 48.3% compared to state-of-the-art solutions, while maintaining the same model accuracy. Yizhou Shi, Qianpiao Ma, Yan Xu 0027, Junlong Zhou, Ming Hu 0003, Yunming Liao, Hongli Xu 0001 |
IEEE Trans. Mob. Comput. | 6 |
| 2026 | Lightweight and Post-Training Structured Pruning for On-Device Large Language ModelsabstractConsidering the hardware-friendly characteristics, structured pruning has emerged as an effective solution to reduce the resource requirements of large language models (LLMs) on resource-constrained devices. Since pruning a certain number of parameters will often reduce the model accuracy, fine-tuning is usually required to recover performance loss. However, fine-tuning relies on high device power and substantial data, making it unsuitable for on-device applications. Recent approaches propose post-training structured pruning techniques that do not require fine-tuning, following three granularities,i.e., neurons, attention heads, and layers. Unfortunately, the previous solutions still suffer from the high memory overhead of substructures' importance evaluation at different granularities and the problem of only applying to models with specific structures, which limit their scope of applications. This paper proposes COMP, a lightweight and general post-training structured hybrid-granularity pruning method designed for on-device applications. COMP begins by assessing the importance of model layers and neurons individually using distributional distance and the matrix condition number. Subsequently, COMP performs fine-grained neuron pruning and simultaneously determines whether to apply coarse-grained layer pruning or not by comparing the outcomes of the two granularity strategies. Furthermore, COMP implements mask tuning to restore model accuracy without additional fine-tuning, minimizing dependency on external data resources and device power. Experimental results demonstrate that COMP performs well across models with various architectures. When pruning 20% of LLaMA-2-13B parameters, COMP requires less than 8 GB of memory, maintaining more than 90% model accuracy. Meanwhile, COMP improves performance by 3.54% compared to LLM-Pruner, while reducing memory overhead by 90%. Zihuai Xu, Yang Xu 0020, Hongli Xu 0001, Yunming Liao, Zuan Xie |
IEEE Trans. Mob. Comput. | 4 |
| 2026 | Heterogeneity-Aware Semi-Supervised Federated Learning With Client-to-Client CollaborationabstractFederated learning (FL) enables multiple clients to collaboratively train models without exposing raw data. However, most FL research assumes that all clients have fully labeled data, which is impractical for many real-world applications. To this end, we focus on semi-supervised FL (SSFL), where data samples of each client are partially labeled. Existing SSFL methods often ignore three inherent characteristics of FL: limited communication resources, data heterogeneity and system heterogeneity, which severely hinder convergence stability and efficiency. This paper proposes E-FedAC, a novel SSFL framework that tackles these challenges through alternating client-to-client (C2C) collaboration. Specifically, E-FedAC employs a two-stage clustering strategy in each training round while concurrently considering differences in client computing/communication capabilities to mitigate the impact of intra-cluster stragglers. The first stage employs similarity clustering to group clients based on local data distribution, which allows similar clients to share knowledge and generate high-quality pseudo-labels for unlabeled data. In the second stage, clients are re-grouped using a dissimilarity clustering strategy to approximate the IID setting at the cluster level, thereby alleviating the bias induced by Non-IID data. E-FedAC utilizes a reinforcement learning algorithm to adaptively determine the average number of intra-cluster iterations across all clusters for both clustering stages, balancing between labeling assistance from similar clients and unbiased optimization from dissimilar clients. Additionally, E-FedAC dynamically optimizes local updating frequencies for each generated cluster based on this average and individual cluster capabilities, further reducing the waiting time across clusters. Extensive evaluations demonstrate that E-FedAC can provide up to 4.02× speedup and 61.74% savings in communication costs compared with existing benchmarks. Yang Xu 0020, Hongli Xu 0001, Yunming Liao, Xitong Fu |
IEEE Trans. Mob. Comput. | 4 |
| 2026 | Heterogeneity-Aware Federated Meta Learning for Personalized Edge DevicesabstractRecently, aided by edge computing (EC), federated meta-learning (FML) has been proposed to adequately utilize personalized data on different edge devices by incorporating meta-learning with federated learning (FL). Unlike conventional FL which conducts several stochastic gradient descent (SGD) steps to directly update the local model in each training round, FML first performs some SGDsteps to obtain a reference model and then performs a meta-step to update the local model along the direction towards the reference model. In this way, FML can identify common patterns and structures across different clients (i.e., edge devices) and quickly adapt to new data distributions, which exhibits its advantage in dealing with statistical heterogeneity inherent in FL. However, existing FML approaches struggle to effectively cope with two key challenges: i) the disparity in dataset informativeness among clients with statistically heterogeneous local data; and ii) the disparity in computing capacities among clients with heterogeneous hardware configurations. Herein, we propose a heterogeneity-aware FML framework, termed HFML, which explores adjusting the execution frequencies of SGD step and meta-step in each training round to address the above statistical and system heterogeneity challenges simultaneously. We introduce usable information (UI) as a new metric in HFML to quantify dataset informativeness, and theoretically analyze the influence of the execution frequencies of SGD steps and meta-steps on model performance and convergence speed. Based on the theoretical analysis, we develop an efficient optimization algorithm to jointly adjust the execution frequencies for different clients to handle the heterogeneity issues, which contributes to accelerating model training and enhancing personalization performance. Experimental results demonstrate that HFML can speed up training by up to 4.02× and improve model accuracy by an average of 11.5% compared to the baselines under heterogeneous settings. Yang Xu 0020, Hongli Xu 0001, Yunming Liao, Liusheng Huang |
IEEE Trans. Mob. Comput. | 4 |
| 2026 | Asynchronous Federated Learning Over Non-IID Data via Over-the-Air ComputationabstractFederated learning (FL) enables training AI models across distributed edge devices (i.e., workers) using local data, while facing challenges including communication resource constraints, edge heterogeneity, and non-IID data. Over-the-air computation (AirComp) has emerged as a promising technique to improve communication efficiency by leveraging the superposition property of a wireless multiple access channel (MAC) for model aggregation. However, over-the-air aggregation requires strict synchronization among edge devices, which is essentially incompatible with the asynchronous FL mechanisms often used to handle edge heterogeneity. To overcome this incompatibility, we propose Air-FedGA, a grouping-based asynchronous FL mechanism via AirComp, where workers are organized into groups for synchronized over-the-air aggregation within each group, while groups asynchronously communicate with the parameter server to update the global model. This design retains the communication efficiency of AirComp while addressing training inefficiency caused by edge heterogeneity. We provide a rigorous convergence analysis for Air-FedGA, theoretically quantifying how the convergence bound depends on several key factors, such as the maximum staleness, the degree of non-IID data among groups, and the AirComp aggregation mean squared error (MSE). Guided by these theoretical insights, we propose power control and worker grouping algorithms to minimize the convergence bound by jointly optimizing the AirComp aggregation MSE and the grouping strategy. We conduct experiments on classical models and datasets, and the results demonstrate that our proposed mechanism and algorithms can accelerate the model training by 1.83-$2.22\times $compared with the state-of-the-art solutions. Qianpiao Ma, Xiaozhu Song, Junlong Zhou, Haibo Wang 0004, Yunming Liao, Jianchun Liu, Hongli Xu 0001 |
IEEE Trans. Netw. | 5 |
| 2026 | SemiSFL: A Novel Semi-Supervised Split Federated Learning System for Efficient Training With Non-IID DataabstractRecently, Split Federated Learning (SFL) advances typical Federated Learning (FL) and offers a feasible solution by splitting a full model into a bottom model (for clients) and a top model (for the server), which alleviates the computation and/or communication burden on clients. However, existing SFL works often assume sufficient labeled data on clients, which is usually impractical. Besides, non-IID data poses another challenge to ensure efficient model training. To our best knowledge, the above two data-related issues have not been simultaneously addressed in SFL. Herein, we review the distinct properties of SFL and develop a novel Semi-supervised SFL system, called SemiSFL. In SemiSFL, a student bottom model and a teacher bottom model are assigned to each client, while the server maintains the top parts and full versions of both the global and teacher models. Cross-entity (i.e., clients and the server) semi-supervised training and server-side supervised training are alternately conducted. During semi-supervised training, we explore regularizing the student features (produced by student bottom models) from different clients with the teacher features (produced by teacher bottom models) before conducting semi-supervised training, so as to improve the generalization ability of student bottom models. Moreover, our theoretical and experimental investigations reveal that the inconsistent training processes on labeled and unlabeled data have an influence on the effectiveness of SemiSFL. To mitigate training inconsistency, we develop an algorithm for dynamically regulating the frequency of performing server-side supervised training. Extensive experiments on benchmark models and datasets show that SemiSFL provides a 3.8× speed-up in training time, reduces the communication cost by about 70.3% while reaching the target accuracy, and achieves up to 5.8% improvement in accuracy under different non-IID settings compared to the state-of-the-art baselines. Yang Xu 0020, Hongli Xu 0001, Yunming Liao |
IEEE Trans. Netw. | 4 |
| 2025 | MPLS: Stacking Diverse Layers Into One Model for Decentralized Federated Learning
Yang Xu 0020, Hongli Xu 0001, Yunming Liao, Zuan Xie |
Euro-Par (1) | 4 |
| 2025 | Federated Fine-Tuning on Heterogeneous Devices with Adaptive Quantization and LoRA DepthsabstractFederated fine-tuning (FedFT) has become a decentralized approach for fine-tuning pre-trained large language models(LLMs). However, due to the immense size of LLMs, there are two critical challenges for efficient FedFT in practice, i.e., resource constraints and system heterogeneity. Though existing approaches employ low-rank adaptation (LoRA) method, to mitigate fine-tuning overhead, they still suffer from high memory consumption and large training latency. According to the characteristics of FedFT, we observe that appropriate quantization and LoRA configurations can reduce resource consumption while maintaining comparable model performance. Consequently, we propose an efficient FedFT framework, termed FedQLoRA, which dynamically adjusts the quantization level and LoRA depth during training. This design enables large models to be fine-tuned on resource-limited devices while reducing training latency and preserving model performance. Extensive experimental results demonstrate that FedQLoRA achieves up to 70.34% reduction in training time and$62.71 \%-77.50 \%$savings in memory consumption compared to baseline methods. Qianshu Wang, Yang Xu 0020, Hongli Xu 0001, Liusheng Huang, Yunming Liao, Jun Liu 0083 |
ICPADS | 5 |
| 2025 | Towards layer-wise quantization for heterogeneous federated clients
Yang Xu 0020, Hongli Xu 0001, Changyu Guo, Yunming Liao |
Comput. Networks | 5 |
| 2025 | LOGO-CL: Accelerating semi-supervised federated learning in edge computing
Yang Xu 0020, Qianshu Wang, Hongli Xu 0001, Yunming Liao, Liusheng Huang, Xin Hang |
Comput. Networks | 4 |
| 2025 | Enhancing Split Federated Learning With Worker Clustering and Feature Compression
Yang Xu 0020, Yunming Liao, Hongli Xu 0001, Chunming Qiao |
IEEE J. Sel. Areas Commun. | 2 |
| 2025 | Enhancing Semi-Supervised Federated Learning With Progressive Training in Heterogeneous Edge ComputingabstractFederated learning (FL) is an efficient distributed learning method that facilitates collaborative model training among multiple edge devices (or clients). However, current research always assumes that clients have access to ground-truth data for training, which is unrealistic in practice because of a lack of expertise. Semi-supervised federated learning (SSFL) has been proposed in many existing works to address this problem, which always adopts a fixed model architecture for training, bringing two main problems with varying amounts of pseudo-labeled data. First, the shallow model cannot have the capability to fit the increasing pseudo-labeled data, leading to poor training performance. Second, the large model suffers from an overfitting problem when exploiting a few labeled data samples in SSFL, and also requires tremendous resource (e.g., computation and communication) costs. To tackle these problems, we propose a novel framework, calledstar, which adopts progressive training to enhance model training in SSFL. Specifically,stargradually increases the model depth through adding the sub-module (e.g., one or several layers) from a shallow model, and performs pseudo-labeling for unlabeled data with a specialized confidence threshold simultaneously. Then, we propose an efficient algorithm to determine the appropriate model depth for each client with varied resource budgets and the proper confidence threshold for pseudo-labeling in SSFL. The experimental results demonstrate the high effectiveness of STAR. For instance,starcan reduce the bandwidth consumption by about 40%, and achieve an average accuracy improvement of around 9.8% compared with the baselines, on CIFAR10. Jianchun Liu, Jun Liu 0083, Hongli Xu 0001, Yunming Liao, Min Chen 0033, Chen Qian 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2025 | Adaptive Parameter-Efficient Federated Fine-Tuning on Heterogeneous DevicesabstractFederated fine-tuning (FedFT) has been proposed to fine-tune the pre-trained language models in a distributed manner. However, there are two critical challenges for efficient FedFT in practical applications,i.e., resource constraints and system heterogeneity. Existing works rely on parameter-efficient fine-tuning methods,e.g., low-rank adaptation (LoRA), but with major limitations. Herein, based on the inherent characteristics of FedFT, we observe that LoRA layers with higher ranks added close to the output help to save resource consumption while achieving comparable fine-tuning performance. Then we propose a novel LoRA-based FedFT framework, termed LEGEND, which faces the difficulty of determining the number of LoRA layers (called, LoRA depth) and the rank of each LoRA layer (called, rank distribution). We analyze the coupled relationship between LoRA depth and rank distribution, and design an efficient LoRA configuration algorithm for heterogeneous devices, thereby promoting fine-tuning efficiency. Extensive experiments are conducted on a physical platform with 80 commercial devices. The results show that LEGEND can achieve a speedup of 1.5-2.8× and save communication costs by about 42.3% when achieving the target accuracy, compared to the advanced solutions. Jun Liu 0083, Yunming Liao, Hongli Xu 0001, Yang Xu 0020, Jianchun Liu, Chen Qian 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2025 | FedSNN: Training Slimmable Neural Network With Federated Learning in Edge ComputingabstractTo provide a flexible tradeoff between inference accuracy and resource requirement at runtime, the slimmable neural network (SNN), a single network executable at different widths with the same deploying and management cost as that of a single model, has been proposed. However, how to effectively train SNN among massive devices in edge computing without revealing their local data remains an open problem. To this end, we leverage a novel distributed machine learning paradigm, i.e., federated learning, to realize effective on-device SNN training. As current FL schemes often train only one model with fixed architecture, and the existing SNN training algorithm is resource-intensive, integrating FL and SNN is non-trivial. Furthermore, two intrinsic features in edge computing, i.e., data and system heterogeneity, exacerbate the difficulty. Motivated by this, we redesign the model distribution, local training, and model aggregation phases in traditional FL, and propose FedSNN, a framework that ensures all widths in SNN can obtain high accuracy with less resource consumption. Specifically, for devices with heterogeneous training capacities and data distributions, the parameter server will distribute each of them with one proper width for adaptive local training guided by their uploaded model features, and their trained models will be weighted-averaged using the proposed multi-width SNN aggregation to improve their statistical utility. Extensive experiments on a distributed testbed show that FedSNN improves the model accuracy by about 2.18%-8.1%, and accelerates training by about$1.31\times $-$6.84\times $, compared with existing solutions. Yang Xu 0020, Yunming Liao, Hongli Xu 0001, Zhiyuan Wang 0002, Lun Wang 0003, Jianchun Liu, Chen Qian 0001 |
IEEE Trans. Netw. | 2 |
| 2025 | PairingFL: Efficient Federated Learning With Model Splitting and Client PairingabstractFederated learning (FL) has recently gained tremendous attention in edge computing and the Internet of Things, due to its capability of enabling clients to perform model training at the network edge or end devices (i.e., clients). However, these end devices are usually resource-constrained without the ability to train large-scale models. In order to accelerate the training of large-scale models on these devices, we incorporate Split Learning (SL) into Federated Learning (FL), and propose a novel FL framework, termedPairingFL. Specifically, we split a full model into a bottom model and a top model, and arrange participating clients into pairs, each of which collaboratively trains the two partial models as one client does in typical FL. Driven by the advantages of SL and FL, PairingFL is able to relax the computation burden on clients and protect model privacy. However, considering the features of system and statistical heterogeneity in edge networks, it is challenging to pair the clients by carefully developing the strategies of client partitioning and matching for efficient model training. To this end, we first theoretically analyze the convergence property of PairingFL, and obtain a convergence upper bound. Guided by this, we then design a greedy and efficient algorithm, which makes the joint decision of client partitioning and matching, so as to well balance the trade-off between convergence rate and model accuracy. The performance of PairingFL is evaluated through extensive simulation experiments. The experimental results demonstrate that PairingFL can speed up the training process by$4.6\times $compared to baselines when achieving the corresponding convergence accuracy. Ji Qi 0005, Yang Xu 0020, Yunming Liao, Hongli Xu 0001, Lun Wang 0003 |
IEEE Trans. Netw. | 4 |
| 2025 | Enhancing Federated Learning Through Layer-Wise Aggregation Over Non-IID DataabstractNowadays, federated learning (FL) has been widely adopted to train deep neural networks (DNNs) among massive devices without revealing their local data in edge computing (EC). To relieve the communication bottleneck of the central server in FL, hierarchical federated learning (HFL), which leverages edge servers as intermediaries to perform model aggregation among devices in proximity, comes into being. Nevertheless, the existing HFL systems may not perform training effectively due to bandwidth constraints and non-IID issues on devices. To conquer these challenges, we introduce anHFL system with device-edgeassignment andlayer selection, namely Heal. Specifically, Heal organizes all the devices into a hierarchical structure (i.e., device-edge assignment) and enables each device to forward only a sub-model with several valuable layers for aggregation (i.e., layer selection). This processing procedure is called layer-wise aggregation. To further save communication resource and improve the convergence performance, we then design an iteration-based algorithm to optimize the development of our layer-wise aggregation strategy by considering the data distribution as well as resource constraints among devices. Extensive experiments on both the physical platform and the simulated environment show that Heal accelerates DNN training by about 1.4–12.5×, and reduces the network traffic consumption by about 31.9–64.1%, compared with the existing HFL systems. Yang Xu 0020, Zhiyuan Wang 0002, Hongli Xu 0001, Yunming Liao |
IEEE Trans. Serv. Comput. | 5 |
| 2024 | Federated Semi-Supervised Learning with Local and Global Updating Frequency OptimizationabstractFederated learning (FL) has gained Important attention for training deep learning models across clients in edge computing. The scarcity of labeled data poses critical challenges in professional fields (e.g., medical diagnosis) for FL, motivating the development of federated semi-supervised learning (FSSL) to exploit unlabeled data. The existing works of FSSL always assign fixed values for global updating frequency and/or local updating frequency during training. Though the previous works can make full use of unlabeled data, they may still hinder the improvement of model performance on heterogeneous clients, due to that the two hyper-parameters are coupled and have a joint impact on training process. To this end, we propose a novel FSSL framework, named LOGO, to enhance training efficiency by jointly optimizing local and global updating frequencies. We analyze the coupled impact of these two hyper-parameters on model training. Then, we develop a multi-armed bandit (MAB) based online algorithm to adaptively determine diverse local updating frequencies for heterogeneous clients as well as appropriate global updating frequency, so as to address system heterogeneity and improve training efficiency. The performance of LOGO is evaluated through extensive simulation experiments. The experimental results demonstrate that LOGO achieves the model training speedup by 2.3 × and reduces communication cost by 44%, compared to the baselines. Xin Hang, Yang Xu 0020, Hongli Xu 0001, Yunming Liao, Lun Wang 0003 |
CCGrid | 4 |
| 2024 | MergeSFL: Split Federated Learning with Feature Merging and Batch Size RegulationabstractRecently, federated learning (FL) has emerged as a popular technique for edge AI to mine valuable knowledge in edge computing (EC) systems. To boost the performance of AI applications, large-scale models have received increasing attention due to their excellent generalized abilities. However, training and transmitting large-scale models will incur significant computing and communication burden on the resource-constrained workers, and the exchange of entire models may violate model privacy. To relax the burden of workers and protect model privacy, split federated learning (SFL) has been released by integrating both data and model parallelism. Despite resource limitations, SFL also faces two other critical challenges in EC systems, i.e., statistical heterogeneity and system heterogeneity. In order to address these challenges, we propose a novel SFL framework, termed MergeSFL, by incorporating feature merging and batch size regulation in SFL. Concretely, feature merging aims to merge the features from workers into a mixed feature sequence, which is approximately equivalent to the features derived from IID data and is employed to promote model accuracy. While batch size regulation aims to assign diverse and suitable batch sizes for heterogeneous workers to improve training efficiency. Moreover, MergeSFL explores to jointly optimize these two strategies upon their coupled relationship to better enhance the performance of SFL. Extensive experiments are conducted on a physical platform with 80 NVIDIA Jetson edge devices, and the experimental results show that MergeSFL can improve the final model accuracy by 5.82% to 26.22%, with a speedup by about 1.39x to 4.14x, compared to the baselines. Yunming Liao, Yang Xu 0020, Hongli Xu 0001, Lun Wang 0003, Chunming Qiao |
ICDE | 1 |
| 2024 | Accelerating Hierarchical Federated Learning with Model Splitting in Edge ComputingabstractRecently, Hierarchical Federated Learning (HFL) stands out as a cutting-edge approach to efficiently learn knowledge from massive data on edge devices or clients. To alleviate the computation/communication burden of training large-scale models on resource-constrained clients, Hierarchical Split Federated Learning (HSFL), which splits an entire model into a top and bottom (sub-)model and offloads the training of the top model to the edge server, has been proposed. Nonetheless, there are two key issues, i.e., system heterogeneity and dynamic contexts, hindering the application of effective HSFL. In response, we present an efficient HSFL method, termed AdaHSFL, which introduces intra-cluster gradient feedback regulation and inter-cluster updating frequencies optimization to enhance training efficiency. Concretely, intra-cluster gradient feedback regulation enables edge servers to instantly process incoming smashed data and calculate gradients using top model copies on background threads, while inter-cluster updating frequencies optimization adjusts edge cluster updating frequencies to align the training duration across all clusters with that of the fastest cluster. Furthermore, AdaHSFL explores to jointly implement these two strategies to eliminate the idle waiting time both intra- and intercluster incurred by synchronization barriers. Rigorous performance evaluation demonstrates that AdaHSFL can improve the accuracy by $3.9 \%-25.5 \%$ within given time budgets, compared with the baselines. Xiangnan Wang, Yang Xu 0020, Hongli Xu 0001, Yunming Liao, Ji Qi 0005 |
ICPADS | 5 |
| 2024 | ParallelSFL: A Novel Split Federated Learning Framework Tackling Heterogeneity IssuesabstractMobile devices contribute more than half of the world's web traffic, providing massive and diverse data for powering various federated learning (FL) applications. In order to avoid the communication bottleneck on the parameter server (PS) and accelerate the training of large-scale models on resource-constraint workers in edge computing (EC) system, we propose a novel split federated learning (SFL) framework, termed ParallelSFL. Concretely, we split an entire model into a bottom submodel and a top submodel, and divide participating workers into multiple clusters, each of which collaboratively performs the SFL training procedure and exchanges entire models with the PS. However, considering the statistical and system heterogeneity in edge systems, it is challenging to arrange suitable workers to specific clusters for efficient model training. To address these challenges, we carefully develop an effective clustering strategy by optimizing a utility function related to training efficiency and model accuracy. Specifically, ParallelSFL partitions workers into different clusters under the heterogeneity restrictions, thereby promoting model accuracy as well as training efficiency. Meanwhile, ParallelSFL assigns diverse and appropriate local updating frequencies for each cluster to further address system heterogeneity. Extensive experiments are conducted on a physical platform with 80 NVIDIA Jetson devices, and the experimental results show that ParallelSFL can reduce the traffic consumption by at least 21%, speed up the model training by at least 1.36X, and improve model accuracy by at least 5% in heterogeneous scenarios, compared to the baselines. Yunming Liao, Yang Xu 0020, Hongli Xu 0001, Liusheng Huang, Chunming Qiao |
MobiCom | 1 |
| 2024 | Decentralized Federated Learning With Adaptive Configuration for Heterogeneous ParticipantsabstractData generated at the network edge can be processed locally by leveraging the paradigm of edge computing (EC). Aided by EC, decentralized federated learning (DFL), which overcomes the single-point-of-failure problem in the parameter server based federated learning, is becoming a practical and popular approach for machine learning over distributed data. However, DFL faces two critical challenges,i.e., system heterogeneity and statistical heterogeneity introduced by edge devices. To ensure fast convergence with the existence of slow edge devices, we present an efficient DFL method, termed FedHP, which integrates adaptive control of both local updating frequency and network topology to better support the heterogeneous participants. We establish a theoretical relationship between local updating frequency and network topology regarding model training performance and obtain a convergence upper bound. Upon the convergence bound, we propose an optimization algorithm that adaptively determines local updating frequencies and constructs the network topology, so as to speed up convergence and improve the model accuracy. We evaluate the performance of FedHP through extensive simulation and testbed experiments. Evaluation results show that the proposed FedHP can reduce the completion time by about 51% and improve model accuracy by at least 5% in heterogeneous scenarios, compared with the baselines. Yunming Liao, Yang Xu 0020, Hongli Xu 0001, Lun Wang 0003, Chen Qian 0001, Chunming Qiao |
IEEE Trans. Mob. Comput. | 1 |
| 2024 | Overcoming Noisy Labels and Non-IID Data in Edge Federated LearningabstractFederated learning (FL) enables edge devices to cooperatively train models without exposing their raw data. However, implementing a practical FL system at the network edge mainly faces three challenges: label noise, data non-IIDness, and device heterogeneity, which seriously harm model performance and slow down convergence speed. Unfortunately, none of the existing works tackle all three challenges simultaneously. To this end, we develop a novel FL system, called Aorta, which features adaptive dataset construction and aggregation weightassignment. On each client, Aorta first calibrates potentially noisy labels and then constructs a training dataset with low noise, balanced distribution, and proper size. To fully utilize limited data on clients, we propose a global model guided method to select clean data and progressively correct noisy labels. To achieve balanced class distribution and proper dataset size, we propose a distribution-and-capability-aware data augmentation method to generate local training data. On the server, Aorta assigns aggregation weights based on the quality of local models to ensure that high-quality models have a greater influence on the global model. The model quality is measured through its cosine similarity with a benchmark model, which is trained on a clean and balanced dataset. We conduct extensive experiments on four datasets with various settings, including different noise types/ratios and non-IID types/levels. Compared to the baselines, Aorta improves model accuracy up to 9.8% on the datasets with moderate noise and non-IIDness, while providing a speedup of 4.2× on average when achieving the same target accuracy. Yang Xu 0020, Yunming Liao, Lun Wang 0003, Hongli Xu 0001, Zhida Jiang, Wuyang Zhang |
IEEE Trans. Mob. Comput. | 2 |
| 2024 | Ferrari: A Personalized Federated Learning Framework for Heterogeneous Edge ClientsabstractFederated semi-supervised learning (FSSL) has been proposed to address the insufficient labeled data problem by training models with pseudo-labeling. In previous FSSL systems, a single global model is always trained without an equivalent generalization ability for the clients under the non-IID setting. Accordingly, model personalization methods have been proposed to overcome this problem. Intuitively, seeking labeling assistance from other clients with similar data distribution,i.e., model migration, can effectively improve the personalization on the clients with scarce labeled data. However, previous works require to migrate a pre-fixed number of models among the clients, causing unnecessary resource waste and accuracy degradation due to resource heterogeneity. Considering that the number of model migrations and the quality of pseudo-labels have a significant impact on the training performance (e.g., efficiency and accuracy), we propose a novel personalized FSSL system, called Ferrari, to boost the efficiency of pseudo-labeling and training accuracy through adaptive model migrations among the clients. Specifically, Ferrari first generates the similarity-based ranking using a Gaussian KD-Tree, considering the varied data distributions among the clients. Combined with the ranking and clients' heterogeneous resource constraints, Ferrari then adaptively determines the proper model migration policy and confidence thresholds for high-quality pseudo-labeling and personalized training for clients. Extensive experiments on a physical platform show that Ferrari provides a 1.2$\sim 5.5\times$speedup without sacrificing model accuracy, compared to existing methods. Jianchun Liu, Hongli Xu 0001, Lun Wang 0003, Chen Qian 0001, Yunming Liao |
IEEE Trans. Mob. Comput. | 6 |
| 2024 | Asynchronous Decentralized Federated Learning for Heterogeneous DevicesabstractData generated at the network edge can be processed locally by leveraging the emerging technology of Federated Learning (FL). However, non-IID local data will lead to degradation of model accuracy and the heterogeneity of edge nodes inevitably slows down model training efficiency. Moreover, to avoid the potential communication bottleneck in the parameter-server-based FL, we concentrate on the Decentralized Federated Learning (DFL) that performs distributed model training in Peer-to-Peer (P2P) manner. To address these challenges, we propose an asynchronous DFL system by incorporating neighbor selection and gradient push, termed AsyDFL. Specifically, we require each edge node to push gradients only to a subset of neighbors for resource efficiency. Herein, we first give a theoretical convergence analysis of AsyDFL under the complicated non-IID and heterogeneous scenario, and further design a priority-based algorithm to dynamically select neighbors for each edge node so as to achieve the trade-off between communication cost and model performance. We evaluate the performance of AsyDFL through extensive experiments on a physical platform with 30 NVIDIA Jetson edge devices. Evaluation results show that AsyDFL can reduce the communication cost by 57% and the completion time by about 35% for achieving the same test accuracy, and improve model accuracy by at least 6% under the non-IID scenario, compared to the baselines. Yunming Liao, Yang Xu 0020, Hongli Xu 0001, Min Chen 0033, Lun Wang 0003, Chunming Qiao |
IEEE/ACM Trans. Netw. | 1 |
| 2024 | Accelerating Federated Learning With Data and Model Parallelism in Edge ComputingabstractRecently, edge AI has been launched to mine and discover valuable knowledge at network edge. Federated Learning, as an emerging technique for edge AI, has been widely deployed to collaboratively train models on many end devices in data-parallel fashion. To alleviate the computation/communication burden on the resource-constrained workers (e.g., end devices) and protect user privacy, Spilt Federated Learning (SFL), which integrates both data parallelism and model parallelism in Edge Computing (EC), is becoming a practical and popular approach for model training over distributed data. However, apart from the resource limitation, SFL still faces two other critical challenges in EC, i.e., system heterogeneity and context dynamics. To overcome these challenges, we present an efficient SFL method, named AdaSFL, which controls both local updating frequency and batch size to better accelerate model training. We theoretically analyze the model convergence rate and obtain a convergence upper bound regarding local updating frequency given a fixed batch size. Upon this, we develop a control algorithm to determine adaptive local updating frequency and diverse batch sizes for heterogeneous workers to enhance the training efficiency. The experimental results show that AdaSFL can reduce the completion time by about 43% and the network traffic consumption by about 31% for achieving the similar test accuracy, compared to the baselines. Yunming Liao, Yang Xu 0020, Hongli Xu 0001, Lun Wang 0003, Chunming Qiao |
IEEE/ACM Trans. Netw. | 1 |
| 2024 | YOGA: Adaptive Layer-Wise Model Aggregation for Decentralized Federated LearningabstractTraditional Federated Learning (FL) is a promising paradigm that enables massive edge clients to collaboratively train deep neural network (DNN) models without exposing raw data to the parameter server (PS). To avoid the bottleneck on the PS, Decentralized Federated Learning (DFL), which utilizes peer-to-peer (P2P) communication without maintaining a global model, has been proposed. Nevertheless, DFL still faces two critical challenges, i.e., limited communication bandwidth and not independent and identically distributed (non-IID) local data, thus hindering efficient model training. Existing works commonly assume full model aggregation at periodic intervals, i.e., clients periodically collect models from peers. To reduce the communication cost, these methods allow clients to collect model(s) from selected peers, but often result in a significant degradation of model accuracy when dealing with non-IID data. Alternatively, the layer-wise aggregation mechanism has been proposed to alleviate communication overhead under the PS architecture, but its potential in DFL remains rarely explored yet. To this end, we propose an efficient DFL framework YOGA that adaptively performs layer-wise model aggregation and training. Specifically, YOGA first generates the ranking of layers in the model according to the learning speed and layer-wise divergence. Combining with the layer ranking and peers’ status information (i.e., data distribution and communication capability), we propose the max-match (MM) algorithm to generate the proper layer-wise model aggregation policy for the clients. Extensive experiments on DNN models and datasets show that YOGA saves communication cost by about 45% without sacrificing the model performance compared with the baselines, and provides 1.53-$3.5\times $speedup on the physical platform. Jun Liu 0083, Jianchun Liu, Hongli Xu 0001, Yunming Liao, Zhiyuan Wang 0002, Qianpiao Ma |
IEEE/ACM Trans. Netw. | 4 |
| 2024 | Federated Learning With Experience-Driven Model Migration in Heterogeneous Edge NetworksabstractTo approach the challenges of non-IID data and limited communication resource raised by the emerging federated learning (FL) in mobile edge computing (MEC), we propose an efficient framework, calledFedMigr, which integrates a deep reinforcement learning (DRL) based model migration strategy into the pioneer FL algorithmFedAvg. According to the data distribution and resource budgets, ourFedMigrwill intelligently guide one client to forward its local model to another client after local updating, before directly sending the local models to the server for global aggregation as inFedAvg. Intuitively, migrating a local model from one client to another is equivalent to training the model over more data from different clients, alleviating the influence of non-IID issue. To this end, we propose an experience-driven method to make proper decisions for model migrations while satisfying the resource constraints. We also prove thatFedMigrcan help to reduce the parameter divergences between different local models and the global model from a theoretical perspective under the non-IID setting. Extensive experiments on three popular benchmark datasets demonstrate thatFedMigrcan achieve an average accuracy improvement of around 13%, and reduce bandwidth consumption for global communication by 42% on average, compared with the baselines. Jianchun Liu, Shilong Wang 0002, Hongli Xu 0001, Yang Xu 0020, Yunming Liao, Jinyang Huang, He Huang 0001 |
IEEE/ACM Trans. Netw. | 5 |
| 2023 | Adaptive Configuration for Heterogeneous Participants in Decentralized Federated LearningabstractData generated at the network edge can be processed locally by leveraging the paradigm of edge computing (EC). Aided by EC, decentralized federated learning (DFL), which overcomes the single-point-of-failure problem in the parameter server (PS) based federated learning, is becoming a practical and popular approach for machine learning over distributed data. However, DFL faces two critical challenges, i.e., system heterogeneity and statistical heterogeneity introduced by edge devices. To ensure fast convergence with the existence of slow edge devices, we present an efficient DFL method, termed FedHP, which integrates adaptive control of both local updating frequency and network topology to better support the heterogeneous participants. We establish a theoretical relationship between local updating frequency and network topology regarding model training performance and obtain a convergence upper bound. Upon this, we propose an optimization algorithm, that adaptively determines local updating frequencies and constructs the network topology, so as to speed up convergence and improve the model accuracy. Evaluation results show that the proposed FedHP can reduce the completion time by about 51% and improve model accuracy by at least 5% in heterogeneous scenarios, compared with the baselines. Yunming Liao, Yang Xu 0020, Hongli Xu 0001, Lun Wang 0003, Chen Qian 0001 |
INFOCOM | 1 |
| 2023 | Adaptive Control of Local Updating and Model Compression for Efficient Federated LearningabstractData generated at the network edge can be processed locally by leveraging the paradigm of Edge Computing (EC). Aided by EC, Federated Learning (FL) has been becoming a practical and popular approach for distributed machine learning over locally distributed data. However, FL faces three critical challenges, i.e., resource constraint, system heterogeneity and context dynamics in EC. To address these challenges, we present a training-efficient FL method, termedFedLamp, by optimizing both theLocal updating frequency andmodel compression ratio in the resource-constrained EC systems. We theoretically analyze the model convergence rate and obtain a convergence upper bound related to the local updating frequency and model compression ratio. Upon the convergence bound, we propose a control algorithm, that adaptively determines diverse and appropriate local updating frequencies and model compression ratios for different edge nodes, so as to reduce the waiting time and enhance the training efficiency. We evaluate the performance ofFedLampthrough extensive simulation and testbed experiments. Evaluation results show thatFedLampcan reduce the traffic consumption by 63% and the completion time by about 52% for achieving the similar test accuracy, compared to the baselines. Yang Xu 0020, Yunming Liao, Hongli Xu 0001, Zhen-guo Ma, Lun Wang 0003, Jianchun Liu |
IEEE Trans. Mob. Comput. | 2 |
| 2022 | Enhancing Federated Learning with Intelligent Model Migration in Heterogeneous Edge ComputingabstractTo approach the challenges of non-IID data and limited communication resource raised by the emerging federated learning (FL) in mobile edge computing (MEC), we propose an efficient framework, called FedMigr, which integrates a deep reinforcement learning (DRL) based model migration strategy into the pioneer FL algorithm FedAvg. According to the data distribution and resource constraints, our FedMigr will intelligently guide one client to forward its local model to another client after local updating, rather than directly sending the local models to the server for global aggregation as in FedAvg. Intuitively, migrating a local model from one client to another is equivalent to training it over more data from different clients, contributing to alleviating the influence of non-IID issue. We prove that FedMigr can help to reduce the parameter divergences between different local models and the global model from a theoretical perspective, even over local datasets with non-IID settings. Extensive experiments on three popular benchmark datasets demonstrate that FedMigr can achieve an average accuracy improvement of around 13%, and reduce bandwidth consumption for global communication by 42% on average, compared with the baselines. Jianchun Liu, Yang Xu 0020, Hongli Xu 0001, Yunming Liao, Zhiyuan Wang 0002, He Huang 0001 |
ICDE | 4 |
| 2022 | Decentralized Federated Learning with Data Feature Transmission and Neighbor SelectionabstractIn edge computing (EC), federated learning (FL) has been widely used to cooperatively train machine learning models and protect privacy over the traditional approaches. In order to avoid the possible communication bottleneck and single point failure in parameter server (PS) framework, we mainly focus on decentralized federated learning (DFL) in which each worker disseminates information through peer-to-peer (P2P) communication network. However, the high communication cost and the non-IID local data hinder the efficient training of DFL. In this paper, we propose a mechanism (termed DFL-DF) to achieve efficient communication and improve the performance on non-IID data. Specifically, each worker will exchange data features with neighbors rather than models (or gradients). Different from the traditional fixed topology methods, each worker will communicate with its neighbors selected at each round. Extensive experiment results show the high effectiveness of DFL-DF. Specifically, DFL-DF can reduce communication cost by 29.4%-82.06% and improve test accuracy 2.5%-9.1% under communication cost constraints, compared with the traditional model exchanging methods. Wenxiao Lou, Yang Xu 0020, Hongli Xu 0001, Yunming Liao |
ICPADS | 4 |