Jianchun Liu

dblp:250/0421 · DBLP profile ↗
← Back
63ranked-venue papers
13as first author
59since 2021 · last 2026
0000-0002-1764-9303ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 51 · 12 first-author · 47 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Security and privacy · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CoSine: Enhancing LLM Serving via Collaborative and Decoupled Speculative Inference
Luyao Gao, Jianchun Liu, Xichong Zhang, Guoju Gao, Yunming Liao
INFOCOM2
2026 FLUDE: An Efficient Federated Learning Framework with Undependable Devices
Shilong Wang 0002, Jianchun Liu, Hongli Xu 0001, Chunming Qiao
INFOCOM2
2026 AdaMoA: Enhancing Mixture-of-Agents via Task-Adaptive Architecture Optimization
Wenhui Lv, Hongli Xu 0001, Jianchun Liu, Zihuai Xu
IWQoS4
2026 Caesar: Optimizing Federated Learning via Low-deviation Compression
abstract
Compression is an efficient way to relieve the tremendous communication overhead of federated learning (FL) systems. However, for the existing works, the information loss under compression will lead to unexpected model/gradient deviation for the FL training, significantly degrading the training performance, especially under the challenges of data heterogeneity and model obsolescence. To strike a delicate trade-off between model accuracy and traffic cost, we propose Caesar, a novel FL framework with a low-deviation compression approach. For the global model download, we design a greedy method to optimize the compression ratio for each device based on the staleness of the local model, ensuring a precise initial model for local training. Regarding the local gradient upload, we utilize the device's local data properties (i.e., sample volume and label distribution) to quantify its local gradient's importance, which then guides the determination of the gradient compression ratio. We have implemented Caesar, on two physical platforms with 40 smartphones and 80 NVIDIA Jetson devices. Extensive results show that Caesar, can reduce the traffic costs by about 25.54%þicksim37.88% when achieving the same target accuracy compared to the compression-based baselines, while incurring only a 0.68% degradation in final test accuracy relative to the full-precision communication.
Jiaming Yan, Jianchun Liu, Hongli Xu 0001, Zhen-guo Ma, Shilong Wang 0002
KDD (1)2
2026 MSFramework: Multi-stage similarity-based key flow identification in high-speed networks
Guoju Gao, Yu-e Sun, He Huang 0001, Jianchun Liu, Haibo Wang 0004, Yang Du 0006
Comput. Networks5
2026 Enhancing federated unlearning using catastrophic forgetting in heterogeneous Industrial Internet of Things
Zhen-guo Ma, Yanjing Sun, Hongli Xu 0001, Jianchun Liu, Yang Xu 0020, Yafei Lu
Comput. Commun.5
2026 Identifying Who You Are No Matter What You Write Through Abstracting Handwriting Style
abstract
With the increasing use of electronic devices, online handwriting verification has become crucial for biometricsbased identity authentication. Traditional methods, which rely on content-dependent verification of the writer's name, are vulnerable to forgery. This paper introduces a content-independent handwriting authentication system, Ph-Wri, designed for commodity smartphones. The core innovation is a multi-path attention feature fusion network that combines both static features (image of the handwritten text) and dynamic features (time-dependent properties during writing), to abstract the handwriting style instead of specific content for recognition, enabling robust user authentication. To extract handwriting style from dynamic writing features, we propose a polarity-aware attention strategy during training. This strategy incorporates Style Channel Attention (SCA) to capture direction-sensitive stylistic features, and Trajectory Spatial Attention (TSA) to highlight key handwriting trajectory regions. In the fine-tuning stage, the Correlation-Aware Attention (CAA) module models inter-channel structural correlations, mitigating the influence of content and enhancing style-consistent representations. By linking content-independent handwriting style to user identity, the system achieves accurate authentication. Extensive experiments on both the self-built CIEHD dataset and the public BiosecurID dataset demonstrate exceptional performance, achieving a 99% Verification Accuracy on CIEHD. Compared to state-of-theart methods that utilize only static or dynamic data, Ph-Wri significantly reduces the Equal Error Rate, showcasing the effectiveness and practicality of the proposed approach.
Jinyang Huang, Yuanhao Feng, Feng-Qi Cui, Xiang Zhang 0011, Zhi Liu 0002, Xin Liu 0104, Jianchun Liu, Fusang Zhang, Meng Li 0006
IEEE Trans. Dependable Secur. Comput.7
2026 FedQuad: Adaptive Layer-Wise LoRA Deployment and Activation Quantization for Federated Fine-Tuning
abstract
Federated fine-tuning (FedFT) provides an effective paradigm for fine-tuning large language models (LLMs) in privacy-sensitive scenarios. However, practical deployment remains challenging due to the limited resources on end devices. Existing methods typically utilize parameter-efficient fine-tuning (PEFT) techniques, such as Low-Rank Adaptation (LoRA), to substantially reduce communication overhead. Nevertheless, significant memory usage for activation storage and computational demands from full backpropagation remain major barriers to efficient deployment on resource-constrained end devices. Moreover, substantial resource heterogeneity across devices results in severe synchronization bottlenecks, diminishing the overall fine-tuning efficiency. To address these issues, we propose FedQuad, a novel LoRA-based FedFT framework that adaptively adjusts the LoRA depth (the number of consecutive tunable LoRA layers from the output) according to devices' computational power, while employing activation quantization to reduce memory overhead, thereby enabling efficient deployment on resource-constrained devices. Specifically, FedQuad first identifies the feasible and efficient combinations of LoRA depth and the number of activation quantization layers based on device-specific resource constraints. Subsequently, FedQuad employs a greedy strategy to select the optimal configurations for each device, effectively accommodating system heterogeneity. Extensive experiments demonstrate that FedQuad achieves a 1.4–5.3× convergence acceleration compared to state-of-the-art baselines when reaching target accuracy, highlighting its efficiency and deployability in resource-constrained and heterogeneous end-device environments.
Jianchun Liu, Rukuo Li, Hongli Xu 0001, Qianpiao Ma, Jiaming Yan, Liusheng Huang
IEEE Trans. Mob. Comput.1
2026 Toward Communication-Efficient Decentralized Federated Graph Learning Over Non-IID Data
abstract
Decentralized Federated Graph Learning (DFGL) overcomes the potential bottlenecks of the parameter server in FGL. However, extensive cross-worker communication of graph node embeddings during DFGL training introduces substantial communication costs. To improve communication efficiency, constructing sparse network topologies or applying graph sampling are potential methods. In this paper, we first reveal the bidirectional coupling between network topology construction and graph sampling, underscoring the necessity of their joint optimization. Motivated by this insight, we proposeDuplex, a unified framework that co-optimizes these two components by explicitly modeling their interdependent relationship, thereby significantly reducing communication costs while enhancing training performance in DFGL.Duplexformulates the decision-making process as a coordinated configuration$\langle \mathbf {A}, \mathbf {R} \rangle$, where$\bf {A}$is the adjacency matrix of the network topology and$\bf {R}$denotes the set of graph sampling ratios for workers. However, determining proper coordinated configurations to achieve optimal communication efficiency and training performance (e.g., model accuracy and convergence rate) is challenging due to several practical issues,e.g., statistical heterogeneity and dynamic network conditions. To overcome these challenges,Duplexintroduces a novel learning-driven algorithm to adaptively determine optimal network topologies and graph sampling ratios for workers. Experimental results demonstrate thatDuplexreduces completion time by 20.1%–48.8% and communication costs by 16.7%–37.6% to achieve target accuracy, while improving accuracy by 3.3%–7.9% under identical resource budgets compared to baselines.
Shilong Wang 0002, Jianchun Liu, Hongli Xu 0001, Chenxia Tang, Qianpiao Ma, Liusheng Huang
IEEE Trans. Mob. Comput.2
2026 Accelerating Decentralized Federated Learning With Probabilistic Communication in Heterogeneous Edge Computing
abstract
Decentralized federated learning (DFL) has gained popularity for training machine learning models on massive data in edge computing, as it avoids the potential bottleneck of conventional parameter server architectures. However, the existing DFL solutions typically use deterministic topologies that struggle with both system heterogeneity and non-IID local data, resulting in high bandwidth costs and slow convergence rates. In this paper, we propose a novel mechanism called Communication-efficient Decentralized Federated Learning (CedFL) to accelerate model training. InCedFL, each worker will communicate with each of its neighbors (i.e., model exchange) according to a certain probability at each epoch, so as to reduce bandwidth consumption. To this end, we then propose an efficient algorithm to adaptively determine the optimal probability for each worker pair according to real-time system situations (e.g., data distribution and bandwidth resource). Our proposed mechanism has been extensively tested on classical models and datasets, and the results demonstrate its high effectiveness.CedFLhas been shown to reduce completion time for model training by approximately 55% and improve test accuracy by 11% under the bandwidth constraint, compared to state-of-the-art solutions.
Jianchun Liu, Jiaming Yan, Hongli Xu 0001, Lun Wang 0003, Zhiyuan Wang 0002, Jinyang Huang, Chunming Qiao
IEEE Trans. Netw.1
2026 Asynchronous Federated Learning Over Non-IID Data via Over-the-Air Computation
abstract
Federated learning (FL) enables training AI models across distributed edge devices (i.e., workers) using local data, while facing challenges including communication resource constraints, edge heterogeneity, and non-IID data. Over-the-air computation (AirComp) has emerged as a promising technique to improve communication efficiency by leveraging the superposition property of a wireless multiple access channel (MAC) for model aggregation. However, over-the-air aggregation requires strict synchronization among edge devices, which is essentially incompatible with the asynchronous FL mechanisms often used to handle edge heterogeneity. To overcome this incompatibility, we propose Air-FedGA, a grouping-based asynchronous FL mechanism via AirComp, where workers are organized into groups for synchronized over-the-air aggregation within each group, while groups asynchronously communicate with the parameter server to update the global model. This design retains the communication efficiency of AirComp while addressing training inefficiency caused by edge heterogeneity. We provide a rigorous convergence analysis for Air-FedGA, theoretically quantifying how the convergence bound depends on several key factors, such as the maximum staleness, the degree of non-IID data among groups, and the AirComp aggregation mean squared error (MSE). Guided by these theoretical insights, we propose power control and worker grouping algorithms to minimize the convergence bound by jointly optimizing the AirComp aggregation MSE and the grouping strategy. We conduct experiments on classical models and datasets, and the results demonstrate that our proposed mechanism and algorithms can accelerate the model training by 1.83-$2.22\times $compared with the state-of-the-art solutions.
Qianpiao Ma, Xiaozhu Song, Junlong Zhou, Haibo Wang 0004, Yunming Liao, Jianchun Liu, Hongli Xu 0001
IEEE Trans. Netw.6
2026 Scalable, Low-Latency, and Hi-Precision Congestion Control in RDMA Datacenter Networks
Sun Xu, Bodong Yan, Yangming Zhao, Jianchun Liu, Hongli Xu 0001
IEEE Trans. Netw.4
2025 Top-nσ: Eliminating Noise in Logit Space for Robust Token Sampling of LLM
abstract
Large language models (LLMs) rely heavily on sampling methods to generate diverse and highquality text.While existing sampling methods like top-p and min-p have identified the detrimental effects of low-probability tails in LLMs' outputs, they still fail to effectively distinguish between diversity and noise.This limitation stems from their reliance on probability-based metrics that are inherently sensitive to temperature scaling.Through empirical and theoretical analysis, we make two key discoveries: (1) the pre-softmax logits exhibit a clear statistical separation between informative tokens and noise, and (2) we prove the mathematical equivalence of min-p and top-(1-p) under uniform distribution over logits.These findings motivate the design of top-nσ, a novel sampling method that identifies informative tokens by eliminating noise directly in logit space.Unlike existing methods that become unstable at high temperatures, top-nσ achieves temperature-invariant token selection while preserving output diversity.Extensive experiments across reasoning and creative writing tasks demonstrate that our method consistently outperforms existing approaches, with particularly significant improvements in high-temperature settings.
Chenxia Tang, Jianchun Liu, Hongli Xu 0001, Liusheng Huang
ACL (1)2
2025 Tackling Non-IID Graphs via Decoupled Structure and Feature in Federated Graph Learning
Longwen Wang, Jianchun Liu, Xianjun Gao, Jinyang Huang
DASFAA (3)2
2025 Towards High-Performance and Compatible RDMA Networks with Receiver-Based and Fine-Grained Congestion Control
Jianchun Liu, Hongli Xu 0001, Yangming Zhao, Zhuolong Yu
ICCCN2
2025 Many Hands Make Light Work: Accelerating Edge Inference via Multi-Client Collaborative Caching
abstract
Edge inference is a technology that enables real-time data processing and analysis on clients near the data source. To ensure compliance with the Service-Level Objectives (SLOs), such as a 30% latency reduction target, caching is usually adopted to reduce redundant computations in inference tasks on stream data. Due to task and data correlations, sharing cache information among clients can improve the inference performance. However, the non-independent and identically distributed (non-IID) nature of data across different clients and the long-tail distributions, where some classes have significantly more samples than others, will reduce cache hit ratios and increase latency. To address the aforementioned challenges, we propose an efficient inference framework, CoCa, which leverages a multi-client collaborative caching mechanism to accelerate edge inference. On the client side, the model is pre-set with multiple cache layers to achieve a quick inference. During inference, the model performs sequential lookups at cache layers activated by the edge server. On the server side, CoCa uses a two-dimensional global cache to periodically aggregate information from clients, mitigating the effects of non-IID data. For client cache allocation, CoCa first evaluates the importance of classes based on how frequently and recently their samples have been accessed. CoCa then selects frequently recurring classes to address long-tail distribution challenges. Finally, CoCa dynamically activates cache layers to balance lookup overhead and accuracy. Extensive experiments demonstrate that CoCa reduces inference latency by 23.0% to 45.2% on the VGG, ResNet and AST models with a slight loss of accuracy.
Wenyi Liang, Jianchun Liu, Hongli Xu 0001, Chunming Qiao, Liusheng Huang
ICDE2
2025 Accelerating End-Cloud Collaborative Inference via Near Bubble-Free Pipeline Optimization
Luyao Gao, Jianchun Liu, Hongli Xu 0001, Sun Xu, Qianpiao Ma, Liusheng Huang
INFOCOM2
2025 Towards Lightweight Traffic Forecasting in RDMA Networks: Design and Application
Bodong Yan, Sun Xu, Bingyi Liu, Jianchun Liu
INFOCOM6
2025 Air-FedGA: A Grouping Asynchronous Federated Learning Mechanism Exploiting Over-The-Air Computation
abstract
Federated learning (FL) is a new paradigm to train AI models over distributed edge devices (i.e., workers) using their local data, while confronting various challenges including communication resource constraints, edge heterogeneity and data Non-IID. Over-the-air computation (AirComp) is a promising technique to achieve efficient utilization of communication resource for model aggregation by leveraging the superposition property of a wireless multiple access channel (MAC). However, AirComp requires strict synchronization among edge devices, which is hard to achieve in heterogeneous scenarios. In this paper, we propose an AirComp-based grouping asynchronous federated learning mechanism (Air-FedGA), which combines the advantages of AirComp and asynchronous FL to address the communication and heterogeneity challenges. Specifically, AirFedGA organizes workers into groups and performs over-theair aggregation within each group, while groups asynchronously communicate with the parameter server to update the global model. In this way, Air-FedGA accelerates the FL model training by over-the-air aggregation, while relaxing the synchronization requirement of this aggregation technology. We theoretically prove the convergence of Air-FedGA. We formulate a training time minimization problem for Air-FedGA and propose the power control and worker grouping algorithm to solve it, which jointly optimizes the power scaling factors at edge devices, the denoising factors at the parameter server, as well as the worker grouping strategy. We conduct experiments on classical models and datasets, and the results demonstrate that our proposed mechanism and algorithm can speed up FL model training by$\mathbf{29.9\% - 71.6\%}$compared with the state-of-the-art solutions.
Qianpiao Ma, Junlong Zhou, Xiangpeng Hou, Jianchun Liu, Hongli Xu 0001, Jianeng Miao, Qingmin Jia
IPDPS4
2025 FRACTAL: Data-Aware Clustering and Communication Optimization for Decentralized Federated Learning
abstract
Decentralized federated learning (DFL) is a promising technique to enable distributed machine learning over edge nodes without relying on a centralized parameter server. However, existing DFL network topologies, such as fully connected, partially connected, or lower-tier hierarchical topology often struggle to effectively address the unique challenges presented by edge networks, including edge heterogeneity, communication resource constraint, and data Non-IID. In order to tackle these challenges, we propose a data-aware clustering algorithm, called FRACTAL, to construct a multi-tier hierarchical topology in a bottomup manner taking into consideration both data distribution and communication efficiency for DFL. We theoretically explore the quantitative relationship between the convergence bound of multi-tier FL and the data distribution among each-tier servers. To further improve communication efficiency and address edge heterogeneity, we deploy a time-sharing communication scheduling algorithm within each fractal unit (the basic structure in FRACTAL consisting of multiple nodes and an aggregator), called magic mirror method (MMM), to determine the optimal order of model distributing and uploading for nodes. We conduct extensive experiments on the classical models and datasets to evaluate the performance of FRACTAL, and the results show that FRACTAL can significantly accelerate the DFL model training by 48.6%- 72.3% compared with the state-of-the-art solutions.
Qianpiao Ma, Jianchun Liu, Hongli Xu 0001, Qingmin Jia, Renchao Xie
IEEE Trans. Big Data2
2025 Enhancing Semi-Supervised Federated Learning With Progressive Training in Heterogeneous Edge Computing
abstract
Federated learning (FL) is an efficient distributed learning method that facilitates collaborative model training among multiple edge devices (or clients). However, current research always assumes that clients have access to ground-truth data for training, which is unrealistic in practice because of a lack of expertise. Semi-supervised federated learning (SSFL) has been proposed in many existing works to address this problem, which always adopts a fixed model architecture for training, bringing two main problems with varying amounts of pseudo-labeled data. First, the shallow model cannot have the capability to fit the increasing pseudo-labeled data, leading to poor training performance. Second, the large model suffers from an overfitting problem when exploiting a few labeled data samples in SSFL, and also requires tremendous resource (e.g., computation and communication) costs. To tackle these problems, we propose a novel framework, calledstar, which adopts progressive training to enhance model training in SSFL. Specifically,stargradually increases the model depth through adding the sub-module (e.g., one or several layers) from a shallow model, and performs pseudo-labeling for unlabeled data with a specialized confidence threshold simultaneously. Then, we propose an efficient algorithm to determine the appropriate model depth for each client with varied resource budgets and the proper confidence threshold for pseudo-labeling in SSFL. The experimental results demonstrate the high effectiveness of STAR. For instance,starcan reduce the bandwidth consumption by about 40%, and achieve an average accuracy improvement of around 9.8% compared with the baselines, on CIFAR10.
Jianchun Liu, Jun Liu 0083, Hongli Xu 0001, Yunming Liao, Min Chen 0033, Chen Qian 0001
IEEE Trans. Mob. Comput.1
2025 Adaptive Parameter-Efficient Federated Fine-Tuning on Heterogeneous Devices
abstract
Federated fine-tuning (FedFT) has been proposed to fine-tune the pre-trained language models in a distributed manner. However, there are two critical challenges for efficient FedFT in practical applications,i.e., resource constraints and system heterogeneity. Existing works rely on parameter-efficient fine-tuning methods,e.g., low-rank adaptation (LoRA), but with major limitations. Herein, based on the inherent characteristics of FedFT, we observe that LoRA layers with higher ranks added close to the output help to save resource consumption while achieving comparable fine-tuning performance. Then we propose a novel LoRA-based FedFT framework, termed LEGEND, which faces the difficulty of determining the number of LoRA layers (called, LoRA depth) and the rank of each LoRA layer (called, rank distribution). We analyze the coupled relationship between LoRA depth and rank distribution, and design an efficient LoRA configuration algorithm for heterogeneous devices, thereby promoting fine-tuning efficiency. Extensive experiments are conducted on a physical platform with 80 commercial devices. The results show that LEGEND can achieve a speedup of 1.5-2.8× and save communication costs by about 42.3% when achieving the target accuracy, compared to the advanced solutions.
Jun Liu 0083, Yunming Liao, Hongli Xu 0001, Yang Xu 0020, Jianchun Liu, Chen Qian 0001
IEEE Trans. Mob. Comput.5
2025 FedACS: An Adaptive Client Selection Framework for Communication-Efficient Federated Graph Learning
abstract
Federated graph learning (FGL) has been proposed to collaboratively train the increasing graph data with graph neural networks (GNNs) in a recommendation system. Nevertheless, implementing an efficient recommendation system with FGL still faces two primary challenges, i.e., limited communication bandwidth and non-IID local graph data. Existing works typically reduce communication frequency or transmission amount, which may suffer significant performance degradation under non-IID settings. Furthermore, some researchers propose to share the underlying structure among clients, which brings massive communication cost. To this end, we propose an efficient FGL framework, named FedACS, which adaptively selects a subset of clients for model training, to alleviate communication overhead and non-IID issues simultaneously. In FedACS, the global GNN model learns significant hidden edges and the structure of graph data among selected clients, enhancing recommendation efficiency. This capability distinguishes it from the traditional FL client selection methods. To optimize the client selection process, we introduce a multi-armed bandit (MAB) based algorithm to select participating clients according to the resource budgets and the training performance (i.e., RMSE). Experimental results indicate that FedACS improves RMSE by 5.4% over baselines with the same resource budget and reduces communication costs by up to 70.7% to achieve the same RMSE performance.
Hongli Xu 0001, Xianjun Gao, Jianchun Liu, Qianpiao Ma, Liusheng Huang
IEEE Trans. Mob. Comput.3
2025 FedSNN: Training Slimmable Neural Network With Federated Learning in Edge Computing
abstract
To provide a flexible tradeoff between inference accuracy and resource requirement at runtime, the slimmable neural network (SNN), a single network executable at different widths with the same deploying and management cost as that of a single model, has been proposed. However, how to effectively train SNN among massive devices in edge computing without revealing their local data remains an open problem. To this end, we leverage a novel distributed machine learning paradigm, i.e., federated learning, to realize effective on-device SNN training. As current FL schemes often train only one model with fixed architecture, and the existing SNN training algorithm is resource-intensive, integrating FL and SNN is non-trivial. Furthermore, two intrinsic features in edge computing, i.e., data and system heterogeneity, exacerbate the difficulty. Motivated by this, we redesign the model distribution, local training, and model aggregation phases in traditional FL, and propose FedSNN, a framework that ensures all widths in SNN can obtain high accuracy with less resource consumption. Specifically, for devices with heterogeneous training capacities and data distributions, the parameter server will distribute each of them with one proper width for adaptive local training guided by their uploaded model features, and their trained models will be weighted-averaged using the proposed multi-width SNN aggregation to improve their statistical utility. Extensive experiments on a distributed testbed show that FedSNN improves the model accuracy by about 2.18%-8.1%, and accelerates training by about$1.31\times $-$6.84\times $, compared with existing solutions.
Yang Xu 0020, Yunming Liao, Hongli Xu 0001, Zhiyuan Wang 0002, Lun Wang 0003, Jianchun Liu, Chen Qian 0001
IEEE Trans. Netw.6
2025 Adaptive Local Update and Neural Composition for Accelerating Federated Learning in Heterogeneous Edge Networks
abstract
Federated Learning (FL) enables distributed clients to collaboratively train models without exposing their private data. However, it is difficult to implement efficient FL due to limited resources. Most existing works compress the transmitted gradients or prune the global model to reduce the resource cost, but leave the compressed or pruned parameters under-optimized, which degrades the training performance. To address this issue, the neural composition technique constructs size-adjustable models by composing low-rank tensors, allowing every parameter in the global model to learn the knowledge from all clients. Nevertheless, some tensors can only be optimized by a small fraction of clients, thus the global model may get insufficient training, leading to a long completion time, especially in heterogeneous edge scenarios. To this end, we enhance the neural composition technique, enabling all parameters to be fully trained. Further, we propose a lightweight FL framework, called Heroes, with enhanced neural composition and adaptive local update. A greedy-based algorithm is designed to adaptively assign the proper tensors and local update frequencies for participating clients according to their heterogeneous capabilities and resource budgets. On this basis, we further propose an extension of Heroes, termed AdaHeroes, which further improves the training performance under the statistical heterogeneity scenario based on an adaptive client selection strategy. Extensive experiments demonstrate that Heroes can reduce traffic consumption by about 72.46% and provide up to$2.76\times $speedup compared to the baselines. Furthermore, with the setting of statistical heterogeneity, AdaHeroes can improve the test accuracy by about 4.77% compared with Heroes and the baselines.
Jianchun Liu, Jiaming Yan, Ji Qi 0005, Hongli Xu 0001, Shilong Wang 0002, Chunming Qiao, Liusheng Huang
IEEE Trans. Netw.1
2024 Heroes: Lightweight Federated Learning with Neural Composition and Adaptive Local Update in Heterogeneous Edge Networks
abstract
Federated Learning (FL) enables distributed clients to collaboratively train models without exposing their private data. However, it is difficult to implement efficient FL due to limited resources. Most existing works compress the transmitted gradients or prune the global model to reduce the resource cost, but leave the compressed or pruned parameters under-optimized, which degrades the training performance. To address this issue, the neural composition technique constructs size-adjustable models by composing low-rank tensors, allowing every parameter in the global model to learn the knowledge from all clients. Nevertheless, some tensors can only be optimized by a small fraction of clients, thus the global model may get insufficient training, leading to a long completion time, especially in heterogeneous edge scenarios. To this end, we enhance the neural composition technique, enabling all parameters to be fully trained. Further, we propose a lightweight FL framework, called Heroes, with enhanced neural composition and adaptive local update. A greedy-based algorithm is designed to adaptively assign the proper tensors and local update frequencies for participating clients according to their heterogeneous capabilities and resource budgets. Extensive experiments demonstrate that Heroes can reduce traffic consumption by about 72.05% and provide up to 2.97× speedup compared to the baselines.
Jiaming Yan, Jianchun Liu, Shilong Wang 0002, Hongli Xu 0001
INFOCOM2
2024 Towards Communication-Efficient Federated Graph Learning: An Adaptive Client Selection Perspective
abstract
Federated graph learning (FGL) has been proposed to collaboratively train the increasing graph data with graph neural networks (GNNs) in a recommendation system, aggregating the features of graph nodes and edges among these nodes. Nevertheless, implementing an efficient recommendation system with FGL still faces two primary challenges, i.e., limited communication bandwidth and non-IID local graph data. Existing works typically reduce communication frequency or transmission amount, which may suffer significant performance degradation under non-IID settings. Furthermore, some researchers propose to share the underlying structure information among all clients, which brings massive communication cost. To this end, we propose an efficient FGL framework, named FedACS, which adaptively selects a subset of clients for model training, to alleviate communication overhead and non-IID issues simultaneously. In FedACS, the global GNN model can learn significant hidden edges and the structure of graph data among selected clients, enhancing recommendation efficiency. This capability distinguishes it from the traditional FL client selection methods. To optimize the client selection process, we introduce a multi-armed bandit (MAB) based algorithm to select participating clients according to the resource budgets and the training performance (i.e., RMSE) under different data distributions. Experimental results show that, given the same resource budget, FedACS achieves the RMSE improvement of 5.4% over the baselines. Besides, when achieving the same RMSE performance, FedACS saves up to approximately 70.7% communication cost, compared with the baselines.
Xianjun Gao, Jianchun Liu, Hongli Xu 0001, Qianpiao Ma, Lun Wang 0003
IWQoS2
2024 LHCC: Low-Latency and Hi-Precision Congestion Control in RDMA Datacenter Networks
abstract
Congestion Control (CC) plays a vital role in deploying lossless datacenter networks based on Remote Direct Memory Access (RDMA). A high-performance CC scheme should provide low-latency and precise feedback to congestion events. However, no existing CC schemes achieved both features simultaneously. In this paper, we propose LHCC, a Low-latency and Hi-precision Congestion Control scheme for RDMA datacenter networks. LHCC uses out-band signaling to notify the network status and hence a packet sender can detect congestion events within an RTT. In addition, LHCC adjusts packet sending rate by taking into consideration all queues along the entire path that a packet has gone through. Accordingly, it provides a more precise CC compared with existing schemes especially when there are multiple bottlenecks in the networks. We build the LHCC prototype on a real testbed carrying NVIDIA BlueField-3 NICs and AGM39D FPGAs. Both testbed experiments and extensive simulations show that LHCC can reduce the Flow Completion Time (FCT) slow down and reduce the buffer usage (i.e., reduce the queue lengths) by up to 62.5% and 58%, respectively, compared with the state-of-the-art high-precision CC scheme, HPCC.
Bodong Yan, Yangming Zhao, Sun Xu, Jianchun Liu, Hongli Xu 0001
IWQoS4
2024 Dynamic Staleness Control for Asynchronous Federated Learning in Decentralized Topology
Qianpiao Ma, Jianchun Liu, Qingmin Jia, Xiaomao Zhou, Yujiao Hu, Renchao Xie
WASA (2)2
2024 FedCD: A Hybrid Federated Learning Framework for Efficient Training With IoT Devices
abstract
With billions of IoT devices producing vast data globally, privacy and efficiency challenges arise in AI applications. Federated learning (FL) has been widely adopted to train deep neural networks (DNNs) without privacy leakage. Existing centralized and decentralized FL architectures have limitations, including memory burden, huge bandwidth pressure and non-IID data issues. This paper introduces a novel hybrid FL framework, named FedCD, merging the benefits of both centralized and decentralized FL architectures. FedCD strategically distributes the model based on layer sizes and consensus distances (i.e., the deviation between the local models and the global average models), effectively relieving network bandwidth pressures and accelerating training speed even under the non-IID setting. This method significantly mitigates resource constraints and improves model accuracy, offering a promising solution to the challenges in distributed machine learning. Extensive experiment results show the high effectiveness of FedCD. The total completion time of FedCD is reduced by 16.3%-53% and the average accuracy improvement is 1.85% compared to the baselines.
Jianchun Liu, Pengcheng Qu, Sun Xu, Zhi Liu 0002, Qianpiao Ma, Jinyang Huang
IEEE Internet Things J.1
2024 KeystrokeSniffer: An Off-the-Shelf Smartphone Can Eavesdrop on Your Privacy From Anywhere
abstract
With mobile phones becoming increasingly prevalent and embedding high-quality microphones, attackers have the ability to employ these microphones to eavesdrop user’s keyboard input. However, existing work usually assumes that keystroke eavesdropping is performed against known environments and victims, which inevitably makes attack systems lack generalization. To reveal the real threat of the acoustic signal-based attack strategy, this paper proposes a keystroke eavesdropping algorithm called KeystrokeSniffer, which is robust to unknown input environments and unknown victims. In particular, to mimic the real input environment of victims, an environment estimation algorithm is first designed by extracting the timbre-related characteristics to predict the keyboard type and identifying large-size key data from collected unlabeled samples to estimate the 3D microphone coordinates. Then, by imitating unknown environments and victim data, this algorithm achieves effective keystroke eavesdropping with a small training set. By further considering the commonalities of different keystroke habits, a robust feature extraction method that reflects the keystroke location is adopted to reduce the impact of individual input habits. Extensive experimental results using various commodity smartphones indicate that the scheme is capable of predicting keyboard input accurately under different unknown scenarios. Specifically, even when both the victims and keyboards are unknown, KeystrokeSniffer can still achieve high Top-5 accuracy, reaching 79.5% in predicting keystrokes and 96.7% in predicting meaningful words, which demonstrates KeystrokeSniffer has excellent generalization capabilities. By setting different parameter values of various impact factors, e.g., noise and hand length factors, the strong robustness of the system is demonstrated, which proves that KeystrokeSniffer can violate privacy in real situations.
Jinyang Huang, Jia-Xuan Bai, Xiang Zhang 0011, Zhi Liu 0002, Yuanhao Feng, Jianchun Liu, Xiao Sun 0003, Mianxiong Dong, Meng Li 0006
IEEE Trans. Inf. Forensics Secur.6
2024 PhyFinAtt: An Undetectable Attack Framework Against PHY Layer Fingerprint-Based WiFi Authentication
abstract
WiFi connection has been suffering from MAC forgery attacks due to the loose authentication mechanism between access points (APs) and clients. To address this problem, the physical (PHY) layer information-based fingerprint has been adopted for safe WiFi authentication. Since such a fingerprint is constant and unique for each specific network interface card (NIC), it can effectively prevent MAC forgery attacks. However, the PHY layer information-based fingerprint is still vulnerable to malicious attacks as it is extracted from Channel State Information (CSI), and its stability can be affected by the wireless environment. In this paper, we propose a novel undetectable attack framework, called PhyFinAtt, base on which the attacker can undermine the stability of the PHY layer-based authentication fingerprints through human movement and further attack the WiFi authentication protocols. Specifically, we first demonstrate that human movement at a designated location can affect the PHY fingerprint. We then illustrate the impact of human movement on the PHY fingerprint and the relationship between the movement and the channel quality to ensure that the PHY fingerprint is destroyed by the movement in an undetected way without affecting normal communication. Extensive experiments in real-world scenarios show that our proposed attack can effectively disrupt the stability of the PHY fingerprints and significantly degrade the performance of the authentication protocols based on such fingerprints. To the best of our knowledge, this is the first study on effective attacks against the PHY information-based WiFi authentication protocols. Furthermore, we also present a practical defense mechanism without involving any additional equipment to mitigate attacks similar to PhyFinAtt.
Jinyang Huang, Bin Liu 0016, Chenglin Miao, Xiang Zhang 0011, Jianchun Liu, Lu Su 0001, Zhi Liu 0002, Yu Gu 0003
IEEE Trans. Mob. Comput.5
2024 Semi-Supervised Decentralized Machine Learning With Device-to-Device Cooperation
abstract
The massive data from mobile and embedded devices have huge potential for training machine learning models. Decentralized machine learning (DML) can avoid the inherent bottleneck of the parameter server (PS) by collaboratively training models in a device-to-device (D2D) fashion. However, the previous DML works often assume that the local data are fully annotated with ground-truth labels, which is unrealistic for many Internet of Things (IoT) applications. This arises a new practical DML scenario, namely semi-supervised DML, where the local data of distributed workers are partially labeled in the D2D network. The existing semi-supervised learning techniques are proposed for standalone or the PS architecture, which ignore the impact of D2D topology on the performance of semi-supervised learning. Thus, they cannot adequately leverage the unlabeled data of decentralized workers, leading to performance degradation. Herein, we propose a novel framework, called SSD, to address the problem of semi-supervised DML by exploiting D2D cooperation. The key insight behind SSD is that neighbor selection has a crucial impact on pseudo-label quality and communication overhead. In SSD, each worker adaptively selects its neighbors with high-quality models and similar data distribution under communication resource constraints, which helps to generate high-confidence pseudo-labels for local unlabeled data and further boosts the DML performance. Extensive empirical evaluations on both testbed and simulated environments show that SSD significantly outperforms other baselines.
Zhida Jiang, Yang Xu 0020, Hongli Xu 0001, Zhiyuan Wang 0002, Jianchun Liu, Chunming Qiao
IEEE Trans. Mob. Comput.5
2024 Computation and Communication Efficient Federated Learning With Adaptive Model Pruning
abstract
Federated learning (FL) has emerged as a promising distributed learning paradigm that enables a large number of mobile devices to cooperatively train a model without sharing their raw data. The iterative training process of FL incurs considerable computation and communication overhead. The workers participating in FL are usually heterogeneous and the workers with poor capabilities may become the bottleneck of model training. To address the challenges of resource overhead and system heterogeneity, this article proposes an efficient FL framework, called FedMP, that improves both computation and communication efficiency over heterogeneous workers through adaptive model pruning. We theoretically analyze the impact of pruning ratio on training performance, and employ a Multi-Armed Bandit based online learning algorithm to adaptively determine different pruning ratios for heterogeneous workers, even without any prior knowledge of their capabilities. As a result, each worker in FedMP can train and transmit the sub-model that fits its own capabilities, accelerating the training process without hurting model accuracy. To prevent the diverse structures of pruned models from affecting the training convergence, we further present a new parameter synchronization scheme, called Residual Recovery Synchronous Parallel (R2SP). Besides, our proposed framework can be extended to the peer-to-peer (P2P) setting. Extensive experiments on physical devices demonstrate that FedMP is effective for different heterogeneous scenarios and data distributions, and can provide up to 4.1× speedup compared to the existing FL methods.
Zhida Jiang, Yang Xu 0020, Hongli Xu 0001, Zhiyuan Wang 0002, Jianchun Liu, Chen Qian 0001, Chunming Qiao
IEEE Trans. Mob. Comput.5
2024 Finch: Enhancing Federated Learning With Hierarchical Neural Architecture Search
abstract
Federated learning (FL) has been widely adopted to train machine learning models over massive data in edge computing. Most works of FL employ pre-defined model architectures on all participating clients for model training. However, these pre-defined architectures may not be the optimal choice for the FL setting since manually designing a high-performance neural architecture is complicated and burdensome with intense human expertise and effort, which easily makes the model training fall into the local suboptimal solution. To this end, Neural Architecture Search (NAS) has been applied to FL to address this critical issue. Unfortunately, the search space of existing federated NAS approaches is extraordinarily large, resulting in unacceptable completion time on the resource-constrained edge clients, especially under the non-independent and identically distributed (non-IID) setting. In order to remedy this, we propose a novel framework, calledFinch, which adopts hierarchical neural architecture search to enhance federated learning. InFinch, we first divide the clients into several clusters according to the data distribution. Then, some subnets are sampled from a pre-trained supernet and allocated to the specific client clusters for searching the optimal model architecture in parallel, so as to significantly accelerate the process of model searching and training. The extensive experimental results demonstrate the high effectiveness of our proposed framework. Specifically,Finchcan reduce the completion time by about 30.6%, and achieve an average accuracy improvement of around 9.8% compared with the baselines.
Jianchun Liu, Jiaming Yan, Hongli Xu 0001, Zhiyuan Wang 0002, Jinyang Huang, Yang Xu 0020
IEEE Trans. Mob. Comput.1
2024 FedUC: A Unified Clustering Approach for Hierarchical Federated Learning
abstract
Federated learning (FL) is an effective approach to train models collaboratively among distributed edge nodes (i.e., workers) while facing three crucial challenges, edge heterogeneity, resource constraint, and Non-IID data. Under the parameter server (PS) architecture, a single parameter server may become the system bottleneck and cannot well deal with the edge heterogeneity, while the peer-to-peer (P2P) architecture causes significant communication consumption to achieve satisfactory training performance. To this end, hierarchical aggregation (HA) architecture is proposed to cluster workers to tackle the edge heterogeneity and reduce communication consumption for FL. However, the existing researches on HA architecture cannot provide a unified clustering approach for various inter-cluster aggregation patterns (e.g., centralized or decentralized structure, synchronous or asynchronous mode). In this paper, we explore the quantitative relationship between the convergence bounds of different inter-cluster patterns and several factors, e.g., data distribution, frequency of clusters participating in inter-cluster aggregation (for asynchronous modes), and inter-cluster topology (for decentralized structures). Based on the convergence bounds, we design a unified clustering algorithm FedUC to organize workers for different patterns. Experimental results on classical models and datasets show that FedUC can greatly accelerate the model training of different patterns by 1.79-7.39× compared with the state-of-the-art clustering methods.
Qianpiao Ma, Yang Xu 0020, Hongli Xu 0001, Jianchun Liu, Liusheng Huang
IEEE Trans. Mob. Comput.4
2024 Like Attracts Like: Personalized Federated Learning in Decentralized Edge Computing
abstract
The emerging Personalized Federated Learning (PFL) methods aim to produce personalized models for different users, so as to keep track of their individualized requirements in Edge Computing (EC). The centralized PFL methods may suffer from the communication bottleneck and single point of failure. As an alternative solution, the decentralized PFL (DPFL) methods are performed in a Peer-to-Peer (P2P) manner, and collaboratively train personalized models by model aggregation among congenial devices. However, these DPFL methods may incur high communication cost and low resource utilization induced by large-scale models. Herein, we take the communication constraint and heterogeneity into consideration and propose to realize communication-efficient DPFL with adaptive model pruning and neighbor selection. We theoretically analyze the convergence of the proposed DPFL method, and study the impacts of both model pruning and neighbor selection on training performance. Furthermore, we propose an efficient algorithm that combines model pruning and neighbor selection to achieve a trade-off between model quality and communication cost. Extensive simulation and testbed experiments on real-world datasets are conducted. The experimental results demonstrate that the proposed algorithm can improve the test accuracy by at most 13% and save the traffic consumption by 45.4% on average compared with the existing PFL methods.
Zhen-guo Ma, Yang Xu 0020, Hongli Xu 0001, Jianchun Liu, Yinxing Xue
IEEE Trans. Mob. Comput.4
2024 FedLC: Accelerating Asynchronous Federated Learning in Edge Computing
abstract
Federated Learning (FL) has been widely adopted to process the enormous data in the application scenarios like Edge Computing (EC). However, the commonly-used synchronous mechanism in FL may incur unacceptable waiting time for heterogeneous devices, leading to a great strain on the devices' constrained resources. In addition, the alternative asynchronous FL is known to suffer from the model staleness, which will lead to performance degradation of the trained model, especially onnon-i.i.d.data. In this paper, we design a novel asynchronous FL mechanism, named FedLC, to handle thenon-i.i.d.issue in EC by enabling the local collaboration among edge devices. Specifically, apart from uploading the local model directly to the server, each device will transmit its gradient to the other devices with different data distributions for local collaboration, which can improve the model generality. We theoretically analyze the convergence rate of FedLC and obtain the quantitative relationship between convergence bound and local collaboration. We design an efficient algorithm utilizing demand-list to determine the set of devices receiving gradients from each device. To handle the model staleness, we further assign different learning rates for various devices according to their participation frequency. The extensive experimental results demonstrate the effectiveness of our proposed mechanism.
Yang Xu 0020, Zhen-guo Ma, Hongli Xu 0001, Suo Chen, Jianchun Liu, Yinxing Xue
IEEE Trans. Mob. Comput.5
2024 Enhancing Federated Learning With Server-Side Unlabeled Data by Adaptive Client and Data Selection
abstract
Federated learning (FL) has been widely applied to collaboratively train deep learning (DL) models on massive end devices (i.e., clients). Due to the limited storage capacity and high labeling cost, the data on each client may be insufficient for model training. Conversely, in cloud datacenters, there exist large-scale unlabeled data, which are easy to collect from public access (e.g., social media). Herein, we propose theAda-FedSemisystem, which leverages both on-device labeled data and in-cloud unlabeled data to boost the performance of DL models. In each round, local models are aggregated to produce pseudo-labels for the unlabeled data, which are utilized to enhance the global model. Considering that the number of participating clients and the quality of pseudo-labels will have a significant impact on the training performance, we introduce a multi-armed bandit (MAB) based online algorithm to adaptively determine the participating fraction and confidence threshold. Besides, to alleviate the impact of stragglers, we assign local models of different depths for heterogeneous clients. Extensive experiments on benchmark models and datasets show that given the same resource budget, the model trained by Ada-FedSemi achieves 3%$\sim$14.8% higher test accuracy than that of the baseline methods. When achieving the same test accuracy, Ada-FedSemi saves up to 48% training cost, compared with the baselines. Under the scenario with heterogeneous clients, the proposed HeteroAda-FedSemi can further speed up the training process by$1.3\times \sim 1.5\times$.
Yang Xu 0020, Lun Wang 0003, Hongli Xu 0001, Jianchun Liu, Zhiyuan Wang 0002, Liusheng Huang
IEEE Trans. Mob. Comput.4
2024 Peaches: Personalized Federated Learning With Neural Architecture Search in Edge Computing
abstract
In edge computing (EC), federated learning (FL) enables numerous distributed devices (or workers) to collaboratively train AI models without exposing their local data. Most works of FL adopt a predefined architecture on all participating workers for model training. However, since workers' local data distributions vary heavily in EC, the predefined architecture may not be the optimal choice for every worker. It is also unrealistic to manually design a high-performance architecture for each worker, which requires intense human expertise and effort. In order to tackle this challenge, neural architecture search (NAS) has been applied in FL to automate the architecture design process. Unfortunately, the existing federated NAS frameworks often suffer from the difficulties of system heterogeneity and resource limitation. To remedy this problem, we present a novel framework, termedPeaches, to achieve efficient searching and training in the resource-constrained EC system. Specifically, the local model of each worker is stacked by base cell and personal cell, where the base cell is shared by all workers to capture the common knowledge and the personal cell is customized for each worker to fit the local data. We determine the number of base cells, shared by all workers, according to the bandwidth budget on the parameters server. Besides, to relieve the data and system heterogeneity, we find the optimal number of personal cells for each worker based on its computing capability. In addition, we gradually prune the search space during training to mitigate the resource consumption. We evaluate the performance ofPeachesthrough extensive experiments, and the results show thatPeachescan achieve an average accuracy improvement of about 6.29% and up to 3.97× speed up compared with the baselines.
Jiaming Yan, Jianchun Liu, Hongli Xu 0001, Zhiyuan Wang 0002, Chunming Qiao
IEEE Trans. Mob. Comput.2
2024 Ferrari: A Personalized Federated Learning Framework for Heterogeneous Edge Clients
abstract
Federated semi-supervised learning (FSSL) has been proposed to address the insufficient labeled data problem by training models with pseudo-labeling. In previous FSSL systems, a single global model is always trained without an equivalent generalization ability for the clients under the non-IID setting. Accordingly, model personalization methods have been proposed to overcome this problem. Intuitively, seeking labeling assistance from other clients with similar data distribution,i.e., model migration, can effectively improve the personalization on the clients with scarce labeled data. However, previous works require to migrate a pre-fixed number of models among the clients, causing unnecessary resource waste and accuracy degradation due to resource heterogeneity. Considering that the number of model migrations and the quality of pseudo-labels have a significant impact on the training performance (e.g., efficiency and accuracy), we propose a novel personalized FSSL system, called Ferrari, to boost the efficiency of pseudo-labeling and training accuracy through adaptive model migrations among the clients. Specifically, Ferrari first generates the similarity-based ranking using a Gaussian KD-Tree, considering the varied data distributions among the clients. Combined with the ranking and clients' heterogeneous resource constraints, Ferrari then adaptively determines the proper model migration policy and confidence thresholds for high-quality pseudo-labeling and personalized training for clients. Extensive experiments on a physical platform show that Ferrari provides a 1.2$\sim 5.5\times$speedup without sacrificing model accuracy, compared to existing methods.
Jianchun Liu, Hongli Xu 0001, Lun Wang 0003, Chen Qian 0001, Yunming Liao
IEEE Trans. Mob. Comput.2
2024 YOGA: Adaptive Layer-Wise Model Aggregation for Decentralized Federated Learning
abstract
Traditional Federated Learning (FL) is a promising paradigm that enables massive edge clients to collaboratively train deep neural network (DNN) models without exposing raw data to the parameter server (PS). To avoid the bottleneck on the PS, Decentralized Federated Learning (DFL), which utilizes peer-to-peer (P2P) communication without maintaining a global model, has been proposed. Nevertheless, DFL still faces two critical challenges, i.e., limited communication bandwidth and not independent and identically distributed (non-IID) local data, thus hindering efficient model training. Existing works commonly assume full model aggregation at periodic intervals, i.e., clients periodically collect models from peers. To reduce the communication cost, these methods allow clients to collect model(s) from selected peers, but often result in a significant degradation of model accuracy when dealing with non-IID data. Alternatively, the layer-wise aggregation mechanism has been proposed to alleviate communication overhead under the PS architecture, but its potential in DFL remains rarely explored yet. To this end, we propose an efficient DFL framework YOGA that adaptively performs layer-wise model aggregation and training. Specifically, YOGA first generates the ranking of layers in the model according to the learning speed and layer-wise divergence. Combining with the layer ranking and peers’ status information (i.e., data distribution and communication capability), we propose the max-match (MM) algorithm to generate the proper layer-wise model aggregation policy for the clients. Extensive experiments on DNN models and datasets show that YOGA saves communication cost by about 45% without sacrificing the model performance compared with the baselines, and provides 1.53-$3.5\times $speedup on the physical platform.
Jun Liu 0083, Jianchun Liu, Hongli Xu 0001, Yunming Liao, Zhiyuan Wang 0002, Qianpiao Ma
IEEE/ACM Trans. Netw.2
2024 Federated Learning With Experience-Driven Model Migration in Heterogeneous Edge Networks
abstract
To approach the challenges of non-IID data and limited communication resource raised by the emerging federated learning (FL) in mobile edge computing (MEC), we propose an efficient framework, calledFedMigr, which integrates a deep reinforcement learning (DRL) based model migration strategy into the pioneer FL algorithmFedAvg. According to the data distribution and resource budgets, ourFedMigrwill intelligently guide one client to forward its local model to another client after local updating, before directly sending the local models to the server for global aggregation as inFedAvg. Intuitively, migrating a local model from one client to another is equivalent to training the model over more data from different clients, alleviating the influence of non-IID issue. To this end, we propose an experience-driven method to make proper decisions for model migrations while satisfying the resource constraints. We also prove thatFedMigrcan help to reduce the parameter divergences between different local models and the global model from a theoretical perspective under the non-IID setting. Extensive experiments on three popular benchmark datasets demonstrate thatFedMigrcan achieve an average accuracy improvement of around 13%, and reduce bandwidth consumption for global communication by 42% on average, compared with the baselines.
Jianchun Liu, Shilong Wang 0002, Hongli Xu 0001, Yang Xu 0020, Yunming Liao, Jinyang Huang, He Huang 0001
IEEE/ACM Trans. Netw.1
2024 Adaptive Block-Wise Regularization and Knowledge Distillation for Enhancing Federated Learning
abstract
Federated Learning (FL) is a distributed model training framework that allows multiple clients to collaborate on training a global model without disclosing their local data in edge computing (EC) environments. However, FL usually faces statistical heterogeneity (e.g., non-IID data) and system heterogeneity (e.g., computing and communication capabilities), resulting in poor model training performance. To deal with the above two challenges, we propose an efficient FL framework, named FedBR, which integrates the idea of block-wise regularization and knowledge distillation (KD) into the pioneering FL algorithm FedAvg, for resource-constrained edge computing. Specifically, we first divide the model into multiple blocks according to the layer order of deep neural network (DNN). The server only sends some consecutive model blocks instead of an entire model to clients for communication efficiency. Then, the clients make use of knowledge distillation to absorb the knowledge of global model blocks to alleviate statistical heterogeneity during local training. We provide a theoretical convergence guarantee for FedBR and show that the convergence bound will decrease as the increasing number of model blocks sent by the server. Besides, since the increasing number of model blocks brings more computing and communication costs, we design a heuristic algorithm (GMBS) to determine the appropriate number of model blocks for clients according to their varied data distributions, computing, and communication capabilities. Extensive experimental results show that FedBR can reduce the bandwidth consumption by about 31%, and achieve an average accuracy improvement of around 5.6% compared with the baselines under heterogeneous settings.
Jianchun Liu, Qingmin Zeng, Hongli Xu 0001, Yang Xu 0020, Zhiyuan Wang 0002, He Huang 0001
IEEE/ACM Trans. Netw.1
2024 FAST: Enhancing Federated Learning Through Adaptive Data Sampling and Local Training
abstract
The emerging paradigm of federated learning (FL) strives to enable devices to cooperatively train models without exposing their raw data. In most cases, the data across devices are non-independently and identically distributed in FL. Thus, the local models trained over different data distributions will inevitably deviate from the global optima, which induces optimization inconsistency and even hurts global convergence. Moreover, the resource-constrained devices with heterogeneous training capacities (e.g., computing and communication) further slow down the convergence rate. To this end, we introduce anFL framework withadaptive datasampling and localtraining, namely FAST. Specifically, even without devices’ private data distributions, FAST enables each device to sample different rates of data points from each of its local classes to rebuild a dataset for training, thus adjusting the convergence direction of the aggregated global model to be closer to the global optima. The theoretical analysis shows that the convergence bound depends on the sampling rates as well as the number of local iterations executed on the sampled data. To achieve resource-effective and convergence-guaranteed FL, we then design an online learning algorithm that jointly optimizes the data sampling and local training strategies so as to encourage the decrease of global loss under the given time budget. Extensive experiments on physical and simulated environments show that, FAST improves the model accuracy by about 1.55%-6.78% given the same time budget, and accelerates training by about 1.39-5.89× with the same target accuracy, compared with the baselines.
Zhiyuan Wang 0002, Hongli Xu 0001, Yang Xu 0020, Zhida Jiang, Jianchun Liu, Suo Chen
IEEE Trans. Parallel Distributed Syst.5
2023 FedCD: A Hybrid Centralized-Decentralized Architecture for Efficient Federated Learning
abstract
With billions of IoT devices producing vast data globally, privacy and efficiency challenges arise in AI applications. Federated learning (FL) has been widely adopted to train deep neural networks (DNNs) without privacy leakage. Existing centralized and decentralized FL architectures have limitations, including memory burden, huge bandwidth pressure and non-IID data issues. This paper introduces a novel framework, named FedCD, merging the benefits of both centralized and decentralized FL architectures. FedCD strategically distributes the model based on layer sizes and consensus distances (measuring the deviation between the local models and the global average models), effectively relieving network bandwidth pressures and accelerating training speed even under the non-IID setting. This method significantly mitigates resource constraints and improves model accuracy, offering a promising solution to the challenges in distributed machine learning. Extensive experiment results show the high effectiveness of FedCD. The total completion time of FedCD is reduced by 16.3%-53% and the average accuracy improvement is 1.85% compared to the existing FL systems.
Pengcheng Qu, Jianchun Liu, Zhiyuan Wang 0002, Qianpiao Ma, Jinyang Huang
ICPADS2
2023 Enhanced Federated Learning with Adaptive Block-wise Regularization and Knowledge Distillation
abstract
Federated Learning (FL) has emerged as an efficient distributed model training framework that enables multiple clients cooperatively to train a global model without exposing their local data in edge computing (EC). However, FL usually faces statistical heterogeneity (e.g., non-IID data) and system heterogeneity (e.g., computing and communication capabilities), resulting in poor model training performance. To deal with the above two challenges, we propose an efficient FL framework, named FedBR, which integrates the idea of block-wise regularization and knowledge distillation (KD) into the pioneer FL algorithm FedAvg, for resource-constrained edge computing. Besides, we design a heuristic algorithm (GMBS) to determine the appropriate number of model blocks for clients according to their varied data distributions, computing, and communication capabilities. Extensive experimental results show that FedBR can reduce the time cost by 19.5% and the communication cost by 27% on average compared with the other three baselines when achieving the target testing accuracy under heterogeneous settings.
Qingmin Zeng, Jianchun Liu, Hongli Xu 0001, Zhiyuan Wang 0002, Yang Xu 0020, Yangming Zhao
IWQoS2
2023 CoopFL: Accelerating federated learning with DNN partitioning and offloading in heterogeneous edge computing
Zhiyuan Wang 0002, Hongli Xu 0001, Yang Xu 0020, Zhida Jiang, Jianchun Liu
Comput. Networks5
2023 Adaptive Asynchronous Federated Learning in Resource-Constrained Edge Computing
abstract
Federated learning (FL) has been widely adopted to train machine learning models over massive data in edge computing. However, machine learning faces critical challenges, e.g., data imbalance, edge dynamics, and resource constraints, in edge computing. The existing FL solutions cannot well cope with data imbalance or edge dynamics, and may cause high resource cost. In this paper, we propose an adaptive asynchronous federated learning (AAFL) mechanism. To deal with edge dynamics, a certain fraction$\alpha$of all local updates will be aggregated by their arrival order at the parameter server in each epoch. Moreover, the system can intelligently vary the number of local updated models for global model aggregation in different epochs with network situations. We then propose experience-driven algorithms based on deep reinforcement learning (DRL) to adaptively determine the optimal value of$\alpha$in each epoch for two cases of AAFL, single learning task and multiple learning tasks, so as to achieve less completion time of training under resource constraints. Extensive experiments on the classical models and datasets show high effectiveness of the proposed algorithms. Specifically, AAFL can reduce the completion time by about 70 percent and improve the learning accuracy by about 28 percent under resource constraints, compared with the state-of-the-art solutions.
Jianchun Liu, Hongli Xu 0001, Lun Wang 0003, Yang Xu 0020, Chen Qian 0001, Jinyang Huang, He Huang 0001
IEEE Trans. Mob. Comput.1
2023 Accelerating Federated Learning With Cluster Construction and Hierarchical Aggregation
abstract
Federated learning (FL) has emerged in edge computing to address the limited bandwidth and privacy concerns of traditional cloud-based training. However, the existing FL mechanisms may lead to a long training time and consume massive communication resources. In this paper, we propose an efficient FL mechanism, namely FedCH, to accelerate FL in heterogeneous edge computing. Different from existing works which adopt the pre-defined system architecture and train models in a synchronous or asynchronous manner, FedCH will construct a special cluster topology and perform hierarchical aggregation for training. Specifically, FedCH arranges all clients into multiple clusters based on their heterogeneous training capacities. The clients in one cluster synchronously forward their local updates to the cluster header for aggregation, while all cluster headers take the asynchronous method for global aggregation. Our analysis shows that the convergence bound depends on the number of clusters and the training epochs. We propose efficient algorithms to determine the optimal number of clusters with resource budgets and then construct the cluster topology to address the client heterogeneity. Extensive experiments on both physical platform and simulated environment show that FedCH reduces the completion time by 49.5-79.5% and the network traffic by 57.4-80.8%, compared with the existing FL mechanisms.
Zhiyuan Wang 0002, Hongli Xu 0001, Jianchun Liu, Yang Xu 0020, He Huang 0001, Yangming Zhao
IEEE Trans. Mob. Comput.3
2023 Adaptive Control of Local Updating and Model Compression for Efficient Federated Learning
abstract
Data generated at the network edge can be processed locally by leveraging the paradigm of Edge Computing (EC). Aided by EC, Federated Learning (FL) has been becoming a practical and popular approach for distributed machine learning over locally distributed data. However, FL faces three critical challenges, i.e., resource constraint, system heterogeneity and context dynamics in EC. To address these challenges, we present a training-efficient FL method, termedFedLamp, by optimizing both theLocal updating frequency andmodel compression ratio in the resource-constrained EC systems. We theoretically analyze the model convergence rate and obtain a convergence upper bound related to the local updating frequency and model compression ratio. Upon the convergence bound, we propose a control algorithm, that adaptively determines diverse and appropriate local updating frequencies and model compression ratios for different edge nodes, so as to reduce the waiting time and enhance the training efficiency. We evaluate the performance ofFedLampthrough extensive simulation and testbed experiments. Evaluation results show thatFedLampcan reduce the traffic consumption by 63% and the completion time by about 52% for achieving the similar test accuracy, compared to the baselines.
Yang Xu 0020, Yunming Liao, Hongli Xu 0001, Zhen-guo Ma, Lun Wang 0003, Jianchun Liu
IEEE Trans. Mob. Comput.6
2022 Enhancing Federated Learning with Intelligent Model Migration in Heterogeneous Edge Computing
abstract
To approach the challenges of non-IID data and limited communication resource raised by the emerging federated learning (FL) in mobile edge computing (MEC), we propose an efficient framework, called FedMigr, which integrates a deep reinforcement learning (DRL) based model migration strategy into the pioneer FL algorithm FedAvg. According to the data distribution and resource constraints, our FedMigr will intelligently guide one client to forward its local model to another client after local updating, rather than directly sending the local models to the server for global aggregation as in FedAvg. Intuitively, migrating a local model from one client to another is equivalent to training it over more data from different clients, contributing to alleviating the influence of non-IID issue. We prove that FedMigr can help to reduce the parameter divergences between different local models and the global model from a theoretical perspective, even over local datasets with non-IID settings. Extensive experiments on three popular benchmark datasets demonstrate that FedMigr can achieve an average accuracy improvement of around 13%, and reduce bandwidth consumption for global communication by 42% on average, compared with the baselines.
Jianchun Liu, Yang Xu 0020, Hongli Xu 0001, Yunming Liao, Zhiyuan Wang 0002, He Huang 0001
ICDE1
2022 Enhancing Federated Learning with In-Cloud Unlabeled Data
abstract
Federated learning (FL) has been widely applied to collaboratively train deep learning (DL) models on massive end devices (i.e., clients). Due to the limited storage capacity and high labeling cost, there are always insufficient data stored and annotated on each client. Conversely, in cloud datacenters, there exist large-scale unlabeled data, which are easy to collect from public access (e.g., social media). Herein, upon the federated semi-supervised learning (FSSL) technology, we propose the Ada-FedSemi system, which leverages both on-device labeled data and in-cloud unlabeled data to boost the performance of DL models. Given the limited communication and massive quantity of the clients, in each training round, we decide to select partial clients to participate in FL, and their local models are aggregated by the parameter server (PS) to produce pseudo-labels for the unlabeled data, which are utilized to enhance the global model. Considering that the number of participating clients and the quality of pseudo-labels will have a significant impact on the training performance (e.g., efficiency and accuracy), we introduce a multi-armed bandit (MAB) based online algorithm to adaptively determine the participating fraction and confidence threshold during federated model training. Extensive experiments on benchmark models and datasets show that, given the same resource budget, the model trained by Ada-FedSemi achieves 3%-14.8 % higher test accuracy than that of the baseline methods. Besides, when achieving the same test accuracy, Ada-FedSemi saves up to 48% training cost, compared with the baselines.
Lun Wang 0003, Yang Xu 0020, Hongli Xu 0001, Jianchun Liu, Zhiyuan Wang 0002, Liusheng Huang
ICDE4
2022 Joint Data Collection and Resource Allocation for Distributed Machine Learning at the Edge
abstract
Under the paradigm of edge computing, the enormous data generated at the network edge can be processed locally. To make full utilization of these widely distributed data, we focus on an edge computing system that conducts distributed machine learning using gradient-descent based approaches. To ensure the system’s performance, there are two major challenges: how to collect data from multiple data source nodes for training jobs and how to allocate the limited resources on each edge server among these jobs. In this paper, we jointly consider the two challenges for distributed training (without service requirement), aiming to maximize the system throughput while ensuring the system’s quality of service (QoS). Specifically, we formulate the joint problem as a mixed-integer non-linear program, which is NP-hard, and propose an efficient approximation algorithm. Furthermore, we take service placement into consideration for diverse training jobs and propose an approximation algorithm. We also analyze that our proposed algorithm can achieve the constant bipartite approximation under many practical situations. We build a test-bed to evaluate the effectiveness of our proposed algorithm in a practical scenario. Extensive simulation results and testing results show that the proposed algorithms can improve the system throughput 56-69 percent compared with the conventional algorithms.
Min Chen 0033, Haichuan Wang, Zeyu Meng, Hongli Xu 0001, Yang Xu 0020, Jianchun Liu, He Huang 0001
IEEE Trans. Mob. Comput.6
2022 SAFE-ME: Scalable and Flexible Policy Enforcement in Middlebox Networks
abstract
The past decades have seen a proliferation of middlebox deployment in various scenarios, including backbone networks and cloud networks. Since flows have to traverse specific service function chains (SFCs) for security and performance enhancement, it becomes much complex for SFC routing due to routing loops, traffic dynamics and scalability requirement. The existing SFC routing solutions may consume many resources (e.g., TCAM) on the data plane and lead to massive overhead on the control plane, which decrease the scalability of middlebox networks. Due to SFC requirement and potential routing loops, solutions like traditional default paths (e.g., using ECMP) that are widely used in non-middlebox networks will no longer be feasible. In this paper, we present and implement a scalable and flexible middlebox policy enforcement (SAFE-ME) system to minimize the TCAM usage and control overhead. To this end, we design the smart tag operations for construction of default SFC paths with less TCAM rules in the data plane, and present lightweight SFC routing update with less control overhead for dealing with traffic dynamics in the control plane. We implement our solution and evaluate its performance with experiments on both physical platform (Pica8) and Programming Protocol-independent Packet Processors (P4) based data plane, as well as large-scale simulations. Both experimental and simulation results show that SAFE-ME can greatly improve scalability (e.g., TCAM cost, update delay, and control overhead) in middlebox networks, especially for large-scale clouds. For example, our system can reduce the control traffic overhead by about 85% while achieving almost the similar middlebox load, compared with state-of-the-art solutions.
Hongli Xu 0001, Peng Xi, Gongming Zhao, Jianchun Liu, Chen Qian 0001, Liusheng Huang
IEEE/ACM Trans. Netw.4
2021 Resource-Efficient Federated Learning with Hierarchical Aggregation in Edge Computing
abstract
Federated learning (FL) has emerged in edge computing to address limited bandwidth and privacy concerns of traditional cloud-based centralized training. However, the existing FL mechanisms may lead to long training time and consume a tremendous amount of communication resources. In this paper, we propose an efficient FL mechanism, which divides the edge nodes into K clusters by balanced clustering. The edge nodes in one cluster forward their local updates to cluster header for aggregation by synchronous method, called cluster aggregation, while all cluster headers perform the asynchronous method for global aggregation. This processing procedure is called hierarchical aggregation. Our analysis shows that the convergence bound depends on the number of clusters and the training epochs. We formally define the resource-efficient federated learning with hierarchical aggregation (RFL-HA) problem. We propose an efficient algorithm to determine the optimal cluster structure (i.e., the optimal value of K) with resource constraints and extend it to deal with the dynamic network conditions. Extensive simulation results obtained from our study for different models and datasets show that the proposed algorithms can reduce completion time by 34.8%-70% and the communication resource by 33.8%-56.5% while achieving a similar accuracy, compared with the well-known FL mechanisms.
Zhiyuan Wang 0002, Hongli Xu 0001, Jianchun Liu, He Huang 0001, Chunming Qiao, Yangming Zhao
INFOCOM3
2021 Communication-efficient asynchronous federated learning in resource-constrained edge computing
Jianchun Liu, Hongli Xu 0001, Yang Xu 0020, Zhen-guo Ma, Zhiyuan Wang 0002, Chen Qian 0001, He Huang 0001
Comput. Networks1
2021 Achieving high reliability and throughput in software defined networks
Xuwei Yang, Hongli Xu 0001, Jianchun Liu, Chen Qian 0001, Xingpeng Fan, He Huang 0001, Haibo Wang 0004
Comput. Networks3
2021 Incremental Server Deployment for Software-Defined NFV-Enabled Networks
abstract
Network Function Virtualization (NFV) is a new paradigm to enable service innovation through virtualizing traditional network functions. To construct a new NFV-enabled network, there are two critical requirements: minimizing server deployment cost and satisfying switch resource constraints. However, prior work mostly focuses on the server deployment cost, while ignoring the switch resource constraints (e.g., switch's flow-table size). It thus results in a large number of rules on switches and leads to massive control overhead. To address this challenge, we propose an incremental server deployment (INSD) problem for construction of scalable NFV-enabled networks. We prove that the INSD problem is NP-Hard, and there is no polynomial-time algorithm with approximation ratio of (1- ϵ)· ln m, where ϵ is an arbitrarily small value and m is the number of requests in the network. We then present an efficient algorithm with an approximation ratio of 2 · H(q · p), where q is the number of VNF's categories and p is the maximum number of requests through a switch. We evaluate the performance of our algorithm with experiments on physical platform (Pica8), Open vSwitches, and large-scale simulations. Both experimental results and simulation results show high scalability of the proposed algorithm. For example, our solution can reduce the control and rule overhead by about 88% with about 5% additional server deployment, compared with the existing solutions.
Jianchun Liu, Hongli Xu 0001, Gongming Zhao, Chen Qian 0001, Xingpeng Fan, Xuwei Yang, He Huang 0001
IEEE/ACM Trans. Netw.1
2020 Incremental Server Deployment for Scalable NFV-enabled Networks
abstract
Network Function Virtualization (NFV) is a new paradigm to enable service innovation through virtualizing traditional network functions. To construct a new NFV-enabled network, there are two critical requirements: minimizing server deployment cost and satisfying switch resource constraints. However, prior work mostly focuses on the server deployment cost, while ignoring the switch resource constraints (e.g., switch's flow-table size). It thus results in a large number of rules on switches and leads to massive control overhead. To address this challenge, we propose an incremental server deployment (INSD) problem for construction of scalable NFV-enabled networks. We prove that the INSD problem is NP-Hard, and there is no polynomial-time algorithm with approximation ratio of (1- ε) ·ln m, where ε is an arbitrarily small value and m is the number of requests in the network. We then present an efficient algorithm with an approximation ratio of 2 · H(q · p)1, where q is the number of VNF's categories and p is the maximum number of requests through a switch. We evaluate the performance of our algorithm with experiments on physical platform (Pica8), Open vSwitches, and large-scale simulations. Both experiment and simulation results show high scalability of the proposed algorithm. For example, our solution can reduce the control and rule overhead by about 88% with about 5% additional server deployment, compared with the existing solutions.
Jianchun Liu, Hongli Xu 0001, Gongming Zhao, Chen Qian 0001, Xingpeng Fan, Liusheng Huang
INFOCOM1
2020 PrePass: Load balancing with data plane resource constraints using commodity SDN switches
Haibo Wang 0004, Hongli Xu 0001, Chen Qian 0001, Juncheng Ge, Jianchun Liu, He Huang 0001
Comput. Networks5
2019 SAFE-ME: Scalable and Flexible Middlebox Policy Enforcement with Software Defined Networking
abstract
The past decades have seen a proliferation of middlebox deployment in various networks, including backbone networks and datacenters. Since network flows have to traverse specific service function chains (SFCs) for security and performance enhancement, it becomes much complex for SFC routing due to routing loops, traffic dynamics and scalability requirement. The existing SFC routing solutions may consume many resources (e.g., TCAM) on the data plane and lead to massive overhead on the control plane, which decrease the scalability of middlebox networks. Due to SFC requirement and potential routing loops, solutions like traditional default paths (e.g., using ECMP) that are widely used in non-middlebox networks will no longer be feasible. In this paper, we present and implement a scalable and flexible middlebox policy enforcement (SAFE-ME) system to minimize the TCAM usage and control overhead. To this end, we design the smart tag operations for construction of default SFC paths with less TCAM rules in the data plane, and present lightweight SFC routing update with less control overhead for dealing with traffic dynamics in the control plane. We implement our solution and evaluate its performance with experiments on both physical platform (Pica8) and Open vSwitch (OVS), as well as large-scale simulations. Both experimental and simulation results show that SAFE-ME can greatly improve scalability (e.g., TCAM cost, update delay, and control overhead) in middlebox networks. For example, our system can reduce the control traffic overhead by about 83% while achieving almost the similar middlebox load, compared with state-of-the-art solutions.
Gongming Zhao, Hongli Xu 0001, Jianchun Liu, Chen Qian 0001, Juncheng Ge, Liusheng Huang
ICNP3
2019 Reducing controller response time with hybrid routing in software defined networks
Hongli Xu 0001, Jianchun Liu, Chen Qian 0001, He Huang 0001, Chunming Qiao
Comput. Networks2