VLDB 2026 Research / reviewers in the wild / expert
Zhixiong Chen 0003
dblp:13/5568-3
· DBLP profile ↗
30ranked-venue papers
20as first author
27since 2021 · last 2026
0000-0003-4183-8857ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 28 · 18 first-author · 25 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Resilient Hierarchical Split Federated Learning over Resource-Limited Wireless Communication Systems
Chunfeng Xie, Zhixiong Chen 0003, Wenqiang Yi, Hyundong Shin, Arumugam Nallanathan |
ICC | 2 |
| 2026 | Collaborative Edge Inference for Large Language Models with Speculative Decoding
Bingjie Zhu, Zhixiong Chen 0003, Hyundong Shin, Arumugam Nallanathan |
ICC | 2 |
| 2026 | A Large Language Model-Based Decision Transformer Approach for UAV Data Collection
Zhixiong Chen 0003, Hyundong Shin, Arumugam Nallanathan |
WCNC | 1 |
| 2026 | Edge Inference for Large Language Models With Pipeline Parallelism and BatchingabstractEdge computing enables distributed inference for computation-intensive applications. However, the autoregressive nature and large model size of large language models (LLMs) pose challenges for their deployment in wireless edge networks. Existing edge inference methods mainly assume an equivalent delay for each token generation step, which fails to capture the dynamic computational and memory overhead incurred during the decoding process. This paper proposes a latency-sensitive wireless edge inference framework for LLMs, where tasks are grouped into multiple batches and processed in parallel by partitioning the LLM into multiple pipeline stages across heterogeneous edge GPUs. An accurate latency model is established, where the latency of each token generation step increases during the autoregressive generation process. Based on this model, the end-to-end inference latency is minimized, which is formulated as a joint optimization problem of bandwidth allocation, model partitioning, and batch scheduling subject to heterogeneous GPU memory constraints. To solve this NP-hard problem with coupled variables, we develop a polynomial-time alternating optimization algorithm that iteratively optimizes model partitioning and batch scheduling via dynamic programming. The closed-form solutions of wireless bandwidth allocation are derived. Extensive simulations show that our approach reduces latency by up to 42.1% versus state-of-the-art baselines across diverse edge scenarios. Jie Jiang 0019, Zhixiong Chen 0003, Hyundong Shin, Arumugam Nallanathan |
IEEE Trans. Commun. | 2 |
| 2026 | Distribution Deviation-Aware Split Federated Learning in Resource-Limited Wireless NetworksabstractThe escalating complexity of deep neural networks introduces substantial challenges to deploying federated learning (FL) in resource-limited edge environments. To address these limitations, split federated learning (SFL) has emerged as a promising paradigm, alleviating client-side computational and communication burdens via strategic model splitting, and periodically aggregating client-side and server-side models consistent with the principles of FL. Nevertheless, existing SFL frameworks encounter significant performance degradation arising from data heterogeneity and imbalance, client heterogeneity, as well as constrained wireless resources. To overcome these issues, this paper introduces a novel data distribution deviation-aware split federated learning (DA-SFL) framework. DA-SFL dynamically adjusts aggregation weights according to the deviation of clients’ data distributions from a global distribution, effectively mitigating biases induced by data imbalance and heterogeneity. Furthermore, we theoretically establish the convergence bound of DA-SFL under a non-convex loss function setting, demonstrating that minimizing the data deviation in each training round enhances learning efficacy. Motivated by this, we formulate a mixed-integer nonlinear programming to optimize learning performance under long-term energy constraints. Leveraging the Lyapunov optimization framework, we decompose the problem into a series of tractable subproblems in each learning round, and propose efficient algorithms to find the client scheduling, adaptive cut layer selection, bandwidth allocation, and aggregation weighting policies. Extensive experimental evaluations conducted on Fashion-MNIST, CIFAR-10, and CINIC-10 datasets across diverse scenarios of data heterogeneity and imbalance demonstrate that DA-SFL significantly outperforms baselines regarding test accuracy, time and energy efficiency, while exhibiting notable robustness and scalability. Chunfeng Xie, Zhixiong Chen 0003, Wenqiang Yi, Hyundong Shin, Arumugam Nallanathan |
IEEE Trans. Commun. | 2 |
| 2026 | Efficient LLM Inference Over Heterogeneous Edge Networks With Speculative Decoding
Bingjie Zhu, Zhixiong Chen 0003, Hyundong Shin, Arumugam Nallanathan |
IEEE Trans. Commun. | 2 |
| 2026 | Enabling Efficient Large Language Model Inference Over Wireless Networks With CachingabstractWith the proliferation of large language models (LLMs), cloud-based LLM serving mechanisms may cause network congestion and high serving delay. Edge computing offers a solution to alleviate backhaul pressure and reduce serving delay by deploying LLMs on edge servers and providing LLM inference services in users’ proximity. However, user accuracy requirements vary over time, and mismatches between these requirements and the deployed LLMs at the edge may lead to inefficient resource usage and increased serving delay. To address this, we formulate a joint LLM caching, inference task scheduling, and network resource allocation problem to minimize LLM serving delay under unknown time-varying user accuracy requirements. To solve the problem, we first derive closed-form solutions for optimal computation and communication resource allocation under any LLM caching and task scheduling policies. Then, we employ an improved branch-and-bound algorithm to obtain optimal task scheduling policies under any LLM caching strategies. Finally, we propose an improved double deep Q-network (DDQN)-based algorithm to determine the LLM caching decisions. It incorporates a state coding and action aggregation (SCAA) mechanism within the deep neural networks (DNNs) of the traditional DDQN. The SCAA-DNNs involve an input-layer gating mechanism to encode users’ request states for LLMs and a two-layer output architecture that dynamically aggregates LLM caching actions to generate the corresponding state-action values, thereby improving learning efficiency and accelerating convergence in large discrete action spaces. Experimental results show that the proposed scheme could rapidly converge and reduce average user delay by up to 20.8% compared to benchmarks. Bingjie Zhu, Zhixiong Chen 0003, Hyundong Shin, Arumugam Nallanathan |
IEEE Trans. Wirel. Commun. | 2 |
| 2025 | DevSFL: Deviation-Aware Split Federated Learning in Resource-Constrained Wireless NetworksabstractIn mobile wireless networks, data heterogeneity and resource constraints cause performance degradation in machine learning tasks on edge clients. To alleviate these issues, we propose a novel deviation-aware split federated learning (DevSFL) framework, which adopts an adaptive aggregation weight determination method for mitigating the effects of data heterogeneity across local datasets and improving overall learning performance. Leveraging Lyapunov optimization, we formulate a comprehensive optimization problem including client scheduling, cut layer selection, bandwidth allocation, and weight decisionmaking to enhance resource utilization and energy efficiency. To tackle this problem, we employ a sample average approximation based algorithm and a dichotomy method for optimizing cut layer selection and bandwidth allocation policies, respectively. Furthermore, a set expansion algorithm is employed to find the optimal client subset. Additionally, we introduce a deviationaware algorithm specifically designed to refine the weighting policy. Comparative analysis with benchmark schemes reveals that our proposed DevSFL framework not only achieves higher accuracy within fewer rounds but also significantly reduces the time required to reach a predefined accuracy level, thereby demonstrating the effectiveness of our proposed algorithms. Chunfeng Xie, Zhixiong Chen 0003, Wenqiang Yi, Hyundong Shin, Arumugam Nallanathan |
ICC | 2 |
| 2025 | Joint Caching and Inference for Large Language Models in Wireless NetworksabstractTo reduce the serving delay of large language model (LLM)-based applications, the edge-based LLM serving mechanism offers a promising solution by caching LLMs at the edge to provide LLM inference services closer to users. Motivated by this, we propose an edge-based LLM caching and inference framework to support low-delay LLM-based services. Based on the framework, we formulate a joint LLM caching, inference task scheduling, and computation resource allocation optimization problem to minimize LLM serving delay, where time-varying LLM popularity is considered. Given an LLM caching policy, we first obtain the optimal solution for the computation resource allocation and task scheduling by using traditional optimization methods. Then, we propose an improved double deep Q-network (IDDQN) algorithm that effectively learns the optimal LLM caching strategy under unknown LLM popularity. The IDDQN algorithm integrates a state coding and action aggregation (SCAA) mechanism in the deep neural network structure, enabling it to efficiently capture users' preferences for LLMs and mitigate the slow convergence issues due to the large action space. Simulation results indicate that the proposed scheme achieves both lower average user delay and faster convergence than other benchmarks. Bingjie Zhu, Zhixiong Chen 0003, Hyundong Shin, Arumugam Nallanathan |
ICC | 2 |
| 2025 | Generative Diffusion Model-Based Variational Inference for MIMO Channel EstimationabstractEfficient and accurate channel estimation with low pilot overhead is essential for massive multiple-input multiple-output (MIMO) wireless communication systems to achieve high spectral and energy efficiency. This work proposes a novel variational inference method for channel estimation by utilizing the generative diffusion model as a prior. Specifically, we first train a generative diffusion model to learn the score, i.e., the gradient of the log-prior distribution, of MIMO channels to serve as a prior in the channel estimation process. The training process is unsupervised and does not rely on specific pilot structures and signal-to-noise ratios (SNRs). Thus, the learned prior is generalizable and can be directly used for channel estimation under different pilot signals and SNRs without requiring re-training. Then, we propose a variational inference method to infer the posterior distribution of the MIMO channel under given pilots and received measurements by incorporating the learned prior. Finally, we estimate the MIMO channels by sampling from the derived posterior distribution. Our simulations under various wireless propagation environments and antenna architectures demonstrate that the proposed approach achieves over 5 dB reduction in normalized mean square error and faster channel recovery compared to state-of-the-art channel estimators. Additionally, the proposed approach exhibits robust estimation performance when the test channel distribution shifts from the training distribution, even outperforming the benchmarks without distribution shifts. Zhixiong Chen 0003, Hyundong Shin, Arumugam Nallanathan |
IEEE Trans. Commun. | 1 |
| 2025 | Adaptive Semi-Asynchronous Federated Learning Over Wireless NetworksabstractOwing to the heterogeneous computation and communication capabilities among clients, the synchronous model aggregation in wireless federated learning (FL) is susceptible to the straggler effect and exhibits low learning efficiency, while asynchronous aggregation encounters delayed gradients that lead to convergence errors and learning performance degradation. To address these obstacles, this work proposes an adaptive semi-asynchronous FL (ASAFL) approach to incorporate the strengths of synchronous and asynchronous FL while mitigating their inherent drawbacks. Specifically, the edge server dynamically adjusts the synchronous degree, i.e., the number of local gradients aggregated in each round, to strike a balance between learning latency and accuracy. Recognizing that data heterogeneity among clients may induce biased global model updating, we propose calibrating the global update by leveraging historical gradients received at the edge server from clients. Following that, we theoretically investigate the impact of synchronous degrees in different rounds on the convergence bound of ASAFL. The results imply that allocating more learning time to the later learning stages to increase the synchronous degree contributes to better learning performance. Based on this, we develop an adaptive synchronous degree control and resource allocation algorithm to enhance the learning performance of FL while adhering to the overall learning latency and wireless resources constraint. Numerical results on the MNIST and CIFAR-10 datasets demonstrate that the proposed approach is capable of attaining faster convergence speed and higher learning accuracy compared to the benchmark FL algorithms. Zhixiong Chen 0003, Wenqiang Yi, Hyundong Shin, Arumugam Nallanathan |
IEEE Trans. Commun. | 1 |
| 2025 | Tackling Class Imbalance and Client Heterogeneity for Split Federated Learning in Wireless NetworksabstractAs the complexity of deep neural networks escalates, traditional federated learning (FL) frameworks increasingly struggle since the training overhead of the full model is costly for resource-limited clients. In addition, the class imbalance among local datasets and client heterogeneity may lead to significant deterioration in learning performance. To address these challenges, we first propose a novel wireless split federated learning (SFL) framework to enhance learning efficiency and performance in resource-constrained networks, which adaptively splits the global model between the clients and server to alleviate the computation burden for clients. Then, we theoretically analyze how the client sampling and wireless network parameters impact on the convergence bound. Based on the analysis, we identify the extent of class imbalance that significantly impacts learning performance. Inspired by this, we formulate an optimization problem to strike a balance between latency and performance by jointly optimizing the client selection, model splitting, and bandwidth allocation policies. To solve this problem, we introduce a latency and class imbalance-aware double greedy algorithm to obtain client scheduling policy. Additionally, bisection-enabled optimal bandwidth allocation and model splitting algorithms are developed to adaptively determine bandwidth allocation and model splitting policies, respectively. Extensive experimental results demonstrate that our approach significantly reduces latency and enhances learning performance. Chunfeng Xie, Zhixiong Chen 0003, Wenqiang Yi, Hyundong Shin, Arumugam Nallanathan |
IEEE Trans. Wirel. Commun. | 2 |
| 2024 | Gradient Compensation Enabled Federated Learning for Unreliable Wireless LinksabstractWireless federated learning (FL) faces significant challenges due to limited wireless resources and unreliable channels. To cope with these challenges, this work proposes a gradient compensation-based FL approach (FL-GC), in which the edge server estimates the local gradients of transmission failure and unselected clients by first-order Taylor approximation based on previously received local gradients. We then theoretically analyze the convergence bound, which reveals that selecting clients with large local gradient staleness helps reduce the estimation error and improve learning performance. Based on this, we jointly optimize the client selection and resource allocation strategies to enhance the FL performance under resource-limited wireless networks. Simulation results under a typical data heterogeneity scenario demonstrate the efficacy of our proposed scheme in mitigating the adverse effects of unreliable transmission and limited resources. It improves 7.34% model accuracy compared to the considered benchmarks and is able to save 42.5% training time to achieve the target accuracy. Zhixiong Chen 0003, Wenqiang Yi, Yun Hee Kim, Arumugam Nallanathan |
GLOBECOM | 1 |
| 2024 | Multi-Agent Reinforcement Learning-Based Digital Twin Migration Over Wireless NetworksabstractTo reduce the synchronization latency in digital twin (DT)-enabled wireless edge networks, the DT migration provides an efficient roaming solution among edge servers by following users' trajectories. In this work, we formulate a joint DT migration, communication and computation resource management problem to minimize the data synchronization latency, where the time-varying network states and user mobility are considered. By decoupling edge servers under a deterministic migration strategy, we first derive the optimal communication and computation resource management policies at each server using convex optimization methods. For the DT migration problem between different servers, we transform it as a decentralized partially observable Markov decision process (Dec-POMDP). Then, we propose a novel agent-contribution-enabled multiagent reinforcement learning (AC-MARL) algorithm to enable distributed DT migration for users, in which the counterfactual baseline method is adopted to characterize the contribution of each agent and facilitate cooperation among agents. Simulation results show that the proposed DT migration scheme is able to reduce 30% data synchronization latency for users compared to the benchmark schemes. Zhixiong Chen 0003, Wenqiang Yi, Arumugam Nallanathan |
ICC | 1 |
| 2024 | Fast Wireless Federated Learning with Adaptive Synchronous Degree ControlabstractThis work proposes an adaptive semi-asynchronous federated learning (FL) approach, namely ASAFL, to incorporate the strengths of synchronous and asynchronous FL while mitigating their inherent drawbacks. Specifically, the edge server dynamically adjusts the synchronous degree, i.e., the number of local gradients aggregated in each round, to strike a balance between learning latency and accuracy. Recognizing that data heterogeneity among clients may induce biased global model updating, we propose calibrating the global update by leveraging historical gradients received at the edge server from clients. Following that, we experimentally revealed that allocating more learning time to the later learning stages to increase the synchronous degree contributes to better learning performance. Inspired by this, we develop an adaptive synchronous degree control and resource allocation algorithm to enhance the learning performance of FL while adhering to the overall learning latency and wireless resources constraint. Numerical results demonstrate that the proposed approach is capable of attaining faster convergence speed and higher learning accuracy compared to the benchmark FL algorithms. Zhixiong Chen 0003, Wenqiang Yi, Hyundong Shin, Arumugam Nallanathan |
VTC Spring | 1 |
| 2024 | Cost-Efficient Cooperative Video Caching Over Edge NetworksabstractCooperative caching has emerged as an efficient way to alleviate backhaul traffic and enhance user experience by proactively prefetching popular videos at the network edge. However, it is challenging to achieve the optimal design of video caching, sharing, and delivery within storage-limited edge networks due to the growing diversity of videos, unpredictable video requirements, and dynamic user preferences. To address this challenge, this work explores cost-efficient cooperative video caching via video compression techniques while considering unknown video popularity. Firstly, we formulate the joint video caching, sharing, and delivery problem to capture a balance between user delay and system operative cost under unknown time-varying video popularity. To solve this problem, we develop a two-layer decentralized reinforcement learning algorithm, which effectively reduces the action space and tackles the coupling among video caching, sharing, and delivery decisions compared to the conventional algorithms. Specifically, the outer layer produces the optimal decisions for video caching and communication resource allocation by employing a multi-agent deep deterministic policy gradient algorithm. Meanwhile, the optimal video sharing and computation resource allocation are determined in each agent’s inner layer using the alternating optimization algorithm. Numerical results show that the proposed algorithm outperforms benchmarks in terms of the cache hit rate, delay of users and system operative cost, and effectively strikes a trade-off between system operative cost and users’ delay. Bingjie Zhu, Wenqiang Yi, Zhixiong Chen 0003, Arumugam Nallanathan |
IEEE Internet Things J. | 4 |
| 2024 | Efficient Wireless Federated Learning With Partial Model AggregationabstractThe data heterogeneity across clients and the limited communication resources, e.g., bandwidth and energy, are two of the main bottlenecks for wireless federated learning (FL). To tackle these challenges, we first devise a novel FL framework with partial model aggregation (PMA). This approach aggregates the lower layers of neural networks, responsible for feature extraction, at the parameter server while keeping the upper layers, responsible for complex pattern recognition, at clients for personalization. The proposed PMA-FL is able to address the data heterogeneity and reduce the transmitted information in wireless channels. Then, we derive a convergence bound of the framework under a non-convex loss function setting to reveal the role of unbalanced data size in the learning performance. On this basis, we maximize the scheduled data size to minimize the global loss function through jointly optimize the client selection, bandwidth allocation, computation and communication time division policies with the assistance of Lyapunov optimization. Our analysis reveals that the optimal time division is achieved when the communication and computation parts of PMA-FL have the same power. We also develop a bisection method to solve the optimal bandwidth allocation policy and use the set expansion algorithm to address the client scheduling policy. Compared with the benchmark schemes, the proposed PMA-FL improves 3.13% and 11.8% absolute accuracy on two typical datasets with heterogeneous data distribution settings, i.e., MINIST and CIFAR-10, respectively. In addition, the proposed joint dynamic client selection and resource management approach achieve slightly higher accuracy than the considered benchmarks, but they provide a satisfactory energy and time reduction: 29% energy or 20% time reduction on the MNIST; and 25% energy or 12.5% time reduction on the CIFAR-10. Zhixiong Chen 0003, Wenqiang Yi, Hyundong Shin, Arumugam Nallanathan, Geoffrey Ye Li |
IEEE Trans. Commun. | 1 |
| 2024 | Robust Federated Learning for Unreliable and Resource-Limited Wireless NetworksabstractFederated learning (FL) is an efficient and privacy-preserving distributed learning paradigm that enables massive edge devices to train machine learning models collaboratively. Although various communication schemes have been proposed to expedite the FL process in resource-limited wireless networks, the unreliable nature of wireless channels was less explored. In this work, we propose a novel FL framework, namely FL with gradient recycling (FL-GR), which recycles the historical gradients of unscheduled and transmission-failure devices to improve the learning performance of FL. To reduce the hardware requirements for implementing FL-GR in the practical network, we develop a memory-friendly FL-GR that is equivalent to FL-GR but requires low memory of the edge server. We then theoretically analyze how the wireless network parameters affect the convergence bound of FL-GR, revealing that minimizing the average square of local gradients’ staleness (AS-GS) helps improve the learning performance. Based on this, we formulate a joint device scheduling, resource allocation and power control optimization problem to minimize the AS-GS for global loss minimization. To solve the problem, we first derive the optimal power control policy for devices and transform the AS-GS minimization problem into a bipartite graph matching problem. Through detailed analysis, we further transform the bipartite matching problem into an equivalent linear program which is convenient to solve. Extensive simulation results on three real-world datasets (i.e., MNIST, CIFAR-10, and CIFAR-100) verified the efficacy of the proposed methods. Compared to the FL algorithms without gradient recycling, FL-GR is able to achieve higher accuracy and fast convergence speed. In addition, the proposed device scheduling and resource allocation algorithm also outperforms the benchmarks in accuracy and convergence speed. Zhixiong Chen 0003, Wenqiang Yi, Yuanwei Liu, Arumugam Nallanathan |
IEEE Trans. Wirel. Commun. | 1 |
| 2024 | Exploring Representativity in Device Scheduling for Wireless Federated LearningabstractExisting device scheduling works in wireless federated learning (FL) mainly focused on selecting the devices with maximum gradient norm or loss function and require all devices to perform local training in each round. This may produce extra training costs and schedule devices with similar data statistics, thus degrading learning performance. To mitigate these problems, we first theoretically characterize the convergence behaviour of the considered FL system, finding that the learning performance is degraded by the difference between the aggregated gradient of scheduled devices and the full participation gradient. Inspired by this, we propose to find a subset of representative devices and the corresponding pre-device stepsizes to approximate the full participation aggregated gradient. Considering the limited wireless bandwidth, we formulate a problem to capture the trade-off between representativity and latency by optimizing device scheduling and bandwidth allocation policies. Our analysis reveals optimal bandwidth allocation is achieved when all scheduled devices have the same latency. Then, by proving the non-monotone submodularity of the problem, we develop a double greedy algorithm to solve the device scheduling policy. To avoid the local training of unscheduled devices, we utilize the historical gradient information of devices to estimate the current gradient for device scheduling design. Compared to existing scheduling algorithms, the proposed representativity-aware device scheduling algorithm improves 6.7% and 4.02% accuracies on two typical datasets under heterogeneous local data distributions, i.e., MNIST and CIFAR-10, respectively. In addition, the proposed latency- and representativity-aware scheduling algorithm saves over 16% and 12% training time for MNIST and CIFAR-10 datasets than the scheduling algorithms based on either latency and representativity individually. Zhixiong Chen 0003, Wenqiang Yi, Arumugam Nallanathan |
IEEE Trans. Wirel. Commun. | 1 |
| 2024 | Adaptive Model Pruning for Communication and Computation Efficient Wireless Federated LearningabstractMost existing wireless federated learning (FL) studies focused on homogeneous model settings where devices train identical local models. In this setting, the devices with poor communication and computation capabilities may delay the global model update and degrade the performance of FL. Moreover, in the homogenous model settings, the scale of the global model is restricted by the device with the lowest capability. To tackle these challenges, this work proposes an adaptive model pruning-based FL (AMP-FL) framework, where the edge server dynamically generates sub-models by pruning the global model for devices’ local training to adapt their heterogeneous computation capabilities and time-varying channel conditions. Since the involvement of diverse structures of devices’ sub-models in the global model updating may negatively affect the training convergence, we propose compensating for the gradients of pruned model regions by devices’ historical gradients. We then introduce an age of information (AoI) metric to characterize the staleness of local gradients and theoretically analyze the convergence behaviour of AMP-FL. The convergence bound suggests scheduling devices with large AoI of gradients and pruning the model regions with small AoI for devices to improve the learning performance. Inspired by this, we define a new objective function, i.e., the average AoI of local gradients, to transform the inexplicit global loss minimization problem into a tractable one for device scheduling, model pruning, and resource block (RB) allocation design. Through detailed analysis, we derive the optimal model pruning strategy and transform the RB allocation problem into equivalent linear programming that can be effectively solved. Experimental results demonstrate the effectiveness and superiority of the proposed approaches. The proposed AMP-FL is capable of achieving 1.9x and 1.6x speed up for FL on MNIST and CIFAR-10 datasets in comparison with the FL schemes with homogeneous model settings. Zhixiong Chen 0003, Wenqiang Yi, Hyundong Shin, Arumugam Nallanathan |
IEEE Trans. Wirel. Commun. | 1 |
| 2023 | Efficient Wireless Federated Learning with Adaptive Model PruningabstractFor wireless federated learning (FL), this work proposes an adaptive model pruning-based FL (AMP-FL) frame-work, where the edge server dynamically generates sub-models by pruning the global model to adapt devices' heterogeneous computation capabilities and time-varying wireless channel conditions. To mitigate the negative effect of different structures of sub-models on learning convergence, this work designs a new compensating strategy for the pruned regions of sub-models via historical gradients. Since the freshness of gradients dominates the convergence speed, this work also defines an age of information (AoI) metric to characterize the staleness of the regions of the local gradients. Based on the compensating strategy, we formulate a joint device scheduling, model pruning, and resource block allocation optimization problem to minimize the average AoI for local gradients. To solve this problem, we theoretically derive an optimal model pruning scheme. After that, we transform the original problem into equivalent linear programming that can be solved with polynomial time complexity. Simulation results on the CIFAR-IO dataset show that the proposed AMP-FL outperforms the benchmark schemes with faster convergence speed and over 7% learning accuracy improvement. Zhixiong Chen 0003, Wenqiang Yi, Sangarapillai Lambotharan, Arumugam Nallanathan |
GLOBECOM | 1 |
| 2023 | Communication-Efficient Federated Learning with Heterogeneous DevicesabstractThe conventional model aggregation-based federated learning (FL) approaches require all local models to have the same architecture and fail to support practical scenarios with heterogeneous local models. Moreover, the frequent model exchange is costly for resource-limited wireless networks since modern deep neural networks usually have over-million parameters. To tackle these challenges, we first propose a novel knowledge-aided FL (KFL) framework, which aggregates light high-level data features, namely knowledge, in the per-round learning process. The KFL allows devices to design their machine learning models independently and reduces the communication overhead in the training process. We then experimentally show that different temporal device scheduling patterns lead to considerably different learning performance. With this insight, we formulate a stochastic optimization problem for joint device scheduling and bandwidth allocation under limited devices' energy budgets and develop an efficient online algorithm to achieve an energy-learning trade-off in the learning process. Experimental results on the CIFAR-10 dataset show that the proposed KFL can reduce over 87% communication overhead while achieving better learning performance than the baselines. In addition, the proposed device scheduling algorithm converges faster than benchmark scheduling schemes. Zhixiong Chen 0003, Wenqiang Yi, Yuanwei Liu, Arumugam Nallanathan |
ICC | 1 |
| 2023 | Is Partial Model Aggregation Energy-Efficient for Federated Learning Enabled Wireless Networks?abstractThis work aims to address two of the main challenges for federated learning (FL), i.e., the limited communication resources and the data heterogeneity across devices. To this end, we first devise a novel FL framework with partial model aggregation (PMA), which only aggregates the lower layers of neural networks responsible for feature extraction while the upper layers corresponding to complex pattern recognition remain at devices for personalization. This design is able to address the data heterogeneity and reduce the transmitted information in wireless channels. Then, we maximize the scheduled data sample volume by joint optimizing the device scheduling, bandwidth allocation, computation and communication time division. Specifically, our analysis reveals that the optimal time division is achieved when the communication and computation parts of PMA-FL have the same power. We also develop a bisection method to solve the optimal bandwidth allocation policy and use the set expansion algorithm to address the optimal device scheduling. Experimental results on the CIFAR-10 dataset show that the proposed PMA-FL improves 11.6% accuracy compared with the state-of-art benchmarks, and the proposed joint dynamic device scheduling and resource optimization approach achieves slightly higher accuracy than the considered benchmarks but reduced 25% energy or 12.5% time budgets. Zhixiong Chen 0003, Wenqiang Yi, Arumugam Nallanathan, Geoffrey Ye Li |
ICC | 1 |
| 2023 | Convergence Analysis for Wireless Federated Learning with Gradient RecyclingabstractHow to tackle the unreliability in wireless channels is critical for federated learning (FL). To solve this problem, we propose a novel FL framework, namely FL with gradient recycling (FL-GR), which recycles the historical gradients of unscheduled and transmission-failure devices to improve the learning performance of FL. Based on the proposed FL-GR, we theoretically analyze how the wireless network parameters affect the convergence bound of FL-GR, revealing that scheduling devices with large staleness and increasing their transmit power in each round helps improve learning performance. Simulation results on MNIST and CIFAR-10 show that FL-GR is able to achieve higher accuracy and fast convergence speed than conventional FL algorithms without gradient recycling. Zhixiong Chen 0003, Wenqiang Yi, Yuanwei Liu, Arumugam Nallanathan |
IWCMC | 1 |
| 2023 | Knowledge-Aided Federated Learning for Energy-Limited Wireless NetworksabstractThe conventional model aggregation-based federated learning (FL) approach requires all local models to have the same architecture, which fails to support practical scenarios with heterogeneous local models. Moreover, the frequent model exchange is costly for resource-limited wireless networks since modern deep neural networks usually have over a million parameters. To tackle these challenges, we first propose a novel knowledge-aided FL (KFL) framework, which aggregates light high-level data features, namely knowledge, in the per-round learning process. This framework allows devices to design their machine-learning models independently and reduces the communication overhead in the training process. We then theoretically analyze the convergence bound of the proposed framework under a non-convex loss function setting, revealing that scheduling more data volume in each round helps to improve the learning performance. In addition, large data volume should be scheduled in early rounds if the total scheduled data volume during the entire learning course is fixed. Inspired by this, we define a new objective function, i.e., the weighted scheduled data sample volume, to transform the inexplicit global loss minimization problem into a tractable one for device scheduling, bandwidth allocation, and power control. To deal with unknown time-varying wireless channels, we transform the considered problem into a deterministic problem for each round with the assistance of the Lyapunov optimization framework. Then, we derive the optimal bandwidth allocation and power control solution by convex optimization techniques. We also develop an efficient online device scheduling algorithm to achieve an energy-learning trade-off in the learning process. Experimental results on two typical datasets (i.e., MNIST and CIFAR-10) under highly heterogeneous local data distributions show that the proposed KFL is capable of reducing over 99% communication overhead while achieving better learning performance than the conventional model aggregation-based algorithms. In addition, the proposed device scheduling algorithm converges faster than the benchmark scheduling schemes. Zhixiong Chen 0003, Wenqiang Yi, Yuanwei Liu, Arumugam Nallanathan |
IEEE Trans. Commun. | 1 |
| 2022 | Dynamic Task Software Caching-Assisted Computation Offloading for Multi-Access Edge ComputingabstractIn multi-access edge computing (MEC), most existing task software caching works focus on statically caching data at the network edge, which may hardly preserve high reusability due to the time-varying user requests in practice. To this end, this work considers dynamic task software caching at the MEC server to assist users’ task execution. Specifically, we formulate a joint task software caching update (TSCU) and computation offloading (COMO) problem to minimize users’ energy consumption while guaranteeing delay constraints, where the limited cache size and computation capability of the MEC server, as well as the time-varying task demand of users are investigated. This problem is proved to be non-deterministic polynomial-time hard, so we transform it into two sub-problems according to their temporal correlations, i.e., the real-time COMO problem and the Markov decision process-based TSCU problem. We first model the COMO problem as a multi-user game and propose a decentralized algorithm to address its Nash equilibrium solution. We then propose a double deep Q-network (DDQN)-based method to solve the TSCU policy. To reduce the computation complexity and convergence time, we provide a new design for the deep neural network (DNN) in DDQN, named state coding and action aggregation (SCAA). In SCAA-DNN, we introduce a dropout mechanism in the input layer to code users’ activity states. Additionally, at the output layer, we devise a two-layer architecture to dynamically aggregate caching actions, which is able to solve the huge state-action space problem. Simulation results show that the proposed solution outperforms existing schemes, saving over 12% energy, and converges with fewer training episodes. Zhixiong Chen 0003, Wenqiang Yi, Atm Shafiul Alam, Arumugam Nallanathan |
IEEE Trans. Commun. | 1 |
| 2021 | Code Caching-Assisted Computation Offloading and Resource Allocation for Multi-User Mobile Edge ComputingabstractUtilizing the data caching technology to reduce data transmission is a promising technique for improving the performance of mobile edge computing (MEC), because the delay and energy consumption produced by data transmission constitute the dominant cost of task execution in MEC. Besides, computation tasks generally consist of input parameters, executive codes, and computation results. The executive codes are fixed and can output difference computation results under different input parameters. Motivated by this, we consider to proactively cache executive codes of tasks at the MEC server to reduce the weighted sum of task execution delay and users’ energy consumption. Aiming at establishing optimal system design, we formulate the problem as a non-linear programming problem which involves jointly optimizing the executive code caching strategy, computation offloading decision, wireless resource allocation, and computing resource allocation. We propose to find the optimal solution by employing an alternating optimization framework. The optimal wireless resource and computing resource allocation problem are firstly addressed by utilizing convex optimization technology. Then, a dynamic programming-based algorithm has been developed to achieve the optimal executive code caching and computation offloading strategies. Extensive simulation results show that the proposed scheme operates well and can substantially reduce the system cost over other benchmark schemes. Zhixiong Chen 0003, Zhaokun Zhou, Chen Chen 0037 |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2020 | Dynamic Task Caching and Computation Offloading for Mobile Edge ComputingabstractMobile edge computing (MEC) provides information technology and cloud-computing capabilities within the network edge close to mobile users, thereby addressing computing demand for users. However, the energy consumption of data uploading from users to the MEC server makes it hard to meet users' demand in some specific applications, i.e., interactive gaming and augmented reality. Motivated by this, we integrate a task caching mechanism into computation offloading technique. Specifically, it allows the MEC server to proactively cache some tasks and users to offload their tasks to the MEC server. Since the limited storage capacity and task demands for users are changing dynamically, which tasks are cached has to be judiciously decided to maximize the MEC system performance. The objective of this paper is to minimize the system cost, which is defined as the average total user energy consumption of all time slots. By formulating the problem as an integer-programming, we propose to find the optimal solution with two steps. Through which we have obtained the optimal online computation offloading and task cache update strategy. Simulation results show that in comparison with the other two baselines, the proposed scheme can effectively reduce the system cost. Zhixiong Chen 0003, Zhaokun Zhou |
GLOBECOM | 1 |
| 2019 | Integrated Task Caching, Computation Offloading and Resource Allocation for Mobile Edge ComputingabstractApplications with more sensitive delay and larger data volumes, such as interactive gaming and augmented reality, have become popular recently. Computation offloading is expected as a promising technique to meet low latency for mobile users. However, computation offloading requires communication between mobile users and the mobile edge computing (MEC) server, the delay and energy consumption caused by the transmission are considerable expenses for users. Motivated by this, we consider joint computation offloading and task caching optimization in a cellular network where users can proactively cache and offload their tasks at the MEC server. The objective of this paper is to minimize the system cost, which is defined as the weighted sum of task execution delay and energy consumption for all users. By formulating the problem as mixed-integer non-linear programming, we propose to find the optimal solution by three steps. Through which we have obtained the optimal computing resource allocation, the optimal task caching scheme and an algorithm which yields the optimal computation offloading scheme. Simulation results show that in comparison to the other three benchmark methods, the proposed scheme can effectively reduce the system cost. Zhixiong Chen 0003, Zhengchuan Chen, Yunjian Jia |
GLOBECOM | 1 |
| 2019 | Residual Energy-Aware Caching in Mobile D2D Cellular NetworkabstractCaching popular contents at the mobile devices is a promising technique to alleviate the backhaul data rate demand. Since both file placement and data exchange among mobile devices consume energy, the energy status of devices has a significant effect on the caching utility of the whole system. This work considers the caching optimization in a cellular network where mobile devices are served by one base station (BS). As the devices can collect the file segments from the local storage, via device-to-device (D2D) links, and via a cellular link, we aim at minimizing the percentage of file segment that should be collected from the BS by optimizing the file placement scheme at devices to improve caching performance. Due to the difficulty of solving the optimal caching problem, we propose a residual energy-aware file placement algorithm based on the popularity distribution of contents and causality of energy arrival. Simulation results show that in comparison to other two conventional caching methods, the proposed algorithm can effectively reduce the percentage of file segments that collected from the BS. Zhixiong Chen 0003, Zhengchuan Chen, Yunjian Jia, Liang Liang 0002 |
ICC | 1 |