Dingzhu Wen

dblp:176/2716 · DBLP profile ↗
← Back
53ranked-venue papers
13as first author
44since 2021 · last 2026
0000-0003-0538-5811ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 48 · 13 first-author · 40 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Robust Edge Inference with Graph Neural Networks under Channel Aging
Wenjie Long, Zhanwei Wang, Mingyao Cui, Dingzhu Wen, Min Sheng
ICC5
2026 Parameter-efficient Large AI Model Co-inference at Multi-cluster Edge Networks
Zhonghao Lyu, Xiaowen Cao 0001, Dingzhu Wen, Yuanhao Cui, Zhaohui Yang 0001, Jie Xu 0002, Shuguang Cui
ICC3
2026 FedLoDrop: Federated LoRA With Dropout for Generalized LLM Fine-Tuning
abstract
Fine-tuning (FT) large language models (LLMs) is crucial for adapting general-purpose models to specific tasks, enhancing accuracy and relevance with minimal resources. To further enhance generalization ability while reducing training costs, this paper proposes Federated LoRA with Dropout (FedLoDrop), a new framework that applies dropout to the rows and columns of the trainable matrix in Federated LoRA. A generalization error bound and convergence analysis under sparsity regularization are obtained, which elucidate the fundamental trade-off between underfitting and overfitting. The error bound reveals that a higher dropout rate increases model sparsity, thereby lowering the upper bound of pointwise hypothesis stability (PHS). While this reduces the gap between empirical and generalization errors, it also incurs a higher empirical error, which, together with the gap, determines the overall generalization error. On the other hand, though dropout reduces communication costs, deploying FedLoDrop at the network edge still faces challenges due to limited network resources. To address this issue, an optimization problem is formulated to minimize the upper bound of the generalization error, by jointly optimizing the dropout rate and resource allocation subject to the latency and per-device energy consumption constraints. To solve this problem, a branch-and-bound (B&B)-based method is proposed to obtain its globally optimal solution. Moreover, to reduce the high computational complexity of the B&B-based method, a penalized successive convex approximation (P-SCA)-based algorithm is proposed to efficiently obtain its high-quality suboptimal solution. Finally, numerical results demonstrate the effectiveness of the proposed approach in mitigating overfitting and improving the generalization capability.
Sijing Xie, Dingzhu Wen, Changsheng You, Qimei Chen, Mehdi Bennis, Kaibin Huang
IEEE J. Sel. Areas Commun.2
2026 Joint Sensing, Communication, and Computation for Vertical Federated Edge Learning in Edge Perception Networks
abstract
Combining wireless sensing and edge intelligence, edge perception networks enable intelligent data collection and processing at the network edge. However, traditional sample partition based horizontal federated edge learning (HFEEL) struggles to effectively fuse complementary multi-view information from distributed devices. To address this limitation, we propose a vertical federated edge learning (VFEEL) framework tailored for feature-partitioned sensing data. In this paper, we consider an integrated sensing, communication, and computation (ISCC)-enabled edge perception network, where multiple edge devices utilize wireless signals to sense environmental information for updating their local models, and the edge server aggregates feature embeddings via over-the-air computation (AirComp) for global model training. First, we analyze the convergence behavior of the ISCC-enabled VFEEL in terms of the loss function degradation in the presence of wireless sensing noise and aggregation distortions during AirComp. Then, to accelerate convergence, we aim to optimize the batch size, sensing power, and transmission power control at edge devices as well as the denoising factors at the edge server under limited network constraints on overall energy consumption and per-round latency. Due to the tight coupling of variables, the problem is non-convex. To address this problem, we design an alternating optimization-based algorithm to efficiently obtain a high-quality solution. Numerical results are conducted based on a human motion recognition task to verify that the proposed ISCC-enabled VFEEL algorithm achieves higher accuracy compared with other benchmarking schemes including ISCC-enabled HFEEL approach.
Xiaowen Cao 0001, Dingzhu Wen, Suzhi Bi, Yuanhao Cui, Guangxu Zhu, Han Hu 0003, Yonina C. Eldar
IEEE Trans. Mob. Comput.2
2026 Task-Oriented Multimodal Edge Intelligence via Integrated Sensing-Communication-Computation
Yinghui He, Zhong Ye, Dingzhu Wen, Guanding Yu
IEEE Trans. Wirel. Commun.4
2026 Decentralized Integration of Sensing-Communication-Computation for Multi-Task Edge AI Inference
abstract
Collaborative artificial intelligence (AI) inference has effectively deployed well-trained AI models at the network edge to empower immersive intelligent services such as autonomous driving and smart cities. This paper proposes an integrated sensing-computation-communication (ISCC) scheme for decentralized multi-task collaborative inference systems. The proposed scheme connects multiple devices via device-to-device (D2D) links. Each device first extracts a homogeneous feature vector from the raw sensory data obtained from the same wide view of the source target and then aggregates all local feature vectors using the over-the-air computation (AirComp) technique to complete a specific inference task. To enhance spectrum efficiency, the full-duplex communication technique is adopted, which allows all devices to transmit and receive in the same frequency band. To suppress the self-interference caused by full duplex communications and simultaneously enhance all tasks’ performance, a multi-objective optimization problem is formulated, where discriminant gain is adopted as the inference performance metric. The challenges to solve this problem arise from three aspects: The impact of the self-interference (SI) channel incurred by full-duplex communication, the precoding design of each device, and the coupling among subcarrier allocation, sensing, computation, and communication processes. To tackle this problem, aquadratic transformandweighted bipartite matchingbased alternating maximization approach is proposed. Numerical results based on jointly completing three tasks of human motion classification, human gender recognition, and human age group classification, verify the effectiveness of the proposed method by showing that the proposed method outperforms the state-of-the-art successive convex approximation (SCA) based algorithm.
Chenye Wang, Zeming Zhuang, Dingzhu Wen, Yuanming Shi, Xin Wang 0003
IEEE Trans. Wirel. Commun.3
2026 Integrated Sensing, Communication, and Computation for Over-the-Air Federated Edge Learning
abstract
This paper studies an over-the-air federated edge learning (Air-FEEL) system with integrated sensing, communication, and computation (ISCC), in which one edge server coordinates multiple edge devices to wirelessly sense the objects and use the sensing data to collaboratively train a machine learning model for recognition tasks. In this system, over-the-air computation (AirComp) is employed to enable one-shot model aggregation from edge devices. Under this setup, we analyze the convergence behavior of the ISCC-enabled Air-FEEL in terms of the loss function degradation, by particularly taking into account the wireless sensing noise during the training data acquisition and the AirComp distortions during the over-the-air model aggregation. The result theoretically shows that sensing, communication, and computation compete for network resources to jointly decide the convergence rate. Based on the analysis, we design the ISCC parameters under the target of maximizing the loss function degradation while ensuring the latency and energy budgets in each round. The challenge lies on the tightly coupled processes of sensing, communication, and computation among different devices. To tackle the challenge, we derive a low-complexity ISCC algorithm by alternately optimizing the batch size control and the network resource allocation. It is found that for each device, less sensing power should be consumed if a larger batch of data samples is obtained and vice versa. Besides, with a given batch size, the optimal computation speed of one device is the minimum one that satisfies the latency constraint. Numerical results based on a human motion recognition task verify the theoretical convergence analysis and show that the proposed ISCC algorithm well coordinates the batch size control and resource allocation among sensing, communication, and computation to enhance the learning performance.
Dingzhu Wen, Sijing Xie, Xiaowen Cao 0001, Yuanhao Cui, Jie Xu 0002, Yuanming Shi, Shuguang Cui
IEEE Trans. Wirel. Commun.1
2026 UAV-Assisted Edge Inference With Integrated Sensing, Communication, and Computation
Dingzhu Wen, Guangxu Zhu, Yuan Liu 0001, Yuanming Shi, Honglin Hu
IEEE Trans. Wirel. Commun.1
2026 Federated Dropout: Convergence Analysis and Resource Allocation
abstract
Federated Dropout is an efficient technique to overcome both communication and computation bottlenecks for deploying federated learning at the network edge. In each training round, an edge device only needs to update and transmit a sub-model, which is generated by the typical method of dropout in deep learning, and thus effectively reduces the per-round latency. \textcolor{blue}{However, the theoretical convergence analysis for Federated Dropout is still lacking in the literature, particularly regarding the quantitative influence of dropout rate on convergence}. To address this issue, by using the Taylor expansion method, we mathematically show that the gradient variance increases with a scaling factor of $γ/(1-γ)$, with $γ\in [0, θ)$ denoting the dropout rate and $θ$ being the maximum dropout rate ensuring the loss function reduction. Based on the above approximation, we provide the convergence analysis for Federated Dropout. Specifically, it is shown that a larger dropout rate of each device leads to a slower convergence rate. This provides a theoretical foundation for reducing the convergence latency by making a tradeoff between the per-round latency and the overall rounds till convergence. Moreover, a low-complexity algorithm is proposed to jointly optimize the dropout rate and the bandwidth allocation for minimizing the loss function in all rounds under a given per-round latency and limited network resources. Finally, numerical results are provided to verify the effectiveness of the proposed algorithm.
Sijing Xie, Dingzhu Wen, Changsheng You, Tharmalingam Ratnarajah, Kaibin Huang
IEEE Trans. Wirel. Commun.2
2026 MIMO Over-the-Air Computation for Device-Edge Collaborative Inference
abstract
Device-edge collaborative inference, which deploys well-trained artificial intelligence (AI) models at the network edge via the cooperation of edge devices and edge servers, emerges as a promising technique to provide ubiquitous intelligent services. In this paper, a multiple-input multiple-output (MIMO) over-the-air computation (AirComp) scheme is proposed for the efficient implementation of device-edge collaborative inference. In the considered system, the technique of MIMO AirComp is utilized to aggregate local feature vectors, extracted from noise-corrupted sensory data on devices, at the server to efficiently derive a denoised global one for completing the downstream inference task. Device-edge collaborative inference features a task-oriented property, that concerns the effectiveness and efficiency of the task execution. In this case, the traditional AirComp criterion, i.e., minimum mean square error (MMSE), is not effective, since the same distortion level on different feature elements may have different influences on the inference performance. To this end, this paper directly adopts inference accuracy as the design objective. As the instantaneous inference accuracy is unknown during the design stage, an approximated but tractable metric, called discriminant gain, which measures the discernibility of different classes, is adopted. To maximize the inference accuracy measured by discriminant gain, a MIMO AirComp technique is proposed to jointly optimize all feature elements. The problem is nonconvex because of the complicated form of the objective function and the constraints. The solution based on semidefinite relaxation (SDR) and successive convex approximation (SCA) is employed to design a joint transmit precoding and receive beamforming scheme. Besides, to enhance the robustness of practical AI models in the inference stage, a post-processing design of feature magnitude normalization is proposed. Extensive experiments are conducted based on a practical human motion recognition task, which verifies our theoretical analysis and the superiority of our proposed scheme.
Dingzhu Wen, Li You 0001, Jingjing Wang 0001, Sheng Wu 0001, Yuanming Shi
IEEE Trans. Wirel. Commun.2
2025 Structured IB: Improving Information Bottleneck with Structured Feature Learning
abstract
The Information Bottleneck (IB) principle has emerged as a promising approach for enhancing the generalization, robustness, and interpretability of deep neural networks, demonstrating efficacy across image segmentation, document clustering, and semantic communication. Among IB implementations, the IB Lagrangian method, employing Lagrangian multipliers, is widely adopted. While numerous methods for the optimizations of IB Lagrangian based on variational bounds and neural estimators are feasible, their performance is highly dependent on the quality of their design, which is inherently prone to errors. To address this limitation, we introduce Structured IB, a framework for investigating potential structured features. By incorporating auxiliary encoders to extract missing informative features, we generate more informative representations. Our experiments demonstrate superior prediction accuracy and task-relevant information preservation compared to the original IB Lagrangian method, even with reduced network size.
Youlong Wu, Dingzhu Wen, Yong Zhou 0006, Yuanming Shi
AAAI3
2025 Integrated Sensing-Communication-Computation for Movable Antennas Assisted Multi-Device Edge AI Inference
Dingzhu Wen, Min Fu 0003, Yong Zhou 0006, Yuanming Shi
GLOBECOM4
2025 Hierarchical Federated Learning with Integrated Sensing-Communication-Computation Over Space-Air-Ground Integrated Networks
abstract
Federated learning has achieved significant advancements in edge artificial intelligence (AI) by addressing issues related to data privacy and communication overload. Moreover, hierarchical federated learning over space-air-ground integrated networks (FedSAG), which consists of low-Earth orbit (LEO) satellites, unmanned aerial vehicles (UAVs), and edge devices, aims to provide AI services in sparsely populated regions lacking ground communication infrastructure. However, previous studies have overlooked the essential sensing process required for acquiring training data, potentially compromising training efficiency and model accuracy. In this paper, we propose an integrated sensing-communication-computation (ISCC) enabled FedSAG system, which allows remote edge devices to collect data via wireless sensing and collaboratively train a global model without sharing local data. We then analyse the convergence of the ISCC-enabled FedSAG and formulate two optimization problems. The first aims to minimize sensing variance under energy and time constraints, while the second seeks to reduce transmission energy through optimal route selection between UAVs and LEO satellite. We reformulate the problems to a minimum spanning tree and propose a Chu-Liu-Edmonds algorithm based a two-stage optimization. Simulation results demonstrate that our proposed algorithm significantly enhances convergence performance and reduces energy consumption.
Zhanpeng Yang, Jingyang Zhu, Dingzhu Wen, Yuanming Shi, Wei Chen 0002
ICC4
2025 Robust Multimodal Information Bottleneck for Satellite-to-Ground Task-Oriented Communication
abstract
In this paper, we study satellite-to-ground taskoriented communication for edge inference tasks, where a satellite extracts, fuses and encodes multimodal feature vectors and then sends them to a ground server under inevitable channel noise conditions for downstream processing. However, the multispectral and multi-resolution characteristics of multimodal satellite remote sensing data render traditional multimodal methods inapplicable. To reduce the data redundancy caused by the high-dimensional and complex multimodal vectors generated onboard while retaining key information and enhancing robustness against channel noise. We propose a Robust Multimodal Information Bottleneck (RMIB) framework which considers channel noise and communication bandwidth and introduces a new information bottleneck optimization objective. By applying this objective through end-to-end training, we optimize the feature extraction, fusion and encodes multimodal data into robust and effective feature vector in noisy communication environments by reducing redundancy and enhancing feature discrimination. To tackle the RMIB objective function, we derive a tractable variational upper bound using the Variational Information Bottleneck technique to overcome the computational intractability of mutual information. Experimental results demonstrate that our method not only outperforms baseline techniques in classification accuracy on three datasets but also enhances robustness against channel noise and reduces communication overhead.
Dingzhu Wen, Youlong Wu, Yuanming Shi, Ting Wang 0001
ISCC2
2025 Cost-Efficient Wideband Beam Training: A 3D Controllable Beam Squint Approach
abstract
The widely adopted extremely large-scale multiple-input multiple-output (XL-MIMO) wideband systems are fundamentally constrained by their two-dimensional planar coverage, primarily due to the conventional use of uniform linear arrays (ULAs). This limitation significantly degrades the quality of service in practical three-dimensional (3D) deployment scenarios. Furthermore, XL-MIMO wideband systems inherently suffer from severe beam squint effects caused by non-negligible signal propagation delays, and the narrow beamwidth results in substantial beam training overhead. To address these challenges, we propose a novel cost-efficient frequency-dependent beamforming framework capable of supporting 3D spatially distributed users by leveraging a full-dimensional uniform planar array (UPA) architecture. Based on this design, we introduce a controllable 3D beam squint scheme via a partially-connected time-delay network, alongside a low-dimensional user-centric spatial-frequency scanning codebook to enable rapid and robust beam training. Numerical results demonstrate that the proposed scheme achieves accurate 3D angle estimation with significantly reduced training overhead during the initial access phase. Additionally, the proposed method can maintain high angle estimation accuracy in poor channel conditions.
Ruihuan Wang, Qimei Chen, Qipeng Zheng, Dingzhu Wen, Muhammad Kaleem Awan
PIMRC5
2025 Federated LoRA with Dropout: An Efficient and Overfitting Control Approach for LLM Fine-Tuning
abstract
This paper introduces the Federated LoRA with Dropout (FedLoDrop) framework, designed to enhance generalization performance for downstream tasks at the network edge while simultaneously reducing overhead. Within this framework, we derive a generalization error bound under sparsity regularization, elucidating the theoretical principles that balance underfitting and overfitting. Our analysis shows that a higher dropout rate increases sparsity, lowering the Pointwise Hypothesis Stability (PHS) upper bound and narrowing the gap between empirical and generalization errors. However, this also leads to a higher empirical error, which, together with the gap, contributes to the total generalization error. Consequently, we formulate an optimization problem that jointly considers dropout rate and resource allocation, aiming to minimize the upper bound of the generalization error. Finally, numerical results demonstrate the effectiveness of the proposed approach in mitigating overfitting and enhancing generalization capabilities.
Sijing Xie, Changsheng You, Qimei Chen, Dingzhu Wen
PIMRC4
2025 Integrated Sensing-Communication-Computation Based Online Federated Learning with Limited Cache
abstract
This paper investigates a cache-assisted scheme for training an over-the-air federated edge learning (Air-FEEL) system with integrated sensing, communication, and computing (ISCC) utilizing both cached and real-time sensory data. In each training round, a server coordinates multiple devices to perform wireless sensing and utilize the real-time and cached old sensory data samples for updating the AI model based on the gradient descent method. Subsequently, the technique of over-the-air computation (AirComp) is utilized to aggregate local gradient vectors from all devices at the server for updating the global model. Particularly, this work makes the first attempt to exploit the spare on-device cache to reuse the old sensory data obtained in the previous round for enhancing learning performance. The theoretical analysis of this Air-FEEL framework is analyzed, where the mathematical relation between the convergence rate and the number of cached and real-time sensory data samples, sensing signal-to-noise ratio (SNR), and AirComp SNR is unveiled. Based on this theoretical finding and targeting enhancing the learning performance, a cache-assisted ISCC scheme is proposed to tackle a non-convex and complicated convergence accelerating problem via joint sensing and communication power control, cache and sensory data size allocation, and computation frequency management. Experimental results based on human motion recognition tasks verify the theoretical convergence analysis and show that the proposed cache-assisted ISCC scheme outperforms existing ISCC-based FL schemes without utilizing the spare on-device cache.
Qiaosheng Hu, Dingzhu Wen
WCNC3
2025 Joint Source and Channel Coding for Multi-Modal Satellite-to-Ground Semantic Communications
abstract
This paper presents a novel Joint Source and Channel Coding (JSCC) method for semantic communication to enhance the communication efficiency for transmitting high-resolution multi-modal data from Low Earth Orbit (LEO) satellites to ground stations. On the satellite, a JSCC encoder consisting of neural networks (NNs) is utilized to map the input multi-modal data into a common signal, while an NN-based JSCC decoder at the ground station reconstructs the original input from the received signal. Throughout this process, the common semantic information among each modality is learned to optimize coding space, and a robust coding scheme specific to the satellite downlink channel is developed between the encoder and decoder. Experiments have demonstrated that the proposed method outperforms existing signal modality JSCC approaches, achieving multi-modal transmission with significantly smaller coding space and thereby reducing communication overhead.
Yanbo Yin, Dingzhu Wen, Youlong Wu, Yuanming Shi
WCNC3
2025 Sensing-Communication-Computation Integration for Federated Edge Learning With Controllable Model Dropout
abstract
Federated edge learning (FEEL) is an advanced paradigm in edge artificial intelligence, enabling privacy-preserving collaborative model training through periodic communication between edge devices and a central server. FEEL involves three key processes: 1) sensing; 2) computation; and 3) communication for data acquisition, processing, and exchange, respectively. Due to limited system resources, optimizing each process individually may lead to suboptimal learning performance. This challenge has sparked research into integrated sensing-computation–communication (ISCC) design for enhanced FEEL. While previous work has optimized general learning parameters, such as batch size and computing frequency, there is a lack of customized designs considering the neural network architecture as an optimizable variable in ISCC for FEEL. To close this gap, we introduce a novel design where each device generates a submodel through controllable weight dropout, adding flexibility by directly manipulating the learning process and reducing computation and communication overhead. To guide ISCC resource allocation in this new setting, we present a comprehensive convergence analysis, revealing the tight coupling of sensing, computation, and communication across devices and their impact on FEEL convergence. Building on these theoretical insights, we formulate an ISCC problem aiming to maximize the FEEL convergence rate through joint optimization of variables, such as batch size, sensing power, dropout rate, and communication power. This nonconvex problem is decomposed into two subproblems via alternating optimization: one controls batch size using a sorting algorithm, while the other focuses on ISCC device parameters, transformable into a convex problem solved by successive convex approximation. Extensive experiments using human motion recognition datasets demonstrate the superiority of the proposed design over baseline schemes.
Xiang Jiao, Guangxu Zhu, Wei Jiang 0003, Li Chen 0015, Wu Luo, Dingzhu Wen
IEEE Internet Things J.6
2025 Efficient Collaborative Learning Over Unreliable D2D Network: Adaptive Cluster Head Selection and Resource Allocation
abstract
Recently, decentralized learning has been proposed for model training among mobile devices without center nodes. However, large resource overhead for model aggregation and synchronization would be incurred, which may reduce the learning performance under a given resource budget. To cope with these issues, we propose a novel cluster-based collaborative learning framework over device-to-device (D2D) network, where one device is selected as the cluster head for model aggregation. Within the proposed framework, the learning performance (evaluated by model divergence) and learning latency are analyzed with the consideration of imbalanced data and unreliable D2D communication. Then, an optimization problem is formulated to maximize the learning performance under a given latency constraint by joint cluster head selection and resource allocation. To solve this problem, a lower bound on latency constraint is first obtained for error-free model aggregation. The optimal learning performance is also derived with different degrees of data distribution. After that, an adaptive cluster head selection and resource allocation algorithm is developed for erroneous case by introducing the outage probability. Finally, comprehensive experiments are conducted on well-known models and datasets to illustrate the effectiveness of the proposed algorithm. The results show that our proposal can improve the learning performance while reducing communication and signaling overheads.
Shengli Liu 0002, Chonghe Liu, Dingzhu Wen, Guanding Yu
IEEE Trans. Commun.3
2025 Balancing Straggler Mitigation and Information Protection for Matrix Multiplication in Heterogeneous Multi-Group Networks
abstract
Distributed computing has made it possible to satisfy the demands for large-scale matrix multiplication. A distributed computing system suffers from both straggler problem and information leakage. In a heterogeneous network consisting of multiple worker groups, information leakage can be caused by both intra-group and inter-group collusion. Besides, considering the heterogeneity of worker nodes, stronger nodes are supposed to compute more tasks to provide robustness for stragglers. However, this results in more information being leaked to stronger nodes, contradicting the principle of information protection. In this paper, we propose a multi-group heterogeneous secure coded matrix multiplication (MG-HSCMM) scheme to solve these problems in a heterogeneous multi-group network. By taking the heterogeneity of worker nodes into consideration, the corresponding recovery threshold and security constraint are obtained. To improve the performance of such a network, a low complexity task allocation policy that balances straggler mitigation and information protection is given. Compared with existing schemes, MG-HSCMM can achieve significant performance gain. Numerical simulation results verify the superiority of our proposed scheme.
Li Chen 0015, Dingzhu Wen
IEEE Trans. Commun.3
2025 Clustered Federated Multi-Task Learning: A Communication-and-Computation Efficient Sparse Sharing Approach
abstract
Federated multi-task learning (FMTL) is a promising technology to tackle one of the most severe non-independent and identically distributed (non-IID) data challenge in federated learning (FL), which treats each client as a single task and learns personalized models by exploiting task correlations. However, the transmission of individual task models generally results in a significant amount of communication overhead compared with global model broadcasting. Furthermore, related works mainly focus on FMTLs with default and static relationships among tasks, which obliterates the non-IID data characteristic. To address these issues, we propose a novel Clustered FMTL mechanism via Sparse Sharing (FedSS). Specifically, we introduce an iterative model pruning approach that trains customized client models to deal with the non-IID issue. Thereafter, we divide clients into different tasks according to their model similarities to promote communication efficiency. Based on clustered tasks, we introduce a sparse sharing mechanism that allows clients to share model parameters dynamically among different tasks to further boost the training performance. On the other aspect, the infertile communication resources would degrade the FMTL performance by restricting the personalized model transmissions. Hence, we first theoretically analyze the convergence performance of the proposed FedSS, which quantitatively unveils the relationship between the local model training performance and communication resources. Thereafter, we formulate a communication-and-computation efficient optimization problem via a joint sparsity ratio assignment and bandwidth allocation strategy. Closed-form expressions for the optimal sparsity ratio and bandwidth allocation are derived based on Lyapunov optimization and block coordinate update (BCU) algorithms. Numerical results illustrate that the proposed FedSS outperforms the benchmarks, and achieves an efficient communication and computation performance.
Yuhan Ai, Qimei Chen, Guangxu Zhu, Dingzhu Wen, Hao Jiang 0010
IEEE Trans. Wirel. Commun.4
2025 Dynamic UAV-Assisted Cooperative Edge AI Inference
abstract
Deploying intelligent service and executing inference tasks in the proximity of the edge enable models to access enormous real-time data generated by the edge devices. However, the dilemma of fulfilling service demands with limited resources at edge devices impairs the efficacy of conventional data-oriented communication systems. To achieve a better trade-off between inference accuracy and communication overhead, in this paper, we propose a dynamic unmanned aerial vehicle (UAV)-assisted cooperative edge inference system, where a UAV acts as an edge server to aggregate the wide-view features from mobile sensors through Over-the-Air computation (AirComp) to complete the inference task cooperatively. Discriminant gain, an effective indicator for the inference accuracy, is adopted to realize task-oriented design. To exploit channel diversity and data diversity in the multi-device cooperative edge inference system, we maximize the discriminant gain of the AirComp feature aggregation by jointly optimizing the UAV trajectory and the power allocation policy with respect to the different important levels of feature dimensions. An alternating algorithm and a successive convex approximation (SCA)-based method are then proposed to solve the optimization problem. Numerical simulations further validate the efficacy of the proposed design compared to the baselines.
Jingfeng Huang, Lixiang Lian, Dingzhu Wen, Yong Zhou 0006, Fuzhai Wang, Weichang Wang, Yuanming Shi
IEEE Trans. Wirel. Commun.3
2024 Convergence Analysis for Federated Dropout
abstract
Federated dropout on the weight is an efficient technique to overcome both communication and computation bottlenecks for deploying federated learning at the network edge. However, the theoretical analysis for Federated Dropout is still lacking in the literature, due to the challenge arising from the gradient bias. To address this issue, by using the Taylor expansion method, we mathematically show that the gradient vector with dropout can be approximated as an unbiased estimation of that without dropout; while its gradient variance increases with a scaling factor of γ/(1 − γ), with γ ∈ [0,θ) denoting the dropout rate and θ being the maximum dropout rate ensuring the loss function reduction. Based on the above approximation, we provide the loss function analysis for Federated Dropout. Specifically, it is shown that a larger dropout rate of each device leads to a slower convergence rate. Finally, numerical results are provided to verify the effects of dropout rate on convergence in both underfitting and overfitting scenarios.
Sijing Xie, Dingzhu Wen, Changsheng You, Tharmalingam Ratnarajah, Kaibin Huang
GLOBECOM2
2024 Dynamic Communication in Multi-Agent Reinforcement Learning via Information Bottleneck
abstract
Effective information sharing is essential for multi-agent systems to execute cooperative tasks successfully. Typically, agents within such systems are either stationary or possess unrestricted communication ranges. However, in more complex scenarios where agent mobility is introduced, the communication network’s topology becomes dynamic over time. This dynamism can result in partial communication unreachability among certain agents. Consequently, striking a balance between minimizing overall communication overhead and optimizing task performance becomes a formidable challenge. In this paper, we address the issue of dynamic communication in multi-agent systems. We propose a novel approach that leverages the principle of information bottleneck theory to develop a multi-mean field multi-agent reinforcement learning algorithm called MMIB. Through a series of experiments, we demonstrate the effectiveness of our proposed algorithm in reducing communication overhead while maintaining task performance at a level comparable to other state-of-the-art multi-agent reinforcement learning algorithms.
Jiawei You, Youlong Wu, Dingzhu Wen, Yong Zhou 0006, Yuning Jiang 0002, Yuanming Shi
GLOBECOM3
2024 Satellite Federated Fine-Tuning for Foundation Models: Architecture Design and System Optimization
abstract
With the surge in the number of low earth orbit (LEO) satellites, continuous research has emerged on using satellite data to train artificial intelligence models. On one hand, traditional centralized training on the ground is not feasible due to privacy concerns and limited bandwidth for downloading raw satellite data. On the other hand, due to the limited energy and computational capability of satellites, training directly on satellites suffers from prolonged latency, especially for large models. To alleviate these issues, we propose a novel satellite-ground collaborative federated fine-tuning architecture, where ground stations (GSs) and satellites collaboratively train a global model without the need for data downloads. In this proposed architecture, satellites serve as edge devices and the ground server serves as a coordinator. However, the short satellite-ground communication windows caused by the high mobility of satellites and the substantial intra-orbit data transmission bring special challenges to the transmission process of federated edge learning. To tackle these challenges, we carefully design the satellite-ground collaborative fine-tuning architecture and utilize an optimized ring all-reduce algorithm and network flow algorithm to enhance the intra-orbit and ground-satellite transmissions, respectively. Experimental results demonstrate that our proposed architecture significantly reduces the training time by 40% compared to training solely on satellite.
Peng Yang 0027, Jingyang Zhu, Dingzhu Wen, Ting Wang 0001, Yong Zhou 0006, Yuanming Shi, Chunxiao Jiang
GLOBECOM4
2024 End-to-End Hybrid Beamforming for mmWave Integrated Access and Backhaul with Active Sensing Strategy
abstract
The effectiveness of Millimeter Wave full-duplex (FD) Integrated Access and Backhaul (IAB) system relies on high-dimensional channel estimation with high computational complexity. To avoid high-overhead pilot training, we propose a novel low-complexity end-to-end (E2E) hybrid beamforming strategy for FD mmwave IAB systems using implicit channel state information (CSI). Particularly, the IAB node first dynamically senses spatial channels, where an sensing Transformer block is introduced to actively design the sensing vector. The active sensing strategy can effectively handle the sequential pilot observations with an arbitrary input length. Capitalizing on the implicit channel features extracted by the Transformer, a hybrid beamforming neural network (HBFnet) is further exploited to design the hybrid precoder/combiner of IAB node, thus efficiently mitigating SI while compensating channel fading. Simulation results demonstrate that the proposed scheme outperforms the benchmarks, especially with low pilot overheads.
Sisi Lin, Xiaoxia Xu 0002, Qimei Chen, Dingzhu Wen, Guocao Tao, Hao Jiang 0010
WCNC5
2024 RIS-Assisted Multi-Device Edge AI Inference
abstract
In this paper, we propose a multi-device co-inference system based on a task-oriented over-the-air computation (Air-Comp) via reconfigurable intelligent surface (RIS). Specially, local feature vectors extracted from the real-time noisy sensory data on devices are aggregated over-the-air by exploiting the waveform superposition in a multi-user channel. Then the aggregated features received at the server are fed into an inference model for decision making or control of actuators. Based on the proposed multi-device co-inference system, we jointly optimize the receive signal strength of the device, the beamforming vector, and RIS phase shifts to suppress the sensing and channel noise and maximize the inference accuracy. To solve the problem, we first transform the original problem into a convex difference (d.c.) problem, and convert the d.c. problem from the complex domain to the real domain. Then, we propose a successive convex approximation based approach to solve the problem in the real domain. With the supportive data and results from the application of human motion recognition, we show the proposed scheme achieves a higher inference accuracy then the conventional approaches.
Yijie Mao, Dingzhu Wen, Yong Zhou 0006, Yuanming Shi
WCNC3
2024 Latency-Aware Microservice Deployment for Edge AI Enabled Video Analytics
abstract
Video analytics plays a pivotal role in public safety (e.g., criminal suspect detection, traffic flow count, and illegal parking management), which assists the polices in monitoring all anomalous events in the street. In this paper, we consider the scenario with multiple video analytics applications from a single video stream. However, traditional monolithic architecture based video analytics applications shall seriously increase the response latency due to the resource contention of repetitive components. Therefore, we utilize the microservice architecture based video analytics (MAVA) to share the universal microser-vices in different applications, which shall decrease the response latency by reducing the computation load and increasing the resource utilization. To further achieve fast and accurate video analytics, the video analytics microservices are deployed in the edge closing to the cameras and users, and artificial intelligence (AI) methods are used in the microservices to realize specified functions. Therefore, an edge AI enabled MAVA (EAI-MAVA) architecture is proposed to achieve accurate video analytics in real-time. Furthermore, we formulate a microservice deployment problem to determine the location of each microservice in EAI-MAVA, which minimizes the response latency of all applications by considering the resource demands of microservices and the resource constraints of heterogeneous edge devices. Finally, a greedy-based heuristic algorithm is proposed to solve the non-convex microservice deployment problem, which obtains a sub-optimal solution with small loss of accuracy and reduces the solution time obviously.
Zhanpeng Yang, Xin Liu 0049, Dingzhu Wen, Yong Zhou 0006, Yuanming Shi
WCNC4
2024 Semantic Communication Meets Edge Intelligence: Semantic-Relay-Aided Text Transmissions
abstract
Semantic communication (SemCom) has emerged as a promising technology to improve the spectrum efficiency of next-generation wireless networks, by extracting meaningful content from the data and transmitting relevant semantic information only. However, the existing research usually overlooks the limited computing and storage resources on the mobile devices, which may make it unaffordable to implement resource-demanding deep learning (DL)-based semantic encoders/decoders. Moreover, besides the end-to-end SemCom framework, cooperative SemCom has not been well studied in the existing works, which can further enhance the communication performance. To address these issues, we propose a new architecture in this article, called semantic relay (SemRelay), which acts as an edge server to provide DL-enabled SemCom (DeepSC) services for two categories of edge users, called semantic users (SemUsers) with rich computing resources and conventional users (ConUsers) with limited resources. Two new transmission protocols are proposed for enabling text transmissions from the base station to the SemUsers and ConUsers, respectively, via the SemRelay (edge server). Moreover, an optimization problem is formulated to jointly design the SemRelay transmit power allocation and system bandwidth allocation to maximize the weighted sum-rate of all the users. Although this problem is nonconvex and hence difficult to solve, we propose an efficient algorithm to obtain a high-quality suboptimal solution by applying the block coordinate descent and successive convex approximation techniques. Finally, the numerical results demonstrate the effectiveness of our proposed algorithm and the superior performance of the proposed SemRelay as compared to the traditional decode-and-forward relays, especially in the small bandwidth regime.
Zeyang Hu, Changsheng You, Dingzhu Wen, Yuanhao Cui, Yi Gong 0001, Kaibin Huang
IEEE Internet Things J.4
2024 Joint Device Scheduling and Resource Allocation for ISCC-Based Multiview-Multitask Inference
abstract
This article investigates an integrated sensing-communication-computation (ISCC)-based multiview-multitask (MVMT) edge artificial intelligence inference system. Each device senses a narrow view of a target area and processes the echo signal to generate real-time sensory data. An edge server receives and combines multiple views of data from multiple devices to complete several downstream inference tasks. Compared with existing designs where dedicated sensory data are obtained, transmitted, and processed for each task, this ISCC-based MVMT framework enjoys reduced costs of sensing, on-device computation, and communication overhead due to data sharing among different tasks. The challenges of improving all tasks’ inference accuracy lie in the tight coupling of sensing, communication, and computation among different devices and sensory view competition among different tasks. These two challenges intertwine, making the multitask optimization problem mixed-integer nonconvex programming. To tackle this problem, we propose a joint device scheduling and resource allocation (JDSRA) scheme, which alternatively solves a subproblem of joint device scheduling and time allocation and a subproblem of resource allocation till convergence. Particularly, in addition to a dynamic-programming-based optimal device scheduling algorithm, a low-complexity suboptimal algorithm is proposed based on sorting a derived closed-form indicator, which represents the increase of all tasks’ inference accuracy per time unit consumption. Besides, a low-complexity optimal resource allocation algorithm is proposed by parallelly solving multiple simple convex subproblems. Numerical results based on jointly completing three tasks of human motion recognition, human height recognition, and localization in smart home scenarios are conducted to verify the performance of our proposed schemes.
Diao Wang, Dingzhu Wen, Yinghui He, Qimei Chen, Guangxu Zhu, Guanding Yu
IEEE Internet Things J.2
2024 Collaborative Edge AI Inference Over Cloud-RAN
abstract
In this paper, a cloud radio access network (Cloud-RAN) based collaborative edge AI inference architecture is proposed. Specifically, geographically distributed devices capture real-time noise-corrupted sensory data samples and extract the noisy local feature vectors, which are then aggregated at each remote radio head (RRH) to suppress sensing noise. To realize efficient uplink feature aggregation, we allow each RRH receives local feature vectors from all devices over the same resource blocks simultaneously by leveraging an over-the-air computation (AirComp) technique. Thereafter, these aggregated feature vectors are quantized and transmitted to a central processor (CP) for further aggregation and downstream inference tasks. Our aim in this work is to maximize the inference accuracy via a surrogate accuracy metric called discriminant gain, which measures the discernibility of different classes in the feature space. The key challenges lie on simultaneously suppressing the coupled sensing noise, AirComp distortion caused by hostile wireless channels, and the quantization error resulting from the limited capacity of fronthaul links. To address these challenges, this work proposes a joint transmit precoding, receive beamforming, and quantization error control scheme to enhance the inference accuracy. Extensive numerical experiments demonstrate the effectiveness and superiority of our proposed optimization algorithm compared to various baselines.
Dingzhu Wen, Guangxu Zhu, Qimei Chen, Kaifeng Han, Yuanming Shi
IEEE Trans. Commun.2
2024 Energy-Efficient Optimal Mode Selection for Edge AI Inference via Integrated Sensing-Communication-Computation
abstract
Existing edge inference methods only consider one paradigm, i.e., one of on-device inference, on-server inference, or edge-device cooperative inference. Each paradigm has its pros and cons as well as dominant application scopes. For example, the on-device paradigm is the best choice when the inference task is not computationally intensive, the on-server paradigm is suitable if the communication capacity is strong, and the edge-device cooperative mode should be selected in the scenario of weak on-device communication and computation. However, each paradigm suffers from poor performance if deployed outside of its application scope, thus leading to limited potential and flexibility. This paper proposes an edge AI inference framework, which makes the first attempt to jointly consider the three modes for making full use of their benefits. In addition, sensing for data acquisition is enabled at both the edge server and the device. This can effectively improve the inference accuracy with rich information on the target area from two different views. On the other hand, energy cost minimization turns out to be a key target all over the world and a significant issue in wireless networks. To this end, we target minimizing the system energy cost under a given inference accuracy guarantee and other network resource constraints, by coordinating sensing, communication, and computation in different modes. By optimally solving the optimization problem, an integrated sensing-communication-computation (ISCC) based task-oriented mode selection scheme is proposed. A practical ISCC platform is built and extensive experiments are conducted to verify our theoretical analysis.
Dingzhu Wen, Qimei Chen, Guangxu Zhu, Yuanming Shi
IEEE Trans. Mob. Comput.2
2024 Task-Oriented Over-the-Air Computation for Multi-Device Edge AI
abstract
Edge inference refers to the use of artificial intelligent (AI) models at the network edge to provide mobile devices inference services and thereby enable intelligent services such as auto-driving and Metaverse towards 6G. However, departing from the classic paradigm of data-centric designs, the 6G networks for supporting edge AI features task-oriented techniques that focus on effective and efficient execution of AI task. Targeting end-to-end system performance, such techniques are sophisticated as they aim to seamlessly integrate sensing (data acquisition), communication (data transmission), and computation (data processing). Aligned with the paradigm shift, a task-oriented over-the-air computation (AirComp) scheme is proposed in this paper for multi-device split-inference system. In the considered system, local feature vectors, which are extracted from the real-time noisy sensory data on devices, are aggregated over-the-air by exploiting the waveform superposition in a multiuser channel. Then the aggregated features as received at a server are fed into an inference model with the result used for decision making or control of actuators. To design inference-oriented AirComp, the transmit precoders at edge devices and receive beamforming at edge server are jointly optimized to rein in the aggregation error and maximize the inference accuracy. The problem is made tractable by measuring the inference accuracy using a surrogate metric called discriminant gain, which measures the discernibility of two object classes in the application of object/event classification. It is discovered that the conventional AirComp beamforming design for minimizing the mean square error in generic AirComp with respect to the noiseless case may not lead to the optimal classification accuracy. The reason is due to the overlooking of the fact that feature dimensions have different sensitivity towards aggregation errors and are thus of different importance levels for classification. This issue is addressed in this work via a new task-oriented AirComp scheme designed by directly maximizing the derived discriminant gain. However, the resultant problem of joint transmit precoding and receive beamforming is nonconvex and difficult to solve due to the complicated form of discriminant gain and the coupling between the control variables. We overcome the difficulty using the successive convex approximation. The performance gain of the proposed task-oriented scheme over the conventional schemes is verified by extensive experiments targeting the application of human motion recognition.
Dingzhu Wen, Xiang Jiao, Peixi Liu, Guangxu Zhu, Yuanming Shi, Kaibin Huang
IEEE Trans. Wirel. Commun.1
2024 Task-Oriented Sensing, Computation, and Communication Integration for Multi-Device Edge AI
abstract
This paper studies a new multi-device edge artificial-intelligent (AI) system, which jointly exploits the AI model split inference and integrated sensing and communication (ISAC) to enable low-latency intelligent services at the network edge. In this system, multiple ISAC devices perform radar sensing to obtain multi-view data, and then offload the quantized version of extracted features to a centralized edge server, which conducts model inference based on the cascaded feature vectors. Under this setup and by considering classification tasks, we measure the inference accuracy by adopting an approximate but tractable metric, namely discriminant gain, which is defined as the distance of two classes in the Euclidean feature space under normalized covariance. To maximize the discriminant gain, we first quantify the influence of the sensing, computation, and communication processes on it with a derived closed-form expression. Then, an end-to-end task-oriented resource management approach is developed by integrating the three processes into a joint design. This integrated sensing, computation, and communication (ISCC) design approach, however, leads to a challenging non-convex optimization problem, due to the complicated form of discriminant gain and the device heterogeneity in terms of channel gain, quantization level, and generated feature subsets. Remarkably, the considered non-convex problem can be optimally solved based on the sum-of-ratios method. This gives the optimal ISCC scheme, that jointly determines the transmit power and time allocation at multiple devices for sensing and communication, as well as their quantization bits allocation for computation distortion control. By using human motions recognition as a concrete AI inference task, extensive experiments are conducted to verify the performance of our derived optimal ISCC scheme.
Dingzhu Wen, Peixi Liu, Guangxu Zhu, Yuanming Shi, Jie Xu 0002, Yonina C. Eldar, Shuguang Cui
IEEE Trans. Wirel. Commun.1
2024 Decentralized Over-the-Air Federated Learning by Second-Order Optimization Method
abstract
Federated learning (FL) is an emerging technique that enables privacy-preserving distributed learning. Most related works focus on centralized FL, which leverages the coordination of a parameter server to implement local model aggregation. However, this scheme heavily relies on the parameter server, which could cause scalability, communication, and reliability issues. To tackle these problems, decentralized FL, where information is shared through gossip, starts to attract attention. Nevertheless, current research mainly relies on first-order optimization methods that have a relatively slow convergence rate, which leads to excessive communication rounds in wireless networks. To design communication-efficient decentralized FL, we propose a novel over-the-air decentralized second-order federated algorithm. Benefiting from the fast convergence rate of the second-order method, total communication rounds are significantly reduced. Meanwhile, owing to the low-latency model aggregation enabled by over-the-air computation, the communication overheads in each round can also be greatly decreased. The convergence behavior of our approach is then analyzed. The result reveals an error term, which involves a cumulative noise effect, in each iteration. To mitigate the impact of this error term, we conduct system optimization from the perspective of the accumulative term and the individual term, respectively. Numerical experiments demonstrate the superiority of our proposed approach and the effectiveness of system optimization.
Peng Yang 0027, Yuning Jiang 0002, Dingzhu Wen, Ting Wang 0001, Colin N. Jones, Yuanming Shi
IEEE Trans. Wirel. Commun.3
2024 Over-the-Air Computation Empowered Vertically Split Inference
abstract
To tackle the issue of heterogeneous input raw data samples obtained by different devices and enhance the feature extraction capability of edge devices, we propose a vertically split neural network based edge-device collaborative artificial intelligence (AI) inference framework. The local results calculated by various light-size sub-networks at edge devices are transmitted and aggregated at the server for the downstream inference task. Nevertheless, the transmission of such high-dimensional local results involves severe communication overhead. To resolve this issue, the technique of over-the-air computation (AirComp) is adopted to enable low-latency aggregation. The same entry of all devices’ local results is transmitted over a same wireless resource block and aggregated via the waveform superposition property. Furthermore, to simultaneously support the aggregation of all dimensions of the local results, we consider a broadband channel and leverage orthogonal frequency division multiplexing (OFDM) to divide the system bandwidth into multiple subcarriers which are then assigned for different dimensions. Consequently, an extra degree of freedom is introduced to design the aggregation of all dimensions. We then propose a scheme of joint subcarrier allocation, power allocation, and receiver beamforming to minimize the aggregation distortion and enhance inference performance. Extensive experiments are conducted to verify the superiority of the proposed design over benchmarks.
Peng Yang 0027, Dingzhu Wen, Qunsong Zeng, Yong Zhou 0006, Ting Wang 0001, Haibin Cai, Yuanming Shi
IEEE Trans. Wirel. Commun.2
2024 Integrated Sensing-Communication-Computation for Over-the-Air Edge AI Inference
abstract
Edge-device co-inference refers to deploying well-trained artificial intelligent (AI) models at the network edge under the cooperation of devices and edge servers for providing ambient intelligent services. For enhancing the utilization of limited network resources in edge-device co-inference tasks from a systematic view, we propose a task-oriented scheme of integrated sensing, computation and communication (ISCC) in this work. In this system, all devices sense a target from the same wide view to obtain homogeneous noise-corrupted sensory data, from which the local feature vectors are extracted. All local feature vectors are aggregated at the server using over-the-air computation (AirComp) in a broadband channel with the orthogonal-frequency-division-multiplexing technique for suppressing the sensing and channel noise. The aggregated denoised global feature vector is further input to a server-side AI model for completing the downstream inference task. A novel task-oriented design criterion, called maximum minimum pair-wise discriminant gain, is adopted for classification tasks. It extends the distance of the closest class pair in the feature space, leading to a balanced and enhanced inference accuracy. Under this criterion, a problem of joint sensing power assignment, transmit precoding and receive beamforming is formulated. The challenge lies in three aspects: the coupling between sensing and AirComp, the joint optimization of all feature dimensions’ AirComp aggregation over a broadband channel, and the complicated form of the maximum minimum pair-wise discriminant gain. To solve this problem, a task-oriented ISCC scheme with AirComp is proposed. Experiments based on a human motion recognition task are conducted to verify the advantages of the proposed scheme over the existing scheme and a baseline.
Zeming Zhuang, Dingzhu Wen, Yuanming Shi, Guangxu Zhu, Sheng Wu 0001, Dusit Niyato
IEEE Trans. Wirel. Commun.2
2023 Decentralized Over-the-Air Computation for Edge AI Inference with Integrated Sensing and Communication
abstract
Collaborative artificial intelligent (AI) inference has been an effective approach to deploying well-trained AI models at the network edge for empowering immersive intelligent services such as autonomous driving and smart cities. In this paper, we propose an integrated sensing-computation-communication (ISCC) scheme for decentralized collaborative inference systems. In the proposed scheme, multiple devices connect to each other via device-to-device (D2D) links. Each device first extracts a homogeneous feature vector from the raw sensory data obtained from the same wide view of the source target and then aggregates all local feature vectors using the over-the-air computation technique. To further enhance the spectrum efficiency, the full-duplex technology is utilized to allow all devices to transmit and receive in the same frequency band. This, however, introduces significant self-interference and coupling among different tasks. To address these challenges, a multi-objective optimization-based ISCC approach is proposed.
Zeming Zhuang, Dingzhu Wen, Yuanming Shi
GLOBECOM2
2023 Task-Oriented Sensing, Computation, and Communication Integration for Multi-Device Edge AI
abstract
This paper studies a new multi-device edge artificial-intelligent (AI) system, which jointly exploits the AI model split inference and integrated sensing and communication (ISAC) to enable low-latency intelligent services at the network edge. In this system, multiple ISAC devices perform radar sensing to obtain multi-view data, and then offload the quantized version of extracted features to a centralized edge server, which conducts model inference based on the cascaded feature vectors. Under this setup and by considering classification tasks, we measure the inference accuracy by adopting an approximate but tractable metric, namely discriminant gain, which is defined as the distance of two classes in the Euclidean feature space under normalized covariance. To maximize the discriminant gain, we first quantify the influence of the sensing, computation, and communication processes on it with a derived closed-form expression. Then, an end-to-end task-oriented resource management approach is developed by designing an optimal integrated sensing, computation, and communication (ISCC) scheme. By using human motions recognition as a concrete AI inference task, extensive experiments are conducted to verify the performance of the proposed scheme.
Dingzhu Wen, Peixi Liu, Guangxu Zhu, Yuanming Shi, Jie Xu 0002, Yonina C. Eldar, Shuguang Cui
ICC1
2023 End-to-End Delay Minimization based on Joint Optimization of DNN Partitioning and Resource Allocation for Cooperative Edge Inference
abstract
Cooperative inference in Mobile Edge Computing (MEC), achieved by deploying partitioned Deep Neural Network (DNN) models between resource-constrained user equipments (UEs) and edge servers (ESs), has emerged as a promising paradigm. Firstly, we consider scenarios of continuous Artificial Intelligence (AI) task arrivals, like the object detection for video streams, and utilize a serial queuing model for the accurate evaluation of End-to-End (E2E) delay in cooperative edge inference. Secondly, to enhance the long-term performance of inference systems, we formulate a multi-slot stochastic E2E delay optimization problem that jointly considers model partitioning and multi-dimensional resource allocation. Finally, to solve this problem, we introduce a Lyapunov-guided Multi-Dimensional Optimization algorithm (LyMDO) that decouples the original problem into per-slot deterministic problems, where Deep Reinforcement Learning (DRL) and convex optimization are used for joint optimization of partitioning decisions and complementary resource allocation. Simulation results show that our approach effectively improves E2E delay while balancing long-term resource constraints.
Xinrui Ye, Yanzan Sun, Dingzhu Wen, Guangjin Pan, Shunqing Zhang
VTC Fall3
2023 Task-Oriented Over-the-Air Computation for Multi-Device Edge Split Inference
abstract
A task-oriented over-the-air computation (AirComp) scheme is proposed in this paper for multi-device edge split inference system. In the considered system, local noise-corrupted feature vectors are aggregated at the server via AirComp to generate a denoised one for the subsequent inference task. By considering classification tasks, the transmit precoders at edge devices and receive beamforming at edge server are jointly designed in an effort to rein in the aggregation error and maximize the inference accuracy, which is approximately measured by a surrogate but more tractable metric called discriminant gain. It is found that the conventional AirComp beamforming design for minimizing the mean square error between the aggregated feature vector by AirComp and the ideally aggregated one may not lead to the optimal classification accuracy, as it fails to respect the fact that some feature dimensions are more sensitive to the aggregation error than the others in terms of the classification accuracy. To tackle this issue, a new task-oriented AirComp scheme is proposed for directly maximizing the derived discriminant gain. The superiority of the proposed scheme over the heuristic benchmarks is verified by extensive experimental results based on a concrete inference task of human motion recognition.
Dingzhu Wen, Xiang Jiao, Peixi Liu, Guangxu Zhu, Yuanming Shi, Kaibin Huang
WCNC1
2023 Communication and Energy Efficient Decentralized Learning Over D2D Networks
abstract
Device-to-device (D2D)-assisted decentralized learning has been proposed for mobile devices to collaboratively train artificial intelligence networks without the centralized parameter server. However, a densely connected network will cause large learning latency and energy consumption due to the limited computation and communication resources. In addition, link selection and aggregation weight have a significant impact on the learning performance. To cope with these challenges, we propose a joint computing power adjustment, wireless resource allocation, link selection, and aggregation weight adaptation mechanism to improve both communication and energy efficiencies. Specifically, the learning performances including the convergence rate, per-iteration learning latency, and per-iteration energy consumption are first analyzed. Then, an optimization problem is formulated to minimize the total learning cost, which is defined as the weighted sum of total learning latency and energy consumption. Given a network topology, the computing power and wireless resource allocation are optimized by the alternating optimization algorithm. Moreover, the optimal aggregation weight is obtained by semidefinite programming. With respect to link selection, we propose a tabu search based meta-heuristic algorithm to approximately achieve feasible solutions with a low computational complexity. Finally, extensive experiments demonstrate that the proposed link selection algorithm can significantly reduce the learning cost under the given learning accuracy requirement.
Shengli Liu 0002, Guanding Yu, Dingzhu Wen, Xianfu Chen, Mehdi Bennis, Hongyang Chen 0001
IEEE Trans. Wirel. Commun.3
2021 Adaptive Subcarrier, Parameter, and Power Allocation for Partitioned Edge Learning Over Broadband Channels
abstract
In this paper, we considerpartitioned edge learning(PARTEL), which implements parameter-server training, a well known distributed learning method, in a wireless network. Thereby, PARTEL leverages distributed computation resources at edge devices to train a large-scaleartificial intelligence(AI) model by dynamically partitioning the model into parametric blocks for separated updating at devices. Targeting broadband channels, we consider the joint control of parameter allocation, sub-channel allocation, and transmission power to improve the performance of PARTEL. Specifically, the policies for joint SUbcarrier, Parameter, and POweR allocaTion (SUPPORT) are optimized under the criterion of minimum learning latency. Two cases are considered. First, for the case of decomposable models (e.g., logistic regression), the latency-minimization problem is a mixed-integer program and non-convex. Due to its intractability, we develop a practical solution by integer relaxation and transforming it into an equivalent convex problem of model size maximization under a latency constraint. Thereby, a low-complexity algorithm is designed to compute the SUPPORT policy. Second, consider the case ofdeep neural network(DNN) models which can be trained using PARTEL by introducing some auxiliary variables. This, however, introduces constraints on model partitioning reducing the granularity of parameter allocation. The preceding policy is extended to DNN models by applying the proposed techniques of load rounding and proportional adjustment to rein in latency expansion caused by the load granularity constraints.
Dingzhu Wen, Ki Jun Jeon, Mehdi Bennis, Kaibin Huang
IEEE Trans. Wirel. Commun.1
2020 Accelerating Partitioned Edge Learning via Joint Parameter-and-Bandwidth Allocation
abstract
In this paper, we consider the framework of partitioned edge learning for iteratively training a large-scale model using many resource-constrained devices (called workers). To this end, in each iteration, the model is dynamically partitioned into parametric blocks, which are downloaded to worker groups for updating using their local data. Then, the local updates are uploaded to and cascaded by the server for updating a global model. To reduce resource usage by minimizing the total learning-and-communication latency, this work focuses on the novel joint design of parameter (computation load) and bandwidth allocation (for downloading and uploading). Two design approaches are adopted. First, a practical sequential approach, called partially integrated parameter-and-bandwidth allocation (PABA), yields one scheme, namely parameter aware bandwidth allocation. It allocates the largest bandwidth to the slowest worker. Second, PABA are jointly optimized. Despite its being a nonconvex problem, an efficient and optimal solution algorithm is derived by intelligently nesting a bisection search and solving a convex problem. Experimental results using real data demonstrate that integrating PABA can substantially improve the performance of partitioned edge learning in terms of latency (by e.g., 46%) and accuracy (by e.g., 4%).
Dingzhu Wen, Mehdi Bennis, Kaibin Huang
GLOBECOM1
2020 Scheduling for Cellular Federated Edge Learning With Importance and Channel Awareness
abstract
In cellular federated edge learning (FEEL), multiple edge devices holding local data jointly train a neural network by communicating learning updates with an access point without exchanging their data samples. With very limited communication resources, it is beneficial to schedule the most informative local learning updates. This paper focuses on FEEL with gradient averaging over participating devices in each round of communication. A novel scheduling policy is proposed to exploit both diversity in multiuser channels and diversity in the “importance” of the edge devices' learning updates. First, a new probabilistic scheduling framework is developed to yield unbiased update aggregation in FEEL. The importance of a local learning update is measured by its gradient divergence. If one edge device is scheduled in each communication round, the scheduling policy is derived in closed form to achieve the optimal trade-off between channel quality and update importance. The probabilistic scheduling framework is then extended to allow scheduling multiple edge devices in each communication round. Numerical results obtained using popular models and learning datasets demonstrate that the proposed scheduling policy can achieve faster model convergence and higher learning accuracy than conventional scheduling policies that only exploit a single type of diversity.
Jinke Ren, Yinghui He, Dingzhu Wen, Guanding Yu, Kaibin Huang, Dongning Guo
IEEE Trans. Wirel. Commun.3
2020 Joint Parameter-and-Bandwidth Allocation for Improving the Efficiency of Partitioned Edge Learning
abstract
To leverage data and computation capabilities of mobile devices, machine learning algorithms are deployed at the network edge for training artificial intelligence (AI) models, resulting in the new paradigm of edge learning. In this paper, we consider the framework of partitioned edge learning for iteratively training a large-scale model using many resource-constrained devices (called workers). To this end, in each iteration, the model is dynamically partitioned into parametric blocks, which are downloaded to worker groups for updating using data subsets. Then, the local updates are uploaded to and cascaded by the server for updating a global model. To reduce resource usage by minimizing the total learning-and-communication latency, this work focuses on the novel joint design of parameter (computation load) allocation and bandwidth allocation (for downloading and uploading). Two design approaches are adopted. First, a practical sequential approach, called partially integrated parameter-and-bandwidth allocation (PABA), yields two schemes, namely bandwidth aware parameter allocation and parameter aware bandwidth allocation. The former minimizes the load for the slowest (in computing) of worker groups, each training a same parametric block. The latter allocates the largest bandwidth to the worker being the latency bottleneck. Second, PABA are jointly optimized. Despite it being a nonconvex problem, an efficient and optimal solution algorithm is derived by intelligently nesting a bisection search and solving a convex problem. Experimental results using real data demonstrate that integrating PABA can substantially improve the performance of partitioned edge learning in terms of latency (by e.g., 46%) and accuracy (by e.g., 4% given the latency of 100 seconds).
Dingzhu Wen, Mehdi Bennis, Kaibin Huang
IEEE Trans. Wirel. Commun.1
2019 Reduced-Dimension Design of MIMO AirComp for Data Aggregation in Clustered IoT Networks
abstract
One basic operation of Internet-of-Things (IoT) networks is to acquire a function of distributed data collected from sensors over wireless channels, called wireless data aggregation (WDA). Targeting dense sensors, low-latency WDA poses a design challenge for high-mobility or mission critical IoT applications. A promising solution is a low- latency multi-access scheme, called over-the-air computing (AirComp), that supports simultaneous transmission such that an access point (AP) can estimate and receive a summation-form function of the distributed data by exploiting the waveform- superposition property of multi-access channels. In this work, we propose a multiple-input-multiple-output (MIMO) AirComp framework for an IoT network with clustered multi-antenna sensors and an AP with large receive arrays. The contributions of this work are two-fold. Define the AirComp error as the error in the functional value received at AP due to channel noise. First, under the criterion of minimum error, the optimal receive beamformer at the AP, called decomposed aggregation beamformer (DAB), is shown to have a decomposed architecture: the inner component focuses on channel-dimension reduction and the outer component focuses on joint equalization of the resultant low-dimensional small-scale fading channels. Second, to provision DAB with the required channel state information (CSI), a low-latency channel feedback scheme is proposed by intelligently leveraging the AirComp principle to support simultaneous channel- feedback by sensors.
Dingzhu Wen, Guangxu Zhu, Kaibin Huang
GLOBECOM1
2019 Reduced-Dimension Design of MIMO Over-the-Air Computing for Data Aggregation in Clustered IoT Networks
abstract
One basic operation of Internet-of-Things (IoT) networks is to acquire a function of distributed data collected from sensors over wireless channels, called wireless data aggregation (WDA). In the presence of dense sensors, low-latency WDA poses a design challenge for high-mobility or mission critical IoT applications. A promising solution is a low-latency multi-access scheme, called over-the-air computing (AirComp), that supports simultaneous transmission such that an access point (AP) can estimate and receive a summation-form function of the distributed sensing data by exploiting the waveform-superposition property of a multi-access channel. In this work, we propose a multiple-input-multiple-output (MIMO) AirComp framework for an IoT network with clustered multi-antenna sensors and an AP with large receive arrays. The framework supports low-complexity and low-latency AirComp of a vector-valued function. The contributions of this work are two-fold. Define the AirComp error as the error in the functional value received at AP due to channel noise. First, under the criterion of minimum error, the optimal receive beamformer at the AP, called decomposed aggregation beamformer (DAB), is shown to have a decomposed architecture: the inner component focuses on channel-dimension reduction and the outer component focuses on joint equalization of the resultant low-dimensional small-scale fading channels. In addition, an algorithm is designed to adjust the ranks of individual components of the DAB for a further performance improvement. Second, to provision DAB with the required channel state information (CSI), a low-latency channel feedback scheme is proposed by intelligently leveraging the AirComp principle to support simultaneous channel feedback by sensors. The proposed framework is shown by simulation to substantially reduce AirComp error compared with the existing design without considering channel structures.
Dingzhu Wen, Guangxu Zhu, Kaibin Huang
IEEE Trans. Wirel. Commun.1
2018 Hybrid full-/half-duplex cellular networks: user admission and power control
abstract
We consider a single-cell network with a hybrid full-/half-duplex base station. For the practical scenario with N channels, K uplink users, and M downlink users (max{ K , M } ≤ N ≤ K + M ), we tackle the issue of user admission and power control to simultaneously maximize the user admission number and minimize the total transmit power when guaranteeing the quality-of-service requirement of individual users. We formulate a 0–1 integer programming problem for the joint-user admission and power allocation problem. Because finding the optimal solution of this problem is NP-hard in general, a low-complexity algorithm is proposed by introducing the novel concept of adding dummy users. Simulation results show that the proposed algorithm achieves performance similar to that of branch and bound algorithm and significantly outperforms the random pairing algorithm.
Dingzhu Wen, Caijun Zhong, Guanding Yu
Frontiers Inf. Technol. Electron. Eng.2
2017 Results on Energy- and Spectral-Efficiency Tradeoff in Cellular Networks With Full-Duplex Enabled Base Stations
abstract
In this paper, we address the tradeoff between energy efficiency (EE) and spectral efficiency (SE) for cellular networks with full-duplex (FD) communications enabled base stations. To be backward compatible with legacy LTE systems, it is assumed that user devices still work in the conventional half-duplex (HD) mode. There usually exists residual self-interference (RSI) in FD communications after advanced interference suppression techniques are applied. In this paper, we consider two different RSI models: constant RSI model and linear RSI model. First, the necessary conditions for an FD transceiver to achieve better EE-SE tradeoff than an HD one are derived for both the RSI models. Then, for the constant RSI model, a closed-form EE-SE expression is obtained in the scenario of single pair of users. We further extend our result and prove that EE is a quasi-concave function of SE in the scenario of multiple user pairs. Accordingly, an optimal algorithm to achieve the maximum EE based on the Lagrange dual decomposition technique is developed. For the linear RSI model, the EE-SE relation is difficult to deal with and we develop a heuristic algorithm by decoupling the problem into two sub-problems: power control and resource allocation. Our analysis and algorithms are finally verified by comprehensive numerical results.
Dingzhu Wen, Guanding Yu, Rongpeng Li, Yan Chen 0010, Geoffrey Ye Li
IEEE Trans. Wirel. Commun.1
2016 Energy-efficient mode selection and power control for device-to-device communications
abstract
Device-to-device (D2D) communications have been a promising technique for future long term evolution (LTE) systems. In this paper, we aim to maximize the system energy efficiency (EE) of a communication system including both conventional cellular users and D2D users. We consider two different D2D access cases depending on whether all D2D users can be accessed or not. In both cases, the EE maximization problem is formulated as a combinatorial fractional programming problem. We then use the fractional programming and the branch-and-bound (BnB) method to find the optimal solution. Furthermore, low complexity algorithms are proposed according to the network load. Simulation results demonstrate the effectiveness of the proposed algorithms.
Dingzhu Wen, Guanding Yu, Lukai Xu
WCNC1
2016 Joint user scheduling and channel allocation for cellular networks with full duplex base stations
abstract
Full‐duplex communication (FDC) can potentially double the network capacity by allowing a device to transmit and receive simultaneously on the same frequency band. In this study, a novel resource allocation and user scheduling algorithm is proposed to maximise the network throughput for a cellular network with full‐duplex (FD) base stations (BSs). The authors consider that FDC is utilised at the BS with imperfect self‐interference (SI) cancellation while user devices only work in the traditional half‐duplex (HD) way. In addition, to potentially cancel co‐channel interference caused by other users, the opportunistic interference cancellation (OIC) technique is applied at user side. Since FDC does not always perform better than HD due to residual SI (RSI), a joint mode selection, user scheduling, and channel allocation problem is formulated to maximise the system throughput. The optimisation problem is non‐convex and NP‐hard, thereby a suboptimal heuristic algorithm with low computational complexity is proposed. Numerical results demonstrate that user diversity gain, FD gain, and OIC gain can be achieved by the proposed algorithm, respectively. The performance of FDC depends on the intensity of RSI and the distribution of user devices.
Guanding Yu, Dingzhu Wen, Fengzhong Qu
IET Commun.2