Guangxu Zhu

dblp:140/7624 · DBLP profile ↗
← Back
104ranked-venue papers
16as first author
82since 2021 · last 2026
0000-0001-9532-9201ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 83 · 15 first-author · 62 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Systems, architecture and hardware · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FeedSign: Robust and Communication-Efficient Federated Fine-tuning of Large Models for Edge AI
Zhijie Cai, Haolong Chen, Guangxu Zhu, Qingjiang Shi, Kaibin Huang
ICC3
2026 Sensing Performance Analysis in Cooperative Air-Ground ISAC Networks for LAE
abstract
To support the development of low altitude economy, the air-ground integrated sensing and communication (ISAC) networks need to be constructed to provide reliable and robust communication and sensing services. In this paper, the sensing capabilities in the cooperative air-ground ISAC networks are evaluated in terms of area radar detection coverage probability under a constant false alarm rate, where the distribution of aggregated sensing interferences is analyzed as a key intermediate result. Compared with the analysis based on the strongest interferer approximation, taking the aggregated sensing interference into consideration is better suited for pico-cell scenarios with high base station density. Simulations are conducted to validate the analysis.
Yihang Jiang 0001, Xiaoyang Li 0002, Guangxu Zhu, Xiaowen Cao 0001, Kaifeng Han, Bingpeng Zhou, Xinyi Wang 0002
ICC3
2026 System Design and Convergence Analysis for Decentralized Federated Fine-Tuning on LEO Satellite Networks
Zhigang Yan, Guangxu Zhu, Haoyuan Pan, Nikolaos Pappas, Tse-Tin Chan
ICC2
2026 FedRMamba: Federated Residual Mamba for Multivariate Time-Series Forecasting
abstract
Time series forecasting underpins many real-world services. Recent trends have focused on foundation models inspired by the paradigm of large language models, which rely on large volumes of centralized time-series data across diverse domains. However, such approaches raise significant concerns regarding data privacy. Federated learning (FL) has emerged as a promising paradigm for training unified time-series models using isolated datasets distributed across multiple clients. Nevertheless, existing FL methods face two critical challenges: heterogeneous variables and heterogeneous temporal correlations. To address these issues, we propose FedRMamba, a personalized federated forecasting framework built entirely from Mamba state-space blocks. Each client adopts a residual-coupled architecture, where a global frequency-aware Mamba module captures the common low-frequency structures shared across different variables, while a local patch-wise Mamba module learns personalized high-frequency patterns within the multivariate context. To clearly separate these responsibilities, we introduce a frequency-aware supervision that aligns the global path with low-frequency components and the local path with high-frequency residuals. Additionally, we design a gated fusion mechanism that dynamically combines the low-frequency and high-frequency components for improved prediction. We conduct extensive experiments to evaluate the performance of our proposed framework, demonstrating its effectiveness in handling heterogeneous data in federated settings.
Zhiwei Hu, Liang Zhang 0042, Guangxu Zhu
WWW3
2026 An overview of domain-specific foundation model: key technologies, applications and challenges
Haolong Chen, Hanzhi Chen, Zijian Zhao 0002, Kaifeng Han, Guangxu Zhu, Yichen Zhao, Wei Xu 0001, Qingjiang Shi
Sci. China Inf. Sci.5
2026 Low-complexity hybrid beamforming for multi-cell mmWave massive MIMO: A primitive Kronecker decomposition approach
Guangxu Zhu, Xiaofan Li 0001, Jiancun Fan, Minghua Xia
Signal Process.2
2026 FeedSign: Robust Full-Parameter Federated Fine-Tuning of Large Models With Extremely Low Communication Overhead of One Bit
Zhijie Cai, Haolong Chen, Guangxu Zhu, Qingjiang Shi, Kaibin Huang
IEEE Trans. Mob. Comput.3
2026 Joint Sensing, Communication, and Computation for Vertical Federated Edge Learning in Edge Perception Networks
abstract
Combining wireless sensing and edge intelligence, edge perception networks enable intelligent data collection and processing at the network edge. However, traditional sample partition based horizontal federated edge learning (HFEEL) struggles to effectively fuse complementary multi-view information from distributed devices. To address this limitation, we propose a vertical federated edge learning (VFEEL) framework tailored for feature-partitioned sensing data. In this paper, we consider an integrated sensing, communication, and computation (ISCC)-enabled edge perception network, where multiple edge devices utilize wireless signals to sense environmental information for updating their local models, and the edge server aggregates feature embeddings via over-the-air computation (AirComp) for global model training. First, we analyze the convergence behavior of the ISCC-enabled VFEEL in terms of the loss function degradation in the presence of wireless sensing noise and aggregation distortions during AirComp. Then, to accelerate convergence, we aim to optimize the batch size, sensing power, and transmission power control at edge devices as well as the denoising factors at the edge server under limited network constraints on overall energy consumption and per-round latency. Due to the tight coupling of variables, the problem is non-convex. To address this problem, we design an alternating optimization-based algorithm to efficiently obtain a high-quality solution. Numerical results are conducted based on a human motion recognition task to verify that the proposed ISCC-enabled VFEEL algorithm achieves higher accuracy compared with other benchmarking schemes including ISCC-enabled HFEEL approach.
Xiaowen Cao 0001, Dingzhu Wen, Suzhi Bi, Yuanhao Cui, Guangxu Zhu, Han Hu 0003, Yonina C. Eldar
IEEE Trans. Mob. Comput.5
2026 Cargo UAVs Pick-Up Systems for Low-Altitude Economy With Communication Quality, Battery Energy, and Time Window Constraints
abstract
The rapid development of the low-altitude economy (LAE) has accelerated the deployment of cargo unmanned aerial vehicles (UAVs) for intelligent logistics and delivery services. However, large-scale UAV operations still face multiple practical challenges, including unstable communication connectivity, limited onboard battery energy, and strict customer time-window constraints. To address these issues, this paper investigates the trajectory and task scheduling optimization problem for multi-UAV cooperative cargo pick-up under joint communication, energy, and time-window constraints. We develop a collision-aware cooperative multi-UAV optimization algorithm (CACMO) that integrates a Dueling Deep Q-Network (D3QN) for communication-aware trajectory learning with a simulated annealing (SA) based global task-sequence planner and an explicit inter-UAV conflict-resolution mechanism. The D3QN module enables adaptive trajectory generation in unknown and time-varying radio environments without requiring an a priori radio map, maintaining stable connectivity while reducing flight cost, whereas the SA module determines efficient task orders and enforces safe coordination among multiple UAVs through collision-aware refinement. Simulation results demonstrate that the proposed CACMO algorithm framework achieves an optimal balance between task completion time (1,719 seconds) and user satisfaction (score of 0.9969) under typical operating conditions, delivering a 70–75% reduction in total weighted cost compared to representative baseline methods. Crucially, this substantial improvement is achieved while explicitly enforcing multi-UAV collision avoidance-a critical constraint absent in most baseline methods. The framework maintains zero communication outage and guarantees safe inter-UAV separation throughout the mission while satisfying all energy and time window constraints in realistic urban environments, confirming its robustness and scalability for cooperative multi-UAV logistics operations within the LAE.
Liang Yang 0001, Jiangling Cao, Guangxu Zhu, Weijie Yuan 0001, Hongbo Jiang 0001, Dusit Niyato
IEEE Trans. Mob. Comput.4
2026 A Disentangled Representation Learning Framework for Low-Altitude Network Coverage Prediction
abstract
The expansion of the low-altitude economy has underscored the significance of Low-Altitude Network Coverage (LANC) prediction for designing aerial corridors. While accurate LANC forecasting hinges on the antenna beam patterns of Base Stations (BSs), these patterns are typically proprietary and not readily accessible. Operational parameters of BSs, which inherently contain beam information, offer an opportunity for data-driven low-altitude coverage prediction. However, collecting extensive low-altitude road test data is cost-prohibitive, often yielding only sparse samples per BS. This scarcity results in two primary challenges: imbalanced feature sampling due to limited variability in high-dimensional operational parameters against the backdrop of substantial changes in low-dimensional sampling locations, and diminished generalizability stemming from insufficient data samples. To overcome these obstacles, we introduce a dual strategy comprising expert knowledge-based feature compression and disentangled representation learning. The former reduces feature space complexity by leveraging communications expertise, while the latter enhances model generalizability through the integration of propagation models and distinct subnetworks that capture and aggregate the semantic representations of latent features. Experimental evaluation con firms the efficacy of our framework, yielding a 7% reduction in error compared to the best baseline algorithm. Real-network validations further attest to its reliability, achieving practical prediction accuracy with MAE errors at the 5 dB level.
Zhijie Cai, Nan Qi 0001, Chao Dong 0001, Guangxu Zhu, Haixia Ma, Qihui Wu 0001, Shi Jin 0002
IEEE Trans. Mob. Comput.5
2026 FlexSpec: Frozen Drafts Meet Evolving Targets in Edge-Cloud Collaborative LLM Speculative Decoding
abstract
Deploying large language models (LLMs) in mobile and edge computing environments is constrained by limited on-device resources, scarce wireless bandwidth, and frequent model evolution. Although edge-cloud collaborative inference with speculative decoding (SD) can reduce end-to-end latency by executing a lightweight draft model at the edge and verifying it with a cloud-side target model, existing frameworks fundamentally rely on tight coupling between the two models. Consequently, repeated model synchronization introduces excessive communication overhead, increasing end-to-end latency, and ultimately limiting the scalability of SD in edge environments. To address these limitations, we propose FlexSpec, a communication-efficient collaborative inference framework tailored for evolving edge-cloud systems. The core design of FlexSpec is a shared-backbone architecture that allows a single and static edge-side draft model to remain compatible with a large family of evolving cloud-side target models. By decoupling edge deployment from cloud-side model updates, FlexSpec eliminates the need for edge-side retraining or repeated model downloads, substantially reducing communication and maintenance costs. Furthermore, to accommodate time-varying wireless conditions and heterogeneous device constraints, we develop a channel-aware adaptive speculation mechanism that dynamically adjusts the speculative draft length based on real-time channel state information and device energy budgets. Extensive experiments demonstrate that FlexSpec achieves superior performance compared to conventional SD approaches in terms of inference efficiency.
Yuchen Li 0006, Zhonghao Lyu, Qiyang Li, Hengyi Cai, Lingyong Yan, Shuaiqiang Wang, Jiashu Zhao, Guangxu Zhu, Linghe Kong, Guihai Chen, Haoyi Xiong, Dawei Yin 0001
IEEE Trans. Mob. Comput.10
2026 KNN-MMD: Cross Domain Wireless Sensing via Local Distribution Alignment
abstract
Wireless sensing has recently found widespread applications in diverse environments, including homes, offices, and public spaces. By analyzing patterns in channel state information (CSI), it is possible to infer human actions for tasks such as person identification, gesture recognition, and fall detection. However, CSI is highly sensitive to environmental changes, where even minor alterations can significantly distort the CSI patterns. This sensitivity often leads to performance degradation or outright failure when applying wireless sensing models trained in one environment to another. To address this challenge, Domain Alignment Learning (DAL) has been widely adopted for cross-domain classification tasks, as it focuses on aligning the global distributions of the source and target domains in feature space. Despite its popularity, DAL often neglects inter-category relationships, which can lead to misalignment between categories across domains, even when global alignment is achieved. To overcome these limitations, we propose K-Nearest Neighbors Maximum Mean Discrepancy (KNN-MMD), a novel few-shot method for cross-domain wireless sensing. Our approach begins by constructing a “help set” using K-Nearest Neighbors (KNN) from the target domain, enabling local alignment between the source and target domains within each category using Maximum Mean Discrepancy (MMD). Additionally, we address a key instability issue commonly observed in cross-domain methods, where model performance fluctuates sharply between epochs. Further, most existing methods struggle to determine an optimal stopping point during training due to the absence of labeled data from the target domain. Our method resolves this by excluding the support set from the target domain during training and employing it as a validation set to determine the stopping criterion. We evaluate the effectiveness of the proposed method across several cross-domain Wi-Fi sensing tasks, including gesture recognition, person identification, fall detection, and action recognition, using both a public dataset and a self-collected dataset. In a one-shot scenario, our method achieves accuracy rates of 93.26%, 81.84%, 77.62%, and 75.30% for the respective tasks. The dataset and code are publicly available athttps://github.com/RS2002/KNN-MMD.
Zijian Zhao 0002, Zhijie Cai, Xiaoyang Li 0002, Hang Li 0003, Qimei Chen, Guangxu Zhu
IEEE Trans. Mob. Comput.7
2026 CSI-BERT2: A BERT-Inspired Framework for Efficient CSI Prediction and Classification in Wireless Communication and Sensing
abstract
Channel state information (CSI) is a fundamental component in both wireless communication and sensing systems, enabling critical functions such as radio resource optimization and environmental perception. In wireless sensing, data scarcity and packet loss hinder efficient model training, while in wireless communication, high-dimensional CSI matrices and short coherent times caused by high mobility present challenges in CSI estimation. To address these issues, we propose a unified framework named CSI-BERT2 for CSI prediction and classification tasks, built on our previous work CSI-BERT, which adapts BERT to capture the complex relationships among CSI sequences through a bidirectional self-attention mechanism. We introduce a two-stage training method that first uses a mask language model (MLM) to enable the model to learn general feature extraction from scarce datasets in an unsupervised manner, followed by fine-tuning for specific downstream tasks. Specifically, we extend MLM into a mask prediction model (MPM), which efficiently addresses the CSI prediction task. To further enhance the representation capacity of CSI data, we modify the structure of the original CSI-BERT. We introduce an adaptive re-weighting layer (ARL) to enhance subcarrier representation and a multi-layer perceptron (MLP)-based temporal embedding module to mitigate temporal information loss problem inherent in the original Transformer. Extensive experiments on both real-world collected and simulated datasets demonstrate that CSI-BERT2 achieves state-of-the-art performance across all tasks. Our results further show that CSI-BERT2 generalizes effectively across varying sampling rates and robustly handles discontinuous CSI sequences caused by packet loss-challenges that conventional methods fail to address. The dataset and code are publicly available athttps://github.com/RS2002/CSI-BERT2.
Zijian Zhao 0002, Zhonghao Lyu, Hang Li 0003, Xiaoyang Li 0002, Guangxu Zhu
IEEE Trans. Mob. Comput.6
2026 DK-Root: A Joint Data-and-Knowledge-Driven Framework for Root Cause Analysis of QoE Degradations in Mobile Networks
abstract
Diagnosing the root causes of Quality of Experience (QoE) degradations in operational mobile networks is challenging due to complex cross-layer interactions among kernel performance indicators (KPIs) and the scarcity of reliable expert annotations. Although rule-based heuristics can generate labels at scale, they are noisy and coarse-grained, limiting the accuracy of purely data-driven approaches. To address this, we propose DK-Root, a joint data-and-knowledge-driven framework that unifies scalable weak supervision with precise expert guidance for robust root-cause analysis. DK-Root first pretrains an encoder via contrastive representation learning using abundant rule-based labels while explicitly denoising their noise through a supervised contrastive objective. To supply task-faithful data augmentation, we introduce a class-conditional diffusion model that generates KPIs sequences preserving root-cause semantics, and by controlling reverse diffusion steps, it produces weak and strong augmentations that improve intra-class compactness and inter-class separability. Finally, the encoder and the lightweight classifier are jointly fine-tuned with scarce expert-verified labels to sharpen decision boundaries. Extensive experiments on a real-world, operator-grade dataset demonstrate state-of-the-art accuracy, with DK-Root surpassing traditional ML and recent semi-supervised time-series methods. Ablations confirm the necessity of the conditional diffusion augmentation and the pretrain-finetune design, validating both representation quality and classification gains.
Qizhe Li, Haolong Chen, Jiansheng Li, Shuqi Chai, Yuzhou Hou, Xinhua Shao, Kaifeng Han, Guangxu Zhu
IEEE Trans. Netw.10
2026 A Two-Timescale Resource Allocation Method Based on Deep Reinforcement Learning for 6G Networks
abstract
With the rapid development of artificial intelligence and the dramatic growth of communication services, the sixth-generation (6G) wireless network needs to handle communication tasks more flexibly and efficiently, significantly exacerbating the challenge of resource allocation. For the access network scenarios in 6G networks, the existing single-layer reinforcement learning resource allocation algorithms are hard to satisfy the diverse demands of users due to the complex and variable state space. Therefore, we propose a reinforcement learning-based two-timescale resource allocation scheme, aiming to jointly enhance the quality of service and system resource utilization. The proposed method comprises an upper-layer controller that allocates network resources to lower-layer controllers on a large time scale. Then, lower-layer controllers refine the resources based on user service types on a smaller time scale. To implement the proposed two-timescale allocation scheme, we propose a two-layer reinforcement learning framework consisting of a deep deterministic policy gradient (DDPG) and a dueling deep Q network (Dueling-DQN). Furthermore, recognizing that coupling multiple reinforcement learning processes may slow down algorithm convergence, we employ asynchronous training, transfer learning, and prediction-based action space simplification to expedite the model’s convergence speed. Finally, we build a prototyping network to verify the performance of the proposed small-timescale and the large-timescale allocation algorithms. Our proposed scheme demonstrates significant improvements in both resource utilization and quality of service compared to existing schemes.
Fan Xu 0001, Guangxu Zhu, Hang Li 0003, Xiongyan Tang, Lexi Xu, Guorong Zhou
IEEE Trans. Netw.3
2026 An Energy-Efficient Wireless Communication and Control Co-Design for WNCS
abstract
To facilitate the development of industrial Internet of Things applications, thewireless networked control system(WNCS) is envisioned to support real-time control and communication interactions performed in finite-time manner. A WNCS comprising multiple wirelessly interconnectedsub-systems(SSs) is considered, wherein the sensed state information in each SS is transmitted to the controller via wireless links, thereby enabling timely decision-making processes. Following multiple operation periods of state sensing and transmission, thesystem identification(SI) is performed, leading to the formulation of optimal control policy. To improve the energy efficiency while guaranteeing the SI performance requirement within the allowed decision-making time, the communication and control co-design for WNCS is investigated, where the transmit power, transmission interval length, number of operation periods, coding block-length, and required transmission reliability are jointly optimized. Our investigation demonstrates the interrelationships among effective capacity, energy consumption, and communication parameters. Furthermore, it is found that the optimal communication parameters, such as transmit power and transmission interval length, should be determined by both communication and control requirements. Consequently, it is found that minimizing energy consumption is equivalent to minimize the number of operation periods while guaranteeing the SI performance with defined confidence, which can be effectively addressed by leveraging the non-decreasing property of controllability Gramian. Moreover, the co-design framework is extended to accommodate the scenarios involving link interruptions and overlapping time slots. Simulation results validate the necessity and effectiveness of exploring optimal system operational configurations from the perspective of the proposed co-design. It is also observed that although a 44.2% surge in energy consumption is associated with the proposed relay scheme in the link interruption case, the proposed time scheduling scheme brings a 23.4% reduction in the extra energy expenditure (from 44.2% to 20.8%).
Xiaoyang Li 0002, Guangxu Zhu, Kaibin Huang, Yi Gong 0001, Qinyu Zhang 0001
IEEE Trans. Wirel. Commun.3
2026 Communication-and-Computation Efficient Split Federated Learning in Wireless Networks: Gradient Aggregation and Resource Management
abstract
With the prevalence of emerging artificial intelligence services in next-generation wireless edge networks, Split Federated Learning (SFL), which divides a learning model into server-side and client-side models, has emerged as an appealing technology to deal with the heavy computational burden for network edge clients. However, existing SFL frameworks would frequently upload smashed data and download gradients between the server and each client, leading to severe communication overheads. To address this issue, this work proposes a novel communication-and-computation efficient SFL framework, which allows dynamic model splitting (server- and client-side model cutting point selection) and broadcasting of aggregated smashed data gradients. We theoretically analyze the impact of the cutting point selection on the convergence rate, revealing that model splitting with a smaller client-side model size leads to a better convergence performance and vise versa. Based on the above insights, we formulate an optimization problem to minimize the model convergence rate and latency under the consideration of data privacy via a joint Cutting point selection, Communication and Computation resource allocation (CCC) strategy. To deal with the proposed mixed integer nonlinear programming optimization problem, we develop an algorithm by integrating the Double Deep Q-learning Network (DDQN) with convex optimization methods. Extensive experiments validate our theoretical analyses across various datasets, and the numerical results demonstrate the effectiveness and superiority of the proposed communication-efficient SFL compared with existing schemes, including parallel split learning and traditional SFL mechanisms.
Yipeng Liang, Qimei Chen, Rongpeng Li, Guangxu Zhu, Muhammad Kaleem Awan, Hao Jiang 0010
IEEE Trans. Wirel. Commun.4
2026 UAV-Assisted Edge Inference With Integrated Sensing, Communication, and Computation
Dingzhu Wen, Guangxu Zhu, Yuan Liu 0001, Yuanming Shi, Honglin Hu
IEEE Trans. Wirel. Commun.3
2026 Semantics-Guided Diffusion for Deep Joint Source-Channel Coding in Wireless Image Transmission
abstract
Joint source-channel coding (JSCC) offers a promising avenue for enhancing transmission efficiency by jointly incorporating source and channel statistics into the system design. A key advancement in this area is the deep joint source and channel coding (DeepJSCC) technique that designs a direct mapping of input signals to channel symbols parameterized by a neural network, which can be trained for arbitrary channel models and semantic quality metrics. This paper advances the DeepJSCC framework toward a semantics-aligned, high-fidelity transmission approach, called semantics-guided diffusion DeepJSCC (SGD-JSCC). Existing schemes that integrate diffusion models (DMs) with JSCC face challenges in transforming random generation into accurate reconstruction and adapting to varying channel conditions. SGD-JSCC incorporates two key innovations: (1) utilizing some inherent information that contributes to the semantics of an image, such as text description or edge map, to guide the diffusion denoising process; and (2) enabling seamless adaptability to varying channel conditions with the help of a semantics-guided DM for channel denoising. The DM is guided by diverse semantic information and integrates seamlessly with DeepJSCC. In a slow fading channel, SGD-JSCC dynamically adapts to the instantaneous channel state information (CSI) directly estimated from the channel output, thereby eliminating the need for additional pilot transmissions for channel estimation. In a fast fading channel, we introduce a training-free denoising strategy, allowing SGD-JSCC to effectively adjust to fluctuations in channel gains. Numerical results demonstrate that, guided by semantic information and leveraging the powerful DM, our method outperforms existing DeepJSCC schemes, delivering satisfactory reconstruction performance even at extremely poor channel conditions. The proposed scheme highlights the potential of incorporating diffusion models in future communication systems. The code and pretrained checkpoints will be publicly available at https://github.com/MauroZMJ/SGDJSCC, allowing integration of this scheme with existing DeepJSCC models, without the need for retraining from scratch.
Maojun Zhang, Guangxu Zhu, Richeng Jin, Xiaoming Chen 0001, Deniz Gündüz
IEEE Trans. Wirel. Commun.3
2025 Personalizing Low-Rank Bayesian Neural Networks Via Federated Learning
abstract
To support real-world decision-making, it is crucial for models to be well-calibrated, i.e., to assign reliable confidence estimates to their predictions. Uncertainty quantification is particularly important in personalized federated learning (PFL), as participating clients typically have small local datasets, making it difficult to unambiguously determine optimal model parameters. Bayesian PFL (BPFL) methods can potentially enhance calibration, but they often come with considerable computational and memory requirements due to the need to track the variances of all the individual model parameters. Furthermore, different clients may exhibit heterogeneous uncertainty levels owing to varying local dataset sizes and distributions. To address these challenges, we propose LR-BPFL, a novel BPFL method that learns a global deterministic model along with personalized low-rank Bayesian corrections. To tailor the local model to each client’s inherent uncertainty level, LR-BPFL incorporates an adaptive rank selection mechanism. We evaluate LR-BPFL across a variety of datasets, demonstrating its advantages in terms of calibration, accuracy, as well as computational and memory requirements. The code is available at \url{https://github.com/Bernie0115/LR-BPFL.}
Dongzhu Liu, Osvaldo Simeone, Guanchu Wang, Dimitrios P. Pezaros, Guangxu Zhu
AISTATS6
2025 Label Anything: An Interpretable, High-Fidelity and Prompt-Free Annotator
abstract
Learning-based street scene semantic understanding in autonomous driving (AD) has advanced significantly recently, but the performance of the AD model is heavily dependent on the quantity and quality of the annotated training data. However, traditional manual labeling involves high cost to annotate the vast amount of required data for training robust model. To mitigate this cost of manual labeling, we propose a Label Anything Model (denoted as LAM), serving as an interpretable, high-fidelity, and prompt-free data annotator. Specifically, we firstly incorporate a pretrained Vision Transformer (ViT) to extract the latent features. On top of ViT, we propose a semantic class adapter (SCA) and an optimization-oriented unrolling algorithm (OptOU), both with a quite small number of trainable parameters. SCA is proposed to fuse ViT-extracted features to consolidate the basis of the subsequent automatic annotation. OptOU consists of multiple cascading layers and each layer contains an optimization formulation to align its output with the ground truth as closely as possible, though which OptOU acts as being interpretable rather than learning-based blackbox nature. In addition, training SCA and OptOU requires only a single pre-annotated RGB seed image, owing to their small volume of learnable parameters. Extensive experiments clearly demonstrate that the proposed LAM can generate high-fidelity annotations (almost 100% in mIoU) for multiple real-world datasets (i.e., Camvid, Cityscapes, and Apolloscapes) and CARLA simulation dataset.
Wei-Bin Kou, Guangxu Zhu, Rongguang Ye, Shuai Wang 0004, Ming Tang 0006, Yik-Chung Wu
ICRA2
2025 Opportunistic Collaborative Planning with Large Vision Model Guided Control and Joint Query-Service Optimization
abstract
Navigating autonomous vehicles in open scenarios is a challenge due to the difficulties in handling unseen objects. Existing solutions either rely on small models that struggle with generalization or large models that are resource-intensive. While collaboration between the two offers a promising solution, the key challenge is deciding when and how to engage the large model. To address this issue, this paper proposes opportunistic collaborative planning (OCP), which seamlessly integrates efficient local models with powerful cloud models through two key innovations. First, we propose large vision model guided model predictive control (LVM-MPC), which leverages the cloud for LVM perception and decision making. The cloud output serves as a global guidance for a local MPC, thereby forming a closed-loop perception-to-control system. Second, to determine the best timing for large model query and service, we propose collaboration timing optimization (CTO), including object detection confidence thresholding (ODCT) and cloud forward simulation (CFS), to decide when to seek cloud assistance and when to offer cloud service. Extensive experiments show that the proposed OCP outperforms existing methods in terms of both navigation time and success rate.
Shuai Wang 0004, Wei Xu 0001, Guangxu Zhu, Derrick Wing Kwan Ng, Cheng-Zhong Xu 0001
IROS5
2025 Enhancing Large Vision Model in Street Scene Semantic Understanding through Leveraging Posterior Optimization Trajectory
abstract
To improve the generalization of the autonomous driving (AD) perception model, vehicles need to update the model over time based on the continuously collected data. As time progresses, the amount of data fitted by the AD model expands, which helps to improve the AD model generalization substantially. However, such ever-expanding data is a double-edged sword for the AD model. Specifically, as the fitted data volume grows to exceed the AD model’s fitting capacities, the AD model is prone to under-fitting. To address this issue, we propose to use a pretrained Large Vision Models (LVMs) as backbone coupled with downstream perception head to understand AD semantic information. This design can not only surmount the aforementioned under-fitting problem due to LVMs’ powerful fitting capabilities, but also enhance the perception generalization thanks to LVMs’ vast and diverse training data. On the other hand, to mitigate vehicles’ computational burden of training the perception head while running LVM backbone, we introduce a Posterior Optimization Trajectory (POT)-Guided optimization scheme (POTGui) to accelerate the convergence. Concretely, we propose a POT Generator (POTGen) to generate posterior (future) optimization direction in advance to guide the current optimization iteration, through which the model can generally converge within 10 epochs. Extensive experiments demonstrate that the proposed method improves the performance by over 66.48% and converges faster over 6 times, compared to the existing state-of-the-art approaches.
Wei-Bin Kou, Qingfeng Lin, Ming Tang 0006, Jingreng Lei, Shuai Wang 0004, Rongguang Ye, Guangxu Zhu, Yik-Chung Wu
IROS7
2025 FedEMA: Federated Exponential Moving Averaging with Negative Entropy Regularizer in Autonomous Driving
abstract
Street Scene Semantic Understanding (denoted as S3U) is a crucial but complex task for autonomous driving (AD) vehicles. Their inference models typically face poor generalization due to domain-shift. Federated Learning (FL) has emerged as a promising paradigm for enhancing the generalization of AD models through privacy-preserving distributed learning. However, these FL AD models face significant temporal catastrophic forgetting when deployed in dynamically evolving environments, where continuous adaptation causes abrupt erosion of historical knowledge. This paper proposes Federated Exponential Moving Average (FedEMA), a novel framework that addresses this challenge through two integral innovations: (I) Server-side model’s historical fitting capability preservation via fusing current FL round’s aggregation model and a proposed previous FL round’s exponential moving average (EMA) model; (II) Vehicle-side negative entropy regularization to prevent FL models’ possible overfitting to EMA-introduced temporal patterns. Above two strategies empower FedEMA a dual-objective optimization that balances model generalization and adaptability. In addition, we conduct theoretical convergence analysis for the proposed FedEMA. Extensive experiments both on Cityscapes dataset and Camvid dataset demonstrate FedEMA’s superiority over existing approaches, showing 7.12% higher mean Intersectionover-Union (mIoU).
Wei-Bin Kou, Guangxu Zhu, Bingyang Cheng, Shuai Wang 0004, Ming Tang 0006, Yik-Chung Wu
IROS2
2025 Quantized Analog Beamforming Enabled Multi-task Federated Learning Over-the-air
Jiacheng Yao, Wei Xu 0001, Guangxu Zhu, Zhaohui Yang 0001, Kaibin Huang, Dusit Niyato
VTC2025-Spring3
2025 Energy Efficient Data Processing: Integrated Sensing-Communication-Computation Design
abstract
In space-air-ground-sea networks, the conventional data processing designs separately considering sensing, communication and computation processes lead to severe wastes of radio, energy, and computation resources. To overcome this drawback, an integrated sensing-communication-computation design is pro-posed in this paper, which aims at realizing energy efficient data processing by jointly determining the data offloading ratio together with the sensing and offloading rates according to the processor profiles of mobile devices and servers. It is proved that the data offloading ratio is determined by the server's processor profile, while the string-pulling algorithms are designed to obtain the optimal sensing and offloading rates. Simulations are conducted to verify the effectiveness of the proposed design.
Ziqin Zhou, Xiaoyang Li 0002, Guangxu Zhu, Bingpeng Zhou, Chang Liu 0008, Kaibin Huang
VTC2025-Spring3
2025 Task-Oriented Wireless Communication and Control Co-Design
abstract
Driven by the rapid development of industrial Internet of Things applications, the wireless networked control system (WNCS) is expected to support real-time control-communication interaction performed in finite-time, which is task-oriented. A WNCS composed of multiple wirelessly inter-connected subsystems (SSs) is considered in this paper. The sensed state information in each SS is transmitted to the controller via wireless links for decision-and-control tasks. After multiple operation periods of state sensing and trans-mission, the system identification (SI) is executed and the optimal control (OC) policy is made. The SI requirement for OC is analyzed via system-level synthesis (SLS) based on robust control theory. A communication and control co-design is investigated, aiming to improve the energy efficiency while guaranteeing the SI performance requirement within the allowed decision-making time. The transmit powers at each sensor and controller, transmission interval length as well as the number of operation periods are jointly optimized. Simulations are conducted to validate the performance of the proposed co-design.
Xiaoyang Li 0002, Guangxu Zhu, Bingpeng Zhou, Kaibin Huang, Yi Gong 0001, Qinyu Zhang 0001
WCNC3
2025 RadioGAT: A Model-Based Learning Framework for Radio Map Reconstruction via Graph Attention Networks
abstract
Reconstructing accurate radio maps is crucial for optimizing wireless network performance and managing spectrum efficiently. In real-world scenarios, radio map data, often sparse and incompletely labelled, poses significant challenges to traditional learning techniques. Graph Neural Networks (GNNs) have become instrumental in efficiently reconstructing radio maps (RMR) in such environments by effectively encoding correlations in unstructured data. Existing GNN-based methods, however, are limited as they typically consider only single factors like location, environment, or transmitter characteristics during correlation encoding. To overcome this limitation, we introduce RadioGAT, a propagation model-based approach that comprehensively integrates these factors. We further utilize Graph Attention Networks to enable semi-supervised learning, enhancing the accuracy of radio map reconstruction. Our experimental results demonstrate the superiority and robustness of RadioGAT, particularly at low sampling rates, and highlight the importance of selecting appropriate correlation encoding methods based on the data availability for RMR.
Hang Li 0003, Xiaoyang Li 0002, Guangxu Zhu, Nan Qi 0001, Ming Xiao 0001
WCNC4
2025 TinyFEL: Communication, Computation, and Memory Efficient Tiny Federated Edge Learning via Model Sparse Update
abstract
Federated edge learning (FEL) is regarded as a promising distributed machine learning paradigm to reduce transmission latency and resources as well as preserve raw data privacy by collaboratively training local deep learning models across multiple edge devices. However, with the development of artificial intelligence (AI) technologies, the size of neural network models grows exponentially with their parameters to meet variable application requirements, which poses significant challenges to the computation, communication, and memory abilities of edge devices. Existing designs typically focus on either communication or computation efficiency without caring each device’s memory ability. To deal with the above issues, we first introduce a novel model sparse update enabled tiny FEL (TinyFEL) architecture, which terminates the backpropagation early in local model training processes. Therefore, the proposed TinyFEL can reduce local memory occupation and lessen the communication-and-computation burden. Furthermore, we propose a parameter splitting mechanism instead of transmitting the full model, only a part of updated layers of parameters is transmitted for aggregation, which significantly reduced the communication overheads. Thereafter, we develop a communication and computation latency minimization problem to accelerate the training of TinyFEL. To this end, we theoretically analyze the convergence performance of TinyFEL, which unveils the mathematical relationship among sparse update ratio assignment, device selection, and learning performance. Then, a joint sparse update ratio assignment, device selection, and resource allocation strategy is introduced based on the alternating direction method of multipliers (ADMMs) and block coordinate descent (BCD) algorithms. Numerical results indicate that our proposed TinyFEL can reduce training memory occupation by over 40% than the traditional FEL at the cost of negligible accuracy loss.
Qimei Chen, Yipeng Liang, Guangxu Zhu, Hao Jiang 0010
IEEE Internet Things J.4
2025 Sensing-Communication-Computation Integration for Federated Edge Learning With Controllable Model Dropout
abstract
Federated edge learning (FEEL) is an advanced paradigm in edge artificial intelligence, enabling privacy-preserving collaborative model training through periodic communication between edge devices and a central server. FEEL involves three key processes: 1) sensing; 2) computation; and 3) communication for data acquisition, processing, and exchange, respectively. Due to limited system resources, optimizing each process individually may lead to suboptimal learning performance. This challenge has sparked research into integrated sensing-computation–communication (ISCC) design for enhanced FEEL. While previous work has optimized general learning parameters, such as batch size and computing frequency, there is a lack of customized designs considering the neural network architecture as an optimizable variable in ISCC for FEEL. To close this gap, we introduce a novel design where each device generates a submodel through controllable weight dropout, adding flexibility by directly manipulating the learning process and reducing computation and communication overhead. To guide ISCC resource allocation in this new setting, we present a comprehensive convergence analysis, revealing the tight coupling of sensing, computation, and communication across devices and their impact on FEEL convergence. Building on these theoretical insights, we formulate an ISCC problem aiming to maximize the FEEL convergence rate through joint optimization of variables, such as batch size, sensing power, dropout rate, and communication power. This nonconvex problem is decomposed into two subproblems via alternating optimization: one controls batch size using a sorting algorithm, while the other focuses on ISCC device parameters, transformable into a convex problem solved by successive convex approximation. Extensive experiments using human motion recognition datasets demonstrate the superiority of the proposed design over baseline schemes.
Xiang Jiao, Guangxu Zhu, Wei Jiang 0003, Li Chen 0015, Wu Luo, Dingzhu Wen
IEEE Internet Things J.2
2025 CrossFi: A Cross Domain Wi-Fi Sensing Framework Based on Siamese Network
abstract
In recent years, Wi-Fi sensing has garnered significant attention due to its numerous benefits, such as privacy protection, low cost, and penetration ability. Extensive research has been conducted in this field, focusing on areas, such as gesture recognition, people identification, and fall detection. However, many data-driven methods encounter challenges related to domain shift, where the model fails to perform well in environments different from the training data. One major factor contributing to this issue is the limited availability of Wi-Fi sensing datasets, which makes models learn excessive irrelevant information and over-fit to the training set. Unfortunately, collecting large-scale Wi-Fi sensing datasets across diverse scenarios is a challenging task. To address this problem, we propose CrossFi, a siamese network-based approach that excels in both in-domain scenario and cross-domain scenario, including few-shot, zero-shot scenarios, and even works in few-shot new-class scenario where testing set contains new categories. The core component of CrossFi is a sample-similarity calculation network called CSi-Net, which improves the structure of the siamese network by using an attention mechanism to capture similarity information, instead of simply calculating the distance or cosine similarity. Based on it, we develop an extra Weight-Net that can generate a template for each class, so that our CrossFi can work in different scenarios. Experimental results demonstrate that our CrossFi achieves state-of-the-art performance across various scenarios. In gesture recognition task, our CrossFi achieves an accuracy of 98.17% in in-domain scenario, 91.72% in one-shot cross-domain scenario, 64.81% in zero-shot cross-domain scenario, and 84.75% in one-shot new-class scenario. The code for our model is publicly available athttps://github.com/RS2002/CrossFi.
Zijian Zhao 0002, Zhijie Cai, Xiaoyang Li 0002, Hang Li 0003, Qimei Chen, Guangxu Zhu
IEEE Internet Things J.7
2025 Exploiting Beam Split Effect on Wideband Beam Alignment: A Deep Unfolding Based Posterior Matching Approach
abstract
The massive-antenna wideband millimeter wave (mmWave)/terahertz (THz) systems inevitably suffer from a severe beam split effect due to the non-negligible signal propagation delays, which dramatically reduces communication efficiency. Nevertheless, if the wideband split effect is properly utilized, it can also bring benefits via sensing split directions for channel training. Hence, this paper proposes a novel wideband beam alignment framework with true-time-delayer (TTD) modules, which can fully exploit the controllable split beams for efficient angle-of-arrivals (AoAs) estimation. Moreover, we develop a hierarchical posterior matching (PM) enabled wideband beam alignment approach, which proactively configures the split beams to accelerate the estimation of AoAs posterior probability distributions. To deal with the computational complexity of the predesigned codebook and the insensitivity of the Gaussian distribution assumption in PM, we further introduce a low-complex and high-flexible wideband beam alignment approach based on a deep unfolding mechanism. Numerical results verify that: 1) The proposed framework can significantly improve the AoAs estimation accuracy at the cost of the same pilot overheads. 2) The proposed low-complexity deep unfolding approach outperforms the conventional PM mechanism even in low signal-to-noise-ratio (SNR) scenarios.
Qimei Chen, Xiaoxia Xu 0002, Guangxu Zhu, Hao Jiang 0010
IEEE J. Sel. Areas Commun.4
2025 Energy-Efficient Edge Inference in Integrated Sensing, Communication, and Computation Networks
abstract
Task-oriented integrated sensing, communication, and computation (ISCC) is a key technology for achieving low-latency edge inference and enabling efficient implementation of artificial intelligence (AI) in industrial cyber-physical systems (ICPS). However, the constrained energy supply at edge devices has emerged as a critical bottleneck. In this paper, we propose a novel energy-efficient ISCC framework for AI inference at resource-constrained edge devices, where adjustable split inference, model pruning, and feature quantization are jointly designed to adapt to diverse task requirements. A joint resource allocation design problem for the proposed ISCC framework is formulated to minimize the energy consumption under stringent inference accuracy and latency constraints. To address the challenge of characterizing inference accuracy, we derive an explicit approximation for it by analyzing the impact of sensing, communication, and computation processes on the inference performance. Building upon the analytical results, we propose an iterative algorithm employing alternating optimization to solve the resource allocation problem. In each subproblem, the optimal solutions are available by respectively applying a golden section search method and checking the Karush-Kuhn-Tucker (KKT) conditions, thereby ensuring the convergence to a local optimum of the original problem. Numerical results demonstrate the effectiveness of the proposed ISCC design, showing a significant reduction in energy consumption of up to 40% compared to existing methods, particularly in low-latency scenarios.
Jiacheng Yao, Wei Xu 0001, Guangxu Zhu, Kaibin Huang, Shuguang Cui
IEEE J. Sel. Areas Commun.3
2025 Beamforming Design for Semantic-Bit Coexisting Communication System
abstract
Semantic communication (SemCom) is emerging as a key technology for future sixth-generation (6G) systems. Unlike traditional bit-level communication (BitCom), SemCom directly optimizes performance at the semantic level, leading to superior communication efficiency. Nevertheless, the task-oriented nature of SemCom renders it challenging to completely replace BitCom. Consequently, it is desired to consider a semantic-bit coexisting communication system, where a base station (BS) serves SemCom users (sem-users) and BitCom users (bit-users) simultaneously. Such a system faces severe and heterogeneous inter-user interference. In this context, this paper provides a new semantic-bit coexisting communication framework and proposes a spatial beamforming scheme to accommodate both types of users. Specifically, we consider maximizing the semantic rate for semantic users while ensuring the quality-of-service (QoS) requirements for bit-users. Due to the intractability of obtaining the exact closed-form expression of the semantic rate, a data driven method is first applied to attain an approximated expression via data fitting. With the resulting complex transcendental function, majorization minimization (MM) is adopted to convert the original formulated problem into a multiple-ratio problem, which allows fractional programming (FP) to be used to further transform the problem into an inhomogeneous quadratically constrained quadratic programs (QCQP) problem. Solving the problem leads to a semi-closed form solution with undetermined Lagrangian factors that can be updated by a fixed point algorithm. This method is referred to as the MM-FP algorithm. Additionally, inspired by the semi-closed form solution, we also propose a low-complexity version of the MM-FP algorithm, called the low-complexity MM-FP (LP-MM-FP), which alleviates the need for iterative optimization of beamforming vectors. Extensive simulation results demonstrate that the proposed MM-FP algorithm outperforms conventional beamforming algorithms such as zero-forcing (ZF), maximum ratio transmission (MRT), and weighted minimum mean-square error (WMMSE). Moreover, the proposed LP-MMFP algorithm achieves comparable performance with the WMMSE algorithm but with lower computational complexity.
Maojun Zhang, Guangxu Zhu, Richeng Jin, Xiaoming Chen 0001, Qingjiang Shi, Caijun Zhong, Kaibin Huang
IEEE J. Sel. Areas Commun.2
2025 pFedLVM: A Large Vision Model (LVM)-Driven and Latent Feature-Based Personalized Federated Learning Framework in Autonomous Driving
abstract
Deep learning-based Autonomous Driving (AD) perception models often exhibit poor generalization due to data heterogeneity in an ever domain-shifting environment. While Federated Learning (FL) could improve the generalization of an AD model (known as FedAD system), conventional models often struggle with under-fitting as the amount of accumulated training data progressively increases. To address this issue, instead of conventional small models, employing Large Vision Models (LVMs) in FedAD is a viable option for better learning of representations from a vast volume of data. However, implementing LVMs in FedAD introduces three challenges:(I)the extremely high communication overheads associated with transmitting LVMs between participating vehicles and a central server;(II)lack of computing resource to deploy LVMs on each vehicle;(III)the performance drop due to LVM focusing on shared features but overlooking local vehicle characteristics. To overcome these challenges, we propose pFedLVM, a LVM-Driven, Latent Feature-Based Personalized Federated Learning framework. In this approach, the LVM is deployed only on central server, which effectively alleviates the computational burden on individual vehicles. Furthermore, the exchange between central server and vehicles are the learned features rather than the LVM parameters, which significantly reduces communication overhead. In addition, we utilize both shared features from all participating vehicles and individual characteristics from each vehicle to establish a personalized learning mechanism. This enables each vehicle’s model to learn features from others while preserving its personalized characteristics, thereby outperforming globally shared models trained in general FL. As a demonstration of the proposed pFedLVM, this paper focuses on the semantic segmentation (SSeg) task. Extensive experiments demonstrate that pFedLVM outperforms the existing state-of-the-art approach by 18.47%, 25.60%, 51.03% and 14.19% in terms of mIoU, mF1, mPrecision and mRecall, respectively.
Wei-Bin Kou, Qingfeng Lin, Ming Tang 0006, Rongguang Ye, Yang Leng, Shuai Wang 0004, Guofa Li, Zhenyu Chen 0001, Guangxu Zhu, Yik-Chung Wu
IEEE Trans. Intell. Transp. Syst.10
2025 Fast-Convergent and Communication-Alleviated Heterogeneous Hierarchical Federated Learning in Autonomous Driving
abstract
Street Scene Semantic Understanding (denoted as TriSU) is a complex task for autonomous driving (AD). However, inference model trained from data in a particular geographical region faces poor generalization when applied in other regions due to inter-city data domain-shift. Hierarchical Federated Learning (HFL) offers a potential solution for improving TriSU model generalization by collaborative privacy-preserving training over distributed datasets from different cities. Unfortunately, it suffers from slow convergence because the data from different cities are with disparate statistical properties. Going beyond existing HFL methods, we propose a Gaussian heterogeneous HFL algorithm (FedGau) to address inter-city data heterogeneity so that convergence can be accelerated. In the proposed FedGau algorithm, both single RGB image and RGB dataset are modelled as Gaussian distributions for aggregation weight design. This approach not only differentiates each RGB image by respective statistical distribution, but also exploits the statistics of dataset from each city in addition to the conventionally considered data volume. With the proposed approach, the convergence is accelerated by 35.5%-40.6% compared to existing state-of-the-art (SOTA) HFL methods. On the other hand, to reduce the involved communication resource, we further introduce a novel performance-aware adaptive resource scheduling (AdapRS) policy. Unlike the traditional static resource scheduling policy that exchanges a fixed number of models between two adjacent aggregations, AdapRS adjusts the number of model aggregation at different levels of HFL so that unnecessary communications are minimized. Extensive experiments demonstrate that AdapRS saves 29.65% communication overhead compared to conventional static resource scheduling policy while maintaining almost the same performance.
Wei-Bin Kou, Qingfeng Lin, Ming Tang 0006, Rongguang Ye, Shuai Wang 0004, Guangxu Zhu, Yik-Chung Wu
IEEE Trans. Intell. Transp. Syst.6
2025 Clustered Federated Multi-Task Learning: A Communication-and-Computation Efficient Sparse Sharing Approach
abstract
Federated multi-task learning (FMTL) is a promising technology to tackle one of the most severe non-independent and identically distributed (non-IID) data challenge in federated learning (FL), which treats each client as a single task and learns personalized models by exploiting task correlations. However, the transmission of individual task models generally results in a significant amount of communication overhead compared with global model broadcasting. Furthermore, related works mainly focus on FMTLs with default and static relationships among tasks, which obliterates the non-IID data characteristic. To address these issues, we propose a novel Clustered FMTL mechanism via Sparse Sharing (FedSS). Specifically, we introduce an iterative model pruning approach that trains customized client models to deal with the non-IID issue. Thereafter, we divide clients into different tasks according to their model similarities to promote communication efficiency. Based on clustered tasks, we introduce a sparse sharing mechanism that allows clients to share model parameters dynamically among different tasks to further boost the training performance. On the other aspect, the infertile communication resources would degrade the FMTL performance by restricting the personalized model transmissions. Hence, we first theoretically analyze the convergence performance of the proposed FedSS, which quantitatively unveils the relationship between the local model training performance and communication resources. Thereafter, we formulate a communication-and-computation efficient optimization problem via a joint sparsity ratio assignment and bandwidth allocation strategy. Closed-form expressions for the optimal sparsity ratio and bandwidth allocation are derived based on Lyapunov optimization and block coordinate update (BCU) algorithms. Numerical results illustrate that the proposed FedSS outperforms the benchmarks, and achieves an efficient communication and computation performance.
Yuhan Ai, Qimei Chen, Guangxu Zhu, Dingzhu Wen, Hao Jiang 0010
IEEE Trans. Wirel. Commun.3
2025 Network-Level Performance Analysis for Air-Ground Integrated Sensing and Communication
abstract
To support the development of air-ground integrated sensing and communication (ISAC), network-level performance analysis is needed for providing an essential guide on the network design. Following the widely adopted orthogonal frequency-division multiplexing (OFDM) technology in existing wireless systems, a cooperative air-ground wireless network based on OFDM-ISAC is introduced in this paper, where the ISAC-enabled base stations (BSs) following the two-dimensional homogeneous Poisson point process (HPPP) distribution serve the terrestrial communication users while sensing the aerial targets. In particular, cooperative beamforming schemes are designed for mitigating the interference among ISAC BSs. First, we analyze the communication as well as sensing performances in terms of different metrics including area communication coverage probability, area communication spectral efficiency, area radar detection coverage probability, and average Cramér-Rao Bound. Simulation results are then presented to validate the theoretical analysis and illustrate the effects of key system parameters on the network performance. It is observed that both the communication and sensing (C&S) performances depend on the BS density and height, while the sensing performance also depends on the height of sensing target together with the numbers of OFDM subcarriers and symbols. Moreover, there exists a tradeoff between the C&S performances with respect to the BS density and height. The results of this paper provide useful guidance to the design and implementation of air-ground wireless network for harnessing the dual benefits of ISAC.
Yihang Jiang 0001, Xiaoyang Li 0002, Guangxu Zhu, Kaifeng Han, Kaitao Meng, Chenji Liu, Qingjiang Shi, Rui Zhang 0006
IEEE Trans. Wirel. Commun.3
2025 Communication-and-Energy Efficient Over-the-Air Federated Learning
abstract
Communication and energy efficiencies are two crucial objectives in the pursuit of edge intelligence in 6G networks, and become increasingly important given the prevalence of large model training. Existing designs typically focus on either communication efficiency or energy efficiency due to the fact that improving one objective generally comes at the expense of the other. Over-the-air federated learning (OTA-FL) has recently emerged as a promising approach to enhance both efficiencies through an integrated communication and computation design. Nevertheless, most previous studies on OTA-FL only consider scenarios where the dataset for the entire FL procedure is collected and available prior to training. In real-world applications, devices continuously collect new data in an online manner. This underscores the significance of sample collection through sensing in a practical FL pipeline. We propose to integrate sensing with communication and computation into a joint design to further boost the communication-and-energy efficiencies of OTA-FL. Specifically, we consider a training latency and energy consumption minimization problem with performance guarantees. To this end, we first derive an average training error (ATE) metric to quantify convergence performance. Then, a joint sensing, communication and computation resource allocation strategy is developed based on a deep reinforcement learning (DRL) algorithm that nests convex optimization with a deep Q-network. Extensive experiments are conducted to validate our theoretical analysis, and demonstrate the effectiveness of the proposed design for communication-and-energy efficient FL.
Yipeng Liang, Qimei Chen, Guangxu Zhu, Hao Jiang 0010, Yonina C. Eldar, Shuguang Cui
IEEE Trans. Wirel. Commun.3
2025 Rethinking Resource Management in Edge Learning: A Joint Pre-Training and Fine-Tuning Design Paradigm
abstract
In some applications, edge learning is experiencing a shift in focus from conventional learning from scratch to two-stage learning combining pre-training and task-specific fine-tuning. This paper considers the problem of joint communication and computation resource management in a two-stage edge learning system. In this system, model pre-training is first conducted at an edge server via centralized learning on local pre-stored general data, and then task-specific fine-tuning is performed at edge devices based on the pre-trained model via federated edge learning. For the two-stage learning model, we first analyze the convergence behavior (in terms of the average squared gradient norm bound), which characterizes the impacts of various system parameters, such as the number of learning rounds and batch sizes in the two stages, on the convergence rate. Based on our analytical results, we then propose a joint communication and computation resource management design to minimize an average squared gradient norm bound, subject to constraints on the transmit power, overall system energy consumption, and training delay. The decision variables include the number of learning rounds, batch sizes, clock frequencies, and transmit power control for both pre-training and fine-tuning stages. Finally, numerical results are provided to evaluate the effectiveness of our proposed design. It is shown that the proposed joint resource management over the pre-training and fine-tuning stages well balances the system performance trade-off among the training accuracy, delay, and energy consumption. The proposed design is also shown to effectively leverage the inherent trade-off between pre-training and fine-tuning, which arises from the differences in data distribution between pre-stored general data versus real-time task-specific data, thus efficiently optimizing overall system performance.
Zhonghao Lyu, Yuchen Li 0006, Guangxu Zhu, Jie Xu 0002, H. Vincent Poor, Shuguang Cui
IEEE Trans. Wirel. Commun.3
2025 Integrated Sensing, Computation, and Communication for UAV-Assisted Federated Edge Learning
abstract
Federated edge learning (FEEL) enables privacy-preserving model training through periodic communication between edge devices and the server. Unmanned Aerial Vehicle (UAV)-mounted edge devices are particularly advantageous for FEEL due to their flexibility and mobility in efficient data collection. In UAV-assisted FEEL, sensing, computation, and communication are coupled and compete for limited onboard resources, and UAV deployment also affects sensing and communication performance. Therefore, the joint design of UAV deployment and resource allocation is crucial to achieving the optimal training performance. In this paper, we address the problem of joint UAV deployment design and resource allocation for FEEL via a concrete case study of human motion recognition based on wireless sensing. We first analyze the impact of UAV deployment on the sensing quality and identify a threshold value for the sensing elevation angle that guarantees a satisfactory quality of data samples. Due to the non-ideal sensing channels, we consider the probabilistic sensing model, where the successful sensing probability of each UAV is determined by its position. Then, we derive the upper bound of the FEEL training loss as a function of the sensing probability. Theoretical results suggest that the convergence rate can be improved if UAVs have a uniform successful sensing probability. Based on this analysis, we formulate a training time minimization problem by jointly optimizing UAV deployment, integrated sensing, computation, and communication (ISCC) resources under a desirable optimality gap constraint. To solve this challenging mixed-integer non-convex problem, we apply the alternating optimization technique, and propose the bandwidth, batch size, and position optimization (BBPO) scheme to optimize these three decision variables alternately. Simulation results demonstrate that our BBPO scheme outperforms other baseline schemes regarding convergence rate and testing accuracy. The simulation implementation is available at https://github.com/TheaSherlock/ISCC-UAV.
Guangxu Zhu, Wei Xu 0001, Man Hon Cheung, Tat-Ming Lok, Shuguang Cui
IEEE Trans. Wirel. Commun.2
2025 Multipath Information Fusion-Boosted Vehicle State Detection, Reflector Positioning, and Channel Estimation for 6G ISAC Systems
abstract
We are interested in multiple-input-multiple-output (MIMO) orthogonal frequency-division multiplexing (OFDM) communication-based vehicle state detection (VSD) (including vehicle location, velocity and pose angle) in multipath interference scenarios, towards 6G integrated sensing and communications. Yet, communication-based VSD is challenging, since its signals undergo multipath interference and random fading, while channel state and reflector locations are even unknown, in addition to vehicle state. To address these challenges, a novel multipath information fusion-assisted VSD scheme is devised to smartly aggregate geometric knowledge from both direct and reflection paths, thus yielding a robust VSD solution against multipath interference. In addition, we propose to divide the complex VSD problem into four subproblems: (i) angle-of-arrival detection, (ii) time-of-flight estimation, (iii) joint reconstruction of angle-of-departure, radial speed and channel state, and (iv) vehicle-and-reflector state detection. An efficient four-step cascaded VSD method is devised by exploiting linearity, quadratic, orthogonality and space-time-domain correlation of MIMO OFDM signals, which finally achieves simultaneous VSD, reflector positioning and channel estimate. It is verified by simulations that our VSD scheme outperforms state-of-the-art baselines due to our specially-tailored problem decoupling and multipath information fusion, which builds a technical foundation for designing environment sensing-assisted communication strategies.
Bingpeng Zhou, Hanglong Chen, Guangxu Zhu, Yue Xiao 0001, Qingjiang Shi
IEEE Trans. Wirel. Commun.4
2024 Integrating Sensing, Communication, and Computation in the Sky
abstract
Unmanned Aerial Vehicle (UAV)-mounted edge devices are particularly advantageous for federated edge learning (FEEL) due to their flexibility and mobility in efficient data collection. In UAV-assisted FEEL, sensing, computation, and communication are coupled and compete for limited onboard resources, and UAV deployment also affects sensing and communication performance. Therefore, the joint design of UAV deployment and resource allocation is crucial to achieving the optimal training performance. In this paper, we address the problem of joint UAV deployment design and resource allocation for FEEL via a concrete case study of human motion recognition based on wireless sensing. Due to the nonideal sensing channels, we consider the probabilistic sensing model. Then, we derive the upper bound of the FEEL training loss as a function of the sensing probability. We formulate a training time minimization problem by jointly optimizing UAV deployment, integrated sensing, computation, and communication (ISCC) resources under a desirable optimality gap constraint. To solve this challenging mixed-integer non-convex problem, we propose our algorithm based on the alternating optimization technique. Simulation results demonstrate that our algorithm outperforms other baselines regarding convergence rate and testing accuracy.
Guangxu Zhu, Wei Xu 0001, Man Hon Cheung, Tat-Ming Lok, Shuguang Cui
ICASSP2
2024 FedRC: A Rapid-Converged Hierarchical Federated Learning Framework in Street Scene Semantic Understanding
abstract
Street Scene Semantic Understanding (denoted as TriSU) is a crucial but complex task for world-wide distributed autonomous driving (AD) vehicles (e.g., Tesla). Its inference model faces poor generalization issue due to inter-city domain-shift. Hierarchical Federated Learning (HFL) offers a potential solution for improving TriSU model generalization, but suffers from slow convergence rate because of vehicles’ surrounding heterogeneity across cities. Going beyond existing HFL works that have deficient capabilities in complex tasks, we propose a rapid-converged heterogeneous HFL framework (FedRC) to address the inter-city data heterogeneity and accelerate HFL model convergence rate. In our proposed FedRC framework, both single RGB image and RGB dataset are modelled as Gaussian distributions in HFL aggregation weight design. This approach not only differentiates each RGB sample instead of typically equalizing them, but also considers both data volume and statistical properties rather than simply taking data quantity into consideration. Extensive experiments on the TriSU task using across-city datasets demonstrate that FedRC converges faster than the state-of-the-art benchmark by 38.7%, 37.5%, 35.5%, and 40.6% in terms of mIoU, mPrecision, mRecall, and mF1, respectively. Furthermore, qualitative evaluations in the CARLA simulation environment confirm that the proposed FedRC framework delivers top-tier performance.
Wei-Bin Kou, Qingfeng Lin, Ming Tang 0006, Shuai Wang 0004, Guangxu Zhu, Yik-Chung Wu
IROS5
2024 Joint Device Scheduling and Resource Allocation for ISCC-Based Multiview-Multitask Inference
abstract
This article investigates an integrated sensing-communication-computation (ISCC)-based multiview-multitask (MVMT) edge artificial intelligence inference system. Each device senses a narrow view of a target area and processes the echo signal to generate real-time sensory data. An edge server receives and combines multiple views of data from multiple devices to complete several downstream inference tasks. Compared with existing designs where dedicated sensory data are obtained, transmitted, and processed for each task, this ISCC-based MVMT framework enjoys reduced costs of sensing, on-device computation, and communication overhead due to data sharing among different tasks. The challenges of improving all tasks’ inference accuracy lie in the tight coupling of sensing, communication, and computation among different devices and sensory view competition among different tasks. These two challenges intertwine, making the multitask optimization problem mixed-integer nonconvex programming. To tackle this problem, we propose a joint device scheduling and resource allocation (JDSRA) scheme, which alternatively solves a subproblem of joint device scheduling and time allocation and a subproblem of resource allocation till convergence. Particularly, in addition to a dynamic-programming-based optimal device scheduling algorithm, a low-complexity suboptimal algorithm is proposed based on sorting a derived closed-form indicator, which represents the increase of all tasks’ inference accuracy per time unit consumption. Besides, a low-complexity optimal resource allocation algorithm is proposed by parallelly solving multiple simple convex subproblems. Numerical results based on jointly completing three tasks of human motion recognition, human height recognition, and localization in smart home scenarios are conducted to verify the performance of our proposed schemes.
Diao Wang, Dingzhu Wen, Yinghui He, Qimei Chen, Guangxu Zhu, Guanding Yu
IEEE Internet Things J.5
2024 Integrating Sensing, Communication, and Power Transfer: Multiuser Beamforming Design
abstract
In the sixth-generation (6G) networks, massive low-power devices are expected to sense environment and deliver tremendous data. To enhance the radio resource efficiency, the integrated sensing and communication (ISAC) technique exploits the sensing and communication functionalities of signals, while the simultaneous wireless information and power transfer (SWIPT) techniques utilizes the same signals as the carriers for both information and power delivery. The further combination of ISAC and SWIPT leads to the advanced technology namely integrated sensing, communication, and power transfer (ISCPT). In this paper, a multi-user multiple-input multiple-output (MIMO) ISCPT system is considered, where a base station equipped with multiple antennas transmits messages to multiple information receivers (IRs), transfers power to multiple energy receivers (ERs), and senses a target simultaneously. The sensing target can be regarded as a point or an extended surface. When the locations of IRs and ERs are separated, the MIMO beamforming designs are optimized to improve the sensing performance while meeting the communication and power transfer requirements. The resultant non-convex optimization problems are solved based on a series of techniques including Schur complement transformation and rank reduction. Moreover, when the IRs and ERs are co-located, the power splitting factors are jointly optimized together with the beamformers to balance the performance of communication and power transfer. To better understand the performance of ISCPT, the target positioning problem is further investigated. Simulations are conducted to verify the effectiveness of our proposed designs, which also reveal a performance tradeoff among sensing, communication, and power transfer.
Ziqin Zhou, Xiaoyang Li 0002, Guangxu Zhu, Jie Xu 0002, Kaibin Huang, Shuguang Cui
IEEE J. Sel. Areas Commun.3
2024 Collaborative Edge AI Inference Over Cloud-RAN
abstract
In this paper, a cloud radio access network (Cloud-RAN) based collaborative edge AI inference architecture is proposed. Specifically, geographically distributed devices capture real-time noise-corrupted sensory data samples and extract the noisy local feature vectors, which are then aggregated at each remote radio head (RRH) to suppress sensing noise. To realize efficient uplink feature aggregation, we allow each RRH receives local feature vectors from all devices over the same resource blocks simultaneously by leveraging an over-the-air computation (AirComp) technique. Thereafter, these aggregated feature vectors are quantized and transmitted to a central processor (CP) for further aggregation and downstream inference tasks. Our aim in this work is to maximize the inference accuracy via a surrogate accuracy metric called discriminant gain, which measures the discernibility of different classes in the feature space. The key challenges lie on simultaneously suppressing the coupled sensing noise, AirComp distortion caused by hostile wireless channels, and the quantization error resulting from the limited capacity of fronthaul links. To address these challenges, this work proposes a joint transmit precoding, receive beamforming, and quantization error control scheme to enhance the inference accuracy. Extensive numerical experiments demonstrate the effectiveness and superiority of our proposed optimization algorithm compared to various baselines.
Dingzhu Wen, Guangxu Zhu, Qimei Chen, Kaifeng Han, Yuanming Shi
IEEE Trans. Commun.3
2024 Energy-Efficient Optimal Mode Selection for Edge AI Inference via Integrated Sensing-Communication-Computation
abstract
Existing edge inference methods only consider one paradigm, i.e., one of on-device inference, on-server inference, or edge-device cooperative inference. Each paradigm has its pros and cons as well as dominant application scopes. For example, the on-device paradigm is the best choice when the inference task is not computationally intensive, the on-server paradigm is suitable if the communication capacity is strong, and the edge-device cooperative mode should be selected in the scenario of weak on-device communication and computation. However, each paradigm suffers from poor performance if deployed outside of its application scope, thus leading to limited potential and flexibility. This paper proposes an edge AI inference framework, which makes the first attempt to jointly consider the three modes for making full use of their benefits. In addition, sensing for data acquisition is enabled at both the edge server and the device. This can effectively improve the inference accuracy with rich information on the target area from two different views. On the other hand, energy cost minimization turns out to be a key target all over the world and a significant issue in wireless networks. To this end, we target minimizing the system energy cost under a given inference accuracy guarantee and other network resource constraints, by coordinating sensing, communication, and computation in different modes. By optimally solving the optimization problem, an integrated sensing-communication-computation (ISCC) based task-oriented mode selection scheme is proposed. A practical ISCC platform is built and extensive experiments are conducted to verify our theoretical analysis.
Dingzhu Wen, Qimei Chen, Guangxu Zhu, Yuanming Shi
IEEE Trans. Mob. Comput.5
2024 Joint Compression and Deadline Optimization for Wireless Federated Learning
abstract
Federated edge learning(FEEL) is a popular distributed learning framework for privacy-preserving at the edge, in which densely distributed edge devices periodically exchange model-updates with the server to complete the global model training. Due to limited bandwidth and uncertain wireless environment, FEEL may impose heavy burden to the current communication system. In addition, under the common FEEL framework, the server needs to wait for the slowest device to complete the update uploading before starting the aggregation process, leading to the straggler issue that causes prolonged communication time. In this paper, we propose to accelerate FEEL from two aspects: i.e., 1) performing data compression on the edge devices and 2) setting a deadline on the edge server to exclude the straggler devices. However, undesired gradient compression errors and transmission outage are introduced by the aforementioned operations respectively, affecting the convergence of FEEL as well. In view of these practical issues, we formulate a training time minimization problem, with the compression ratio and deadline to be optimized. To this end, an asymptotically unbiased aggregation scheme is first proposed to ensure zero optimality gap after convergence, and the impact of compression error and transmission outage on the overall training time are quantified through convergence analysis. Then, the formulated problem is solved in an alternating manner, based on which, the noveljoint compression and deadline optimization(JCDO) algorithm is derived. Numerical experiments for different use cases in FEEL including image classification and autonomous driving show that the proposed method is nearly 30X faster than the vanilla FedSGD algorithm, and outperforms the state-of-the-art schemes.
Maojun Zhang, Yang Li 0049, Dongzhu Liu, Richeng Jin, Guangxu Zhu, Caijun Zhong, Tony Q. S. Quek
IEEE Trans. Mob. Comput.5
2024 RadioGAT: A Joint Model-Based and Data-Driven Framework for Multi-Band Radiomap Reconstruction via Graph Attention Networks
abstract
Multi-band radiomap reconstruction (MB-RMR) is a key component in wireless communications for tasks such as spectrum management and network planning. However, traditional machine-learning-based MB-RMR methods, which rely heavily on simulated data or complete structured ground truth, face significant deployment challenges. These challenges stem from the differences between simulated and actual data, as well as the scarcity of real-world measurements. To address these challenges, our study presents RadioGAT, a novel framework based on Graph Attention Network (GAT) tailored for MB-RMR within a single area, eliminating the need for multi-region datasets. RadioGAT innovatively merges model-based spatial-spectral correlation encoding with data-driven radiomap generalization, thus minimizing the reliance on extensive data sources. The framework begins by transforming sparse multi-band data into a graph structure through an innovative encoding strategy that leverages radio propagation models to capture the spatial-spectral correlation inherent in the data. This graph-based representation not only simplifies data handling but also enables tailored label sampling during training, significantly enhancing the framework’s adaptability for deployment. Subsequently, The GAT is employed to generalize the radiomap information across various frequency bands. Extensive experiments using raytracing datasets based on real-world environments have demonstrated RadioGAT’s enhanced accuracy in supervised learning settings and its robustness in semi-supervised scenarios. These results underscore RadioGAT’s effectiveness and practicality for MB-RMR in environments with limited data availability.
Songyang Zhang 0002, Hang Li 0003, Xiaoyang Li 0002, Lexi Xu, Haigao Xu, Hui Mei, Guangxu Zhu, Nan Qi 0001, Ming Xiao 0001
IEEE Trans. Wirel. Commun.8
2024 Semantic Communications for Image Recovery and Classification via Deep Joint Source and Channel Coding
abstract
With the recent advancements in edge artificial intelligence (AI), future sixth-generation (6G) networks need to support new AI tasks such as classification and clustering apart from data recovery. Motivated by the success of deep learning, the semantic-aware and task-oriented communications with deep joint source and channel coding (JSCC) have emerged as new paradigm shifts in 6G from the conventional data-oriented communications with separate source and channel coding (SSCC). However, most existing works focused on the deep JSCC designs for one task of data recovery or AI task execution independently, which cannot be transferred to other unintended tasks. Differently, this paper investigates the JSCC semantic communications to support multi-task services, by performing the image data recovery and classification task execution simultaneously. First, we propose a new end-to-end deep JSCC framework by unifying the coding rate reduction maximization and the mean square error (MSE) minimization in the loss function. Here, the coding rate reduction maximization facilitates the learning of discriminative features for enabling to perform classification tasks directly in the feature space, and the MSE minimization helps the learning of informative features for high-quality image data recovery. Next, to further improve the robustness against variational wireless channels, we propose a new gated deep JSCC design, in which a gated net is incorporated for adaptively pruning the output features to adjust their dimensions based on channel conditions. Finally, we present extensive numerical experiments to validate the performance of our proposed deep JSCC designs as compared to various benchmark schemes. It is shown that our proposed designs simultaneously provide efficient multi-task services, and the proposed gated deep JSCC framework efficiently reduces the communication overhead with only marginal performance loss. It is also shown that performing the classification task on the feature space via coding rate reduction maximization is able to better defend the label corruption than the traditional label-fitting methods.
Zhonghao Lyu, Guangxu Zhu, Jie Xu 0002, Bo Ai 0001, Shuguang Cui
IEEE Trans. Wirel. Commun.2
2024 Task-Oriented Over-the-Air Computation for Multi-Device Edge AI
abstract
Edge inference refers to the use of artificial intelligent (AI) models at the network edge to provide mobile devices inference services and thereby enable intelligent services such as auto-driving and Metaverse towards 6G. However, departing from the classic paradigm of data-centric designs, the 6G networks for supporting edge AI features task-oriented techniques that focus on effective and efficient execution of AI task. Targeting end-to-end system performance, such techniques are sophisticated as they aim to seamlessly integrate sensing (data acquisition), communication (data transmission), and computation (data processing). Aligned with the paradigm shift, a task-oriented over-the-air computation (AirComp) scheme is proposed in this paper for multi-device split-inference system. In the considered system, local feature vectors, which are extracted from the real-time noisy sensory data on devices, are aggregated over-the-air by exploiting the waveform superposition in a multiuser channel. Then the aggregated features as received at a server are fed into an inference model with the result used for decision making or control of actuators. To design inference-oriented AirComp, the transmit precoders at edge devices and receive beamforming at edge server are jointly optimized to rein in the aggregation error and maximize the inference accuracy. The problem is made tractable by measuring the inference accuracy using a surrogate metric called discriminant gain, which measures the discernibility of two object classes in the application of object/event classification. It is discovered that the conventional AirComp beamforming design for minimizing the mean square error in generic AirComp with respect to the noiseless case may not lead to the optimal classification accuracy. The reason is due to the overlooking of the fact that feature dimensions have different sensitivity towards aggregation errors and are thus of different importance levels for classification. This issue is addressed in this work via a new task-oriented AirComp scheme designed by directly maximizing the derived discriminant gain. However, the resultant problem of joint transmit precoding and receive beamforming is nonconvex and difficult to solve due to the complicated form of discriminant gain and the coupling between the control variables. We overcome the difficulty using the successive convex approximation. The performance gain of the proposed task-oriented scheme over the conventional schemes is verified by extensive experiments targeting the application of human motion recognition.
Dingzhu Wen, Xiang Jiao, Peixi Liu, Guangxu Zhu, Yuanming Shi, Kaibin Huang
IEEE Trans. Wirel. Commun.4
2024 Task-Oriented Sensing, Computation, and Communication Integration for Multi-Device Edge AI
abstract
This paper studies a new multi-device edge artificial-intelligent (AI) system, which jointly exploits the AI model split inference and integrated sensing and communication (ISAC) to enable low-latency intelligent services at the network edge. In this system, multiple ISAC devices perform radar sensing to obtain multi-view data, and then offload the quantized version of extracted features to a centralized edge server, which conducts model inference based on the cascaded feature vectors. Under this setup and by considering classification tasks, we measure the inference accuracy by adopting an approximate but tractable metric, namely discriminant gain, which is defined as the distance of two classes in the Euclidean feature space under normalized covariance. To maximize the discriminant gain, we first quantify the influence of the sensing, computation, and communication processes on it with a derived closed-form expression. Then, an end-to-end task-oriented resource management approach is developed by integrating the three processes into a joint design. This integrated sensing, computation, and communication (ISCC) design approach, however, leads to a challenging non-convex optimization problem, due to the complicated form of discriminant gain and the device heterogeneity in terms of channel gain, quantization level, and generated feature subsets. Remarkably, the considered non-convex problem can be optimally solved based on the sum-of-ratios method. This gives the optimal ISCC scheme, that jointly determines the transmit power and time allocation at multiple devices for sensing and communication, as well as their quantization bits allocation for computation distortion control. By using human motions recognition as a concrete AI inference task, extensive experiments are conducted to verify the performance of our derived optimal ISCC scheme.
Dingzhu Wen, Peixi Liu, Guangxu Zhu, Yuanming Shi, Jie Xu 0002, Yonina C. Eldar, Shuguang Cui
IEEE Trans. Wirel. Commun.3
2024 Integrated Sensing-Communication-Computation for Over-the-Air Edge AI Inference
abstract
Edge-device co-inference refers to deploying well-trained artificial intelligent (AI) models at the network edge under the cooperation of devices and edge servers for providing ambient intelligent services. For enhancing the utilization of limited network resources in edge-device co-inference tasks from a systematic view, we propose a task-oriented scheme of integrated sensing, computation and communication (ISCC) in this work. In this system, all devices sense a target from the same wide view to obtain homogeneous noise-corrupted sensory data, from which the local feature vectors are extracted. All local feature vectors are aggregated at the server using over-the-air computation (AirComp) in a broadband channel with the orthogonal-frequency-division-multiplexing technique for suppressing the sensing and channel noise. The aggregated denoised global feature vector is further input to a server-side AI model for completing the downstream inference task. A novel task-oriented design criterion, called maximum minimum pair-wise discriminant gain, is adopted for classification tasks. It extends the distance of the closest class pair in the feature space, leading to a balanced and enhanced inference accuracy. Under this criterion, a problem of joint sensing power assignment, transmit precoding and receive beamforming is formulated. The challenge lies in three aspects: the coupling between sensing and AirComp, the joint optimization of all feature dimensions’ AirComp aggregation over a broadband channel, and the complicated form of the maximum minimum pair-wise discriminant gain. To solve this problem, a task-oriented ISCC scheme with AirComp is proposed. Experiments based on a human motion recognition task are conducted to verify the advantages of the proposed scheme over the existing scheme and a baseline.
Zeming Zhuang, Dingzhu Wen, Yuanming Shi, Guangxu Zhu, Sheng Wu 0001, Dusit Niyato
IEEE Trans. Wirel. Commun.4
2023 CHA-Sens: An End-to-End Comprehensive Residual Convolution Framework for CSI-based Human Activity Sensing
abstract
Channel state information (CSI)-based human activity recognition (HAR) receives increasing research interests due to its broad applications such as human-computer interaction, health care, and security surveillance. Deep learning (DL) methods have been widely adopted on CSI-based HAR tasks to extract features automatically, overcoming the complexity and unstableness of manual feature extraction process. However, many DL approaches fail to customize the designed model structure with the input CSI tensor shape, applying DL models recklessly. In addition, some researchers utilize attention mechanism yet independently along temporal, spatial, or frequency dimension. To address these issues, we propose an end-to-end comprehensive residual convolution framework, namely CHA-Sens, for general CSI-based human activity sensing. CHA-Sens consists of several comprehensive residual convolution modules (CRCM) that feature adaptive kernel size and stride, regional parameter-free attention mechanism and shortcut of identity mapping. Extensive experiments are conducted on three public CSI datasets for recognizing both single human activity (SHA) and human-to-human interaction (HHI) to show the superiority of the proposed design over other state-of-the-art benchmarks.
Fujia Zhou, Wei Zhang 0100, Guangxu Zhu, Hang Li 0003, Qingjiang Shi
CSCWD3
2023 Bayesian Over-the-Air FedAvg via Channel Driven Stochastic Gradient Langevin Dynamics
abstract
The recent development of scalable Bayesian inference methods has renewed interest in the adoption of Bayesian learning as an alternative to conventional frequentist learning that offers improved model calibration via uncertainty quantification. Recently, federated averaging Langevin dynamics (FALD) was introduced as a variant of federated averaging that can efficiently implement distributed Bayesian learning in the presence of noiseless communications. In this paper, we propose wireless FALD (WFALD), a novel protocol that realizes FALD in wireless systems by integrating over-the-air computation and channel-driven sampling for Monte Carlo updates. Unlike prior work on wireless Bayesian learning, WFALD enables (i) multiple local updates between communication rounds; and (ii) stochastic gradients computed by mini-batch. A convergence analysis is presented in terms of the 2- Wasserstein distance between the samples produced by WFALD and the targeted global posterior distribution. Analysis and experiments show that, when the signal-to-noise ratio is sufficiently large, channel noise can be fully repurposed for Monte Carlo sampling, thus entailing no loss in performance.
Dongzhu Liu, Osvaldo Simeone, Guangxu Zhu
GLOBECOM4
2023 Federated Edge Learning via Integrated Sensing, Computation, and Communication
abstract
Sensing, computation, and communication (SC2) are highly coupled processes in federated edge learning (FEEL) and need to be jointly designed in a task-oriented manner for pursuing the best FEEL performance under the stringent resource constraints at edge devices. However, this remains an open problem as there is a lack of theoretical understanding on how the SC2resources jointly affect the FEEL performance. In this paper, we address the problem of joint SC2resource allocation for FEEL via a concrete case study of human motion recognition based on wireless sensing. Specifically, the joint SC2resource allocation problem is cast to maximize the convergence speed of FEEL, under the constraints on training time and energy supply of each edge device. Solving this problem entails solving two subproblems in order: the first one reduces to determining a joint sensing and communication resource allocation that maximizes the total number of samples sensed during the entire training process; the second one concerns the partition of the total number of sensed samples over communication rounds to determine the batch size at each round for convergence speed maximization. Finally, extensive simulation results are provided to validate the superiority of the proposed scheme over several baseline schemes.
Peixi Liu, Guangxu Zhu, Shuai Wang 0004, Miaowen Wen, Wu Luo, H. Vincent Poor, Shuguang Cui
ICC2
2023 Task-Oriented Sensing, Computation, and Communication Integration for Multi-Device Edge AI
abstract
This paper studies a new multi-device edge artificial-intelligent (AI) system, which jointly exploits the AI model split inference and integrated sensing and communication (ISAC) to enable low-latency intelligent services at the network edge. In this system, multiple ISAC devices perform radar sensing to obtain multi-view data, and then offload the quantized version of extracted features to a centralized edge server, which conducts model inference based on the cascaded feature vectors. Under this setup and by considering classification tasks, we measure the inference accuracy by adopting an approximate but tractable metric, namely discriminant gain, which is defined as the distance of two classes in the Euclidean feature space under normalized covariance. To maximize the discriminant gain, we first quantify the influence of the sensing, computation, and communication processes on it with a derived closed-form expression. Then, an end-to-end task-oriented resource management approach is developed by designing an optimal integrated sensing, computation, and communication (ISCC) scheme. By using human motions recognition as a concrete AI inference task, extensive experiments are conducted to verify the performance of the proposed scheme.
Dingzhu Wen, Peixi Liu, Guangxu Zhu, Yuanming Shi, Jie Xu 0002, Yonina C. Eldar, Shuguang Cui
ICC3
2023 Communication Resources Constrained Hierarchical Federated Learning for End-to-End Autonomous Driving
abstract
While federated learning (FL) improves the generalization of end-to-end autonomous driving by model aggregation, the conventional single-hop FL (SFL) suffers from slow convergence rate due to long-range communications among vehicles and cloud server. Hierarchical federated learning (HFL) overcomes such drawbacks via introduction of mid-point edge servers. However, the orchestration between constrained communication resources and HFL performance becomes an urgent problem. This paper proposes an optimization-based Communication Resource Constrained Hierarchical Federated Learning (CRCHFL) framework to minimize the generalization error of the autonomous driving model using hybrid data and model aggregation. The effectiveness of the proposed CRCHFL is evaluated in the Car Learning to Act (CARLA) simulation platform. Results show that the proposed CRCHFL both accelerates the convergence rate and enhances the generalization of federated learning autonomous driving model. Moreover, under the same communication resource budget, it outperforms the HFL by 10.33% and the SFL by 12.44%.
Wei-Bin Kou, Shuai Wang 0004, Guangxu Zhu, Bin Luo 0004, Yingxian Chen, Derrick Wing Kwan Ng, Yik-Chung Wu
IROS3
2023 UKFWiTr: A Single-link Indoor Tracking Method Based on WiFi CSI
abstract
The indoor location based services are fascinating in many applications such as commercial recommendation, surveillance, and navigation. In this paper, we propose a high-precision indoor single-link passive tracking method based on Unscented Kalman Filter (UKF) using WiFi channel state information (CSI), namely UKFWiTr. In this method, both the CSI-quotient and Space-Alternating Generalized Expectation-maximization algorithm are used to estimate Doppler frequency shift and Time-of-Flight. Then, an Arrival-of-Angle optimization method and a tracking accuracy improvement method both based on UKF are put forward in UKFWiTr. The experimental results show that the average tracking error in different environments is less than 1.3m, and can even achieve 0.49m in particular scenarios.
Jiachen Wang 0007, Hang Li 0003, Xiaoyang Li 0002, Chao Shen 0004, Guangxu Zhu
WCNC6
2023 Task-Oriented Over-the-Air Computation for Multi-Device Edge Split Inference
abstract
A task-oriented over-the-air computation (AirComp) scheme is proposed in this paper for multi-device edge split inference system. In the considered system, local noise-corrupted feature vectors are aggregated at the server via AirComp to generate a denoised one for the subsequent inference task. By considering classification tasks, the transmit precoders at edge devices and receive beamforming at edge server are jointly designed in an effort to rein in the aggregation error and maximize the inference accuracy, which is approximately measured by a surrogate but more tractable metric called discriminant gain. It is found that the conventional AirComp beamforming design for minimizing the mean square error between the aggregated feature vector by AirComp and the ideally aggregated one may not lead to the optimal classification accuracy, as it fails to respect the fact that some feature dimensions are more sensitive to the aggregation error than the others in terms of the classification accuracy. To tackle this issue, a new task-oriented AirComp scheme is proposed for directly maximizing the derived discriminant gain. The superiority of the proposed scheme over the heuristic benchmarks is verified by extensive experimental results based on a concrete inference task of human motion recognition.
Dingzhu Wen, Xiang Jiao, Peixi Liu, Guangxu Zhu, Yuanming Shi, Kaibin Huang
WCNC4
2023 Pushing AI to wireless network edge: an overview on integrated sensing, communication, and computation towards 6G
Guangxu Zhu, Zhonghao Lyu, Xiang Jiao, Peixi Liu, Mingzhe Chen, Jie Xu 0002, Shuguang Cui
Sci. China Inf. Sci.1
2023 Integrated Sensing, Communication, and Computation Over-the-Air: MIMO Beamforming Design
abstract
To support the unprecedented growth of the Internet of Things (IoT) applications, tremendous data need to be collected by the IoT devices and delivered to the server for further computation. By utilizing the same signals for both radar sensing and data transmission, theintegrated sensing and communication(ISAC) technique enables simultaneous data collection and delivery in the physical layer. By exploiting the analog-wave addition property in a multi-access channel,over-the-air computation(AirComp) has been proposed as a communication approach that also enables function computation. The promising performances of ISAC and AirComp motivate the current work on developing a framework calledintegrated sensing, communication, and computation over-the-air(ISCCO). Two schemes are designed to supportmultiple-input-multiple-output(MIMO) ISCCO simultaneously, namely theseparated and sharedschemes. The separated scheme splits antenna array for radar sensing and AirComp, while all the antennas transmit a joint waveform for both radar sensing and AirComp in the shared scheme. The performance of radar sensing is evaluated by themean squared error(MSE) of the estimated target response matrix, while the MSE of the estimated function is adopted as the metric to evaluate the performance of the coupled communication and computation in AirComp. The design challenge of MIMO ISCCO lies in the joint optimization of beamformers at both the IoT devices and the server, which results in a non-convex problem. To solve this problem, an algorithmic solution based on the technique of semidefinite relaxation is proposed. The results reveal that the beamformer at each sensor needs to account for supporting dual-functional signals in the shared scheme, while dedicated beamformers for sensing and AirComp are needed to mitigate the mutual interference between the two functionalities in the separated scheme. The application of ISCCO on target location estimation is further demonstrated via simulation.
Xiaoyang Li 0002, Fan Liu 0005, Ziqin Zhou, Guangxu Zhu, Shuai Wang 0004, Kaibin Huang, Yi Gong 0001
IEEE Trans. Wirel. Commun.4
2023 Energy Efficient Wireless Crowd Labeling: Joint Annotator Clustering and Power Control
abstract
The unprecedented growth of mobile data traffic has fueled the deployment of artificial intelligence (AI) at the network edge, while distilling the intelligence from raw data by machine learning requires tremendous labelling effort. To overcome this challenge, wireless crowd labelling (WCL) is proposed for efficient data labelling by exploiting billions of available mobile annotators and the multicasting property of wireless channels. A WCL system is considered in this paper where unlabelled data (objects) are multicast via fading channels to different clusters of annotators for repetition labelling to improve the accuracy. Given the desired labelling accuracy, the superposition coding technique together with the repetition labelling scheme give rise to a new tradeoff between radio-and-annotator resource consumption. Building on such tradeoff, the annotator clustering and transmit power control are jointly optimized to maximize the labelling throughput (i.e., the number of labelled objects) or minimize the power consumption, resulting in NP-hard integer programming problems. To solve these problems, the optimal structure of annotator clustering is derived by exploiting the property that the power allocation for multicasting objects tends to compensate for the worst channel among the annotators in each cluster. Based on such structure, the throughput maximization problem can be recognized as a longest-path problem and solved by means of branch-and-bound, while the power minimization problem can be recasted to a shortest-path problem and solved by means of forward dynamic programming. The solution approaches can be further simplified when the channels are symmetric by merging the same nodes and cutting the identical paths in the path graph. In addition, exact polices are derived for the special cases where either the annotators or power are constrained. Last, simulation results are presented to demonstrate the performance of our proposed joint designs.
Xiaoyang Li 0002, Guangxu Zhu, Kaiming Shen, Kaifeng Han, Kaibin Huang, Yi Gong 0001
IEEE Trans. Wirel. Commun.2
2023 Joint Maneuver and Beamforming Design for UAV-Enabled Integrated Sensing and Communication
abstract
This paper studies the unmanned aerial vehicle (UAV)-enabled integrated sensing and communication (ISAC), in which UAVs are dispatched as aerial dual-functional access points (APs) that can exploit the UAV maneuver control and strong line-of-sight (LoS) air-to-ground (A2G) links for efficient communication and sensing. In particular, we consider that one UAV-AP, equipped with a vertically placed uniform linear array (ULA), sends combined information and sensing signals to communicate with multiple users and at the same time sense potential targets at interested areas on the ground. Under this setup, we consider two scenarios with quasi-stationary and fully mobile UAVs, in which the UAV is deployed at an optimizable location over the whole ISAC mission period and can fly over different locations during the ISAC mission period, respectively. For the two scenarios, our objective is to jointly design the UAV maneuver (deployment location or flight trajectory) and the transmit beamforming, for maximizing the weighted sum-rate throughput of communication users, while ensuring the sensing beampattern gain requirements, subject to the transmit power and flight constraints. However, due to the ULA consideration at the UAV, the two formulated problems are highly non-convex and very difficult to be optimally solved, as the UAV’s location/trajectory variables are involved on the exponent parts of each entry in the steering vectors, and are closely coupled with the transmit beamforming vectors. To tackle this issue, we propose efficient algorithms to find their suboptimal but high-quality solutions, by using various techniques from convex and non-convex optimization. Finally, numerical results are provided to validate the superiority of our proposed designs as compared to various benchmark schemes with heuristic maneuver designs. It is shown that the joint maneuver and transmit beamforming design efficiently balances the inherent tradeoff between sensing and communication with regards to different beampattern gain thresholds.
Zhonghao Lyu, Guangxu Zhu, Jie Xu 0002
IEEE Trans. Wirel. Commun.2
2022 Learning and Energy Efficient Edge Intelligence: Data Partition and Rate Control
abstract
The rapid development of artificial intelligence together with the powerful computation capabilities of the advanced edge servers make it possible to deploy learning tasks at the wireless network edge, which is dubbed as edge intelligence (EI). The communication bottleneck between the data resource and the server results in deteriorated learning performance as well as tremendous energy consumption. To tackle this challenge, we explore a new paradigm called learning-and-energy-efficient (LEE) EI, which simultaneously maximizes the learning accuracies and energy efficiencies of multiple tasks via data partition and rate control. Mathematically, this results in a multi-objective optimization problem. Moreover, the continuous varying rates introduce infinite variables, which further complicates the problem. To solve this complex problem, the number of variables is reduced to a finite level by exploiting the optimality of constant-rate transmission in each epoch, based on which a string-pulling (SP) algorithm is proposed to obtain the numerical values. The performance of the proposed joint data partition and rate control design is examined by experiments based on public datasets.
Xiaoyang Li 0002, Shuai Wang 0004, Guangxu Zhu, Ziqin Zhou, Kaibin Huang, Yi Gong 0001
ICC3
2022 Joint Trajectory and Beamforming Design for UAV-Enabled Integrated Sensing and Communication
abstract
This paper studies the unmanned aerial vehicle (UAV)-enabled integrated sensing and communication (ISAC), in which UAVs are dispatched as aerial dual-functional access points (APs) that can exploit the UAV maneuver control and strong line- of-sight (LoS) aerial-to-ground (A2G) links for efficient ISAC. Particularly, we consider a scenario with one UAV-AP equipped with a vertically placed uniform linear array (ULA), which sends combined information and sensing signals to communicate with multiple users and at the same time sense potential targets on the ground. Our objective is to jointly design the UAV trajectory and transmit beamforming to maximize the average weighted sum-rate throughput of communication users over the whole period, subject to the sensing beampattern gain requirements and transmit power constraints over different time slots, as well as practical flight constraints. While the above problem is challenging to solve, we propose an efficient algorithm by adopting the alternating optimization together with the successive convex approximation (SCA) and semidefinite relaxation (SDR). Numerical results are provided to validate the superiority of our proposed designs as compared to various benchmark schemes with heuristic trajectory designs.
Zhonghao Lyu, Guangxu Zhu, Jie Xu 0002
ICC2
2022 Accelerating Edge Intelligence via Integrated Sensing and Communication
abstract
Realizing edge intelligence consists of sensing, communication, training, and inference stages. Conventionally, the sensing and communication stages are executed sequentially, which results in excessive amount of dataset generation and uploading time. This paper proposes to accelerate edge intelligence via integrated sensing and communication (ISAC). As such, the sensing and communication stages are merged so as to make the best use of the wireless signals for the dual purpose of dataset generation and uploading. However, ISAC also introduces additional interference between sensing and communication functionalities. To address this challenge, this paper proposes a classification error minimization formulation to design the ISAC beamforming and time allocation. The globally optimal solution is derived via the rank-1 guaranteed semidefinite relaxation, and performance analysis is performed to quantify the ISAC gain over that of conventional edge intelligence. Simulation results are provided to verify the effectiveness of the proposed ISAC-assisted edge intelligence system. Interestingly, we find that ISAC is always beneficial, when the duration of generating a sample is more than the duration of uploading a sample. Otherwise, the ISAC gain can vanish or even be negative. Nevertheless, we still derive a sufficient condition, under which a positive ISAC gain is feasible.
Tong Zhang 0026, Shuai Wang 0004, Fan Liu 0005, Guangxu Zhu, Rui Wang 0007
ICC5
2022 Training Time Minimization in Quantized Federated Edge Learning under Bandwidth Constraint
abstract
In this paper, the training time minimization problem is investigated in a quantized FEEL system, where the heterogeneous edge devices send quantized gradients to the edge server via orthogonal channels. In particular, a stochastic quantization scheme is adopted for compression of uploaded gradients, which can reduce the burden of per-round communication but may come at the cost of increasing number of communication rounds. The intrinsic trade-off between the number of communication rounds and per-round latency is characterized. Specifically, we analyze the convergence behavior of the quantized FEEL in terms of the optimality gap. Constrained by total bandwidth, the training time minimization problem is formulated as a joint quantization level and bandwidth allocation optimization problem. To this end, an algorithm based on alternating optimization is proposed, which alternatively solves the subproblem of quantization optimization via successive convex approximation and the subproblem of bandwidth allocation via bisection search. With different learning tasks and models, the validation of our analysis and the near-optimal performance of the proposed algorithm are demonstrated by the experimental results.
Peixi Liu, Jiamo Jiang, Guangxu Zhu, Lei Cheng 0003, Wei Jiang 0003, Wu Luo, Zhiqin Wang
WCNC3
2022 Toward Tailored Models on Private AIoT Devices: Federated Direct Neural Architecture Search
abstract
Neural networks often encounter various stringent resource constraints while deploying on edge devices. To tackle these problems with less human efforts, automated machine learning becomes popular in finding various neural architectures that fit diverse Artificial Intelligence of Things (AIoT) scenarios. Recently, to prevent the leakage of private information while enable automated machine intelligence, there is an emerging trend to integrate federated learning and neural architecture search (NAS). Although promising as it may seem, the coupling of difficulties from both tenets makes the algorithm development quite challenging. In particular, how to efficiently search the optimal neural architecture directly from massive nonindependent and identically distributed (non-IID) data among AIoT devices in a federated manner is a hard nut to crack. In this article, to tackle this challenge, by leveraging the advances in ProxylessNAS, we propose a federated direct neural architecture search (FDNAS) framework that allows for hardware-friendly NAS from non-IID data across devices. To further adapt to both various data distributions and different type of devices with heterogeneous embedded hardware platforms, inspired by meta-learning, a cluster federated direct neural architecture search (CFDNAS) framework is proposed to achieve device-aware NAS, in the sense that each device can learn a tailored deep learning model for its particular data distribution and hardware constraint. Extensive experiments on non-IID data sets have shown the state-of-the-art accuracy–efficiency tradeoffs achieved by the proposed solution in the presence of both data and device heterogeneity.
Xiaoming Yuan 0002, Qianyun Zhang 0001, Guangxu Zhu, Lei Cheng 0003, Ning Zhang 0007
IEEE Internet Things J.4
2022 Transmission Power Control for Over-the-Air Federated Averaging at Network Edge
abstract
Over-the-air computation(AirComp) has emerged as a new analog power-domainnon-orthogonal multiple access(NOMA) technique for low-latency model/gradient-updatesaggregation in federated edge learning(FEEL). By integrating communication and computation into a joint design, AirComp can significantly enhance the communication efficiency, but at the cost of aggregation errors caused by channel fading and noise. This paper studies a particular type of FEEL with federated averaging (FedAvg) and AirComp-based model-update aggregation, namelyover-the-airFedAvg (Air-FedAvg). We investigate the transmission power control in Air-FedAvg to combat against the AirComp aggregation errors for enhancing the training accuracy and accelerating the training speed. Towards this end, we first analyze the convergence behavior (in terms of the optimality gap) of Air-FedAvg with aggregation errors at different outer iterations. Then, to enhance the training accuracy, we minimize the optimality gap by jointly optimizing the transmission power control at edge devices and the denoising factors at edge server, subject to a series of power constraints at individual edge devices. Furthermore, to accelerate the training speed, we also minimize the training latency of Air-FedAvg with a given targeted optimality gap, in which learning hyper-parameters including the numbers of outer iterations and local training epochs are optimized jointly with the power control. Finally, numerical results show that the proposed transmission power control policy achieves significantly faster convergence speed for Air-FedAvg, as compared with benchmark policies with fixed power transmission or per-iterationmean squared error(MSE) minimization. It is also shown that the Air-FedAvg achieves an order-of-magnitude shorter training latency than the conventional FedAvg with digitalorthogonal multiple access(OMA-FedAvg).
Xiaowen Cao 0001, Guangxu Zhu, Jie Xu 0002, Shuguang Cui
IEEE J. Sel. Areas Commun.2
2022 Optimized Power Control Design for Over-the-Air Federated Edge Learning
abstract
Over-the-air federated edge learning(Air-FEEL) has emerged as a communication-efficient solution to enable distributed machine learning over edge devices by using their data locally to preserve the privacy. By exploiting the waveform superposition property of wireless channels, Air-FEEL allows the “one-shot” over-the-air aggregation of gradient-updates to enhance the communication efficiency, but at the cost of a compromised learning performance due to the aggregation errors caused by channel fading and noise. This paper investigates the transmission power control to combat against such aggregation errors in Air-FEEL. Different from conventional power control designs (e.g., to minimize the individualmean squared error(MSE) of the over-the-air aggregation at each round), we consider a new power control design aiming at directly maximizing the convergence speed. Towards this end, we first analyze the convergence behavior of Air-FEEL (in terms of the optimality gap) subject to aggregation errors at different communication rounds. It is revealed that if the aggregation estimates are unbiased, then the training algorithm would converge exactly to the optimal point with mild conditions; while if they are biased, then the algorithm would converge with an error floor determined by the accumulated estimate bias over communication rounds. Next, building upon the convergence results, we optimize the power control to directly minimize the derived optimality gaps under the cases without and with unbiased aggregation constraints, subject to a set of average and maximum power constraints at individual edge devices. We transform both problems into convex forms, and obtain their structured optimal solutions, both appearing in a form of regularized channel inversion, by using the Lagrangian duality method. Finally, numerical results show that the proposed power control policies achieve significantly faster convergence for Air-FEEL, as compared with benchmark policies with fixed power transmission or conventional MSE minimization.
Xiaowen Cao 0001, Guangxu Zhu, Jie Xu 0002, Zhiqin Wang, Shuguang Cui
IEEE J. Sel. Areas Commun.2
2022 Training time minimization for federated edge learning with optimized gradient quantization and bandwidth allocation
abstract
Training a machine learning model with federated edge learning (FEEL) is typically time consuming due to the constrained computation power of edge devices and the limited wireless resources in edge networks. In this study, the training time minimization problem is investigated in a quantized FEEL system, where heterogeneous edge devices send quantized gradients to the edge server via orthogonal channels. In particular, a stochastic quantization scheme is adopted for compression of uploaded gradients, which can reduce the burden of per-round communication but may come at the cost of increasing the number of communication rounds. The training time is modeled by taking into account the communication time, computation time, and the number of communication rounds. Based on the proposed training time model, the intrinsic trade-off between the number of communication rounds and per-round latency is characterized. Specifically, we analyze the convergence behavior of the quantized FEEL in terms of the optimality gap. Furthermore, a joint data-and-model-driven fitting method is proposed to obtain the exact optimality gap, based on which the closed-form expressions for the number of communication rounds and the total training time are obtained. Constrained by the total bandwidth, the training time minimization problem is formulated as a joint quantization level and bandwidth allocation optimization problem. To this end, an algorithm based on alternating optimization is proposed, which alternatively solves the subproblem of quantization optimization through successive convex approximation and the subproblem of bandwidth allocation by bisection search. With different learning tasks and models, the validation of our analysis and the near-optimal performance of the proposed optimization algorithm are demonstrated by the simulation results.
Peixi Liu, Jiamo Jiang, Guangxu Zhu, Lei Cheng 0003, Wei Jiang 0003, Wu Luo, Zhiqin Wang
Frontiers Inf. Technol. Electron. Eng.3
2022 Data Partition and Rate Control for Learning and Energy Efficient Edge Intelligence
abstract
The rapid development of artificial intelligence together with the powerful computation capabilities of the advanced edge servers make it possible to deploy learning tasks at the wireless network edge, which is dubbed as edge intelligence (EI). The communication bottleneck between the data resource and the server results in deteriorated learning performance as well as tremendous energy consumption. To tackle this challenge, we explore a new paradigm called learning-and-energy-efficient (LEE) EI, which simultaneously maximizes the learning accuracies and energy efficiencies of multiple tasks via data partition and rate control. Mathematically, this results in a multi-objective optimization problem. Moreover, the continuously varying communication rates introduce infinite variables, which further complicates the problem. To solve this complex problem, we consider the case with infinite server buffer capacity and one-shot data arrival at sensor. First, the number of variables is reduced to a finite level by exploiting the optimality of constant-rate transmission in each epoch. Second, the optimal solution of the multi-objective problem is found by applying the stratified sequencing or merging of objectives. By assuming higher priority of learning efficiency in stratified sequencing, the optimal data partition is derived in closed form by the Lagrange method, while the optimal rate control is proved to have the structure of directional water filling (DWF), based on which a string-pulling (SP) algorithm is proposed to obtain the numerical values. The DWF structure of rate control is also proved to be optimal in merging of objectives, which combines different objectives in a weighted manner. By exploiting the optimal rate changing properties, the SP algorithm is further extended to tackle the more challenging cases with limited server buffer capacity or bursty data arrival at sensor. The performance of the proposed joint data partition and rate control design is examined by extensive experiments based on public datasets.
Xiaoyang Li 0002, Shuai Wang 0004, Guangxu Zhu, Ziqin Zhou, Kaibin Huang, Yi Gong 0001
IEEE Trans. Wirel. Commun.3
2022 Communication-Efficient Federated Edge Learning via Optimal Probabilistic Device Scheduling
abstract
Federated edge learning (FEEL) is a popular distributed learning framework that allows privacy-preserving collaborative model training via periodic learning-updates communication between edge devices and server. Due to the constrained bandwidth, only a subset of devices can be selected to upload their updates at each training iteration. This has led to an active research area in FEEL studying the optimal device scheduling policy for communication time minimization. However, owing to the difficulty in quantifying the exact communication time, prior work in this area can only tackle the problem partially and indirectly by minimizing either the iteration rounds or per-round latency, while the total communication time is determined by both metrics. To close this research gap, we make the first attempt in this paper to formulate and solve the communication time minimization problem. We first derive a tight bound to approximate the remaining communication time through cross-disciplinary effort that combines the learning theory for convergence rate analysis and communication theory for per-round latency analysis. Building on the novel analytical result, an optimized probabilistic device scheduling policy is derived in closed-form by solving the approximate communication time minimization problem. It is found that the optimized policy gradually turns its priority from suppressing the remaining communication rounds to reducing per-round latency as the training process evolves. Extensive experiments based on real-world dataset and a use case on collaborative 3D objective detection in autonomous driving are provided to verify the superiority of the proposed policy over three benchmark policies based on the indirect solution approaches.
Maojun Zhang, Guangxu Zhu, Shuai Wang 0004, Jiamo Jiang, Qing Liao 0001, Caijun Zhong, Shuguang Cui
IEEE Trans. Wirel. Commun.2
2022 Turning Channel Noise Into an Accelerator for Over-the-Air Principal Component Analysis
abstract
The enormous data distributed at the network edge and ubiquitous connectivity have led to the emergence of the new paradigm of distributed machine learning and large-scale data analytics. Distributed principal component analysis (PCA) concerns finding a low-dimensional subspace that contains the most important information of high-dimensional data distributed over the network edge. The subspace is useful for distributed data compression and feature extraction. This work advocates the application of over-the-air federated learning to efficient implementation of distributed PCA in a wireless network under a data-privacy constraint, termed AirPCA. The design features the exploitation of the waveform-superposition property of a multi-access channel to realize over-the-air aggregation of local subspace updates computed and simultaneously transmitted by devices to a server, thereby reducing the multi-access latency. The original drawback of this class of techniques, namely channel-noise perturbation to uncoded analog modulated signals, is turned into a mechanism for escaping from saddle points during stochastic gradient descent (SGD) in the AirPCA algorithm. As a result, the convergence of the AirPCA algorithm is accelerated. To materialize the idea, descent speeds in different types of descent regions are analyzed mathematically using martingale theory by accounting for wireless propagation and techniques including broadband transmission, over-the-air aggregation, channel fading and noise. The results reveal the accelerating effect of noise in saddle regions and the opposite effect in other types of regions. The insight and results are applied to designing an online scheme for adapting receive signal power to the type of current descent region. Specifically, the scheme amplifies the noise effect in saddle regions by reducing signal power and applies the power savings to suppressing the effect in other regions. From experiments using real datasets, such power control is found to accelerate convergence while achieving the same convergence accuracy as in the ideal case of centralized PCA.
Guangxu Zhu, Rui Wang 0007, Vincent K. N. Lau, Kaibin Huang
IEEE Trans. Wirel. Commun.2
2021 Optimized Power Control for Over-the-Air Federated Edge Learning
abstract
Over-the-air federated edge learning (Air-FEEL) is a communication-efficient solution for privacy-preserving distributed learning over wireless networks. Air-FEEL allows "one-shot" over-the-air aggregation of gradient/model-updates by exploiting the waveform superposition property of wireless channels, and thus promises an extremely low aggregation latency that is independent of the network size. However, such communication efficiency may come at a cost of learning performance degradation due to the aggregation error caused by the non-uniform channel fading over devices and noise perturbation. Prior work adopted channel inversion power control (or its variants) to reduce the aggregation error by aligning the channel gains, which, however, could be highly suboptimal in deep fading scenarios due to the noise amplification. To overcome this issue, we investigate the power control optimization for enhancing the learning performance of Air-FEEL. Towards this end, we first analyze the convergence behavior of the Air-FEEL by deriving the optimality gap of the loss-function under any given power control policy. Then we optimize the power control to minimize the optimality gap for accelerating convergence, subject to a set of average and maximum power constraints at edge devices. The problem is generally non-convex and challenging to solve due to the coupling of power control variables over different devices and iterations. To tackle this challenge, we develop an efficient algorithm by jointly exploiting the successive convex approximation (SCA) and trust region methods. Numerical results show that the optimized power control policy achieves significantly faster convergence than the benchmark policies such as channel inversion and uniform power transmission.
Xiaowen Cao 0001, Guangxu Zhu, Jie Xu 0002, Shuguang Cui
ICC2
2021 Privacy-Preserving Neural Architecture Search Across Federated IoT Devices
abstract
While deploying on edge devices, deep learning mod-els often encounter various strict resource constraints. Automated machine learning becomes popular in finding various neural architectures that fit diverse Internet of Things (IoT) scenarios to handle these problems with less human efforts. Recently, there is an emerging trend to integrate federated learning and Neural Architecture Search (NAS) to prevent private data leakage while enabling automated machine learning. The algorithm development is quite challenging because of the coupling of difficulties from both tenets, although promising as it may seem. Especially, it is a hard nut to efficiently search the optimal neural architecture directly from massive non-Independent and Identically Distributed (non-IID) data among IoT devices in a federated manner. In this paper, by leveraging the advances in ProxylessNAS, we propose a Federated Direct Neural Architecture Search (FDNAS) framework that allows hardware-friendly NAS from non-IID data across devices to tackle the challenge. Extensive experiments on non-IID datasets demonstrate the state-of-the-art accuracy-efficiency trade-offs achieved by proposed methods.
Xiaoming Yuan 0002, Qianyun Zhang 0001, Guangxu Zhu, Lei Cheng 0003, Ning Zhang 0007
TrustCom4
2021 Symbiotic Sensing and Communications Towards 6G: Vision, Applications, and Technology Trends
abstract
Driven by the vision of intelligent connection of everything and digital twin towards 6G, a myriad of new applications, such as immersive extended reality, autonomous driving, holographic communications, intelligent industrial internet, will emerge in the near future, holding the promise to revolutionize the way we live and work. These trends inspire a novel technical design principle that seamlessly integrates two originally decoupled functionalities, i.e., wireless communication and sensing, into one system in a symbiotic way, which is dubbed symbiotic sensing and communications (SSaC), to endow the wireless network with the capability to “see” and “talk” to the physical world simultaneously. Noting that the term SSaC is used instead of ISAC (integrated sensing and communications) because the word “symbiotic/symbiosis” is more inclusive and can better accommodate different integration levels and evolution stages of sensing and communications. Aligned with this understanding, this article makes the first attempts to clarify the concept of SSaC, illustrate its vision, envision the three-stage evolution roadmap, namely neutralism, commensalism, and mutualism of SaC. Then, three categories of applications of SSaC are introduced, followed by detailed description of typical use cases in each category. Finally, we summarize the major performance metrics and key enabling technologies for SSaC.
Zhiqin Wang, Kaifeng Han, Jiamo Jiang, Zhiqing Wei, Guangxu Zhu, Zhiyong Feng 0001, Jianmin Lu, Chunwei Meng
VTC Fall5
2021 Cooperative Interference Management for Over-the-Air Computation Networks
abstract
Recently, over-the-air computation (AirComp) has emerged as an efficient solution for access points (APs) to aggregate distributed data from many edge devices (e.g., sensors) by exploiting the waveform superposition property of multiple access (uplink) channels. While prior work focuses on the single-cell setting where inter-cell interference is absent, this article considers a multi-cell AirComp network limited by such interference and investigates the optimal policies for controlling devices' transmit power to minimize the mean squared errors (MSEs) in aggregated signals received at different APs. First, we consider the scenario of centralized multi-cell power control. To quantify the fundamental AirComp performance tradeoff among different cells, we characterize the Pareto boundary of the multi-cell MSE region by minimizing the sum MSE subject to a set of constraints on individual MSEs. Though the sum-MSE minimization problem is non-convex and its direct solution intractable, we show that this problem can be optimally solved via equivalently solving a sequence of convex second-order cone program (SOCP) feasibility problems together with a bisection search. This results in an efficient algorithm for computing the optimal centralized multi-cell power control, which optimally balances the interference-and-noise-induced errors and the signal misalignment errors unique for AirComp. Next, we consider the other scenario of distributed power control, e.g., when there lacks a centralized controller. In this scenario, we introduce a set of interference temperature (IT) constraints, each of which constrains the maximum total inter-cell interference power between a specific pair of cells. Accordingly, each AP only needs to individually control the power of its associated devices for single-cell MSE minimization, but subject to a set of IT constraints on their interference to neighboring cells. By optimizing the IT levels, the distributed power control is shown to provide an alternative method for characterizing the same multi-cell MSE Pareto boundary as the centralized counterpart. Building on this result, we further propose an efficient algorithm for different APs to cooperate in iteratively updating the IT levels to achieve a Pareto-optimal MSE tuple, by pairwise information exchange. Last, simulation results demonstrate that cooperative power control using the proposed algorithms can substantially reduce the sum MSE of AirComp networks compared with the conventional single-cell approaches.
Xiaowen Cao 0001, Guangxu Zhu, Jie Xu 0002, Kaibin Huang
IEEE Trans. Wirel. Commun.2
2021 Wireless Data Acquisition for Edge Learning: Data-Importance Aware Retransmission
abstract
By deploying machine-learning algorithms at the network edge, edge learning can leverage the enormous real-time data generated by billions of mobile devices to train AI models, which enable intelligent mobile applications. In this emerging research area, one key direction is to efficiently utilize radio resources for wireless data acquisition to minimize the latency of executing a learning task at an edge server. Along this direction, we consider the specific problem of retransmission decision in each communication round to ensure both reliability and quantity of those training data for accelerating model convergence. To solve the problem, a new retransmission protocol called data-importance aware automatic-repeat-request (importance ARQ) is proposed. Unlike the classic ARQ focusing merely on reliability, importance ARQ selectively retransmits a data sample based on its uncertainty which helps learning and can be measured using the model under training. Underpinning the proposed protocol is a derived elegant communication-learning relation between two corresponding metrics, i.e., signal-to-noise ratio (SNR) and data uncertainty. This relation facilitates the design of a simple threshold based policy for importance ARQ. The policy is first derived based on the classic classifier model of support vector machine (SVM), where the uncertainty of a data sample is measured by its distance to the decision boundary. The policy is then extended to the more complex model of convolutional neural networks (CNN) where data uncertainty is measured by entropy. Extensive experiments have been conducted for both the SVM and CNN using real datasets with balanced and imbalanced distributions. Experimental results demonstrate that importance ARQ effectively copes with channel fading and noise in wireless data acquisition to achieve faster model convergence than the conventional channel-aware ARQ. The gain is more significant when the dataset is imbalanced.
Dongzhu Liu, Guangxu Zhu, Qunsong Zeng, Jun Zhang 0004, Kaibin Huang
IEEE Trans. Wirel. Commun.2
2021 One-Bit Over-the-Air Aggregation for Communication-Efficient Federated Edge Learning: Design and Convergence Analysis
abstract
Federated edge learning (FEEL) is a popular framework for model training at an edge server using data distributed at edge devices (e.g., smart-phones and sensors) without compromising their privacy. In the FEEL framework, edge devices periodically transmit high-dimensional stochastic gradients to the edge server, where these gradients are aggregated and used to update a global model. When the edge devices share the same communication medium, the multiple access channel (MAC) from the devices to the edge server induces a communication bottleneck. To overcome this bottleneck, an efficient broadband analog transmission scheme has been recently proposed, featuring the aggregation of analog modulated gradients (or local models) via the waveform-superposition property of the wireless medium. However, the assumed linear analog modulation makes it difficult to deploy this technique in modern wireless systems that exclusively use digital modulation. To address this issue, we propose in this work a novel digital version of broadband over-the-air aggregation, called one-bit broadband digital aggregation (OBDA). The new scheme features one-bit gradient quantization followed by digital quadrature amplitude modulation (QAM) at edge devices and over-the-air majority-voting based decoding at edge server. We provide a comprehensive analysis of the effects of wireless channel hostilities (channel noise, fading, and channel estimation errors) on the convergence rate of the proposed FEEL scheme. The analysis shows that the hostilities slow down the convergence of the learning process by introducing a scaling factor and a bias term into the gradient norm. However, we show that all the negative effects vanish as the number of participating devices grows, but at a different rate for each type of channel hostility.
Guangxu Zhu, Deniz Gündüz, Kaibin Huang
IEEE Trans. Wirel. Commun.1
2020 One-Bit Over-the-Air Aggregation for Communication-Efficient Federated Edge Learning
abstract
To mitigate the multi-access latency in federated edge learning, an efficient broadband analog transmission scheme has been recently proposed, featuring the aggregation of analog modulated gradients via the waveform-superposition property of the wireless medium. However, the assumed linear analog modulation makes it difficult to deploy this technique in modern wireless systems that exclusively use digital modulation. To address this issue, we propose in this work a novel digital version of broadband over-the-air aggregation, called one-bit broadband digital aggregation. The new scheme features one-bit gradient quantization followed by digital modulation at the edge devices and a simple threshold-based decoding at the edge server. We develop a comprehensive analysis framework for quantifying the effects of wireless channel hostilities (channel noise and fading) on the convergence rate. The analysis shows that the hostilities slow down the convergence of the learning process by introducing a scaling factor and a bias term into the gradient norm. However, all the negative effects vanish as the number of devices grows, but at a different rate for each type of channel hostility.
Guangxu Zhu, Deniz Gündüz, Kaibin Huang
GLOBECOM1
2020 Spectrum Allocation in Wireless Networks for Crowd Labelling
abstract
The massive sensing data generated by Internet-of-Things will provide fuel for ubiquitous artificial intelligence (AI), while tremendous labels are required for AI model training via supervised learning. To tackle this challenge, a novel framework of wireless crowd labelling is proposed that downloads data to many imperfect mobile annotators for repetition labelling by exploiting multicasting in wireless networks. The integration of the rate-distortion theory and the principle of repetition labelling gives rise to a new tradeoff between radio-and-annotator resources under a constraint on labelling accuracy. Aiming at maximizing the labelling throughput, this work focuses on optimizing the joint annotator-and-spectrum allocation (JASA). To develop an efficient solution approach, an optimal sequential annotator-clustering scheme is derived. Thereby, the optimal JASA policy can be found by an efficient tree search.
Xiaoyang Li 0002, Guangxu Zhu, Kaiming Shen, Yi Gong 0001, Kaibin Huang
ICASSP2
2020 Optimized Power Control for Over-the-Air Computation in Fading Channels
abstract
Over-the-air computation (AirComp) of a function (e.g., averaging) has recently emerged as an efficient multiple-access scheme for fast aggregation of distributed data at mobile devices (e.g., sensors) at a fusion center (FC) over wireless channels. To realize reliable AirComp in practice, it is crucial to adaptively control the devices' transmit power for coping with channel distortion to achieve the desired magnitude alignment of simultaneous signals. In this paper, we solve the power control problem. Our objective is to minimize the computation error by jointly optimizing the transmit power at devices and a signal scaling factor (called denoising factor) at the FC, subject to individual average power constraints at devices. The problem is generally non-convex due to the coupling of the transmit powers at devices and denoising factor at the FC. To tackle the challenge, we first consider the special case with static channels, for which we derive the optimal solution in closed form. The derived power control exhibits a threshold-based structure: if the product of the channel quality and power budget for each device, called quality indicator, exceeds an optimized threshold, this device applies channel-inversion power control; otherwise, it performs full power transmission. We proceed to consider the general case with time-varying channels. To solve the more challenging non-convex power control problem, we use the Lagrange-duality method via exploiting its “time-sharing” property. The derived power control exhibits a regularized channel inversion structure, where the regularization balances the tradeoff between the signal-magnitude alignment and noise suppression. Moreover, for the special case with only one device being power limited, we show that the power control for the power-limited device has an interesting channel-inversion water-filling structure, while those for other devices (with sufficiently large power budgets) reduce to channel-inversion power control. Numerical results show that the derived power control significantly reduces the computation error as compared with the conventional designs.
Xiaowen Cao 0001, Guangxu Zhu, Jie Xu 0002, Kaibin Huang
IEEE Trans. Wirel. Commun.2
2020 Joint Annotator-and-Spectrum Allocation in Wireless Networks for Crowd Labeling
abstract
The massive sensing data generated by Internet-of-Things will provide fuel for ubiquitous artificial intelligence (AI), automating the operations of our society ranging from transportation to healthcare. The implementation of ubiquitous AI, however, entails labelling of an enormous amount of data prior to the training of AI models via supervised learning. To tackle this challenge, we explore a new direction called wireless crowd labelling, which involves downloading data to many imperfect mobile annotators for repetition labelling with an aim of exploiting multicasting in wireless networks. In this cross-disciplinary area, the rate-distortion theory and the principle of repetition labelling for accuracy improvement together give rise to a new tradeoff between radio-and-annotator resources under a constraint on labelling accuracy. Building on the tradeoff and aiming at maximizing the labelling throughput, this work focuses on the joint optimization of encoding rate, annotator clustering, and sub-channel allocation, which results in an NP-hard integer programming problem. To devise an efficient solution approach, we establish an optimal sequential annotator-clustering scheme based on the order of decreasing signal-to-noise ratios, thereby allowing the optimal solution to be found by an efficient tree search. This solution can be further simplified when the channels are symmetric. Alternatively, the optimization problem can be recognized as a knapsack problem, which can be efficiently solved in pseudo-polynomial time by means of dynamic programming. In addition, the optimal polices are derived for the annotator constrained and spectrum constrained cases. Last, simulation results are presented to demonstrate the significant throughput gains based on the optimal solution compared with decoupled allocation of the two types of resources.
Xiaoyang Li 0002, Guangxu Zhu, Kaiming Shen, Wei Yu 0001, Yi Gong 0001, Kaibin Huang
IEEE Trans. Wirel. Commun.2
2020 Broadband Analog Aggregation for Low-Latency Federated Edge Learning
abstract
To leverage rich data distributed at the network edge, a new machine-learning paradigm, called edge learning, has emerged where learning algorithms are deployed at the edge for providing intelligent services to mobile users. While computing speeds are advancing rapidly, the communication latency is becoming the bottleneck of fast edge learning. To address this issue, this work is focused on designing a low-latency multi-access scheme for edge learning. To this end, we consider a popular privacy-preserving framework, federated edge learning (FEEL), where a global AI-model at an edge-server is updated by aggregating (averaging) local models trained at edge devices. It is proposed that the updates simultaneously transmitted by devices over broadband channels should be analog aggregated “over-the-air” by exploiting the waveform-superposition property of a multi-access channel. Such broadband analog aggregation (BAA) results in dramatical communication-latency reduction compared with the conventional orthogonal access (i.e., OFDMA). In this work, the effects of BAA on learning performance are quantified targeting a single-cell random network. First, we derive two tradeoffs between communication-and-learning metrics, which are useful for network planning and optimization. The power control (“truncated channel inversion”) required for BAA results in a tradeoff between the update-reliability [as measured by the receive signal-to-noise ratio (SNR)] and the expected update-truncation ratio. Consider the scheduling of cell-interior devices to constrain path loss. This gives rise to the other tradeoff between the receive SNR and fraction of data exploited in learning. Next, the latency-reduction ratio of the proposed BAA with respect to the traditional OFDMA scheme is proved to scale almost linearly with the device population. Experiments based on a neural network and a real dataset are conducted for corroborating the theoretical results.
Guangxu Zhu, Kaibin Huang
IEEE Trans. Wirel. Commun.1
2019 Optimal Power Control for Over-the-Air Computation
abstract
Over-the-air computation (AirComp) of a function (e.g., averaging) has recently emerged as an efficient multi-access scheme for fast aggregation of distributed data at devices (e.g., sensors) to fusion centers (FCs) over wireless channels. To realize reliable AirComp in practice, it is crucial to control the devices' transmit power for coping with channel distortion to achieve the desired magnitude alignment of simultaneous signals. % to strike a balance between enforcing signal-magnitude alignment for overcoming heterogenous channel fading and suppressing noise. In this paper, we study the power control problem for AirComp over fading channels. Our objective is to minimize the computation error by jointly optimizing the transmit power at devices and a signal scaling factor at the FC, called denoising factor, subject to the individual average power constraints at devices. The problem is generally non-convex due to the coupling of transmit power over devices and denoising factor. To optimally solve this problem, we apply the Lagrange duality method via exploiting its ''time-sharing'' property. The derived optimal power control exhibits a regularized channel inversion structure where the regularization has the function of balancing the tradeoff between the signal-magnitude alignment and noise suppression. Moreover, for the special case that only one device is power-limited, we show that the optimal power control for the power-limited device has an interesting channel-inversion water- filling structure, while those for other devices (with sufficiently large power budgets) reduce to channel-inversion power control over all fading states. Numerical results show that the optimal power control remarkably reduces the computation error as compared with other heuristic designs.
Xiaowen Cao 0001, Guangxu Zhu, Jie Xu 0002, Kaibin Huang
GLOBECOM2
2019 Reduced-Dimension Design of MIMO AirComp for Data Aggregation in Clustered IoT Networks
abstract
One basic operation of Internet-of-Things (IoT) networks is to acquire a function of distributed data collected from sensors over wireless channels, called wireless data aggregation (WDA). Targeting dense sensors, low-latency WDA poses a design challenge for high-mobility or mission critical IoT applications. A promising solution is a low- latency multi-access scheme, called over-the-air computing (AirComp), that supports simultaneous transmission such that an access point (AP) can estimate and receive a summation-form function of the distributed data by exploiting the waveform- superposition property of multi-access channels. In this work, we propose a multiple-input-multiple-output (MIMO) AirComp framework for an IoT network with clustered multi-antenna sensors and an AP with large receive arrays. The contributions of this work are two-fold. Define the AirComp error as the error in the functional value received at AP due to channel noise. First, under the criterion of minimum error, the optimal receive beamformer at the AP, called decomposed aggregation beamformer (DAB), is shown to have a decomposed architecture: the inner component focuses on channel-dimension reduction and the outer component focuses on joint equalization of the resultant low-dimensional small-scale fading channels. Second, to provision DAB with the required channel state information (CSI), a low-latency channel feedback scheme is proposed by intelligently leveraging the AirComp principle to support simultaneous channel- feedback by sensors.
Dingzhu Wen, Guangxu Zhu, Kaibin Huang
GLOBECOM2
2019 MIMO Over-the-Air Computation for High-Mobility Multimodal Sensing
abstract
In future Internet-of-Things networks, sensors or even access points can be mounted on ground/aerial vehicles for smart-city surveillance or environment monitoring. For such high-mobility sensing, it is impractical to collect data from a large population of sensors using any traditional orthogonal multi-access scheme due to the excessive latency. To tackle the challenge, a technique called over-the-air computation (AirComp) was recently developed to enable a data-fusion center to receive a desired function of sensing data from concurrent sensor transmissions, by exploiting the superposition property of a multi-access channel. This paper aims at further developing multiple-input-multiple output (MIMO) AirComp for enabling high-mobility multimodal sensing. Specifically, we design MIMO-AirComp equalization and channel feedback techniques for spatially multiplexing multifunction computation. Given the objective of minimizing the computation error, a close-to-optimal equalizer is derived in closed-form using differential geometry. The solution can be computed as the weighted centroid of points on a Grassmann manifold, where each point represents the subspace spanned by the channel matrix of a sensor. As a by-product, the problem of MIMO-AirComp equalization is proved to have the same form as the classic problem of multicast beamforming, establishing the AirComp-multicasting duality. Its significance lies in making the said Grassmannian-centroid solution transferable to the latter problem which otherwise is solved using the computation-intensive semidefinite relaxation method. Last, building on the AirComp architecture, an efficient channel-feedback technique is designed for direct acquisition of the equalizer at the access point from simultaneous feedback by all sensors. This overcomes the difficulty of provisioning orthogonal feedback channels for many sensors.
Guangxu Zhu, Kaibin Huang
IEEE Internet Things J.1
2019 Wirelessly Powered Data Aggregation for IoT via Over-the-Air Function Computation: Beamforming and Power Control
abstract
As a revolution in networking, the Internet of Things (IoT) aims at automating the operations of our societies by connecting and leveraging an enormous number of distributed devices (e.g., sensors and actuators). One design challenge is efficient wireless data aggregation (WDA) over the dense IoT devices. This can enable a series of the IoT applications ranging from latency-sensitive high-mobility sensing to data-intensive distributed machine learning. Over-the-air (function) computation (AirComp) has emerged to be a promising solution that merges computing and communication by exploiting analog-wave addition in the air. Another IoT design challenge is battery recharging for dense sensors which can be tackled by wireless power transfer (WPT). The coexisting of AirComp and WPT in the IoT system calls for their integration to enhance the performance and efficiency of WDA. This motivates the current work on developing the wirelessly powered AirComp (WP-AirComp) framework by jointly optimizing wireless power control, energy and (data) aggregation beamforming to minimize the AirComp error. To derive a practical solution, we recast the non-convex joint optimization problem into the equivalent outer and inner sub-problems for (inner) wireless power control and energy beamforming, and (outer) the efficient aggregation beamforming, respectively. The former is solved in closed form while the latter is efficiently solved using the semidefinite relaxation technique. The results reveal that the optimal energy beams point to the dominant Eigen-directions of the WPT channels, and the optimal power allocation tends to equalize the close-loop (down-link WPT and up-link AirComp) effective channels of different sensors. The simulation demonstrates that the controlling WPT provides additional design dimensions for substantially reducing the AirComp error.
Xiaoyang Li 0002, Guangxu Zhu, Yi Gong 0001, Kaibin Huang
IEEE Trans. Wirel. Commun.2
2019 Reduced-Dimension Design of MIMO Over-the-Air Computing for Data Aggregation in Clustered IoT Networks
abstract
One basic operation of Internet-of-Things (IoT) networks is to acquire a function of distributed data collected from sensors over wireless channels, called wireless data aggregation (WDA). In the presence of dense sensors, low-latency WDA poses a design challenge for high-mobility or mission critical IoT applications. A promising solution is a low-latency multi-access scheme, called over-the-air computing (AirComp), that supports simultaneous transmission such that an access point (AP) can estimate and receive a summation-form function of the distributed sensing data by exploiting the waveform-superposition property of a multi-access channel. In this work, we propose a multiple-input-multiple-output (MIMO) AirComp framework for an IoT network with clustered multi-antenna sensors and an AP with large receive arrays. The framework supports low-complexity and low-latency AirComp of a vector-valued function. The contributions of this work are two-fold. Define the AirComp error as the error in the functional value received at AP due to channel noise. First, under the criterion of minimum error, the optimal receive beamformer at the AP, called decomposed aggregation beamformer (DAB), is shown to have a decomposed architecture: the inner component focuses on channel-dimension reduction and the outer component focuses on joint equalization of the resultant low-dimensional small-scale fading channels. In addition, an algorithm is designed to adjust the ranks of individual components of the DAB for a further performance improvement. Second, to provision DAB with the required channel state information (CSI), a low-latency channel feedback scheme is proposed by intelligently leveraging the AirComp principle to support simultaneous channel feedback by sensors. The proposed framework is shown by simulation to substantially reduce AirComp error compared with the existing design without considering channel structures.
Dingzhu Wen, Guangxu Zhu, Kaibin Huang
IEEE Trans. Wirel. Commun.2
2018 Automatic Recognition of Space-Time Constellations by Learning on the Grassmann Manifold
abstract
Recent breakthroughs in machine learning especially artificial intelligence shift the paradigm of wireless communication towards intelligence radios. One of their core operations is automatic modulation recognition (AMR). Existing research focuses on coherent modulation schemes such as QAM, PSK and FSK. The AMR of (non- coherent) space-time modulation remains an uncharted area despite its wide deployment in modern multiple-input-multiple-output (MIMO) systems. The scheme using a so called Grassmann constellation (comprising unitary matrices) enables rate- enhancement using multi-antennas and blind detection. In this work, we propose an AMR approach for Grassmann constellation based on data clustering, which differs from traditional AMR based on classification using a modulation database. The approach allows algorithms for clustering on the Grassmann manifold (or the Grassmannian), such as Grassmann K-means, originally developed for computer vision to be applied to AMR. In this paper, the maximum- likelihood (ML) Grassmann constellation detection is proved to be equivalent to clustering on the Grassmannian. Thereby, a well-known machine-learning result that was originally established only for the Euclidean space is rediscovered for the Grassmannian.
Guangxu Zhu, Jiayao Zhang 0001, Kaibin Huang
GLOBECOM2
2018 MIMO Over-the-Air Computation: Beamforming Optimization on the Grassmann Manifold
abstract
To support future IoT networks with dense sensor connectivity, a technique called over-the-air computation (Air-Comp) was recently developed to enable a data-fusion center to receive a desired function (e.g., mean value) of sensing data from concurrent sensor transmissions. This is made possible by exploiting the superposition property of a multi-access channel. This work aims at further developing AirComp for next-generation multi-antenna multi-modal sensor networks where a multi-modal sensor monitors multiple environmental parameters such as temperature, pollution and humidity. To be specific, we design beamforming techniques for AirComp of multiple functions, each corresponding to a particular sensing-data type. Given the objective of minimizing sum mean-squared error of computed functions, the optimization of receive beamforming for multi-function AirComp is a NP-hard problem. The approximate problem based on tightening transmission-power constraints, however, is shown to be solvable using differential geometry. The solution is proved to be the weighted centroid of points on a Grassmann manifold, where each point represents the subspace spanned by the channel matrix of a sensor. Simulation results demonstrate the effectiveness of the proposed solution.
Guangxu Zhu, Li Chen 0015, Kaibin Huang
GLOBECOM1
2018 Inference From Randomized Transmissions by Many Backscatter Sensors
abstract
Attaining the vision of Smart Cities requires the deployment of an enormous number of sensors for monitoring various conditions of the environment. Backscatter sensors have emerged to be a promising solution due to the uninterruptible energy supply and relative simple hardwares. On the other hand, backscatter sensors with limited signal processing capabilities are unable to support conventional algorithms for multiple access and channel training. Thus, the key challenge in designing backscatter sensor networks is to enable readers to accurately detect sensing values given simple ALOHA random access, primitive transmission schemes, and no knowledge of channel states. We tackle this challenge by proposing the novel framework of backscatter sensing (BackSense) featuring random encoding at sensors and statistical inference at readers. Specifically, assuming the on/off keying for backscatter transmissions, the practical random encoding scheme causes the on/off transmission of a sensor to follow a distribution parameterized by the sensing values. Facilitated by the scheme, statistical inference algorithms are designed to enable a reader to infer sensing values from randomized transmissions by multiple sensors. The specific design procedure involves the construction of Bayesian networks, namely deriving conditional distributions for relating unknown parameters and variables to signals observed by the reader. Then based on the Bayesian networks and the well-known expectation-maximization principle, inference algorithms are derived to recover sensing values. Simulation of the BackSense system demonstrates high accuracy in reader inference despite the mentioned limitations of backscatter sensors, which grows with increasing numbers of received symbols and reader antennas.
Guangxu Zhu, Seung-Woo Ko 0001, Kaibin Huang
IEEE Trans. Wirel. Commun.1
2017 Beamforming via Kronecker Decomposition for Interference Cancellation in the Analog Domain
abstract
The integration of two complementary technologies, millimeter-wave (mmWave) communications and massive multiple-input multiple-output (MIMO), will play a key role in enabling gigabit access in 5G systems. However, implementing mmWave massive MIMO using the traditional fully digital architecture will lead to prohibitive hardware complexity as it requires a massive number of RF chains matching antennas in number. To address this issue, the hybrid beamforming architecture has been recently proposed for efficient implementation of mmWave massive MIMO. Specifically, large-scale MIMO beamforming is implemented in the analog domain, called analog beamforming, that exploits the sparsity in mmWave channels for dramatic dimension reduction for digital MIMO signal processing. The typical phase-array implementation of analog beamforming introduces the uni-modulus constraints on the beamforming coefficients and renders the classic MIMO techniques unsuitable. This motivates the novel design framework, called Kronecker analog beamforming, proposed in this paper for multi-cell multiuser massive MIMO systems over mmWave channels characterized by sparse propagation paths. The framework relies on the decomposition of analog beamforming vectors and path observation vectors into Kronecker products of factor vectors with uni-modulus elements. Exploiting the properties of Kronecker product, different factors of the analog beamformer are designed for either nulling interference paths or coherently combining data paths. Thereby, Kronecker analog beamforming achieves interference nulling and signal enhancement both in the analog domain as well as dimension reduction for digital beamforming.
Guangxu Zhu, Kaibin Huang, Vincent K. N. Lau, Bin Xia 0001, Xiaofan Li 0001
GLOBECOM1
2017 Hybrid Beamforming via the Kronecker Decomposition for the Millimeter-Wave Massive MIMO Systems
abstract
Millimeter-wave (mmWave) massive multiple-input multiple-output (MIMO) seamlessly integrates two wireless technologies, mmWave communications and massive MIMO, which provides spectrums with tens of GHz of total bandwidth and supports aggressive space division multiple access using large-scale arrays. Though it is a promising solution for next-generation systems, the realization of mmWave massive MIMO faces several practical challenges. In particular, implementing massive MIMO in the digital domain requires hundreds to thousands of radio frequency chains and analog-to-digital converters matching the number of antennas. Furthermore, designing these components to operate at the mmWave frequencies is challenging and costly. These motivated the recent development of the hybrid-beamforming architecture, where MIMO signal processing is divided for separate implementation in the analog and digital domains, called the analog and digital beamforming, respectively. Analog beamforming using a phase array introduces uni-modulus constraints on the beamforming coefficients. They render the conventional MIMO techniques unsuitable and call for new designs. In this paper, we present a systematic design framework for hybrid beamforming for multi-cell multiuser massive MIMO systems over mmWave channels characterized by sparse propagation paths. The framework relies on the decomposition of analog beamforming vectors and path observation vectors into Kronecker products of factors being uni-modulus vectors. Exploiting properties of Kronecker mixed products, different factors of the analog beamformer are designed for either nulling interference paths or coherently combining data paths. Furthermore, a channel estimation scheme is designed for enabling the proposed hybrid beamforming. The scheme estimates the angles-of-arrival (AoA) of data and interference paths by analog beam scanning and data-path gains by analog beam steering. The performance of the channel estimation scheme is analyzed. In particular, the AoA spectrum resulting from beam scanning, which displays the magnitude distribution of paths over the AoA range, is derived in closed form. It is shown that the inter-cell interference level diminishes inversely with the array size, the square root of pilot sequence length, and the spatial separation between paths, suggesting different ways of tackling pilot contamination.
Guangxu Zhu, Kaibin Huang, Vincent K. N. Lau, Bin Xia 0001, Xiaofan Li 0001
IEEE J. Sel. Areas Commun.1
2016 Analog spatial decoupling for tackling the near-far problem in wirelessly powered communications
abstract
A practical architecture for wirelessly powered communications (WPC) with dedicated power beacons (PBs) deployed in existing cellular networks, called PB-assisted WPC, is considered in this paper. Assuming those PBs can access to backhaul network and perform simultaneous wireless information and power transfer (SWIPT) to the energy constrained users, the near-far problem in the PB-assisted WPC system is first identified. Specifically, the significant difference of the received power of SWIPT signal (from a PB) and information transfer (IT) signal (from a base station) due to different transmission ranges leads to extremely small signal-to-quantization-noise ratio (SQNR) for the IT signal after quantization of the mixed signals. To retrieve the information carried by the SWIPT and IT signals respectively, it is essential to decouple the strong SWIPT and the weak IT signals in analog domain. To this end, a novel technique called analog spatial decoupling using only simple components such as phase shifters and adders is proposed in this paper. In particular, for the single-PB case, the optimal Fourier based and Hadamard based schemes are proposed for implementing the analog spatial decoupling. For the multiple-PB case, the corresponding design problem is more challenging, making it hard to extend the solution for the single-PB counterpart. To tackle this problem, a systematic solution approach is proposed for analog decoupling with multiple PBs.
Guangxu Zhu, Kaibin Huang
ICC1
2016 Analog Spatial Cancellation for Tackling the Near-Far Problem in Wirelessly Powered Communications
abstract
The implementation of wireless power transfer in wireless communication systems opens up a new research area, known as wirelessly powered communications (WPC). In next-generation heterogeneous networks where ultradense small-cell base stations are deployed, simultaneous-wireless-information-and-power-transfer (SWIPT) is feasible over short ranges. One challenge for designing a WPC system is the severe near-far problem where a user attempts to decode an information-transfer (IT) signal in the presence of extremely strong SWIPT signals. Jointly quantizing the mixed signals causes the IT signal to be completely corrupted by quantization noise, and thus the SWIPT signals have to be suppressed in the analog domain. This motivates the design of a framework in this paper for analog spatial cancellation in a multiantenna WPC system. In the framework, an analog circuit consisting of simple phase shifters and adders is adapted to cancel the SWIPT signals by multiplying it with a cancellation matrix having unit-modulus elements and full rank, where the full rank retains the spatial-multiplexing gain of the IT channel. The unit-modulus constraints render the conventional zero-forcing method unsuitable. Therefore, this paper presents a novel systematic approach for constructing cancellation matrices. For the single-SWIPT-interferer case, the matrices are obtained as truncated Fourier/Hadamard matrices after compensating for propagation phase shifts over the SWIPT channel. For the more challenging multiple-SWIPT-interferer case, it is proposed that each row of the cancellation matrix is constructed as a Kronecker product of component vectors, with each component vectors designed to null the signal from a corresponding SWIPT interferer similarly as in the preceding case.
Guangxu Zhu, Kaibin Huang
IEEE J. Sel. Areas Commun.1
2015 Wireless powered dual-hop multiple antenna relay transmission in the presence of interference
abstract
This paper investigates the impact of the multiple antenna and co-channel interference (CCI) on the outage performance of a dual-hop amplify-and-forward energy harvesting relaying network. The energy constrained relay is powered by radio frequency signals and employs the power splitting receiver architecture. To exploit the benefit of multiple antennas, two different linear processing schemes are investigated, namely, Maximum ratio combining/maximal ratio transmission (MRC/ MRT) and Minimum mean-square error/MRT (MMSE/MRT). For both schemes, a new closed-form outage lower bound and a simple high signal-to-noise ratio outage approximation are derived, respectively. Also, the achievable diversity order is quantified. In addition, we study the optimal power splitting ratio which minimizes the outage probability. Our results show that, by increasing the energy harvesting capability, the implementation of multiple antennas significantly improves the systems performance. Moreover, CCI could be potentially exploited to boost the performance, while how much performance gain can be obtained depends on the choice of the linear processing scheme.
Guangxu Zhu, Caijun Zhong, Himal A. Suraweera, George K. Karagiannidis, Zhaoyang Zhang 0001, Theodoros A. Tsiftsis
ICC1
2015 Wireless Information and Power Transfer in Relay Systems With Multiple Antennas and Interference
abstract
In this paper, an energy harvesting dual-hop relaying system without/with the presence of co-channel interference (CCI) is investigated. Specifically, the energy constrained multi-antenna relay node is powered by either the information signal of the source or via the signal receiving from both the source and interferer. In particular, we first study the outage probability and ergodic capacity of an interference free system, and then extend the analysis to an interfering environment. To exploit the benefit of multiple antennas, three different linear processing schemes are investigated, namely, 1) Maximum ratio combining/maximum ratio transmission (MRC/MRT), 2) Zero-forcing/MRT (ZF/MRT) and 3) Minimum mean-square error/MRT (MMSE/MRT). For all schemes, both the systems outage probability and ergodic capacity are studied, and the achievable diversity order is also presented. In addition, the optimal power splitting ratio minimizing the outage probability is characterized. Our results show that the implementation of multiple antennas increases the energy harvesting capability, hence, significantly improves the systems performance.
Guangxu Zhu, Caijun Zhong, Himal A. Suraweera, George K. Karagiannidis, Zhaoyang Zhang 0001, Theodoros A. Tsiftsis
IEEE Trans. Commun.1
2014 Linear processing for dual-hop AF relay systems with interference: Outage probability analysis
abstract
This paper investigates the impact of different linear processing techniques on the outage performance of dual-hop amplify-and-forward (AF) relaying systems with co-channel interference (CCI) at the multiple antenna relay. Specifically, three heuristic linear precoding schemes are proposed to combat the detrimental effect of CCI, namely, 1) Maximum ratio combining/maximal ratio transmission (MRC/MRT), 2) Zero-forcing/MRT (ZF/MRT), 3) Minimum mean-square error/MRT (MMSE/MRT). New exact outage expressions as well as simple high signal-to-noise ratio (SNR) outage approximations are derived for all three schemes. Our results demonstrate that while the MRC/MRT and the MMSE/MRT schemes achieve a full diversity order of N, the ZF/MRT scheme can only achieve a diversity order of N-M, where N is the number of relay antennas and M is the number of interferers. In addition, it is observed that the MMSE/MRT scheme always achieves the best outage performance. The ZF/MRT scheme outperforms the MRC/MRT scheme in the low SNR regime, and exhibits an inferior performance compared to the MRC/MRT scheme in the high SNR regime.
Guangxu Zhu, Caijun Zhong, Himal A. Suraweera, Zhaoyang Zhang 0001, Chau Yuen
ICC1
2014 Ergodic Capacity Comparison of Different Relay Precoding Schemes in Dual-Hop AF Systems With Co-Channel Interference
abstract
In this paper, we analyze the ergodic capacity of a dual-hop amplify-and-forward relaying system, where the relay is equipped with multiple antennas and subject to co-channel interference and the additive white Gaussian noise. Specifically, we consider three heuristic precoding schemes, where the relay first applies the: 1) maximal-ratio combining (MRC); 2) zero-forcing (ZF); and 3) minimum mean-squared error (MMSE) principle to combine the signal from the source, and then steers the transformed signal toward the destination with the maximum ratio transmission (MRT) technique. For the MRC/MRT and MMSE/MRT schemes, we present new tight analytical upper and lower bounds for the ergodic capacity, while for the ZF/MRT scheme, we derive a new exact analytical ergodic capacity expression. Moreover, we make a comparison among all three schemes, and our results reveal that, in terms of the ergodic capacity performance, the MMSE/MRT scheme always has the best performance and the ZF/MRT scheme is slightly inferior, while the MRC/MRT scheme is always the worst one. Finally, the asymptotic behavior of ergodic capacity for the three proposed schemes are characterized in large N scenario, where N is the number of relay antennas. Our results reveal that, in the large N regime, both the ZF/MRT and MMSE/MRT schemes have perfect interference cancellation capability, which is not possible with the MRC/MRT scheme.
Guangxu Zhu, Caijun Zhong, Himal A. Suraweera, Zhaoyang Zhang 0001, Chau Yuen, Rui Yin 0001
IEEE Trans. Commun.1
2014 Outage Probability of Dual-Hop Multiple Antenna AF Systems with Linear Processing in the Presence of Co-Channel Interference
abstract
This paper considers a dual-hop amplify-and-forward (AF) relaying system where the relay is equipped with multiple antennas, while the source and the destination are equipped with a single antenna. Assuming that the relay is subjected to co-channel interference (CCI) and additive white Gaussian noise (AWGN) while the destination is corrupted by AWGN only, we propose three heuristic relay precoding schemes to combat the CCI, namely, 1) Maximum ratio combining/maximal ratio transmission (MRC/MRT), 2) Zero-forcing/MRT (ZF/MRT), 3) Minimum mean-square error/MRT (MMSE/MRT). We derive new exact outage expressions as well as simple high signal-to-noise ratio (SNR) outage approximations for all three schemes. Our findings suggest that both the MRC/MRT and the MMSE/MRT schemes achieve a full diversity of N, while the ZF/MRT scheme achieves a diversity order of N-M, where N is the number of relay antennas and M is the number of interferers. In addition, we show that the MMSE/MRT scheme always achieves the best outage performance, and the ZF/MRT scheme outperforms the MRC/MRT scheme in the low SNR regime, while becomes inferior to the MRC/MRT scheme in the high SNR regime. Finally, in the large N regime, we show that both the ZF/MRT and MMSE/MRT schemes are capable of completely eliminating the CCI, while perfect interference cancelation is not possible with the MRC/MRT scheme.
Guangxu Zhu, Caijun Zhong, Himal A. Suraweera, Zhaoyang Zhang 0001, Chau Yuen
IEEE Trans. Wirel. Commun.1