Yuanming Shi

dblp:129/1008 · DBLP profile ↗
← Back
184ranked-venue papers
15as first author
124since 2021 · last 2026
0000-0002-1418-7465ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 140 · 11 first-author · 111 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 3 since 2021Systems, architecture and hardware · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Theory of computation · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Communication-Efficient Over-the-Air Federated Fine-Tuning with Heterogeneous LoRA
Shushan He, Yuanming Shi, Yong Zhou 0006
ICC3
2026 Asynchronous Satellite Federated Learning with Intermittent Ground-to-Satellite Links
Ruanjun Li, Jingyang Zhu, Yong Zhou 0006, Yuanming Shi, Linling Kuang, Chunxiao Jiang
ICC4
2026 Unseen Cost of Space Computing: Quantifying LEO Battery Aging via Physics-Driven Modeling
abstract
Low Earth Orbit (LEO) satellite constellations in the 6G era are evolving into intelligent in-orbit computational platforms, forming Space Computing Power Networks (SCPNs) to deliver global-scale computing services. However, the intensive computation within SCPN incurs a significant "unseen cost": the frequent charge-discharge cycles accelerate the physical degradation of satellites’ life-limiting and high-cost batteries, thereby threatening the long-term operational viability of such a system. Existing approaches, often relying on indirect metrics like Depth of Discharge (DoD) and neglecting the complex, nonlinear degradation process of battery aging, fail to accurately quantify this cost. To address this, we introduce a high-fidelity, physics-driven model that quantitatively links computational workload parameters to the nonlinear battery degradation. Building on this model, we formulate a degradation-aware scheduling problem and analyze heuristic policies across different energy regimes. Simulations reveal that the optimal strategy should be adaptive: in solar-rich conditions, a myopic policy maximizing instantaneous solar utilization is superior, whereas under energy scarcity, a reactive policy leveraging real-time battery state significantly extends lifetime.
Jingyang Zhu, Yuanming Shi, Khaled Ben Letaief
ICC4
2026 Service Function Chain Routing in LEO Networks Using Shortest-Path Delay Statistical Stability
abstract
Low Earth orbit (LEO) satellite constellations have become a critical enabler for global coverage, utilizing numerous satellites orbiting Earth at high speeds. By decomposing complex network services into lightweight service functions, network function virtualization (NFV) transforms global network services into diverse service function chains (SFCs), coordinated by resource-constrained LEOs. However, the dynamic topology of satellite networks, marked by highly variable inter-satellite link delays, poses significant challenges for designing efficient routing strategies that ensure reliable and low-latency communication. Many existing routing methods suffer from poor scalability and degraded performance, limiting their practical implementation. To address these challenges, this paper proposes a novel SFC routing approach that leverages the statistical properties of network link states to mitigate instability caused by instantaneous modeling in dynamic satellite networks. Through comprehensive simulations on end-to-end shortest-path propagation delays in LEO networks, we identify and validate the statistical stability of multi-hop routes. Building on this insight, we introduce the Stability-Aware Multi-Stage Graph Routing (SA-MSGR) algorithm, which incorporates pre-computed average delays into a multi-stage graph optimization framework. Extensive simulations demonstrate the superior performance of SA-MSGR, achieving significantly lower and more predictable end-to-end SFC delays compared to representative baseline strategies.
Yuanming Shi, Khaled Ben Letaief
WCNC3
2026 DMT-PPO: Weight-Adaptive Multiobjective Task Offloading With Dynamic Preference Learning in Heterogeneous Edge Computing
Honggang Yuan, Yuxiang Deng, Qin Li 0002, Ting Wang 0001, Yuanming Shi
IEEE Internet Things J.7
2026 Fairness-Aware Joint Source-Channel Coding for Robust Task-Oriented Communication
abstract
Learning-based joint source-channel coding (JSCC) is widely used in task-oriented communication, which aims to extract and transmit only task-relevant information to improve communication efficiency. However, the learning-empowered algorithms in task-oriented communication may lead to information leakage on sensitive attributes and cause discrimination towards specific groups, resulting in fairness issues in social equity. Meanwhile, directly adopting fair representation learning techniques in the source encoder of communication systems poses significant challenges: First, the favorable fairness-utility tradeoff in the encoded feature representations would be deteriorated by channel noise and dynamic variations. Second, the inherent separation of source and channel design precludes the efficiency offered by JSCC for end-to-end transmission. To address these issues, we propose a task-oriented JSCC communication scheme, namely Fair-RIB, that achieves efficient encoding and inference while preserving group fairness. Our approach leverages an information bottleneck-based framework that maximizes the task utility information while limiting the sensitive information leakage to ensure fairness, and adopts a hypernetwork-parametrization mechanism to adapt to varying channel conditions. We also provide theoretical bounds for fairness guarantees by fully exploiting the characteristics of the channel noise, and introduce a selective noise injection mechanism to better manage the fairness-utility tradeoff. To overcome the intractability of the high-dimensional mutual information terms, we adopt variational approximations to derive a tractable upper bound for objective optimization. Experiments on benchmark tabular and image datasets demonstrate the superiority of our framework in achieving a fairness-utility tradeoff and the adaptability to channel variations.
Youlong Wu, Songjie Xie, Shuai Ma 0002, Yuanming Shi, Meixia Tao
IEEE J. Sel. Areas Commun.5
2026 Throughput-Optimized Service Routing for Microservice Flows in LEO Satellite Networks
abstract
Satellite-based microservice systems have emerged as a promising architecture for enabling scalable and flexible service deployment in Low Earth Orbit (LEO) satellite networks, which are increasingly recognized as a key solution to meet the rising demand for global communication and computation, especially in remote and underserved regions. However, routing microservices efficiently in such systems presents major challenges due to the dynamic topology, intermittent connectivity, and unstable link conditions inherent to satellite constellations. These issues become even more severe under high traffic loads. To address these challenges, we propose Service Pressure, a novel routing algorithm specifically designed for satellite-based microservice systems. Service Pressure comprises two key components: first, the construction of an augmented subgraph to model the complex execution and data transmission dependencies of microservices; second, a distributed service routing algorithm that utilizes queue backlogs. This combination enables the algorithm to effectively handle high throughput and adapt to the dynamic network conditions of satellite constellations. By optimizing resource utilization, minimizing latency, and balancing load across satellite nodes, Service Pressure ensures efficient and stable service orchestration even under fluctuating traffic conditions. Extensive simulations demonstrate that our approach significantly outperforms existing routing methods, particularly in terms of throughput, latency, and stability. Service Pressure offers a significant advancement in satellite microservice routing, making it ideal for next-generation space-ground integrated networks.
Xindi He, Ting Wang 0001, Yuanming Shi, Xin Liu 0049
IEEE Trans. Mob. Comput.3
2026 Federated Linear Bandit Learning via UAV Aided Over-the-Air Computation
abstract
This paper investigates federated contextual linear bandit learning in a wireless network with a central server and multiple devices. To reduce communication latency, devices interact with the server via over-the-air computation (AirComp) over noisy, fading channels, where signal distortion can occur due to channel imperfections. Departing from traditional AirComp designs for static networks, we propose a novel federated bandit learning framework that leverages unmanned aerial vehicles (UAVs) as mobile servers to aggregate data from distributed IoT devices. To optimize this system, we employ a block coordinate descent method combined with the alternating direction method of multipliers (BCD-ADMM), jointly optimizing the UAV trajectory, receive normalization factor, and transmission power to minimize the time-averaged mean square error (MSE) of AirComp. Our approach addresses the challenge of decentralized data across multiple devices, enabling secure and efficient collaboration without direct data sharing. Theoretical analysis establishes an upper bound on the algorithm's regret, affirming the framework's scalability and robustness against noise. Simulation results support these findings, highlighting notable performance improvements in federated bandit learning with UAV-assisted AirComp.
Junkai Qian, Yuning Jiang 0002, Xin Liu 0049, Ting Wang 0001, Yuanming Shi, Colin N. Jones
IEEE Trans. Mob. Comput.6
2026 Microservice Deployment in Space Computing Power Networks Via Robust Reinforcement Learning
abstract
With the growing demand for Earth observation, it is important to provide reliable real-time remote sensing inference services to meet the low-latency requirements. The Space Computing Power Network (Space-CPN) offers a promising solution by providing onboard computing and extensive coverage capabilities for real-time inference. This paper presents a remote sensing artificial intelligence applications deployment framework designed for Low Earth Orbit satellite constellations to achieve real-time inference performance. The framework employs the microservice architecture, decomposing monolithic inference tasks into reusable, independent modules to address high latency and resource heterogeneity. This distributed approach enables optimized microservice deployment, minimizing resource utilization while meeting quality of service and functional requirements. We introduce Robust Optimization to the deployment problem to address data uncertainty. Additionally, we model the Robust Optimization problem as a Partially Observable Markov Decision Process and propose a robust reinforcement learning algorithm to handle the semi-infinite Quality of Service constraints. Our approach yields sub-optimal solutions that minimize accuracy loss while maintaining acceptable computational costs. Simulation results demonstrate the effectiveness of our framework.
Yuning Jiang 0002, Xin Liu 0049, Yuanming Shi, Chunxiao Jiang, Linling Kuang
IEEE Trans. Mob. Comput.4
2026 Zeroth-Order Federated Fine-Tuning for Large AI Models in Resource-Constrained Wireless Networks
abstract
Large artificial intelligence (AI) models have demonstrated impressive performance in a wide range of fields. Despite their versatility, adapting large AI models to specific downstream applications often requires fine-tuning on decentralized and privacy-sensitive data, posing significant challenges in resource-constrained wireless networks. In this paper, we propose a novel zeroth-order federated fine-tuning framework for efficient fine-tuning large AI models to alleviate the computation, communication, and memory bottleneck issues. Specifically, to address the computation limitation on edge devices, we adopt the split learning architecture, hosting the most computation-intensive component of the large AI model on the edge server. Besides, we employ a memory-efficient zeroth-order fine-tuning algorithm to further reduce the GPU memory consumption. Furthermore, we conduct a rigorous convergence analysis to illustrate how device scheduling influences the learning performance. Based on the analysis, we formulate a global loss minimization problem that jointly optimizes device scheduling, transmit power, and receive beamforming under the average transmission latency constraint. To tackle this complex mixed-integer nonlinear programming problem, we apply Lyapunov theory to break down the long-term optimization problem into multiple subproblems, followed by designing an effective online algorithm. Simulation results show that the proposed framework can lower the GPU memory consumption by up to 84% and GPU hours by 39% compared to the baselines, while achieving comparable accuracy.
Tianle Wang 0015, Yong Zhou 0006, Yuanming Shi, Nan Cheng 0001, Hangguan Shan
IEEE Trans. Wirel. Commun.3
2026 Decentralized Integration of Sensing-Communication-Computation for Multi-Task Edge AI Inference
abstract
Collaborative artificial intelligence (AI) inference has effectively deployed well-trained AI models at the network edge to empower immersive intelligent services such as autonomous driving and smart cities. This paper proposes an integrated sensing-computation-communication (ISCC) scheme for decentralized multi-task collaborative inference systems. The proposed scheme connects multiple devices via device-to-device (D2D) links. Each device first extracts a homogeneous feature vector from the raw sensory data obtained from the same wide view of the source target and then aggregates all local feature vectors using the over-the-air computation (AirComp) technique to complete a specific inference task. To enhance spectrum efficiency, the full-duplex communication technique is adopted, which allows all devices to transmit and receive in the same frequency band. To suppress the self-interference caused by full duplex communications and simultaneously enhance all tasks’ performance, a multi-objective optimization problem is formulated, where discriminant gain is adopted as the inference performance metric. The challenges to solve this problem arise from three aspects: The impact of the self-interference (SI) channel incurred by full-duplex communication, the precoding design of each device, and the coupling among subcarrier allocation, sensing, computation, and communication processes. To tackle this problem, aquadratic transformandweighted bipartite matchingbased alternating maximization approach is proposed. Numerical results based on jointly completing three tasks of human motion classification, human gender recognition, and human age group classification, verify the effectiveness of the proposed method by showing that the proposed method outperforms the state-of-the-art successive convex approximation (SCA) based algorithm.
Chenye Wang, Zeming Zhuang, Dingzhu Wen, Yuanming Shi, Xin Wang 0003
IEEE Trans. Wirel. Commun.4
2026 Integrated Sensing, Communication, and Computation for Over-the-Air Federated Edge Learning
abstract
This paper studies an over-the-air federated edge learning (Air-FEEL) system with integrated sensing, communication, and computation (ISCC), in which one edge server coordinates multiple edge devices to wirelessly sense the objects and use the sensing data to collaboratively train a machine learning model for recognition tasks. In this system, over-the-air computation (AirComp) is employed to enable one-shot model aggregation from edge devices. Under this setup, we analyze the convergence behavior of the ISCC-enabled Air-FEEL in terms of the loss function degradation, by particularly taking into account the wireless sensing noise during the training data acquisition and the AirComp distortions during the over-the-air model aggregation. The result theoretically shows that sensing, communication, and computation compete for network resources to jointly decide the convergence rate. Based on the analysis, we design the ISCC parameters under the target of maximizing the loss function degradation while ensuring the latency and energy budgets in each round. The challenge lies on the tightly coupled processes of sensing, communication, and computation among different devices. To tackle the challenge, we derive a low-complexity ISCC algorithm by alternately optimizing the batch size control and the network resource allocation. It is found that for each device, less sensing power should be consumed if a larger batch of data samples is obtained and vice versa. Besides, with a given batch size, the optimal computation speed of one device is the minimum one that satisfies the latency constraint. Numerical results based on a human motion recognition task verify the theoretical convergence analysis and show that the proposed ISCC algorithm well coordinates the batch size control and resource allocation among sensing, communication, and computation to enhance the learning performance.
Dingzhu Wen, Sijing Xie, Xiaowen Cao 0001, Yuanhao Cui, Jie Xu 0002, Yuanming Shi, Shuguang Cui
IEEE Trans. Wirel. Commun.6
2026 UAV-Assisted Edge Inference With Integrated Sensing, Communication, and Computation
Dingzhu Wen, Guangxu Zhu, Yuan Liu 0001, Yuanming Shi, Honglin Hu
IEEE Trans. Wirel. Commun.5
2026 Joint Source-Channel Coding for Task-Oriented Broadcast Communications: An Information Bottleneck Approach With Rate Splitting
abstract
To support efficient and accurate multi-task inference in edge environments, we propose a task-oriented broadcast communication system that enables an edge transmitter to serve multiple edge devices with heterogeneous inference tasks. The proposed system adopts a two-phase design inspired by Marton’s channel coding with rate splitting and the information bottleneck principle. In the first phase, a common feature vector is extracted to capture the shared information across tasks. In the second phase, task-specific private feature vectors are generated conditioned on the common feature to preserve unique task-relevant information. To facilitate interference-robust task execution, our scheme leverages the intrinsic structural alignment between the task correlations and broadcast channel properties; specifically, the common and private features are mapped directly to Marton’s common and private codewords. A variational approximation method is introduced to optimize the feature extraction process in both phases, allowing for compact and informative representations while reducing redundant data transmission. Extensive experiments on a real-world multi-label dataset demonstrate that the proposed method achieves superior inference accuracy and robustness over wireless networks, compared to traditional digital compression and deep learning-based joint source-channel coding schemes. These results confirm the potential of task-oriented design for scalable and reliable edge intelligence.
Youlong Wu, Jingfeng Huang, Yuanming Shi, Shuai Ma 0002, Kai Niu 0001, Meixia Tao, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.3
2026 Enabling Symbol-Level mmWave Radar-Backscatter Communication
abstract
This paper presents mmDFRBC, a symbol-level millimeter-wave (mmWave) backscatter communication system that reuses commercial mmWave FMCW radar infrastructure as the dual-function access point (AP) without hardware modification. We propose M-ary Frequency Shift Modulation (MFSM), a lightweight encoding scheme that enables the tag to modulate information using orthogonal frequency blocks over adjacent chirps, achieving symbol-level modulation with data rates exceeding kbps. To robustly extract the modulation signals from strong radar sensing clutter, we introduce a Coherent Cancellation Demodulation (CCD) method that exploits the coherence difference between sensing signals and modulated reflections. We develop a soft synchronization strategy for operating asynchronously and requiring no synchronization or downconversion circuits at the tag, supporting kbps-level data rates with low power and cost. We implement mmDFRBC using a commercial mmWave radar and verify its performance across static and mobile scenarios, achieving BERs below$10^{-3}$over 4 meters, with strong resilience to radar clutter interference. mmDFRBC supports data rates up to 50 kbps, highlighting DFRBC’s potential to enable high-performance DFRC systems using existing infrastructure.
Zeming Yang, Fengyuan Zhu 0001, Yuanming Shi, Yong Zhou 0006, Xiaohua Tian
IEEE Trans. Wirel. Commun.5
2026 MIMO Over-the-Air Computation for Device-Edge Collaborative Inference
abstract
Device-edge collaborative inference, which deploys well-trained artificial intelligence (AI) models at the network edge via the cooperation of edge devices and edge servers, emerges as a promising technique to provide ubiquitous intelligent services. In this paper, a multiple-input multiple-output (MIMO) over-the-air computation (AirComp) scheme is proposed for the efficient implementation of device-edge collaborative inference. In the considered system, the technique of MIMO AirComp is utilized to aggregate local feature vectors, extracted from noise-corrupted sensory data on devices, at the server to efficiently derive a denoised global one for completing the downstream inference task. Device-edge collaborative inference features a task-oriented property, that concerns the effectiveness and efficiency of the task execution. In this case, the traditional AirComp criterion, i.e., minimum mean square error (MMSE), is not effective, since the same distortion level on different feature elements may have different influences on the inference performance. To this end, this paper directly adopts inference accuracy as the design objective. As the instantaneous inference accuracy is unknown during the design stage, an approximated but tractable metric, called discriminant gain, which measures the discernibility of different classes, is adopted. To maximize the inference accuracy measured by discriminant gain, a MIMO AirComp technique is proposed to jointly optimize all feature elements. The problem is nonconvex because of the complicated form of the objective function and the constraints. The solution based on semidefinite relaxation (SDR) and successive convex approximation (SCA) is employed to design a joint transmit precoding and receive beamforming scheme. Besides, to enhance the robustness of practical AI models in the inference stage, a post-processing design of feature magnitude normalization is proposed. Extensive experiments are conducted based on a practical human motion recognition task, which verifies our theoretical analysis and the superiority of our proposed scheme.
Dingzhu Wen, Li You 0001, Jingjing Wang 0001, Sheng Wu 0001, Yuanming Shi
IEEE Trans. Wirel. Commun.6
2026 Integrated Sensing, Computation, and Communication Enabled Federated Edge Learning
abstract
To support ambient intelligence with federated edge learning (FEEL) over resource-constrained wireless networks, it is essential to jointly design and optimize the sensing, computation, and communication processes. In this paper, we propose an integrated sensing, computation, and communication (ISCC) enabled FEEL framework, where each edge device performs wireless sensing to enrich local datasets, executes local model training with accumulated local datasets, and transmits updated local gradients for global model aggregation. Via analyzing the convergence of ISCC-enabled FEEL, we explicitly characterize the impact of newly sensed dataset size in each training round on the optimality gap. Due to the coupling of the sensing, computation, and communication processes, we formulate a long-term optimality gap minimization problem involving the joint optimization of newly sensed dataset size, computation frequency, communication bandwidth, and transmit power. By leveraging Lyapunov optimization, we develop an online optimization algorithm, where, at each iteration, the optimization variables are all derived in closed-form. Moreover, we prove that the proposed algorithm achieves its asymptotic optimal performance and conduct simulations to show the superiority of the proposed ISCC-enabled FEEL.
Yong Zhou 0006, Qiaochu An, Zhibin Wang 0003, Hangguan Shan, Yuanming Shi
IEEE Trans. Wirel. Commun.5
2026 Robust Information Bottleneck for Satellite Edge Inference Over MIMO Channel
Jielin Zhu, Jingyang Zhu, Youlong Wu, Ting Wang 0001, Yuanming Shi, Wei Chen 0002, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.6
2026 Satellite Federated Fine-Tuning for Foundation Models in Space Computing Power Networks
abstract
Advancements in artificial intelligence and low-earth orbit satellites have promoted the application of large remote sensing foundation models (FMs) for various downstream tasks. However, direct downloading of these models for fine-tuning on the ground is impeded by privacy concerns and limited bandwidth. Satellite federated learning (FL) offers a solution by enabling model fine-tuning directly on-board satellites and aggregating model updates without data downloading. Nevertheless, for large FMs, the computational capacity of satellites is insufficient to support effective on-board fine-tuning in traditional satellite FL frameworks. To address these challenges, we propose a satellite-ground collaborative federated fine-tuning framework. The key of the framework lies in how to reasonably decompose and allocate model components to alleviate insufficient on-board computation capabilities. During fine-tuning, satellites exchange intermediate results with ground stations or other satellites for forward propagation and back propagation, which brings communication challenges due to the special communication topology of space transmission networks, such as intermittent satellite-ground communication, short duration of satellite-ground communication windows, and unstable inter-orbit inter-satellite links. To reduce transmission delays, we further introduce tailored communication strategies that integrate both communication and computing resources. Specifically, we propose a parallel intra-orbit communication strategy, a topology-aware satellite-ground communication strategy, and a latency-minimization inter-orbit communication strategy to reduce space communication costs. Simulation results demonstrate significant reductions in training time to 33% of on-board training time.
Jingyang Zhu, Ting Wang 0001, Yuanming Shi, Chunxiao Jiang, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.4
2025 Structured IB: Improving Information Bottleneck with Structured Feature Learning
abstract
The Information Bottleneck (IB) principle has emerged as a promising approach for enhancing the generalization, robustness, and interpretability of deep neural networks, demonstrating efficacy across image segmentation, document clustering, and semantic communication. Among IB implementations, the IB Lagrangian method, employing Lagrangian multipliers, is widely adopted. While numerous methods for the optimizations of IB Lagrangian based on variational bounds and neural estimators are feasible, their performance is highly dependent on the quality of their design, which is inherently prone to errors. To address this limitation, we introduce Structured IB, a framework for investigating potential structured features. By incorporating auxiliary encoders to extract missing informative features, we generate more informative representations. Our experiments demonstrate superior prediction accuracy and task-relevant information preservation compared to the original IB Lagrangian method, even with reduced network size.
Youlong Wu, Dingzhu Wen, Yong Zhou 0006, Yuanming Shi
AAAI5
2025 Latency-Minimal Decentralized LAM Training with Looped Transformers in Heterogeneous LEO Satellite Constellations
abstract
The traditional approach of transmitting sensing data to ground stations for training large AI models (LAMs) is becoming increasingly impractical due to unstable ground-to-satellite links (GSLs) and growing privacy concerns. Fortunately, the rapid expansion of Low Earth Orbit (LEO) satellite constellations offers a new solution. The increasing number of LEO satellites constitutes a distributed computing network, enabling on-orbit computation without the need for data transmission to ground stations. However, the limited computational resources on satellites result in substantial latency when training edge LAMs directly on a single satellite, and the memory capacity is insufficient to accommodate large models for training. To address these challenges, we propose a latency-optimized decentralized training framework that enables collaborative in-orbit LAM pretraining across heterogeneous LEO satellites without the need for downloading data to the ground. In the proposed architecture, we adopt a looped Transformer architecture that reduces memory usage through cross-layer parameter sharing. To improve training efficiency, we introduce a hybrid parallel training strategy that combines intra-group pipeline parallelism and inter-group data parallelism through satellite grouping. Furthermore, to accommodate heterogeneous satellite capabilities, we formulate an optimization problem for workload balancing and develop a dynamic programming–based strategy to minimize overall system latency. Extensive experiments on a heterogeneous satellite testbed demonstrate that our method significantly outperforms existing baselines in both training speed and memory efficiency.
He Xian, Honggang Yuan, Ting Wang 0001, Yuanming Shi
GLOBECOM5
2025 Integrated Sensing-Communication-Computation for Movable Antennas Assisted Multi-Device Edge AI Inference
Dingzhu Wen, Min Fu 0003, Yong Zhou 0006, Yuanming Shi
GLOBECOM7
2025 Adaptive Task-Oriented Communication with Fairness Guarantees
abstract
Learning-based joint source-channel coding (JSCC) is widely used in task-oriented communication, which aims to extract and transmit only task-relevant information to improve communication efficiency. However, the learning-empowered algorithms in task-oriented communication may bring potential bias towards sensitive groups, and the adaptability to dynamic channel conditions still remains a challenge. To address these issues, we propose a task-oriented communication scheme that achieves efficient encoding and inference while preserving group fairness. Our approach leverages an information bottleneckbased framework that maximizes the task utility information while limiting the dependence of the inference result on the sensitive attribute and adopts a hypernetwork-parametrization mechanism to adapt to varying channel conditions. We also provide a theoretical bound for fairness guarantee and design a noise injection module to control the fairness-utility tradeoff. Experiments on benchmark datasets demonstrate the superiority of our framework in achieving a fairness-utility tradeoff and the adaptability to channel variations.
Songjie Xie, Yuanming Shi, Youlong Wu, Meixia Tao
ICC3
2025 MIMO Over-The-Air Federated Learning With Spiking Neural Network Via Lattice Code
abstract
Spiking neural networks (SNNs) have emerged as an energy-efficient alternative to the traditional artificial neural networks (ANNs) which are compute-intensive. This paper proposes a novel MIMO over-the-air federated learning scheme trained on SNNs using lattice code. Based on the lattice structure, we design a reliable transceiver with lattice quantizer that can combat the noise and interference from the devices. We further derive a convergence analysis of the proposed method considering the nondifferentiable spikes of SNNs. The experimental results verify that the proposed method is effective by showing that the proposed method can achieve comparable accuracy to the ideal benchmarks and outperform the existing approach by employing a small number of antennas at the server and devices. We also show that SNNs are$23.08 \times$more energy-efficient than ANNs.
Chenye Wang, Youlong Wu, Ting Wang 0001, Yuanming Shi
ICC4
2025 Hierarchical Federated Learning with Integrated Sensing-Communication-Computation Over Space-Air-Ground Integrated Networks
abstract
Federated learning has achieved significant advancements in edge artificial intelligence (AI) by addressing issues related to data privacy and communication overload. Moreover, hierarchical federated learning over space-air-ground integrated networks (FedSAG), which consists of low-Earth orbit (LEO) satellites, unmanned aerial vehicles (UAVs), and edge devices, aims to provide AI services in sparsely populated regions lacking ground communication infrastructure. However, previous studies have overlooked the essential sensing process required for acquiring training data, potentially compromising training efficiency and model accuracy. In this paper, we propose an integrated sensing-communication-computation (ISCC) enabled FedSAG system, which allows remote edge devices to collect data via wireless sensing and collaboratively train a global model without sharing local data. We then analyse the convergence of the ISCC-enabled FedSAG and formulate two optimization problems. The first aims to minimize sensing variance under energy and time constraints, while the second seeks to reduce transmission energy through optimal route selection between UAVs and LEO satellite. We reformulate the problems to a minimum spanning tree and propose a Chu-Liu-Edmonds algorithm based a two-stage optimization. Simulation results demonstrate that our proposed algorithm significantly enhances convergence performance and reduces energy consumption.
Zhanpeng Yang, Jingyang Zhu, Dingzhu Wen, Yuanming Shi, Wei Chen 0002
ICC5
2025 SAI: Latency-Aware Satellite Edge LAM Inference with Looped Transformer
abstract
The rapid advancements in computing and communication capabilities of Low Earth Orbit (LEO) satellites have made it feasible to execute complex and collaborative inorbit computation missions. Transformer-based large AI models (LAMs), known for their exceptional performance in in-context learning (ICL) and prompt-based reasoning, have attracted significant attention, providing powerful intelligence across sectors such as industry and aerospace. However, the significant parameter volume of LAMs poses a substantial challenge for direct deployment on satellites with constrained computing power and energy provision. To address this, the looped Transformer model reduces parameter requirements through layerwise parameter sharing, achieving performance comparable to vanilla Transformer-based LAMs in ICL tasks. Despite this efficiency, the limited and heterogeneous space-borne computing and storage capabilities complicate the orchestration for balanced workload allocation during multi-satellite cooperation. In this paper, we propose SAI, a collaborative multi-satellite space AI system that exploits the memory efficiency of the looped Transformer and the inherent parallelism in batch data processing. SAI enables accelerated on-satellite inference by integrating heterogeneous onboard resources and introducing a novel hybrid approach combining data and pipeline parallelism. This approach supports cross-satellite cooperation with parallelism planning and asynchronous inter-batch overlapping, significantly reducing inference latency and enhancing resource efficiency. Furthermore, SAI optimizes inference latency by formulating it as a shortest-path problem, effectively solved via Dijkstras algorithm. Extensive evaluations demonstrate SAIs superior performance in reducing inference latency and runtime memory usage compared to existing baselines.
Honggang Yuan, Yuning Jiang 0002, Xin Liu 0049, Yuanming Shi, Ting Wang 0001
ICC5
2025 ECMSA: Dual-Agent Learning-Based Edge Caching with Multi-Strategy Adaptation in Dynamic Environments
abstract
With the proliferation of mobile devices and IoT applications, edge caching has become vital for mitigating network congestion and enhancing user Quality of Experience (QoE). However, traditional caching policies, such as Least Frequently Used (LFU), First-In-First-Out (FIFO), and Least Recently Used (LRU), often struggle to perform effectively in highly dynamic and heterogeneous environments, particularly when content sizes vary significantly. Moreover, existing approaches, whether AI-driven or heuristic-based, typically adopt a single caching strategy, which inherently limits their flexibility and adaptability. To address these limitations, we propose ECMSA, a learning-based multi-strategy edge caching algorithm that integrates a reinforcement learning-driven proactive caching strategy with three conventional reactive caching strategies. Specifically, ECMSA operates in two stages: First, it generates four candidate cache lists—three derived from traditional caching policies (LFU, FIFO, LRU) and one produced by our self-attention-enhanced Deep Deterministic Policy Gradient (Atten-Actor DDPG)-based proactive caching strategy. Next, it employs another Atten-Actor DDPG agent to dynamically select the optimal strategy in real time, leveraging current state features. This dual-agent framework enables continuous learning and adaptation of caching decisions, effectively optimizing content placement and update policies in response to evolving user demands. Extensive experiments conducted on both synthetic and real-world datasets demonstrate that ECMSA achieves 15-17% higher cache-hit ratios and 16-22% lower latency than baseline methods under constrained cache capacities and diverse content sizes. Furthermore, ECMSA exhibits strong robustness and generalization ability, allowing it to rapidly adapt to unseen environments.
Ting Wang 0001, Lu Yang 0003, Yuanming Shi, Haibin Cai
ICPADS4
2025 Robust Multimodal Information Bottleneck for Satellite-to-Ground Task-Oriented Communication
abstract
In this paper, we study satellite-to-ground taskoriented communication for edge inference tasks, where a satellite extracts, fuses and encodes multimodal feature vectors and then sends them to a ground server under inevitable channel noise conditions for downstream processing. However, the multispectral and multi-resolution characteristics of multimodal satellite remote sensing data render traditional multimodal methods inapplicable. To reduce the data redundancy caused by the high-dimensional and complex multimodal vectors generated onboard while retaining key information and enhancing robustness against channel noise. We propose a Robust Multimodal Information Bottleneck (RMIB) framework which considers channel noise and communication bandwidth and introduces a new information bottleneck optimization objective. By applying this objective through end-to-end training, we optimize the feature extraction, fusion and encodes multimodal data into robust and effective feature vector in noisy communication environments by reducing redundancy and enhancing feature discrimination. To tackle the RMIB objective function, we derive a tractable variational upper bound using the Variational Information Bottleneck technique to overcome the computational intractability of mutual information. Experimental results demonstrate that our method not only outperforms baseline techniques in classification accuracy on three datasets but also enhances robustness against channel noise and reduces communication overhead.
Dingzhu Wen, Youlong Wu, Yuanming Shi, Ting Wang 0001
ISCC4
2025 Collaborative Multi-Device Edge Inference for Vision-Language Models with Speculative Decoding
abstract
Large language models (LLMs) have demonstrated remarkable success across various domains. However, the substantial computational and memory demands pose significant challenges for deploying LLMs at the network edge. To address this issue, the existing studies mainly focus on either reducing the size of LLMs or distributing LLMs across multiple devices. However, the former approach suffers from performance degradation, while the latter incurs high communication cost due to the transmission of high-dimensional intermediate features. To tackle these issues, we propose a novel collaborative framework to support vision-language model (VLM) inference at the network edge, expanding speculative decoding into a multi-device scenario. In this framework, each edge device first utilizes its small VLM to generate draft tokens in an auto-regressive manner. These tokens are then transmitted to an edge server, where they are corrected in parallel by a large VLM. By benefiting from speculative decoding, the number of calls to the large VLM is reduced without degrading the inference performance. Furthermore, we minimize the average latency of inference tasks by developing a deep reinforcement learning algorithm to optimize the number of draft tokens generated at each iteration. Simulation results confirm that the proposed algorithm achieves a lower average latency compared to other baselines.
Luteng Qiao, Jiawei Shao, Yong Zhou 0006, Yuanming Shi, Xuelong Li 0001, Khaled Ben Letaief
PIMRC4
2025 Topology-Aware Routing for Federated Learning Over Multi-Layer Satellite Networks
abstract
Recent advancements in space computing power networks, particularly the integration of onboard computing capabilities in Low Earth Orbit (LEO) satellites, have paved the way for federated learning (FL) in satellite networks. Despite its potential, satellite FL faces unique challenges, such as the dynamic nature of satellite networks and the instability of inter-orbit communication links, which complicate global model aggregation. To address these challenges, we explore FL over multi-layer satellite networks, incorporating LEO, Medium Earth Orbit (MEO), and Geostationary Earth Orbit (GEO) satellites. Specifically, by modeling the dynamic network as a series of time-varying graph snapshots, we propose a novel topology-aware FL framework. To optimize the aggregation routing in the multi-layer satellite network, we leverage the directed minimum spanning tree (DMST) problem in graph theory and introduce a communication-efficient satellite aggregation routing algorithm (CESAR), which effectively reduces communication overhead and aggregation delays, ensuring efficient training and model updates across the satellite network. Extensive experimental results validate the efficacy of the proposed framework, demonstrating its potential to overcome the inherent challenges of satellite FL and significantly advance the capabilities of multi-layer satellite networks.
Ruanjun Li, Jingyang Zhu, Yijie Mao, Yuanming Shi, Ting Wang 0001, Chunxiao Jiang
WCNC4
2025 Joint Source and Channel Coding for Multi-Modal Satellite-to-Ground Semantic Communications
abstract
This paper presents a novel Joint Source and Channel Coding (JSCC) method for semantic communication to enhance the communication efficiency for transmitting high-resolution multi-modal data from Low Earth Orbit (LEO) satellites to ground stations. On the satellite, a JSCC encoder consisting of neural networks (NNs) is utilized to map the input multi-modal data into a common signal, while an NN-based JSCC decoder at the ground station reconstructs the original input from the received signal. Throughout this process, the common semantic information among each modality is learned to optimize coding space, and a robust coding scheme specific to the satellite downlink channel is developed between the encoder and decoder. Experiments have demonstrated that the proposed method outperforms existing signal modality JSCC approaches, achieving multi-modal transmission with significantly smaller coding space and thereby reducing communication overhead.
Yanbo Yin, Dingzhu Wen, Youlong Wu, Yuanming Shi
WCNC5
2025 Satellite edge artificial intelligence with large models: architectures and technologies
Yuanming Shi, Jingyang Zhu, Chunxiao Jiang, Linling Kuang, Khaled Ben Letaief
Sci. China Inf. Sci.1
2025 Joint Task Offloading and Energy Harvesting in Space-Air-Ground-Integrated MEC Networks
abstract
The rapid development of space-air–ground integrated technology has laid a foundation for achieving wide area wireless communication coverage, but its applicability is limited due to the limited computational capacity and energy storage of the terminal devices. This article thus proposes a space-air–ground integrated mobile-edge computing (MEC) task offloading and computing resource allocation (SIMOC) algorithm. First, a space-air–ground integrated MEC system is designed in which the terminal devices in the system can offload tasks to air- and space-based servers while performing energy harvesting. Second, an optimization objective of maximizing the task execution benefit minus the sum cost of task offloading and execution is established, for which the problem is transformed into a time-slot-based minimization problem of queue-stability minus revenue using Lyapunov optimization. Finally, the energy harvesting and task offloading subproblems of the queue-stability-minus-revenue problem are solved, respectively, to maximize the total revenue while ensuring device stability. Simulation results show that the SIMOC algorithm can reduce the all-task completion time by up to 98.16% compared with only local task execution algorithm in the absence of newly added tasks and shows good performance in handling newly added tasks. Meanwhile, the SIMOC algorithm has better performance compared to the particle swarm optimization algorithm.
Yuexia Zhang 0001, Yunong Yang, Sheng Wu 0001, Yuanming Shi, Jiangzhou Wang
IEEE Internet Things J.5
2025 Federated Edge Learning for 6G: Foundations, Methodologies, and Applications
abstract
Artificial intelligence (AI) is envisioned to be natively integrated into the sixth-generation (6G) mobile networks to support a diverse range of intelligent applications. Federated edge learning (FEEL) emerges as a vital enabler of this vision by leveraging the sensing, communication, and computation capabilities of geographically dispersed edge devices to collaboratively train AI models without sharing raw data. This article explores the pivotal role of FEEL in advancing both the “wireless for AI” and “AI for wireless” paradigms, thereby facilitating the realization of scalable, adaptive, and intelligent 6G networks. We begin with a comprehensive overview of learning architectures, models, and algorithms that form the foundations of FEEL. We, then, establish a novel task-oriented communication principle to examine key methodologies for deploying FEEL in dynamic and resource-constrained wireless environments, focusing on device scheduling, model compression, model aggregation, and resource allocation. Furthermore, we investigate the domain-specific optimizations of FEEL to facilitate its promising applications, ranging from wireless air-interface technologies to mobile and the Internet of Things (IoT) services. Finally, we highlight key future research directions for enhancing the design and impact of FEEL in 6G.
Meixia Tao, Yong Zhou 0006, Yuanming Shi, Jianmin Lu, Shuguang Cui, Jianhua Lu, Khaled Ben Letaief
Proc. IEEE3
2025 Coded Computing for Multi-Cluster Distributed Computations
abstract
Distributed computing, which leverages distributed storage and computing resources, is a promising paradigm for handling large-scale computational tasks. However, its potential is often hindered by high communication latency due to limited network bandwidth. In this paper, we study the computation-communication tradeoff of multi-cluster MapReduce systems where a central server connects to multiple clusters, each comprising a set of workers that jointly perform a MapReduce task. Workers can exchange information directly within their cluster (inner-cluster communication) or indirectly through the central server (cross-cluster communication). To reduce the communication load, we propose a nested coded distributed computing (CDC) scheme that is feasible for the heterogeneous scenario where different clusters could have arbitrary numbers of workers and computation loads. It is shown that our scheme can greatly reduce communication load compared to all existing schemes, and could achieve the optimal cross-cluster communication load. In addition, the proposed scheme can significantly reduce the computational complexity of the conventional CDC schemes, whose computational complexity exponentially increases with the computation load.
Youlong Wu, Haoyang Hu, Xiyu Song 0001, Shuai Ma 0002, Yuanming Shi
IEEE Trans. Commun.6
2025 Brain-Inspired Decentralized Satellite Learning in Space Computing Power Networks
Peng Yang 0027, Ting Wang 0001, Haibin Cai, Yuanming Shi, Chunxiao Jiang, Linling Kuang
IEEE Trans. Mob. Comput.4
2025 GIRP: Energy-Efficient QoS-Oriented Microservice Resource Provisioning via Multi-Objective Multi-Task Reinforcement Learning
abstract
Microservice architecture has revolutionized web service development by facilitating loosely coupled and independently developable components distributed as containers or virtual machines. While existing studies emphasize end-to-end latency, this paper investigates energy-efficient quality-of-service (QoS)-oriented microservice provisioning, focusing on both QoS satisfaction and power consumption (PC) conservation. We propose the Green and Intelligent Resource Provision (GIRP) architecture, integrating a data-driven energy-latency-aware resource allocation and scheduling manager to balance latency and PC. To reconcile the trade-offs involved, a dual-objective optimization problem is formulated to minimize latency and energy use by selecting proper servers, allocating CPU cores, and determining service replicas. To address challenges with discrete variables, dual objectives, and implicit mappings, we leverage a model-free deep deterministic policy gradient-based reinforcement learning algorithm. Specifically, we develop a multi-task agent via the Multi-gate Mixture-of-Experts model to simultaneously make two separate actions regarding CPU core numbers and service replica numbers, followed by a single-task agent to determine service scheduling. Extensive experiments on the DeathStarBenchmark testbed validate GIRP’s effectiveness, demonstrating approximately 52% resource savings and a 43% reduction in PC compared to leading methods like Sinan, Firm, and heuristic-based algorithms. These results highlight GIRP’s capability to optimize microservice orchestration by balancing end-to-end latency and power efficiency.
Honggang Yuan, Ting Wang 0001, Min Fu 0003, Yuanming Shi
IEEE Trans. Mob. Comput.4
2025 Hierarchical Learning and Computing Over Space-Ground Integrated Networks
abstract
Space-ground integrated networks hold great promise for providing global connectivity, particularly in remote areas where large amounts of valuable data are generated by Internet of Things (IoT) devices, but lacking terrestrial communication infrastructure. The massive data is conventionally transferred to the cloud server for centralized artificial intelligence (AI) models training, raising huge communication overhead and privacy concerns. To address this, we propose a hierarchical learning and computing framework, which leverages the low-latency characteristic of low-earth-orbit (LEO) satellites and the global coverage of geostationary-earth-orbit (GEO) satellites, to provide global aggregation services for locally trained models on ground IoT devices. Due to the time-varying nature of satellite network topology and the energy constraints of LEO satellites, efficiently aggregating the received local models from ground devices on LEO satellites is highly challenging. By leveraging the predictability of inter-satellite connectivity, modeling the space network as a directed graph, we formulate a network energy minimization problem for model aggregation, which turns out to be aDirected Steiner Tree (DST)problem. We propose a topology-aware energy-efficient routing (TAEER) algorithm to solve theDSTproblem by finding a minimum spanning arborescence on a substitute directed graph. Extensive simulations under real-world space-ground integrated network settings demonstrate that the proposed TAEER algorithm significantly reduces energy consumption and outperforms benchmarks.
Jingyang Zhu, Yuanming Shi, Yong Zhou 0006, Chunxiao Jiang, Linling Kuang
IEEE Trans. Mob. Comput.2
2025 Dynamic UAV-Assisted Cooperative Edge AI Inference
abstract
Deploying intelligent service and executing inference tasks in the proximity of the edge enable models to access enormous real-time data generated by the edge devices. However, the dilemma of fulfilling service demands with limited resources at edge devices impairs the efficacy of conventional data-oriented communication systems. To achieve a better trade-off between inference accuracy and communication overhead, in this paper, we propose a dynamic unmanned aerial vehicle (UAV)-assisted cooperative edge inference system, where a UAV acts as an edge server to aggregate the wide-view features from mobile sensors through Over-the-Air computation (AirComp) to complete the inference task cooperatively. Discriminant gain, an effective indicator for the inference accuracy, is adopted to realize task-oriented design. To exploit channel diversity and data diversity in the multi-device cooperative edge inference system, we maximize the discriminant gain of the AirComp feature aggregation by jointly optimizing the UAV trajectory and the power allocation policy with respect to the different important levels of feature dimensions. An alternating algorithm and a successive convex approximation (SCA)-based method are then proposed to solve the optimization problem. Numerical simulations further validate the efficacy of the proposed design compared to the baselines.
Jingfeng Huang, Lixiang Lian, Dingzhu Wen, Yong Zhou 0006, Fuzhai Wang, Weichang Wang, Yuanming Shi
IEEE Trans. Wirel. Commun.7
2025 Federated Fine-Tuning for Pre-Trained Foundation Models Over Wireless Networks
abstract
Pre-trained foundation models (FMs), with extensive number of neurons, are key to advancing next-generation intelligence services, where personalizing these models requires massive amount of task-specific data and computational resources. The prevalent solution involves centralized processing at the edge server, which, however, raises privacy concerns due to the transmission of raw data. Instead, federated fine-tuning (FedFT) is an emerging privacy-preserving fine-tuning (FT) paradigm for personalized pre-trained foundation models. In particular, by integrating low-rank adaptation (LoRA) with federated learning (FL), federated LoRA enables the collaborative FT of a global model with edge devices, achieving comparable learning performance to full FT while training fewer parameters over distributed data and preserving raw data privacy. However, the limited radio resources and computation capabilities of edge devices pose significant challenges for deploying 3 LoRA over wireless networks. To this paper, we propose a split federated LoRA framework, which deploys the computationally-intensive encoder of a pre-trained model at the edge server, while keeping the embedding and task modules at the edge devices. The information exchanges between these modules occur over wireless networks. Building on this split framework, the paper provides a rigorous analysis of the upper bound of the convergence gap for the wireless federated LoRA system. This analysis reveals the weighted impact of the number of edge devices participating in FedFT over all rounds, motivating the formulation of a long-term upper bound minimization problem. To address the long-term constraint, we decompose the formulated long-term mixed-integer programming (MIP) problem into sequential sub-problems using the Lyapunov technique. We then develop an online algorithm for effective device scheduling and bandwidth allocation. Simulation results demonstrate the effectiveness of the proposed online algorithm in enhancing learning performance.
Yong Zhou 0006, Yuanming Shi, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.3
2025 Learning to Beamform for Integrated Sensing and Communication: A Graph Neural Network With Implicit Projection Approach
abstract
Integrated sensing and communication (ISAC), as an important usage scenario of 6G, is capable of seamlessly integrating wireless sensing and communication for their mutual benefit. Taking full advantage of ISAC heavily relies on effectively solving resource allocation problems, which, however, are generally high-dimensional and non-convex, resulting in the optimization-based algorithms exhibiting high computation complexity and the traditional learning-based algorithms returning infeasible solutions. In this paper, we consider an ISAC scenario featured by multiple communication users and multiple sensing targets, aiming to develop an efficient and scalable algorithm that optimizes the radar transmit beampattern under the communication performance constraint. To this end, we propose a graph neural network (GNN) with implicit projection framework, where GNN captures the intricate interactions between communication users and sensing targets and meanwhile enables the joint optimization of communication and sensing beamforming matrices, and the projection module is applied to ensure the feasibility of the beamforming matrices design. Via capturing the permutation equivalence for communication matrices and the permutation invariance for the sensing matrix, the scalability of the proposed algorithm is guaranteed. Simulation results show that the proposed algorithm significantly reduces the computation complexity compared to the baselines, and achieves excellent algorithmic scalability and constraint satisfaction.
Yong Zhou 0006, Yuanming Shi, Nan Cheng 0001
IEEE Trans. Wirel. Commun.4
2024 Federated Low-Rank Adaptation for Large Language Model Fine-Tuning Over Wireless Networks
abstract
Low-rank adaptation (LoRA) is an emerging fine-tuning method for personalized large language models (LLMs) due to its capability of achieving comparable learning performance to full fine-tuning by training a much smaller number of parameters. Federated fine-tuning (FedFT) combines LoRA with federated learning (FL) to enable collaborative fine-tuning of a global model with edge devices, leveraging distributed data while ensuring privacy. However, limited radio resources and computation capabilities of edge devices pose critical challenges on deploying FedFT over wireless networks. In this paper, we propose a split FedFT framework to separately deploy the computationally-intensive encoder of a pre-trained model at the edge server while reserving the embedding and the task modules at the edge devices, where the information exchanges between these modules are carried out over wireless networks. By exploiting the low-rank property of LoRA, the proposed FedFT framework reduces communication overhead by aggregating the gradient of the task module with respect to the output of a low-rank matrix. To enhance learning performance under stringent resource constraints, we formulate a joint device scheduling and bandwidth allocation problem while considering average transmission delay. By applying the Lyapunov technique, we decompose the formulated long-term mixed-integer programming (MIP) problem into sequential subproblems, followed by developing an online algorithm for effective device scheduling and bandwidth allocation. Simulation results demonstrate the effectiveness of our proposed online algorithm in enhancing learning performance.
Yong Zhou 0006, Yuanming Shi, Khaled Ben Letaief
GLOBECOM3
2024 Dynamic Communication in Multi-Agent Reinforcement Learning via Information Bottleneck
abstract
Effective information sharing is essential for multi-agent systems to execute cooperative tasks successfully. Typically, agents within such systems are either stationary or possess unrestricted communication ranges. However, in more complex scenarios where agent mobility is introduced, the communication network’s topology becomes dynamic over time. This dynamism can result in partial communication unreachability among certain agents. Consequently, striking a balance between minimizing overall communication overhead and optimizing task performance becomes a formidable challenge. In this paper, we address the issue of dynamic communication in multi-agent systems. We propose a novel approach that leverages the principle of information bottleneck theory to develop a multi-mean field multi-agent reinforcement learning algorithm called MMIB. Through a series of experiments, we demonstrate the effectiveness of our proposed algorithm in reducing communication overhead while maintaining task performance at a level comparable to other state-of-the-art multi-agent reinforcement learning algorithms.
Jiawei You, Youlong Wu, Dingzhu Wen, Yong Zhou 0006, Yuning Jiang 0002, Yuanming Shi
GLOBECOM6
2024 Satellite Federated Fine-Tuning for Foundation Models: Architecture Design and System Optimization
abstract
With the surge in the number of low earth orbit (LEO) satellites, continuous research has emerged on using satellite data to train artificial intelligence models. On one hand, traditional centralized training on the ground is not feasible due to privacy concerns and limited bandwidth for downloading raw satellite data. On the other hand, due to the limited energy and computational capability of satellites, training directly on satellites suffers from prolonged latency, especially for large models. To alleviate these issues, we propose a novel satellite-ground collaborative federated fine-tuning architecture, where ground stations (GSs) and satellites collaboratively train a global model without the need for data downloads. In this proposed architecture, satellites serve as edge devices and the ground server serves as a coordinator. However, the short satellite-ground communication windows caused by the high mobility of satellites and the substantial intra-orbit data transmission bring special challenges to the transmission process of federated edge learning. To tackle these challenges, we carefully design the satellite-ground collaborative fine-tuning architecture and utilize an optimized ring all-reduce algorithm and network flow algorithm to enhance the intra-orbit and ground-satellite transmissions, respectively. Experimental results demonstrate that our proposed architecture significantly reduces the training time by 40% compared to training solely on satellite.
Peng Yang 0027, Jingyang Zhu, Dingzhu Wen, Ting Wang 0001, Yong Zhou 0006, Yuanming Shi, Chunxiao Jiang
GLOBECOM7
2024 Over-the-Air Computation Assisted Federated Learning with Progressive Training
abstract
Federated learning (FL) with progressive training is a promising privacy-preserving and communication-efficient framework for edge intelligence applications. Specifically, by partitioning the global model into multiple sub-models and dividing the FL training into multiple stages, FL with progressive training enables the gradual training of a large model, thereby significantly reducing the transmission overhead without compromising learning performance. However, implementing FL with progressive training over wireless networks is hindered by the limited radio and energy resources. To address these issues, we adopt over-the-air computation (AirComp) to support FL with progressive training over wireless networks. By balancing the tradeoff between the AirComp transmission distortion and the transition efficiency of progressive training, we formulate a mixed-integer optimization problem with energy and power constraints, which is further decomposed into several subproblems via Lyapunov optimization. Subsequently, we develop a low computational-complexity algorithm that jointly optimizes transmit power, receive beamforming, and transition indicator in an alternating manner. Simulation results demonstrate the effectiveness of our optimization algorithm in improving the learning performance of the considered FL system.
Qiaochu An, Zhibin Wang 0003, Yuanming Shi, Yong Zhou 0006
ICC5
2024 Microservice Deployment for Satellite Edge AI Inference via Deep Reinforcement Learning
abstract
Artificial intelligence (AI) is critical in evolving 5G and developing 6G networks, running on edge devices, and solving resource management challenges. The burgeoning number of edge devices draws attention to the potential of low-earth orbit (LEO) satellite networks with their onboard computing capabilities for edge inference. This paper explores LEO scenarios where multiple remote sensing edge AI inference tasks concurrently process data from a single source. However, due to there being parts with the same functions between different AI applications, traditional monolithic edge AI architecture must be deployed repeatedly and falls short in efficiently harnessing the heterogeneous resources of LEO satellite networks. To solve this problem, we utilize the microservice architecture to decouple a single AI application into several independent microservices to reuse these same functions. However, due to the high latency caused by multiple microservices’ communication, we need to design a deployment strategy to fully utilize resources to reduce the service latency. We present a microservice deployment model to minimize the total service latency across all AI applications and meet resource constraints with the constraints of hardware, energy, and memory limitations. This latency optimization problem is rewritten as a Markov decision process (MDP) to effectively deal with the challenge posed by the time-varying transmission rate caused by satellite mobility. To increase the training data utilization, we employ a Proximal Policy Optimization (PPO) based reinforcement learning algorithm to meet the dynamic environment challenge. Finally, we obtain a sub-optimal solution with minimal accuracy loss and an acceptable solution time.
Hei Victor Cheng, Zhanpeng Yang, Xin Liu 0049, Yuning Jiang 0002, Yong Zhou 0006, Yuanming Shi
PIMRC7
2024 Delay Minimization for NOMA-Assisted Federated Learning
abstract
Federated learning (FL) enables multiple users to collaboratively train a shared model while protecting user privacy. In this paper, we investigate the transmission delay minimization problem for non-orthogonal multiple access (NOMA)-assisted FL. We analyze the convergence rate of heterogeneous quantized FL to demonstrate that the minimum quantization level among scheduled users is crucial in controlling the trade-off between the number of training rounds and the transmission delay of each round. Based on the convergence analysis, we formulate a delay minimization problem for NOMA-assisted FL and propose a communication-efficient heterogeneous compression NOMA scheme for FL. Subsequently, we develop a block coordinate descent (BCD)-based algorithm that jointly optimizes the sub channel allocation, power allocation, and quan-tization level for each scheduled user. Results reveal that our proposed algorithm significantly reduces the transmission delay while achieving the same learning performance compared with conventional FL algorithms.
Dong Zheng 0003, Zhibin Wang 0003, Qiaochu An, Yuanming Shi, Yong Zhou 0006
WCNC5
2024 RIS-Assisted Multi-Device Edge AI Inference
abstract
In this paper, we propose a multi-device co-inference system based on a task-oriented over-the-air computation (Air-Comp) via reconfigurable intelligent surface (RIS). Specially, local feature vectors extracted from the real-time noisy sensory data on devices are aggregated over-the-air by exploiting the waveform superposition in a multi-user channel. Then the aggregated features received at the server are fed into an inference model for decision making or control of actuators. Based on the proposed multi-device co-inference system, we jointly optimize the receive signal strength of the device, the beamforming vector, and RIS phase shifts to suppress the sensing and channel noise and maximize the inference accuracy. To solve the problem, we first transform the original problem into a convex difference (d.c.) problem, and convert the d.c. problem from the complex domain to the real domain. Then, we propose a successive convex approximation based approach to solve the problem in the real domain. With the supportive data and results from the application of human motion recognition, we show the proposed scheme achieves a higher inference accuracy then the conventional approaches.
Yijie Mao, Dingzhu Wen, Yong Zhou 0006, Yuanming Shi
WCNC5
2024 Latency-Aware Microservice Deployment for Edge AI Enabled Video Analytics
abstract
Video analytics plays a pivotal role in public safety (e.g., criminal suspect detection, traffic flow count, and illegal parking management), which assists the polices in monitoring all anomalous events in the street. In this paper, we consider the scenario with multiple video analytics applications from a single video stream. However, traditional monolithic architecture based video analytics applications shall seriously increase the response latency due to the resource contention of repetitive components. Therefore, we utilize the microservice architecture based video analytics (MAVA) to share the universal microser-vices in different applications, which shall decrease the response latency by reducing the computation load and increasing the resource utilization. To further achieve fast and accurate video analytics, the video analytics microservices are deployed in the edge closing to the cameras and users, and artificial intelligence (AI) methods are used in the microservices to realize specified functions. Therefore, an edge AI enabled MAVA (EAI-MAVA) architecture is proposed to achieve accurate video analytics in real-time. Furthermore, we formulate a microservice deployment problem to determine the location of each microservice in EAI-MAVA, which minimizes the response latency of all applications by considering the resource demands of microservices and the resource constraints of heterogeneous edge devices. Finally, a greedy-based heuristic algorithm is proposed to solve the non-convex microservice deployment problem, which obtains a sub-optimal solution with small loss of accuracy and reduces the solution time obviously.
Zhanpeng Yang, Xin Liu 0049, Dingzhu Wen, Yong Zhou 0006, Yuanming Shi
WCNC6
2024 Federated Reinforcement Learning for Electric Vehicles Charging Control on Distribution Networks
abstract
With the growing popularity of electric vehicles (EVs), maintaining power grid stability has become a significant challenge. To address this issue, EV charging control strategies have been developed to manage the switch between vehicle-to-grid (V2G) and grid-to-vehicle (G2V) modes for EVs. In this context, multiagent deep reinforcement learning (MADRL) has proven its effectiveness in EV charging control. However, existing MADRL-based approaches fail to consider the natural power flow of EV charging/discharging in the distribution network and ignore driver privacy. To deal with these problems, this article proposes a novel approach that combines multi-EV charging/discharging with a radial distribution network (RDN) operating under optimal power flow (OPF) to distribute power flow in real time. A mathematical model is developed to describe the RDN load. The EV charging control problem is formulated as a Markov decision process (MDP) to find an optimal charging control strategy that balances V2G profits, RDN load, and driver anxiety. To effectively learn the optimal EV charging control strategy, a federated deep reinforcement learning algorithm named FedSAC is further proposed. Comprehensive simulation results demonstrate the effectiveness and superiority of our proposed algorithm in terms of the diversity of the charging control strategy, the power fluctuations on RDN, the convergence efficiency, and the generalization ability.
Junkai Qian, Yuning Jiang 0002, Xin Liu 0049, Ting Wang 0001, Yuanming Shi, Wei Chen 0002
IEEE Internet Things J.6
2024 Over-the-Air Computation for 6G: Foundations, Technologies, and Applications
abstract
The rapid advancement of artificial intelligence technologies has given rise to diversified intelligent services, which place unprecedented demands on massive connectivity and gigantic data aggregation. However, the scarce radio resources and stringent latency requirement make it challenging to meet these demands. To tackle these challenges, over-the-air computation (AirComp) emerges as a potential technology. Specifically, AirComp seamlessly integrates the communication and computation procedures through the superposition property of multiple-access channels, which yields a revolutionary multiple-access paradigm shift from “compute-after-communicate” to “compute-when-communicate”. By this means, AirComp enables spectral-efficient and low-latency wireless data aggregation by allowing multiple devices to occupy the same channel for transmission. In this paper, we aim to present the recent advancement of AirComp in terms of foundations, technologies, and applications. The mathematical form and communication design are introduced as the foundations of AirComp, and the critical issues of AirComp over different network architectures are then discussed along with the review of existing literature. The technologies employed for the analysis and optimization on AirComp are reviewed from the information theory and signal processing perspectives. Moreover, we present the existing studies that tackle the practical implementation issues in AirComp systems, and elaborate the applications of AirComp in Internet of Things and edge intelligent networks. Finally, potential research directions are highlighted to motivate the future development of AirComp.
Zhibin Wang 0003, Yapeng Zhao, Yong Zhou 0006, Yuanming Shi, Chunxiao Jiang, Khaled Ben Letaief
IEEE Internet Things J.4
2024 Latency Minimization for Wireless Federated Learning With Heterogeneous Local Model Updates
abstract
In this article, we study the latency minimization problem for a wireless federated learning (FL) system with heterogeneous computation capability, where different edge devices perform different numbers of local model updates in each communication round. We formulate a total latency minimization problem with probabilistic device selection, taking into account both the communication and computation latency in the whole FL procedure. However, it is highly challenging to optimally solve this problem due to the coupling issues of model convergence and latency minimization problem caused by the heterogeneity of local model updates. Through convergence analysis, we reveal that decoupling the resource allocation variables from the model convergence is essential to reduce the problem to a single-round latency minimization problem. To solve this simplified problem, we propose an alternating optimization scheme to jointly consider communication and computation resource allocation and mitigate the straggler effect. We prove that the resulting subproblems, i.e., bandwidth and computation capacity allocation, are both convex and can be optimally solved in closed form, respectively. Simulation results show that compared with the baseline scheme that allocates the communication and computation resources equally across edge devices, the proposed scheme can achieve up to 47.04% single-round latency reduction.
Jingyang Zhu, Yuanming Shi, Min Fu 0003, Yong Zhou 0006, Youlong Wu, Liqun Fu 0001
IEEE Internet Things J.2
2024 Over-the-Air Federated Learning and Optimization
abstract
Federated learning (FL), as an emerging distributed machine learning paradigm, allows a mass of edge devices to collaboratively train a global model while preserving privacy. In this tutorial, we focus on FL via over-the-air computation (AirComp), which is proposed to reduce the communication overhead for FL over wireless networks at the cost of compromising in the learning performance due to model aggregation error arising from channel fading and noise. We first provide a comprehensive study on the convergence of AirComp-based FEDAVG (AIRFEDAVG) algorithms under both strongly convex and non-convex settings with constant and diminishing learning rates in the presence of data heterogeneity. Through convergence and asymptotic analysis, we characterize the impact of aggregation error on the convergence bound and provide insights for system design with convergence guarantees. Then we derive convergence rates for AIRFEDAVG algorithms for strongly convex and non-convex objectives. For different types of local updates that can be transmitted by edge devices (i.e., local model, gradient, and model difference), we reveal that transmitting local model in AIRFEDAVG may cause divergence in the training procedure. In addition, we consider more practical signal processing schemes to improve the communication efficiency and further extend the convergence analysis to different forms of model aggregation error caused by these signal processing schemes. Extensive simulation results under different settings of objective functions, transmitted local information, and communication schemes verify the theoretical conclusions.
Jingyang Zhu, Yuanming Shi, Yong Zhou 0006, Chunxiao Jiang, Wei Chen 0002, Khaled Ben Letaief
IEEE Internet Things J.2
2024 On Exploiting Network Topology for Hierarchical Coded Multi-Task Learning
abstract
Distributed multi-task learning (MTL) is a learning paradigm where distributed users simultaneously learn multiple tasks by leveraging the correlations among tasks. However, distributed MTL suffers from a more severe communication bottleneck than single-task learning as more than one models need to be transmitted in the communication phase. To address this issue, we investigate the hierarchical MTL system where distributed users wish to jointly learn different learning models orchestrated by a central server with the help of multiple relays. We propose a coded distributed computing scheme for hierarchical MTL systems that jointly exploits the network topology and relays’ computing capability to create coded multicast opportunities to improve communication efficiency. We theoretically prove that the proposed scheme can significantly reduce the communication loads both in the uplink and downlink transmissions between relays and the server. To further illustrate the optimality of the proposed scheme, we derive information-theoretic lower bounds on the minimum uplink and downlink communication loads and prove that the gaps between achievable upper bounds and lower bounds are within the minimum number of connected users among all relays. In particular, when the network topology can be delicately designed, the proposed scheme can achieve the information-theoretic optimal communication loads. Experiments on real-world datasets show that our proposed scheme can greatly reduce the overall training time compared to the conventional hierarchical MTL scheme.
Haoyang Hu, Minquan Cheng, Shuai Ma 0002, Yuanming Shi, Youlong Wu
IEEE Trans. Commun.5
2024 Collaborative Edge AI Inference Over Cloud-RAN
abstract
In this paper, a cloud radio access network (Cloud-RAN) based collaborative edge AI inference architecture is proposed. Specifically, geographically distributed devices capture real-time noise-corrupted sensory data samples and extract the noisy local feature vectors, which are then aggregated at each remote radio head (RRH) to suppress sensing noise. To realize efficient uplink feature aggregation, we allow each RRH receives local feature vectors from all devices over the same resource blocks simultaneously by leveraging an over-the-air computation (AirComp) technique. Thereafter, these aggregated feature vectors are quantized and transmitted to a central processor (CP) for further aggregation and downstream inference tasks. Our aim in this work is to maximize the inference accuracy via a surrogate accuracy metric called discriminant gain, which measures the discernibility of different classes in the feature space. The key challenges lie on simultaneously suppressing the coupled sensing noise, AirComp distortion caused by hostile wireless channels, and the quantization error resulting from the limited capacity of fronthaul links. To address these challenges, this work proposes a joint transmit precoding, receive beamforming, and quantization error control scheme to enhance the inference accuracy. Extensive numerical experiments demonstrate the effectiveness and superiority of our proposed optimization algorithm compared to various baselines.
Dingzhu Wen, Guangxu Zhu, Qimei Chen, Kaifeng Han, Yuanming Shi
IEEE Trans. Commun.6
2024 Energy-Efficient Optimal Mode Selection for Edge AI Inference via Integrated Sensing-Communication-Computation
abstract
Existing edge inference methods only consider one paradigm, i.e., one of on-device inference, on-server inference, or edge-device cooperative inference. Each paradigm has its pros and cons as well as dominant application scopes. For example, the on-device paradigm is the best choice when the inference task is not computationally intensive, the on-server paradigm is suitable if the communication capacity is strong, and the edge-device cooperative mode should be selected in the scenario of weak on-device communication and computation. However, each paradigm suffers from poor performance if deployed outside of its application scope, thus leading to limited potential and flexibility. This paper proposes an edge AI inference framework, which makes the first attempt to jointly consider the three modes for making full use of their benefits. In addition, sensing for data acquisition is enabled at both the edge server and the device. This can effectively improve the inference accuracy with rich information on the target area from two different views. On the other hand, energy cost minimization turns out to be a key target all over the world and a significant issue in wireless networks. To this end, we target minimizing the system energy cost under a given inference accuracy guarantee and other network resource constraints, by coordinating sensing, communication, and computation in different modes. By optimally solving the optimization problem, an integrated sensing-communication-computation (ISCC) based task-oriented mode selection scheme is proposed. A practical ISCC platform is built and extensive experiments are conducted to verify our theoretical analysis.
Dingzhu Wen, Qimei Chen, Guangxu Zhu, Yuanming Shi
IEEE Trans. Mob. Comput.6
2024 Online Optimization for Over-the-Air Federated Learning With Energy Harvesting
abstract
Federated learning (FL) is recognized as a promising privacy-preserving distributed machine learning paradigm, given its potential to enable collaborative model training among distributed devices without sharing their raw data. However, supporting FL over wireless networks confronts the critical challenges of periodically executing power-hungry training tasks on energy-constrained devices and transmitting high-dimensional model updates over spectrum-limited channels. In this paper, we reap the benefits of both energy harvesting (EH) and over-the-air computation (AirComp) to alleviate the battery limitation by harvesting ambient energy to support both the training and transmission of local models, and to achieve low-latency model aggregation by concurrently transmitting local gradients via AirComp. We characterize the convergence of the proposed FL by deriving an upper bound of the expected optimality gap, revealing that the convergence depends on the accumulated errors due to partial device participation and model distortion, both of which further depend on dynamic energy levels. To accelerate the convergence, we formulate a joint AirComp transceiver design and device scheduling problem, which is then tackled by developing an efficient Lyapunov-based online optimization algorithm. Simulations demonstrate that, by appropriately scheduling devices and allocating energy across multiple communication rounds, our proposed algorithm achieves a much better learning performance than benchmarks.
Qiaochu An, Yong Zhou 0006, Zhibin Wang 0003, Hangguan Shan, Yuanming Shi, Mehdi Bennis
IEEE Trans. Wirel. Commun.5
2024 Federated Learning via Unmanned Aerial Vehicle
abstract
Federated learning (FL) has emerged as a promising alternative to centralized machine learning for exploiting large amounts of data generated by networks while ensuring data privacy. Unlike previous FL works that rely on terrestrial base stations, this paper studies an unmanned aerial vehicle (UAV)-assisted FL system where a UAV collects local models from distributed ground devices. By leveraging the UAV’s high altitude and mobility, it can proactively establish short-distance line-of-sight links with devices to mitigate the communication straggler effect and improve communication efficiency in FL. Specifically, we present the convergence analysis of FL without convexity assumptions, demonstrating the effect of device scheduling on the global gradients. Based on the derived convergence bound, we aim to minimize the completion time of FL training by jointly optimizing device scheduling, UAV trajectory, and time allocation. This problem explicitly incorporates the devices’ energy budgets, dynamic channel conditions, and convergence accuracy of FL constraints. Despite the non-convexity of the formulated problem, we exploit its structure to decompose it into two sub-problems and further derive the closed-form solutions via the Lagrange dual ascent method. Simulation results show that the proposed design significantly improves the tradeoff between completion time and test accuracy compared to existing benchmarks.
Min Fu 0003, Yuanming Shi, Yong Zhou 0006
IEEE Trans. Wirel. Commun.2
2024 Features Disentangled Semantic Broadcast Communication Networks
abstract
Single-user semantic communications have attracted extensive research recently, but multi-user semantic broadcast communication (BC) is still in its infancy. In this paper, we propose a practical robust features-disentangled multi-user semantic BC framework, where the transmitter includes a feature selection module and each user has a feature completion module. Instead of broadcasting all extracted features, the semantic encoder extracts the disentangled semantic features, and then only the users’ intended semantic features are selected for broadcasting, which can further improve the transmission efficiency. Within this framework, we further investigate two information-theoretic metrics, including the ultimate compression rate under both the distortion and perception constraints, and the achievable rate region of the semantic BC. Furthermore, to realize the proposed semantic BC framework, we design a lightweight robust semantic BC network by exploiting a supervised autoencoder (AE), which can controllably disentangle sematic features. Moreover, we design the first hardware proof-of-concept prototype of the semantic BC network, where the proposed semantic BC network can be implemented in real time. Simulations and experiments demonstrate that the proposed robust semantic BC network can significantly improve transmission efficiency.
Shuai Ma 0002, Zhi Zhang 0003, Youlong Wu, Hang Li 0003, Guangming Shi, Dahua Gao, Yuanming Shi, Shiyin Li, Naofal Al-Dhahir
IEEE Trans. Wirel. Commun.7
2024 Vertical Federated Learning Over Cloud-RAN: Convergence Analysis and System Optimization
abstract
Vertical federated learning (FL) is a collaborative machine learning framework that enables devices to learn a global model from the feature-partition datasets without sharing local raw data. However, as the number of the local intermediate outputs is proportional to the training samples, it is critical to develop communication-efficient techniques for wireless vertical FL to support high-dimensional model aggregation with full device participation. In this paper, we propose a novel cloud radio access network (Cloud-RAN) based vertical FL system to enable fast and accurate model aggregation by leveraging over-the-air computation (AirComp) and alleviating communication straggler issue with cooperative model aggregation among geographically distributed edge servers. However, the model aggregation error caused by AirComp and quantization errors caused by the limited fronthaul capacity degrade the learning performance for vertical FL. To address these issues, we characterize the convergence behavior of the vertical FL algorithm considering both uplink and downlink transmissions. To improve the learning performance, we establish a system optimization framework by joint transceiver and fronthaul quantization design, for which successive convex approximation and alternate convex search based system optimization algorithms are developed. We conduct extensive simulations to demonstrate the effectiveness of the proposed system architecture and optimization framework for vertical FL.
Yuanming Shi, Shuhao Xia, Yong Zhou 0006, Yijie Mao, Chunxiao Jiang, Meixia Tao
IEEE Trans. Wirel. Commun.1
2024 Federated Edge Learning With Differential Privacy: An Active Reconfigurable Intelligent Surface Approach
abstract
Federated edge learning (FL) has become an unprecedented machine learning paradigm that enables distributed training across multiple edge devices without sharing their private data. Nevertheless, recent privacy eavesdropping attacks have raised severe privacy concerns, which make FL untrustworsthy and thus hinder the wide deployment of FL in emerging high-stake applications, such as vehicular networks and healthcare industry. Fortunately, differential privacy (DP) provides a flexible approach by introducing additional randomness to the released model updates so that the eavesdroppers cannot divulge any private information. However, the injected perturbation ensures privacy at the expense of learning accuracy and communication cost, yielding an accuracy-privacy-communication dilemma. In this article, we propose an active reconfigurable intelligent surface (RIS) approach to tackle the dilemma in differentially private FL, which is achieved by exploiting the reconfigurability of active RIS to address the heterogeneous wireless links and privacy concerns, as well as the waveform superposition property with over-the-air computation (AirComp) for low-latency model aggregation. We comprehensively analyze the convergence behavior and systematic privacy guarantee of the active RIS-enabled differentially private FL system, followed by proposing a two-step online power adaptation scheme to minimize the learning optimality gap while satisfying the systematic privacy and power constraints by jointly designing the transmit scalar and artificial noise at the edge devices and the reflection beamforming pattern at the active RIS. Simulation results validate our theoretical achievements and demonstrate the advancements of active RIS in addressing the accuracy-privacy-communication dilemma in differentially private FL.
Yuanming Shi, Youlong Wu
IEEE Trans. Wirel. Commun.1
2024 Satellite Federated Edge Learning: Architecture Design and Convergence Analysis
abstract
The proliferation of low-earth-orbit (LEO) satellite networks leads to the generation of vast volumes of remote sensing data which is traditionally transferred to the ground server for centralized processing, raising privacy and bandwidth concerns. Federated edge learning (FEEL), as a distributed machine learning approach, has the potential to address these challenges by sharing only model parameters instead of raw data. Although promising, the dynamics of LEO networks, characterized by the high mobility of satellites and short ground-to-satellite link (GSL) duration, pose unique challenges for FEEL. Notably, frequent model transmission between the satellites and ground incurs prolonged waiting time and large transmission latency. This paper introduces a novel FEEL algorithm, named FEDMEGA, tailored to LEO mega-constellation networks. By integrating inter-satellite links (ISL) for intra-orbit model aggregation, the proposed algorithm significantly reduces the usage of low datarate and intermittent GSL. Our proposed method includes a ring all-reduce based intra-orbit aggregation mechanism, coupled with a network flow-based transmission scheme for global model aggregation, which enhances transmission efficiency. Theoretical convergence analysis is provided to characterize the algorithm performance. Extensive simulations show that our FEDMEGA algorithm outperforms existing satellite FEEL algorithms, exhibiting an approximate 30% improvement in convergence rate.
Yuanming Shi, Jingyang Zhu, Yong Zhou 0006, Chunxiao Jiang, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.1
2024 Green Federated Learning Over Cloud-RAN With Limited Fronthaul Capacity and Quantized Neural Networks
abstract
In this paper, we propose an energy-efficient federated learning (FL) framework for the energy-constrained devices over cloud radio access network (Cloud-RAN), where each device adopts quantized neural networks (QNNs) to train a local FL model and transmits the quantized model parameter to the remote radio heads (RRHs). Each RRH receives the signals from devices over the wireless link and forwards the signals to the server via the fronthaul link. We rigorously develop an energy consumption model for the local training at devices through the use of QNNs and communication models over Cloud-RAN. Based on the proposed energy consumption model, we formulate an energy minimization problem that optimizes the fronthaul rate allocation, device transmit power allocation, and QNN precision levels while satisfying the limited fronthaul capacity constraint and ensuring the convergence of the proposed FL model to a target accuracy. To solve this problem, we analyze the convergence rate and propose efficient algorithms based on the alternative optimization technique. Simulation results show that the proposed FL framework can significantly reduce energy consumption compared to other conventional approaches. We draw the conclusion that the proposed framework holds great potential for achieving a sustainable and environmentally-friendly FL in Cloud-RAN.
Yijie Mao, Ting Wang 0001, Yuanming Shi
IEEE Trans. Wirel. Commun.4
2024 Over-the-Air Federated Graph Learning
abstract
Message-passing graph neural network (MPGNN) shows tremendous promise in modeling complex networks by capturing the interaction among vertices via the messaging-passing mechanism. However, the dimension of MPGNN is tied to the size of vertices in the graph, which varies from graph to graph, resulting in dimension mismatch that hinders the utilization of graph data distributed at the network edge. To address this issue, we in this paper leverage the attention mechanism to project the graph representation of MPGNNs into a unified space and apply over-the-air computation (AirComp) to support federated graph learning (FGL) over wireless networks. By explicitly deriving the upper bound on the convergence of over-the-air FGL, we formulate a long-term transmission distortion minimization problem, which is further decomposed into a series of online optimization problems by using Lyapunov optimization. We further propose a deep reinforcement learning based algorithm to optimize the AirComp transceiver, where the analytical expression of transmit power is exploited in the action design to reduce the searching space and also enhance the training performance. Simulations demonstrate that, compared to the benchmarks, the proposed algorithm attains two orders of magnitude acceleration in the inference stage, while exhibiting enhanced robustness and improving learning performance.
Yong Zhou 0006, Yuanming Shi
IEEE Trans. Wirel. Commun.3
2024 Decentralized Over-the-Air Federated Learning in Full-Duplex MIMO Networks
abstract
Decentralized federated learning (FL) is capable of enabling efficient and robust collaborative model training with device-to-device (D2D) communications. However, most existing studies on decentralized FL employ half-duplex communication to achieve time-division model aggregation, which is inefficient in scenarios with massive geographically dispersed devices. To address this issue, we in this paper propose decentralized over-the-air FL (DOAFL) with full-duplex (FD) communication, where over-the-air computation (AirComp) and FD communication are fused together to enable parallel model exchange and aggregation, and antenna arrays are leveraged to suppress residual self-interference (SI). Specifically, we first conduct the convergence analysis for DOAFL to characterize the influence of the consensus error introduced by residual SI, channel fading, and receiver noise on the learning performance. Subsequently, we formulate a joint communication and computation (JC2) optimization problem with an objective to increase both the accuracy and time efficiency of the model training, followed by developing a JC2 design algorithm to efficiently optimize transceiver beamforming and computing frequencies. Simulation results verify the superiority of our proposed DOAFL in terms of training latency, residual SI suppression, and learning performance under low energy budgets.
Zhibin Wang 0003, Yong Zhou 0006, Yuanming Shi
IEEE Trans. Wirel. Commun.3
2024 Task-Oriented Over-the-Air Computation for Multi-Device Edge AI
abstract
Edge inference refers to the use of artificial intelligent (AI) models at the network edge to provide mobile devices inference services and thereby enable intelligent services such as auto-driving and Metaverse towards 6G. However, departing from the classic paradigm of data-centric designs, the 6G networks for supporting edge AI features task-oriented techniques that focus on effective and efficient execution of AI task. Targeting end-to-end system performance, such techniques are sophisticated as they aim to seamlessly integrate sensing (data acquisition), communication (data transmission), and computation (data processing). Aligned with the paradigm shift, a task-oriented over-the-air computation (AirComp) scheme is proposed in this paper for multi-device split-inference system. In the considered system, local feature vectors, which are extracted from the real-time noisy sensory data on devices, are aggregated over-the-air by exploiting the waveform superposition in a multiuser channel. Then the aggregated features as received at a server are fed into an inference model with the result used for decision making or control of actuators. To design inference-oriented AirComp, the transmit precoders at edge devices and receive beamforming at edge server are jointly optimized to rein in the aggregation error and maximize the inference accuracy. The problem is made tractable by measuring the inference accuracy using a surrogate metric called discriminant gain, which measures the discernibility of two object classes in the application of object/event classification. It is discovered that the conventional AirComp beamforming design for minimizing the mean square error in generic AirComp with respect to the noiseless case may not lead to the optimal classification accuracy. The reason is due to the overlooking of the fact that feature dimensions have different sensitivity towards aggregation errors and are thus of different importance levels for classification. This issue is addressed in this work via a new task-oriented AirComp scheme designed by directly maximizing the derived discriminant gain. However, the resultant problem of joint transmit precoding and receive beamforming is nonconvex and difficult to solve due to the complicated form of discriminant gain and the coupling between the control variables. We overcome the difficulty using the successive convex approximation. The performance gain of the proposed task-oriented scheme over the conventional schemes is verified by extensive experiments targeting the application of human motion recognition.
Dingzhu Wen, Xiang Jiao, Peixi Liu, Guangxu Zhu, Yuanming Shi, Kaibin Huang
IEEE Trans. Wirel. Commun.5
2024 Task-Oriented Sensing, Computation, and Communication Integration for Multi-Device Edge AI
abstract
This paper studies a new multi-device edge artificial-intelligent (AI) system, which jointly exploits the AI model split inference and integrated sensing and communication (ISAC) to enable low-latency intelligent services at the network edge. In this system, multiple ISAC devices perform radar sensing to obtain multi-view data, and then offload the quantized version of extracted features to a centralized edge server, which conducts model inference based on the cascaded feature vectors. Under this setup and by considering classification tasks, we measure the inference accuracy by adopting an approximate but tractable metric, namely discriminant gain, which is defined as the distance of two classes in the Euclidean feature space under normalized covariance. To maximize the discriminant gain, we first quantify the influence of the sensing, computation, and communication processes on it with a derived closed-form expression. Then, an end-to-end task-oriented resource management approach is developed by integrating the three processes into a joint design. This integrated sensing, computation, and communication (ISCC) design approach, however, leads to a challenging non-convex optimization problem, due to the complicated form of discriminant gain and the device heterogeneity in terms of channel gain, quantization level, and generated feature subsets. Remarkably, the considered non-convex problem can be optimally solved based on the sum-of-ratios method. This gives the optimal ISCC scheme, that jointly determines the transmit power and time allocation at multiple devices for sensing and communication, as well as their quantization bits allocation for computation distortion control. By using human motions recognition as a concrete AI inference task, extensive experiments are conducted to verify the performance of our derived optimal ISCC scheme.
Dingzhu Wen, Peixi Liu, Guangxu Zhu, Yuanming Shi, Jie Xu 0002, Yonina C. Eldar, Shuguang Cui
IEEE Trans. Wirel. Commun.4
2024 Federated Learning With Massive Random Access
abstract
In this paper, we propose an online federated learning framework with massive random access, aiming to learn a sequence of global models using local data that are sequentially collected by massive edge devices. As only a subset of devices is capable of collecting data and performing local model update at any specific moment, the communication pattern between the edge server and devices is random and sporadic, which is referred to assporadic local updates. This motivates us to adopt a two-phase grant-free random access scheme that consists of the activity detection and model transmission phases to facilitate efficient communication between the edge server and devices. We first provide the regret analysis for online federated learning, and derive the optimality gap in terms of successful transmission probabilities. Then, we characterize the achievable transmission rate of each active device using random matrix theory and establish the relationship between the pilot length and the outage probability. Furthermore, we propose an optimal pilot length design by minimizing the optimality gap. To validate our scheme, we provide comprehensive experimental results that demonstrate the superiority of the proposed scheme over traditional schemes in various online tasks.
Shuhao Xia, Yuanming Shi, Yong Zhou 0006, Youlong Wu, Lin Yang 0011, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.2
2024 Decentralized Over-the-Air Federated Learning by Second-Order Optimization Method
abstract
Federated learning (FL) is an emerging technique that enables privacy-preserving distributed learning. Most related works focus on centralized FL, which leverages the coordination of a parameter server to implement local model aggregation. However, this scheme heavily relies on the parameter server, which could cause scalability, communication, and reliability issues. To tackle these problems, decentralized FL, where information is shared through gossip, starts to attract attention. Nevertheless, current research mainly relies on first-order optimization methods that have a relatively slow convergence rate, which leads to excessive communication rounds in wireless networks. To design communication-efficient decentralized FL, we propose a novel over-the-air decentralized second-order federated algorithm. Benefiting from the fast convergence rate of the second-order method, total communication rounds are significantly reduced. Meanwhile, owing to the low-latency model aggregation enabled by over-the-air computation, the communication overheads in each round can also be greatly decreased. The convergence behavior of our approach is then analyzed. The result reveals an error term, which involves a cumulative noise effect, in each iteration. To mitigate the impact of this error term, we conduct system optimization from the perspective of the accumulative term and the individual term, respectively. Numerical experiments demonstrate the superiority of our proposed approach and the effectiveness of system optimization.
Peng Yang 0027, Yuning Jiang 0002, Dingzhu Wen, Ting Wang 0001, Colin N. Jones, Yuanming Shi
IEEE Trans. Wirel. Commun.6
2024 One-Bit Byzantine-Tolerant Distributed Learning via Over-the-Air Computation
abstract
Distributed learning has become a promising computational parallelism paradigm that enables a wide scope of intelligent applications from the Internet of Things (IoT) to autonomous driving and the healthcare industry. This paper studies distributed learning in wireless data center networks, which contain a central edge server and multiple edge workers to collaboratively train a shared global model and benefit from parallel computing. However, the distributed nature causes the vulnerability of the learning process to faults and adversarial attacks from Byzantine edge workers, as well as the severe communication and computation overhead induced by the periodical information exchange process. To achieve fast and reliable model aggregation in the presence of Byzantine attacks, we develop a signed stochastic gradient descent (SignSGD)-based Hierarchical Vote framework via over-the-air computation (AirComp), where one voting process is performed locally at the wireless edge by taking advantage of Bernoulli coding while the other is operated over-the-air at the central edge server by utilizing the waveform superposition property of the multiple-access channels. We comprehensively analyze the proposed framework on the impacts including Byzantine attacks and the wireless environment (channel fading and receiver noise), followed by characterizing the convergence behavior under non-convex settings. Simulation results validate our theoretical achievements and demonstrate the robustness of our proposed framework in the presence of Byzantine attacks and receiver noise.
Youlong Wu, Yuning Jiang 0002, Yuanming Shi
IEEE Trans. Wirel. Commun.4
2024 Over-the-Air Computation Empowered Vertically Split Inference
abstract
To tackle the issue of heterogeneous input raw data samples obtained by different devices and enhance the feature extraction capability of edge devices, we propose a vertically split neural network based edge-device collaborative artificial intelligence (AI) inference framework. The local results calculated by various light-size sub-networks at edge devices are transmitted and aggregated at the server for the downstream inference task. Nevertheless, the transmission of such high-dimensional local results involves severe communication overhead. To resolve this issue, the technique of over-the-air computation (AirComp) is adopted to enable low-latency aggregation. The same entry of all devices’ local results is transmitted over a same wireless resource block and aggregated via the waveform superposition property. Furthermore, to simultaneously support the aggregation of all dimensions of the local results, we consider a broadband channel and leverage orthogonal frequency division multiplexing (OFDM) to divide the system bandwidth into multiple subcarriers which are then assigned for different dimensions. Consequently, an extra degree of freedom is introduced to design the aggregation of all dimensions. We then propose a scheme of joint subcarrier allocation, power allocation, and receiver beamforming to minimize the aggregation distortion and enhance inference performance. Extensive experiments are conducted to verify the superiority of the proposed design over benchmarks.
Peng Yang 0027, Dingzhu Wen, Qunsong Zeng, Yong Zhou 0006, Ting Wang 0001, Haibin Cai, Yuanming Shi
IEEE Trans. Wirel. Commun.7
2024 Federated Multi-Task Learning with Non-Stationary and Heterogeneous Data in Wireless Networks
abstract
Federated multi-task learning (FMTL) is a promising edge learning framework to fit the data with non-independent and non-identical distribution (non-i.i.d.) by leveraging the statistical correlations among the personalized models. For many practical applications in wireless communications, the sensory data are not only heterogeneous but also non-stationary due to the mobility of terminals and the randomness of link connections. The non-stationary heterogeneous data may lead to model divergence and staleness in the training stage and poor test accuracy in the inference stage. In this paper, we shall develop an adaptive FMTL framework, which works well with non-stationary data. We further propose to optimize the model updating and cluster splitting schemes in the training stage to accelerate model convergence. We also design a low-complexity model selection and pruning schemes in both the training and inference stages to select the best model for fitting the current data and delete redundant models, respectively. The proposed framework is validated in the edge learning model, namely, the linear regression problem for indoor localization in wireless networks and GNN for wireless power control problems. Numerical results demonstrate that the proposed framework can accelerate the model training convergence and reduce the computation complexity while ensuring model accuracy.
Hongwei Zhang 0006, Meixia Tao, Yuanming Shi, Xiaoyan Bi, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.3
2024 Adaptive Partitioning and Placement for Two-Layer Collaborative Caching in Mobile Edge Computing Networks
abstract
With the explosive growth in demands for mobile video services, the focus of cellular networks is evolving from the core network to the edge network to facilitate resource-intensive services and mitigate backhaul burdens. However, the overlapping in coverage and the user similarity in preferences lead to substantial cache redundancy. In addition, the mobility of users poses a significant challenge to smart devices for device-to-device (D2D) communications, which affects the content richness and hit ratio of the edge cache. In this paper, an adaptive cache partitioning and placement (ACPP) strategy is proposed for mobile edge computing (MEC) networks to minimize the average cost of content access. A practical two-layer collaborative caching model is presented, which comprises 5G base station (gNB) clusters and D2D caching. Besides, a public and private cache partitioning method is designed for gNBs to improve the content richness of local cache, and a static and dynamic cache partitioning method is developed for user devices to address the varying mobility patterns. Simulation results demonstrate the effectiveness of the proposed ACPP strategy in delivering content at a lower average cost, and achieve a better hit ratio with a relatively high energy consumption at user ends.
Yingxue Zhao, Ailing Xiao, Sheng Wu 0001, Chunxiao Jiang, Linling Kuang, Yuanming Shi
IEEE Trans. Wirel. Commun.6
2024 Integrated Sensing-Communication-Computation for Over-the-Air Edge AI Inference
abstract
Edge-device co-inference refers to deploying well-trained artificial intelligent (AI) models at the network edge under the cooperation of devices and edge servers for providing ambient intelligent services. For enhancing the utilization of limited network resources in edge-device co-inference tasks from a systematic view, we propose a task-oriented scheme of integrated sensing, computation and communication (ISCC) in this work. In this system, all devices sense a target from the same wide view to obtain homogeneous noise-corrupted sensory data, from which the local feature vectors are extracted. All local feature vectors are aggregated at the server using over-the-air computation (AirComp) in a broadband channel with the orthogonal-frequency-division-multiplexing technique for suppressing the sensing and channel noise. The aggregated denoised global feature vector is further input to a server-side AI model for completing the downstream inference task. A novel task-oriented design criterion, called maximum minimum pair-wise discriminant gain, is adopted for classification tasks. It extends the distance of the closest class pair in the feature space, leading to a balanced and enhanced inference accuracy. Under this criterion, a problem of joint sensing power assignment, transmit precoding and receive beamforming is formulated. The challenge lies in three aspects: the coupling between sensing and AirComp, the joint optimization of all feature dimensions’ AirComp aggregation over a broadband channel, and the complicated form of the maximum minimum pair-wise discriminant gain. To solve this problem, a task-oriented ISCC scheme with AirComp is proposed. Experiments based on a human motion recognition task are conducted to verify the advantages of the proposed scheme over the existing scheme and a baseline.
Zeming Zhuang, Dingzhu Wen, Yuanming Shi, Guangxu Zhu, Sheng Wu 0001, Dusit Niyato
IEEE Trans. Wirel. Commun.3
2023 Federated Linear Bandit Learning via Over-the-air Computation
abstract
In this paper, we investigate federated contextual linear bandit learning within a wireless system that comprises a server and multiple devices. Each device interacts with the environment, selects an action based on the received reward, and sends model updates to the server. The primary objective is to minimize cumulative regret across all devices within a finite time horizon. To reduce the communication overhead, devices communicate with the server via over-the-air computation (AirComp) over noisy fading channels, where the channel noise may distort the signals. In this context, we propose a customized federated linear bandits scheme, where each device transmits an analog signal, and the server receives a superposition of these signals distorted by channel noise. A rigorous mathematical analysis is conducted to determine the regret bound of the proposed scheme. Both theoretical analysis and numerical experiments demonstrate the competitive performance of our proposed scheme in terms of regret bounds in various settings.
Yuning Jiang 0002, Xin Liu 0049, Ting Wang 0001, Yuanming Shi
GLOBECOM5
2023 Multi-Task-Oriented Broadcast for Edge AI Inference via Information Bottleneck
abstract
In this paper, we consider a task-oriented communication paradigm over multi-user broadcast channels for multi-task edge AI inference, where an edge transmitter performs feature extraction and broadcasts the encoded features to multiple edge devices, each of which conducts a specific inference task based on the received signals. However, due to the heterogeneity of communication links, it still remains an open problem to balance the trade-off between inference accuracy and robustness to channel noise over broadcast channels. To address this issue, we propose a task-oriented broadcast framework guided by information bottleneck (IB) for feature extraction and broadcasting, yielding a two-phase network training strategy, where the former aims to extract task-relevant features for each task, and the latter is utilized to compress the extracted features and perform robust broadcasting. Simulation results demonstrate that the proposed task-oriented broadcast framework can achieve a better trade-off between inference accuracy and robustness under various channel conditions than the benchmarks.
Youlong Wu, Shuai Ma 0002, Yuanming Shi
GLOBECOM4
2023 Decentralized Over-the-Air Computation for Edge AI Inference with Integrated Sensing and Communication
abstract
Collaborative artificial intelligent (AI) inference has been an effective approach to deploying well-trained AI models at the network edge for empowering immersive intelligent services such as autonomous driving and smart cities. In this paper, we propose an integrated sensing-computation-communication (ISCC) scheme for decentralized collaborative inference systems. In the proposed scheme, multiple devices connect to each other via device-to-device (D2D) links. Each device first extracts a homogeneous feature vector from the raw sensory data obtained from the same wide view of the source target and then aggregates all local feature vectors using the over-the-air computation technique. To further enhance the spectrum efficiency, the full-duplex technology is utilized to allow all devices to transmit and receive in the same frequency band. This, however, introduces significant self-interference and coupling among different tasks. To address these challenges, a multi-objective optimization-based ISCC approach is proposed.
Zeming Zhuang, Dingzhu Wen, Yuanming Shi
GLOBECOM3
2023 Task-Oriented Sensing, Computation, and Communication Integration for Multi-Device Edge AI
abstract
This paper studies a new multi-device edge artificial-intelligent (AI) system, which jointly exploits the AI model split inference and integrated sensing and communication (ISAC) to enable low-latency intelligent services at the network edge. In this system, multiple ISAC devices perform radar sensing to obtain multi-view data, and then offload the quantized version of extracted features to a centralized edge server, which conducts model inference based on the cascaded feature vectors. Under this setup and by considering classification tasks, we measure the inference accuracy by adopting an approximate but tractable metric, namely discriminant gain, which is defined as the distance of two classes in the Euclidean feature space under normalized covariance. To maximize the discriminant gain, we first quantify the influence of the sensing, computation, and communication processes on it with a derived closed-form expression. Then, an end-to-end task-oriented resource management approach is developed by designing an optimal integrated sensing, computation, and communication (ISCC) scheme. By using human motions recognition as a concrete AI inference task, extensive experiments are conducted to verify the performance of the proposed scheme.
Dingzhu Wen, Peixi Liu, Guangxu Zhu, Yuanming Shi, Jie Xu 0002, Yonina C. Eldar, Shuguang Cui
ICC4
2023 STAR-RIS Empowered Full Duplex Cooperative Rate Splitting
abstract
Cooperative rate-splitting (CRS), which takes the advantages of cooperative user relaying and rate-splitting multiple access (RSMA), has gained recognition for its potential in enhancing the spectral efficiency, user fairness, and coverage of wireless networks. In this work, to further strengthen the performance of CRS, we propose a simultaneously transmitting and reflecting reconfigurable intelligent surface (STAR-RIS)-enabled full-duplex (FD) CRS transmission framework. By jointly designing the active beamforming, common rate allocation, and the STAR-RIS passive transmission and reflection beamforming to maximize the worst case rate among users, we show that the proposed STAR-RIS assisted FD CRS transmission scheme considerably enhances user fairness compared to conventional FD CRS approaches.
Kangchun Zhao, Yijie Mao, Yuanming Shi
VTC Fall3
2023 Reconfigurable Intelligent Surface Empowered Rate-Splitting Multiple Access for Simultaneous Wireless Information and Power Transfer
abstract
Rate-splitting multiple access (RSMA) and reconfigurable intelligent surface (RIS) have been both recognized as promising techniques for 6G. The benefits of combining the two techniques to enhance the spectral and energy efficiency have been recently exploited in communication-only networks. Inspired by the recent advances, in this work we investigate the use of RIS empowered RSMA for simultaneous wireless information and power transfer (SWIPT) with one transmitter concurrently sending information to multiple information receivers (IRs) and transferring energy to multiple energy receivers (ERs). Specifically, we jointly optimize the transmit beamformers and the RIS reflection coefficients to maximize the weighted sum-rate (WSR) of IRs under the harvested energy constraint of ERs and the transmit power constraint. An alternating optimization and successive convex approximation (SCA)-based optimization framework is then raised to address the problem. Numerical results demonstrate that by marrying the benefits of RSMA and RIS, the proposed RIS empowered RSMA achieves a better tradeoff between the WSR of IRs and energy harvested at ERs. In addition, the rate region of RSMA almost coincides with that of SDMA+RIS especially when the two users have similar weights. Therefore, we conclude that RIS empowered RSMA is a promising strategy for SWIPT.
Chengzhong Tian, Yijie Mao, Kangchun Zhao, Yuanming Shi, Bruno Clerckx
WCNC4
2023 Task-Oriented Over-the-Air Computation for Multi-Device Edge Split Inference
abstract
A task-oriented over-the-air computation (AirComp) scheme is proposed in this paper for multi-device edge split inference system. In the considered system, local noise-corrupted feature vectors are aggregated at the server via AirComp to generate a denoised one for the subsequent inference task. By considering classification tasks, the transmit precoders at edge devices and receive beamforming at edge server are jointly designed in an effort to rein in the aggregation error and maximize the inference accuracy, which is approximately measured by a surrogate but more tractable metric called discriminant gain. It is found that the conventional AirComp beamforming design for minimizing the mean square error between the aggregated feature vector by AirComp and the ideally aggregated one may not lead to the optimal classification accuracy, as it fails to respect the fact that some feature dimensions are more sensitive to the aggregation error than the others in terms of the classification accuracy. To tackle this issue, a new task-oriented AirComp scheme is proposed for directly maximizing the derived discriminant gain. The superiority of the proposed scheme over the heuristic benchmarks is verified by extensive experimental results based on a concrete inference task of human motion recognition.
Dingzhu Wen, Xiang Jiao, Peixi Liu, Guangxu Zhu, Yuanming Shi, Kaibin Huang
WCNC5
2023 Latency Minimization for Wireless Federated Learning with Heterogeneous Local Updates
abstract
In this paper, we study the latency minimization problem for a wireless federated learning (FL) system with heterogeneous computation capability, where different edge devices perform different numbers of local updates in each communication round. We formulate a total latency minimization problem, taking into account both the communication and computation latency in the whole FL procedure. We reveal that decoupling the resource allocation variables from the model convergence is essential to reduce the problem to a single-round latency minimization problem. To solve this simplified problem, we propose an alternating optimization scheme to jointly consider communication and computation resource allocation and mitigate the straggler effect. We prove that the resulting sub-problems, i.e., bandwidth and computation capacity allocation, are both convex and can be optimally solved in closed form, respectively. Simulations show that compared with the baseline scheme that allocates the communication and computation resources equally across edge devices, the proposed scheme can achieve single-round latency reduction.
Jingyang Zhu, Yuanming Shi, Min Fu 0003, Yong Zhou 0006, Youlong Wu, Liqun Fu 0001
WCNC2
2023 Trustworthy Federated Learning via Blockchain
abstract
The safety-critical scenarios of artificial intelligence (AI), such as autonomous driving, Internet of Things, smart healthcare, etc., have raised critical requirements of trustworthy AI to guarantee the privacy and security with reliable decisions. As a nascent branch for trustworthy AI, federated learning (FL) has been regarded as a promising privacy preserving framework for training a global AI model over collaborative devices. However, security challenges still exist in the FL framework, e.g., Byzantine attacks from malicious devices, and model tampering attacks from malicious server, which will degrade or destroy the accuracy of trained global AI model. In this article, we shall propose a decentralized blockchain-based FL (B-FL) architecture by using a secure global aggregation algorithm to resist malicious devices, and deploying a practical Byzantine fault tolerance consensus protocol with high effectiveness and low energy consumption among multiple edge servers to prevent model tampering from the malicious server. However, to implement B-FL system at the network edge, multiple rounds of cross-validation in blockchain consensus protocol will induce long training latency. We thus formulate a network optimization problem that jointly considers bandwidth and power allocation for the minimization of long-term average training latency consisting of progressive learning rounds. We further propose to transform the network optimization problem as a Markov decision process and leverage the deep reinforcement learning (DRL)-based algorithm to provide high system performance with low computational complexity. Simulation results demonstrate that B-FL can resist malicious attacks from edge devices and servers, and the training latency of B-FL can be significantly reduced by the DRL-based algorithm compared with the baseline algorithms.
Zhanpeng Yang, Yuanming Shi, Yong Zhou 0006, Kai Yang 0006
IEEE Internet Things J.2
2023 Gradient and Channel Aware Dynamic Scheduling for Over-the-Air Computation in Federated Edge Learning Systems
abstract
To satisfy the expected plethora of computation-heavy applications, federated edge learning (FEEL) is a new paradigm featuring distributed learning to carry the capacities of low-latency and privacy-preserving. To further improve the efficiency of wireless data aggregation and model learning, over-the-air computation (AirComp) is emerging as a promising solution by using the superposition characteristics of wireless channels. However, the fading and noise of wireless channels can cause aggregate distortions in AirComp enabled federated learning. In addition, the quality of collected data and energy consumption of edge devices may also impact the accuracy and efficiency of model aggregation as well as convergence. To solve these problems, this work proposes a dynamic device scheduling mechanism, which can select qualified edge devices to transmit their local models with a proper power control policy so as to participate the model training at the server in federated learning via AirComp. In this mechanism, the data importance is measured by the gradient of local model parameter, channel condition and energy consumption of the device jointly. In particular, to fully use distributed datasets and accelerate the convergence rate of federated learning, the local updates of unselected devices are also retained and accumulated for future potential transmission, instead of being discarded directly. Furthermore, the Lyapunov drift-plus-penalty optimization problem is formulated for searching the optimal device selection strategy. Simulation results validate that the proposed scheduling mechanism can achieve higher test accuracy and faster convergence rate, and is robust against different channel conditions.
Jun Du 0001, Bingqing Jiang, Chunxiao Jiang, Yuanming Shi, Zhu Han 0001
IEEE J. Sel. Areas Commun.4
2023 Robust Information Bottleneck for Task-Oriented Communication With Digital Modulation
abstract
Task-oriented communications, mostly using learning-based joint source-channel coding (JSCC), aim to design a communication-efficient edge inference system by transmitting task-relevant information to the receiver. However, only transmitting task-relevant information without introducing any redundancy may cause robustness issues in learning due to the channel variations, and the JSCC which directly maps the source data into continuous channel input symbols poses compatibility issues on existing digital communication systems. In this paper, we address these two issues by first investigating the inherent tradeoff between the informativeness of the encoded representations and the robustness to information distortion in the received representations, and then propose a task-oriented communication scheme with digital modulation, named discrete task-oriented JSCC (DT-JSCC), where the transmitter encodes the features into a discrete representation and transmits it to the receiver with the digital modulation scheme. In the DT-JSCC scheme, we develop a robust encoding framework, named robust information bottleneck (RIB), to improve the communication robustness to the channel variations, and derive a tractable variational upper bound of the RIB objective function using the variational approximation to overcome the computational intractability of mutual information. The experimental results demonstrate that the proposed DT-JSCC achieves better inference performance than the baseline methods with low communication latency, and exhibits robustness to channel variations due to the applied RIB framework.
Songjie Xie, Shuai Ma 0002, Ming Ding 0001, Yuanming Shi, MingJian Tang 0001, Youlong Wu
IEEE J. Sel. Areas Commun.4
2023 POBO: Safe and optimal resource management for cloud microservices
Hengquan Guo, Hongchen Cao, Jingzhu He, Xin Liu 0049, Yuanming Shi
Perform. Evaluation5
2023 Reconfigurable Intelligent Surfaces Empowered Green Wireless Networks With User Admission Control
abstract
Reconfigurable intelligent surface (RIS) has emerged as a cost-effective and energy-efficient technique for 6G. By adjusting the phase shifts of passive reflecting elements, RIS is capable of suppressing the interference and combining the desired signals constructively at receivers, thereby significantly enhancing the performance of communication system. In this paper, we consider a green multi-user multi-antenna cellular network, where multiple RISs are deployed to provide energy-efficient communication service to end users. We jointly optimize the phase shifts of RISs, beamforming of the base stations, and the active RIS set with the aim of minimizing the power consumption of the base station (BS) and RISs subject to the quality of service (QoS) constraints of users and the transmit power constraint of the BS. However, the problem is mixed combinatorial and non-convex, and there is a potential infeasibility issue when the QoS constraints cannot be guaranteed by all users. To deal with the infeasibility issue, we further investigate a user admission control problem to jointly optimize the transmit beamforming, RIS phase shifts, and the admitted user set. A unified alternating optimization (AO) framework is then proposed to solve both the power minimization and user admission control problems. Specifically, we first decompose the original non-convex problem into several rank-one constrained optimization subproblems via matrix lifting. A difference-of-convex (DC) algorithm is then developed to solve each decomposed subproblem. The proposed AO framework efficiently minimizes the power consumption of wireless networks as well as user admission control when the QoS constraints cannot be guaranteed by all users. To further address the complexity-sensitive issue for practical implementation, we propose an alternative low-complexity beamforming and RISs phase shifts design algorithm based on zero-forcing (ZF) to enable the green cellular networks.
Jinglian He, Yijie Mao, Yong Zhou 0006, Ting Wang 0001, Yuanming Shi
IEEE Trans. Commun.5
2023 Communication-Efficient Coded Computing for Distributed Multi-Task Learning
abstract
Distributed multi-task learning (MTL) can jointly learn multiple models and achieve better generalization performance by exploiting relevant information between the tasks. However, distributed MTL suffers from communication bottlenecks, in particular for large-scale learning with a massive number of tasks. This paper considers distributed MTL systems where distributed workers wish to learn different models orchestrated by a central server. To mitigate communication bottlenecks both in the uplink and downlink, we propose coded computing schemes for flexible and fixed data placements, respectively. Our schemes can significantly reduce communication loads by exploiting workers’ local information and creating multicast opportunities for both the server and workers. Moreover, we establish information-theoretic lower bounds on the optimal downlink and uplink communication loads, and prove the approximate optimality of the proposed schemes. For flexible data placement, our scheme achieves theoptimaldownlink communication load, and theorder optimaluplink communication load that is smaller than 2 times of the information-theoretic optimum. For fixed data placement, the gaps between our communication load and the optimum are within the minimum computation load among all workers, regardless of the number of workers. Experiments demonstrate that our schemes can significantly speed up the training process compared to the traditional approach.
Haoyang Hu, Youlong Wu, Yuanming Shi, Chunxiao Jiang, Wei Zhang 0001
IEEE Trans. Commun.3
2023 Multi-Agent Reinforcement Learning for Dynamic Resource Management in 6G in-X Subnetworks
abstract
The 6G network enables a subnetwork-wide evolution, resulting in a “network of subnetworks”. However, due to the dynamic mobility of wireless subnetworks, the data transmission of intra-subnetwork and inter-subnetwork will inevitably interfere with each other, which poses a great challenge to radio resource management. Moreover, most existing approaches require the instantaneous channel gain between subnetworks, which are usually difficult to be collected. To tackle these issues, in this paper we propose a novel effective intelligent radio resource management method using multi-agent deep reinforcement learning (MARL), which only needs the sum of received power, named received signal strength indicator (RSSI), on each channel instead of channel gains. However, to directly separate individual interference from RSSI is an almost impossible thing. To this end, we further propose a novel MARL architecture, named GA-Net, which integrates a hard attention layer to model the importance distribution of inter-subnetwork relationships based on RSSI and excludes the impact of unrelated subnetworks, and employs a graph attention network with a multi-head attention layer to exact the features and calculate their weights that will impact individual throughput. Experimental results prove that our proposed framework significantly outperforms both traditional and MARL-based methods in various aspects.
Ting Wang 0001, Qiang Feng 0004, Chenhui Ye, Tao Tao 0004, Lu Wang 0002, Yuanming Shi, Mingsong Chen 0001
IEEE Trans. Wirel. Commun.7
2023 UAV-Assisted Multi-Cluster Over-the-Air Computation
abstract
In this paper, we study unmanned aerial vehicles (UAVs) assisted wireless data aggregation (WDA) in multi-cluster networks, where multiple UAVs simultaneously perform different WDA tasks via over-the-air computation (AirComp) without terrestrial base stations. This work focuses on maximizing the minimum amount of WDA tasks performed by each cluster by optimizing the UAV trajectory and transceiver design as well as cluster scheduling and association, while considering the WDA accuracy requirement. Such a joint design is critical for interference management in multi-cluster AirComp networks, via enhancing the signal quality between each UAV and its associated cluster for signal alignment while reducing the inter-cluster interference between each UAV and its non-associated clusters. Although it is generally challenging to optimally solve the formulated non-convex mixed-integer nonlinear programming, an efficient iterative algorithm as a compromise approach is developed by exploiting bisection and block coordinate descent methods, yielding an optimal transceiver solution in each iteration. The optimal binary variables and a suboptimal trajectory are obtained by using the dual method and successive convex approximation, respectively. Simulations show the considerable performance gains of the proposed design over benchmarks and the superiority of deploying multiple UAVs in increasing the number of performed tasks while reducing access delays.
Min Fu 0003, Yong Zhou 0006, Yuanming Shi, Chunxiao Jiang, Wei Zhang 0001
IEEE Trans. Wirel. Commun.3
2023 Task-Oriented Explainable Semantic Communications
abstract
Semantic communications utilize the transceiver computing resources to alleviate scarce transmission resources, such as bandwidth and energy. Although the conventional deep learning (DL) based designs may achieve certain transmission efficiency, the uninterpretability issue of extracted features is the major challenge in the development of semantic communications. In this paper, we propose an explainable and robust semantic communication framework by incorporating the well-established bit-level communication system, which not only extracts and disentangles features into independent and semantically interpretable features, but also only selects task-relevant features for transmission, instead of all extracted features. Based on this framework, we derive the optimal input for rate-distortion-perception theory, and derive both lower and upper bounds on the semantic channel capacity. Furthermore, based on the$\beta $-variational autoencoder ($\beta $-VAE), we propose a practical explainable semantic communication system design, which simultaneously achieves semantic features selection and is robust against semantic channel noise. We further design a real-time wireless mobile semantic communication proof-of-concept prototype. Our simulations and experiments demonstrate that our proposed explainable semantic communications system can significantly improve transmission efficiency, and also verify the effectiveness of our proposed robust semantic transmission scheme.
Shuai Ma 0002, Weining Qiao, Youlong Wu, Hang Li 0003, Guangming Shi, Dahua Gao, Yuanming Shi, Shiyin Li, Naofal Al-Dhahir
IEEE Trans. Wirel. Commun.7
2023 A Graph Neural Network Learning Approach to Optimize RIS-Assisted Federated Learning
abstract
Over-the-air federated learning (FL) is a promising privacy-preserving edge artificial intelligence paradigm, where over-the-air computation enables spectral-efficient model aggregation by achieving simultaneous communication and aggregation. However, due to limited transmit power, the performance of over-the-air FL is limited by the device with the worst channel condition toward the edge server. In this paper, we leverage reconfigurable intelligent surface (RIS) to mitigate the communication bottleneck of over-the-air FL and explicitly characterize the corresponding convergence upper bound. The convergence analysis illustrates the detrimental impact of the accumulated aggregation error over all rounds and inspires us to formulate a time-average transmission distortion minimization problem by jointly optimizing the transceiver and RIS phase-shifts. To reduce the computation complexity and enhance the model aggregation accuracy, we develop a graph neural network (GNN) based learning algorithm to directly map channel coefficients to the optimized network parameters. By exploiting permutation equivalence and invariance properties of graphs, the parameter dimension of the proposed algorithm is independent of the number of edge devices, which reduces the computational complexity and improves the algorithmic scalability. Simulations show that the proposed algorithm speeds up the computation by three orders of magnitude compared to the baselines, while achieving performance superiority and algorithmic robustness.
Yong Zhou 0006, Yinan Zou, Qiaochu An, Yuanming Shi, Mehdi Bennis
IEEE Trans. Wirel. Commun.5
2023 Online Client Selection for Asynchronous Federated Learning With Fairness Consideration
abstract
Federated learning (FL) leverages the private data and computing power of multiple clients to collaboratively train a global model. Many existing FL algorithms over wireless networks adopting synchronous model aggregation suffer from the straggler issue, due to the heterogeneity of local computing power and channel conditions. To address this issue, we in this paper advocate an asynchronous FL framework with adaptive client selection for training latency minimization, taking into account the client availability and long-term fairness. We consider a practical scenario, where the channel conditions and the locally available computing power are not known in prior. This makes the client selection problem challenging, as the training latency consists of the uplink/downlink transmission time and the local training time. To this end, we tackle the asynchronous client selection problem in an online manner by converting the latency minimization problem into a multi-armed bandit problem, and leverage the upper confidence bound policy and virtual queue technique in Lyapunov optimization to solve the problem. We theoretically show that the proposed algorithm achieves sub-linear regret performance, ensures long-term fairness, and guarantees training convergence. Results show that the proposed algorithm can reduce the training time by up to 50% when compared to the baseline algorithms.
Hongbin Zhu, Yong Zhou 0006, Hua Qian, Yuanming Shi, Xu Chen 0004, Yang Yang 0001
IEEE Trans. Wirel. Commun.4
2022 Communication-Efficient Device Scheduling via Over-the-Air Computation for Federated Learning
abstract
Artificial intelligence (AI) is expected as a revo-lutionary technology to be widely used in Internet-of- Things (IoT) networks for computationally intensive tasks. However, the traditional centralized training framework imposes large latency, network burdens and high risk of privacy disclosure. As a promising distributed solution, federated learning involves the collaborative model training among edge devices, with the orchestration of a server to carry the capacities of low-latency and privacy preservation for AI -driven networks. To further improve the communication efficiency, over-the-air computation (AirComp) is capable of computing while transmitting data by exploiting the superposition property of wireless channels to harness the interference. However, gradient aggregation suffers from channel distortion induced by channel fading and noise, which may degrade the training performance. Moreover, it is beneficial to schedule the informative edge devices in federated learning under limited energy resources. In this work, we propose a dynamic device scheduling scheme for AirComp enabled federated learning systems. In this scheme, a proper number of qualified edge devices with channel inversion based power control are scheduled to participate the model training, where local updates diversity, channel condition and energy consumption are exploited jointly. Inspired by the Lyapunov drift-plus-penalty method, we formulate the optimization problem to attain the device selection strategy. Simulation results validate that the proposed scheme can achieve a close-to-optimal test accuracy with fast convergence rate, and present good performance of robustness under different channel conditions.
Bingqing Jiang, Jun Du 0001, Chunxiao Jiang, Yuanming Shi, Zhu Han 0001
GLOBECOM4
2022 Robust Wirtinger Flow Algorithm for Channel Coded Blind Demixing
abstract
As applications of Internet-of-things (IoT) rapidly expand, unscheduled multiple user access with low latency and low cost communication is attracting growing more interests. To recover the multiple uplink signals without strict access control under dynamic co-channel interference environment, the problem of blind demixing emerges as an important obstacle for us to overcome. Without channel state information, successful blind demixing can recover multiple user signals more effectively by leveraging prior information on signal characteristics such as constellations and distribution. This work studies how forward error correction (FEC) codes in Galois Field can generate more effective blind demixing algorithms. We propose a constrained Wirtinger flow algorithm by defining a valid signal set based on FEC codewords. Specifically, targeting the popular polar codes for FEC of short IoT packets, we introduce signal projections within iterations of Wirtinger Flow based on FEC code infor-mation. Simulation results demonstrate stronger robustness of the proposed algorithm against noise and practical obstacles and also faster convergence rate compared to regular Wirtinger flow algorithm.
Amin Jalali 0005, Yuanming Shi, Zhi Ding 0001
ICC2
2022 Federated Multi-Task Learning with Non-Stationary Heterogeneous Data
abstract
Federated multi-task learning (FMTL) is a promising edge learning framework to fit the data with non-independent and non-identical distribution (non-i.i.d.) by exploiting the correlations of personalized models. In many practical systems, the sensory data distribution in wireless systems is not only heterogeneous but also non-stationary due to the mobility of terminals and the randomness of link connections. The non-stationary heterogeneous data may lead to model divergence and staleness in the training stage and poor accuracy in the inference stage. In this paper, we design an adaptive FMTL framework, which can work in a non-stationary environment. We propose to optimize the model update scheme and cluster splitting scheme in the training stage to accelerate model convergencse when the training data are non-stationary. We further design a low-complexity model selection scheme in both the training and the inference stages to choose the best model for fitting the current data. The proposed framework is validated in two scenarios, linear regression and graph neural network (GNN)-based power control in wireless device-to-device (D2D) networks. Both sets of numerical results demonstrate that the proposed framework can accelerate the model training convergence and reduce the computation complexity while ensuring model accuracy.
Hongwei Zhang 0006, Meixia Tao, Yuanming Shi, Xiaoyan Bi
ICC3
2022 Beamforming Design for Integrated Sensing and SWIPT System
abstract
In order to achieve high spectrum utilization, it is necessary to consider the employment of the Integrated Sensing and Communication for the future wireless communication systems. On the other hand, since the radio signal for wireless communication can carry energy at the same time, simultaneous wireless information and power transfer (SWIPT) becomes a popular and helpful technology to enhance the energy transmission efficiency. In order to further improve the spectrum efficiency, we propose to design an integrated sensing and SWIPT (iS2WIPT) system, where a base station transmits combined signals to perform downlink multiuser communication, multiuser energy harvesting and radar target sensing simultaneously. We further formulate the problem as designing the transmit beamforming to minimize the beampattern matching error subject to the given total transmit power budget and the quality-of-service (QoS) constraints at all users, i.e., the individual signal-to-interference-plus-noise ratio (SINR) requirements at information users (IUs) and energy harvesting requirements at energy users (EUs). So as to solve the intractable nonconvex problem, we provide a difference-of-convex-functions (DC) representation for the problem and further solve it with global convergence guarantees. Numerical results clearly show that our proposed DC algorithm outperforms the classical semidefinite relaxation algorithm significantly under different situations. What is more, the results show that there is a tradeoff between the radar sensing performance and the QoS requirements of downlink users (both IUs and EUs).
Xiangyu Zeng 0001, Lukuan Xing, Youlong Wu, Yuanming Shi
PIMRC4
2022 RIS-Assisted Over-the-Air Computation in Millimeter Wave Communication Networks
abstract
Over-the-air computation (AirComp) and millimeter wave (mmWave) communications have the feasibility to perform fast wireless data aggregation (WDA) by allowing simultaneous transmissions and providing abundant spectral resources, respectively. However, AirComp is limited by the link with the worst channel condition, while mmWave communications are vulnerable to the blockages. To address these issues, this paper proposes to leverage reconfigurable intelligent surface (RIS) aided AirComp for WDA in mmWave communication networks. To enhance the system performance, we formulate an optimization problem to minimize the mean-squared error (MSE) of WDA by jointly optimizing the receive beamforming vector of the access point, the transmit scalars of devices, and the phase-shift matrix of the RIS. To this end, we derive the closed-form expression of transmit scalars and then propose a Riemannian conjugate gradient algorithm, which can efficiently tackle the unit-modulus constraints with a low computational complexity. Compared to the baseline algorithms, simulation results reveal that the proposed algorithm achieves a faster convergence rate and a smaller MSE.
Zhibin Wang 0003, Hongbin Zhu, Yuanming Shi, Yong Zhou 0006
VTC Spring4
2022 Faster Activity and Data Detection in Massive Random Access: A Multiarmed Bandit Approach
abstract
This article investigates the grant-free random access mechanism for massive Internet of Things (IoT) devices. By embedding the data symbols in the signature sequences, joint device activity detection and data decoding can be achieved, which, however, significantly increases the computational complexity. Coordinate descent algorithms that enjoy a low per-iteration complexity have been employed to solve this detection problem, but previous works typically employ a random coordinate selection policy which leads to slow convergence. In this article, we develop multiarmed bandit (MAB) approaches for more efficient detection via coordinate descent, which achieves a delicate tradeoff betweenexplorationandexploitationin coordinate selection. Specifically, we first propose a bandit-based strategy, i.e., Bernoulli sampling, to speed up the convergence rate of coordinate descent, by learning which coordinates will result in more aggressive descent of thenonconvex objective function. To further improve the convergence rate, an inner MAB problem is established to learn the exploration policy of Bernoulli sampling. Both convergence rate analysis and simulation results are provided to show that the proposed bandit-based algorithms enjoy faster convergence rates with a lower time complexity compared with the state-of-the-art algorithm. Furthermore, our proposed algorithms are generally applicable to different scenarios, e.g., massive random access with low-precision analog-to-digital converters (ADCs).
Jialin Dong, Jun Zhang 0004, Yuanming Shi, Hui Wang 0011
IEEE Internet Things J.3
2022 Differentially Private Federated Learning via Reconfigurable Intelligent Surface
abstract
Federated learning (FL), as a disruptive machine learning (ML) paradigm, enables the collaborative training of a global model over decentralized local data sets without sharing them. It spans a wide scope of applications from the Internet of Things (IoT) to biomedical engineering and drug discovery. To support low-latency and high-privacy FL over wireless networks, in this article, we propose a reconfigurable intelligent surface (RIS)-empowered over-the-air FL system to alleviate the dilemma between learning accuracy and privacy. This is achieved by simultaneously exploiting the channel propagation reconfigurability with RIS for boosting the received signal power, as well as the waveform superposition property with over-the-air computation (AirComp) for fast model aggregation. By considering a practical scenario, where high-dimensional local model updates are transmitted across multiple communication blocks, we characterize the convergence behaviors of the differentially private federated optimization algorithm. We further formulate a system optimization problem to optimize the learning accuracy while satisfying privacy and power constraints via the joint design of transmit power, artificial noise, and phase shifts at RIS, for which a two-step alternating minimization framework is developed. Simulation results validate our systematic, theoretical, and algorithmic achievements and demonstrate that RIS can achieve a better tradeoff between privacy and accuracy for over-the-air FL systems.
Yong Zhou 0006, Youlong Wu, Yuanming Shi
IEEE Internet Things J.4
2022 Edge Artificial Intelligence for 6G: Vision, Enabling Technologies, and Applications
abstract
The thriving of artificial intelligence (AI) applications is driving the further evolution of wireless networks. It has been envisioned that 6G will be transformative and will revolutionize the evolution of wireless from “connected things” to “connected intelligence”. However, state-of-the-art deep learning and big data analytics based AI systems require tremendous computation and communication resources, causing significant latency, energy consumption, network congestion, and privacy leakage in both of the training and inference processes. By embedding model training and inference capabilities into the network edge, edge AI stands out as a disruptive technology for 6G to seamlessly integrate sensing, communication, computation, and intelligence, thereby improving the efficiency, effectiveness, privacy, and security of 6G networks. In this paper, we shall provide our vision for scalable and trustworthy edge AI systems with integrated design of wireless communication strategies and decentralized machine learning models. New design principles of wireless networks, service-driven resource allocation optimization methods, as well as a holistic end-to-end system architecture to support edge AI will be described. Standardization, software and hardware platforms, and application scenarios are also discussed to facilitate the industrialization and commercialization of edge AI systems.
Khaled Ben Letaief, Yuanming Shi, Jianmin Lu, Jianhua Lu
IEEE J. Sel. Areas Commun.2
2022 Interference Management for Over-the-Air Federated Learning in Multi-Cell Wireless Networks
abstract
Federated learning (FL) over resource-constrained wireless networks has recently attracted much attention. However, most existing studies consider one FL task in single-cell wireless networks and ignore the impact of downlink/uplink inter-cell interference on the learning performance. In this paper, we investigate FL over a multi-cell wireless network, where each cell performs a different FL task and over-the-air computation (AirComp) is adopted to enable fast uplink gradient aggregation. We conduct convergence analysis of AirComp-assisted FL systems, taking into account the inter-cell interference in both the downlink and uplink model/gradient transmissions, which reveals that the distorted model/gradient exchanges induce a gap to hinder the convergence of FL. We characterize the Pareto boundary of the error-induced gap region to quantify the learning performance trade-off among different FL tasks, based on which we formulate an optimization problem to minimize the sum of error-induced gaps in all cells. To tackle the coupling between the downlink and uplink transmissions as well as the coupling among multiple cells, we propose a cooperative multi-cell FL optimization framework to achieve efficient interference management for downlink and uplink transmission design. Results demonstrate that our proposed algorithm achieves much better average learning performance over multiple cells than non-cooperative baseline schemes.
Zhibin Wang 0003, Yong Zhou 0006, Yuanming Shi, Weihua Zhuang
IEEE J. Sel. Areas Commun.3
2022 A Proximal Iteratively Reweighted Approach for Efficient Network Sparsification
abstract
The huge size of deep neural networks makes it difficult to deploy on the embedded platforms with limited computation resources directly. In this article, we propose a novel trimming approach to determine the redundant parameters of the trained deep neural network in a layer-wise manner to produce a compact neural network. This is achieved by minimizing a nonconvex sparsity-inducing term of the network parameters while maintaining the response close to the original one. We present a proximal iteratively reweighted method to resolve the resulting nonconvex model, which approximates the nonconvex objective by a weighted l1 norm of the network parameters. Moreover, to alleviate the computational burden, we develop a novel termination criterion during the subproblem solution, significantly reducing the total pruning time. Global convergence analysis and a worst-case O(1/k) ergodic convergence rate for our proposed algorithm is established. Numerical experiments demonstrate the proposed approach is efficient and reliable.
Hao Wang 0045, Yuanming Shi, Jun Lin 0001
IEEE Trans. Computers3
2022 Sparse and Low-Rank Optimization for Pliable Index Coding via Alternating Projection
abstract
Pliable index coding (PICOD) has recently been regarded as a promising solution that exploits the coding advantage to improve communication efficiency of content-type systems (e.g., recommendation system), where clients are pliable and are interested in receiving any new message that they do not have. PICOD aims to find an effective coding strategy that satisfies the demands of all clients with the minimum number of transmissions. However, most of the previous works mainly provided theoretical understanding on PICOD in special instances based on greedy algorithms. In contrast, in this paper, we present a flexible sparse and low-rank matrix modeling approach to minimize the number of transmissions for the general PICOD problems. This is achieved by establishing generalized pliable alignment conditions to guarantee the requirements of all clients. As the resulting non-convex problem is highly intractable, we further develop an alternating pursuit framework to detect the rank of the matrix to be recovered by using the rank-increasing strategy. To address the feasibility-detection issues in the existing methods, we propose an alternating projection algorithm, which admits closed-form expressions and avoids excessive sparsity inducing. Moreover, we establish the global convergence of the alternating projection algorithm with random initial points. Simulation results demonstrate that the proposed alternating pursuit algorithm significantly reduces the number of transmissions compared to the state-of-the-art methods.
Min Fu 0003, Tao Jiang 0016, Hayoung Choi, Yong Zhou 0006, Yuanming Shi
IEEE Trans. Commun.5
2022 Decentralized Multi-Agent Power Control in Wireless Networks With Frequency Reuse
abstract
Many of the existing optimization-based transmit power control algorithms suffer from high computational complexity and require instantaneous global channel state information (CSI), both of which hinder their practical implementation. In this paper, we consider a wireless network with multiple transmitter-receiver pairs, where each transmitter only has access to its local CSI fed back from its intended receiver and does not require local CSI exchange with its neighboring transmitters. In such a network scenario, we propose a deep reinforcement learning based decentralized multi-agent power control (DEC-MAPC) algorithm for sum-rate maximization, where each transmitter acts as an intelligent agent. By leveraging the value decomposition technique, we establish a nonlinear mapping from the local reward of each agent to the global reward. Such a design allows each agent to independently control its transmit power based on its local CSI while enabling global collaboration among the agents. The proposed algorithm is scalable to large-scale networks as only local CSI is required, and is robust to the channel and interference variations via interacting with the environment. Simulation results show that the proposed DEC-MAPC algorithm with local CSI achieves comparable sum-rate performance with the centralized optimization algorithms with global CSI, while significantly reducing the computational complexity.
Jun Zong, Yong Zhou 0006, Yuanming Shi, Vincent W. S. Wong 0001
IEEE Trans. Commun.4
2022 UAV Aided Over-the-Air Computation
abstract
Different from the existing works that focus on transceiver design of over-the-air computation (AirComp) over static networks, we in this paper consider an unmanned aerial vehicle (UAV) aided AirComp system, where the UAV as a flying base station aggregates data from mobile sensors. The trajectory design of the UAV provides an additional degree of freedom to improve the performance of AirComp. We aim to minimize the time-averaged mean-squared error (MSE) of AirComp by jointly optimizing the UAV trajectory, receive normalizing factors, and sensors’ transmit power. To this end, we first propose a novel and equivalent problem transformation by introducing intermediate variables. This reformulation leads to a convex subproblem when fixing any other two blocks of variables, thereby enabling efficient algorithm design based on the principle of block coordinate descent and alternating direction method of multipliers (ADMM) techniques. In particular, we derive the optimal closed-form solutions for normalizing factors and intermediate variables optimization subproblems. We also recast the convex trajectory design subproblem into an ADMM form and obtain the closed-form expressions for each variable updating. Simulation results show that the proposed algorithm achieves a smaller time-averaged MSE while reducing the simulation time by orders of magnitude compared to state-of-the-art algorithms.
Min Fu 0003, Yong Zhou 0006, Yuanming Shi, Wei Chen 0002, Rui Zhang 0006
IEEE Trans. Wirel. Commun.3
2022 Reconfigurable Intelligent Surface Assisted Massive MIMO With Antenna Selection
abstract
Antenna selection is capable of reducing the hardware complexity of massive multiple-input multiple-output (MIMO) networks at the cost of certain performance degradation. Reconfigurable intelligent surface (RIS) has emerged as a cost-effective technique that can enhance the spectrum-efficiency of wireless networks by reconfiguring the propagation environment. By employing RIS to compensate for the performance loss due to antenna selection, in this paper we propose a new network architecture, i.e., RIS-assisted massive MIMO system with antenna selection, to enhance the system performance while enjoying a low hardware cost. This is achieved by maximizing the channel capacity via joint antenna selection and passive beamforming while taking into account the cardinality constraint of active antennas and the unit-modulus constraints of all RIS elements. However, the formulated problem turns out to be highly intractable due to the non-convex constraints and coupled optimization variables, for which an alternating optimization framework is provided, yielding antenna selection and passive beamforming subproblems. The computationally efficient submodular optimization algorithms are developed to solve the antenna selection subproblem under different channel state information assumptions. The iterative algorithms based on block coordinate descent are further proposed for the passive beamforming design by exploiting the unique problem structures. Moreover, the proposed algorithms are feasible to any finite number of antennas, and thus can be applicable in both ordinary MIMO and massive MIMO settings. Experimental results will demonstrate the algorithmic advantages and desirable performance of the proposed algorithms for RIS-assisted massive MIMO systems with antenna selection.
Jinglian He, Kaiqiang Yu, Yuanming Shi, Yong Zhou 0006, Wei Chen 0002, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.3
2022 Algorithm Unrolling for Massive Access via Deep Neural Networks With Theoretical Guarantee
abstract
Massive access is a critical design challenge of Internet of Things (IoT) networks. In this paper, we consider the grant-free uplink transmission of an IoT network with a multiple-antenna base station (BS) and a large number of single-antenna IoT devices. Taking into account the sporadic nature of IoT devices, we formulate the joint activity detection and channel estimation (JADCE) problem as a group-sparse matrix estimation problem. This problem can be solved by applying the existing compressed sensing techniques, which however either suffer from high computational complexities or lack of algorithm robustness. To this end, we propose a novel algorithm unrolling framework based on the deep neural network to simultaneously achieve low computational complexity and high robustness for solving the JADCE problem. Specifically, we map the original iterative shrinkage thresholding algorithm (ISTA) into an unrolled recurrent neural network (RNN), thereby improving the convergence rate and computational efficiency through end-to-end training. Moreover, the proposed algorithm unrolling approach inherits the structure and domain knowledge of the ISTA, thereby maintaining the algorithm robustness, which can handle non-Gaussian preamble sequence matrix in massive access. With rigorous theoretical analysis, we further simplify the unrolled network structure by reducing the redundant training parameters. Furthermore, we prove that the simplified unrolled deep neural network structures enjoy a linear convergence rate. Extensive simulations based on various preamble signatures show that the proposed unrolled networks outperform the existing methods in terms of the convergence rate, robustness and estimation accuracy.
Yandong Shi, Hayoung Choi, Yuanming Shi, Yong Zhou 0006
IEEE Trans. Wirel. Commun.3
2022 Federated Learning via Intelligent Reflecting Surface
abstract
Over-the-air computation (AirComp) based federated learning (FL) is capable of achieving fast model aggregation by exploiting the waveform superposition property of multiple-access channels. However, the model aggregation performance is severely limited by the unfavorable wireless propagation channels. In this paper, we propose to leverage intelligent reflecting surface (IRS) to achieve fast yet reliable model aggregation for AirComp-based FL. To optimize the learning performance, we present the convergence analysis of our proposed IRS-assisted AirComp-based FL system, based on which we propose to maximize the number of scheduled devices of each communication round under certain mean-squared error (MSE) requirements. To tackle the formulated highly-intractable problem, we propose a two-step optimization framework. Specifically, we induce the sparsity of device selection in the first step, followed by solving a series of MSE minimization problems to find the maximum feasible device set in the second step. We then propose an alternating optimization framework, supported by the difference-of-convex programming for low-rank optimization, to efficiently design the aggregation beamformers at the BS and phase shifts at the IRS. Simulation results demonstrate that our proposed algorithm and the deployment of an IRS can achieve a higher FL prediction accuracy than the baseline schemes.
Zhibin Wang 0003, Jiahang Qiu, Yong Zhou 0006, Yuanming Shi, Liqun Fu 0001, Wei Chen 0002, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.4
2022 Over-the-Air Federated Learning via Second-Order Optimization
abstract
Federated learning (FL) is a promising learning paradigm that can tackle the increasingly prominent isolated data islands problem while keeping users’ data locally with privacy and security guarantees. However, FL could result in task-oriented data traffic flows over wireless networks with limited radio resources. To design communication-efficient FL, most of the existing studies employ the first-order federated optimization approach that has a slow convergence rate. This however results in excessive communication rounds for local model updates between the edge devices and edge server. To address this issue, in this paper, we instead propose a novel over-the-air second-order federated optimization algorithm to simultaneously reduce the communication rounds and enable low-latency global model aggregation. This is achieved by exploiting the waveform superposition property of a multi-access channel to implement the distributed second-order optimization algorithm over wireless networks. The convergence behavior of the proposed algorithm is further characterized, which reveals a linear-quadratic convergence rate with an accumulative error term in each iteration. We thus propose a system optimization approach to minimize the accumulated error gap by joint device selection and beamforming design. Numerical results demonstrate the system and communication efficiency compared with the state-of-the-art approaches.
Peng Yang 0027, Yuning Jiang 0002, Ting Wang 0001, Yong Zhou 0006, Yuanming Shi, Colin N. Jones
IEEE Trans. Wirel. Commun.5
2021 Capacity Region of Intelligent Reflecting Surface Aided Wireless Networks via Active Learning
abstract
Intelligent Reflecting Surface (IRS) is a promising technology that is able to manipulate the wireless propagation channels via smartly adjusting the signal reflection. With continuous phase shifts, IRS has been shown to be effective in enlarging the achievable rate region. In this paper, we investigate the achievable rate region of a IRS-aided multi-user interference channel, where the phase shifts at the IRS can only take a finite number of discrete values. We formulate a multi-objective optimization problem (MOOP) to characterize the achievable rate region. The commonly adopted approaches such as the rate profile method fail to solve MOOP with optimization variables. Although the exhaustive search method can obtain the Pareto-optimal solutions, it suffers from high computational complexity. To this end, we propose a computationally efficient active learning algorithm via Gaussian process (GP). By modeling the objectives of MOOP as a draw from a GP distribution with only a few randomly computed rate-tuples, the active learning algorithm can quickly dominate the non-optimal points and find Pareto-optimal points without calculating rate-tuples. Numerical simulations demonstrate that the achievable rate region of IRS-aided interference channel is much larger than that without IRS and the proposed active learning framework obtains near-optimal Pareto solutions with a much lower computational complexity than the traditional exhaustive search algorithm.
Yandong Shi, Min Fu 0003, Yong Zhou 0006, Yuanming Shi
GLOBECOM4
2021 Learning Proximal Operator Methods for Massive Connectivity in IoT Networks
abstract
Grant-free random access has the potential to sup-port massive connectivity in Internet of Things (IoT) networks, where joint activity detection and channel estimation (JADCE) is a key issue that needs to be tackled. The existing methods for JADCE usually suffer from one of the following limitations: high computational complexity, ineffective in inducing sparsity, and incapable of handling complex matrix estimation. To mitigate all the aforementioned limitations, we in this paper develop an effective unfolding neural network framework built upon the proximal operator method to tackle the JADCE problem in IoT networks, where the base station is equipped with multiple antennas. Specifically, the JADCE problem is formulated as a group-sparse-matrix estimation problem, which is regularized by non-convex minimax concave penalty (MCP). This problem can be iteratively solved by using the proximal operator method, based on which we develop a unfolding neural network structure by parameterizing the algorithmic iterations. By further exploiting the coupling structure among the training parameters as well as the analytical computation, we develop two additional unfolding structures to reduce the training complexity. We prove that the proposed algorithm achieves a linear convergence rate. Results show that our proposed three unfolding structures not only achieve a faster convergence rate but also obtain a higher estimation accuracy than the baseline methods.
Yinan Zou, Yong Zhou 0006, Yuanming Shi, Xu Chen 0004
GLOBECOM3
2021 UAV-Assisted Over-the-Air Computation
abstract
Over-the-air computation (AirComp) provides a promising way to support ultrafast aggregation of distributed data. However, its performance cannot be guaranteed in long-distance transmission due to the distortion induced by the channel fading and noise. To unleash the full potential of AirComp, this paper proposes to use a low-cost unmanned aerial vehicle (UAV) acting as a mobile base station to assist AirComp systems. Specifically, due to its controllable high-mobility and high-altitude, the UAV can move sufficiently close to the sensors to enable line-of-sight transmission and adaptively adjust all the links' distances, thereby enhancing the signal magnitude alignment and noise suppression. Our goal is to minimize the time-averaging mean-square error for AirComp by jointly optimizing the UAV trajectory, the scaling factor at the UAV, and the transmit power at the sensors, under constraints on the UAV’s predetermined locations and flying speed, sensors’ average and peak power limits. However, due to the highly coupled optimization variables and time-dependent constraints, the resulting problem is non-convex and challenging. We thus propose an efficient iterative algorithm by applying the block coordinate descent and successive convex optimization techniques. Simulation results verify the convergence of the proposed algorithm and demonstrate the performance gains and robustness of the proposed design compared with benchmarks.
Min Fu 0003, Yong Zhou 0006, Yuanming Shi, Ting Wang 0001, Wei Chen 0002
ICC3
2021 Fast Convergence Algorithm for Analog Federated Learning
abstract
In this paper, we consider federated learning (FL) over a noisy fading multiple access channel (MAC), where an edge server aggregates the local models transmitted by multiple end devices through over-the-air computation (AirComp). To realize efficient analog federated learning over wireless channels, we propose an AirComp-based FedSplit algorithm, where a threshold-based device selection scheme is adopted to achieve reliable local model uploading. In particular, we analyze the performance of the proposed algorithm and prove that the proposed algorithm linearly converges to the optimal solutions under the assumption that the objective function is strongly convex and smooth. We also characterize the robustness of proposed algorithm to the ill-conditioned problems, thereby achieving fast convergence rates and reducing communication rounds. A finite error bound is further provided to reveal the relationship between the convergence behavior and the channel fading and noise. Our algorithm is theoretically and experimentally verified to be much more robust to the ill-conditioned problems with faster convergence compared with other benchmark FL algorithms.
Shuhao Xia, Jingyang Zhu, Yong Zhou 0006, Yuanming Shi, Wei Chen 0002
ICC5
2021 Saddle Point Approximation Based Delay Analysis for Wireless Federated Learning
abstract
Wireless federated learning (FL) holds the potential of preserving data privacy and reducing network traffic congestion, thereby attracting much recent attention. Due to the fading nature of wireless channels, wireless FL suffers from the random delay in each uplink and downlink transmission. As a result, how to analyze the overall random delay of a FL task over wireless fading channels remains open. To solve this challenging problem, we present a saddle point approximation based approach to obtain the distribution of the delay caused by communication in wireless FL systems. In particular, we obtain the uplink delay distribution and the downlink delay distribution by Lugannani-Rice formula. The overall delay distribution is then obtained through the convolution of those two distributions and the generating function. Simulation results demonstrate that the theoretical results provide accurate characterizations for the empirical results, which corroborates the validity of the analysis in this paper.
Longwei Yang, Xin Guo 0008, Yuanming Shi, Haiming Wang 0002, Wei Chen 0002
ICC4
2021 Communication-Efficient Quantized SGD for Learning Polynomial Neural Network
abstract
This paper establishes the convergence rates for fitting a polynomial neural network with quadratic activation function via the mini-batch Stochastic Gradient Descent (SGD) algorithm. Specifically, we focus on the parallel implementation of calculating mini-batch gradients on a distributed computing platform. We first illustrate that the SGD converges at a linear rate to the optimal solution, and the convergence rate can be characterized as a function of mini-batch sizes. Next, we deploy the SGD with a distributed approach across multiple processors, where the partial mini-batch gradient is calculated and quantized to send to a master processor in each iteration, yielding a Quantized Stochastic Gradient Descent (QSGD) algorithm. This scheme can effectively reduce the communication overhead by the quantization strategy. Furthermore, we reveal that QSGD provably maintains a similar convergence rate of SGD to a globally optimal solution while significantly reduces the communication cost. In particular, the number of bits required for quantization and the mini-batch size affect the convergence rate of QSGD.
Zhanpeng Yang, Yong Zhou 0006, Youlong Wu, Yuanming Shi
IPCCC4
2021 Over-the-Air Decentralized Federated Learning
abstract
In this paper, we consider decentralized federated learning (FL) over wireless networks, where over-the-air computation (AirComp) is adopted to facilitate the local model consensus in a device-to-device (D2D) communication manner. However, the AirComp-based consensus phase brings the additive noise in each algorithm iterate and the consensus needs to be robust to wireless network topology changes, which introduce a coupled and novel challenge of establishing the convergence for wireless decentralized FL algorithm. To facilitate consensus phase, we propose an AirComp-based DSGD with gradient tracking and variance reduction (DSGT-VR) algorithm, where both precoding and decoding strategies are developed for D2D communication. Furthermore, we prove that the proposed algorithm converges linearly and establish the optimality gap for strongly convex and smooth loss functions, taking into account the channel fading and noise. The theoretical result shows that the additional error bound in the optimality gap depends on the number of devices. Extensive simulations verify the theoretical results and show that the proposed algorithm outperforms other benchmark decentralized FL algorithms over wireless networks.
Yandong Shi, Yong Zhou 0006, Yuanming Shi
ISIT3
2021 Robust Design for Reconfigurable Intelligent Surface Assisted Over-the-Air Computation
abstract
Distributed data aggregation is a critical design aspect in future Internet-of-Things (IoT) networks. Over-the-air computation (AirComp) is capable of achieving ultra-fast data aggregation by exploiting the superposition property of wireless channel. However, the performance of AirComp, measured by the mean-squared-error (MSE), is generally restricted by the unfavorable channel conditions and relies on the availability of perfect channel state information (CSI). In this paper, we propose to use reconfigurable intelligent surface (RIS) to assist the wireless data aggregation in IoT networks via AirComp in the presence of imperfect CSI. By taking into account the constraints of the transmit power at the devices and the unit modulus of the RIS, we formulate an optimization problem to jointly optimize the transmit power of IoT devices, the beamforming vector at the access point, and the phase-shift matrix at the RIS under the expectation-based channel uncertainty model. We present an alternating optimization method to solve this nonconvex problem. In each iteration, the transmit power and the receive beamformer are updated according to Karush-Kuhn-Tucker conditions and a closed-form solution, respectively. Moreover, we also develop a difference-of-convex algorithm to tackle the nonconvex rank-one constraint in the problem of optimizing the phase-shift matrix. Simulation results illustrate the robustness of the proposed algorithm in terms of minimizing the AirComp distortion.
Qiaochu An, Yong Zhou 0006, Yuanming Shi
WCNC3
2021 Wireless-Powered Over-the-Air Computation in Intelligent Reflecting Surface-Aided IoT Networks
abstract
Fast wireless data aggregation and efficient battery recharging are two critical design challenges of Internet-of-Things (IoT) networks. Over-the-air computation (AirComp) and energy beamforming (EB) turn out to be two promising techniques that can address these two challenges, necessitating the design of wireless-powered AirComp. However, due to severe channel propagation, the energy harvested by IoT devices may not be sufficient to support AirComp. In this article, we propose to leverage the intelligent reflecting surface (IRS) that is capable of dynamically reconfiguring the propagation environment to drastically enhance the efficiency of both downlink EB and uplink AirComp in IoT networks. Due to the coupled problems of downlink EB and uplink AirComp, we further propose the joint design of energy and aggregation beamformers at the access point, downlink/uplink phase-shift matrices at the IRS, and transmit power at the IoT devices, to minimize the mean-squared error (MSE), which quantifies the AirComp distortion. However, the formulated problem is a highly intractable nonconvex quadratic programming problem. To solve this problem, we first obtain the closed-form expressions of the energy beamformer and the device transmit power, and then develop an alternating optimization framework based on difference-of-convex programming to design the aggregation beamformers and IRS phase-shift matrices. Simulation results demonstrate the performance gains of the proposed algorithm over the baseline methods and show that deploying an IRS can significantly reduce the MSE of AirComp.
Zhibin Wang 0003, Yuanming Shi, Yong Zhou 0006, Ning Zhang 0007
IEEE Internet Things J.2
2021 Nonconvex and Nonsmooth Sparse Optimization via Adaptively Iterative Reweighted Methods
Hao Wang 0045, Fan Zhang 0060, Yuanming Shi
J. Glob. Optim.3
2021 Delay Analysis of Wireless Federated Learning Based on Saddle Point Approximation and Large Deviation Theory
abstract
Federated learning (FL) is a collaborative machine learning paradigm, which enables deep learning model training over a large volume of decentralized data residing in mobile devices without accessing clients’ private data. Driven by the ever increasing demand for model training of mobile applications or devices, a vast majority of FL tasks are implemented over wireless fading channels. Due to the time-varying nature of wireless channels, however, random delay occurs in both the uplink and downlink transmissions of FL. How to analyze the overall time consumption of a wireless FL task, or more specifically, a FL’s delay distribution, becomes a challenging but important open problem, especially for delay-sensitive model training. In this paper, we present a unified framework to calculate the approximate delay distributions of FL over arbitrary fading channels. Specifically, saddle point approximation, extreme value theory (EVT), and large deviation theory (LDT) are jointly exploited to find the approximate delay distribution along with its tail distribution, which characterizes the quality-of-service of a wireless FL system. Simulation results will demonstrate that our approximation method achieves a small approximation error, which vanishes with the increase of training accuracy.
Longwei Yang, Xin Guo 0008, Yuanming Shi, Haiming Wang 0002, Wei Chen 0002, Khaled Ben Letaief
IEEE J. Sel. Areas Commun.4
2021 Graph Neural Networks for Scalable Radio Resource Management: Architecture Design and Theoretical Analysis
abstract
Deep learning has recently emerged as a disruptive technology to solve challenging radio resource management problems in wireless networks. However, the neural network architectures adopted by existing works suffer from poor scalability and generalization, and lack of interpretability. A long-standing approach to improve scalability and generalization is to incorporate the structures of the target task into the neural network architecture. In this paper, we propose to apply graph neural networks (GNNs) to solve large-scale radio resource management problems, supported by effective neural network architecture design and theoretical analysis. Specifically, we first demonstrate that radio resource management problems can be formulated as graph optimization problems that enjoy a universal permutation equivariance property. We then identify a family of neural networks, namedmessage passing graph neural networks(MPGNNs). It is demonstrated that they not only satisfy the permutation equivariance property, but also can generalize to large-scale problems, while enjoying a high computational efficiency. For interpretablity and theoretical guarantees, we prove the equivalence between MPGNNs and a family of distributed optimization algorithms, which is then used to analyze the performance and generalization of MPGNN-based methods. Extensive simulations, with power control and beamforming as two examples, demonstrate that the proposed method, trained in an unsupervised manner with unlabeled samples, matches or even outperforms classic optimization-based algorithms without domain-specific knowledge. Remarkably, the proposed method is highly scalable and can solve the beamforming problem in an interference channel with 1000 transceiver pairs within 6 milliseconds on a single GPU.
Yifei Shen 0004, Yuanming Shi, Jun Zhang 0004, Khaled Ben Letaief
IEEE J. Sel. Areas Commun.2
2021 Over-the-Air Computation via Reconfigurable Intelligent Surface
abstract
Over-the-air computation (AirComp) is a disruptive technique for fast wireless data aggregation in Internet of Things (IoT) networks via exploiting the waveform superposition property of multiple-access channels. However, the performance of AirComp is bottlenecked by the worst channel condition among all links between the IoT devices and the access point. In this paper, a reconfigurable intelligent surface (RIS) assisted AirComp system is proposed to boost the received signal power and thus mitigate the performance bottleneck by reconfiguring the propagation channels. With an objective to minimize the AirComp distortion, we propose a joint design of AirComp transceivers and RIS phase-shifts, which however turns out to be a highly intractable non-convex programming problem. To this end, we develop a novel alternating minimization framework in conjunction with the successive convex approximation technique, which is proved to converge monotonically. To reduce the computational complexity, we transform the subproblem in each alternation as a smooth convex-concave saddle point problem, which is then tackled by proposing a Mirror-Prox method that only involves a sequence of closed-form updates. Simulations show that the computation time of the proposed algorithm can be two orders of magnitude smaller than that of the state-of-the-art algorithms, while achieving a similar distortion performance.
Wenzhi Fang, Yuning Jiang 0002, Yuanming Shi, Yong Zhou 0006, Wei Chen 0002, Khaled Ben Letaief
IEEE Trans. Commun.3
2021 Reconfigurable Intelligent Surface Empowered Downlink Non-Orthogonal Multiple Access
abstract
Power-domain non-orthogonal multiple access (NOMA) has become a promising technology to exploit the new dimension of the power domain to enhance the spectral efficiency of wireless networks. However, most existing NOMA schemes rely on the strong assumption that users’ channel gains are quite different, which may be invalid in practice. To unleash the potential of power-domain NOMA, we propose a reconfigurable intelligent surface (RIS)-empowered NOMA scheme to introduce desirable channel gain differences among the users by adjusting the phase shifts at the RIS. Our goal is to minimize the total transmit power by jointly optimizing the beamforming vectors at the base station, the phase-shift matrix at the RIS, and user ordering. To address challenge due to the highly coupled optimization variables, we present an alternating optimization framework to decompose the non-convex bi-quadratically constrained quadratic problem under a specific user ordering into two rank-one constrained matrices optimization problems via matrix lifting. To accurately detect the feasibility of the non-convex rank-one constraints and improve performance by avoiding early stopping in the alternating optimization procedure, we equivalently represent the rank-one constraint as the difference between nuclear norm and spectral norm. A difference-of-convex (DC) algorithm is further developed to solve the resulting DC programs via successive convex relaxation, followed by establishing the convergence of the proposed DC-based alternating optimization method. We further propose an efficient user ordering scheme with closed-form expressions, considering both the channel conditions and users’ target data rates. Simulation results validate the ability of an RIS in enlarging the channel-gain difference when the users’ original channel conditions are similar and the superiority of the proposed DC-based alternating optimization method in reducing the total transmit power.
Min Fu 0003, Yong Zhou 0006, Yuanming Shi, Khaled Ben Letaief
IEEE Trans. Commun.3
2020 Stochastic Beamforming for Reconfigurable Intelligent Surface Aided Over-the-Air Computation
abstract
Over-the-air computation (AirComp) is a promising technology that is capable of achieving fast data aggregation in Internet of Things (IoT) networks. The mean-squared error (MSE) performance of AirComp is bottlenecked by the unfavorable channel conditions. This limitation can be mitigated by deploying a reconfigurable intelligent surface (RIS), which reconfigures the propagation environment to facilitate the receiving power equalization. The achievable performance of RIS relies on the availability of accurate channel state information (CSI), which however is generally difficult to be obtained. In this paper, we consider an RIS-aided AirComp IoT network, where an access point (AP) aggregates sensing data from distributed devices. Without assuming any prior knowledge on the underlying channel distribution, we formulate a stochastic optimization problem to maximize the probability that the MSE is below a certain threshold. The formulated problem turns out to be non-convex and highly intractable. To this end, we propose a data-driven approach to jointly optimize the receive beamforming vector at the AP and the phase-shift vector at the RIS based on historical channel realizations. After smoothing the objective function by adopting the sigmoid function, we develop an alternating stochastic variance reduced gradient (SVRG) algorithm with a fast convergence rate to solve the problem. Simulation results demonstrate the effectiveness of the proposed algorithm and the importance of deploying an RIS in reducing the MSE outage probability.
Wenzhi Fang, Min Fu 0003, Kunlun Wang 0001, Yuanming Shi, Yong Zhou 0006
GLOBECOM4
2020 Age of Aggregated Information: Timely Status Update with Over-the-Air Computation
abstract
Fast wireless data aggregation is a critical design challenge in Internet-of-Things (IoT) networks. In this paper, we consider a real-time status update IoT network, where an access point (AP) aims to aggregate data from multiple IoT devices using over-the-air computation (AirComp). To evaluate the freshness of the aggregated data at the AP, we propose the metric of age of aggregated information (AoAI), extended from the age of information (AoI), which is defined as the time elapsed since the generation of the latest valid aggregated data received at the AP. An aggregated status update is considered to be valid if the AirComp distortion, quantified by the mean-squarederror (MSE), is smaller than a pre-determined threshold. We formulate a constrained Markov decision process (MDP) problem for minimizing the average AoAI subject to the average transmit power constraint of each IoT device. The formulated constrained MDP problem is then reformulated as an unconstrained MDP problem by using the Lagrangian approach. By analyzing the structure of the MDP, we propose a state aggregation procedure to reduce the computational complexity. We further propose both offline and online scheduling algorithms to solve the problem. Simulation results show that the proposed algorithms significantly outperform the baseline algorithm with a fixed scheduling threshold in terms of the AoAI, and also strike a good balance between the AoAI and the total power consumption.
Jie Li 0002, Yong Zhou 0006, He Henry Chen, Yuanming Shi
GLOBECOM4
2020 Bandit Sampling for Faster Activity and Data Detection in Massive Random Access
abstract
This paper considers the grant-free random access scheme in IoT networks with a massive number of devices. By embedding the data symbols in the signature sequences, joint device activity detection, and data decoding can be achieved, which, however, significantly increases the computational complexity. Coordinate descent algorithms, with a low per-iteration complexity, have been employed to solve the detection problem, but previous works typically employ a random coordinate selection policy which leads to slow convergence. This paper develops a bandit based strategy, i.e., bandit sampling, to speed up the convergence of coordinate descent. We exploit a multi-armed bandit algorithm to learn which coordinates will result in more aggressive descent of the objective function. Both convergence rate analysis and simulation results are provided to show that the proposed algorithm enjoys a faster convergence rate with a lower time complexity compared with the state-of-the-art algorithm.
Jialin Dong, Jun Zhang 0004, Yuanming Shi
ICASSP3
2020 Intelligent Reflecting Surface for Massive Device Connectivity: Joint Activity Detection and Channel Estimation
abstract
Intelligent Reflecting Surface (IRS) has been a promising solution to enhance wireless networks both spectral-efficiently and energy-efficiently. This paper considers an IRS-assisted the Internet of Things network for massive connectivity. We aim to solve the IRS-related activity detection and channel estimation problem which has not been studied before. In this paper, we formulate the IRS-related activity detection and channel estimation problem as sparse matrix factorization, matrix completion and multiple measurement vector problem and, we propose a three-stage framework based on the approximate message passing. Simulation results verify the effectiveness of the proposed algorithm.
Shuhao Xia, Yuanming Shi
ICASSP2
2020 Reconfigurable Intelligent Surface Assisted Non-Orthogonal Unicast and Broadcast Transmission
abstract
Layered-division-multiplexing (LDM) is a spectrum-efficient physical-layer technology that can simultaneously support multiple services with diversified quality of service (QoS) requirements. In this paper, we propose a reconfigurable intelligent surface (RIS) assisted LDM system, where the base station (BS) simultaneously transmits non-orthogonal unicast and broadcast messages to multiple users. Our goal is to minimize the transmit power of the BS, while taking into account the QoS requirements of all messages and the unit modulus constraint of the RIS. To this end, we formulate a joint phase-shift matrix optimization as well as unicast and broadcast beamformer design problem. However, the formulated problem is a non-convex bi-quadratic programming problem. After utilizing alternating optimization and matrix lifting techniques, we transform the problem into an alternating sequence of rank-constrained semidefinite programming (SDP) problems. By introducing a difference-of-convex (DC) representation for the rank-one constraints, we develop an efficient DC algorithm to solve the low-rank optimization problem. Simulation results demonstrate the performance gains of the proposed algorithm over the state-of-art methods in reducing the BS transmit power.
Qiaochu An, Yuanming Shi, Yong Zhou 0006
VTC Spring2
2020 Phase Retrieval via Difference of Convex Programming
abstract
In this paper, we consider the convolutional phase retrieval problem, which is a crucial and challenging problem in signal processing and wireless communication. Our goal is to recover a signal from phaseless measurements. To address the challenge of phaseless measurements, we recast the recovery problem into a low-rank matrix optimization problem via matrix lifting. To further exactly detect the feasibility of resulting fixed-rank constraint, we propose a novel difference of convex functions (DC) representation for the rank function by exploiting the difference between trace norm and spectral norm, followed by presenting an efficient algorithm to solve the resulting DC programming. The simulation results demonstrate that our proposed algorithm outperforms the existing convex relaxation methods in terms of signal recovery sample complexity and the noise robustness.
Jinglian He, Min Fu 0003, Kaiqiang Yu, Yuanming Shi
VTC Spring4
2020 Coordinated Passive Beamforming for Distributed Intelligent Reflecting Surfaces Network
abstract
Intelligent reflecting surface (IRS) is a proposing technology in 6G to enhance the performance of wireless networks by smartly reconfiguring the propagation environment with a large number of passive reflecting elements. However, current works mainly focus on single IRS-empowered wireless networks, where the channel rank deficiency problem has emerged. In this paper, we propose a distributed IRS-empowered communication network architecture, where multiple source-destination pairs communicate through multiple distributed IRSs. We further contribute to maximize the achievable sum-rates in this network via jointly optimizing the transmit power vector at the sources and the phase shift matrix with passive beamforming at all distributed IRSs. Unfortunately, this problem turns out to be non-convex and highly intractable, for which an alternating approach is developed via solving the resulting fractional programming problems alternatively. In particular, the closed-form expressions are proposed for coordinated passive beamforming at IRSs. The numerical results will demonstrate the algorithmic advantages and desirable performances of the distributed IRS-empowered communication network.
Jinglian He, Kaiqiang Yu, Yuanming Shi
VTC Spring3
2020 Reconfigurable Intelligent Surface Enhanced Cognitive Radio Networks
abstract
The cognitive radio (CR) network is a promising network architecture that meets the requirement of enhancing scarce radio spectrum utilization. Meanwhile, reconfigurable intelligent surfaces (RIS) is a promising solution to enhance the energy and spectrum efficiency of wireless networks by properly altering the signal propagation via tuning a large number of passive reflecting units. In this paper, we investigate the downlink transmit power minimization problem for the RIS-enhanced single-cell cognitive radio (CR) network coexisting with a single-cell primary radio (PR) network by jointly optimizing the transmit beamformers at the secondary user (SU) transmitter and the phase shift matrix at the RIS. The investigated problem is a highly intractable due to the coupled optimization variables and unit modulus constraint, for which an alternative minimization framework is presented. Furthermore, a novel difference-of-convex (DC) algorithm is developed to solve the resulting non-convex quadratic program by lifting it into a low-rank matrix optimization problem. We then represent non-convex rank-one constraint as a DC function by exploiting the difference between trace norm and spectral norm. The simulation results validate that our proposed algorithm outperforms the existing state-of-the-art methods.
Jinglian He, Kaiqiang Yu, Yong Zhou 0006, Yuanming Shi
VTC Fall4
2020 Noisy Demixing: Convex Relaxation Meets Nonconvex Optimization
abstract
This paper focuses on the noisy demixing problem for robust recovery of a sequence of source signals and impulse responses from a sum of their convolution with additional noise. There are two main prevalent paradigms for this nonconvex estimation problem. One is leveraging the convex relaxation approach to provide good theoretical sample complexity guarantees. Another one is based on the nonconvex optimization method to enjoy good computational complexity. However, both methods are explored separately at this stage. Instead, we shall develop a method bridging convex relaxation with nonconvex optimization in a rigorous theoretical way. In fact, we find that the solution of the convex relaxation approach and the critical point of nonconvex optimization method can be almost the same. Based on our work, the good theoretical sample complexity guarantee of convex relaxation approach can be applied to nonconvex optimization method. And on the other hand, the certain stability guarantees from nonconvex optimization method can be propagated to convex relaxation approach.
Shaoming Huang, Yong Zhou 0006, Yuanming Shi
VTC Fall3
2020 Multigroup Multicast Transmission via Intelligent Reflecting Surface
abstract
Intelligent reflecting surface (IRS) has recently attracted increasing research interests due to its great potential in enhancing the energy and spectrum efficiency for wireless networks. In this paper, we shall investigate the downlink transmit power optimization problem for an IRS-assisted multigroup multicast system, where the beamformers at the base station (BS) and the phase shifts at the IRS are jointly optimized. We take into account the quality-of-service (QoS) requirements of all users in each multicasting group as well as the constant-envelope reflection of all IRS elements. The formulated problem turns out to be a non-convex quadratically constrained bi-quadratic programming problem, due to the intricate coupling between the optimization variables and the non-convex constant-envelope constraints. To this end, we present the alternating optimization method with matrix lifting to decouple the optimization variables, followed by ensuring the feasibility of the rank-one constraints via introducing the difference-of-convex (DC) function representation. We develop an effective alternating algorithm to solve the joint optimization problem. Extensive simulation results show the superiority of the proposed algorithm in reducing the transmit power of multigroup multicast systems.
Qiaochu An, Yuanming Shi, Yong Zhou 0006
VTC Fall3
2020 Wirelessly Powered Data Aggregation via Intelligent Reflecting Surface Assisted Over-the-Air Computation
abstract
Fast wireless data aggregation and efficient battery recharging are two critical design challenges of Internet of Things (IoT) networks. Over-the-air computation (AirComp) and energy beamforming (EB) are two promising techniques that can tackle these two challenges. In this paper, we propose to leverage the intelligent reflecting surface (IRS) to drastically enhance the efficiency of both downlink EB and uplink AirComp in IoT networks by exploiting the passive beamforming gains at the IRS. Due to the coupled downlink EB and uplink AirComp, we propose the joint design of energy and aggregation beamformers at the access point, downlink/uplink phase-shift matrices at the IRS, and transmit power at the IoT devices to minimize the mean-squared-error (MSE), which quantifies the AirComp distortion. However, the formulated problem is a highly intractable nonconvex quadratic programming problem. To this end, we first obtain the closed-form expressions of the energy beamformer and the transmit power, and then propose an efficient algorithm that alternatively updates other variables using semidefinite relaxation (SDR) to solve the problem. Simulation results demonstrate the performance gains of the proposed algorithm over the baseline methods and show that deploying an IRS can significantly reduce the MSE of AirComp.
Zhibin Wang 0003, Yuanming Shi, Yong Zhou 0006
VTC Spring2
2020 Towards Reconfigurable Intelligent Surfaces Powered Green Wireless Networks
abstract
The adoption of reconfigurable intelligent surface (RIS) in wireless networks can enhance the spectrum- and energy-efficiency by controlling the propagation environment. Although the RIS does not consume any transmit power, the circuit power of the RIS cannot be ignored, especially when the number of reflecting elements is large. In this paper, we propose the joint design of beamforming vectors at the base station, active RIS set, and phase-shift matrices at the active RISs to minimize the network power consumption, including the RIS circuit power consumption, while taking into account each user's target data rate requirement and each reflecting element's constant modulus constraint. However, the formulated problem is a mixed-integer quadratic programming (MIQP) problem, which is NP-hard. To this end, we present an alternating optimization method, which alternately solves second order cone programming (SOCP) and MIQP problems to update the optimization variables. Specifically, the MIQP problem is further transformed into a semidefinite programming problem by applying binary relaxation and semidefinite relaxation. Finally, an efficient algorithm is developed to solve the problem. Simulation results show that the proposed algorithm significantly reduces the network power consumption and reveal the importance of taking into account the RIS circuit power consumption.
Min Fu 0003, Yuanming Shi, Yong Zhou 0006
WCNC3
2020 Energy-Efficient Processing and Robust Wireless Cooperative Transmission for Edge Inference
abstract
Edge machine learning can deliver low-latency and private artificial intelligent (AI) services for mobile devices by leveraging computation and storage resources at the network edge. This article presents an energy-efficient edge processing framework to execute deep learning inference tasks at the edge computing nodes whose wireless connections to mobile devices are prone to channel uncertainties. Aimed at minimizing the sum of computation and transmission power consumption with probabilistic Quality-of-Service (QoS) constraints, we formulate the joint inference tasking and the downlink beamforming problem that is characterized by a group sparse objective function. We provide a statistical learning-based robust optimization approach to approximate the highly intractable probabilistic-QoS constraints by nonconvex quadratic constraints, which are further reformulated as matrix inequalities with a rank-one constraint via matrix lifting. We design a reweighted power minimization approach by iteratively reweighted ℓ1minimization with difference-of-convex-functions (DC) regularization and updating weights, where the reweighted approach is adopted for enhancing group sparsity whereas the DC regularization is designed for inducing rank-one solutions. The numerical results demonstrate that the proposed approach outperforms other state-of-the-art approaches.
Kai Yang 0006, Yuanming Shi, Wei Yu 0001, Zhi Ding 0001
IEEE Internet Things J.2
2020 An Algebraic-Geometric Approach for Linear Regression Without Correspondences
abstract
Linear regression without correspondences is the problem of performing a linear regression fit to a dataset for which the correspondences between the independent samples and the observations are unknown. Such a problem naturally arises in diverse domains such as computer vision, data mining, communications and biology. In its simplest form, it is tantamount to solving a linear system of equations, for which the entries of the right hand side vector have been permuted. This type of data corruption renders the linear regression task considerably harder, even in the absence of other corruptions, such as noise, outliers or missing entries. Existing methods are either applicable only to noiseless data or they are very sensitive to initialization or they work only for partially shuffled data. In this paper we address these issues via an algebraic geometric approach, which uses symmetric polynomials to extract permutation-invariant constraints that the parameters ξ* ∈ Rnof the linear regression model must satisfy. This naturally leads to a polynomial system of n equations in n unknowns, which contains ξ* in its root locus. Using the machinery of algebraic geometry we prove that as long as the independent samples are generic, this polynomial system is always consistent with at most n! complex roots, regardless of any type of corruption inflicted on the observations. The algorithmic implication of this fact is that one can always solve this polynomial system and use its most suitable root as initialization to the Expectation Maximization algorithm. To the best of our knowledge, the resulting method is the first working solution for small values of n able to handle thousands of fully shuffled noisy observations in milliseconds.
Manolis C. Tsakiris, Liangzu Peng, Aldo Conca, Laurent Kneip, Yuanming Shi, Hayoung Choi
IEEE Trans. Inf. Theory5
2020 Ranking from Crowdsourced Pairwise Comparisons via Smoothed Riemannian Optimization
abstract
Social Internet of Things has recently become a promising paradigm for augmenting the capability of humans and devices connected in the networks to provide services. In social Internet of Things network, crowdsourcing that collects the intelligence of the human crowd has served as a powerful tool for data acquisition and distributed computing. To support critical applications (e.g., a recommendation system and assessing the inequality of urban perception), in this article, we shall focus on the collaborative ranking problems for user preference prediction from crowdsourced pairwise comparisons. Based on the Bradley--Terry--Luce (BTL) model, a maximum likelihood estimation (MLE) is proposed via low-rank approach in order to estimate the underlying weight/score matrix, thereby predicting the ranking list for each user. A novel regularized formulation with the smoothed surrogate of elementwise infinity norm is proposed in order to address the unique challenge of the coupled the non-smooth elementwise infinity norm constraint and non-convex low-rank constraint in the MLE problem. We solve the resulting smoothed rank-constrained optimization problem via developing the Riemannian trust-region algorithm on quotient manifolds of fixed-rank matrices, which enjoys the superlinear convergence rate. The admirable performance and algorithmic advantages of the proposed method over the state-of-the-art algorithms are demonstrated via numerical results. Moreover, the proposed method outperforms state-of-the-art algorithms on large collaborative filtering datasets in both success rate of inferring preference and normalized discounted cumulative gain.
Jialin Dong, Kai Yang 0006, Yuanming Shi
ACM Trans. Knowl. Discov. Data3
2020 LORM: Learning to Optimize for Resource Management in Wireless Networks With Few Training Samples
abstract
Effective resource management plays a pivotal role in wireless networks, which, unfortunately, typically results in challenging mixed-integer nonlinear programming (MINLP) problems. Machine learning-based methods have recently emerged as a disruptive way to obtain near-optimal performance for MINLPs with affordable computational complexity. There have been some attempts in applying such methods to resource management in wireless networks, but these attempts require huge amounts of training samples and lack the capability to handle constrained problems. Furthermore, they suffer from severe performance deterioration when the network parameters change, which commonly happens and is referred to as thetask mismatchproblem. In this paper, to reduce the sample complexity and address the feasibility issue, we propose a framework of Learning to Optimize for Resource Management (LORM). In contrast to the end-to-end learning approach adopted in previous studies, LORM learns the optimal pruning policy in the branch-and-bound algorithm for MINLPs via a sample-efficient method, namely,imitation learning. To further address the task mismatch problem, we develop a transfer learning method via self-imitation in LORM, namedLORM-TL, which can quickly adapt a pre-trained machine learning model to the new task with only a few additionalunlabeledtraining samples. Numerical simulations demonstrate that LORM outperforms specialized state-of-the-art algorithms and achieves near-optimal performance, while providing significant speedup compared with the branch-and-bound algorithm. Moreover, LORM-TL, by relying on a few unlabeled samples, achieves comparable performance with the model trained from scratch with sufficient labeled samples.
Yifei Shen 0004, Yuanming Shi, Jun Zhang 0004, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.2
2020 Federated Learning via Over-the-Air Computation
abstract
The stringent requirements for low-latency and privacy of the emerging high-stake applications with intelligent devices such as drones and smart vehicles make the cloud computing inapplicable in these scenarios. Instead, edge machine learning becomes increasingly attractive for performing training and inference directly at network edges without sending data to a centralized data center. This stimulates a nascent field termed as federated learning for training a machine learning model on computation, storage, energy and bandwidth limited mobile devices in a distributed manner. To preserve data privacy and address the issues of unbalanced and non-IID data points across different devices, the federated averaging algorithm has been proposed for global model aggregation by computing the weighted average of locally updated model at each selected device. However, the limited communication bandwidth becomes the main bottleneck for aggregating the locally computed updates. We thus propose a novel over-the-air computation based approach for fast global model aggregation via exploring the superposition property of a wireless multiple-access channel. This is achieved by joint device selection and beamforming design, which is modeled as a sparse and low-rank optimization problem to support efficient algorithms design. To achieve this goal, we provide a difference-of-convex-functions (DC) representation for the sparse and low-rank function to enhance sparsity and accurately detect the fixed-rank constraint in the procedure of device selection. A DC algorithm is further developed to solve the resulting DC program with global convergence guarantees. The algorithmic advantages and admirable performance of the proposed methodologies are demonstrated through extensive numerical results.
Kai Yang 0006, Tao Jiang 0016, Yuanming Shi, Zhi Ding 0001
IEEE Trans. Wirel. Commun.3
2019 Blind Demixing via Wirtinger Flow with Random Initialization
abstract
This paper concerns the problem of demixing a series of source signals from the sum of bilinear measurements. This problem spans diverse areas such as communication, imaging processing, machine learning, etc. However, semidefinite programming for blind demixing is prohibitive to large-scale problems due to high computational complexity and storage cost. Although several efficient algorithms have been developed recently that enjoy the benefits of fast convergence rates and even regularization free, they still call for spectral initialization. To find simple initialization approach that works equally well as spectral initialization, we propose to solve blind demixing problem via Wirtinger flow with random initialization, which yields a natural implementation. To reveal the efficiency of this algorithm, we provide the global convergence guarantee concerning randomly initialized Wirtinger flow for blind demixing. Specifically, it shows that with sufficient samples, the iterates of randomly initialized Wirtinger flow can enter a local region that enjoys strong convexity and strong smoothness within a few iterations at the first stage. At the second stage, iterates of randomly initialized Wirtinger flow further converge linearly to the ground truth.
Jialin Dong, Yuanming Shi
AISTATS2
2019 Over-the-Air Computation via Intelligent Reflecting Surfaces
abstract
Over-the-air computation (AirComp) becomes a promising approach for fast wireless data aggregation via exploiting the superposition property in a multiple access channel. To further overcome the unfavorable signal propagation conditions for AirComp, in this paper, we propose an intelligent reflecting surface (IRS) aided AirComp system to build controllable wireless environments, thereby boosting the received signal power significantly. This is achieved by smartly tuning the phase shifts for the incoming electromagnetic waves at IRS, resulting in reconfigurable signal propagations. Unfortunately, it turns out that the joint design problem for AirComp transceivers and IRS phase shifts becomes a highly intractable nonconvex bi-quadratic programming problem, for which a novel alternating difference-of-convex (DC) programming algorithm is developed. This is achieved by providing a novel DC function representation for the rank-one constraint in the low-rank matrix optimization problem via matrix lifting. Simulation results demonstrate the algorithmic advantages and admirable performance of the proposed approaches compared with the state-of-art solutions.
Tao Jiang 0016, Yuanming Shi
GLOBECOM2
2019 Randomized Sketching Based Beamforming for Massive MIMO
abstract
Massive MIMO system yields significant improvements in spectral and energy efficiency for future wireless communication systems. The regularized zero-forcing (RZF) beamforming is able to provide good performance with the capability of achieving numerical stability and robustness to the channel uncertainty. However, in massive MIMO systems, the matrix inversion operation in RZF beamforming becomes computationally expensive. To address this computational issue, we shall propose a novel randomized sketching based RZF beamforming approach with low computational latency. This is achieved by solving a linear system via randomized sketching based on the preconditioned Richard iteration, which guarantees high quality approximations to the optimal solution. We theoretically prove that the sequence of approximations obtained iteratively converges to the exact RZF beamforming matrix linearly fast as the number of iterations increases. Also, it turns out that the system sum-rate for such sequence of approximations converges to the exact one at a linear convergence rate. Our simulation results verify our theoretical findings.
Hayoung Choi, Tao Jiang 0016, Weijing Li, Yuanming Shi
GLOBECOM4
2019 Sparse Blind Demixing for Low-latency Signal Recovery in Massive Iot Connectivity
abstract
Internet-of-Things (IoT) networks are envisioned to typically include a massive number of devices with sporadic and low-latency uplink service needs. This paper presents a blind demixing approach to support the data recovery of multiple simultaneous and unscheduled device transmissions without a priori channel state information (CSI). The proposed joint receiver leverages the group sparse bilinear characteristics of the underlying problem that involves active device detection and data recovery. We exploit the manifold geometry of rank-one matrices in the lifted bilinear equation and apply smoothed ℓ1/ℓ2-norm to induce the group sparsity for active device detection. We further develop a smoothed Riemannian algorithm to solve the sparse blind demixing optimization problem. Numerical results demonstrate the algorithmic advantage and desirable performance of the proposed algorithm.
Jialin Dong, Yuanming Shi, Zhi Ding 0001
ICASSP2
2019 Pliable Data Shuffling for On-device Distributed Learning
abstract
Dataset reshuffling across mobile devices allows for speeding up on-device distributed machine learning, which however requires significant communication bandwidth. In this paper, we propose a pliable data shuffling approach to significantly reduce the communication cost for on-device distributed learning via joint data placement and transmission design. This is achieved by establishing the novel interference alignment conditions and diversity constraints for data shuffling to improve the statistical learning performance. Unfortunately, the presented pliable data shuffling problem is a highly intractable mixed combinatorial optimization problem, for which a novel sparse and low-rank framework is developed, supported by the computationally efficient difference-of-convex (DC) algorithm. Numerical results demonstrate that the proposed pliable data shuffling is able to significantly reduce the communication bandwidth while achieving desirable learning performance.
Tao Jiang 0016, Kai Yang 0006, Yuanming Shi
ICASSP3
2019 Layer-wise Deep Neural Network Pruning via Iteratively Reweighted Optimization
abstract
The huge number of parameters of deep neural network makes it difficult to deploy on embedded devices with limited hardware, computation, storage and energy resources. In this paper, we shall propose a log-sum minimization approach to prune a trained network layer by layer thereby improving the network compression ratio. Specifically, this is achieved by enhancing sparsity for network parameters such that the output of the network after pruning is consistent with the original one. We further present an iteratively reweighted algorithm to solve the nonconvex and nonsmooth log-sum minimization problem with general convex constraints. Furthermore, we show the existence of the cluster points for the iterates and the global convergence of the proposed iteratively reweighted algorithm. Numerical experiments demonstrate that the proposed approach is able to significantly prune the trained neural network while preserving the prediction accuracy.
Tao Jiang 0016, Yuanming Shi, Hao Wang 0045
ICASSP3
2019 Algebraically-initialized Expectation Maximization for Header-free Communication
abstract
Towards low-latency communication for short-packet transmission, this paper tackles the problem of shuffled linear regression for large-scale wireless sensor networks with header-free communication by using results from algebraic geometry as well as an alternating optimization scheme. The shuffled linear regression problem is to solve a linear system with shuffled entries of the right hand side vector. However, solving the shuffled linear system requires high computational cost. The key idea of our approach is to eliminate the shuffled structure via symmetric polynomials, which leads to a system of polynomial equations. Considering one of the solutions of the resulting polynomial system as an initialization to the Expectation Maximization algorithm, we propose the Algebraically-Initialized Expectation Maximization algorithm. Computational experiments with synthetic data show that our proposed algorithm is extensively efficient, and it performs well even with noise.
Liangzu Peng, Xuming Song, Manolis C. Tsakiris, Hayoung Choi, Laurent Kneip, Yuanming Shi
ICASSP6
2019 Learning Shallow Neural Networks via Provable Gradient Descent with Random Initialization
abstract
This paper presents the provable gradient descent algorithm with random initialization for learning a two-layer neural network with quadratic activation functions. Specifically, we focus on the under-parameterized regime where the number of hidden units is smaller than the dimension of the inputs. We reveal that the randomly initialized gradient descent for the nonconvex neural network training problem is able to enter a local region that enjoys strong convexity and strong smoothness within a few iterations, and then provably converges to a globally optimal model at a linear rate.
Shuhao Xia, Yuanming Shi
ICASSP2
2019 Transfer Learning for Mixed-Integer Resource Allocation Problems in Wireless Networks
abstract
Effective resource allocation plays a pivotal role in wireless networks. Unfortunately, typical resource allocation problems are mixed-integer nonlinear programming (MINLP) problems, which are NP-hard. Machine learning based methods recently emerge as a disruptive way to obtain near-optimal performance for MINLP problems with affordable computational complexity. However, they suffer from severe performance deterioration when the network parameters change, which commonly happens in practice and can be characterized as the task mismatch issue. In this paper, we propose a transfer learning method via self-imitation, to address this issue for effective resource allocation in wireless networks. It is based on a general “learning to optimize” framework for solving MINLP problems. A unique advantage of the proposed method is that it can tackle the task mismatch issue with a few additional unlabeled training samples, which is especially important when transferring to large-size problems. Numerical experiments demonstrate that the proposed method, with much less training time, achieves comparable performance with the model trained from scratch based on sufficient labeled samples. To the best of our knowledge, this is the first work that applies transfer learning for resource allocation in wireless networks.
Yifei Shen 0004, Yuanming Shi, Jun Zhang 0004, Khaled Ben Letaief
ICC2
2019 Federated Learning Based on Over-the-Air Computation
abstract
The rapid growth in storage capacity and computational power of mobile devices is making it increasingly attractive for devices to process data locally instead of risking privacy by sending them to the cloud or networks. This reality has stimulated a novel federated learning framework for training statistical machine learning models on mobile devices directly using decentralized data. However, communication bandwidth remains a bottleneck for globally aggregating the locally computed updates. This work presents a novel model aggregation approach by exploiting the natural signal superposition of wireless multiple-access channel. This over-the-air computation is achieved by joint device selection and receiver beamforming design to improve the statistical learning performance. To tackle the difficult mixed combinatorial optimization problem with nonconvex quadratic constraints, we propose a novel sparse and low-rank modeling approach and develop an efficient difference-of-convex-function (DC) algorithm. Our results demonstrate the algorithm's ability to aggregate results from more devices to deliver superior learning performance.
Kai Yang 0006, Tao Jiang 0016, Yuanming Shi, Zhi Ding 0001
ICC3
2019 Stochastic Submodular Maximization for Scalable Network Adaptation in Dense Cloud-RAN
abstract
We propose a stochastic submodular maximization approach to enable scalable network adaptation in dense cloud radio access networks (Cloud-RAN) without any prior knowledge of the underlying channel distribution. This is achieved by exploiting the submodular and monotone characteristics of the objective function, followed by equivalently lifting the discrete problem into the continuous domain via multilinear extension. Although maximizing the continuous submodular function is still nonconvex, the stochastic projected gradient method is able to provide strong approximation guarantees to the global maxima. As the channel distribution is unknown, we propose to access the unbiased estimate of gradient only based on the historically collected channel samples, thereby significantly reducing the channel signaling overhead. We further provide a fast way to compute the unbiased estimate of gradient by exploiting the algebraic structure in the continuous submodular function. Therefore, the proposed stochastic submodular maximization based network adaptation framework enjoys the benefits of low computational complexity and low channel signaling overhead. Simulation results demonstrate the algorithmic advantages and desirable performances of the proposed methods for network adaptation in dense Cloud-RAN.
Kaiqiang Yu, Jinglian He, Yuanming Shi
ICC3
2019 Sparse Blind Demixing for Low-Latency Wireless Random Access with Massive Connectivity
abstract
Massive connectivity has become a critical requirement for Internet-of-Things (IoT) networks, where a large number of devices need to connect to an access-point sporadically. Moreover, low-latency communication and sporadic device traffic are essential to support intelligent services in IoT networks. In this paper, to support low-latency communication for massive devices with sporadic traffic, we present a sparse blind demixing to simultaneously detect the active devices and decode multiple source signals without a priori channel state information in multi- in-multi-out (MIMO) networks. To address the unique challenges of bilinear measurements and sporadic device activity detection, we recast the estimation problem as a sparse and low-rank optimization problem via matrix lifting. We further propose a difference-of-convex-functions (DC) representation for the rank function to guarantee the exact rank constraint, followed by ignoring the non-convex group sparse function. This is achieved by exploiting the difference between nuclear norm and the convex Ky Fan k- norm for a rank function representation. We then develop an efficient DC algorithm to solve the resulting non-convex DC program without regularization parameter. Numerical results demonstrate that the proposed DC approach is able to exactly recover the ground truth signals with reduced sample sizes, as well as achieve better performance against noise compared with the existing convex methods.
Min Fu 0003, Jialin Dong, Yuanming Shi
VTC Fall3
2019 Blind Deconvolution Meets Phase Retrieval in Optical Wireless Communications
abstract
Optical wireless communication becomes a key enabling technology for achieving ultra-high data rate requirements in beyond 5G systems. In this paper, to reduce both channel signaling overhead and hardware cost in optical wireless communications, we present a blind deconvolutional phase retrieval approach to recover the source signals from phaseless measurements without a priori channel information. To deal with the coupled challenges of phaseless measurements and bilinear signaling model, we recast the signal recovery problem into a rank-one matrices recovery problem via matrix lifting, followed by relaxing each phaseless matrix measurement into its convex hull. We further propose a difference-of-convex-functions (DC) programming algorithm to solve the low-rank matrix optimization problem. This is achieved by proposing the DC representation for the rank function based on the convex Ky Fan k-norm, thereby exactly detecting the fixed-rank constraints. The numerical results demonstrate that the proposed DC approach outperforms the state-of-the-art methods in terms of signal recovery performance and the robustness to the noise.
Min Fu 0003, Yuanming Shi
VTC Fall2
2019 On-Device Federated Learning via Second-Order Optimization with Over-the-Air Computation
abstract
Federated learning becomes a promising approach for preserving privacy by keeping user data locally. The basic idea is that a central server iteratively aggregates distributed local models trained directly on mobile users' local datasets to form a high-quality global model by computing the weighted sum of the locally updated models. However, the communication cost becomes the main bottleneck as a large number of communication rounds are involved in the federated learning procedure. We propose to update local models by second- order optimization methods with fast convergence rates, thereby significantly reducing the communication rounds for global model updates. Furthermore, the over-the-air computation technique is adopted to improve communication efficiency for model aggregation by utilizing the superposition property of wireless channels. A nonconvex low-rank beamforming approach is then developed to support over- the-air computation via difference-of-convex-functions (DC) programming. Through extensive experiments, we reveal that the proposed DC algorithm is able to significantly minimize the aggregation error, and the second- order methods are quite robust to the model aggregation errors.
Sheng Hua, Kai Yang 0006, Yuanming Shi
VTC Fall3
2019 Deep Learning Tasks Processing in Fog-RAN
abstract
Recently, the demand on performing intelligent tasks for mobile devices with low-latency is ever- increasing, while the requirements of intensive computation and large storage size impedes the deployment of deep learning models directly on devices. Fog radio access network (Fog-RAN) offers a promising solution by integrating the computing power and storage of edge processing nodes (e.g., base stations). In this paper, we propose a joint deep learning task selection and downlink transmit beamforming approach to improve the communication efficiency while achieving green computing, i.e., minimizing the sum of computation power consumption for deep learning tasks and downlink transmit power consumption. For efficient algorithm design, we exploit the group sparsity structure of the aggregated beamforming vector and propose a log-sum based group sparse beamforming framework. An iteratively reweighted ℓ1algorithm is developed to solve the nonconvex and nonsmooth log-sum minimization problem. Furthermore, we derive the global convergence result of the iteratively reweighted ℓ1algorithm, which shows that it converges to the first-order stationary point from any feasible starting point. Simulation results demonstrate that our log-sum based group sparse beamforming approach for deep learning tasks processing is more efficient than state- of-the-art approaches in terms of power consumption.
Sheng Hua, Kai Yang 0006, Gao Yin, Yuanming Shi, Hao Wang 0045
VTC Fall5
2019 Joint Activity Detection and Channel Estimation for IoT Networks: Phase Transition and Computation-Estimation Tradeoff
abstract
Massive device connectivity is a crucial communication challenge for Internet of Things (IoT) networks, which consist of a large number of devices with sporadic traffic. In each coherence block, the serving base station needs to identify the active devices and estimate their channel state information for effective communication. By exploiting the sparsity pattern of data transmission, we develop a structured group sparsity estimation method to simultaneously detect the active devices and estimate the corresponding channels. This method significantly reduces the signature sequence length while supporting massive IoT access. To determine the optimal signature sequence length, we study the phase transition behavior of the group sparsity estimation problem. Specifically, user activity can be successfully estimated with a high probability when the signature sequence length exceeds a threshold; otherwise, it fails with a high probability. The location and width of the phase transition region are characterized via the theory of conic integral geometry. We further develop a smoothing method to solve the high-dimensional structured estimation problem with a given limited time budget. This is achieved by sharply characterizing the convergence rate in terms of the smoothing parameter, signature sequence length and estimation accuracy, yielding a tradeoff between the estimation accuracy and computational cost. Numerical results are provided to illustrate the accuracy of our theoretical results and the benefits of smoothing techniques.
Tao Jiang 0016, Yuanming Shi, Jun Zhang 0004, Khaled Ben Letaief
IEEE Internet Things J.2
2019 Blind Demixing for Low-Latency Communication
abstract
In next-generation wireless networks, low-latency communication is critical to support emerging diversified applications, e.g., tactile Internet and virtual reality. In this paper, a novel blind demixing approach is developed to reduce the channel signaling overhead, thereby supporting low-latency communication. Specifically, we develop a low-rank approach to recover the original information only based on the single observed vector without any channel estimation. To address the unique challenges of multiple non-convex rank-one constraints, the quotient manifold geometry of the product of complex symmetric rank-one matrices is exploited. This is achieved by equivalently reformulating the original problem that uses complex asymmetric matrices to the one that uses Hermitian positive semidefinite matrices. We further generalize the geometric concepts of the complex product manifold via element-wise extension of the geometric concepts of the individual manifolds. The scalable Riemannian optimization algorithms, i.e., the Riemannian gradient descent algorithm and the Riemannian trust-region algorithm, are then developed to solve the blind demixing problem efficiently with low iteration complexity and low iteration cost. The statistical analysis shows that the Riemannian gradient descent with spectral initialization is guaranteed to linearly converge to the ground truth signals provided sufficient measurements. In addition, the Riemannian trust-region algorithm is provable to converge to an approximate local minimum from the arbitrary initialization point. Numerical experiments have been carried out in settings with different types of encoding matrices to demonstrate the algorithmic advantages, performance gains, and sample efficiency of the Riemannian optimization algorithms.
Jialin Dong, Kai Yang 0006, Yuanming Shi
IEEE Trans. Wirel. Commun.3
2019 Generalized Low-Rank Optimization for Topological Cooperation in Ultra-Dense Networks
abstract
Network densification is a natural way to support dense mobile applications under stringent requirements, such as ultra-low latency, ultra-high data rate, and massive connecting devices. Severe interference in ultra-dense networks poses a key bottleneck. Sharing channel state information (CSI) and messages across transmitters can potentially alleviate the interferences and improve the system performance. Most existing works on interference coordination require significant CSI signaling overhead and are impractical in the ultra-dense networks. This paper investigates the topological cooperation to manage interferences in message sharing based only on the network connectivity information. In particular, we propose a generalized low-rank optimization approach in a complex field to maximize the achievable degrees of freedom (DoFs) by establishing interference alignment conditions for the topological cooperation. To tackle the challenges of poor structure and non-convex rank function, we develop the Riemannian optimization algorithms to solve a sequence ofcomplexfixed-rank subproblems through a rank growth strategy. By exploiting the non-compact Stiefel manifold formed by the set of complex full column rank matrices, we develop the Riemannian optimization algorithms to solve the complex fixed-rank optimization problem by applying the semidefinite lifting technique and the Burer–Monteiro factorization approach. The numerical results demonstrate the computational efficiency and higher DoFs achieved by the proposed algorithms.
Kai Yang 0006, Yuanming Shi, Zhi Ding 0001
IEEE Trans. Wirel. Commun.2
2018 Low-Rank Optimization for Data Shuffling in Wireless Distributed Computing
abstract
Wireless distributed computing presents new opportunities to execute intelligent tasks on mobile devices for low-latency applications, by wirelessly aggregating the computation and storage resources among mobile devices. However, for low-latency applications, the key bottleneck lies in the exchange of intermediate results among mobile devices for data shuffling. To improve communication efficiency therein, we establish a novel interference alignment condition by exploiting the locally computed intermediate values as side information. The low-rank optimization model is further developed to maximize the achieved degrees-of-freedom (DoFs). Unfortunately, existing convex relaxation based approach fails to yield satisfied performance due to the poor structure in the formulated low-rank optimization problem, for which we develop a novel difference-of-convex (DC) programming based algorithm. We show that this new approach can significantly improve communication efficiency and the achievable DoF is independent of the number of mobile devices.
Kai Yang 0006, Yuanming Shi, Zhi Ding 0001
ICASSP2
2018 Comparing Massive Networks via Moment Matrices
abstract
In this paper, a novel similarity measure for comparing massive complex networks based on moment matrices is proposed. We consider the corresponding adjacency matrix of a graph as a real random variable of the algebraic probability space with a state. It is shown that the spectral distribution of the matrix can be expressed as a unique discrete probability measure. Then we use the geodesic distance between positive definite moment matrices for comparing massive networks. It is proved that this distance is graph invariant and sub-structure invariant. Numerical simulations demonstrate that the proposed method outperforms state-of-art method in collaboration network classification and its computational cost is extremely cheap.
Hayoung Choi, Yifei Shen 0004, Yuanming Shi
ISIT3
2018 Nonconvex Demixing from Bilinear Measurements
abstract
We consider the problem of demixing a sequence of source signals from the sum of bilinear measurements. It is a generalized mathematical model of blind demixing with deconvolution, which has wide applications in communication, image processing and dictionary learning, etc. However, state-of-art algorithms for blind demixing either fail to scale to large problem sizes or require proper regularization with tedious algorithmic parameters for optimality guarantees. To address the limitations of exiting methods, we propose a provable nonconvex demixing procedure via Wirtinger flow, much like vanilla gradient descent, to harness the benefits of regularization free, fast convergence rate, and optimality guarantees. This is achieved by exploiting the benign geometry of blind demixing, thereby revealing that Wirtinger flow enforces the iterates in the region of strong convexity and qualified level of smoothness.
Jialin Dong, Yuanming Shi
ISIT2
2018 Permuted Linear Model for Header-Free Communication via Symmetric Polynomials
abstract
We present a linear model with an unknown permutation matrix for header-free communication in massive Internet-of-Things (IoT) networks, thereby supporting low-latency communication. Besides the header-free communication application, the permuted linear model also has many applications in matching and correspondence estimation problems. To solve this permuted linear system, the key idea is to convert it into a system of polynomial equations via power sum symmetric polynomials. In our proposed method, specific assumptions (e.g., low-rankness and sparsity) about the permuted linear model is not needed. The closed form of the solution is established for a special case. Computational results show that the proposed method achieves good performance.
Xuming Song, Hayoung Choi, Yuanming Shi
ISIT3
2018 Phase Transitions of Massive Device Connectivity via Convex Geometry
abstract
Massive device connectivity is a crucial communication requirement for Internet of Things (IoT) networks consisting of a large number of devices with sporadic traffic communications. In each coherence time interval, base station (BS) needs to identify the active devices and estimate the channel state information, thereby supporting communication services for the active IoT devices. By exploiting the sparsity pattern in device activity, we develop a group-structured sparsity estimation approach to simultaneously detect the active devices and estimate the wireless channels. This significantly reduces the signature sequence length while supporting massive connectivity with sporadic traffic communications. Specifically, we adopt the convex geometry approach to characterize the phase transition behaviors of the group-structured sparsity estimation problem in complex field. The developed results provide guidelines for choosing appropriate signature sequence length in practice. Numerical results are provided to illustrate the accuracy of our theoretical results.
Tao Jiang 0016, Yuanming Shi
VTC Fall2
2018 Generalized Low-Rank Matrix Completion via Nonconvex Schatten $p$-Norm Minimization
abstract
In this paper, we present a generalized low-rank matrix completion (LRMC) model for topological interference management (TIM), thereby maximizing the achievable degrees of freedom (DoFs) only based on the network connectivity information. Unfortunately, contemporary convex relaxation approaches, e.g, nuclear norm minimization, fail to return low-rank solutions, due to the poor structures in the generalized low-rank model. Most existing nonconvex approaches, however, often need the optimal rank as prior information, which is unavailable in our setting. We thus propose a novel nonconvex relaxation approach with the nonconvex Schatten p-norm to provide a tight approximation for the rank function. A smooth function is formulated to approximate the nonsmooth and nonconvex objective, then an Iteratively Reweighted Least Squares (IRLS- p) method is employed to handle the nonconvexity of the model, which iteratively minimizes the weighted Frobenius norm models of smoothed subproblems while driving the smoothing parameter to 0. We further improve the efficiency by proposing an Iteratively Adaptively Reweighted Least Squares (IARLS- p) algorithm, which uses an adaptively updating strategy for the smoothing parameters in each iteration. Numerical results exhibit the ability of the proposed algorithm to find low-rank solutions, that is, it can achieve higher DoFs in most cases.
Qiong Wu 0005, Fan Zhang 0060, Hao Wang 0045, Yuanming Shi
VTC Fall4
2018 $L_{2}$-Box Optimization for Green Cloud-RAN via Network Adaptation
abstract
In this paper, we propose a novel reformulation of the Mixed Integer Programming (MIP) problem for solving the Cloud Radio Access Network (Cloud-RAN) power consumption minimization problem, and present an l2-box technique to reformulate the MIP problem into an exact and continuous model for recasting the binary constraints into a box with an l2sphere constraint. A Majorization-Minimization (MM) dual ascent algorithm is proposed for solving the reformulated problem, which leads to solving a sequence of Difference of Convex (DC) subproblems handled by an inexact MM algorithm. After obtaining the final solution, we use it as the initial result of the bi-section Group Sparse Beamforming (GSBF) algorithm to promote the group-sparsity of beamformers, rather than using the weighted l1/l2-norm. Simulation results indicate that the new method outperforms the bi-section GSBF algorithm in achieving smaller network power consumption, especially in sparser cases, i.e., Cloud-RANs with a lot of Remote Radio Heads (RRHs) but fewer users.
Fan Zhang 0060, Qiong Wu 0005, Hao Wang 0045, Yuanming Shi
VTC Fall4
2018 Blind demixing for low-latency communication
abstract
In the next generation wireless networks, low-latency communication is critical to support emerging diversified applications, e.g., Tactile Internet and Virtual Reality. In this paper, a novel blind demixing approach is developed to reduce the channel signaling overhead, thereby supporting low-latency communication. Specifically, we develop a low-rank approach to recover the original information only based on a single observed vector without any channel estimation. Unfortunately, this problem turns out to be a highly intractable non-convex optimization problem due to the multiple non-convex rank-one constraints. To address the unique challenges, the quotient manifold geometry of product of complex asymmetric rank-one matrices is exploited by equivalently reformulating original complex asymmetric matrices to the Hermitian positive semidefinite matrices. We further generalize the geometric concepts of the complex product manifolds via element-wise extension of the geometric concepts of the individual manifolds. A scalable Riemannian trust-region algorithm is then developed to solve the blind demixing problem efficiently with fast convergence rates and low iteration cost. Numerical results will demonstrate the algorithmic advantages and admirable performance of the proposed algorithm compared with the state-of-art methods.
Jialin Dong, Kai Yang 0006, Yuanming Shi
WCNC3
2018 Massive CSI Acquisition for Dense Cloud-RANs With Spatial-Temporal Dynamics
abstract
Dense cloud radio access networks (cloud-RANs) provide a promising way to enable scalable connectivity and handle diversified service requirements for massive mobile devices. To fully exploit the performance gains of dense cloud-RANs, channel state information of both the signal link and interference links is required. However, with limited radio resources for training, the channel estimation problem in dense cloud-RANs becomes a high-dimensional estimation problem, i.e., the number of measurements will be typically smaller than the dimension of the channel. In this paper, we shall develop a generic high-dimensional structured channel estimation framework for dense cloud-RANs, which is based on a convex structured regularizing formulation. Observing that the wireless channel possesses ample exploitable statistical characteristics, we propose to convert the available spatial and temporal prior information into appropriate convex regularizers. Simulation results demonstrate that exploiting the spatial and temporal dynamics can achieve good estimation performance even with limited training resources. The alternating direction method of multipliers algorithm is further adopted to solve the resultant large-scale high-dimensional channel estimation problems. The proposed framework thus enjoys modeling flexibility, low training overhead, and computation cost scalability.
Xuan Liu 0005, Yuanming Shi, Jun Zhang 0004, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.2
2018 Enhanced Group Sparse Beamforming for Green Cloud-RAN: A Random Matrix Approach
abstract
Group sparse beamforming is a general framework to minimize the network power consumption for cloud radio access networks, which, however, suffers high computational complexity. In particular, a complex optimization problem needs to be solved to obtain the remote radio head (RRH) ordering criterion in each transmission block, which will help to determine the active RRHs and the associated fronthaul links. In this paper, we propose innovative approaches to reduce the complexity of this key step in group sparse beamforming. Specifically, we first develop a smoothed ℓp-minimization approach with the iterative reweighted-ℓ2algorithm to return a Karush-Kuhn- Tucker (KKT) point solution, as well as enhance the capability of inducing group sparsity in the beamforming vectors. By leveraging the Lagrangian duality theory, we obtain closedform solutions at each iteration to reduce the computational complexity. The well-structured solutions provide opportunities to apply the large-dimensional random matrix theory to derive deterministic approximations for the RRH ordering criterion. Such an approach helps to guide the RRH selection only based on the statistical channel state information, which does not require frequent update, thereby significantly reducing the computation overhead. Simulation results shall demonstrate the performance gains of the proposed ℓp-minimization approach, as well as the effectiveness of the large system analysis-based framework for computing the RRH ordering criterion.
Yuanming Shi, Jun Zhang 0004, Wei Chen 0002, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.1
2017 Massive CSI acquisition in dense cloud-RAN with spatial and temporal prior information
abstract
In this paper, we shall develop a generic channel estimation framework based on the convex formulation for dense cloud radio access networks (Cloud-RAN). Due to the training resource constraint and the large number of transmit antennas, the pilot length is smaller than the antenna number, and thus channel estimation becomes an ill-posed inverse problem. By observing that the wireless channel possesses ample exploitable statistical characteristics, we propose to convert the available spatial and temporal prior information into appropriate convex regularizing functions, yielding convex optimization formulations for the underdetermined channel estimation problem. Simulation results demonstrate that exploiting the prior information of large-scale fading and temporal correlation can achieve good estimation performance even with limited training resources. The alternating direction method of multipliers (ADMM) algorithm is further adopted to solve the resultant large-scale channel estimation problems. The proposed framework is, therefore, scalable to the overhead of prior information and the computation cost for large network sizes.
Xuan Liu 0005, Yuanming Shi, Jun Zhang 0004, Khaled Ben Letaief
ICC2
2017 A Sparse and Low-Rank Optimization Framework for Network Topology Control in Dense Fog-RAN
abstract
In this paper, we propose a sparse and low-rank optimization approach for network topology control in the partially connected fog radio access network (Fog-RAN). In this model, the sparsity of the modeling matrix represents the number of non-connected interference links, while the rank of the matrix models the achievable symmetric degrees-of-freedom (DoF) allocations. This model helps find the network topologies with the maximum number of allowed connected interference links. However, the sparse and low-rank optimization problem turns out to be highly intractable due to the non-convex sparsity objective and non- convex fixed-rank rank constraint. To address the coupled challenges in the objective and constraint, we propose a smoothed Riemannian optimization framework by exploiting the quotient manifold geometry of fixed-rank matrices, followed by a smoothed sparsity inducing surrogate. The proposed Rie- mannian algorithm has much lower computational cost compared with state-of-art matrix factorization parameterized methods. Simulation results further demonstrate the appealing sparsity and low-rankness tradeoff in the proposed model, thereby guiding the network deployment in dense Fog-RAN.
Yuanming Shi, Bamdev Mishra, Xuan Liu 0005, Wei Chen 0002
VTC Spring1
2017 Layered Group Sparse Beamforming for Cache-Enabled Green Wireless Networks
abstract
The exponential growth of mobile data traffic is driving the deployment of dense wireless networks, which will not only impose heavy backhaul burdens, but also generate considerable power consumption. Introducing caches to the wireless network edge is a potential and cost-effective solution to address these challenges. In this paper, we will investigate the problem of minimizing the network power consumption of cache-enabled wireless networks, consisting of the base station (BS) and backhaul power consumption. The objective is to develop efficient algorithms that unify adaptive BS selection, backhaul content assignment, and multicast beamforming, while taking account of user QoS requirements and backhaul capacity limitations. To address the NP-hardness of the network power minimization problem, we first propose a generalized layered group sparse beamforming (LGSBF) modeling framework, which helps to reveal the layered sparsity structure in the beamformers. By adopting the reweighted ℓ1/ℓ2-norm technique, we further develop a convex approximation procedure for the LGSBF problem, followed by a three-stage iterative LGSBF framework to induce the desired sparsity structure in the beamformers. Simulation results validate the effectiveness of the proposed algorithm in reducing the network power consumption, and demonstrate that caching plays a more significant role in networks with higher user densities and less power-efficient backhaul links.
Xi Peng 0006, Yuanming Shi, Jun Zhang 0004, Khaled Ben Letaief
IEEE Trans. Commun.2
2017 Topological Interference Management With User Admission Control via Riemannian Optimization
abstract
Topological interference management (TIM) provides a promising way to manage interference only based on the network connectivity information. Previous works on the TIM problem mainly focus on using the index coding approach and graph theory to establish conditions of network topologies to achieve the feasibility of topological interference management. In this paper, we propose a novel user admission control approach via sparse and low-rank optimization to maximize the number of admitted users for achieving the feasibility of topological interference management. However, the resulting sparse and low-rank optimization problem is non-convex and highly intractable, for which the conventional convex relaxation approaches are inapplicable, e.g., a simple ℓ1-norm relaxation approach yields the objective unbounded and non-convex. To assist efficient algorithms design for the formulated rank-constrained (i.e., degrees-of-freedom (DoFs) allocation) ℓ0-norm maximization (i.e., user capacity maximization) problem, we propose a novel non-convex but smoothed ℓ1-regularized minimization approach to induce sparsity pattern with bounded objective values. We further develop a Riemannian trust-region algorithm to solve the resulting rank-constrained smooth non-convex optimization problem via exploiting the quotient manifold of fixed-rank matrices. Simulation results demonstrate the effectiveness and optimality of the proposed Riemannian algorithm to maximize the number of admitted users for topological interference management.
Yuanming Shi, Bamdev Mishra, Wei Chen 0002
IEEE Trans. Wirel. Commun.1
2016 Low-Rank Matrix Completion for Mobile Edge Caching in Fog-RAN via Riemannian Optimization
abstract
The upcoming big data era requires tremendous computation and storage resources for communications. By pushing computation and storage to network edges, fog radio access networks (Fog- RAN) can effectively increase network throughput and reduce transmission latency. Furthermore, we can exploit the benefits of cache enabled architecture in Fog-RAN to deliver contents with less latency. Radio access units (RAUs) need content delivery from fog servers through wireline links whereas multiple mobile devices need contents from RAUs wirelessly. This work proposes a unified low-rank matrix completion (LRMC) approach to solving the content delivery problem in both wireline and wireless parts of Fog-RAN. To attain a low caching latency, we present a high precision approach with Riemannian trust-region method to solve the challenging LRMC problem by exploiting the quotient manifold geometry of fixed-rank matrices. Numerical results show that the new approach has a faster convergence rate, is able to achieve optimal results, and outperforms other state-of-art algorithms.
Kai Yang 0006, Yuanming Shi, Zhi Ding 0001
GLOBECOM2
2016 Computation offloading in cloud-RAN based mobile cloud computing system
abstract
The cloud radio access network (Cloud-RAN) based mobile cloud computing (MCC) system offers a promising solution to offload computation-intensive tasks from battery and computation capability limited mobile devices to a cloud service provider via wireless transmission, thereby reducing the energy consumption and latency for mobile devices. However, the extra energy and latency caused by wireless transmission for offloading may offset the gains of computation offloading. In this paper, we focus on minimizing the network energy consumption while satisfying the delay requirements by jointly optimizing the communication resources (i.e., uplink and downlink beamforming design), computation resources (i.e., computation capability) and offloading decision (i.e., offload or not offload). Unfortunately, this problem turns out to be an NP-hard non-convex mixed integer non-linear programming (MINLP) problem. To resolve this challenge, the principle of uplink and downlink duality is exploited to aid efficient beamforming design. Furthermore, an effective iterative algorithm with polynomial time complexity is proposed in a holistic way to guide the offloading decision. Simulation results will demonstrate that the proposed iterative algorithm achieves near-optimal energy efficiency in Cloud-RAN based mobile cloud computing system.
Jinkun Cheng, Yuanming Shi, Bo Bai 0001, Wei Chen 0002
ICC2
2016 Statistical group sparse beamforming for green Cloud-RAN via large system analysis
abstract
In this paper, we develop a statistical group sparse beamforming framework to minimize the network power consumption for green cloud radio access networks (Cloud-RANs). It will promote group sparsity structures in the beamforming vectors, which will provide a good indicator for remote radio head (RRH) ordering to enable adaptive RRH selection for power saving. In contrast to the previous works that depend heavily on instantaneous channel state information (CSI), the proposed algorithm only depends on the long-term channel state attenuation for RRH ordering, which does not require frequent update, thereby significantly reducing the computation overhead. This is achieved by developing a smoothed ℓp-minimization approach to induce group sparsity in beamforming vectors, followed by an iterative reweighted-ℓ2algorithm via the principles of the majorization-minimization (MM) algorithm and the Lagrangian duality theory. With the well-structured closed-form solutions at each iteration, we further leverage the large-dimensional random matrix theory to derive deterministic approximations for the squared ℓ2-norm of the induced group sparse beamforming vectors in the large system regimes. The deterministic approximation results only depend on statistical CSI and will guide the RRH ordering. Simulation results demonstrate the near-optimal performance of the proposed algorithm, even in finite systems.
Yuanming Shi, Jun Zhang 0004, Khaled Ben Letaief
ISIT1
2016 Smoothed Lp-Minimization for Green Cloud-RAN With User Admission Control
abstract
The cloud radio access network (Cloud-RAN) has recently been proposed as one of the cost-effective and energy-efficient techniques for 5G wireless networks. By moving the signal processing functionality to a single baseband unit (BBU) pool, centralized signal processing and resource allocation are enabled in cloud-RAN, thereby providing the promise of improving the energy efficiency via effective network adaptation and interference management. In this paper, we propose a holistic sparse optimization framework to design green cloud-RAN by taking into consideration the power consumption of the fronthaul links, multicast services, as well as user admission control. Specifically, we first identify the sparsity structures in the solutions of both the network power minimization and user admission control problems, which call for adaptive remote radio head (RRH) selection and user admission. However, finding the optimal sparsity structures turns out to be NP-hard, with the coupled challenges of the ℓ0-norm-based objective functions and the nonconvex quadratic QoS constraints due to multicast beamforming. In contrast to the previous works on convex but nonsmooth sparsity inducing approaches, e.g., the group sparse beamforming algorithm based on the mixed ℓ1/ℓ2-norm relaxation, we adopt the nonconvex but smoothed ℓp-minimization (02algorithm is developed, which will converge to a Karush-Kuhn-Tucker (KKT) point of the relaxed smoothed ℓp-minimization problem from the SDR technique. We illustrate the effectiveness of the proposed algorithms with extensive simulations for network power minimization and user admission control in multicast cloud-RAN.
Yuanming Shi, Jinkun Cheng, Jun Zhang 0004, Bo Bai 0001, Wei Chen 0002, Khaled Ben Letaief
IEEE J. Sel. Areas Commun.1
2016 Low-Rank Matrix Completion for Topological Interference Management by Riemannian Pursuit
abstract
In this paper, we present a flexible low-rank matrix completion (LRMC) approach for topological interference management (TIM) in the partially connected $K$-user interference channel. No channel state information (CSI) is required at the transmitters except the network topology information. The previous attempt on the TIM problem is mainly based on its equivalence to the index coding problem, but so far only a few index coding problems have been solved. In contrast, in this paper, we present an algorithmic approach to investigate the achievable degrees-of-freedom (DoFs) by recasting the TIM problem as an LRMC problem. Unfortunately, the resulting LRMC problem is known to be NP-hard, and the main contribution of this paper is to propose a Riemannian pursuit (RP) framework to detect the rank of the matrix to be recovered by iteratively increasing the rank. This algorithm solves a sequence of fixed-rank matrix completion problems. To address the convergence issues in the existing fixed-rank optimization methods, the quotient manifold geometry of the search space of fixed-rank matrices is exploited via Riemannian optimization. By further exploiting the structure of the low-rank matrix varieties, i.e., the closure of the set of fixed-rank matrices, we develop an efficient rank increasing strategy to find good initial points in the procedure of rank pursuit. Simulation results demonstrate that the proposed RP algorithm achieves a faster convergence rate and higher achievable DoFs for the TIM problem compared with the state-of-the-art methods.
Yuanming Shi, Jun Zhang 0004, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.1
2015 Group sparse beamforming for multicast green Cloud-RAN via parallel semidefinite programming
abstract
The Cloud radio access network (Cloud-RAN) has great potentials to improve energy efficiency and increase capacity of wireless networks. In this paper, we investigate multicast beamforming design for network power minimization of Cloud-RAN, which is shown to be a highly intractable non-convex mixed integer non-linear programming problem. To provide an efficient solution to this highly complicated problem, we propose a three-stage algorithm based on the group-sparsity inducing norm, which minimizes network power by coordinated multicast beamforming and adaptively selecting active remote radio heads (RRHs). In particular, a novel quadratic variational weighted ℓ1=ℓ2-norm aided alternating algorithm is proposed to exploit the group-sparsity structure of the beamforming vector, thereby guiding the active RRH set selection. Given the selected RRH set, multicast beamforming is performed to minimize the network power consumption. Furthermore, to enhance the computation efficiency upon utilizing the shared computing resources in the cloud center, we employ the alternating direction method of multipliers (ADMM) algorithm to solve the resulting semidefinite programming problems in parallel. Extensive simulation results will demonstrate the effectiveness of the proposed multicast group sparse beamforming algorithm.
Jinkun Cheng, Yuanming Shi, Bo Bai 0001, Wei Chen 0002, Jun Zhang 0004, Khaled Ben Letaief
ICC2
2015 Low-rank matrix completion via Riemannian pursuit for topological interference management
abstract
This paper considers the topological interference management problem in a partially connected K-user interference channel, where no channel state information at transmitters (CSIT) is available beyond the network topology knowledge. Due to the practical CSI assumption, this problem has recently received enough attention. In particular, it has been established that the topological interference management problem, in terms of degrees of freedom (DoF), is equivalent to the index coding problem with linear schemes. However, so far only a few index coding problems have been solved, and thus there is a lack of a systematic way to characterize optimal DoF of an arbitrary network topology. In this paper, we present a low-rank matrix completion (LRMC) approach to find linear solutions to maximize the achievable symmetric DoF for any given network topology. To decode the desired messages at each receiver, we also propose an LRMC based channel acquisition scheme, which can obtain interference-free measurements of the desired channel at each receiver while minimizing the pilot training length. To address the NP-hardness of the non-convex rank objective function in the resulting LRMC problem, we further present a Riemannian pursuit (RP) algorithm to solve it efficiently. This algorithm alternatively performs fixed-rank optimization using Riemannian optimization and rank increase by exploiting the manifold structure of the fixed-rank matrices. The LRMC approach aided by the RP algorithms not only recovers the existing optimal DoF results but also provides insights for general network topologies.
Yuanming Shi, Jun Zhang 0004, Khaled Ben Letaief
ISIT1
2014 Scalable coordinated beamforming for dense wireless cooperative networks
abstract
To meet the ever growing demand for both high throughput and uniform coverage in future wireless networks, dense network deployment will be ubiquitous, for which cooperation among the access points is critical. Considering the computational complexity of designing coordinated beamformers for dense networks, low-complexity and suboptimal precoding strategies are often adopted. However, it is not clear how much performance loss will be caused. To enable optimal coordinated beamforming, in this paper, we propose a framework to design a scalable beamforming algorithm based on the alternative direction method of multipliers (ADMM). Specifically, we first propose to apply the matrix stuffing technique to transform the original optimization problem to an equivalent ADMM-compliant problem, which is much more efficient than the widely-used modeling framework CVX. We will then propose to use the ADMM algorithm, a.k.a. the operator splitting method, to solve the transformed ADMM-compliant problem efficiently. In particular, the subproblems of the ADMM algorithm at each iteration can be solved with closed-forms and in parallel. Simulation results show that the proposed techniques can result in significant computational efficiency compared to the state-of-the-art interior-point solvers. Furthermore, the simulation results demonstrate that the optimal coordinated beamforming can significantly improve the system performance compared to sub-optimal zero forcing beamforming.
Yuanming Shi, Jun Zhang 0004, Khaled Ben Letaief
GLOBECOM1
2014 CSI overhead reduction with stochastic beamforming for cloud radio access networks
abstract
Cloud radio access network (Cloud-RAN) is a promising network architecture to meet the explosive growth of the mobile data traffic. In this architecture, as all the baseband signal processing is shifted to a single baseband unit (BBU) pool, interference management can be efficiently achieved through coordinated beamforming, which, however, often requires full channel state information (CSI). In practice, the overhead incurred to obtain full CSI will dominate the available radio resource. In this paper, we propose a unified framework for the CSI overhead reduction and downlink coordinated beamforming. Motivated by the channel heterogeneity phenomena in large-scale wireless networks, we first propose a novel CSI acquisition scheme, called compressive CSI acquisition, which will obtain instantaneous CSI of only a subset of all the channel links and statistical CSI for the others, thus forming the mixed CSI at the BBU pool. This subset is determined by the statistical CSI. Then we propose a new stochastic beamforming framework to minimize the total transmit power while guaranteeing quality-of-service (QoS) requirements with the mixed CSI. Simulation results show that the proposed CSI acquisition scheme with stochastic beamforming can significantly reduce the CSI overhead while providing performance close to that with full CSI.
Yuanming Shi, Jun Zhang 0004, Khaled Ben Letaief
ICC1
2014 Group Sparse Beamforming for Green Cloud-RAN
abstract
A cloud radio access network (Cloud-RAN) is a network architecture that holds the promise of meeting the explosive growth of mobile data traffic. In this architecture, all the baseband signal processing is shifted to a single baseband unit (BBU) pool, which enables efficient resource allocation and interference management. Meanwhile, conventional powerful base stations can be replaced by low-cost low-power remote radio heads (RRHs), producing a green and low-cost infrastructure. However, as all the RRHs need to be connected to the BBU pool through optical transport links, the transport network power consumption becomes significant. In this paper, we propose a new framework to design a green Cloud-RAN, which is formulated as a joint RRH selection and power minimization beamforming problem. To efficiently solve this problem, we first propose a greedy selection algorithm, which is shown to provide near-optimal performance. To further reduce the complexity, a novel group sparse beamforming method is proposed by inducing the group-sparsity of beamformers using the weighted ℓ1/ℓ2-norm minimization, where the group sparsity pattern indicates those RRHs that can be switched off. Simulation results will show that the proposed algorithms significantly reduce the network power consumption and demonstrate the importance of considering the transport link power consumption.
Yuanming Shi, Jun Zhang 0004, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.1
2012 Coordinated relay beamforming for amplify-and-forward two-hop interference networks
abstract
Relaying is a promising technique to extend coverage and improve throughput in wireless networks, but its performance is degraded in the presence of co-channel interference. In this paper, we consider coordinated relay beamforming to suppress interference and improve the date rates of two-hop interference networks. We first propose optimal coordinated relay beamforming algorithms to characterize the achievable rate region and maximize the sum-rate. By imposing a constraint on the desired signals, a low-complexity iterative algorithm is then proposed to maximize the sum-rate. Through performance comparison, we show that the proposed relaying strategy provides a promising tradeoff between complexity and performance. To further reduce design complexity, we propose a new interference management scheme, interference neutralization, to cancel the interferences over the air at the second hop. We show that this scheme yields a closed-form solution for the beamforming design and provides good performance especially at high signal-to-noise ratio (SNR).
Yuanming Shi, Jun Zhang 0004, Khaled Ben Letaief
GLOBECOM1