Jun Zhang 0004

dblp:z/JunZhang4 · DBLP profile ↗
← Back
224ranked-venue papers
23as first author
111since 2021 · last 2026
0000-0002-5222-1898ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 159 · 9 first-author · 76 since 2021Artificial intelligence and machine learning · 27 · 4 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 2 first-author · 15 since 2021Databases, data management, data science and information retrieval · 14 · 9 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 4 first-author · 4 since 2021Systems, architecture and hardware · 6 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 GaussianImage++: Boosted Image Representation and Compression with 2D Gaussian Splatting
abstract
Implicit neural representations (INRs) have achieved remarkable success in image representation and compression, but they require substantial training time and memory. Meanwhile, recent 2D Gaussian Splatting (GS) methods (\textit{e.g.}, GaussianImage) offer promising alternatives through efficient primitive-based rendering. However, these methods require excessive Gaussian primitives to maintain high visual fidelity. To exploit the potential of GS-based approaches, we present GaussianImage++, which utilizes limited Gaussian primitives to achieve impressive representation and compression performance. Firstly, we introduce a distortion-driven densification mechanism. It progressively allocates Gaussian primitives according to signal intensity. Secondly, we employ context-aware Gaussian filters for each primitive, which assist in the densification to optimize Gaussian primitives based on varying image content. Thirdly, we integrate attribute-separated learnable scalar quantizers and quantization-aware training, enabling efficient compression of primitive attributes. Experimental results demonstrate the effectiveness of our method. In particular, GaussianImage++ outperforms GaussianImage and INRs-based COIN in representation and compression performance while maintaining real-time decoding and low memory usage.
Xingtong Ge, Tongda Xu, Dailan He, Jun Zhang 0004, Yan Wang 0105
AAAI6
2026 VIL2C: Value-of-Information Aware Low-Latency Communication for Multi-Agent Reinforcement Learning
abstract
Inter-agent communication serves as an effective mechanism for enhancing performance in collaborative multi-agent reinforcement learning (MARL) systems. However, the inherent communication latency in practical systems induces both action decision delays and outdated information sharing, impeding MARL performance gains, particularly in time-critical applications like autonomous driving. In this work, we propose a Value-of-Information aware Low-latency Communication (VIL2C) scheme that proactively adjusts the latency distribution to mitigate its effects in MARL systems. Specifically, we define a Value of Information (VoI) metric to quantify the importance of delayed messages on the recipient agent's decision. We then design a VoI aware resource allocation method that dynamically prioritizes message transmission based on each delayed message's importance. Moreover, we propose a progressive message reception mechanism to adaptively adjust the reception duration based on received messages. We derive the optimized VoI aware resource allocation and theoretically prove the performance advantage of the proposed VIL2C scheme. Extensive experiments demonstrate that VIL2C outperforms existing approaches under various communication conditions. These gains are attributed to the low-latency transmission of high-VoI messages via resource allocation and the elimination of unnecessary waiting periods via adaptive reception duration.
Zhuo Sun 0002, Yao Zhang 0005, Zhiwen Yu 0001, Bin Guo 0001, Jun Zhang 0004
AAAI6
2026 Focus-dLLM: Accelerating Long-Context Diffusion LLM Inference via Confidence-Guided Context Focusing
abstract
Lingkun Long, Yushi Huang, Shihao Bai, Ruihao Gong, Jun Zhang, Ao Zhou, Jianlei Yang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Lingkun Long, Yushi Huang, Shihao Bai, Ruihao Gong, Jun Zhang 0004, Jianlei Yang 0001
ACL (1)5
2026 Rethinking Mutual Coupling in Movable Antenna MIMO Systems
abstract
Movable antenna (MA) systems have emerged as a promising technology for future wireless communication systems. The movement of antennas gives rise to mutual coupling (MC) effects, which have been previously ignored and can be exploited to enhance the capacity of multiple-input multiple-output (MIMO) systems. To this end, we first model an MA-enabled point-to-point MIMO communication system with MC effects using a circuit-theoretic framework. The capacity maximization problem is then formulated as a non-concave optimization problem and solved via a block coordinate ascent (BCA)-based algorithm. The subproblem of optimizing MA positions is challenging due to the presence of the analytically intractable MC matrices. To overcome this difficulty, we develop a trust region method (TRM)-based algorithm to optimize MA positions, wherein Sylvester equations are employed to compute the derivatives of the inverse square roots of the MC matrices. Simulation results show significant capacity gains from leveraging MC effects, primarily due to customizable MC matrices and superdirectivity.
Tianyi Liao, Wei Guo 0030, Shenghui Song 0001, Jun Zhang 0004, Khaled Ben Letaief
ICC5
2026 Modular Foundation Model Inference at the Edge: Network-Aware Microservice Optimization
Juan Zhu, Shenghui Song 0001, Jun Zhang 0004, Khaled Ben Letaief
ICC4
2026 Multi-Modal Data Driven Virtual Base Station Construction for Massive MIMO Beam Alignment
abstract
Massive multiple-input multiple-output (MIMO) is a key enabler for the high data rates required by the sixth-generation networks, yet its performance hinges on effective beam management with low training overhead. This paper proposes an interpretable framework to tackle beam alignment in mixed line-of-sight (LoS) and non-line-of-sight (NLoS) propagation environments. Our approach utilizes multimodal data to construct virtual base stations (VBSs), which are geometrically defined as mirror images of the base station across reflecting surfaces reconstructed from 3D LiDAR points. These VBSs provide a sparse and spatial representation of the dominant features of the wireless environment. Based on the constructed VBSs, we develop a VBS-assisted beam alignment scheme comprising coarse channel reconstruction followed by partial beam training. Numerical results demonstrate that the proposed method achieves near-optimal performance in terms of spectral efficiency.
Yijie Bian, Wei Guo 0030, Jie Yang 0035, Shenghui Song 0001, Jun Zhang 0004, Shi Jin 0002, Khaled Ben Letaief
WCNC5
2026 Joint Beamforming and Antenna Position Optimization for Fluid Antenna-Assisted MU-MIMO Networks
Tianyi Liao, Wei Guo 0030, Hengtao He, Shenghui Song 0001, Jun Zhang 0004, Khaled Ben Letaief
IEEE J. Sel. Areas Commun.5
2026 Semantic Communications With World Models
Peiwen Jiang, Jiajia Guo 0001, Chao-Kai Wen, Shi Jin 0002, Jun Zhang 0004
IEEE Trans. Commun.5
2026 ReconX: Reconstruct Any Scene From Sparse Views With Video Diffusion Model
abstract
Advancements in 3D scene reconstruction have transformed 2D images from the real world into 3D models, producing realistic 3D results from hundreds of input photos. Despite great success in dense-view reconstruction scenarios, rendering a detailed scene from sparse views is still an ill-posed optimization problem, often resulting in artifacts and distortions in unseen areas. In this paper, we propose ReconX, a novel 3D scene reconstruction paradigm that reframes the ambiguous reconstruction problem as a temporal generation task. The key insight is to unleash the strong generative prior of large pre-trained video diffusion models for sparse-view reconstruction. Nevertheless, it is challenging to preserve 3D view consistency when directly generating video frames from pre-trained models. To address this issue, given limited input views, the proposed ReconX first constructs a global point cloud and encodes it into a contextual space as the 3D structure condition. Guided by the condition, the video diffusion model then synthesizes video frames that are detail-preserved and exhibit a high degree of 3D consistency, ensuring the coherence of the scene from various perspectives. Finally, we recover the 3D scene from the generated video through a confidence-aware 3D Gaussian Splatting optimization scheme. Extensive experiments on various real-world datasets show the superiority of ReconX over state-of-the-art methods in terms of quality and generalizability.
Fangfu Liu, Wenqiang Sun, Hanyang Wang 0003, Yikai Wang 0001, Haowen Sun 0004, Junliang Ye, Jun Zhang 0004, Yueqi Duan
IEEE Trans. Image Process.7
2026 Task-Oriented Feature Compression for Multimodal Understanding via Device-Edge Co-Inference
abstract
With the rapid development of large multimodal models (LMMs), multimodal understanding applications are emerging. As most LMM inference requests originate from edge devices with limited computational capabilities, the predominant inference pipeline involves directly forwarding the input data to an edge server which handles all computations. However, this approach introduces high transmission latency due to limited uplink bandwidth of edge devices and significant computation latency caused by the prohibitive number of visual tokens, thus hindering delay-sensitive tasks and degrading user experience. To address this challenge, we propose a task-oriented feature compression (TOFC) method for multimodal understanding in a device-edge co-inference framework, where visual features are merged by clustering and encoded by a learnable and selective entropy model before feature projection. Specifically, we employ density peaks clustering based on$K$nearest neighbors to reduce the number of visual features, thereby minimizing both data transmission and computational complexity. Subsequently, a learnable entropy model with hyperprior is utilized to encode and decode merged features, further reducing transmission overhead. To enhance compression efficiency, multiple entropy models are adaptively selected based on the characteristics of the visual features, enabling a more accurate estimation of the probability distribution. Comprehensive experiments on seven visual question answering benchmarks validate the effectiveness of the proposed TOFC method. Results show that TOFC achieves up to 52% reduction in data transmission overhead and 63% reduction in system latency while maintaining identical task performance, compared with neural compression ELIC.
Zhening Liu 0001, Jiashu Lv, Jiawei Shao, Yufei Jiang, Jun Zhang 0004, Xuelong Li 0001
IEEE Trans. Mob. Comput.6
2026 MobileROS: A Wireless-Native Robot Operating System for Mobile Robotics
abstract
The increasing deployment of mobile robots in dynamic outdoor environments necessitates robotic systems capable of maintaining reliability amidst fluctuating wireless connectivity. While the Robot Operating System (ROS) has established itself as the de facto standard for such networked robotics, its abstraction of communication as an opaque, besteffort utility creates a critical bottleneck: it fails to leverage physical layer (PHY) information, resulting in degraded performance and unreliable execution in fluctuating networks. To address this, this paper presents MobileROS, a wireless-native robot operating system that transforms wireless communication from an external service into a core system resource. Grounded in the Symbiotic Paradigm, MobileROS establishes a bidirectional exchange where network conditions inform robotic decisions and mission requirements guide network resource allocation. Based on service mesh principles and domain-driven design, our architecture implements a Hub-Engines-Cells (HEC) model. It features a central Hub for global optimization, three specialized engines (the Radio Information Engine, the Cross Domain Engine, and the Physical Adaptive Engine) for crosslayer intelligence, and distributed Cells as functional units. A key mechanism, Application-Driven Bidirectional Dynamic Slicing, allows robots to actively reconfigure network resources based on semantic urgency, transforming the robot from a passive observer into an active network controller. We systematically evaluate MobileROS across three cities (London, Hong Kong, and Shenzhen) in five scenarios: distributed visual SLAM, cross-domain LiDAR perception, V2X autonomous driving, hybrid multi-robot collaboration against WebRTC baselines, and partition recovery validating CAP-theorem-aware failsafe mechanisms. Results demonstrate that MobileROS maintains significantly more stable performance than standard ROS in mobile wireless deployments.We provide implementation details athttps://github.com/MobileROS.
Boyi Liu 0003, Qianyi Zhang, Yongguang Lu, Jianhao Jiao, Jagmohan Chauhan, Wen Wu 0003, Jun Zhang 0004, Dimitrios Kanoulas
IEEE Trans. Robotics7
2026 Dynamics-Aware Gaussian Splatting Streaming Toward Fast On-the-Fly 4D Reconstruction
abstract
The recent development of 3D Gaussian splatting (3DGS) has led to great interest in 4D dynamic spatial reconstruction. Existing approaches mainly rely on full-length multi-view videos, while there has been limited exploration of online reconstruction methods that enable on-the-fly training and per-timestep streaming. Current 3DGS-based streaming methods treat the Gaussian primitives uniformly and constantly renew the densified Gaussians. Thus, they overlook the difference between dynamic and static features and neglect the temporal continuity of the scene. To address these limitations, we propose a novel pipeline for iterative streamable 4D dynamic spatial reconstruction. It comprises three stages: a selective inheritance stage that retains priors from previous timesteps to preserve the temporal continuity, a dynamics-aware shift stage that distinguishes dynamic and static primitives and employs distinct strategies to optimize their movements, and an error-guided densification stage that efficiently identifies Gaussians requiring densification to accommodate emerging objects. Our method achieves state-of-the-art performance in online 4D reconstruction, demonstrating compact storage, the fastest on-the-fly training speed, and superior representation quality.
Zhening Liu 0001, Yingdong Hu, Jiawei Shao, Zehong Lin, Jun Zhang 0004
IEEE Trans. Vis. Comput. Graph.7
2026 Position-Aided Semantic Communication for Efficient Image Transmission: Design, Implementation, and Experimental Results
Peiwen Jiang, Chao-Kai Wen, Shi Jin 0002, Jun Zhang 0004
IEEE Trans. Wirel. Commun.4
2026 Neural Representation for Wireless Radiation Field Reconstruction: A 3D Gaussian Splatting Approach
abstract
Wireless channel modeling plays a pivotal role in designing, analyzing, and optimizing wireless communication systems. Nevertheless, developing an effective channel modeling approach has been a long-standing challenge. This issue has been escalated due to denser network deployment, larger antenna arrays, and broader bandwidth in next-generation networks. To address this challenge, we put forth WRF-GS, a novel framework for channel modeling based on wireless radiation field (WRF) reconstruction using 3D Gaussian splatting (3D-GS). WRF-GS employs 3D Gaussian primitives and neural networks to capture the interactions between the environment and radio signals, enabling efficient WRF reconstruction and visualization of the propagation characteristics. The reconstructed WRF can then be used to synthesize the spatial spectrum for comprehensive wireless channel characterization. While WRF-GS demonstrates remarkable effectiveness, it faces limitations in capturing high-frequency signal variations caused by complex multipath effects. To overcome these limitations, we propose WRF-GS+, an enhanced framework that integrates electromagnetic wave physics into the neural network design. WRF-GS+ leverages deformable 3D Gaussians to model both static and dynamic components of the WRF, significantly improving its ability to characterize signal variations. In addition, WRF-GS+ accelerates the splatting process by simplifying the 3D-GS modeling operation and reducing sample complexity. Experimental results demonstrate that both WRF-GS and WRF-GS+ outperform baselines for spatial spectrum synthesis, including ray tracing and other deep-learning approaches. Notably, WRF-GS+ achieves state-of-the-art performance in the received signal strength indication (RSSI) and channel state information (CSI) prediction tasks, surpassing existing methods by more than 0.7 dB and 3.36 dB, respectively. The code is available at https://github.com/wenchaozheng/WRF-GSplus.
Chaozheng Wen, Jingwen Tong, Yingdong Hu, Zehong Lin, Jun Zhang 0004
IEEE Trans. Wirel. Commun.5
2026 MUSE-FM: Multi-Task Environment-Aware Foundation Model for Wireless Communications
abstract
Recent advancements in foundation models (FMs) have attracted increasing attention in the wireless communication domain. Leveraging the powerful multi-task learning capability, FMs hold the promise of unifying multiple tasks of wireless communication with a single framework. Nevertheless, existing wireless FMs face limitations in the uniformity to address multiple tasks with diverse inputs/outputs across different communication scenarios. In this paper, we propose a MUlti-taSk Environment-aware FM (MUSE-FM) with a unified architecture to handle multiple tasks in wireless communications, while effectively incorporating scenario information. Specifically, to achieve task uniformity, we propose a unified prompt-guided data encoder-decoder pair to handle data with heterogeneous formats and distributions across different tasks. Besides, we integrate the environmental context as a multi-modal input, which serves as prior knowledge of environment and channel distributions and facilitates cross-scenario feature extraction. Simulation results illustrate that the proposed MUSE-FM outperforms existing methods for various tasks, and its prompt-guided encoder-decoder pair facilitates few-shot adaptation to new task configurations. Moreover, the incorporation of environment information improves the ability to adapt to different scenarios.
Tianyue Zheng, Jiajia Guo 0001, Linglong Dai, Shi Jin 0002, Jun Zhang 0004
IEEE Trans. Wirel. Commun.5
2026 Federated Prompt-Based Decision Transformer for Resource Allocation of Customized VR Streaming in Mobile Edge Computing
abstract
This paper investigates resource allocation for providing heterogeneous users with customized virtual reality (VR) streaming services in a mobile edge computing (MEC) system. We introduce a quality of experience (QoE) metric that considers system latency, user attention levels, and preferred resolutions to measure user experience based on the Weber-Fechner Law. A QoE maximization problem is then formulated for resource allocation to optimize user experience. It is cast as a reinforcement learning problem, aiming to learn a generalized policy applicable across diverse user environments of different MEC servers. To solve the problem, we propose a FedPromptDT framework, which employs federated learning (FL) and prompt-based generative sequence modeling to pre-train a common decision model across MEC servers. FL addresses the issue of insufficient local MEC data while protecting user privacy during offline training. Meanwhile, by integrating user-environment cues and user-preferred allocation, the design of prompts enhances the model’s adaptability to various user environments during online execution. Extensive experimental evaluations demonstrate that FedPromptDT outperforms baseline methods, exhibiting remarkable adaptability and maintaining superior performance across various user environments.
Tailin Zhou, Jiadong Yu, Jun Zhang 0004, Danny H. K. Tsang
IEEE Trans. Wirel. Commun.3
2025 CAMSIC: Content-aware Masked Image Modeling Transformer for Stereo Image Compression
abstract
Existing learning-based stereo image codec adopt sophisticated transformation with simple entropy models derived from single image codecs to encode latent representations. However, those entropy models struggle to effectively capture the spatial-disparity characteristics inherent in stereo images, which leads to suboptimal rate-distortion results. In this paper, we propose a stereo image compression framework, named CAMSIC. CAMSIC independently transforms each image to latent representation and employs a powerful decoder-free Transformer entropy model to capture both spatial and disparity dependencies, by introducing a novel content-aware masked image modeling (MIM) technique. Our content-aware MIM facilitates efficient bidirectional interaction between prior information and estimated tokens, which naturally obviates the need for an extra Transformer decoder. Experiments show that our stereo image codec achieves state-of-the-art rate-distortion performance on two stereo image datasets Cityscapes and InStereo2K with fast encoding and decoding speed.
Shenyuan Gao, Zhening Liu 0001, Jiawei Shao, Xingtong Ge, Dailan He, Tongda Xu, Yan Wang 0105, Jun Zhang 0004
AAAI9
2025 Reinforcement Learning with Intrinsically Motivated Feedback Graph for Lost-sales Inventory Control
abstract
Reinforcement learning (RL) has proven to be well-performed and versatile in inventory control (IC). However, further improvement of RL algorithms in the IC domain is impeded by two limitations of online experience. First, online experience is expensive to acquire in real-world applications. With the low sample efficiency nature of RL algorithms, it would take extensive time to collect enough data and train the RL policy to convergence. Second, online experience may not reflect the true demand due to the lost-sales phenomenon typical in IC, which makes the learning process more challenging. To address the above challenges, we propose a training framework that combines reinforcement learning with feedback graph (RLFG) and intrinsically motivated exploration (IME) to boost sample efficiency. In particular, we first leverage the MDP structure inherent in lost-sales IC problems and design the feedback graph (FG) tailored to lost-sales IC problems to generate abundant side experiences aiding in RL updates. Then we conduct a rigorous theoretical analysis of how the designed FG reduces the sample complexity of RL methods. Guided by these insights, we design an intrinsic reward to direct the RL agent to explore to the state-action space with more side experiences, further exploiting FG’s capability. Experimental results on single-item, multi-item, and multi-echelon environments demonstrate that our method greatly improves the sample efficiency of applying RL in IC. Our code is available at \url{https://github.com/Ziffer-byakuya/RLIMFG4IC}
Zifan Liu, Shibo Chen 0002, Gen Li 0011, Jiashuo Jiang, Jun Zhang 0004
AISTATS6
2025 Fluid Antenna-Assisted MU-MIMO Systems with Decentralized Baseband Processing
Tianyi Liao, Wei Guo 0030, Hengtao He, Shenghui Song 0001, Jun Zhang 0004, Khaled Ben Letaief
GLOBECOM5
2025 Accurate and Fast Channel Estimation for Fluid Antenna Systems with Diffusion Models
abstract
Fluid antenna systems (FAS) offer enhanced spatial diversity for next-generation wireless systems. However, acquiring accurate channel state information (CSI) remains challenging due to the large number of reconfigurable ports and the limited availability of radio-frequency (RF) chains— particularly in high-dimensional FAS scenarios. To address this challenge, we propose an efficient posterior sampling-based channel estimator that leverages a diffusion model (DM) with a simplified U-Net architecture to capture the spatial correlation structure of two-dimensional FAS channels. The DM is initially trained offline in an unsupervised way and then applied online as a learned implicit prior to reconstruct CSI from partial observations via posterior sampling through a denoising diffusion restoration model (DDRM). To accelerate the online inference, we introduce a skipped sampling strategy that updates only a subset of latent variables during the sampling process, thereby reducing the computational cost with minimal accuracy degradation. Simulation results demonstrate that the proposed approach achieves significantly higher estimation accuracy and over 20× speedup compared to state-of-the-art compressed sensing-based methods, highlighting its potential for practical deployment in high-dimensional FAS.
Erqiang Tang, Wei Guo 0030, Hengtao He, Shenghui Song 0001, Jun Zhang 0004, Khaled Ben Letaief
GLOBECOM5
2025 Remote Training in Task-Oriented Communication: Supervised or Self-Supervised with Fine-Tuning?
abstract
Task-oriented communication focuses on extracting and transmitting only the information relevant to specific tasks, effectively minimizing communication overhead. Most existing methods prioritize reducing this overhead during inference, often assuming feasible local training or minimal training communication resources. However, in real-world wireless systems with dynamic connection topologies, training models locally for each new connection is impractical, and task-specific information is often unavailable before establishing connections. Therefore, minimizing training overhead and enabling label-free, task-agnostic pre-training before the connection establishment are essential for effective task-oriented communication. In this paper, we tackle these challenges by employing a mutual information maximization approach grounded in self-supervised learning and information-theoretic analysis. We propose an efficient strategy that pre-trains the transmitter in a task-agnostic and label-free manner, followed by joint fine-tuning of both the transmitter and receiver in a task-specific, label-aware manner. Simulation results show that our proposed method reduces training communication overhead to about half that of full-supervised methods using the SGD optimizer, demonstrating significant improvements in training efficiency.
Hengtao He, Shenghui Song 0001, Jun Zhang 0004, Khaled Ben Letaief
ICC5
2025 Distributed on-Device LLM Inference with Over-the-Air Computation
abstract
Large language models (LLMs) have achieved remarkable success across various artificial intelligence tasks. However, their enormous sizes and computational demands pose significant challenges for the deployment on edge devices. To address this issue, we present a distributed on-device LLM inference framework based on tensor parallelism, which partitions neural network tensors (e.g., weight matrices) of LLMs among multiple edge devices for collaborative inference. Nevertheless, tensor parallelism involves frequent all-reduce operations to aggregate intermediate layer outputs across participating devices during inference, resulting in substantial communication overhead. To mitigate this bottleneck, we propose an over-the-air computation method that leverages the analog superposition property of wireless multipleaccess channels to facilitate fast all-reduce operations. To minimize the average transmission mean-squared error, we investigate joint model assignment and transceiver optimization, which can be formulated as a mixed-timescale stochastic non-convex optimization problem. Then, we develop a mixed-timescale algorithm leveraging semidefinite relaxation and stochastic successive convex approximation methods. Comprehensive simulation results will show that the proposed approach significantly reduces inference latency while improving accuracy. This makes distributed ondevice LLM inference practical for resource-constrained edge devices.
Hengtao He, Shenghui Song 0001, Jun Zhang 0004, Khaled Ben Letaief
ICC4
2025 Dimensionx: Create Any 3D and 4D Scenes From a Single Image With Decoupled Video Diffusion
Wenqiang Sun, Fangfu Liu, Zilong Chen, Yueqi Duan, Jun Zhu 0001, Jun Zhang 0004, Yikai Wang 0001
ICCV7
2025 MEGA: Memory-Efficient 4D Gaussian Splatting for Dynamic Scenes
abstract
4D Gaussian Splatting (4DGS) has recently emerged as a promising technique for capturing complex dynamic 3D scenes with high fidelity. It utilizes a 4D Gaussian representation and a GPU-friendly rasterizer, enabling rapid rendering speeds. Despite its advantages, 4DGS faces significant challenges, notably the requirement of millions of 4D Gaussians, each with extensive associated attributes, leading to substantial memory and storage cost. This paper introduces a memory-efficient framework for 4DGS. We streamline the color attribute by decomposing it into a per-Gaussian direct color component with only 3 parameters and a shared lightweight alternating current color predictor. This approach eliminates the need for spherical harmonics coefficients, which typically involve up to 144 parameters in classic 4DGS, thereby creating a memory-efficient 4D Gaussian representation. Furthermore, we introduce an entropy-constrained Gaussian deformation technique that uses a deformation field to expand the action range of each Gaussian and integrates an opacity-based entropy loss to limit the number of Gaussians, thus forcing our model to use as few Gaussians as possible to fit a dynamic scene well. With simple half-precision storage and zip compression, our framework achieves a storage reduction by approximately 190$\times$ and 125$\times$ on the Technicolor and Neural 3D Video datasets, respectively, compared to the original 4DGS. Meanwhile, it maintains comparable rendering speeds and scene representation quality, setting a new standard in the field. Code is available at https://github.com/Xinjie-Q/MEGA.
Zhening Liu 0001, Yifan Zhang 0004, Xingtong Ge, Dailan He, Tongda Xu, Yan Wang 0105, Zehong Lin, Shuicheng Yan, Jun Zhang 0004
ICCV10
2025 GI-GS: Global Illumination Decomposition on Gaussian Splatting for Inverse Rendering
abstract
We present GI-GS, a novel inverse rendering framework that leverages 3D Gaussian Splatting (3DGS) and deferred shading to achieve photo-realistic novel view synthesis and relighting. In inverse rendering, accurately modeling the shading processes of objects is essential for achieving high-fidelity results. Therefore, it is critical to incorporate global illumination to account for indirect lighting that reaches an object after multiple bounces across the scene. Previous 3DGS-based methods have attempted to model indirect lighting by characterizing indirect illumination as learnable lighting volumes or additional attributes of each Gaussian, while using baked occlusion to represent shadow effects. These methods, however, fail to accurately model the complex physical interactions between light and objects, making it impossible to construct realistic indirect illumination during relighting. To address this limitation, we propose to calculate indirect lighting using efficient path tracing with deferred shading. In our framework, we first render a G-buffer to capture the detailed geometry and material properties of the scene. Then, we perform physically-based rendering (PBR) only for direct lighting. With the G-buffer and previous rendering results, the indirect lighting can be calculated through a lightweight path tracing. Our method effectively models indirect lighting under any given lighting conditions, thereby achieving better novel view synthesis and competitive relighting. Quantitative and qualitative results show that our GI-GS outperforms existing baselines in both rendering quality and efficiency. Project page: https://stopaimme.github.io/GI-GS-site/.
Hongze Chen, Zehong Lin, Jun Zhang 0004
ICLR3
2025 Exponential Topology-enabled Scalable Communication in Multi-agent Reinforcement Learning
abstract
In cooperative multi-agent reinforcement learning (MARL), well-designed communication protocols can effectively facilitate consensus among agents, thereby enhancing task performance. Moreover, in large-scale multi-agent systems commonly found in real-world applications, effective communication plays an even more critical role due to the escalated challenge of partial observability compared to smaller-scale setups. In this work, we endeavor to develop a scalable communication protocol for MARL. Unlike previous methods that focus on selecting optimal pairwise communication links—a task that becomes increasingly complex as the number of agents grows—we adopt a global perspective on communication topology design. Specifically, we propose utilizing the exponential topology to enable rapid information dissemination among agents by leveraging its small-diameter and small-size properties. This approach leads to a scalable communication protocol, named ExpoComm. To fully unlock the potential of exponential graphs as communication topologies, we employ memory-based message processors and auxiliary tasks to ground messages, ensuring that they reflect global information and benefit decision-making. Extensive experiments on large-scale cooperative benchmarks, including MAgent and Infrastructure Management Planning, demonstrate the superior performance and robust zero-shot transferability of ExpoComm compared to existing communication strategies. The code is publicly available at [https://github.com/LXXXXR/ExpoComm](https://github.com/LXXXXR/ExpoComm).
Chenjia Bai, Jun Zhang 0004
ICLR4
2025 HarmoniCa: Harmonizing Training and Inference for Better Feature Caching in Diffusion Transformer Acceleration
abstract
Diffusion Transformers (DiTs) excel in generative tasks but face practical deployment challenges due to high inference costs. Feature caching, which stores and retrieves redundant computations, offers the potential for acceleration. Existing learning-based caching, though adaptive, overlooks the impact of the prior timestep. It also suffers from misaligned objectives-*aligned predicted noise vs. high-quality images*-between training and inference. These two discrepancies compromise both performance and efficiency. To this end, we *harmonize* training and inference with a novel learning-based *caching* framework dubbed **HarmoniCa**. It first incorporates *Step-Wise Denoising Training* (SDT) to ensure the continuity of the denoising process, where prior steps can be leveraged. In addition, an *Image Error Proxy-Guided Objective* (IEPO) is applied to balance image quality against cache utilization through an efficient proxy to approximate the image error. Extensive experiments across $8$ models, $4$ samplers, and resolutions from $256\times256$ to $2K$ demonstrate superior performance and speedup of our framework. For instance, it achieves over $40\\%$ latency reduction (*i.e.*, $2.07\times$ theoretical speedup) and improved performance on PixArt-$\alpha$. Remarkably, our *image-free* approach reduces training time by $25\\%$ compared with the previous method. Our code is available at https://github.com/ModelTC/HarmoniCa.
Yushi Huang, Ruihao Gong, Jing Liu 0048, Jinyang Guo 0002, Xianglong Liu 0001, Jun Zhang 0004
ICML8
2025 C2IQL: Constraint-Conditioned Implicit Q-learning for Safe Offline Reinforcement Learning
abstract
Safe offline reinforcement learning aims to develop policies that maximize cumulative rewards while satisfying safety constraints without the need for risky online interaction. However, existing methods often struggle with the out-of-distribution (OOD) problem, leading to potentially unsafe and suboptimal policies. To address this issue, we first propose Constrained Implicit Q-learning (CIQL), a novel algorithm designed to avoid the OOD problem. In particular, CIQL expands the implicit update of reward value functions to constrained settings and then estimates cost value functions under the same implicit policy. Despite its advantages, the further performance improvement of CIQL is still hindered by the inaccurate discounted approximations of constraints. Thus, we further propose Constraint-Conditioned Implicit Q-learning (C2IQL). Building upon CIQL, C2IQL employs a cost reconstruction model to derive non-discounted cumulative costs from discounted values and incorporates a flexible, constraint-conditioned mechanism to accommodate dynamic safety constraints. Experiment results on DSRL benchmarks demonstrate the superiority of C2IQL compared to baseline methods in achieving higher rewards while guaranteeing safety constraints under different threshold conditions.
Zifan Liu, Jun Zhang 0004
ICML3
2025 WRF-GS: Wireless Radiation Field Reconstruction with 3D Gaussian Splatting
Chaozheng Wen, Jingwen Tong, Yingdong Hu, Zehong Lin, Jun Zhang 0004
INFOCOM5
2025 Exploring Selective Layer Fine-Tuning in Federated Learning
abstract
Federated learning (FL) has emerged as a promising paradigm for fine-tuning foundation models using distributed data in a privacy-preserving manner. Under limited computational resources, clients often find it more practical to fine-tune a selected subset of layers, rather than the entire model, based on their task-specific data. In this study, we provide a thorough theoretical exploration of selective layer fine-tuning in FL, emphasizing a flexible approach that allows the clients to adjust their selected layers according to their local data and resources. We theoretically demonstrate that the layer selection strategy has a significant impact on model convergence in two critical aspects: the importance of selected layers and the heterogeneous choices across clients. Drawing from these insights, we further propose a strategic layer selection method that utilizes local gradients and regulates layer selections across clients. Extensive experiments on both image and text datasets demonstrate the effectiveness of the proposed strategy compared with several baselines, highlighting its advances in identifying critical layers that adapt to the client heterogeneity and training dynamics in FL.
Yuchang Sun 0003, Yuexiang Xie, Bolin Ding, Yaliang Li, Jun Zhang 0004
ISIT5
2025 Graph Neural Network Enhanced Retrieval for Question Answering of Large Language Models
abstract
Zijian Li, Qingyan Guo, Jiawei Shao, Lei Song, Jiang Bian, Jun Zhang, Rui Wang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Zijian Li 0023, Qingyan Guo, Jiawei Shao, Lei Song 0001, Jiang Bian 0002, Jun Zhang 0004, Rui Wang 0028
NAACL (Long Papers)6
2025 Multimodal Deep Learning-Empowered Beam Prediction in Future THz ISAC Systems
abstract
Integrated sensing and communication (ISAC) systems operating at terahertz (THz) bands are envisioned to enable both ultra-high data-rate communication and precise environmental awareness for next-generation wireless networks. However, the narrow width of THz beams makes them prone to misalignment and necessitates frequent beam prediction in dynamic environments. Multimodal sensing, which integrates complementary modalities such as camera images, positional data, and radar measurements, has recently emerged as a promising solution for proactive beam prediction. Nevertheless, existing multimodal approaches typically employ static fusion architectures that cannot adjust to varying modality reliability and contributions, thereby degrading predictive performance and robustness. To address this challenge, we propose a novel and efficient multimodal mixtureof-experts (MoE) deep learning framework for proactive beam prediction in THz ISAC systems. The proposed multimodal MoE framework employs multiple modality-specific expert networks to extract representative features from individual sensing modalities, and dynamically fuses them using adaptive weights generated by a gating network according to the instantaneous reliability of each modality. Simulation results in realistic vehicle-to-infrastructure (V2I) scenarios demonstrate that the proposed MoE framework outperforms traditional static fusion methods and unimodal baselines in terms of prediction accuracy and adaptability, highlighting its potential in practical THz ISAC systems with ultra-massive multiple-input multiple-output (MIMO).
Hengtao He, Shenghui Song 0001, Jun Zhang 0004, Khaled Ben Letaief
PIMRC5
2025 Fractional Delay and Doppler Estimation for OTFS Systems with Doppler Squint Effect
abstract
Orthogonal time frequency space (OTFS) modulation is a promising technology for mitigating severe Doppler effects in high-mobility scenarios. However, existing OTFS channel estimation methods neglect the Doppler Squint Effect (DSE), which incurs serious performance loss. In this paper, we propose a channel estimation algorithm based on Newton's method to accurately estimate fractional delay and Doppler in OTFS systems with DSE. In particular, we obtain the maximum delay and Doppler grid spacing for codebook design to guarantee the convergence of the algorithm. Additionally, we derive the Cramér-Rao lower bound (CRLB) for testing channel parameter estimation performance of our proposed algorithm. Simulation results demonstrate that our proposed algorithm outperforms the orthogonal matching pursuit (OMP) algorithm in terms of normalized mean square error (NMSE), surpassing the Newtonized OMP algorithm with traditional dictionary matrix and approaching the CRLB performance.
Meiying Zhang, Ruoxiao Cao, Hengtao He, Shenghui Song 0001, Jun Zhang 0004, Khaled Ben Letaief
WCNC6
2025 Tackling Distribution Shifts in Task-Oriented Communication With Information Bottleneck
abstract
Task-oriented communication aims to extract and transmit task-relevant information to significantly reduce the communication overhead and transmission latency. However, theunpredictabledistribution shifts between training and test data, includingdomain shiftandsemantic shift, can dramatically undermine the system performance. In order to tackle these challenges, it is crucial to ensure that the encoded features can generalize todomain-shifteddata and detectsemantic-shifteddata, while remaining compact for transmission. In this paper, we propose a novel approach based on the information bottleneck (IB) principle and invariant risk minimization (IRM) framework. The proposed method aims to extract compact and informative features that possess high capability for effectivedomain-shift generalizationand accuratesemantic-shift detectionwithout any knowledge of the test data during training. Specifically, we propose an invariant feature encoding approach based on the IB principle and IRM framework fordomain-shiftgeneralization, which aims to find the causal relationship between the input data and task result by minimizing the complexity and domain dependence of the encoded feature. Furthermore, we enhance the task-oriented communication with the label-dependent feature encoding approach forsemantic-shift detectionwhich achieves joint gains in IB optimization and detection performance. To avoid the intractable computation of the IB-based objective, we leverage variational approximation to derive a tractable upper bound for optimization. Extensive simulation results on image classification tasks demonstrate that the proposed scheme outperforms state-of-the-art approaches and achieves a better rate-distortion tradeoff.
Jiawei Shao, Hengtao He, Shenghui Song 0001, Jun Zhang 0004, Khaled Ben Letaief
IEEE J. Sel. Areas Commun.5
2025 Intelligent Channel Allocation for IEEE 802.11be Multi-Link Operation: When MAB Meets LLM
abstract
WiFi networks have achieved remarkable success in enabling seamless communication and data exchange worldwide. The IEEE 802.11be standard, known as WiFi 7, introduces Multi-Link Operation (MLO), a groundbreaking feature that enables devices to establish multiple simultaneous connections across different bands and channels. While MLO promises substantial improvements in network throughput and latency reduction, it presents significant challenges in channel allocation, particularly in dense network environments. Current research has predominantly focused on performance analysis and throughput optimization within static WiFi 7 network configurations. In contrast, this paper addresses the dynamic channel allocation problem in dense WiFi 7 networks with MLO capabilities. We formulate this challenge as a combinatorial optimization problem, leveraging a novel network performance analysis mechanism. Given the inherent lack of prior network information, we model the problem within a Multi-Armed Bandit (MAB) framework to enable online learning of optimal channel allocations. Our proposed Best-Arm Identification-enabled Monte Carlo Tree Search (BAI-MCTS) algorithm includes rigorous theoretical analysis, providing upper bounds for both sample complexity and error probability. To further reduce sample complexity and enhance generalizability across diverse network scenarios, we put forth LLM-BAI-MCTS, an intelligent algorithm for the dynamic channel allocation problem by integrating the Large Language Model (LLM) into the BAI-MCTS algorithm. Numerical results demonstrate that the BAI-MCTS algorithm achieves a convergence rate approximately 50.44% faster than the state-of-the-art algorithms when reaching 98% of the optimal value. Notably, the convergence rate of the LLM-BAI-MCTS algorithm increases by over 63.32% in dense networks.
Shumin Lian, Jingwen Tong, Jun Zhang 0004, Liqun Fu 0001
IEEE J. Sel. Areas Commun.3
2025 Toward Real-Time Edge AI: Model-Agnostic Task-Oriented Communication With Visual Feature Alignment
abstract
Task-oriented communication presents a promising approach to improve the communication efficiency of edge inference systems by optimizing learning-based modules to extract and transmit relevant task information. However, real-time applications face practical challenges, such as incomplete coverage and potential malfunctions of edge servers. This situation necessitates cross-model communication between different inference systems, enabling edge devices from one service provider to collaborate effectively with edge servers from another. Independent optimization of diverse edge systems often leads to incoherent feature spaces, which hinders the cross-model inference for existing task-oriented communication. To facilitate and achieve effective cross-model task-oriented communication, this study introduces a novel framework that utilizes shared anchor data across diverse systems. This approach addresses the challenge of feature alignment in both server-based and on-device scenarios. In particular, by leveraging the linear invariance of visual features, we propose efficient server-based feature alignment techniques to estimate linear transformations using encoded anchor data features. For on-device alignment, we exploit the angle-preserving nature of visual features and propose to encode relative representations with anchor data to streamline cross-model communication without additional alignment procedures during the inference. The experimental results on computer vision benchmarks demonstrate the superior performance of the proposed feature alignment approaches in cross-model task-oriented communications. The runtime and computation overhead analysis further confirm the effectiveness of the proposed feature alignment approaches in real-time applications.
Songjie Xie, Hengtao He, Shenghui Song 0001, Jun Zhang 0004, Khaled Ben Letaief
IEEE J. Sel. Areas Commun.4
2025 Cell-Free Massive MIMO Detection: A Distributed Expectation Propagation Approach
abstract
Cell-free massive MIMO is one of the core technologies for next-generation wireless networks. It is expected to bring enormous benefits, including ultra-high reliability, data throughput, energy efficiency, and uniform coverage. However, the radically distributed architecture of cell-free massive MIMO necessitates new paradigms for transceiver design, especially by exploiting efficient distributed processing algorithms. In this paper, we propose a distributed expectation propagation (EP) detector for cell-free massive MIMO, which consists of two modules: a nonlinear module at the central processing unit (CPU) and a linear module at each access point (AP). The turbo principle in iterative channel decoding is utilized to compute and pass the extrinsic information between the two modules. An analytical framework is provided to characterize the asymptotic performance of the proposed EP detector with a large number of antennas. Furthermore, a distributed iterative channel estimation and data detection (ICD) algorithm is developed to handle the practical scenario with imperfect channel state information (CSI). Simulation results will show that the proposed method outperforms existing detectors for cell-free massive MIMO systems in terms of the bit-error rate and the developed theoretical analysis can be utilized as an asymptotic lower bound. Finally, it is shown that with imperfect CSI, the proposed ICD algorithm can significantly improve the system performance and reduce the pilot overhead.
Hengtao He, Xianghao Yu, Jun Zhang 0004, Shenghui Song 0001, Ross Murch, Khaled Ben Letaief
IEEE Trans. Mob. Comput.3
2025 How to Collaborate: Towards Maximizing the Generalization Performance in Cross-Silo Federated Learning
abstract
Federated learning (FL) has attracted vivid attention as a privacy-preserving distributed learning framework. In this work, we focus on cross-silo FL, where clients become the model owners after training and are only concerned about the model's generalization performance on their local data. Due to the data heterogeneity issue, asking all the clients to join a single FL training process may result in model performance degradation. To investigate the effectiveness of collaboration, we first derive a generalization bound for each client when collaborating with others or when training independently. We show that the generalization performance of a client can be improved by collaborating with other clients that have more training data and similar data distributions. Our analysis allows us to formulate a client utility maximization problem by partitioning clients into multiple collaborating groups. Ahierarchicalclustering-basedcollaborativetraining (HCCT) scheme is then proposed, which does not need to fix in advance the number of groups. We further analyze the convergence of HCCT for general non-convex loss functions which unveils the effect of data similarity among clients. Extensive simulations show that HCCT achieves better generalization performance than baseline schemes, whereas it degenerates to independent training and conventional FL in specific scenarios.
Yuchang Sun 0001, Marios Kountouris, Jun Zhang 0004
IEEE Trans. Mob. Comput.3
2025 Achieving Linear Speedup in Asynchronous Federated Learning With Heterogeneous Clients
abstract
Federated learning (FL) is an emerging distributed training paradigm that aims to learn a common global model without exchanging or transferring the data that are stored locally at different clients. The Federated Averaging (FedAvg)-based algorithms have gained substantial popularity in FL to reduce the communication overhead, where each client conducts multiple localized iterations before communicating with a central server. In this paper, we focus on FL where the clients have diverse computation and/or communication capabilities. Under this circumstance, FedAvg can be less efficient since it requires all clients that participate in the global aggregation in a round to initiate iterations from thelatestglobal model, and thus the synchronization among fast clients andstraggler clientscan severely slow down the overall training process. To address this issue, we propose an efficient asynchronous federated learning (AFL) framework calledDelayed Federated Averaging (DeFedAvg). In DeFedAvg, the clients are allowed to perform local training with different stale global models at their own paces. Theoretical analyses demonstrate that DeFedAvg achieves asymptotic convergence rates that are on par with the results of FedAvg for solving nonconvex problems. More importantly, DeFedAvg is the first AFL algorithm that provably achieves the desirablelinear speedupproperty, which indicates its high scalability. Additionally, we carry out extensive numerical experiments using real datasets to validate the efficiency and scalability of our approach when training deep neural networks.
Zijian Li 0023, Shi Jin 0002, Jun Zhang 0004
IEEE Trans. Mob. Comput.4
2025 Cost-Efficient FEC Scheme for Time-Sensitive Multi-Hop Transmissions in Overlay Networks
abstract
In pursuit of low latency, real-time communication (RTC) service providers usually use multi-hop overlay links worldwide to bypass congested links, especially for medium- and long-distance transmissions. In such multi-hop long-distance transmission scenarios, utilizing retransmission to recover lost packets can result in increased end-to-end latency. Therefore, Forward Error Correction (FEC) is viewed as a promising way to solve the loss problem. However, for multi-hop overlay transmission, existing FEC schemes either introduce a non-negligible processing delay at each hop or reduce the processing delay at the cost of a high coefficient overhead. In this work, we propose a multi-hop FEC scheme, i.e., FEC-OEM, which considers both processing delay and coefficient overhead. FEC-OEM is designed based on two observations we obtained from measurements. First, coefficient overhead can only be reduced through an implicit transmission way. Therefore, we design a modulation-based recoding module that enables implicit coefficient transmission and hop-by-hop recoding at the same time. Second, using on-the-fly computation is a promising way to reduce processing delay. Accordingly, we design an elimination method to make the modulation-based recoding can be carried out on-the-fly. Real-world experiments demonstrate that FEC-OEM can reduce the processing delay by up to 88% without increasing the coefficient overhead compared to state-of-the-art schemes. We also use FEC-OEM to transmit packets for applications with different loss tolerances, and the results show that FEC-OEM can improve the QoE more effectively than state-of-the-art coding schemes.
Chao Xu 0015, Hui Wang 0011, Jilong Wang 0001, Jun Zhang 0004
IEEE Trans. Mob. Comput.6
2025 Orchestrating Joint Offloading and Scheduling for Low-Latency Edge SLAM
abstract
Visual Simultaneous Localization and Mapping (vSLAM) is a prevailing technology for many emerging robotic applications. Achieving real-time SLAM on mobile robotic systems with limited computational resources is challenging because the complexity of SLAM algorithms increases over time. This restriction can be lifted by offloading computations to edge servers, forming the emerging paradigm ofedge-assisted SLAM. Nevertheless, the exogenous and stochastic input processes affect the dynamics of the edge-assisted SLAM system. Moreover, the requirements of clients on SLAM metrics change over time, exerting implicit and time-varying effects on the system. In this paper, we aim to push the limit beyond existing edge-assist SLAM by proposing a new architecture that can handle the input-driven processes and also satisfy clients’ implicit and time-varying requirements. The key innovations of our work involve a regional feature prediction method for importance-aware local data processing, a configuration adaptation policy that integrates data compression/decompression and task offloading, and an input-dependent learning framework for task scheduling with constraint satisfaction. Extensive experiments prove that our architecture improves pose estimation accuracy and saves up to 47% of communication costs compared with a popular edge-assisted SLAM system, as well as effectively satisfies the clients’ requirements.
Yao Zhang 0005, Yuyi Mao, Hui Wang 0011, Zhiwen Yu 0001, Song Guo 0001, Jun Zhang 0004, Liang Wang 0017, Bin Guo 0001
IEEE Trans. Mob. Comput.6
2025 Low-Complexity CSI Feedback for FDD Massive MIMO Systems via Learning to Optimize
abstract
In frequency-division duplex (FDD) massive multiple-input multiple-output (MIMO) systems, the growing number of base station antennas leads to prohibitive feedback overhead for downlink channel state information (CSI). To address this challenge, state-of-the-art (SOTA) fully data-driven deep learning (DL)-based CSI feedback schemes have been proposed. However, the high computational complexity and memory requirements of these methods hinder their practical deployment on resource-constrained devices like mobile phones. To solve the problem, we propose a model-driven DL-based CSI feedback approach by integrating the wisdom of compressive sensing and learning to optimize (L2O). Specifically, only a linear learnable projection is adopted at the encoder side to compress the CSI matrix, thereby significantly cutting down the user-side complexity and memory expenditure. On the other hand, the decoder incorporates two specially designed components, i.e., a learnable sparse transformation and an element-wise L2O reconstruction module. The former is developed to learn a sparse basis for CSI within the angular domain, which explores channel sparsity effectively. The latter shares the same long short term memory (LSTM) network across all elements of the optimization variable, eliminating the retraining cost when problem scale changes. Simulation results show that the proposed method achieves a comparable performance with the SOTA CSI feedback scheme but with much-reduced complexity, and enables multiple-rate feedback.
Hengtao He, Shenghui Song 0001, Jun Zhang 0004, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.4
2024 Task-Aware Encoder Control for Deep Video Compression
abstract
Prior research on deep video compression (DVC) for machine tasks typically necessitates training a unique codec for each specific task, mandating a dedicated decoder per task. In contrast, traditional video codecs employ a flexible encoder controller, enabling the adaptation of a single codec to different tasks through mechanisms like mode prediction. Drawing inspiration from this, we introduce an innovative encoder controller for deep video compression for machines. This controller features a mode prediction and a Group of Pictures (GoP) selection module. Our approach centralizes control at the encoding stage, allowing for adaptable encoder adjustments across different tasks, such as detection and tracking, while maintaining compatibility with a standard pre-trained DvC decoder. Empirical evidence demonstrates that our method is applica-ble across multiple tasks with various existing pre-trained Dv'Cs. Moreover, extensive experiments demonstrate that our method outperforms previous DVC by about 25% bi-trate for different tasks, with only one pre-trained decoder.
Xingtong Ge, Jixiang Luo, Tongda Xu, Guo Lu, Dailan He, Yan Wang 0105, Jun Zhang 0004, Hongwei Qin
CVPR9
2024 Boosting Neural Representations for Videos with a Conditional Decoder
abstract
Implicit neural representations (INRs) have emerged as a promising approach for video storage and processing, showing remarkable versatility across various video tasks. However, existing methods often fail to fully leverage their representation capabilities, primarily due to inadequate alignment of intermediate features during target frame decoding. This paper introduces a universal boosting framework for current implicit video representation approaches. Specifically, we utilize a conditional decoder with a temporal-aware affine transform module, which uses the frame index as a prior condition to effectively align intermediate features with target frames. Besides, we introduce a sinusoidal NeRV-like block to generate diverse intermediate features and achieve a more balanced parameter distribution, thereby enhancing the model's capacity. With a high-frequency information-preserving reconstruction loss, our approach successfully boosts multiple baseline INRs in the reconstruction quality and convergence speed for video regression, and exhibits superior inpainting and interpolation results. Further, we integrate a consistent entropy minimization technique and develop video codecs based on these boosted INRs. Experiments on the UVG dataset confirm that our enhanced codecs significantly outperform baseline INRs and offer competitive rate-distortion performance compared to traditional and learning-based codecs. Code is available at htt ps://github.com/Xin j ieQ/Boosting-NeRV.
Dailan He, Xingtong Ge, Tongda Xu, Yan Wang 0105, Hongwei Qin, Jun Zhang 0004
CVPR8
2024 Bidirectional Stereo Image Compression with Cross-Dimensional Entropy Model
Zhening Liu 0001, Jiawei Shao, Zehong Lin, Jun Zhang 0004
ECCV (8)5
2024 GaussianImage: 1000 FPS Image Representation and Compression by 2D Gaussian Splatting
Xingtong Ge, Tongda Xu, Dailan He, Yan Wang 0080, Hongwei Qin, Guo Lu, Jun Zhang 0004
ECCV (9)9
2024 Quantization and Privacy Noise Co-Design for Utility-Privacy-Communication Trade-off in Federated Learning
abstract
This study addresses the core challenges in federated learning (FL), namely achieving optimal model utility, safeguarding local data privacy, and maintaining efficient communication. While previous research has focused on either the privacy-utility or communication-utility trade-offs, the investigation of simultaneously considering utility, privacy protection, and communication efficiency has been largely overlooked. In this paper, we propose a novel training framework for FL that combines communication efficiency and differential privacy. Specifically, we employ quantization and binomial noise on model updates to enhance privacy protection and communication efficiency concurrently. Through convergence and privacy analysis, we formulate an optimization problem that maximizes model utility while adhering to privacy and communication constraints. Additionally, we introduce an adaptive algorithm to determine key system parameters, including the level of quantization and privacy noise. Simulation results validate the effectiveness of our proposed FL framework and parameter optimization algorithm.
Lumin Liu, Yuyi Mao, Jun Zhang 0004, Shenghui Song 0001, Khaled Ben Letaief
GLOBECOM3
2024 Decentralizing Coherent Joint Transmission Precoding Via Deterministic Equivalents
abstract
In order to control the inter-cell interference for a multi-cell multi-user multiple-input multiple-output network, we consider the precoder design for coordinated multi-point with downlink coherent joint transmission. To avoid costly information exchange among the cooperating base stations in a centralized precoding scheme, we propose a decentralized one by considering the power minimization problem. By approximating the inter-cell interference using the deterministic equivalents, this problem is decoupled to sub-problems which are solved in a decentralized manner at different base stations. Simulation results demonstrate the effectiveness of our proposed decentralized precoding scheme, where only 2 ∼ 7% more transmit power is needed compared with the optimal centralized precoder.
Yuhao Liu 0005, Xinyu Bian, Yuyi Mao, Jun Zhang 0004
ICASSP7
2024 Learning Bayes-Optimal Channel Estimation for Holographic MIMO in Unknown EM Environments
abstract
Holographic MIMO (HMIMO) has recently been recognized as a promising enabler for future 6G systems through the use of an ultra-massive number of antennas in a compact space to exploit the propagation characteristics of the electromagnetic (EM) channel. Nevertheless, the promised gain of HMIMO could not be fully unleashed without an efficient means to estimate the high-dimensional channel. Bayes-optimal estimators typically necessitate either a large volume of supervised training samples or a priori knowledge of the true channel distribution, which could hardly be available in practice due to the enormous system scale and the complicated EM environments. It is thus important to design a Bayes-optimal estimator for the HMIMO channels in arbitrary and unknown EM environments, free of any supervision or priors. This work proposes a self-supervised minimum mean-square-error (MMSE) channel estimation algorithm based on powerful machine learning tools, i.e., score matching and principal component analysis. The training stage requires only the pilot signals, without knowing the spatial correlation, the ground-truth channels, or the received signal-to-noise-ratio. Simulation results will show that, even being totally self-supervised, the proposed algorithm can still approach the performance of the oracle MMSE method with an extremely low complexity, making it a competitive candidate in practice.
Hengtao He, Xianghao Yu, Shenghui Song 0001, Jun Zhang 0004, Ross Murch, Khaled Ben Letaief
ICC5
2024 Individual Contributions as Intrinsic Exploration Scaffolds for Multi-agent Reinforcement Learning
abstract
In multi-agent reinforcement learning (MARL), effective exploration is critical, especially in sparse reward environments. Although introducing global intrinsic rewards can foster exploration in such settings, it often complicates credit assignment among agents. To address this difficulty, we propose Individual Contributions as intrinsic Exploration Scaffolds (ICES), a novel approach to motivate exploration by assessing each agent’s contribution from a global view. In particular, ICES constructs exploration scaffolds with Bayesian surprise, leveraging global transition information during centralized training. These scaffolds, used only in training, help to guide individual agents towards actions that significantly impact the global latent state transitions. Additionally, ICES separates exploration policies from exploitation policies, enabling the former to utilize privileged global information during training. Extensive experiments on cooperative benchmark tasks with sparse rewards, including Google Research Football (GRF) and StarCraft Multi-agent Challenge (SMAC), demonstrate that ICES exhibits superior exploration capabilities compared with baselines. The code is publicly available at https://github.com/LXXXXR/ICES.
Zifan Liu, Shibo Chen 0002, Jun Zhang 0004
ICML4
2024 Poster Abstract: LLM-Slice: Dedicated Wireless Network Slicing for Large Language Models
abstract
The rapid adoption of large language models (LLMs) presents new challenges for existing network architectures due to significant peak traffic and high communication uncertainty. Traditional wireless networks struggle to support efficiently, leading to intolerable response delays, disconnections, and resource wastage. To address these issues, we propose LLM-Slice, the first system to provide dedicated communication slices for LLMs within a wireless network environment. By creating LLM-specific network slices, LLM-Slice efficiently binds services with communication resources. Based on user equipment (UE) requests and a permissions database, the system registers specific slices to offer controllable LLM services, integrating a downlink resource control module to optimize response speed, enhance resource utilization, and reduce disconnections. By deploying and validating in a real UE-gNB-CN environment, numerical results demonstrate that LLM-Slice significantly improves response speed and resource efficiency, providing a novel solution for fast and controllable LLM access in wireless networks.
Boyi Liu 0003, Jingwen Tong, Jun Zhang 0004
SenSys3
2024 Communication-Learning Co- Design for Over-the-Air Federated Distillation
abstract
The rapid proliferation of artificial intelligence (AI) services gives rise to the development of federated learning (FL), enabling the cooperative learning among wireless devices (WDs) with only local model parameters communicated. Nevertheless, the current emergence of large AI models renders the existing FL approaches inefficient, due to the huge communication overhead. In this paper, we propose a novel over-the-air federated distillation (FD) framework by synergizing the strength of FL and knowledge distillation to avoid the heavy local model transmission. Instead of sharing model parameters, only WDs' model outputs, referred to as knowledge, are shared and aggregated over-the-air by exploiting the superposition property of the multiple-access channel. Accordingly, we study the communication-learning co-design in over-the-air FD, aiming to maximize the learning convergence rate while meeting the power constraints of the transceivers. The main challenge lies in the intractability of the learning performance analysis, as well as the non-convex nature and the optimization spanning the whole FD training period. To tackle this problem, we propose an efficient algorithm to jointly optimize the transmit power of the WDs, estimator for over-the-air aggregation, and receiver beamforming per training round. Numerical results demonstrate that the proposed over-the-air FD achieves significant communication overhead reduction, with only a slight compensation of testing accuracy compared to conventional FL benchmarks.
Zihao Hu, Jia Yan 0003, Ying-Jun Angela Zhang, Jun Zhang 0004, Khaled Ben Letaief
VTC Spring4
2024 Newtonized Near-Field Channel Estimation for Ultra-Massive MIMO Systems
abstract
To meet the stringent requirements of future communication systems, ultra-massive multiple-input and multiple-output (UM-MIMO) technology has garnered significant attention as a key enabling technology for 6G. However, the deployment of UM-MIMO introduces new challenges, particularly the near-field effect. In this paper, by leveraging the unique characteristics of near-field channels, we propose a novel near-field channel estimation algorithm based on the Newton's method. We also design a near-field codebook that meets the requirements for convergence guarantee. Our algorithm overcomes the limitations of existing approaches by offering a low-complexity, tuning-free, and convergence-guaranteed solution. Simulation results show that our proposed algorithm outperforms state-of-the-art baselines in terms of estimation accuracy, establishing its effectiveness in near-field channel estimation for UM-MIMO systems.
Ruoxiao Cao, Hengtao He, Xianghao Yu, Shenghui Song 0001, Jun Zhang 0004, Yi Gong 0001, Khaled Ben Letaief
WCNC6
2024 Communication-Efficient Federated Distillation: Theoretical Analysis and Performance Enhancement
abstract
Federated learning (FL) is a promising paradigm for privacy-preserving deep learning using data distributed on Internet of Things devices. Traditional model sharing-based methods, e.g., federated averaging (FedAvg), suffer from high communication overhead and difficulty in accommodating heterogeneous model architectures. Federated distillation (FD) is a recently proposed alternative to enable communication-efficient and robust FL, as well as heterogeneous client models. However, there is a lack of theoretical understanding of FD-based methods, and their design guidelines remain elusive. This article presents a generic meta-algorithm for FD, generalizing most existing FD training algorithms. By studying a linear classification problem, we show that, with sufficient distillation samples, the training performance of the meta-algorithm is the same as the vanilla FedAvg. To guide the algorithm design and improve communication efficiency, we further investigate the binary classification problem with a Gaussian mixture model, which shows that more distillation data and sampling data with higher confidence improve the training performance. Furthermore, we propose an effective distillation data sampling technique to improve the performance of the FD-meta algorithm, which also reduces communication overhead. Simulations on the benchmark data sets validate the theoretical findings and demonstrate that our proposed algorithm effectively reduces the communication overhead while achieving a satisfactory performance.
Lumin Liu, Jun Zhang 0004, Shenghui Song 0001, Khaled Ben Letaief
IEEE Internet Things J.2
2024 Green Edge AI: A Contemporary Survey
abstract
Artificial intelligence (AI) technologies have emerged as pivotal enablers across a multitude of industries, including consumer electronics, healthcare, and manufacturing, largely due to their significant resurgence over the past decade. The transformative power of AI is primarily derived from the utilization of deep neural networks (DNNs), which require extensive data for training and substantial computational resources for processing. Consequently, DNN models are typically trained and deployed on resource-rich cloud servers. However, due to potential latency issues associated with cloud communications, deep learning (DL) workflows (e.g., DNN training and inference) are increasingly being transitioned to wireless edge networks in proximity to end-user devices (EUDs). This shift is designed to support latency-sensitive applications and has given rise to a new paradigm of edge AI, which will play a critical role in upcoming sixth-generation (6G) networks to support ubiquitous AI applications. Despite its considerable potential, edge AI faces substantial challenges, mostly due to the dichotomy between the resource limitations of wireless edge networks and the resource-intensive nature of DL. Specifically, the acquisition of large-scale data, as well as the training and inference processes of DNNs, can rapidly deplete the battery energy of EUDs. This necessitates an energy-conscious approach to edge AI to ensure both optimal and sustainable performance. In this article, we present a contemporary survey on green edge AI. We commence by analyzing the principal energy consumption components of edge AI systems to identify the fundamental design principles of green edge AI. Guided by these principles, we then explore energy-efficient design methodologies for the three critical tasks in edge AI systems, including training data acquisition, edge training, and edge inference. Finally, we underscore potential future research directions to further enhance the energy efficiency (EE) of edge AI.
Yuyi Mao, Xianghao Yu, Kaibin Huang, Ying-Jun Angela Zhang, Jun Zhang 0004
Proc. IEEE5
2024 Grant-Free Massive Random Access With Retransmission: Receiver Optimization and Performance Analysis
abstract
There is an increasing demand of massive machine-type communication (mMTC) to provide scalable access for a large number of devices, which has prompted extensive investigation on grant-free massive random access (RA) in 5G and beyond wireless networks. Although many efficient signal processing algorithms have been developed, the limited radio resource for pilot transmission in grant-free massive RA systems makes accurate user activity detection and channel estimation challenging, which thereby compromises the communication reliability. In this paper, we adopt retransmission as a means to improve the quality of service (QoS) for grant-free massive RA. Specifically, by jointly leveraging the user activity correlation between adjacent transmission blocks and the historical channel estimation results, we first develop an activity-correlation-aware receiver for grant-free massive RA systems with retransmission based on the correlated approximate message passing (AMP) algorithm. Then, we analyze the performance of the proposed receiver, including the user activity detection, channel estimation, and data error, by resorting to the state evolution of the correlated AMP algorithm and the random matrix theory (RMT). Our analysis admits a tight closed-form approximation for frame error rate (FER) evaluation. Simulation results corroborate our theoretical analysis and demonstrate the effectiveness of the proposed receiver for grant-free massive RA with retransmission, compared with a conventional design that disregards the critical user activity correlation.
Xinyu Bian, Yuyi Mao, Jun Zhang 0004
IEEE Trans. Commun.3
2024 FedCiR: Client-Invariant Representation Learning for Federated Non-IID Features
abstract
Federated learning (FL) is a distributed learning paradigm that maximizes the potential of data-driven models for edge devices without sharing their raw data. However, devices often have non-independent and identically distributed ( non-IID) data, meaning their local data distributions can vary significantly. The heterogeneity in input data distributions across devices, commonly referred to as the feature shift problem, can adversely impact the training convergence and accuracy of the global model. To analyze the intrinsic causes of the feature shift problem, we develop a generalization error bound in FL, which motivates us to propose FedCiR, a client-invariant representation learning framework that enables clients to extract informative and client-invariant features. Specifically, we improve the mutual information term between representations and labels to encourage representations to carry essential classification knowledge, and diminish the mutual information term between the client set and representations conditioned on labels to promote representations of clients to be client-invariant. We further incorporate two regularizers into the FL framework to bound the mutual information terms with an approximate global representation distribution to compensate for the absence of the ground-truth global representation distribution, thus achieving informative and client-invariant feature extraction. To achieve global representation distribution approximation, we propose a data-free mechanism performed by the server without compromising privacy. Extensive experiments demonstrate the effectiveness of our approach in achieving client-invariant representation learning and solving the data heterogeneity issue.
Zijian Li 0023, Zehong Lin, Jiawei Shao, Yuyi Mao, Jun Zhang 0004
IEEE Trans. Mob. Comput.5
2024 Feature Matching Data Synthesis for Non-IID Federated Learning
abstract
Federated learning (FL) has emerged as a privacy-preserving paradigm that trains neural networks on edge devices without collecting data at a central server. However, FL encounters an inherent challenge in dealing with non-independent and identically distributed (non-IID) data among devices. To address this challenge, this paper proposes a hard feature matching data synthesis (HFMDS) method to share auxiliary data besides local models. Specifically, synthetic data are generated by learning the essential class-relevant features of real samples and discarding the redundant features, which helps to effectively tackle the non-IID issue. For better privacy preservation, we propose a hard feature augmentation method to transfer real features towards the decision boundary, with which the synthetic data not only improve the model generalization but also erase the information of real features. By integrating the proposed HFMDS method with FL, we present a novel FL framework with data augmentation to relieve data heterogeneity. The theoretical analysis highlights the effectiveness of our proposed data synthesis method in solving the non-IID challenge. Simulation results further demonstrate that our proposed HFMDS-FL algorithm outperforms the baselines in terms of accuracy, privacy preservation, and complexity saving on various benchmark datasets.
Zijian Li 0023, Yuchang Sun 0001, Jiawei Shao, Yuyi Mao, Hui Wang 0011, Jun Zhang 0004
IEEE Trans. Mob. Comput.6
2024 Live Migration of Video Analytics Applications in Edge Computing
abstract
In order to schedule resources efficiently or maintain applications' continuity for mobile customers, edge platforms often need to adaptively migrate the applications on them. However, our measurement shows that existing migration solutions cannot solve the issue of migrating video analytics applications in edge computing because the memory states of video analytics applications have different characteristics from other applications. We conduct a breakdown analysis of the memory states of video analytics applications, and propose to treat three types of states separately with three different techniques,i.e., warm-up, sync, and replay, to minimize the negative influence of migrations on application performance. Based on this idea, we implement a prototype system in which two new components,i.e.,state storeandsidecar, are designed to achieve near-transparent live migration with minimal application code modifications. Evaluation experiments demonstrate that the time of application interruption caused by migrating a video analytics application with our solution is less than 405ms, and our solution does not consume much resources.
Chenghao Rong, Hui Wang 0011, Jilong Wang 0001, Yipeng Zhou, Jun Zhang 0004
IEEE Trans. Mob. Comput.5
2024 MimiC: Combating Client Dropouts in Federated Learning by Mimicking Central Updates
abstract
Federated learning (FL) is a promising framework for privacy-preserving collaborative learning, where model training tasks are distributed to clients and only the model updates need to be collected at a server. However, when being deployed at mobile edge networks, clients may have unpredictable availability and drop out of the training process, which hinders the convergence of FL. This paper tackles such a critical challenge. Specifically, we first investigate the convergence of the classical FedAvg algorithm with arbitrary client dropouts. We find that with the common choice of a decaying learning rate, FedAvg may oscillate around a stationary point of the global loss function in the worst case, which is caused by the divergence between the aggregated and desired central update. Motivated by this new observation, we then design a novel training algorithm named MimiC, where the server modifies each received model update based on the previous ones. The proposed modification of the received model updates mimics the imaginary central update irrespective of dropout clients. The theoretical analysis of MimiC shows that divergence between the aggregated and central update diminishes with proper learning rates, leading to its convergence. Simulation results further demonstrate that MimiC maintains stable convergence performance and learns better models than the baseline methods
Yuchang Sun 0001, Yuyi Mao, Jun Zhang 0004
IEEE Trans. Mob. Comput.3
2024 From Learning to Analytics: Improving Model Efficacy With Goal-Directed Client Selection
abstract
Federated learning (FL) is an appealing paradigm for learning a global model among distributed clients while preserving data privacy. Driven by the demand for high-quality user experiences, evaluating the well-trained global model after the FL process is crucial. In this paper, we propose a closed-loop model analytics framework that allows for effective evaluation of the trained global model using clients' local data. To address the challenges posed by system and data heterogeneities in the FL process, we study agoal-directedclient selection problem based on the model analytics framework by selecting a subset of clients for the model training. This problem is formulated as a stochastic multi-armed bandit (SMAB) problem. We first put forth a quick initial upper confidence bound (Quick-Init UCB) algorithm to solve this SMAB problem under the federated analytics (FA) framework. Then, we further propose a belief propagation-based UCB (BP-UCB) algorithm under the democratized analytics (DA) framework. Moreover, we derive two regret upper bounds for the proposed algorithms, which increase logarithmically over the time horizon. The numerical results demonstrate that the proposed algorithms achieve nearly optimal performance, with a gap of less than 1.44% and 3.12% under the FA and DA frameworks, respectively.
Jingwen Tong, Liqun Fu 0001, Jun Zhang 0004, Zhu Han 0001
IEEE Trans. Mob. Comput.4
2024 A Federated Online Restless Bandit Framework for Cooperative Resource Allocation
abstract
Restless multi-armed bandits (RMABs) have been widely utilized to address resource allocation problems with Markov reward processes (MRPs). Existing works often assume that the dynamics of MRPs are known prior, which makes the RMAB problem solvable from an optimization perspective. Nevertheless, an efficient learning-based solution for RMABs with unknown system dynamics remains an open problem. In this paper, we fill this gap by investigating a cooperative resource allocation problem with unknown system dynamics of MRPs. This problem can be modeled as a multi-agent online RMAB problem, where multiple agents collaboratively learn the system dynamics while maximizing their accumulated rewards. We devise a federated online RMAB framework to mitigate the communication overhead and data privacy issue by adopting the federated learning paradigm. Based on this framework, we put forth a Federated Thompson Sampling-enabled Whittle Index (FedTSWI) algorithm to solve this multi-agent online RMAB problem. The FedTSWI algorithm enjoys a high communication and computation efficiency, and a privacy guarantee. Moreover, we derive a regret upper bound for the FedTSWI algorithm. Finally, we demonstrate the effectiveness of the proposed algorithm on the case of online multi-user multi-channel access. Numerical results show that the proposed algorithm achieves a fast convergence rate of$\mathcal {O}(\sqrt{T\log (T)})$and better performance compared with baselines. More importantly, its sample complexity reduces sublinearly with the number of agents.
Jingwen Tong, Liqun Fu 0001, Jun Zhang 0004, Khaled Ben Letaief
IEEE Trans. Mob. Comput.4
2024 Learning Decentralized Traffic Signal Controllers With Multi-Agent Graph Reinforcement Learning
abstract
This paper considers optimal traffic signal control in smart cities, which has been taken as a complex networked system control problem. Given the interacting dynamics among traffic lights and road networks, attaining controller adaptivity and scalability stands out as a primary challenge. Capturing the spatial-temporal correlation among traffic lights under the framework of Multi-Agent Reinforcement Learning (MARL) is a promising solution. Nevertheless, existing MARL algorithms ignore effective information aggregation which is fundamental for improving the learning capacity of decentralized agents. In this paper, we design a new decentralized control architecture with improved environmental observability to capture the spatial-temporal correlation. Specifically, we first develop atopology-aware information aggregationstrategy to extract correlation-related information from unstructured data gathered in the road network. Particularly, we transfer the road network topology into a graph shift operator by forming a diffusion process on the topology, which subsequently facilitates the construction of graph signals. A diffusion convolution module is developed, forming a new MARL algorithm, which endows agents with the capabilities of graph learning. Extensive experiments based on both synthetic and real-world datasets verify that our proposal outperforms existing decentralized algorithms.
Yao Zhang 0005, Zhiwen Yu 0001, Jun Zhang 0004, Liang Wang 0017, Tom H. Luan, Bin Guo 0001, Chau Yuen
IEEE Trans. Mob. Comput.3
2024 Understanding and Improving Model Averaging in Federated Learning on Heterogeneous Data
abstract
Model averaging is a widely adopted technique in federated learning (FL) that aggregates multiple client models to obtain a global model. Remarkably, model averaging in FL yields a superior global model, even when client models are trained with non-convex objective functions and on heterogeneous local datasets. However, the rationale behind its success remains poorly understood. To shed light on this issue, we first visualize the loss landscape of FL over client and global models to illustrate their geometric properties. The visualization shows that the client models encompass the global model within a common basin, and interestingly, the global model may deviate from the basin's center while still outperforming the client models. To gain further insights into model averaging in FL, we decompose the expected loss of the global model into five factors related to the client models. Specifically, our analysis reveals that the global model loss after early training mainly arises fromi)the client model's loss on non-overlapping data between client datasets and the global dataset andii)the maximum distance between the global and client models. Based on the findings from our loss landscape visualization and loss decomposition, we propose utilizing iterative moving averaging (IMA) on the global model at the late training phase to reduce its deviation from the expected minimum, while constraining client exploration to limit the maximum distance between the global and client models. Our experiments demonstrate that incorporating IMA into existing FL methods significantly improves their accuracy and training speed on various heterogeneous data setups of benchmark datasets. Code is available athttps://github.com/TailinZhou/FedIMA.
Tailin Zhou, Zehong Lin, Jun Zhang 0004, Danny H. K. Tsang
IEEE Trans. Mob. Comput.3
2024 FedFA: Federated Learning With Feature Anchors to Align Features and Classifiers for Heterogeneous Data
abstract
Federated learning allows multiple clients to collaboratively train a model without exchanging their data, thus preserving data privacy. Unfortunately, it suffers significant performance degradation due to heterogeneous data at clients. Common solutions involve designing an auxiliary loss to regularize weight divergence or feature inconsistency during local training. However, we discover that these approaches fall short of the expected performance because they ignore the existence of avicious cyclebetween feature inconsistency and classifier divergence across clients. Thisvicious cyclecauses client models to be updated in inconsistent feature spaces with more diverged classifiers. To break thevicious cycle, we propose a novel framework namedFederated learning withFeatureAnchors(FedFA). FedFA utilizes feature anchors to align features and calibrate classifiers across clients simultaneously. This enables client models to be updated in a shared feature space with consistent classifiers during local training. Theoretically, we analyze the non-convex convergence rate of FedFA. We also demonstrate that the integration of feature alignment and classifier calibration in FedFA brings avirtuous cyclebetween feature and classifier updates, which breaks thevicious cycleexisting in current approaches. Extensive experiments show that FedFA significantly outperforms existing approaches on various classification datasets under label distribution skew and feature distribution skew.
Tailin Zhou, Jun Zhang 0004, Danny H. K. Tsang
IEEE Trans. Mob. Comput.2
2024 Data-Driven Online Resource Allocation for User Experience Improvement in Mobile Edge Clouds
abstract
As the cloud is pushed to the edge of the network, resource allocation for user experience improvement in mobile edge clouds (MEC) is increasingly important and faces multiple challenges. This paper studies quality of experience (QoE)-oriented resource allocation in MEC while considering user diversity, limited resources, and the complex relationship between allocated resources and user experience. We introduce a closed-loop online resource allocation (CORA) framework to tackle this problem. It learns the objective function of resource allocation from the historical dataset and updates the learned model using the online testing results. Due to the learned objective model is typically non-convex and challenging to solve in real-time, we leverage the Lyapunov optimization to decouple the long-term average constraint and apply the prime-dual method to solve this decoupled resource allocation problem. Thereafter, we put forth a data-driven optimal online queue resource allocation (OOQRA) algorithm and a data-driven robust OQRA (ROQRA) algorithm for homogenous and heterogeneous user cases, respectively. Moreover, we provide a rigorous convergence analysis for the OOQRA algorithm. We conduct extensive experiments to evaluate the proposed algorithms using the synthesis and YouTube datasets. Numerical results validate the theoretical analysis and demonstrate that the user complaint rate is reduced by up to 100% and 18% in the synthesis and YouTube datasets, respectively.
Liqun Fu 0001, Jingwen Tong, Tongtong Lin, Jun Zhang 0004
IEEE Trans. Wirel. Commun.4
2024 Message Passing Meets Graph Neural Networks: A New Paradigm for Massive MIMO Systems
abstract
As one of the core technologies for 5G systems, massive multiple-input multiple-output (MIMO) introduces dramatic capacity improvements along with very high beamforming and spatial multiplexing gains. When developing efficient physical layer algorithms for massive MIMO systems, message passing is one promising candidate owing to its superior performance. However, as their computational complexity increases dramatically with the problem size, the state-of-the-art message passing algorithms cannot be directly applied to future 6G systems, where an exceedingly large number of antennas are expected to be deployed. To address this issue, we propose a model-driven deep learning (DL) framework, namely the AMP-GNN for massive MIMO transceiver design, by considering thelow complexityof the AMP algorithm andadaptabilityof GNNs. Specifically, the structure of the AMP-GNN network is customized by unfolding the approximate message passing (AMP) algorithm and introducing a graph neural network (GNN) module into it. The permutation equivariance property of AMP-GNN is proved, which enables the AMP-GNN to learn more efficiently and to adapt to different numbers of users. We also reveal the underlying reason why GNNs improve the AMP algorithm from the perspective of expectation propagation, which motivates us to amalgamate various GNNs with different message passing algorithms. In the simulation, we take the massive MIMO detection to exemplify that the proposed AMP-GNN significantly improves the performance of the AMP detector, achieves comparable performance as the state-of-the-art DL-based MIMO detectors, and presents strong robustness to various mismatches.
Hengtao He, Xianghao Yu, Jun Zhang 0004, Shenghui Song 0001, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.3
2024 Task-Oriented Communication for Edge Video Analytics
abstract
With the development of artificial intelligence (AI) techniques and the increasing popularity of camera-equipped devices, many edge video analytics applications are emerging, calling for the deployment of computation-intensive AI models at the network edge. Edge inference is a promising solution to move computation-intensive workloads from low-end devices to a powerful edge server for video analytics, but device-server communications will remain a bottleneck due to limited bandwidth. This paper proposes a task-oriented communication framework for edge video analytics, where multiple devices collect the visual sensory data and transmit the informative features to an edge server for processing. To enable low-latency inference, this framework removes video redundancy in spatial and temporal domains and transmits minimal information that is essential for the downstream task, rather than reconstructing the videos on the edge server. Specifically, it extracts compact task-relevant features based on the deterministic information bottleneck (IB) principle, which characterizes a tradeoff between the informativeness of the features and the communication cost. As the features of consecutive frames are temporally correlated, we propose a temporal entropy model (TEM) to reduce the bitrate by taking the previous features as side information in feature encoding. To further improve the inference performance, we build a spatial-temporal fusion module on the server to integrate features of the current and previous frames for joint inference. Extensive experiments on video analytics tasks evidence that the proposed framework effectively encodes task-relevant information of video data and achieves a better rate-performance tradeoff than existing methods.
Jiawei Shao, Jun Zhang 0004
IEEE Trans. Wirel. Commun.3
2024 Channel and Gradient-Importance Aware Device Scheduling for Over-the-Air Federated Learning
abstract
Federated learning (FL) is a popular privacy-preserving distributed training scheme, where multiple devices collaborate to train machine learning models by uploading local model updates. To improve communication efficiency, over-the-air computation (AirComp) has been applied to FL, which leverages analog modulation to harness the superposition property of radio waves such that numerous devices can upload their model updates concurrently for aggregation. However, the uplink channel noise incurs considerable model aggregation distortion, which is critically determined by the device scheduling and compromises the learned model performance. In this paper, we propose a probabilistic device scheduling framework for over-the-air FL, namedPO-FL, to mitigate the negative impact of channel noise, where each device is scheduled according to a certain probability and its model update is reweighted using this probability in aggregation. We prove the unbiasedness of this aggregation scheme and demonstrate the convergence of PO-FL on both convex and non-convex loss functions. Our convergence bounds unveil that the device scheduling affects the learning performance through thecommunication distortionandglobal update variance. Based on the convergence analysis, we further develop a channel and gradient-importance aware algorithm to optimize the device scheduling probabilities in PO-FL. Extensive simulation results show that the proposed PO-FL framework with channel and gradient-importance awareness achieves faster convergence and produces better models than baseline methods.
Yuchang Sun 0001, Zehong Lin, Yuyi Mao, Shi Jin 0002, Jun Zhang 0004
IEEE Trans. Wirel. Commun.5
2024 Stochastic Coded Federated Learning: Theoretical Analysis and Incentive Mechanism Design
abstract
Federated learning (FL) has achieved great success as a privacy-preserving distributed training paradigm, where many edge devices collaboratively train a machine learning model by sharing the model updates instead of the raw data with a server. However, the heterogeneous computational and communication resources of edge devices give rise to stragglers that significantly decelerate the training process. To mitigate this issue, we propose a novel FL framework named stochastic coded federated learning (SCFL) that leverages coded computing techniques. In SCFL, before the training process starts, each edge device uploads a privacy-preserving coded dataset to the server, which is generated by adding Gaussian noise to the projected local dataset. During training, the server computes gradients on the global coded dataset to compensate for the missing model updates of the straggling devices. We design a gradient aggregation scheme to ensure that the aggregated model update is an unbiased estimate of the desired global update. Moreover, this aggregation scheme enables periodical model averaging to improve the training efficiency. We characterize the tradeoff between the convergence performance and privacy guarantee of SCFL. In particular, a more noisy coded dataset provides stronger privacy protection for edge devices but results in learning performance degradation. We further develop a contract-based incentive mechanism to coordinate such a conflict. The simulation results show that SCFL learns a better model within the given time and achieves a better privacy-performance tradeoff than the baseline methods. In addition, the proposed incentive mechanism grants better training performance than the conventional Stackelberg game approach.
Yuchang Sun 0001, Jiawei Shao, Yuyi Mao, Jun Zhang 0004
IEEE Trans. Wirel. Commun.5
2023 Generalized Relation Modeling for Transformer Tracking
abstract
Compared with previous two-stream trackers, the recent one-stream tracking pipeline, which allows earlier interaction between the template and search region, has achieved a remarkable performance gain. However, existing one-stream trackers always let the template interact with all parts inside the search region throughout all the encoder layers. This could potentially lead to target-background confusion when the extracted feature representations are not sufficiently discriminative. To alleviate this issue, we propose a generalized relation modeling method based on adaptive token division. The proposed method is a generalized formulation of attention-based relation modeling for Transformer tracking, which inherits the merits of both previous two-stream and one-stream pipelines whilst enabling more flexible relation modeling by selecting appropriate search tokens to interact with template tokens. An attention masking strategy and the Gumbel-Softmax technique are introduced to facilitate the parallel computation and end-to-end learning of the token division module. Extensive experiments show that our method is superior to the two-stream and one-stream pipelines and achieves state-of-the-art performance on six challenging benchmarks with a real-time running speed. Code and models are publicly available at https://github.com/Little-Podi/GRM.
Shenyuan Gao, Chunluan Zhou, Jun Zhang 0004
CVPR3
2023 Joint Activity-Delay Detection and Channel Estimation for Asynchronous Massive Random Access
abstract
Most existing studies on joint activity detection and channel estimation for grant-free massive random access (RA) systems assume perfect synchronization among all active users, which is hard to achieve in practice. Therefore, this paper considers asynchronous grant-free massive RA systems and develops novel algorithms for joint user activity detection, synchronization delay detection, and channel estimation. In particular, the framework of orthogonal approximate message passing (OAMP) is first utilized to deal with the non-independent and identically distributed (i.i.d.) pilot matrix in asynchronous grant-free massive RA systems, and an OAMP-based algorithm capable of leveraging the common sparsity among the received pilot signals from multiple base station antennas is developed. To reduce the computational complexity, a memory AMP (MAMP)-based algorithm is further proposed that eliminates the matrix inversions in the OAMP-based algorithm. Simulation results demonstrate the effectiveness of the two proposed algorithms over the baseline methods. Besides, the MAMP-based algorithm reduces 37% of the computations while maintaining comparable detection/estimation accuracy, compared with the OAMP-based algorithm.
Xinyu Bian, Yuyi Mao, Jun Zhang 0004
GLOBECOM3
2023 Task-Oriented Communication with Out-of-Distribution Detection: An Information Bottleneck Framework
abstract
Task-oriented communication is an emerging paradigm for next-generation communication networks, which extracts and transmits task-relevant information, instead of raw data, for downstream applications. Most existing deep learning (DL)-based task-oriented communication systems adopt a closed-world assumption, assuming either the same data distribution for training and testing, or the system could have access to a large out-of-distribution (OoD) dataset for retraining. However, in practical open-world scenarios, task-oriented communication systems will be exposed to unknown OoD data. The powerful approximation ability of learning methods may force the task-oriented communication systems to overfit the training data (i.e., in-distribution data). Therefore, these systems tend to provide overconfident judgments when encountering OoD data. Based on the information bottleneck (IB) framework, we propose a class conditional IB (CCIB) approach to address this problem, supported by information-theoretical insights. The idea is to extract distinguishable features from in-distribution data while keeping their compactness and informativeness. It is achieved by imposing the class conditional latent prior distribution and enforcing the latent of different classes to be far away from each other. Simulation results shall demonstrate that the proposed approach detects OoD data more efficiently than the baselines and state-of-the-art approaches, without compromising the rate-distortion tradeoff.
Hengtao He, Jiawei Shao, Shenghui Song 0001, Jun Zhang 0004, Khaled Ben Letaief
GLOBECOM6
2023 Binary Federated Learning with Client-Level Differential Privacy
abstract
Federated learning (FL) is a privacy-preserving collaborative learning framework, and differential privacy can be applied to further enhance its privacy protection. Existing FL systems typically adopt Federated Average (FedAvg) as the training algorithm and implement differential privacy with a Gaussian mechanism. However, the inherent privacy-utility trade-off in these systems severely degrades the training performance if a tight privacy budget is enforced. Besides, the Gaussian mechanism requires model weights to be of high-precision. To improve communication efficiency and achieve a better privacy-utility trade-off, we propose a communication-efficient FL training algorithm with differential privacy guarantee. Specifically, we propose to adopt binary neural networks (BNNs) and introduce discrete noise in the FL setting. Binary model parameters are uploaded for higher communication efficiency and discrete noise is added to achieve the client-level differential privacy protection. The achieved performance guarantee is rigorously proved, and it is shown to depend on the level of discrete noise. Experimental results based on MNIST and Fashion-MNIST datasets will demonstrate that the proposed training algorithm achieves client-level privacy protection with performance gain while enjoying the benefits of low communication overhead from binary model updates.
Lumin Liu, Jun Zhang 0004, Shenghui Song 0001, Khaled Ben Letaief
GLOBECOM2
2023 Blind Performance Prediction for Deep Learning Based Ultra-Massive MIMO Channel Estimation
abstract
Reliability is of paramount importance for the physical layer of wireless systems due to its decisive impact on end-to-end performance. However, the uncertainty of prevailing deep learning (DL)-based physical layer algorithms is hard to quantify due to the black-box nature of neural networks. This limitation is a major obstacle that hinders their practical deployment. In this paper, we attempt to quantify the uncertainty of an important category of DL-based channel estimators. An efficient statistical method is proposed to make blind predictions for the mean squared error of the DL-estimated channel solely based on received pilots, without knowledge of the ground-truth channel, the prior distribution of the channel, or the noise statistics. The complexity of the blind performance prediction is low and scales only linearly with the number of antennas. Simulation results for ultra-massive multiple-input multiple-output (UM-MIMO) channel estimation with a mixture of far-field and near-field paths are provided to verify the accuracy and efficiency of the proposed method.
Hengtao He, Xianghao Yu, Shenghui Song 0001, Jun Zhang 0004, Khaled Ben Letaief
ICC5
2023 Sparse Mixture-of-Experts are Domain Generalizable Learners
Bo Li 0080, Yifei Shen 0004, Yezhen Wang, Jiawei Ren 0001, Tong Che, Jun Zhang 0004, Ziwei Liu 0002
ICLR7
2023 LDMIC: Learning-based Distributed Multi-view Image Coding
Jiawei Shao, Jun Zhang 0004
ICLR3
2023 Low-complexity Deep Video Compression with A Distributed Coding Architecture
abstract
Prevalent predictive coding-based video compression methods rely on a heavy encoder to reduce temporal redundancy, which makes it challenging to deploy them on resource-constrained devices. Since the 1970s, distributed source coding theory has indicated that independent encoding and joint decoding with side information (SI) can achieve high-efficient compression of correlated sources. This has inspired a distributed coding architecture aiming at reducing the encoding complexity. However, traditional distributed coding methods suffer from a substantial performance gap to predictive coding ones. Inspired by the great success of learning-based compression, we propose the first end-to-end distributed deep video compression framework to improve the rate-distortion performance. A key ingredient is an effective SI generation module at the decoder, which helps to effectively exploit inter-frame correlations without computation-intensive encoder-side motion estimation and compensation. Experiments show that our method significantly outperforms conventional distributed video coding and H.264. Meanwhile, it enjoys 6 ∼ 7× encoding speedup against DVC [1] with comparable compression performance. Code is released at https://github.com/Xinjie-Q/Distributed-DVC.
Jiawei Shao, Jun Zhang 0004
ICME3
2023 GNN-Enhanced Approximate Message Passing for Massive/Ultra-Massive MIMO Detection
abstract
Efficient massive/ultra-massive multiple-input multiple-output (MIMO) detection algorithms with satisfactory performance and low complexity are critical to meet the high throughput and ultra-low latency requirements in 5G and beyond communications, given the extremely large number of antennas. In this paper, we propose a low complexity graph neural network (GNN) enhanced approximate message passing (AMP) algorithm, AMP-GNN, for massive/ultra-massive MIMO detection. The structure of the neural network is customized by unfolding the AMP algorithm and introducing the GNN module for multiuser interference cancellation. Numerical results will show that the proposed AMP-GNN significantly improves the performance of the AMP detector and achieves comparable performance as the state-of-the-art deep learning-based MIMO detectors but with reduced computational complexity. Furthermore, it presents strong robustness to the change of the number of users.
Hengtao He, Alva Kosasih, Xianghao Yu, Jun Zhang 0004, Shenghui Song 0001, Wibowo Hardjawana, Khaled Ben Letaief
WCNC4
2023 Joint Activity Detection, Channel Estimation, and Data Decoding for Grant-Free Massive Random Access
abstract
In the massive machine-type communication (mMTC) scenario, a large number of devices with sporadic traffic need to access the network on limited radio resources. Recently, grant-free random access has emerged as a promising mechanism for this challenging scenario, but its potential has not been fully unleashed. In particular, the available auxiliary information has not been fully exploited, including the common sparsity pattern in the received pilot and data signal, as well as the channel decoding information. This article develops advanced receivers in a holistic manner to improve the massive access performance by jointly designing activity detection, channel estimation, and data decoding. To tackle the algorithmic and computational challenges, a turbo structure is adopted at the joint receiver. For performance enhancement, all the received symbols are utilized to jointly estimate the channel state, user activity, and soft data symbols, which effectively exploits the common sparsity pattern. Meanwhile, the extrinsic information from the channel decoder will assist the joint channel estimation and data detection. To reduce the complexity, a low-cost side information (SI)-aided receiver is also proposed, where the channel decoder provides SI to update the estimates on whether a user is active or not. Simulation results show that the turbo receiver is able to reduce the activity detection, channel estimation, and data decoding errors effectively, supporting twice as many active users compared with a separate design that disregards the common sparsity. In addition, the SI-aided receiver notably outperforms the conventional methods with a relatively low complexity.
Xinyu Bian, Yuyi Mao, Jun Zhang 0004
IEEE Internet Things J.3
2023 Monocular 3-D Object Detection Based on Depth-Guided Local Convolution for Smart Payment in D2D Systems
abstract
3-D object detection from mobile phones in Device-to-Device (D2D) system provides a new smart payment tool for the next generation of fintech, which is more flexible and efficient than the traditional barcode. In this article, we propose a monocular 3-D object detection method based on depth-guided local convolution. The method combines the information of RGB image mode and depth mode by using a convolution kernel through depth image and works on a single RGB image locally. According to the multiscale input information, the convolution kernel is adaptively adjusted to capture the target objects of different scales, so as to improve the performance of 3-D object detection. In addition, we use the soft-non-maximum suppression algorithm instead of traditional non-maximum suppression to select the best prediction box. In order to further improve the accuracy of 3-D object detection, the depth estimation network and 3-D object detection network are jointly trained in this method to make the two networks constrain each other and achieve the best performance.
Jun Li 0036, Yongbin Gao, Huixing Wang, Yier Yan, Bo Huang 0014, Jun Zhang 0004, Wei Wang 0030
IEEE Internet Things J.7
2023 Semi-Decentralized Federated Edge Learning With Data and Device Heterogeneity
abstract
Federated edge learning (FEEL) emerges as a privacy-preserving paradigm to effectively train deep learning models from the distributed data in 6G networks. Nevertheless, the limited coverage of a single edge server results in an insufficient number of participating client nodes, which may impair the learning performance. In this paper, we investigate a novel FEEL framework, namelysemi-decentralized federated edge learning(SD-FEEL), where multiple edge servers collectively coordinate a large number of client nodes. By exploiting the low-latency communication among edge servers for efficient model sharing, SD-FEEL incorporates more training data, while enjoying lower latency compared with conventional federated learning. We detail the training algorithm for SD-FEEL with three steps, including local model update, intra-cluster, and inter-cluster model aggregations. The convergence of this algorithm is proved on non-independent and identically distributed data, which reveals the effects of key parameters and provides design guidelines. Meanwhile, the heterogeneity of edge devices may cause the straggler effect and deteriorate the convergence speed of SD-FEEL. To resolve this issue, we propose an asynchronous training algorithm with a staleness-aware aggregation scheme, of which, the convergence is also analyzed. The simulations demonstrate the effectiveness and efficiency of the proposed algorithms for SD-FEEL and corroborate our analysis.
Yuchang Sun 0001, Jiawei Shao, Yuyi Mao, Hui Wang 0011, Jun Zhang 0004
IEEE Trans. Netw. Serv. Manag.5
2023 Hierarchical Federated Learning With Quantization: Convergence Analysis and System Design
abstract
Federated learning (FL) is a powerful distributed machine learning framework where a server aggregates models trained by different clients without accessing their private data. Hierarchical FL, with a client-edge-cloud aggregation hierarchy, can effectively leverage both the cloud server’s access to many clients’ data and the edge servers’ closeness to the clients to achieve a high communication efficiency. Neural network quantization can further reduce the communication overhead during model uploading. To fully exploit the advantages of hierarchical FL, an accurate convergence analysis with respect to the key system parameters is needed. Unfortunately, existing analysis is loose and does not consider model quantization. In this paper, we derive a tighter convergence bound for hierarchical FL with quantization. The convergence result leads to practical guidelines for important design problems such as the client-edge aggregation and edge-client association strategies. Based on the obtained analytical results, we optimize the two aggregation intervals and show that the client-edge aggregation interval should slowly decay while the edge-cloud aggregation interval needs to adapt to the ratio of the client-edge and edge-cloud propagation delay. Simulation results shall verify the design guidelines and demonstrate the effectiveness of the proposed aggregation strategy.
Lumin Liu, Jun Zhang 0004, Shenghui Song 0001, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.2
2023 Task-Oriented Communication for Multidevice Cooperative Edge Inference
abstract
This paper investigates task-oriented communication for multi-device cooperative edge inference, where a group of distributed low-end edge devices transmit the extracted features of local samples to a powerful edge server for inference. While cooperative edge inference can overcome the limited sensing capability of a single device, it substantially increases the communication overhead and may incur excessive latency. To enable low-latency cooperative inference, we propose a learning-based communication scheme that optimizes local feature extraction and distributed feature encoding in a task-oriented manner, i.e., to remove data redundancy and transmit information that is essential for the downstream inference task rather than reconstructing the data samples at the edge server. Specifically, we leverage Tishby’s information bottleneck (IB) principle (Tishby et al., 2000) to extract the task-relevant feature at each edge device, and adopt the distributed information bottleneck (DIB) framework of Aguerri and Zaidi, 2021, to formalize a single-letter characterization of the optimal rate-relevance tradeoff for distributed feature encoding. To admit flexible control of the communication overhead, we extend the DIB framework to a distributed deterministic information bottleneck (DDIB) objective that explicitly incorporates the representational costs of the encoded features. As the IB-based objectives are computationally prohibitive for high-dimensional data, we adopt variational approximations to make the optimization problems tractable. To compensate for the potential performance loss due to the variational approximations, we also develop a selective retransmission (SR) mechanism to identify the redundancy in the encoded features among multiple edge devices to attain additional communication overhead reduction. Extensive experiments on multi-view image classification and multi-view object recognition tasks evidence that the proposed task-oriented communication scheme achieves a better rate-relevance tradeoff than existing methods.
Jiawei Shao, Yuyi Mao, Jun Zhang 0004
IEEE Trans. Wirel. Commun.3
2023 Graph Neural Networks for Wireless Communications: From Theory to Practice
abstract
Deep learning-based approaches have been developed to solve challenging problems in wireless communications, leading to promising results. Early attempts adopted neural network architectures inherited from applications such as computer vision. They often yield poor performance in large scale networks (i.e., poor scalability) and unseen network settings (i.e., poor generalization). To resolve these issues, graph neural networks (GNNs) have been recently adopted, as they can effectively exploit the domain knowledge, i.e., the graph topology in wireless communications problems. GNN-based methods can achieve near-optimal performance in large-scale networks and generalize well under different system settings, but the theoretical underpinnings and design guidelines remain elusive, which may hinder their practical implementations. This paper endeavors to fill both the theoretical and practical gaps. For theoretical guarantees, we prove that GNNs achieve near-optimal performance in wireless networks with much fewer training samples than traditional neural architectures. Specifically, to solve an optimization problem on an$n$-node graph (where the nodes may represent users, base stations, or antennas), GNNs’ generalization error and required number of training samples are$\mathcal {O}(n)$and$\mathcal {O}(n^{2})$times lower than the unstructured multi-layer perceptrons. For design guidelines, we propose a unified framework that is applicable to general design problems in wireless networks, which includes graph modeling, neural architecture design, and theory-guided performance enhancement. Extensive simulations, which cover a variety of important problems and network settings, verify our theory and the effectiveness of the proposed design framework.
Yifei Shen 0004, Jun Zhang 0004, Shenghui Song 0001, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.2
2022 Error Rate Analysis for Grant-free Massive Random Access with Short-Packet Transmission
abstract
Grant-free massive random access (RA) is a promising protocol to support the massive machine-type communications (mMTC) scenario in 5G and beyond networks. In this paper, we focus on the error rate analysis in grant-free massive RA, which is critical for practical deployment but has not been well studied. We consider a two-phase frame structure, with a pilot transmission phase for activity detection and channel estimation, followed by a data transmission phase with coded data symbols. Considering the characteristics of short-packet transmission, we analyze the block error rate (BLER) in the finite blocklength regime to characterize the data transmission performance. The analysis involves characterizing the activity detection and channel estimation errors as well as applying the random matrix theory (RMT) to analyze the distribution of the post-processing signal-to-noise ratio (SNR). As a case study, the derived BLER expression is further simplified to optimize the pilot length. Simulation results verify our analysis and demonstrate its effectiveness in pilot length optimization.
Xinyu Bian, Yuyi Mao, Jun Zhang 0004
GLOBECOM3
2022 Augmented Deep Unfolding for Downlink Beamforming in Multi-cell Massive MIMO With Limited Feedback
abstract
In limited feedback multi-user multiple-input multiple-output (MU-MIMO) cellular networks, users send quantized information about the channel conditions to the associated base station (BS) for downlink beamforming. However, channel quantization and beamforming have been treated as two separate tasks conventionally, which makes it difficult to achieve global system optimality. In this paper, we propose an augmented deep unfolding (ADU) approach that jointly optimizes the beamforming scheme at the BSs and the channel quantization scheme at the users. In particular, the classic WMMSE beamformer is unrolled and a deep neural network (DNN) is leveraged to preprocess its input to enhance the performance. The variational information bottleneck technique is adopted to further improve the performance when the feedback capacity is strictly restricted. Simulation results demonstrate that the proposed ADU method outperforms all the benchmark schemes in terms of the system average rate.
Xianghao Yu, Jun Zhang 0004, Shenghui Song 0001, Khaled Ben Letaief
GLOBECOM3
2022 Hybrid Far- and Near-Field Channel Estimation for THz Ultra-Massive MIMO via Fixed Point Networks
abstract
Terahertz ultra-massive multiple-input multiple-output (THz UM-MIMO) is envisioned as one of the key enablers of 6G wireless systems. Due to the joint effect of its large array aperture and small wavelength, the near-field region of THz UM-MIMO is greatly enlarged. The high-dimensional channel of such systems thus consists of a stochastic mixture of far and near fields, which renders channel estimation extremely challenging. Previous works based on uni-field assumptions cannot capture the hybrid far- and near-field features, thus suffering significant performance loss. This motivates us to consider hybrid-field channel estimation. We draw inspirations from fixed point theory to develop an efficient deep learning based channel estimator with adaptive complexity and linear convergence guarantee. Built upon classic orthogonal approximate message passing, we transform each iteration into a contractive mapping, comprising a closed-form linear estimator and a neural network based non-linear estimator. A major algorithmic innovation involves applying fixed point iteration to compute the channel estimate while modeling neural networks with arbitrary depth and adapting to the hybrid-field channel conditions. Simulation results verify our theoretical analysis and show significant performance gains over state-of-the-art approaches in the estimation accuracy and convergence rate.
Yifei Shen 0004, Hengtao He, Xianghao Yu, Jun Zhang 0004, Khaled Ben Letaief
GLOBECOM5
2022 Asynchronous Semi-Decentralized Federated Edge Learning for Heterogeneous Clients
abstract
Federated edge learning (FEEL) has drawn much attention as a privacy-preserving distributed learning framework for mobile edge networks. In this work, we investigate a novel semi-decentralized FEEL (SD-FEEL) architecture where multiple edge servers collaborate to incorporate more data from edge devices in training. Despite the low training latency enabled by fast edge aggregation, the device heterogeneity in computational resources deteriorates the efficiency. This paper proposes an asynchronous training algorithm to overcome this issue in SD-FEEL, where edge servers are allowed to independently set deadlines for the associated client nodes and trigger the model aggregation. To deal with different levels of model staleness, we design a staleness-aware aggregation scheme and analyze its convergence. Simulation results demonstrate the effectiveness of our proposed algorithm in achieving faster convergence and better learning performance than synchronous training.
Yuchang Sun 0001, Jiawei Shao, Yuyi Mao, Jun Zhang 0004
ICC4
2022 Communication-Efficient Federated Distillation with Active Data Sampling
abstract
Federated learning (FL) is a promising paradigm to enable privacy-preserving deep learning from distributed data. Most previous works are based on federated average (FedAvg), which, however, faces several critical issues, including a high communication overhead and the difficulty in dealing with heterogeneous model architectures. Federated Distillation (FD) is a recently proposed alternative to enable communication-efficient and robust FL, which achieves orders of magnitude reduction of the communication overhead compared with FedAvg and is flexible to handle heterogeneous models at the clients. However, so far there is no unified algorithmic framework or theoretical analysis for FD-based methods. In this paper, we first present a generic meta-algorithm for FD and investigate the influence of key parameters through empirical experiments. Then, we verify the empirical observations theoretically. Based on the empirical results and theory, we propose a communication-efficient FD algorithm with active data sampling to improve the model performance and reduce the communication overhead. Empirical simulations on benchmark datasets will demonstrate that our proposed algorithm effectively and significantly reduces the communication overhead while achieving a satisfactory performance.
Lumin Liu, Jun Zhang 0004, Shenghui Song 0001, Khaled Ben Letaief
ICC2
2022 How Neural Architectures Affect Deep Learning for Communication Networks?
abstract
In recent years, there has been a surge in applying deep learning to various challenging design problems in communication networks. The early attempts adopt neural architectures inherited from applications such as computer vision, which suffer from poor generalization, scalability, and lack of interpretability. To tackle these issues, domain knowledge has been integrated into the neural architecture design, which achieves near-optimal performance in large-scale networks and generalizes well under different system settings. This paper endeavors to theoretically validate the importance and effects of neural architectures when applying deep learning to communication network design. We prove that by exploiting permutation invariance, a common property in communication networks, graph neural networks (GNNs) converge faster and generalize better than fully connected multi-layer perceptrons (MLPs), especially when the number of nodes (e.g., users, base stations, or antennas) is large. Specifically, we prove that under common assumptions, for a communication network with n nodes, GNNs converge O(n log n) times faster and their generalization error is O(n) times lower, compared with MLPs.
Yifei Shen 0004, Jun Zhang 0004, Khaled Ben Letaief
ICC2
2022 Loading Cost-Aware Model Caching and Request Routing for Cooperative Edge Inference
abstract
Most existing works on edge service caching and request routing fail to consider the influence of the service loading time. Meanwhile, the requests generated by end devices will change dynamically, which means that the caching strategy should adapt accordingly. In this paper, we investigate loading cost-aware joint model caching and request routing with cooperative edge computing, considering both the service loading time and the dynamic user requests. A system throughput maximization problem is formulated, which is proved to be NP-hard. Then, a randomized rounding-based online algorithm with M/(M − 2 ln N)-approximation ratio is proposed to solve it, where M and N are the numbers of end devices and deep neural network (DNN) models, respectively. Extensive experimental results demonstrate that our algorithm achieves more than 42.7% throughput gain than baseline algorithms.
Mianyang Yao, Long Chen 0006, Jun Zhang 0004, Jigang Wu
ICC3
2022 Stochastic Coded Federated Learning with Convergence and Privacy Guarantees
abstract
Federated learning (FL) has attracted much attention as a privacy-preserving distributed machine learning framework, where many clients collaboratively train a machine learning model by exchanging model updates with a parameter server instead of sharing their raw data. Nevertheless, FL training suffers from slow convergence and unstable performance due to stragglers caused by the heterogeneous computational resources of clients and fluctuating communication rates. This paper proposes a coded FL framework to mitigate the straggler issue, namely stochastic coded federated learning (SCFL). In this framework, each client generates a privacy-preserving coded dataset by adding additive noise to the random linear combination of its local data. The server collects the coded datasets from all the clients to construct a composite dataset, which helps to compensate for the straggling effect. In the training process, the server as well as clients perform mini-batch stochastic gradient descent (SGD), and the server adds a make-up term in model aggregation to obtain unbiased gradient estimates. We characterize the privacy guarantee by the mutual information differential privacy (MI-DP) and analyze the convergence performance in federated learning. Besides, we demonstrate a privacy-performance tradeoff of the proposed SCFL method by analyzing the influence of the privacy constraint on the convergence rate. Finally, numerical experiments corroborate our analysis and show the benefits of SCFL in achieving fast convergence while preserving data privacy.
Yuchang Sun 0001, Jiawei Shao, Yuyi Mao, Jun Zhang 0004
ISIT5
2022 DReS-FL: Dropout-Resilient Secure Federated Learning for Non-IID Clients via Secret Data Sharing
abstract
Federated learning (FL) strives to enable collaborative training of machine learning models without centrally collecting clients' private data. Different from centralized training, the local datasets across clients in FL are non-independent and identically distributed (non-IID). In addition, the data-owning clients may drop out of the training process arbitrarily. These characteristics will significantly degrade the training performance. This paper proposes a Dropout-Resilient Secure Federated Learning (DReS-FL) framework based on Lagrange coded computing (LCC) to tackle both the non-IID and dropout problems. The key idea is to utilize Lagrange coding to secretly share the private datasets among clients so that each client receives an encoded version of the global dataset, and the local gradient computation over this dataset is unbiased. To correctly decode the gradient at the server, the gradient function has to be a polynomial in a finite field, and thus we construct polynomial integer neural networks (PINNs) to enable our framework. Theoretical analysis shows that DReS-FL is resilient to client dropouts and provides privacy protection for the local datasets. Furthermore, we experimentally demonstrate that DReS-FL consistently leads to significant performance gains over baseline methods.
Jiawei Shao, Yuchang Sun 0001, Jun Zhang 0004
NeurIPS4
2022 Semi-Decentralized Federated Edge Learning for Fast Convergence on Non-IID Data
abstract
Federated edge learning (FEEL) has emerged as an effective approach to reduce the large communication latency in Cloud-based machine learning solutions, while preserving data privacy. Unfortunately, the learning performance of FEEL may be compromised due to limited training data in a single edge cluster. In this paper, we investigate a novel framework of FEEL, namely semi-decentralized federated edge learning (SD-FEEL). By allowing model aggregation across different edge clusters, SD-FEEL enjoys the benefit of FEEL in reducing the training latency, while improving the learning performance by accessing richer training data from multiple edge clusters. A training algorithm for SD-FEEL with three main procedures in each round is presented, including local model updates, intra-cluster and inter-cluster model aggregations, which is proved to converge on non-independent and identically distributed (non-IID) data. We also characterize the interplay between the network topology of the edge servers and the communication overhead of inter-cluster model aggregation on the training performance. Experiment results corroborate our analysis and demonstrate the effectiveness of SD-FFEL in achieving faster convergence than traditional federated learning architectures. Besides, guidelines on choosing critical hyper-parameters of the training algorithm are also provided.
Yuchang Sun 0001, Jiawei Shao, Yuyi Mao, Hui Wang 0011, Jun Zhang 0004
WCNC5
2022 Faster Activity and Data Detection in Massive Random Access: A Multiarmed Bandit Approach
abstract
This article investigates the grant-free random access mechanism for massive Internet of Things (IoT) devices. By embedding the data symbols in the signature sequences, joint device activity detection and data decoding can be achieved, which, however, significantly increases the computational complexity. Coordinate descent algorithms that enjoy a low per-iteration complexity have been employed to solve this detection problem, but previous works typically employ a random coordinate selection policy which leads to slow convergence. In this article, we develop multiarmed bandit (MAB) approaches for more efficient detection via coordinate descent, which achieves a delicate tradeoff betweenexplorationandexploitationin coordinate selection. Specifically, we first propose a bandit-based strategy, i.e., Bernoulli sampling, to speed up the convergence rate of coordinate descent, by learning which coordinates will result in more aggressive descent of thenonconvex objective function. To further improve the convergence rate, an inner MAB problem is established to learn the exploration policy of Bernoulli sampling. Both convergence rate analysis and simulation results are provided to show that the proposed bandit-based algorithms enjoy faster convergence rates with a lower time complexity compared with the state-of-the-art algorithm. Furthermore, our proposed algorithms are generally applicable to different scenarios, e.g., massive random access with low-precision analog-to-digital converters (ADCs).
Jialin Dong, Jun Zhang 0004, Yuanming Shi, Hui Wang 0011
IEEE Internet Things J.2
2022 Learning Task-Oriented Communication for Edge Inference: An Information Bottleneck Approach
abstract
This paper investigates task-oriented communication for edge inference, where a low-end edge device transmits the extracted feature vector of a local data sample to a powerful edge server for processing. It is critical to encode the data into aninformativeandcompactrepresentation for low-latency inference given the limited bandwidth. We propose a learning-based communication scheme that jointly optimizes feature extraction, source coding, and channel coding in a task-oriented manner, i.e., targeting the downstream inference task rather than data reconstruction. Specifically, we leverage an information bottleneck (IB) framework to formalize a rate-distortion tradeoff between the informativeness of the encoded feature and the inference performance. As the IB optimization is computationally prohibitive for the high-dimensional data, we adopt a variational approximation, namely the variational information bottleneck (VIB), to build a tractable upper bound. To reduce the communication overhead, we leverage a sparsity-inducing distribution as the variational prior for the VIB framework to sparsify the encoded feature vector. Furthermore, considering dynamic channel conditions in practical communication systems, we propose a variable-length feature encoding scheme based on dynamic neural networks to adaptively adjust the activated dimensions of the encoded feature to different channel conditions. Extensive experiments evidence that the proposed task-oriented communication system achieves a better rate-distortion tradeoff than baseline methods and significantly reduces the feature transmission latency in dynamic channel conditions.
Jiawei Shao, Yuyi Mao, Jun Zhang 0004
IEEE J. Sel. Areas Commun.3
2022 Dependency-Aware Computation Offloading for Mobile Edge Computing With Edge-Cloud Cooperation
abstract
Most of existing Multi-access edge computing (MEC) studies consider the remote cloud server as a special edge server, the opportunity of edge-cloud collaboration has not been well exploited. We propose a dependency-aware offloading scheme in MEC with edge-cloud cooperation under task dependency constraints. Each mobile device has a limited budget and has to determine which sub-task should be computed locally or should be sent to the edge or remote cloud. To address this issue, we divide the offloading problem into two application finishing time minimization sub-problems with two different cooperation modes, both of which are proved to be NP-hard. We then devise one greedy algorithm with approximation ratio of$1+\epsilon$for the first mode with edge-cloud cooperation but no edge-edge cooperation. Then we design an efficient greedy algorithm for the second mode, considering both edge-cloud and edge-edge co-operations. Extensive simulation results show that for the first mode, the proposed greedy algorithm achieves near optimal performance for typical task topologies. On average, it outperforms the modified Hermes benchmark algorithm by about$23\%\sim 43.6\%$in terms of application finishing time with given budgets. By further exploiting collaborations among edge servers in the second cooperation mode, the proposed algorithm helps to achieve over 20.3 percent average performance gain on the application finishing time over the first mode under various scenarios. Real-world experiments comply with simulation results.
Long Chen 0006, Jigang Wu, Jun Zhang 0004, Hongning Dai, Mianyang Yao
IEEE Trans. Cloud Comput.3
2022 Towards Dependency-Aware Cache Management for Data Analytics Applications
abstract
Memory caches are being used aggressively in today's data analytics systems such as Spark, Tez, and Piccolo. The significant performance impact of caches and their limited sizes call for efficient cache management in data analytics clusters. However, prevalent data analytics systems employ rather simple cache management policies—notably Least Recently Used (LRU) and Least Frequently Used (LFU)—that areobliviousto the application semantics of data dependency, expressed as directed acyclic graphs (DAGs). Without this knowledge, cache management can, at best, be performed by “guessing” the future data access patterns based on history, which frequently results in inefficient, erroneous caching with a low hit rate and a long response time. Worse still, the lack of data dependency knowledge makes it impossible to retain theall-or-nothingcache property of cluster applications, in that a compute task cannot be sped up unless all the dependent data has been kept in the main memory. In this paper, we propose a novel cache replacement policy, named Least Reference Count (LRC), which exploits the application's data dependency information to optimize the cache management. LRC keeps track of thereference countof each data block, defined as the number of dependent child blocks that have not been computed yet, and always evicts the block with the smallest reference count. Furthermore, we incorporate the all-or-nothing requirement into LRC by coordinately managing the reference counts of all the input data blocks for the same computation. We demonstrate the efficacy of LRC through both empirical analysis and cluster deployments against popular benchmarking workloads. Our Spark implementation shows that, the proposed policies well address the all-or-nothing requirement and significantly improve the cache performance. Compared with LRU and a recently proposed caching policy called MEMTUNE, LRC improves the caching performance of typical workloads in production clusters by 22 and 284 percent, respectively.
Yinghao Yu, Chengliang Zhang, Wei Wang 0030, Jun Zhang 0004, Khaled Ben Letaief
IEEE Trans. Cloud Comput.4
2022 Learn to Communicate With Neural Calibration: Scalability and Generalization
abstract
The conventional design of wireless communication systems typically relies on established mathematical models that capture the characteristics of different communication modules. Unfortunately, such design cannot be easily and directly applied to future wireless networks, which will be characterized by large-scale ultra-dense networks whose design complexity scales exponentially with the network size. Furthermore, such networks will vary dynamically in a significant way, which makes it intractable to develop comprehensive analytical models. Recently, deep learning-based approaches have emerged as potential alternatives for designing complex and dynamic wireless systems. However, existing learning-based methods have limited capabilities to scale with the problem size and to generalize with varying network settings. In this paper, we propose a scalable and generalizable neural calibration framework for future wireless system design, where a neural network is adopted to calibrate the input of conventional model-based algorithms. Specifically, the backbone of a traditional time-efficient algorithm is integrated with deep neural networks to achieve a high computational efficiency, while enjoying enhanced performance. The permutation equivariance property, carried out by the topological structure of wireless systems, is furthermore utilized to develop a generalizable neural network architecture. The proposed neural calibration framework is applied to solve challenging resource management problems in massive multiple-input multiple-output (MIMO) systems. Simulation results will show that the proposed neural calibration approach enjoys significantly improved scalability and generalization compared with the existing learning-based methods.
Yifei Shen 0004, Xianghao Yu, Jun Zhang 0004, Shenghui Song 0001, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.4
2021 How Powerful is Graph Convolution for Recommendation?
abstract
Graph convolutional networks (GCNs) have recently enabled a popular class of algorithms for collaborative filtering (CF). Nevertheless, the theoretical underpinnings of their empirical successes remain elusive. In this paper, we endeavor to obtain a better understanding of GCN-based CF methods via the lens of graph signal processing. By identifying the critical role of smoothness, a key concept in graph signal processing, we develop a unified graph convolution-based framework for CF. We prove that many existing CF methods are special cases of this framework, including the neighborhood-based methods, low-rank matrix factorization, linear auto-encoders, and LightGCN, corresponding to different low-pass filters. Based on our framework, we then present a simple and computationally efficient CF baseline, which we shall refer to as Graph Filter based Collaborative Filtering (GF-CF). Given an implicit feedback matrix, GF-CF can be obtained in a closed form instead of expensive training with back-propagation. Experiments will show that GF-CF achieves competitive or better performance against deep learning-based methods on three well-known datasets, notably with a 70% performance gain over LightGCN on the Amazon-book dataset.
Yifei Shen 0004, Yao Zhang 0009, Jun Zhang 0004, Khaled Ben Letaief, Dongsheng Li 0002
CIKM5
2021 Distributed Expectation Propagation Detection for Cell-Free Massive MIMO
abstract
In cell-free massive MIMO networks, an efficient distributed detection algorithm is of significant importance. In this paper, we propose a distributed expectation propagation (EP) detector for cell-free massive MIMO. The detector is composed of two modules, a nonlinear module at the central processing unit (CPU) and a linear module at the access point (AP). The turbo principle in iterative decoding is utilized to compute and pass the extrinsic information between modules. An analytical framework is then provided to characterize the asymptotic performance of the proposed EP detector with a large number of antennas. Simulation results will show that the proposed method outperforms the distributed detectors in terms of bit-error rate.
Hengtao He, Hanqing Wang 0002, Xianghao Yu, Jun Zhang 0004, Shenghui Song 0001, Khaled Ben Letaief
GLOBECOM4
2021 Neural Calibration for Scalable Beamforming in FDD Massive MIMO with Implicit Channel Estimation
abstract
Channel estimation and beamforming play critical roles in frequency-division duplexing (FDD) massive multiple-input multiple-output (MIMO) systems. However, these two modules have been treated as two stand-alone components, which makes it difficult to achieve a global system optimality. In this paper, we propose a deep learning-based approach that directly optimizes the beamformers at the base station according to the received uplink pilots, thereby, bypassing the explicit channel estimation. Different from the existing fully data-driven approach where all the modules are replaced by deep neural networks (DNNs), a neural calibration method is proposed to improve the scalability of the end-to-end design. In particular, the backbone of conventional time-efficient algorithms, i.e., the least-squares (LS) channel estimator and the zero-forcing (ZF) beamformer, is preserved and DNNs are leveraged to calibrate their inputs for better performance. The permutation equivariance property of the formulated resource allocation problem is then identified to design a low-complexity neural network architecture. Simulation results will show the superiority of the proposed neural calibration method over benchmark schemes in terms of both the spectral efficiency and scalability in large-scale wireless networks.
Yifei Shen 0004, Xianghao Yu, Jun Zhang 0004, Shenghui Song 0001, Khaled Ben Letaief
GLOBECOM4
2021 Communication-Computation Efficient Device-Edge Co-Inference via AutoML
abstract
Device-edge co-inference, which partitions a deep neural network between a resource-constrained mobile device and an edge server, recently emerges as a promising paradigm to support intelligent mobile applications. To accelerate the in-ference process, on-device model sparsification and intermediate feature compression are regarded as two prominent techniques. However, as the on-device model sparsity level and intermediate feature compression ratio have direct impacts on computation workload and communication overhead respectively, and both of them affect the inference accuracy, finding the optimal values of these hyper-parameters brings a major challenge due to the large search space. In this paper, we endeavor to develop an efficient algorithm to determine these hyper-parameters. By selecting a suitable model split point and a pair of encoder/decoder for the intermediate feature vector, this problem is casted as a sequential decision problem, for which, a novel automated machine learning (AutoML) framework is proposed based on deep reinforcement learning (DRL). Experiment results on an image classification task demonstrate the effectiveness of the proposed framework in achieving a better communication-computation trade-off and significant inference speedup against various baseline schemes.
Jiawei Shao, Yuyi Mao, Jun Zhang 0004
GLOBECOM4
2021 Branchy-GNN: A Device-Edge Co-Inference Framework for Efficient Point Cloud Processing
abstract
The recent advancements of three-dimensional (3D) data acquisition devices have spurred a new breed of applications that rely on point cloud data processing. However, processing a large volume of point cloud data brings a significant workload on resource-constrained mobile devices, prohibiting from unleashing their full potentials. Built upon the emerging paradigm of device-edge co-inference, where an edge device extracts and transmits the intermediate feature to an edge server for further processing, we propose Branchy-GNN for efficient graph neural network (GNN) based point cloud processing by leveraging edge computing platforms. In order to reduce the on-device computational cost, the Branchy-GNN adds branch networks for early exiting. Besides, it employs learning-based joint source-channel coding (JSCC) for the intermediate feature compression to reduce the communication overhead. Our experimental results demonstrate that the proposed Branchy-GNN secures a significant latency reduction compared with several benchmark methods.
Jiawei Shao, Yuyi Mao, Jun Zhang 0004
ICASSP4
2021 Supporting More Active Users for Massive Access via Data-assisted Activity Detection
abstract
Massive machine-type communication (mMTC) has been regarded as one of the most important use scenarios in the fifth generation (5G) and beyond wireless networks, which demands scalable access for a large number of devices. While grant-free random access has emerged as a promising mechanism for massive access, its potential has not been fully unleashed. Particularly, the two key tasks in massive access systems, namely, user activity detection and data detection, were handled separately in most existing studies, which ignored the common sparsity pattern in the received pilot and data signal. Moreover, error detection and correction in the payload data provide additional mechanisms for performance improvement. In this paper, we propose a data-assisted activity detection framework, which aims at supporting more active users by reducing the activity detection error, consisting of false alarm and missed detection errors. Specifically, after an initial activity detection step based on the pilot symbols, the false alarm users are filtered by applying energy detection for the data symbols; once data symbols of some active users have been successfully decoded, their effect in activity detection will be resolved via successive pilot interference cancellation, which reduces the missed detection error. Simulation results show that the proposed algorithm effectively increases the activity detection accuracy, and it is able to support ∼20% more active users compared to a conventional method in some sample scenarios.
Xinyu Bian, Yuyi Mao, Jun Zhang 0004
ICC3
2021 Long-term optimization for MEC-enabled HetNets with device-edge-cloud collaboration
Long Chen 0006, Jigang Wu, Jun Zhang 0004
Comput. Commun.3
2021 Graph Neural Networks for Scalable Radio Resource Management: Architecture Design and Theoretical Analysis
abstract
Deep learning has recently emerged as a disruptive technology to solve challenging radio resource management problems in wireless networks. However, the neural network architectures adopted by existing works suffer from poor scalability and generalization, and lack of interpretability. A long-standing approach to improve scalability and generalization is to incorporate the structures of the target task into the neural network architecture. In this paper, we propose to apply graph neural networks (GNNs) to solve large-scale radio resource management problems, supported by effective neural network architecture design and theoretical analysis. Specifically, we first demonstrate that radio resource management problems can be formulated as graph optimization problems that enjoy a universal permutation equivariance property. We then identify a family of neural networks, namedmessage passing graph neural networks(MPGNNs). It is demonstrated that they not only satisfy the permutation equivariance property, but also can generalize to large-scale problems, while enjoying a high computational efficiency. For interpretablity and theoretical guarantees, we prove the equivalence between MPGNNs and a family of distributed optimization algorithms, which is then used to analyze the performance and generalization of MPGNN-based methods. Extensive simulations, with power control and beamforming as two examples, demonstrate that the proposed method, trained in an unsupervised manner with unlabeled samples, matches or even outperforms classic optimization-based algorithms without domain-specific knowledge. Remarkably, the proposed method is highly scalable and can solve the beamforming problem in an interference channel with 1000 transceiver pairs within 6 milliseconds on a single GPU.
Yifei Shen 0004, Yuanming Shi, Jun Zhang 0004, Khaled Ben Letaief
IEEE J. Sel. Areas Commun.3
2021 Wireless Data Acquisition for Edge Learning: Data-Importance Aware Retransmission
abstract
By deploying machine-learning algorithms at the network edge, edge learning can leverage the enormous real-time data generated by billions of mobile devices to train AI models, which enable intelligent mobile applications. In this emerging research area, one key direction is to efficiently utilize radio resources for wireless data acquisition to minimize the latency of executing a learning task at an edge server. Along this direction, we consider the specific problem of retransmission decision in each communication round to ensure both reliability and quantity of those training data for accelerating model convergence. To solve the problem, a new retransmission protocol called data-importance aware automatic-repeat-request (importance ARQ) is proposed. Unlike the classic ARQ focusing merely on reliability, importance ARQ selectively retransmits a data sample based on its uncertainty which helps learning and can be measured using the model under training. Underpinning the proposed protocol is a derived elegant communication-learning relation between two corresponding metrics, i.e., signal-to-noise ratio (SNR) and data uncertainty. This relation facilitates the design of a simple threshold based policy for importance ARQ. The policy is first derived based on the classic classifier model of support vector machine (SVM), where the uncertainty of a data sample is measured by its distance to the decision boundary. The policy is then extended to the more complex model of convolutional neural networks (CNN) where data uncertainty is measured by entropy. Extensive experiments have been conducted for both the SVM and CNN using real datasets with balanced and imbalanced distributions. Experimental results demonstrate that importance ARQ effectively copes with channel fading and noise in wireless data acquisition to achieve faster model convergence than the conventional channel-aware ARQ. The gain is more significant when the dataset is imbalanced.
Dongzhu Liu, Guangxu Zhu, Qunsong Zeng, Jun Zhang 0004, Kaibin Huang
IEEE Trans. Wirel. Commun.4
2021 Blind Data Detection in Massive MIMO via ℓ₃-Norm Maximization Over the Stiefel Manifold
abstract
Massive MIMO has been regarded as a key enabling technique for 5G and beyond networks. Nevertheless, its performance is limited by the large overhead needed to obtain the high-dimensional channel information. To reduce the huge training overhead associated with conventional pilot-aided designs, we propose a novel blind data detection method by leveraging the channel sparsity and data concentration properties. Specifically, we propose a novel$\ell _{3}$-norm-based formulation to recover the data without channel estimation. We prove that the global optimal solution to the proposed formulation can be made arbitrarily close to the transmitted data up to a phase-permutation ambiguity. We then propose an efficient parameter-free algorithm to solve the$\ell _{3}$-norm problem and resolve the phase-permutation ambiguity. We also derive the convergence rate in terms of key system parameters such as the number of transmitters and receivers, the channel noise power, and the channel sparsity level. Numerical experiments will show that the proposed scheme has superior performance with low computational complexity.
Ye Xue, Yifei Shen 0004, Vincent K. N. Lau, Jun Zhang 0004, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.4
2021 Partially-Connected Hybrid Beamforming for Spectral Efficiency Maximization via a Weighted MMSE Equivalence
abstract
Hybrid beamforming (HBF) is an attractive technology for practical massive multiple-input and multiple-output (MIMO) millimeter wave (mmWave) systems. Compared with the fully-connected HBF architecture, the partially-connected one can further reduce the hardware cost and power consumption. However, the special block diagonal structure of its analog beamforming matrix brings additional design challenges. In this paper, we develop effective HBF algorithms for spectral efficiency maximization (SEM) in wideband mmWave massive MIMO systems with the partially-connected architecture. One main contribution is that we prove the equivalence of the SEM problem and a weighted mean square error minimization (WMMSE) problem, which leads to a convenient algorithmic approach to directly tackle the SEM problem. Specifically, we decompose the equivalent WMMSE problem into the hybrid precoding and hybrid combining subproblems, for which both the optimal digital precoder and combiner have closed-form solutions. For the more challenging analog precoder and combiner, we propose an element iteration based algorithm and a manifold optimization based algorithm. Finally, the hybrid precoder and combiner are alternatively updated. The overall HBF algorithms are proved to monotonously increase the spectral efficiency and converge. Furthermore, we also propose modified algorithms with reduced computational complexity and finite-resolution phase shifters. Simulation results demonstrate that the proposed HBF algorithms achieve significant performance gains over conventional algorithms.
Xingyu Zhao 0003, Tian Lin 0004, Yu Zhu 0002, Jun Zhang 0004
IEEE Trans. Wirel. Commun.4
2020 Bandit Sampling for Faster Activity and Data Detection in Massive Random Access
abstract
This paper considers the grant-free random access scheme in IoT networks with a massive number of devices. By embedding the data symbols in the signature sequences, joint device activity detection, and data decoding can be achieved, which, however, significantly increases the computational complexity. Coordinate descent algorithms, with a low per-iteration complexity, have been employed to solve the detection problem, but previous works typically employ a random coordinate selection policy which leads to slow convergence. This paper develops a bandit based strategy, i.e., bandit sampling, to speed up the convergence of coordinate descent. We exploit a multi-armed bandit algorithm to learn which coordinates will result in more aggressive descent of the objective function. Both convergence rate analysis and simulation results are provided to show that the proposed algorithm enjoys a faster convergence rate with a lower time complexity compared with the state-of-the-art algorithm.
Jialin Dong, Jun Zhang 0004, Yuanming Shi
ICASSP2
2020 Client-Edge-Cloud Hierarchical Federated Learning
abstract
Federated Learning is a collaborative machine learning framework to train a deep learning model without accessing clients’ private data. Previous works assume one central parameter server either at the cloud or at the edge. The cloud server can access more data but with excessive communication overhead and long latency, while the edge server enjoys more efficient communications with the clients. To combine their advantages, we propose a client-edge-cloud hierarchical Federated Learning system, supported with a HierFAVG algorithm that allows multiple edge servers to perform partial model aggregation. In this way, the model can be trained faster and better communication-computation trade-offs can be achieved. Convergence analysis is provided for HierFAVG and the effects of key parameters are also investigated, which lead to qualitative design guidelines. Empirical experiments verify the analysis and demonstrate the benefits of this hierarchical architecture in different data distribution scenarios. Particularly, it is shown that by introducing the intermediate edge servers, the model training time and the energy consumption of the end devices can be simultaneously reduced compared to cloud-based Federated Learning.
Lumin Liu, Jun Zhang 0004, Shenghui Song 0001, Khaled Ben Letaief
ICC2
2020 Complete Dictionary Learning via ℓp-norm Maximization
Yifei Shen 0004, Ye Xue, Jun Zhang 0004, Khaled Ben Letaief, Vincent K. N. Lau
UAI3
2020 Mobile Edge Intelligence and Computing for the Internet of Vehicles
abstract
The Internet of Vehicles (IoV) is an emerging paradigm that is driven by recent advancements in vehicular communications and networking. Meanwhile, the capability and intelligence of vehicles are being rapidly enhanced, and this will have the potential of supporting a plethora of new exciting applications that will integrate fully autonomous vehicles, the Internet of Things (IoT), and the environment. These trends will bring about an era of intelligent IoV, which will heavily depend on communications, computing, and data analytics technologies. To store and process the massive amount of data generated by intelligent IoV, onboard processing and cloud computing will not be sufficient due to resource/power constraints and communication overhead/latency, respectively. By deploying storage and computing resources at the wireless network edge, e.g., radio access points, the edge information system (EIS), including edge caching, edge computing, and edge AI, will play a key role in the future intelligent IoV. EIS will provide not only low-latency content delivery and computation services but also localized data acquisition, aggregation, and processing. This article surveys the latest development in EIS for intelligent IoV. Key design issues, methodologies, and hardware platforms are introduced. In particular, typical use cases for intelligent vehicles are illustrated, including edge-assisted perception, mapping, and localization. In addition, various open-research problems are identified.
Jun Zhang 0004, Khaled Ben Letaief
Proc. IEEE1
2020 Achieving Load-Balanced, Redundancy-Free Cluster Caching with Selective Partition
abstract
Data-intensive clusters increasingly rely on in-memory storages to improve I/O performance. However, the routinely observed file popularity skew and load imbalance create hot spots, which significantly degrade the benefits of in- memory caching. Common approaches to tame load imbalance include copying multiple replicas of hot files and creating parity chunks using storage codes. Yet, these techniques either suffer from high memory overhead due to cache redundancy or incur non-trivial encoding/decoding complexity. In this paper, we propose an effective approach to achieve load balancing without cache redundancy or encoding/decoding overhead. Our solution, termed SP-Cache,selectively partitionsfiles based on the loads they contribute and evenly caches those partitions across the cluster. We develop an efficient algorithm to determine the optimal number of partitions for a hot file—too few partitions are incapable of mitigating hot spots, while too many are susceptible to stragglers. We have implemented SP-Cache atop Alluxio, a popular in-memory distributed storage system, and evaluated its performance through EC2 deployment and trace-driven simulations. SP-Cache can quickly react to the changing load by dynamically re-balancing cache servers. Compared to the state-of-the-art solution, SP-Cache reduces the file access latency by up to 40 percent in both the mean and the tail, using 40 percent less memory.
Yinghao Yu, Wei Wang 0030, Renfei Huang, Jun Zhang 0004, Khaled Ben Letaief
IEEE Trans. Parallel Distributed Syst.4
2020 LORM: Learning to Optimize for Resource Management in Wireless Networks With Few Training Samples
abstract
Effective resource management plays a pivotal role in wireless networks, which, unfortunately, typically results in challenging mixed-integer nonlinear programming (MINLP) problems. Machine learning-based methods have recently emerged as a disruptive way to obtain near-optimal performance for MINLPs with affordable computational complexity. There have been some attempts in applying such methods to resource management in wireless networks, but these attempts require huge amounts of training samples and lack the capability to handle constrained problems. Furthermore, they suffer from severe performance deterioration when the network parameters change, which commonly happens and is referred to as thetask mismatchproblem. In this paper, to reduce the sample complexity and address the feasibility issue, we propose a framework of Learning to Optimize for Resource Management (LORM). In contrast to the end-to-end learning approach adopted in previous studies, LORM learns the optimal pruning policy in the branch-and-bound algorithm for MINLPs via a sample-efficient method, namely,imitation learning. To further address the task mismatch problem, we develop a transfer learning method via self-imitation in LORM, namedLORM-TL, which can quickly adapt a pre-trained machine learning model to the new task with only a few additionalunlabeledtraining samples. Numerical simulations demonstrate that LORM outperforms specialized state-of-the-art algorithms and achieves near-optimal performance, while providing significant speedup compared with the branch-and-bound algorithm. Moreover, LORM-TL, by relying on a few unlabeled samples, achieves comparable performance with the model trained from scratch with sufficient labeled samples.
Yifei Shen 0004, Yuanming Shi, Jun Zhang 0004, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.3
2019 Transfer Learning for Mixed-Integer Resource Allocation Problems in Wireless Networks
abstract
Effective resource allocation plays a pivotal role in wireless networks. Unfortunately, typical resource allocation problems are mixed-integer nonlinear programming (MINLP) problems, which are NP-hard. Machine learning based methods recently emerge as a disruptive way to obtain near-optimal performance for MINLP problems with affordable computational complexity. However, they suffer from severe performance deterioration when the network parameters change, which commonly happens in practice and can be characterized as the task mismatch issue. In this paper, we propose a transfer learning method via self-imitation, to address this issue for effective resource allocation in wireless networks. It is based on a general “learning to optimize” framework for solving MINLP problems. A unique advantage of the proposed method is that it can tackle the task mismatch issue with a few additional unlabeled training samples, which is especially important when transferring to large-size problems. Numerical experiments demonstrate that the proposed method, with much less training time, achieves comparable performance with the model trained from scratch based on sufficient labeled samples. To the best of our knowledge, this is the first work that applies transfer learning for resource allocation in wireless networks.
Yifei Shen 0004, Yuanming Shi, Jun Zhang 0004, Khaled Ben Letaief
ICC3
2019 LACS: Load-Aware Cache Sharing with Isolation Guarantee
abstract
Cluster caching has been increasingly deployed in front of cloud storage to improve I/O performance. In shared, multi-tenant environments such as cloud datacenters, cluster caches are constantly contended by many users. Enforcing performance isolation between users hence becomes imperative to cluster caching. A user's caching performance critically depends on two factors: (1) the amount of cache allocation and (2) the load of servers in which its files are cached. However, existing cache sharing policies only provide guarantees on the amount of cache allocation, while remaining agnostic to the load of cache servers. Consequently, "mice" users having files co-located with "elephants" contributing heavy data accesses may experience extremely long latency, hence receiving no isolation. In this paper, we propose a Load-Aware Cache Sharing scheme (LACS) to enforce isolation between users. LACS keeps track of the load contributed by each user and reins back the congestions caused by elephant users by throttling their cache usage and network bandwidth. We have implemented LACS atop Alluxio, a popular cluster caching system. EC2 deployment shows that LACS achieves performance isolation in the presence of elephants, while improving the mean read latency by up to 80.4% (25.3% on average) over the state-of-the-art load balancing technique.
Yinghao Yu, Wei Wang 0030, Jun Zhang 0004, Khaled Ben Letaief
ICDCS3
2019 Connectivity-Aware UAV Path Planning with Aerial Coverage Maps
abstract
Cellular networks are promising to support effective wireless communications for unmanned aerial vehicles (UAVs), which will help to enable various long-range UAV applications. However, these networks are optimized for terrestrial users, and thus do not guarantee seamless aerial coverage. In this paper, we propose to overcome this difficulty by exploiting controllable mobility of UAVs, and investigate connectivity-aware UAV path planning. To explicitly impose communication requirements on UAV path planning, we introduce two new metrics to quantify the cellular connectivity quality of a UAV path. Moreover, aerial coverage maps are used to provide accurate locations of scattered coverage holes in the complicated propagation environment. We formulate the UAV path planning problem as finding the shortest path subject to connectivity constraints. Based on graph search methods, a novel connectivity-aware path planning algorithm with low complexity is proposed. The effectiveness and superiority of our proposed algorithm are demonstrated using the aerial coverage map of an urban section in Virginia, which is built by ray tracing. Simulation results also illustrate a tradeoff between the path length and connectivity quality of UAVs.
Jun Zhang 0004, Shenghui Song 0001, Khaled Ben Letaief
WCNC2
2019 Joint Activity Detection and Channel Estimation for IoT Networks: Phase Transition and Computation-Estimation Tradeoff
abstract
Massive device connectivity is a crucial communication challenge for Internet of Things (IoT) networks, which consist of a large number of devices with sporadic traffic. In each coherence block, the serving base station needs to identify the active devices and estimate their channel state information for effective communication. By exploiting the sparsity pattern of data transmission, we develop a structured group sparsity estimation method to simultaneously detect the active devices and estimate the corresponding channels. This method significantly reduces the signature sequence length while supporting massive IoT access. To determine the optimal signature sequence length, we study the phase transition behavior of the group sparsity estimation problem. Specifically, user activity can be successfully estimated with a high probability when the signature sequence length exceeds a threshold; otherwise, it fails with a high probability. The location and width of the phase transition region are characterized via the theory of conic integral geometry. We further develop a smoothing method to solve the high-dimensional structured estimation problem with a given limited time budget. This is achieved by sharply characterizing the convergence rate in terms of the smoothing parameter, signature sequence length and estimation accuracy, yielding a tradeoff between the estimation accuracy and computational cost. Numerical results are provided to illustrate the accuracy of our theoretical results and the benefits of smoothing techniques.
Tao Jiang 0016, Yuanming Shi, Jun Zhang 0004, Khaled Ben Letaief
IEEE Internet Things J.3
2019 Hybrid Beamforming for Millimeter Wave Systems Using the MMSE Criterion
abstract
Hybrid analog and digital beamforming (HBF) has recently emerged as an attractive technique for millimeter-wave (mmWave) communication systems. It well balances the demand for sufficient beamforming gains to overcome the propagation loss and the desire to reduce the hardware cost and power consumption. In this paper, the mean square error (MSE) is chosen as the performance metric to characterize the transmission reliability. Using the minimum sum-MSE criterion, we investigate the HBF design for broadband mmWave transmissions. To overcome the difficulty of solving the multi-variable design problem, the alternating minimization method is adopted to optimize the hybrid transmit and receive beamformers alternatively. Specifically, a manifold optimization-based HBF algorithm is first proposed, which directly handles the constant modulus constraint of the analog component. Its convergence is then proved. To reduce the computational complexity, we then propose a low-complexity general eigenvalue decomposition-based HBF algorithm in the narrowband scenario and three algorithms via the eigenvalue decomposition and orthogonal matching pursuit methods in the broadband scenario. A particular innovation in our proposed alternating minimization algorithms is a carefully designed initialization method, which leads to a faster convergence. Furthermore, we extend the sum-MSE-based design to that with weighted sum-MSE, which is then connected to the spectral efficiency-based design. Simulation results show that the proposed HBF algorithms achieve a significant performance improvement over existing ones and perform close to full-digital beamforming.
Tian Lin 0004, Jiaqi Cong, Yu Zhu 0002, Jun Zhang 0004, Khaled Ben Letaief
IEEE Trans. Commun.4
2018 Joint Device Caching and Channel Allocation for D2D-Assisted Wireless Content Delivery
abstract
To exploit the potential of content caching and device-to-device (D2D) communication, we propose a user-centric joint device caching and channel assignment (DCA) policy to facilitate content exchanges between user equipments (UEs). The objective is to minimize the average content delivery delay by effectively leveraging D2D communications using as few channels as possible, subject to the UEs' cache capacities and availability of D2D links. This joint design problem is formulated as a nonlinear combinatorial optimization problem which is NP-hard. We first analyze the optimal DCA policy in two special cases. Then, a low-complexity heuristic algorithm is proposed for general cases which alternatively performs greedy device caching and graphcoloring based channel allocating. Simulation results show that the proposed DCA policy can reduce the average content delivery delay by more than half, in contrast to baseline schemes with locally popular caching.
Juan Liu 0002, Bo Bai 0001, Jun Zhang 0004, Khaled Ben Letaief, Youming Li
ICC3
2018 OpuS: Fair and Efficient Cache Sharing for In-Memory Data Analytics
abstract
We study the fair cache allocation problem in shared cloud environments, where many users and applications contend for the main memory to cache shared datasets or files. Unlike other resources such as CPUs and networks, in-memory caches can be non-exclusively shared across many users, e.g., a cached columnar dataset queried by many Spark SQL jobs. This results in a unique challenge of the "free-riding" problem, where a user lies about its caching preferences to trick other users to cache files for it, using their allocated cache space. We show that existing cache allocation policies either suffer from such manipulations or result in poor efficiency. To address this problem, we propose a new cache allocation algorithm, termed OpuS, or Opportunistic Sharing for high efficiency. We show that OpuS provides performance isolation between users and is strategy-proof against "free-riding" manipulations. We have implemented OpuS as a pluggable cache manager in Alluxio, a popular memory-centric filesystem. Cluster deployment and trace-driven simulations demonstrate that OpuS allocates each user a fair share of caches while achieving near-optimal efficiency in cache utilization.
Yinghao Yu, Wei Wang 0030, Jun Zhang 0004, Qizhen Weng 0001, Khaled Ben Letaief
ICDCS3
2018 A Multi-dimension Measurement Study of a Large Scale Campus WiFi Network
abstract
The growing trend of wireless devices and WiFi networks poses significant management challenges to network administrators. Characterizing WiFi user behavior and understanding WiFi network usage pattern are helpful to identify the management challenges so that network administrators could manage WiFi networks more efficiently. In this work, we collect comprehensive datasets, i.e., DHCP dataset, AAA dataset, SNMP dataset of ACs in a large campus WiFi network. We provide a detailed measurement study from multiple dimensions, i.e., server plane, temporal plane, spatial plane and traffic plane. We observe that the WiFi network under study is far from optimal. First, the phenomenon of IP waste is severe due to the isolation between DHCP server and AAA server. Second, current deployment of network infrastructure resources is based on network administrators' experience and it results in that the WiFi performance varies a lot across different areas. Furthermore, we also study the user behavior with different types of devices and in different kinds of buildings. Our observations indicate that the WiFi network could be improved and managed more efficiently from multiple dimensions. We believe that this measurement study is helpful for network administrators and researchers to understand more about large scale WiFi networks.
Congcong Miao, Jilong Wang 0001, Hui Wang 0011, Jun Zhang 0004, Shengchao Liu
LCN4
2018 SP-cache: load-balanced, redundancy-free cluster caching with selective partition
Yinghao Yu, Renfei Huang, Wei Wang 0030, Jun Zhang 0004, Khaled Ben Letaief
SC4
2018 Massive CSI Acquisition for Dense Cloud-RANs With Spatial-Temporal Dynamics
abstract
Dense cloud radio access networks (cloud-RANs) provide a promising way to enable scalable connectivity and handle diversified service requirements for massive mobile devices. To fully exploit the performance gains of dense cloud-RANs, channel state information of both the signal link and interference links is required. However, with limited radio resources for training, the channel estimation problem in dense cloud-RANs becomes a high-dimensional estimation problem, i.e., the number of measurements will be typically smaller than the dimension of the channel. In this paper, we shall develop a generic high-dimensional structured channel estimation framework for dense cloud-RANs, which is based on a convex structured regularizing formulation. Observing that the wireless channel possesses ample exploitable statistical characteristics, we propose to convert the available spatial and temporal prior information into appropriate convex regularizers. Simulation results demonstrate that exploiting the spatial and temporal dynamics can achieve good estimation performance even with limited training resources. The alternating direction method of multipliers algorithm is further adopted to solve the resultant large-scale high-dimensional channel estimation problems. The proposed framework thus enjoys modeling flexibility, low training overhead, and computation cost scalability.
Xuan Liu 0005, Yuanming Shi, Jun Zhang 0004, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.3
2018 Enhanced Group Sparse Beamforming for Green Cloud-RAN: A Random Matrix Approach
abstract
Group sparse beamforming is a general framework to minimize the network power consumption for cloud radio access networks, which, however, suffers high computational complexity. In particular, a complex optimization problem needs to be solved to obtain the remote radio head (RRH) ordering criterion in each transmission block, which will help to determine the active RRHs and the associated fronthaul links. In this paper, we propose innovative approaches to reduce the complexity of this key step in group sparse beamforming. Specifically, we first develop a smoothed ℓp-minimization approach with the iterative reweighted-ℓ2algorithm to return a Karush-Kuhn- Tucker (KKT) point solution, as well as enhance the capability of inducing group sparsity in the beamforming vectors. By leveraging the Lagrangian duality theory, we obtain closedform solutions at each iteration to reduce the computational complexity. The well-structured solutions provide opportunities to apply the large-dimensional random matrix theory to derive deterministic approximations for the RRH ordering criterion. Such an approach helps to guide the RRH selection only based on the statistical channel state information, which does not require frequent update, thereby significantly reducing the computation overhead. Simulation results shall demonstrate the performance gains of the proposed ℓp-minimization approach, as well as the effectiveness of the large system analysis-based framework for computing the RRH ordering criterion.
Yuanming Shi, Jun Zhang 0004, Wei Chen 0002, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.2
2018 Exploiting Mobility in Cache-Assisted D2D Networks: Performance Analysis and Optimization
abstract
Caching popular content at mobile devices, accompanied by device-to-device (D2D) communications, is one promising technology for effective mobile content delivery. User mobility is an important factor when investigating such networks, which unfortunately was largely ignored in most previous works. Preliminary studies have been carried out but the effect of mobility on the caching performance has not been fully understood. In this paper, by explicitly considering users’ contact and inter-contact durations via an alternating renewal process, we first investigate the effect of mobility with a given cache placement. A tractable expression of the data offloading ratio, i.e., the proportion of requested data that can be delivered via D2D links, is derived, which is proved to be increasing with the user moving speed. The analytical results are then used to develop an effective mobility-aware caching strategy to maximize the data offloading ratio. Simulation results are provided to confirm the accuracy of the analytical results and also validate the effect of user mobility. Performance gains of the proposed mobility-aware caching strategy are demonstrated with both stochastic models and real-life data sets. It is observed that the information of the contact durations is critical to design cache placement, especially when they are relatively short or comparable to the inter-contact durations.
Rui Wang 0028, Jun Zhang 0004, Shenghui Song 0001, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.2
2018 A Unified Framework for the Tractable Analysis of Multi-Antenna Wireless Networks
abstract
Densifying networks and deploying more antennas at each access point are two principal ways to boost the capacity of wireless networks. However, the complicated distributions of the signal power and the accumulated interference power, largely induced by various space-time processing techniques, make it highly challenging to quantitatively characterize the performance of multi-antenna networks. In this paper, using tools from stochastic geometry, a unified framework is developed for the analysis of such networks. The major results are two innovative representations of the coverage probability, which make the analysis of multi-antenna networks almost as tractable as the single-antenna case. One is expressed as an ℓ1-induced norm of a Toeplitz matrix, and the other is given in a finite sum form. With a compact representation, the former incorporates many existing analytical results on single- and multi-antenna networks as special cases and leads to tractable expressions for evaluating the coverage probability in both ad hoc and cellular networks. While the latter is more complicated for numerical evaluation, it helps analytically gain key design insights. In particular, it helps prove that the coverage probability of ad hoc networks is a monotonically decreasing convex function of the transmitter density and that there exists a peak value of the coverage improvement when increasing the number of transmit antennas. On the other hand, in multi-antenna cellular networks, it is shown that the coverage probability is independent of the transmitter density and that the outage probability decreases exponentially as the number of transmit antennas increases.
Xianghao Yu, Chang Li 0002, Jun Zhang 0004, Martin Haenggi, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.3
2017 LERC: Coordinated Cache Management for Data-Parallel Systems
abstract
Memory caches are being aggressively used in today's data- parallel frameworks such as Spark, Tez and Storm. By caching input and intermediate data in memory, compute tasks can witness speedup by orders of magnitude. To maximize the chance of in-memory data access, existing cache algorithms, be it recency- or frequency-based, settle on cache hit ratio as the optimization objective. However, unlike the conventional belief, we show in this paper that simply pursuing a higher cache hit ratio of individual data blocks does not necessarily translate into faster task completion in data-parallel environments. A data-parallel task typically depends on multiple input data blocks. Unless all of these blocks are cached in memory, no speedup will result. To capture this all-or-nothing property, we propose a more relevant metric, called effective cache hit ratio. Specifically, a cache hit of a data block is said to be effective if it can speed up a compute task. In order to optimize the effective cache hit ratio, we propose the Least Effective Reference Count (LERC) policy that persists the dependent blocks of a compute task as a whole in memory. We have implemented the LERC policy as a memory manager in Spark and evaluated its performance through Amazon EC2 deployment. Evaluation results demonstrate that LERC helps speed up data-parallel jobs by up to 37% compared with the widely employed least-recently-used (LRU) policy.
Yinghao Yu, Wei Wang 0030, Jun Zhang 0004, Khaled Ben Letaief
GLOBECOM3
2017 Hybrid Precoding in Millimeter Wave Systems: How Many Phase Shifters Are Needed?
abstract
Hybrid precoding has been recently proposed as a cost- effective transceiver solution for millimeter wave (mm- wave) systems. The analog component in such precoders, which is composed of a phase shifter network, is the key differentiating element in contrast to conventional fully digital precoders. While a large number of phase shifters with unquantized phases are commonly assumed in existing works, in practice the phase shifters should be discretized with a coarse quantization, and their number should be reduced to a minimum due to cost and power consideration. In this paper, we propose a new hybrid precoder implementation using a small number of phase shifters with quantized and fixed phases, i.e., a fixed phase shifter (FPS) implementation, which significantly reduces the cost and hardware complexity. In addition, a dynamic switch network is proposed to enhance the spectral efficiency. Based on the proposed FPS implementation, an effective alternating minimization (AltMin) algorithm is developed with closed-form solutions in each iteration. Simulation results show that the proposed algorithm with the FPS implementation outperforms existing ones. More importantly, it needs much fewer phase shifters than existing hybrid precoder proposals, e.g., around 10 fixed phase shifters are sufficient for practically relevant system settings.
Xianghao Yu, Jun Zhang 0004, Khaled Ben Letaief
GLOBECOM2
2017 Massive CSI acquisition in dense cloud-RAN with spatial and temporal prior information
abstract
In this paper, we shall develop a generic channel estimation framework based on the convex formulation for dense cloud radio access networks (Cloud-RAN). Due to the training resource constraint and the large number of transmit antennas, the pilot length is smaller than the antenna number, and thus channel estimation becomes an ill-posed inverse problem. By observing that the wireless channel possesses ample exploitable statistical characteristics, we propose to convert the available spatial and temporal prior information into appropriate convex regularizing functions, yielding convex optimization formulations for the underdetermined channel estimation problem. Simulation results demonstrate that exploiting the prior information of large-scale fading and temporal correlation can achieve good estimation performance even with limited training resources. The alternating direction method of multipliers (ADMM) algorithm is further adopted to solve the resultant large-scale channel estimation problems. The proposed framework is, therefore, scalable to the overhead of prior information and the computation cost for large network sizes.
Xuan Liu 0005, Yuanming Shi, Jun Zhang 0004, Khaled Ben Letaief
ICC3
2017 Mobility increases the data offloading ratio in D2D caching networks
abstract
Caching at mobile devices, accompanied by device-to-device (D2D) communications, is one promising technique to accommodate the exponentially increasing mobile data traffic. While most previous works ignored user mobility, there are some recent works taking it into account. However, the duration of user contact times has been ignored, making it difficult to explicitly characterize the effect of mobility. In this paper, we adopt the alternating renewal process to model the duration of both the contact and inter-contact times, and investigate how the caching performance is affected by mobility. The data offloading ratio, i.e., the proportion of requested data that can be delivered via D2D links, is taken as the performance metric. We first approximate the distribution of the communication time for a given user by beta distribution through moment matching. With this approximation, an accurate expression of the data offloading ratio is derived. For the homogeneous case where the average contact and intercontact times of different user pairs are identical, we prove that the data offloading ratio increases with the user moving speed, assuming that the transmission rate remains the same. Simulation results are provided to show the accuracy of the approximate result, and also validate the effect of user mobility.
Rui Wang 0028, Jun Zhang 0004, Shenghui Song 0001, Khaled Ben Letaief
ICC2
2017 A tractable framework for performance analysis of dense multi-antenna networks
abstract
Densifying the network and deploying more antennas at each access point are two principal ways to boost the capacity of wireless networks. However, due to the complicated distributions of random signal and interference channel gains, largely induced by various space-time processing techniques, it is highly challenging to quantitatively characterize the performance of dense multi-antenna networks. In this paper, using tools from stochastic geometry, a tractable framework is proposed for the analytical evaluation of such networks. The major result is an innovative representation of the coverage probability, as an induced ℓ1-norm of a Toeplitz matrix. This compact representation incorporates lots of existing analytical results on single-and multi-antenna networks as special cases, and its evaluation is almost as simple as the single-antenna case with Rayleigh fading. To illustrate its effectiveness, we apply the proposed framework to investigate two kinds of prevalent dense wireless networks, i.e., physical layer security aware networks and millimeter-wave networks. In both examples, in addition to tractable analytical results of relevant performance metrics, insightful design guidelines are also analytically obtained.
Xianghao Yu, Chang Li 0002, Jun Zhang 0004, Khaled Ben Letaief
ICC3
2017 LRC: Dependency-aware cache management for data analytics clusters
abstract
Memory caches are being aggressively used in today's data-parallel systems such as Spark, Tez, and Piccolo. However, prevalent systems employ rather simple cache management policies — notably the Least Recently Used (LRU) policy — that are oblivious to the application semantics of data dependency, expressed as a directed acyclic graph (DAG). Without this knowledge, memory caching can at best be performed by “guessing” the future data access patterns based on historical information (e.g., the access recency and/or frequency), which frequently results in inefficient, erroneous caching with low hit ratio and a long response time. In this paper, we propose a novel cache replacement policy, Least Reference Count (LRC), which exploits the application-specific DAG information to optimize the cache management. LRC evicts the cached data blocks whose reference count is the smallest. The reference count is defined, for each data block, as the number of dependent child blocks that have not been computed yet. We demonstrate the efficacy of LRC through both empirical analysis and cluster deployments against popular benchmarking workloads. Our Spark implementation shows that, compared with LRU, LRC speeds up typical applications by 60%.
Yinghao Yu, Wei Wang 0030, Jun Zhang 0004, Khaled Ben Letaief
INFOCOM3
2017 Multi-objective resource allocation for mobile edge computing systems
abstract
To enhance the computation capability of mobile devices by offloading computation-demanding tasks to the nearby edge servers. In order to minimize the task latency and the device energy consumption, in this paper, we investigate the multi-objective resource allocation for multi-user MEC systems by adopting the system utility as the performance metric, which is a normalized weighted combination of the time and energy saving achieved by computation offloading. To provide an efficient solution, a low-complexity ranking-based algorithm is proposed based on the modified Newton method and the concept of computation offloading priority. Simulation results show that our proposed algorithm achieves a near-optimal performance and greatly outperforms a baseline algorithm with random spectrum allocation. In addition, it is demonstrated that jointly optimizing the spectrum and computational resource management policy is more critical when the number of MEC users is large.
Yuyi Mao, Jun Zhang 0004, Khaled Ben Letaief
PIMRC3
2017 Joint Task Offloading Scheduling and Transmit Power Allocation for Mobile-Edge Computing Systems
abstract
Mobile-edge computing (MEC) has emerged as a prominent technique to provide mobile services with high computation requirement, by migrating the computation- intensive tasks from the mobile devices to the nearby MEC servers. To reduce the execution latency and device energy consumption, in this paper, we jointly optimize task offloading scheduling and transmit power allocation for MEC systems with multiple independent tasks. A low-complexity sub-optimal algorithm is proposed to minimize the weighted sum of the execution delay and device energy consumption based on alternating minimization. Specifically, given the transmit power allocation, the optimal task off loading scheduling, i.e., to determine the order of offloading, is obtained with the help of flow shop scheduling theory. Besides, the optimal transmit power allocation with a given task offloading scheduling decision will be determined using convex optimization techniques. Simulation results show that task offloading scheduling is more critical when the available radio and computational resources in MEC systems are relatively balanced. In addition, it is shown that the proposed algorithm achieves near-optimal execution delay along with a substantial device energy saving.
Yuyi Mao, Jun Zhang 0004, Khaled Ben Letaief
WCNC2
2017 Coverage Analysis for Millimeter Wave Networks: The Impact of Directional Antenna Arrays
abstract
Millimeter wave (mm-wave) communications is considered a promising technology for 5G networks. Exploiting beamforming gains with large-scale antenna arrays to combat the increased path loss at mm-wave bands is one of the defining features. However, previous works on mm-wave network analysis usually adopted oversimplified antenna patterns for tractability, which can lead to significant deviation from the performance with actual antenna patterns. In this paper, using tools from stochastic geometry, we carry out a comprehensive investigation on the impact of directional antenna arrays in mm-wave networks. We first present a general and tractable framework for coverage analysis with arbitrary distributions for interference power and arbitrary antenna patterns. It is then applied to mm-wave ad hoc and cellular networks, where two sophisticated antenna patterns with desirable accuracy and analytical tractability are proposed to approximate the actual antenna pattern. Compared with previous works, the proposed approximate antenna patterns help to obtain more insights on the role of directional antenna arrays in mm-wave networks. In particular, it is shown that the coverage probabilities of both types of networks increase as a non-decreasing concave function with the antenna array size. The analytical results are verified to be effective and reliable through simulations, and numerical results also show that large-scale antenna arrays are required for satisfactory coverage in mm-wave networks.
Xianghao Yu, Jun Zhang 0004, Martin Haenggi, Khaled Ben Letaief
IEEE J. Sel. Areas Commun.2
2017 Layered Group Sparse Beamforming for Cache-Enabled Green Wireless Networks
abstract
The exponential growth of mobile data traffic is driving the deployment of dense wireless networks, which will not only impose heavy backhaul burdens, but also generate considerable power consumption. Introducing caches to the wireless network edge is a potential and cost-effective solution to address these challenges. In this paper, we will investigate the problem of minimizing the network power consumption of cache-enabled wireless networks, consisting of the base station (BS) and backhaul power consumption. The objective is to develop efficient algorithms that unify adaptive BS selection, backhaul content assignment, and multicast beamforming, while taking account of user QoS requirements and backhaul capacity limitations. To address the NP-hardness of the network power minimization problem, we first propose a generalized layered group sparse beamforming (LGSBF) modeling framework, which helps to reveal the layered sparsity structure in the beamformers. By adopting the reweighted ℓ1/ℓ2-norm technique, we further develop a convex approximation procedure for the LGSBF problem, followed by a three-stage iterative LGSBF framework to induce the desired sparsity structure in the beamformers. Simulation results validate the effectiveness of the proposed algorithm in reducing the network power consumption, and demonstrate that caching plays a more significant role in networks with higher user densities and less power-efficient backhaul links.
Xi Peng 0006, Yuanming Shi, Jun Zhang 0004, Khaled Ben Letaief
IEEE Trans. Commun.3
2017 Joint Fronthaul Multicast Beamforming and User-Centric Clustering in Downlink C-RANs
abstract
The cloud radio access network (C-RAN) has been deemed a cost-effective architecture for exploiting the capacity benefit of densely deployed radio access points. The low-latency fronthaul data transmission from the central processor to small-cell base stations (SBSs) is a key requirement in C-RANs for which conventional wired fronthaul links will be cost-prohibitive and also inconvenient. Therefore, scalable and low-cost wireless fronthaul solutions have drawn much attention in both industry and academia. In this paper, we propose adopting the multicast beamforming strategy over fronthaul links to deliver each user's message to a cluster of SBSs selected according to the user-centric clustering scheme, which then adopts the joint beamforming technique to cooperatively transmit the signal to the target users. Some approximate techniques are applied to obtain a tractable formulation for this mixed integer nonlinear programming problem, and an iterative algorithm based on the block coordinate update method is proposed accordingly. Then, a binary search based algorithm is developed to preserve the sparsity of beamformers due to the relaxation of the discrete clustering function with the continuous exponential function. Extensive simulation results are provided to show the performance of the proposed algorithms in terms of convergence, power consumption, and weighted sum rate.
Cunqing Hua, Jun Zhang 0004, Cailian Chen, Xin-Ping Guan
IEEE Trans. Wirel. Commun.3
2017 Cache Placement in Fog-RANs: From Centralized to Distributed Algorithms
abstract
To deal with the rapid growth of high-speed and/or ultra-low latency data traffic for massive mobile users, fog radio access networks (Fog-RANs) have emerged as a promising architecture for next-generation wireless networks. In Fog-RANs, the edge nodes and user terminals possess storage, computation and communication functionalities to various degrees, which provide high flexibility for network operation, i.e., from fully centralized to fully distributed operation. In this paper, we study the cache placement problem in Fog-RANs, by taking into account flexible physical-layer transmission schemes and diverse content preferences of different users. We develop both centralized and distributed transmission aware cache placement strategies to minimize users' average download delay subject to the storage capacity constraints. In the centralized mode, the cache placement problem is transformed into a matroid constrained submodular maximization problem, and an approximation algorithm is proposed to find a solution within a constant factor to the optimum. In the distributed mode, a belief propagation-based distributed algorithm is proposed to provide a suboptimal solution, with iterative updates at each BS based on locally collected information. Simulation results show that by exploiting caching and cooperation gains, the proposed transmission aware caching algorithms can greatly reduce the users' average download delay.
Juan Liu 0002, Bo Bai 0001, Jun Zhang 0004, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.3
2017 Stochastic Joint Radio and Computational Resource Management for Multi-User Mobile-Edge Computing Systems
abstract
Mobile-edge computing (MEC) has recently emerged as a prominent technology to liberate mobile devices from computationally intensive workloads, by offloading them to the proximate MEC server. To make offloading effective, the radio and computational resources need to be dynamically managed, to cope with the time-varying computation demands and wireless fading channels. In this paper, we develop an online joint radio and computational resource management algorithm for multi-user MEC systems, with the objective of minimizing the long-term average weighted sum power consumption of the mobile devices and the MEC server, subject to a task buffer stability constraint. Specifically, at each time slot, the optimal CPU-cycle frequencies of the mobile devices are obtained in closed forms, and the optimal transmit power and bandwidth allocation for computation offloading are determined with the Gauss-Seidel method; while for the MEC server, both the optimal frequencies of the CPU cores and the optimal MEC server scheduling decision are derived in closed forms. Besides, a delay-improved mechanism is proposed to reduce the execution delay. Rigorous performance analysis is conducted for the proposed algorithm and its delay-improved version, indicating that the weighted sum power consumption and execution delay obey an [O (1/V) , O (V)] tradeoff with V as a control parameter. Simulation results are provided to validate the theoretical analysis and demonstrate the impacts of various parameters.
Yuyi Mao, Jun Zhang 0004, Shenghui Song 0001, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.2
2017 Mobility-Aware Caching in D2D Networks
abstract
Caching at mobile devices can facilitate device-to-device (D2D) communications, which may significantly improve spectrum efficiency and alleviate the heavy burden on backhaul links. However, most previous works ignored user mobility, thus having limited practical applications. In this paper, we take advantage of the user mobility pattern by the inter-contact times between different users, and propose a mobility-aware caching placement strategy to maximize thedata offloading ratio, which is defined as the percentage of the requested data that can be delivered via D2D links rather than through base stations. Given the NP-hard caching placement problem, we first propose an optimal dynamic programming algorithm to obtain a performance benchmark with much lower complexity than exhaustive search. We then prove that the problem falls in the category of monotone submodular maximization over a matroid constraint, and propose a time-efficient greedy algorithm, which achieves an approximation ratio as$\frac {1}{2}$. Simulation results with real-life data sets will validate the effectiveness of our proposed mobility-aware caching placement strategy. We observe that users moving at either a very low or very high speed should cache the most popular files, while users moving at a medium speed should cache less popular files to avoid duplication.
Rui Wang 0028, Jun Zhang 0004, Shenghui Song 0001, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.2
2016 Power-Delay Tradeoff in Multi-User Mobile-Edge Computing Systems
abstract
Mobile-edge computing (MEC) has recently emerged as a promising paradigm to liberate mobile devices from increasingly intensive computation workloads, as well as to improve the quality of computation experience. In this paper, we investigate the tradeoff between two critical but conflicting objectives in multi-user MEC systems, namely, the power consumption of mobile devices and the execution delay of computation tasks. A power consumption minimization problem with task buffer stability constraints is formulated to investigate the tradeoff, and an online algorithm that decides the local execution and computation offloading policy is developed based on Lyapunov optimization. Specifically, at each time slot, the optimal frequencies of the local CPUs are obtained in closed forms, while the optimal transmit power and bandwidth allocation for computation offloading are determined with the Gauss-Seidel method. Performance analysis is conducted for the proposed algorithm, which indicates that the power consumption and execution delay obeys an [0(1/V), 0(V)] tradeoff with V as a control parameter. Simulation results are provided to validate the theoretical analysis and demonstrate the impacts of various parameters to the system performance.
Yuyi Mao, Jun Zhang 0004, Shenghui Song 0001, Khaled Ben Letaief
GLOBECOM2
2016 Joint Subcarrier and CPU Time Allocation for Mobile Edge Computing
abstract
In mobile edge computing systems, mobile devices can offload compute-intensive tasks to a nearby cloudlet, so as to save energy and extend battery life. Unlike a fully-fledged cloud, a cloudlet is a small-scale datacenter deployed at a wireless access point, and thus is highly constrained by both radio and compute resources. We show in this paper that separately optimizing the allocation of either compute or radio resource - as most existing works did - is highly suboptimal: the congestion of compute resource leads to the waste of radio resource, and vice versa. To address this problem, we propose a joint scheduling algorithm that allocates both radio and compute resources coordinately. Specifically, we consider a cloudlet in an Orthogonal Frequency-Division Multiplexing Access (OFDMA) system with multiple mobile devices, where we study subcarrier allocation for task offloading and CPU time allocation for task execution in the cloudlet. Simulation results show that the proposed algorithm significantly outperforms per-resource optimization, accommodating more offloading requests while achieving salient energy saving.
Yinghao Yu, Jun Zhang 0004, Khaled Ben Letaief
GLOBECOM2
2016 Selective uplink training for massive MIMO systems
abstract
As a promising technique to meet the drastically growing demand for both high throughput and uniform coverage in the fifth generation (5G) wireless networks, massive multiple-input multiple-output (MIMO) systems have attracted significant attention in recent years. However, in massive MIMO systems, as the density of mobile users (MUs) increases, conventional uplink training methods will incur prohibitively high training overhead, which is proportional to the number of MUs. In this paper, we propose a selective uplink training method for massive MIMO systems, where in each channel block only part of the MUs will send uplink pilots for channel training, and the channel states of the remaining MUs are predicted from the estimates in previous blocks, taking advantage of the channels' temporal correlation. We propose an efficient algorithm to dynamically select the MUs to be trained within each block and determine the optimal uplink training length. Simulation results show that the proposed training method provides significant throughput gains compared to the existing methods, while much lower estimation complexity is achieved. It is observed that the throughput gain becomes higher as the MU density increases.
Jun Zhang 0004, Shenghui Song 0001, Khaled Ben Letaief
ICC2
2016 Content caching at the wireless network edge: A distributed algorithm via belief propagation
abstract
Caching popular contents at the edge of wireless networks has recently emerged as a promising technology to improve the quality of service for mobile users, while balancing the peak-to-average transmissions over backhaul links. In contrast to existing works, where a central coordinator is required to design the cache placement strategy, we consider a distributed caching problem which is highly relevant in dense network settings. In the considered scenario, each Base Station (BS) has a cache storage of finite capacity, and each user will be served by one or multiple BSs depending on the employed transmission scheme. A belief propagation based distributed algorithm is proposed to solve the cache placement problem, where the parallel computations are performed by individual BSs based on limited local information and very few messages passed between neighboring BSs. Thus, no central coordinator is required to collect the information of the whole network, which significantly saves signaling overhead. Simulation results show that the proposed low-complexity distributed algorithm can greatly reduce the average download delay by collaborative caching and transmissions.
Juan Liu 0002, Bo Bai 0001, Jun Zhang 0004, Khaled Ben Letaief
ICC3
2016 Cache size allocation in backhaul limited wireless networks
abstract
Caching popular content at base stations is a powerful supplement to existing limited backhaul links for accommodating the exponentially increasing mobile data traffic. Given the limited cache budget, we investigate the cache size allocation problem in cellular networks to maximize the user success probability (USP), taking wireless channel statistics, backhaul capacities and file popularity distributions into consideration. The USP is defined as the probability that one user can successfully download its requested file either from the local cache or via the backhaul link. We first consider a single-cell scenario and derive a closed-form expression for the USP, which helps reveal the impacts of various parameters, such as the file popularity distribution. More specifically, for a highly concentrated file popularity distribution, the required cache size is independent of the total number of files, while for a less concentrated file popularity distribution, the required cache size is in linear relation to the total number of files. Furthermore, we study the multi-cell scenario, and provide a bisection search algorithm to find the optimal cache size allocation. The optimal cache size allocation is verified by simulations, and it is shown to play a more significant role when the file popularity distribution is less concentrated.
Xi Peng 0006, Jun Zhang 0004, Shenghui Song 0001, Khaled Ben Letaief
ICC2
2016 QoS-aware joint mode selection and channel assignment for D2D communications
abstract
Underlaying device-to-device (D2D) communications to a cellular network is considered as a key technique to improve spectral efficiency in 5G networks. For such D2D systems, mode selection and resource allocation have been widely utilized for managing interference. However, previous works allowed at most one D2D link to access the same channel, while mode selection and resource allocation are typically separately designed. In this paper, we jointly optimize the mode selection and channel assignment in a cellular network with underlaying D2D communications, where multiple D2D links may share the same channel. Meanwhile, the QoS requirements for both cellular and D2D links are guaranteed, in terms of Signal-to-Interference-plus-Noise Ratio (SINR). We first propose an optimal dynamic programming (DP) algorithm, which provides a much lower computation complexity compared to exhaustive search and serves as the performance bench mark. A bipartite graph based greedy algorithm is then proposed to achieve a polynomial time complexity. Simulation results will demonstrate the advantage of allowing each channel to be accessed by multiple D2D links in dense D2D networks, as well as, the effectiveness of the proposed algorithms.
Rui Wang 0028, Jun Zhang 0004, Shenghui Song 0001, Khaled Ben Letaief
ICC2
2016 Delay-optimal computation task scheduling for mobile-edge computing systems
abstract
Mobile-edge computing (MEC) emerges as a promising paradigm to improve the quality of computation experience for mobile devices. Nevertheless, the design of computation task scheduling policies for MEC systems inevitably encounters a challenging two-timescale stochastic optimization problem. Specifically, in the larger timescale, whether to execute a task locally at the mobile device or to offload a task to the MEC server for cloud computing should be decided, while in the smaller timescale, the transmission policy for the task input data should adapt to the channel side information. In this paper, we adopt a Markov decision process approach to handle this problem, where the computation tasks are scheduled based on the queueing state of the task buffer, the execution state of the local processing unit, as well as the state of the transmission unit. By analyzing the average delay of each task and the average power consumption at the mobile device, we formulate a power-constrained delay minimization problem, and propose an efficient one-dimensional search algorithm to find the optimal task scheduling policy. Simulation results are provided to demonstrate the capability of the proposed optimal stochastic task scheduling policy in achieving a shorter average execution delay compared to the baseline policies.
Juan Liu 0002, Yuyi Mao, Jun Zhang 0004, Khaled Ben Letaief
ISIT3
2016 Statistical group sparse beamforming for green Cloud-RAN via large system analysis
abstract
In this paper, we develop a statistical group sparse beamforming framework to minimize the network power consumption for green cloud radio access networks (Cloud-RANs). It will promote group sparsity structures in the beamforming vectors, which will provide a good indicator for remote radio head (RRH) ordering to enable adaptive RRH selection for power saving. In contrast to the previous works that depend heavily on instantaneous channel state information (CSI), the proposed algorithm only depends on the long-term channel state attenuation for RRH ordering, which does not require frequent update, thereby significantly reducing the computation overhead. This is achieved by developing a smoothed ℓp-minimization approach to induce group sparsity in beamforming vectors, followed by an iterative reweighted-ℓ2algorithm via the principles of the majorization-minimization (MM) algorithm and the Lagrangian duality theory. With the well-structured closed-form solutions at each iteration, we further leverage the large-dimensional random matrix theory to derive deterministic approximations for the squared ℓ2-norm of the induced group sparse beamforming vectors in the large system regimes. The deterministic approximation results only depend on statistical CSI and will guide the RRH ordering. Simulation results demonstrate the near-optimal performance of the proposed algorithm, even in finite systems.
Yuanming Shi, Jun Zhang 0004, Khaled Ben Letaief
ISIT2
2016 CoDAR: Revealing the Generalized Procedure & Recommending Algorithms of Community Detection
abstract
Community detection has attracted great interest in graph analysis and mining during the past decade, and a great number of approaches have been developed to address this problem. However, the lack of a uniform framework and a reasonable evaluation method makes it a puzzle to analyze, compare and evaluate the extensive work, let alone picking out a best one when necessary. In this paper, we design a tool called CoDAR, which reveals the generalized procedure of community detection and monitors the real-time structural changes of network during the detection process. Moreover, CoDAR adopts 12 recognized metrics and builds a rating model for performance evaluation of communities to recom- mend the best-performing algorithm. Finally, the tool also provides nice interactive windows for display.
Xiang Ying, Chaokun Wang, Jeffrey Xu Yu, Jun Zhang 0004
SIGMOD Conference5
2016 ARQ with adaptive feedback for energy harvesting receivers
abstract
Automatic repeat request (ARQ) is widely used in modern communication systems to improve transmission reliability. In conventional ARQ protocols developed for systems with energy-unconstrained receivers, an acknowledgement/negative-acknowledgement (ACK/NACK) message is fed back when decoding succeeds/fails. Such kind of non-adaptive feedback consumes significant amount of energy, and thus will limit the performance of systems with energy harvesting (EH) receivers. In order to overcome this limitation and to utilize the harvested energy more efficiently, we propose a novel ARQ protocol for EH receivers, where the ACK feedback can be adapted based upon the receiver's EH state. Two conventional ARQ protocols are also considered. By adopting the packet drop probability (PDP) as the performance metric, we formulate the throughput constrained PDP minimization problem for a communication link with a non-EH transmitter and an EH receiver. Optimal reception policies including the sampling, decoding and feedback strategies, are developed for different ARQ protocols. Simulation results will show that the proposed ARQ protocol not only outperforms the conventional ARQs in terms of PDP, but can also achieve a higher throughput.
Yuyi Mao, Jun Zhang 0004, Khaled Ben Letaief
WCNC2
2016 Coverage analysis for dense millimeter wave cellular networks: The impact of array size
abstract
Millimeter wave (mmWave) communications has been considered as a promising technology for 5G cellular networks. Exploiting directional beamforming using antenna arrays to combat path loss is one of the defining features in mmWave cellular networks. However, previous works on mmWave network analysis usually adopt simplified antenna patterns for tractability. In this paper, we show that there are huge discrepancies between the simplified and actual antenna patterns when investigating the coverage probability of mmWave networks. Analytical expressions for the coverage probabilities are derived using tools from stochastic geometry, by considering the actual antenna pattern with the uniform linear array. Moreover, the impact of the array size is investigated, which cannot be revealed from existing results with simplified antenna patterns. Numerical results will show that large-scale antenna arrays are required for satisfactory coverage in mmWave cellular networks. Furthermore, dense mmWave cellular networks are shown to achieve much higher rate coverage than conventional sub-6 GHz cellular systems.
Xianghao Yu, Jun Zhang 0004, Khaled Ben Letaief
WCNC2
2016 Dynamic Computation Offloading for Mobile-Edge Computing With Energy Harvesting Devices
abstract
Mobile-edge computing (MEC) is an emerging paradigm to meet the ever-increasing computation demands from mobile applications. By offloading the computationally intensive workloads to the MEC server, the quality of computation experience, e.g., the execution latency, could be greatly improved. Nevertheless, as the on-device battery capacities are limited, computation would be interrupted when the battery energy runs out. To provide satisfactory computation performance as well as achieving green computing, it is of significant importance to seek renewable energy sources to power mobile devices via energy harvesting (EH) technologies. In this paper, we will investigate a green MEC system with EH devices and develop an effective computation offloading strategy. The execution cost, which addresses both the execution latency and task failure, is adopted as the performance metric. A low-complexity online algorithm is proposed, namely, the Lyapunov optimization-based dynamic computation offloading algorithm, which jointly decides the offloading decision, the CPU-cycle frequencies for mobile execution, and the transmit power for computation offloading. A unique advantage of this algorithm is that the decisions depend only on the current system state without requiring distribution information of the computation task request, wireless channel, and EH processes. The implementation of the algorithm only requires to solve a deterministic problem in each time slot, for which the optimal solution can be obtained either in closed form or by bisection search. Moreover, the proposed algorithm is shown to be asymptotically optimal via rigorous analysis. Sample simulation results shall be presented to corroborate the theoretical analysis as well as validate the effectiveness of the proposed algorithm.
Yuyi Mao, Jun Zhang 0004, Khaled Ben Letaief
IEEE J. Sel. Areas Commun.2
2016 Smoothed Lp-Minimization for Green Cloud-RAN With User Admission Control
abstract
The cloud radio access network (Cloud-RAN) has recently been proposed as one of the cost-effective and energy-efficient techniques for 5G wireless networks. By moving the signal processing functionality to a single baseband unit (BBU) pool, centralized signal processing and resource allocation are enabled in cloud-RAN, thereby providing the promise of improving the energy efficiency via effective network adaptation and interference management. In this paper, we propose a holistic sparse optimization framework to design green cloud-RAN by taking into consideration the power consumption of the fronthaul links, multicast services, as well as user admission control. Specifically, we first identify the sparsity structures in the solutions of both the network power minimization and user admission control problems, which call for adaptive remote radio head (RRH) selection and user admission. However, finding the optimal sparsity structures turns out to be NP-hard, with the coupled challenges of the ℓ0-norm-based objective functions and the nonconvex quadratic QoS constraints due to multicast beamforming. In contrast to the previous works on convex but nonsmooth sparsity inducing approaches, e.g., the group sparse beamforming algorithm based on the mixed ℓ1/ℓ2-norm relaxation, we adopt the nonconvex but smoothed ℓp-minimization (02algorithm is developed, which will converge to a Karush-Kuhn-Tucker (KKT) point of the relaxed smoothed ℓp-minimization problem from the SDR technique. We illustrate the effectiveness of the proposed algorithms with extensive simulations for network power minimization and user admission control in multicast cloud-RAN.
Yuanming Shi, Jinkun Cheng, Jun Zhang 0004, Bo Bai 0001, Wei Chen 0002, Khaled Ben Letaief
IEEE J. Sel. Areas Commun.3
2016 Success Probability and Area Spectral Efficiency in Multiuser MIMO HetNets
abstract
We derive a general and closed-form result for the success probability in downlink multiple-antenna (MIMO) heterogeneous cellular networks (HetNets), utilizing a novel Toeplitz matrix representation. This main result, which is equivalently the signal-to-interference ratio (SIR) distribution, includes multiuser MIMO, single-user MIMO and per-tier biasing for K different tiers of randomly placed base stations (BSs), assuming zero-forcing precoding and perfect channel state information. The large SIR limit of this result admits a simple closed form that is accurate at moderate SIRs, e.g., above 5 dB. These results reveal that the SIR-invariance property of SISO HetNets does not hold for MIMO HetNets; instead the success probability may decrease as the network density increases. We prove that the maximum success probability is achieved by activating only one tier of BSs, while the maximum area spectral efficiency (ASE) is achieved by activating all the BSs. This reveals a unique tradeoff between the ASE and link reliability in multiuser MIMO HetNets. To achieve the maximum ASE while guaranteeing a certain link reliability, we develop efficient algorithms to find the optimal BS densities. It is shown that as the link reliability requirement increases, more BSs and more tiers should be deactivated.
Chang Li 0002, Jun Zhang 0004, Jeffrey G. Andrews, Khaled Ben Letaief
IEEE Trans. Commun.2
2016 Transmit Power Minimization for Wireless Networks With Energy Harvesting Relays
abstract
Energy harvesting (EH) has recently emerged as a key technology for green communications as it can power wireless networks with renewable energy sources. However, directly replacing the conventional non-EH transmitters by EH nodes will be a challenge. In this paper, we propose to deploy extra EH nodes as relays over an existing non-EH network. Specifically, the considered non-EH network consists of multiple source-destination (S-D) pairs. The deployed EH relays will take turns to assist each S-D pair, and energy diversity can be achieved to combat the low-EH rate of each EH relay. To make the best of these EH relays, with the source transmit power minimization as the design objective, we formulate a joint power assignment and relay selection problem, which, however, is NP-hard. We thus propose a general framework to develop efficient suboptimal algorithms, which is mainly based on a sufficient condition for the feasibility of the optimization problem. This condition yields useful design insights and also reveals an energy hardening effect, which provides the possibility to exempt the requirement of noncausal EH information. Simulation results will show that the proposed co-operation strategy can achieve near-optimal performance and provide significant power savings. Compared to the greedy co-operation method that only optimizes the performance of the current transmission block, the proposed strategy can achieve the same performance with much fewer relays, and the performance gap increases with the number of S-D pairs.
Yaming Luo, Jun Zhang 0004, Khaled Ben Letaief
IEEE Trans. Commun.2
2016 Inferring Directions of Undirected Social Ties
abstract
The directionality is a significant but inherent property of social ties, though usually ignored in undirected social networks due to its invisibility. However, we believe most social ties are natively directed, and the perception of directionality can improve our understanding about the network structures and further benefit other tasks upon social networks. In this study, we address the latent tie direction inference problem in undirected social networks. We engage in the investigation of directionality on real-world large-scale directed social networks and summarize our findings using four patterns. Upon that we propose a family of ReDirect approaches, including ReDirect-N, ReDirect-T and ReDirect-One, to inferring the hidden directions of undirected social ties based on the network topology only. ReDirect can incorporate with other predictive tasks, and introduce supervision to improve performance. We also present a simple but effective strategy to construct self-labeled data. Experimental results show that even without external information, our approach can recover the directions of networks effectively. Moreover, we find the ReDirect approaches can benefit the predictive tasks remarkably in an experimental study on link prediction. The ReDirect family can be a beneficial general data preprocess tool for various network analysis tasks by uncovering the hidden directions.
Jun Zhang 0004, Chaokun Wang, Jianmin Wang 0001, Jeffrey Xu Yu, Jun Chen 0004, Changping Wang
IEEE Trans. Knowl. Data Eng.1
2016 Grid Energy Consumption and QoS Tradeoff in Hybrid Energy Supply Wireless Networks
abstract
Hybrid energy supply (HES) wireless networks have recently emerged as a new paradigm to enable green networks, which are powered by both the electric grid and harvested renewable energy. In this paper, we will investigate two critical but conflicting design objectives of HES networks, i.e., the grid energy consumption and quality of service (QoS). Minimizing grid energy consumption by utilizing the harvested energy will make the network environmentally friendly, but the achievable QoS may be degraded due to the intermittent nature of energy harvesting. To investigate the tradeoff between these two aspects, we introduce the total service cost as the performance metric, which is the weighted sum of the grid energy cost and the QoS degradation cost. Base station assignment and power control is adopted as the main strategy to minimize the total service cost, while both cases with non-causal and causal side information are considered. With non-causal side information, a Greedy Assignment algorithm with low complexity and near-optimal performance is proposed. With causal side information, the design problem is formulated as a discrete Markov decision problem. Interesting solution structures are derived, which shall help to develop an efficient monotone backward induction algorithm. To further reduce complexity, a Look-Ahead policy and a Threshold-based Heuristic policy are also proposed. Simulation results shall validate the effectiveness of the proposed algorithms and demonstrate the unique grid energy consumption and QoS tradeoff in HES networks.
Yuyi Mao, Jun Zhang 0004, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.2
2016 Compressed CSI Acquisition in FDD Massive MIMO: How Much Training is Needed?
abstract
Massive multiple-input-multiple-output (MIMO) is a promising technique for providing unprecedented spectral efficiency. However, it has been well recognized that the excessive training overhead required for obtaining the channel side information is a major handicap in frequency-division duplexing (FDD) massive MIMO. Several attempts have been made to reduce this training overhead by exploiting the sparsity structures of massive MIMO channels. So far, however, there has been little discussion about how to exploit the partial support information of these channels to achieve further overhead reductions. Such information, which is a set of indices of the significant elements of a channel vector, can be acquired in advance and hence is an important option to explore. In this paper, we examine the impact on the required training overhead when this information is applied within a weighted ℓ1minimization framework, and analytically show that a sharp estimate of the reduced overhead size can be successfully obtained. Furthermore, we examine how the accuracy of the partial support information impacts the achievable overhead reduction. Numerical results for a wide range of sparsity and partial support information reliability levels are presented to quantify our findings and main conclusions.
Juei-Chin Shen, Jun Zhang 0004, Emad Alsusa, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.2
2016 Low-Rank Matrix Completion for Topological Interference Management by Riemannian Pursuit
abstract
In this paper, we present a flexible low-rank matrix completion (LRMC) approach for topological interference management (TIM) in the partially connected $K$-user interference channel. No channel state information (CSI) is required at the transmitters except the network topology information. The previous attempt on the TIM problem is mainly based on its equivalence to the index coding problem, but so far only a few index coding problems have been solved. In contrast, in this paper, we present an algorithmic approach to investigate the achievable degrees-of-freedom (DoFs) by recasting the TIM problem as an LRMC problem. Unfortunately, the resulting LRMC problem is known to be NP-hard, and the main contribution of this paper is to propose a Riemannian pursuit (RP) framework to detect the rank of the matrix to be recovered by iteratively increasing the rank. This algorithm solves a sequence of fixed-rank matrix completion problems. To address the convergence issues in the existing fixed-rank optimization methods, the quotient manifold geometry of the search space of fixed-rank matrices is exploited via Riemannian optimization. By further exploiting the structure of the low-rank matrix varieties, i.e., the closure of the set of fixed-rank matrices, we develop an efficient rank increasing strategy to find good initial points in the procedure of rank pursuit. Simulation results demonstrate that the proposed RP algorithm achieves a faster convergence rate and higher achievable DoFs for the TIM problem compared with the state-of-the-art methods.
Yuanming Shi, Jun Zhang 0004, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.2
2016 Optimal QoS-Aware Channel Assignment in D2D Communications With Partial CSI
abstract
In this paper, we propose effective channel assignment algorithms for network utility maximization in a cellular network with underlaying device-to-device (D2D) communications. A major innovation is the consideration of partial channel state information (CSI), i.e., the base station (BS) is assumed to be able to acquire “partial” instantaneous CSI of the cellular and D2D links, as well as, the interference links. In contrast to the existing works, multiple D2D links are allowed to share the same channel, and the quality of service (QoS) requirements for both the cellular and D2D links are enforced. We first develop an optimal channel assignment algorithm based on dynamic programming, which enjoys a much lower complexity compared with exhaustive search and will serve as a performance benchmark. To further reduce complexity, we propose a cluster-based sub-optimal channel assignment algorithm. New closed-form expressions for the expected weighted sum rate and the successful transmission probabilities are also derived. Simulation results verify the effectiveness of the proposed algorithms. Moreover, by comparing different partial CSI scenarios, we observe that the CSI of the D2D communication links and the interference links from the D2D transmitters to the BS significantly affects the network performance, while the CSI of the interference links from the BS to the D2D receivers only has a negligible impact.
Rui Wang 0028, Jun Zhang 0004, Shenghui Song 0001, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.2
2015 Backhaul-Aware Caching Placement for Wireless Networks
abstract
As the capacity demand of mobile applications keeps increasing, the backhaul network is becoming a bottleneck to support high quality of experience (QoE) in next-generation wireless networks. Content caching at base stations (BSs) is a promising approach to alleviate the backhaul burden and reduce user-perceived latency. In this paper, we consider a wireless caching network where all the BSs are connected to a central controller via backhaul links. In such a network, users can obtain the required data from candidate BSs if the data are pre-cached. Otherwise, the user data need to be first retrieved from the central controller to local BSs, which introduces extra delay over the backhaul. In order to reduce the download delay, the caching placement strategy needs to be optimized. We formulate such a design problem as the minimization of the average download delay over user requests, subject to the caching capacity constraint of each BS. Different from existing works, our model takes BS cooperation in the radio access into consideration and is fully aware of the propagation delay on the backhaul links. The design problem is a mixed integer programming problem and is highly complicated, and thus we relax the problem and propose a low-complexity algorithm. Simulation results will show that the proposed algorithm can effectively determine the near-optimal caching placement and provide significant performance gains over conventional caching placement strategies.
Xi Peng 0006, Juei-Chin Shen, Jun Zhang 0004, Khaled Ben Letaief
GLOBECOM3
2015 QoS-Aware Channel Assignment for Weighted Sum-Rate Maximization in D2D Communications
abstract
Underlaying device-to-device (D2D) communication links to a cellular network is a promising way to improve spectrum efficiency, for which the cross- link interference should be carefully controlled. Resource allocation has been widely utilized for managing interference in D2D networks. However, most previous works made simple assumptions by either ignoring the reliability requirement of D2D links or not allowing multiple D2D links to share the same channel. In this paper, we propose effective channel assignment algorithms to maximize the weighted sum-rate in a cellular network with underlaying D2D communications, where multiple D2D links are allowed to share the same channel. Meanwhile, the minimum Signal-to- Interference-plus-Noise Ratio (SINR) requirements for both cellular and D2D links are guaranteed. We first provide an optimal algorithm based on dynamic programming (DP) to serve as the performance benchmark, which enjoys much lower complexity compared to exhaustive search. To further reduce complexity, we then propose a cluster-based near-optimal channel assignment algorithm. Simulation results will demonstrate the advantage of allowing multiple D2D links to share the same channel in dense D2D networks, as well as verifying the effectiveness of the proposed algorithms.
Rui Wang 0028, Jun Zhang 0004, Shenghui Song 0001, Khaled Ben Letaief
GLOBECOM2
2015 Hybrid Precoding Design in Millimeter Wave MIMO Systems: An Alternating Minimization Approach
abstract
Millimeter wave (mmWave) communications holds a promise to offer an unprecedented capacity boost for 5G cellular networks. Due to the small wavelength of mmWave signals, mmWave multiple-input-multiple-output (MIMO) systems can leverage large-scale antennas to combat the path loss and rain attenuation via precoding. Different from conventional MIMO systems, mmWave MIMO cannot realize precoding entirely at baseband using digital precoders, as a result of potentially high power consumed by signal mixers and analog-to-digital converters (ADCs). As a cost- effective alternative, a hybrid precoding transceiver architecture for mmWave MIMO systems has received considerable attention. However, the optimal design of such hybrid precoding has not been fully understood. In this paper, an alternating minimization algorithm based on manifold optimization is proposed to design the hybrid precoder, thereby making it comparable in performance to the digital precoder. Numerical results show that our proposed algorithm can significantly outperform existing ones in terms of spectral efficiency and, more importantly, it can achieve the optimal performance in certain cases. Finally, the alternating minimization approach is shown to be generally applicable to precoding design with different hybrid structures, and the corresponding comparison will show interesting design insights for hybrid precoding.
Xianghao Yu, Juei-Chin Shen, Jun Zhang 0004, Khaled Ben Letaief
GLOBECOM3
2015 Group sparse beamforming for multicast green Cloud-RAN via parallel semidefinite programming
abstract
The Cloud radio access network (Cloud-RAN) has great potentials to improve energy efficiency and increase capacity of wireless networks. In this paper, we investigate multicast beamforming design for network power minimization of Cloud-RAN, which is shown to be a highly intractable non-convex mixed integer non-linear programming problem. To provide an efficient solution to this highly complicated problem, we propose a three-stage algorithm based on the group-sparsity inducing norm, which minimizes network power by coordinated multicast beamforming and adaptively selecting active remote radio heads (RRHs). In particular, a novel quadratic variational weighted ℓ1=ℓ2-norm aided alternating algorithm is proposed to exploit the group-sparsity structure of the beamforming vector, thereby guiding the active RRH set selection. Given the selected RRH set, multicast beamforming is performed to minimize the network power consumption. Furthermore, to enhance the computation efficiency upon utilizing the shared computing resources in the cloud center, we employ the alternating direction method of multipliers (ADMM) algorithm to solve the resulting semidefinite programming problems in parallel. Extensive simulation results will demonstrate the effectiveness of the proposed multicast group sparse beamforming algorithm.
Jinkun Cheng, Yuanming Shi, Bo Bai 0001, Wei Chen 0002, Jun Zhang 0004, Khaled Ben Letaief
ICC5
2015 Analysis of area spectral efficiency and link reliability in multiuser MIMO HetNets
abstract
Heterogeneous networks (HetNets) provide an effective way to meet the explosive growth of mobile data traffic. Previous studies have revealed that the successful transmission probability, i.e., the link reliability, of a SISO HetNet is invariant to the base station (BS) density. This indicates that the area spectral efficiency (ASE) can be increased by densifying the network, without sacrificing the link performance. However, in this paper, we shall show that the above invariance property no longer holds in multi-antenna HetNets. More specifically, changing the BS density will affect both the link reliability and the ASE, and there exists a tradeoff between these two important performance metrics. By adopting the Poisson point process to model the BS positions, we develop an exact expression of the successful transmission probability for general multiuser MIMO HetNets. We then use this result to evaluate the tradeoff between the ASE and the link reliability. It is analytically shown that the maximum successful transmission probability of the network is achieved by activating only the tier of BSs with the largest number of antennas per BS, while the maximum ASE is achieved by activating all the BSs. By adjusting the density of each tier of the HetNet, a tradeoff between the link reliability and ASE can be achieved.
Chang Li 0002, Jun Zhang 0004, Shenghui Song 0001, Khaled Ben Letaief
ICC2
2015 Compressed CSI acquisition in FDD massive MIMO with partial support information
abstract
Massive MIMO is a promising technique to provide unprecedented spectral efficiency. However, it has been well recognized that huge training overhead for obtaining channel side information (CSI) is a major handicap in frequency-division duplexing (FDD) massive MIMO. Several attempts have been made to reduce this training overhead by exploiting sparse structures of massive MIMO channels. So far, however, there has been little discussion about how to utilize partial support information of sparse channels to achieve further overhead reduction. This support information, which is a set of indexes of significant elements of a channel vector, actually can be acquired in advance. In this paper, we examine the required training overhead when partial support information is applied within a weighted ℓ1minimization framework and analytically show that a sharp estimate of this overhead size can be successfully obtained. Furthermore, we demonstrate that the accuracy of partial support information plays an important role in determining how much reduction can be achieved. Numerical results shall verify the main conclusions.
Juei-Chin Shen, Jun Zhang 0004, Emad Alsusa, Khaled Ben Letaief
ICC2
2015 Low-rank matrix completion via Riemannian pursuit for topological interference management
abstract
This paper considers the topological interference management problem in a partially connected K-user interference channel, where no channel state information at transmitters (CSIT) is available beyond the network topology knowledge. Due to the practical CSI assumption, this problem has recently received enough attention. In particular, it has been established that the topological interference management problem, in terms of degrees of freedom (DoF), is equivalent to the index coding problem with linear schemes. However, so far only a few index coding problems have been solved, and thus there is a lack of a systematic way to characterize optimal DoF of an arbitrary network topology. In this paper, we present a low-rank matrix completion (LRMC) approach to find linear solutions to maximize the achievable symmetric DoF for any given network topology. To decode the desired messages at each receiver, we also propose an LRMC based channel acquisition scheme, which can obtain interference-free measurements of the desired channel at each receiver while minimizing the pilot training length. To address the NP-hardness of the non-convex rank objective function in the resulting LRMC problem, we further present a Riemannian pursuit (RP) algorithm to solve it efficiently. This algorithm alternatively performs fixed-rank optimization using Riemannian optimization and rank increase by exploiting the manifold structure of the fixed-rank matrices. The LRMC approach aided by the RP algorithms not only recovers the existing optimal DoF results but also provides insights for general network topologies.
Yuanming Shi, Jun Zhang 0004, Khaled Ben Letaief
ISIT2
2015 Joint base station assignment and power control in hybrid energy supply wireless networks
abstract
This paper addresses the joint base station (BS) assignment and power control problem in a hybrid energy supply wireless network, where an energy harvesting BS and a grid-powered BS coordinate to serve a mobile user. In order to minimize the grid energy consumption while maximizing the number of transmitted data packets, we introduce the total service cost over an N-block frame as the performance metric, which is the weighted sum of the grid energy cost and the packet drop cost. With non-causal side information (SI) available at the BSs, including energy SI and channel SI, a Greedy Assignment algorithm with low complexity and near optimal performance is proposed. For the causal SI setting, the design problem is formulated as a discrete Markov decision problem. Interesting solution structures are derived, which help develop an efficient monotone backward induction algorithm. To further reduce the complexity, a heuristic online policy is also proposed. Simulation results shall validate the effectiveness of the proposed policies and demonstrate a unique tradeoff in such networks, i.e., the tradeoff between the grid energy consumption and the provided quality of service.
Yuyi Mao, Jun Zhang 0004, Khaled Ben Letaief
WCNC2
2015 A Lyapunov Optimization Approach for Green Cellular Networks With Hybrid Energy Supplies
abstract
Powering cellular networks with renewable energy sources via energy harvesting (EH) have recently been proposed as a promising solution for green networking. However, with intermittent and random energy arrivals, it is challenging to provide satisfactory quality of service (QoS) in EH networks. To enjoy the greenness brought by EH while overcoming the instability of the renewable energy sources, hybrid energy supply (HES) networks that are powered by both EH and the electric grid have emerged as a new paradigm for green communications. In this paper, we will propose new design methodologies for HES green cellular networks with the help of Lyapunov optimization techniques. The network service cost, which addresses both the grid energy consumption and achievable QoS, is adopted as the performance metric, and it is optimized via base station assignment and power control (BAPC). Our main contribution is a low-complexity online algorithm to minimize the long-term average network service cost, namely, the Lyapunov optimization-based BAPC (LBAPC) algorithm. One main advantage of this algorithm is that the decisions depend only on the instantaneous side information without requiring distribution information of channels and EH processes. To determine the network operation, we only need to solve a deterministic per-time slot problem, for which an efficient inner-outer optimization algorithm is proposed. Moreover, the proposed algorithm is shown to be asymptotically optimal via rigorous analysis. Finally, sample simulation results are presented to verify the theoretical analysis as well as validate the effectiveness of the proposed algorithm.
Yuyi Mao, Jun Zhang 0004, Khaled Ben Letaief
IEEE J. Sel. Areas Commun.2
2015 Community Detection in Social Networks: An In-depth Benchmarking Study with a Procedure-Oriented Framework
abstract
Revealing the latent community structure, which is crucial to understanding the features of networks, is an important problem in network and graph analysis. During the last decade, many approaches have been proposed to solve this challenging problem in diverse ways, i.e. different measures or data structures. Unfortunately, experimental reports on existing techniques fell short in validity and integrity since many comparisons were not based on a unified code base or merely discussed in theory. We engage in an in-depth benchmarking study of community detection in social networks. We formulate a generalized community detection procedure and propose a procedure-oriented framework for benchmarking. This framework enables us to evaluate and compare various approaches to community detection systematically and thoroughly under identical experimental conditions. Upon that we can analyze and diagnose the inherent defect of existing approaches deeply, and further make effective improvements correspondingly. We have re-implemented ten state-of-the-art representative algorithms upon this framework and make comprehensive evaluations of multiple aspects, including the efficiency evaluation, performance evaluations, sensitivity evaluations, etc. We discuss their merits and faults in depth, and draw a set of take-away interesting conclusions. In addition, we present how we can make diagnoses for these algorithms resulting in significant improvements.
Chaokun Wang, Jeffrey Xu Yu, Jun Zhang 0004
Proc. VLDB Endow.4
2015 User-Centric Intercell Interference Nulling for Downlink Small Cell Networks
abstract
Small cell networks are regarded as a promising candidate to meet the exponential growth of mobile data traffic in cellular networks. With a dense deployment of access points, spatial reuse will be improved, and uniform coverage can be provided. However, such performance gains cannot be achieved without effective intercell interference management. In this paper, a novel interference coordination strategy, called user-centric intercell interference nulling, is proposed for small cell networks. A main merit of the proposed strategy is its ability to effectively identify and mitigate the dominant interference for each user. Different from existing works, each user selects the coordinating base stations (BSs) based on the relative distance between the home BS and the interfering BSs, called the interference nulling (IN) range, and thus interference nulling adapts to each user's own interference situation. By adopting a random spatial network model, we derive an approximate expression of the successful transmission probability to the typical user, which is then used to determine the optimal IN range. Simulation results shall confirm the tightness of the approximation, and demonstrate significant performance gains (about 35-40%) of the proposed coordination strategy, compared with the non-coordination case. Moreover, it is shown that the proposed strategy outperforms other interference nulling methods. Finally, the effect of imperfect channel state information (CSI) is investigated, where CSI is assumed to be obtained via limited feedback. It is shown that the proposed coordination strategy still provides significant performance gains even with a moderate number of feedback bits.
Chang Li 0002, Jun Zhang 0004, Martin Haenggi, Khaled Ben Letaief
IEEE Trans. Commun.2
2015 Downlink User Capacity of Massive MIMO Under Pilot Contamination
abstract
Pilot contamination has been regarded as a main limiting factor of time division duplexing (TDD) massive multiple-input-multiple-output (Massive MIMO) systems, as it will make the signal-to-interference-plus-noise ratio (SINR) saturated. However, how pilot contamination will limit the user capacity of downlink Massive MIMO, i.e., the maximum number of users whose SINR targets can be achieved, has not been addressed. This paper provides an explicit expression of the Massive MIMO user capacity in the pilot-contaminated regime where the number of users is larger than the pilot sequence length. This capacity expression characterizes a region within which a set of SINR requirements can be jointly satisfied. The size of this region is fundamentally limited by the pilot sequence length. Furthermore, the scheme for achieving the user capacity, i.e., the uplink pilot training sequences and downlink power allocation, has been identified. Specifically, the generalized Welch bound equality sequences are exploited and it is shown that the power allocated to each user should be proportional to its SINR target. With this capacity-achieving scheme, the SINR requirement of each user can be satisfied and energy-efficient transmission is achieved in the large-antenna-size (LAS) regime. The comparison with two non-capacity-achieving schemes highlights the superiority of our proposed scheme in terms of achieving higher user capacity. Furthermore, for the practical scenario with a finite number of antennas, the actual antenna size required to achieve a significant percentage of the asymptotic performance has been analytically quantified.
Juei-Chin Shen, Jun Zhang 0004, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.2
2014 Learning Temporal Dynamics of Behavior Propagation in Social Networks
abstract
Social influence has been widely accepted to explain people's cascade behaviors and further utilized in many related applications. However, few of existing work studied the direct, microscopic and temporal impact of social influence on people's behaviors in detail. In this paper we concentrate on the behavior modeling and systematically formulate the family of behavior propagation models (BPMs) including the static models (BP and IBP), and their discrete temporal variants (DBP and DIBP). To address the temporal dynamics of behavior propagation over continuous time, we propose a continuous temporal interest-aware behavior propagation model, called CIBP. As a new member of the BPM family, CIBP exploits the continuous-temporal functions (CTFs) to model the fully-continuous dynamic variance of social influence over time. Experiments on real-world datasets evaluated the family of BPMs and demonstrated the effectiveness of our proposed approach.
Jun Zhang 0004, Chaokun Wang, Jianmin Wang 0001
AAAI1
2014 Joint link selection and relay power allocation for energy harvesting relaying systems
abstract
Energy harvesting (EH) has recently been attracting significant attention because of its ability to scavenge environmentally friendly energy. In this paper, we investigate the use of EH relay nodes to improve the quality of service (QoS) for relaying networks. To simplify the hardware design, we adopt a half-duplex selective decode-and-forward (SDF) relay. We propose a joint link selection and relay power allocation strategy to minimize the average outage probability. Both offline and online policies, i.e., with non-causal or causal side information about the energy state and the decoding result at the relay, are investigated by utilizing deterministic and stochastic dynamic programming (DP) algorithms, respectively. Furthermore, to reduce the complexity of the optimal online solution, we propose two low-complexity suboptimal online policies. Simulation results will show that the proposed suboptimal policies outperform the existing policies and achieve near optimal performance.
Yuyi Mao, Jun Zhang 0004, Shenghui Song 0001, Khaled Ben Letaief
GLOBECOM2
2014 User capacity of pilot-contaminated TDD massive MIMO systems
abstract
Pilot contamination has been regarded as a main limiting factor of time division duplexing (TDD) massive multiple-input-multiple-output (Massive MIMO) systems, as it will make the signal-to-interference-plus-noise ratio (SINR) saturated. However, how pilot contamination will limit the user capacity of downlink Massive MIMO, i.e., the maximum number of admissible users, has not been addressed. This paper provides an explicit expression of the Massive MIMO user capacity in the pilot-contaminated regime where the number of users is larger than the pilot sequence length. Furthermore, the scheme for achieving the user capacity, i.e., the uplink pilot training sequence and downlink power allocation, has been identified. By using this capacity-achieving scheme, the SINR requirement of each user can be satisfied and energy-efficient transmission is feasible in the large-antenna-size (LAS) regime. Comparison with two non-capacity-achieving schemes highlights the superiority of our proposed scheme in terms of achieving higher user capacity.
Juei-Chin Shen, Jun Zhang 0004, Khaled Ben Letaief
GLOBECOM2
2014 Scalable coordinated beamforming for dense wireless cooperative networks
abstract
To meet the ever growing demand for both high throughput and uniform coverage in future wireless networks, dense network deployment will be ubiquitous, for which cooperation among the access points is critical. Considering the computational complexity of designing coordinated beamformers for dense networks, low-complexity and suboptimal precoding strategies are often adopted. However, it is not clear how much performance loss will be caused. To enable optimal coordinated beamforming, in this paper, we propose a framework to design a scalable beamforming algorithm based on the alternative direction method of multipliers (ADMM). Specifically, we first propose to apply the matrix stuffing technique to transform the original optimization problem to an equivalent ADMM-compliant problem, which is much more efficient than the widely-used modeling framework CVX. We will then propose to use the ADMM algorithm, a.k.a. the operator splitting method, to solve the transformed ADMM-compliant problem efficiently. In particular, the subproblems of the ADMM algorithm at each iteration can be solved with closed-forms and in parallel. Simulation results show that the proposed techniques can result in significant computational efficiency compared to the state-of-the-art interior-point solvers. Furthermore, the simulation results demonstrate that the optimal coordinated beamforming can significantly improve the system performance compared to sub-optimal zero forcing beamforming.
Yuanming Shi, Jun Zhang 0004, Khaled Ben Letaief
GLOBECOM2
2014 User-centric intercell interference coordination in small cell networks
abstract
Small cell networks provide an effective way to meet the explosive growth of mobile data traffic, which, however, complicates the network structure and makes intercell interference management more challenging. One existing interference management approach is to divide the whole network into disjoint clusters, with base stations (BSs) within each cluster doing interference coordination, but the performance will then be limited by the cluster edge users. In this paper, a novel intercell interference coordination method is proposed from the user's point of view. Each mobile user will request some neighboring BSs for interference avoidance, which is based on the relative distance between the home BS and the interfering BSs, called as the interference coordination (IC) range. In this way, the most critical interfering sources for each user can be suppressed, and thus there will be no edge user. We derive an accurate approximation for the successful transmission probability of a typical user with the proposed interference coordination method, based on which the optimal IC range can be obtained. Simulation results demonstrate a significant performance gain for the proposed method, and also show that it outperforms the existing BS clustering method.
Chang Li 0002, Jun Zhang 0004, Khaled Ben Letaief
ICC2
2014 CSI overhead reduction with stochastic beamforming for cloud radio access networks
abstract
Cloud radio access network (Cloud-RAN) is a promising network architecture to meet the explosive growth of the mobile data traffic. In this architecture, as all the baseband signal processing is shifted to a single baseband unit (BBU) pool, interference management can be efficiently achieved through coordinated beamforming, which, however, often requires full channel state information (CSI). In practice, the overhead incurred to obtain full CSI will dominate the available radio resource. In this paper, we propose a unified framework for the CSI overhead reduction and downlink coordinated beamforming. Motivated by the channel heterogeneity phenomena in large-scale wireless networks, we first propose a novel CSI acquisition scheme, called compressive CSI acquisition, which will obtain instantaneous CSI of only a subset of all the channel links and statistical CSI for the others, thus forming the mixed CSI at the BBU pool. This subset is determined by the statistical CSI. Then we propose a new stochastic beamforming framework to minimize the total transmit power while guaranteeing quality-of-service (QoS) requirements with the mixed CSI. Simulation results show that the proposed CSI acquisition scheme with stochastic beamforming can significantly reduce the CSI overhead while providing performance close to that with full CSI.
Yuanming Shi, Jun Zhang 0004, Khaled Ben Letaief
ICC2
2014 Joint data assignment and beamforming for backhaul limited caching networks
abstract
Caching at wireless access points is a promising approach to alleviate the backhaul burden in wireless networks. In this paper, we consider a cooperative wireless caching network where all the base stations (BSs) are connected to a central controller via backhaul links. In such a network, users can get the required data locally if they are cached at the BSs. Otherwise, the user data need to be assigned from the central controller to BSs via backhaul. In order to reduce the network cost, i.e., the back-haul cost and the transmit power cost, the data assignment for different BSs and the coordinated beamforming to serve different users need to be jointly designed. We formulate such a design problem as the minimization of the network cost, subject to the quality of service (QoS) constraint of each user and the transmit power constraint of each BS. This problem involves mixed-integer programming and is highly complicated. In order to provide an efficient solution, the connection between the data assignment and the sparsity-introducing norm is established. Low-complexity algorithms are then proposed to solve the joint optimization problem, which essentially decouple the data assignment and the transmit power minimization beamforming. Simulation results show that the proposed algorithms can effectively minimize the network cost and provide near optimal performance.
Xi Peng 0006, Juei-Chin Shen, Jun Zhang 0004, Khaled Ben Letaief
PIMRC3
2014 Average throughput analysis of downlink cellular networks with multi-antenna base stations
abstract
Random spatial network models have been recently utilized in the performance analysis and system design for multi-cell networks. Such an approach has been mainly adopted to investigate the outage based system performance, such as the outage probability and outage throughput. However, these performance metrics are defined with a fixed-rate transmission, and cannot characterize the performance of data traffic, which normally adopts rate adaptation. In this paper, we will evaluate the average throughput of a space division multiple access (SDMA) based cellular network by considering stochastically distributed base stations (BSs) and mobile terminals (MTs). The major difficulty for the performance analysis is the complicated distribution of the interference links. We shall provide an analytical framework for evaluating the average throughput by using the Moment Generating Function (MGF) based method. Simulations will show that the proposed method is very accurate. In particular, the analytical result can be utilized to determine the optimum number of MTs to be served in SDMA networks that can maximize the network throughput.
Rui Wang 0028, Jun Zhang 0004, Shenghui Song 0001, Khaled Ben Letaief
PIMRC2
2014 Location-aware spectrum sharing in cognitive radio networks - A semi-matching approach
abstract
Cognitive radio can improve the spectrum efficiency by allowing multiple secondary users to access the idle licensed spectrum, for which efficient and fair spectrum sharing is one of the key challenges. In this paper, we investigate the multiuser multi-channel spectrum allocation problem in cognitive radio networks with the objective of maximizing the minimum throughput among all the cognitive pairs. In the proposed approach, all the secondary users will get the opportunity to access the channel and thus a good fairness can be achieved. We first introduce a weighted bipartite graph model for this design problem. A novel semi-matching based framework is then proposed to provide an efficient suboptimal solution, which only requires statistical channel state information. Simulation results will show that by applying this approach, max-min fairness can be improved for the secondary users.
Fangyong Li, Jun Zhang 0004, Khaled Ben Letaief
WCNC2
2014 Who proposed the relationship?: recovering the hidden directions of undirected social networks
abstract
Together with the sign (positive or negative) and strength (strong or weak), the directionality is also an important property of social ties, though usually ignored in undirected social networks for its invisibility. However, we believe most social ties are natively directed, and the awareness of directionality can improve our understanding about the network structures and further benefit social network analysis and mining tasks. Thus it's appealing to study whether there exist interesting patterns about directionality in social networks and whether we can learn the directions for undirected networks based on these patterns. In this study, we engage in the investigation of directionality patterns on real-world directed social networks and summarize our findings using four consistency hypotheses. Based on these hypotheses, we propose ReDirect, an optimization framework which makes it possible to infer the hidden directions of undirected social ties based on the network topology only. This general framework can incorporate various predictive models under specific scenarios. Furthermore, we show how to improve ReDirect by introducing semi/self-supervision in the framework and how to construct the self-labeled training data using simple but effective heuristics. Experimental results show that even without external information, our approach can recover the directions of networks effectively.
Jun Zhang 0004, Chaokun Wang, Jianmin Wang 0001
WWW1
2014 Inferring Continuous Dynamic Social Influence and Personal Preference for Temporal Behavior Prediction
abstract
It is always attractive and challenging to explore the intricate behavior data and uncover people's motivations, preference and habits, which can greatly benefit many tasks including link prediction, item recommendation, etc. Traditional work usually studies people's behaviors without time information in a static or discrete manner, assuming the underlying factors stay invariant in a long period. However, we believe people's behaviors are dynamic, and the contributing factors including the social influence and personal preference for behaviors are varying continuously over time. Such continuous dynamics convey important knowledge about people's behavior patterns; ignoring them would lead to inaccurate models. In this work, we address the continuous dynamic modeling of temporal behaviors. To model the fully continuous temporal dynamics of behaviors and the underlying factors, we propose the DP-Space, a dynamic preference probability space, which can capture their smooth variation in various shapes over time with flexible basis functions. Upon that we propose a generative dynamic behavior model, ConTyor, which considers the temporal item-adoption behaviors as joint effect of dynamic social influence and varying personal preference over continuous time. We also develop effective inference methods for ConTyor and present its applications. We conduct a comprehensive experimental study using real-world datasets to evaluate the effectiveness of our model and the temporal modeling. Results verify that ConTyor outperforms existing state-of-the-art static and temporal models in behavior predictions. Moreover, in our detailed study on temporal modeling, we show that temporal modeling is superior to static approaches and modeling over continuous time is further better than that over discrete time. We also demonstrate that the ancient behavior data can still become important and beneficial if modeled well.
Jun Zhang 0004, Chaokun Wang, Jianmin Wang 0001, Jeffrey Xu Yu
Proc. VLDB Endow.1
2014 Performance Analysis and System Design for Hierarchical Modulated BICM-ID
abstract
Hierarchical modulation (HM) enables unequal priority transmissions using a signal constellation of non-uniformly spaced constellation points, which is an important property for broadcast systems. Using bit interleaved coded modulation (BICM) and iterative decoding (ID), the performance of HM systems can be further improved. In this paper, we study HM system with BICM-ID or HM-BICM-ID. Since the performance of BICM-ID depends on signal mapping rules, we derive a mapping rule based on distance properties to minimize the bit error rate (BER). Furthermore, we perform the BER performance analysis for HM-BICM-ID based on binary erasure channel (BEC) modeling, which provides BER prediction without time-consuming simulations. Based on the derived BER prediction scheme, we can also perform system design and optimization for applications of HM-BICM-ID. For example, we are able to optimize the constellation priority parameter in a broadcast system to maximize the throughput.
Qiaoyu Li, Jun Zhang 0004, Lin Bai 0001, Jinho Choi 0001
IEEE Trans. Wirel. Commun.2
2014 Throughput and Energy Efficiency Analysis of Small Cell Networks with Multi-Antenna Base Stations
abstract
Small cell networks have recently been proposed as an important evolution path for the next-generation cellular networks. However, with more and more irregularly deployed base stations (BSs), it is becoming increasingly difficult to quantify the achievable network throughput or energy efficiency. In this paper, we develop an analytical framework for downlink performance evaluation of small cell networks, based on a random spatial network model, where BSs and users are modeled as two independent spatial Poisson point processes. A new simple expression of the outage probability is derived, which is analytically tractable and is especially useful with multi-antenna transmissions. This new result is then applied to evaluate the network throughput and energy efficiency. It is analytically shown that deploying more BSs can always increase the network throughput, but the throughput will scale with the BS density first linearly, then logarithmically, and finally converge to a constant. On the other hand, increasing the number of BS antennas can decrease the outage probability exponentially, thus can always increase the network throughput. However, increasing the BS density or the number of transmit antennas will first increase and then decrease the energy efficiency if the non-transmission power or the circuit power consumption is less than certain thresholds, and the optimal BS density and the optimal number of BS antennas can be found. Otherwise, the energy efficiency will always decrease. Simulation results shall demonstrate that our conclusions based on the random network model are general and also hold in a regular grid-based model.
Chang Li 0002, Jun Zhang 0004, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.2
2014 Coordinated 3D Beamforming for Interference Management in Cellular Networks
abstract
We consider downlink transmission in a cellular network consisting of multi-antenna base stations (BSs), single-antenna mobile users, directional antennas each with a vertically adjustable beam, and control signaling delays in feedback and backhaul links. We propose a novel transmission technique in which intercell interference management is performed via coordinating the beamforming jointly in the horizontal and vertical planes of the wireless channel, denoted as coordinated 3D beamforming. In the horizontal plane, we focus on intercell interference cancelation (ICIC) and investigate its performance when the provided channel state information (CSI) is impaired due to delay/mobility. It is demonstrated that the superiority of ICIC over conventional intracell maximum ratio transmission is highly dependent on the CSI accuracy. In the vertical plane, we consider intercell interference control via coordinatively adapting the elevation angle of the BS antenna pattern, denoted as tilt, to the locations of the scheduled users. It is shown that with perfect CSI interference management should be performed in the horizontal plane using ICIC, while at high delay/mobility it should be done in the vertical plane via coordinated tilt adaptation. For intermediate values of delay/mobility, joint interference management in both planes is required.
Nima Seifi, Jun Zhang 0004, Robert W. Heath Jr., Tommy Svensson, Mikael Coldrey
IEEE Trans. Wirel. Commun.2
2014 Group Sparse Beamforming for Green Cloud-RAN
abstract
A cloud radio access network (Cloud-RAN) is a network architecture that holds the promise of meeting the explosive growth of mobile data traffic. In this architecture, all the baseband signal processing is shifted to a single baseband unit (BBU) pool, which enables efficient resource allocation and interference management. Meanwhile, conventional powerful base stations can be replaced by low-cost low-power remote radio heads (RRHs), producing a green and low-cost infrastructure. However, as all the RRHs need to be connected to the BBU pool through optical transport links, the transport network power consumption becomes significant. In this paper, we propose a new framework to design a green Cloud-RAN, which is formulated as a joint RRH selection and power minimization beamforming problem. To efficiently solve this problem, we first propose a greedy selection algorithm, which is shown to provide near-optimal performance. To further reduce the complexity, a novel group sparse beamforming method is proposed by inducing the group-sparsity of beamformers using the weighted ℓ1/ℓ2-norm minimization, where the group sparsity pattern indicates those RRHs that can be switched off. Simulation results will show that the proposed algorithms significantly reduce the network power consumption and demonstrate the importance of considering the transport link power consumption.
Yuanming Shi, Jun Zhang 0004, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.2
2013 Performance analysis of SDMA in multicell wireless networks
abstract
Multi-antenna transmission, or MIMO, is a major enabling technique for broadband cellular networks. The current implementation, however, is mainly for the point-to-point link, and its potential for Space-Division Multiple Access (SDMA) has not been fully exploited. In this paper, we will analytically evaluate the performance of SDMA in multicell networks based on a spatial random network model, where both the base stations (BSs) and users are modeled as two independent Poisson point processes. The main difficulty is the evaluation of the interference distribution, for which we propose a novel BS grouping approach that leads to a closed-form expression for the network area spectral efficiency. We find that the number of active users (U) served with SDMA is critical, as it affects the spatial multiplexing gain, the aggregated interference, and the diversity gain for each user. The optimal value of U can be selected based on our analytical result, with which SDMA is shown to outperform both the single-user beamforming and full-SDMA for which U is the same as the number of BS antennas. In particular, it is shown that the performance gain of SDMA is higher when the BS density is relatively small compared to the user density, but the optimal value of U is almost the same for different scenarios, which is close to half of the BS antenna number.
Chang Li 0002, Jun Zhang 0004, Khaled Ben Letaief
GLOBECOM2
2013 Relay selection for energy harvesting cooperative communication systems
abstract
Energy harvesting (EH) has recently emerged as a promising technique for green communications, as it can power communication systems with renewable energy. In this paper, we investigate how to adopt cooperative relay selection to improve the short-term performance of EH communication systems. The main focus is on how to efficiently utilize the available side information (SI), including channel side information (CSI) and energy side information (ESI). We formulate relay selection problems with either non-causal or causal SI, with an emphasis on the more practical causal case. For this causal SI case, we propose a low-complexity relay selection strategy based on the relative throughput, that is, in each block, the relay with enough energy and with the highest instantaneous throughput compared with the average throughput is selected. This relay selection rule captures the key characteristic of EH systems, namely, each relay should have some chance to be selected so that the harvested energy can be efficiently utilized, and it should be selected only if its throughput is near its own peak. Simulation results will show that the proposed relay selection method provides significant throughput gain over the conventional one which is only based on the current side information.
Yaming Luo, Jun Zhang 0004, Khaled Ben Letaief
GLOBECOM2
2013 Energy efficiency analysis of small cell networks
abstract
Small cell networks have recently been proposed as an important evolution path for the next-generation cellular networks. While such approach has the potential of meeting the growing network throughput requirement, the energy efficiency of small cell networks is of great concern as the base station (BS) density will be significantly increased. The objective of this paper is to analyze the energy efficiency in small cell networks. To do so, we adopt a random spatial network model, where BSs and users are modeled as two independent spatial Poisson point processes (PPPs). We shall derive analytical results for the network energy efficiency, which show that the BS power consumption model plays a critical role. In particular, it will be shown that increasing the BS density can actually improve the energy efficiency if the BS power consumption that is not related to signal transmission is less than a certain threshold. By comparing the cases between single-antenna and multi-antenna BSs, we find that single-antenna BSs provide a higher energy efficiency if the circuit power is larger than a threshold. Simulation results will demonstrate that our conclusions which are based on the random network model also hold in a regular grid-based model.
Chang Li 0002, Jun Zhang 0004, Khaled Ben Letaief
ICC2
2013 Throughput maximization for two-hop energy harvesting communication systems
abstract
Energy harvesting (EH) has recently emerged as a promising technique for green communications. To realize its potential, communication protocols need to be redesigned to combat the randomness of the harvested energy. In this paper, we investigate how to apply relaying to improve the short-term performance of EH communication systems. With an EH source and a non-EH half-duplex relay, we consider the problem of maximizing the achievable rate for a given time duration. The half-duplex constraint at the relay renders the design problem quite challenging, as the source and relay transmission periods should be carefully scheduled. Moreover, the adaptive power allocation is needed at the source to combat the random energy arrivals. A key finding is that the optimal power allocation algorithm, called directional water-filling (DWF), for the single-hop EH system can serve as guideline for the design of a two-hop communication system, as it not only provides an achievable performance upper bound, but also forms the basis to derive the optimal solution for our design problem. Based on a modified energy profile according to the DWF power allocation, we derive key properties of the optimal solution and thereafter propose an efficient algorithm to maximize the throughput. Simulation results show that both scheduling and power allocation optimizations are necessary in two-hop EH communication systems.
Yaming Luo, Jun Zhang 0004, Khaled Ben Letaief
ICC2
2013 LAFT-Explorer: inferring, visualizing and predicting how your social network expands
abstract
The study of social network evolution has attracted many attentions from both the industry and academia. In this paper we demonstrate LaFT-Explorer, a general toolkit for explaining and reproducing the network growth process based on the friendship propagation. LaFT-Explorer presents multiple perspectives for analyzing the network evolution process and structure, including LaFT-Tree, LaFT-Trace and LaFT-Flow. Upon that we build LaFT-Rec, a new visualized interactive friend recommendation service based on the friendship propagation. LaFT-Rec not only shows whom one may make friends with, but also tells the user that why you should make friends with him and how you can reach him. We demonstrate our system built upon the academic social network of DBLP.
Jun Zhang 0004, Chaokun Wang, Yuanchi Ning, Yichi Liu, Jianmin Wang 0001, Philip S. Yu
KDD1
2013 Learning latent friendship propagation networks with interest awareness for link prediction
abstract
It's well known that the transitivity of friendship is a popular sociological principle in social networks. However, it's still unknown that to what extent people's friend-making behaviors follow this principle and to what extent it can benefit the link prediction task.
Jun Zhang 0004, Chaokun Wang, Philip S. Yu, Jianmin Wang 0001
SIGIR1
2013 LaFT-tree: perceiving the expansion trace of one's circle of friends in online social networks
abstract
Many patterns have been discovered to explain and analyze how people make friends. Among them is the triadic closure, supported by the principle of the transitivity of friendship, which means for an individual the friends of her friend are more likely to become her new friends. However, people's motivations under this principle haven't been well studied, and it's still unknown that how this principle works in diverse situations.
Jun Zhang 0004, Chaokun Wang, Jianmin Wang 0001, Philip S. Yu
WSDM1
2013 Lattice Reduction-Based Approximate MAP Detection with Bit-Wise Combining and Integer Perturbed List Generation
abstract
For iterative detection and decoding (IDD) in multiple-input multiple-output (MIMO) systems, the log-likelihood ratio (LLR) of each coded bit can be found by an optimal bit-wise maximum a posteriori probability (MAP) detector. However, since this MAP detector requires a prohibitively high computational complexity, low-complexity suboptimal detectors are desirable. In this paper, lattice reduction (LR)-based MIMO detection is investigated to derive a low-complexity detector that can achieve near MAP performance for IDD. In order to approximate LLR values incorporating the extrinsic information provided by a soft-input soft-output (SISO) decoder, bit-wise LR-based minimum mean square error (MMSE) filters are derived. Furthermore, in order to minimize the performance degradation due to quantization (or rounding) errors in the LR-based detection, a low-complexity integer perturbed list generation method is proposed, where no tree search is used by taking advantage of a near orthogonal channel basis obtained by LR. Through a complexity analysis and simulations, it is shown that the proposed approach achieves near optimal performance, while the complexity is comparable with that of the MMSE soft cancellation method, which is known to be computationally efficient. As a bit-wise detector, a parallel implementation of the proposed method would be straightforward, which lowers the detection delay.
Qiaoyu Li, Jun Zhang 0004, Lin Bai 0001, Jinho Choi 0001
IEEE Trans. Commun.2
2013 Optimal Scheduling and Power Allocation for Two-Hop Energy Harvesting Communication Systems
abstract
Energy harvesting (EH) has recently emerged as a promising technique for green communications. To realize its potential, communication protocols need to be redesigned to combat the randomness of the harvested energy. In this paper, we investigate how to apply relaying to improve the short-term performance of EH communication systems. With an EH source and a non-EH half-duplex relay, we consider two different design objectives: 1) short-term throughput maximization; and 2) transmission completion time minimization. Both problems are joint time scheduling and power allocation problems, rendered quite challenging by the half-duplex constraint at the relay. A key finding is that directional water-filling (DWF), which is the optimal power allocation algorithm for the single-hop EH system, can serve as guideline for the design of two-hop communication systems, as it not only determines the value of the optimal performance, but also forms the basis to derive optimal solutions for both design problems. Based on a relaxed energy profile along with the DWF algorithm, we derive key properties of the optimal solutions for both problems and thereafter propose efficient algorithms. Simulation results will show that both time scheduling and power allocation optimizations are necessary in two-hop EH communication systems.
Yaming Luo, Jun Zhang 0004, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.2
2012 Training optimization for energy harvesting communication systems
abstract
Energy harvesting (EH) has recently emerged as an effective way to solve the lifetime challenge of wireless sensor networks, as it can continuously harvest energy from the environment. Unfortunately, it is challenging to guarantee a satisfactory short-term performance in EH communication systems because the harvested energy is sporadic. In this paper, we consider the channel training optimization problem in EH communication systems, i.e., how to obtain accurate channel state information to improve the communication performance. In contrast to conventional communication systems, the optimization of the training power and training period in EH communication systems is a coupled problem, which makes such optimization very challenging. We shall formulate the optimal training design problem for EH communication systems, and propose two solutions that adaptively adjust the training period and power based on either the instantaneous energy profile or the average energy harvesting rate. Numerical and simulation results will show that training optimization is important in EH communication systems. In particular, it will be shown that for short block lengths, training optimization is critical. In contrast, for long block lengths, the optimal training period is not too sensitive to the value of the block length nor to the energy profile. Therefore, a properly selected fixed training period value can be used.
Yaming Luo, Jun Zhang 0004, Khaled Ben Letaief
GLOBECOM2
2012 Coordinated relay beamforming for amplify-and-forward two-hop interference networks
abstract
Relaying is a promising technique to extend coverage and improve throughput in wireless networks, but its performance is degraded in the presence of co-channel interference. In this paper, we consider coordinated relay beamforming to suppress interference and improve the date rates of two-hop interference networks. We first propose optimal coordinated relay beamforming algorithms to characterize the achievable rate region and maximize the sum-rate. By imposing a constraint on the desired signals, a low-complexity iterative algorithm is then proposed to maximize the sum-rate. Through performance comparison, we show that the proposed relaying strategy provides a promising tradeoff between complexity and performance. To further reduce design complexity, we propose a new interference management scheme, interference neutralization, to cancel the interferences over the air at the second hop. We show that this scheme yields a closed-form solution for the beamforming design and provides good performance especially at high signal-to-noise ratio (SNR).
Yuanming Shi, Jun Zhang 0004, Khaled Ben Letaief
GLOBECOM2
2012 Maximizing energy efficiency in wireless networks with a minimum average throughput requirement
abstract
Given the growing concern over energy consumption and associated global warming, green communication is becoming more and more important. Lots of efforts have been put into investigating energy efficiency based design in wireless systems. Unfortunately, the maximum energy efficiency of a point-to-point link is normally achieved when the transmit power approaches zero, which, however, is not desirable in practical systems due to the low achievable data rate. In this paper, we consider the energy efficiency optimization with a practical power consumption model, where, besides a maximum transmit power constraint, we set a rate constraint (r0) to guarantee an average throughput requirement (Rth). Due to the possible outage in transmission, r0is in general different from Rth, and determining r0for a given Rthis not trivial. We shall derive a closed-form solution for the optimal transmit power that maximizes energy efficiency. We will also demonstrate that a carefully selected rate constraint r0can guarantee the required average throughput Rthand provide the freedom to achieve different tradeoffs between energy efficiency and average throughput.
Chang Li 0002, Shenghui Song 0001, Jun Zhang 0004, Khaled Ben Letaief
WCNC3
2011 Optimizing Training and Feedback for Spatial Intercell Interference Cancellation
abstract
In this paper, we investigate spatial intercell interference cancellation - an efficient technique to mitigate intercell interference in multicell networks. We consider a practical model for channel state information (CSI), where the transmit CSI is acquired through downlink training and uplink feedback. Due to the requirement of channel information from multiple base stations, the training and feedback design is quite different from conventional single-cell processing systems. We optimize training and feedback, where both analog and digital feedback is considered. For analog feedback, it is shown that the downlink training optimization provides a more significant performance gain than feedback optimization; while conversely for digital feedback over a finite-rate feedback channel, the feedback bit allocation is more important than the training optimization.
Jun Zhang 0004, Jeffrey G. Andrews, Khaled Ben Letaief
GLOBECOM1
2011 Location-Based Joint Relay Selection and Channel Allocation for Cognitive Radio Networks
abstract
In cognitive radio networks (CRNs), dynamic spectrum access has been demonstrated as an effective way to improve the spectrum utilization. Spectrum holes can be exploited not only in certain time slots or frequency bands, but also at particular locations. In relay assisted CRNs, one relay at a certain location can help to identify and provide different spectrum holes over multiple channels. In this paper, a multi-dimensional combinatorial optimization problem is formulated for joint relay selection and channel allocation. We propose a weighted bipartite graph model and a minimum weighted assignment approach to efficiently get the optimal solution of the considered problem. Simulation results show that by applying this approach, spectrum efficiency, relay selection diversity and power efficiency can be improved simultaneously for the cognitive users. Besides, only the statistical channel state information is needed and the allocation results can be computed efficiently by using the proposed approach.
Fangyong Li, Bo Bai 0001, Jun Zhang 0004, Khaled Ben Letaief
GLOBECOM3
2011 On the Accuracy of the Wyner Model in Downlink Cellular Networks
abstract
Compared to real cellular systems where users are spatially distributed and interference levels vary by several orders of magnitude over a cell, in the Wyner model user locations are fixed and the interference intensity is characterized by a single fixed parameter. Although it is a fairly extreme simplification, the Wyner model has been extensively used to analyze cellular networks. Does it capture some of the main trends of such networks or not? In this study of downlink cellular networks, we show that from an outage point of view, the Wyner model is highly inaccurate since outage is primarily a function of user location. However, in the case of average throughput, the Wyner model may in some special cases be an acceptable simplification if the interference parameter is set appropriately. In particular, we show that it is relatively accurate in terms of the average throughput for CDMA systems with single-cell processing and perfect channel inversion, and for the sum throughput of multicell processing with equal transmit power per user. In short, the Wyner model appears to be a reasonable approximation for SINR mean-based metrics like sum and average throughput for certain scenarios, but is unreasonable in nearly all cases for SINR tail-based metrics like outage probability.
Jiaming Xu 0002, Jun Zhang 0004, Jeffrey G. Andrews
ICC2
2011 Multi-Mode Transmission for the MIMO Broadcast Channel with Imperfect Channel State Information
abstract
This paper proposes an adaptive multi-mode transmission strategy to improve the spectral efficiency achieved in the multiple-input multiple-output (MIMO) broadcast channel with delayed and quantized channel state information. The adaptive strategy adjusts the number of active users, denoted as the transmission mode, to balance transmit array gain, spatial division multiplexing gain, and residual inter-user interference. Accurate closed-form approximations are derived for the achievable rates for different modes, which help identify the active mode that maximizes the average sum throughput for given feedback delay and channel quantization error. The proposed transmission strategy can be easily combined with round-robin scheduling to serve a large number of users. As instantaneous channel information is not exploited, the proposed algorithm cannot provide multiuser diversity gain, but it is still able to provide throughput gain over single-user MIMO at moderate signal-to-noise ratio. In addition, it has a light feedback overhead and only requires feedback of instantaneous channel state information from a small number of users. In the system with a feedback load constraint, it is shown that the proposed algorithm provides performance close to that achieved by opportunistic scheduling with instantaneous feedback from a large number of users.
Jun Zhang 0004, Marios Kountouris, Jeffrey G. Andrews, Robert W. Heath Jr.
IEEE Trans. Commun.1
2011 On the Accuracy of the Wyner Model in Cellular Networks
abstract
The Wyner model has been widely used to model and analyze cellular networks due to its simplicity and analytical tractability. Its key aspects include fixed user locations and the deterministic and homogeneous interference intensity. While clearly a significant simplification of a real cellular system, which has random user locations and interference levels that vary by several orders of magnitude over a cell, a common presumption by theorists is that the Wyner model nevertheless captures the essential aspects of cellular interactions. But is this true? To answer this question, we compare the Wyner model to a model that includes random user locations and fading. We consider both uplink and downlink transmissions and both outage-based and average-based metrics. For the uplink, for both metrics, we conclude that the Wyner model is in fact quite accurate for systems with a sufficient number of simultaneous users, e.g., a CDMA system. Conversely, it is broadly inaccurate otherwise. Turning to the downlink, the Wyner model becomes inaccurate even for systems with a large number of simultaneous users. In addition, we derive an approximation for the main parameter in the Wyner model - the interference intensity term, which depends on the path loss exponent.
Jiaming Xu 0002, Jun Zhang 0004, Jeffrey G. Andrews
IEEE Trans. Wirel. Commun.2
2010 When Does the Wyner Model Accurately Describe an Uplink Cellular Network?
abstract
The Wyner model has been widely used to model and analyze cellular networks due to its simplicity and analytical tractability. The key aspects of this model are fixed user location and deterministic and homogeneous interference intensity. While clearly a significant simplification of a real cellular system, which has random user locations and interference levels that can vary by several orders of magnitude over a cell, a common presumption is that the Wyner model nevertheless captures the essential aspects of cellular interactions. But is this true? In this study of uplink cellular networks, we argue that the Wyner model is only accurate for systems with a sufficient number of simultaneous users. Therefore, it is a reasonable abstraction for CDMA multicell networks but quite inaccurate for those employing TDMA. With single-cell signal processing, the Wyner model fails to capture the fact that intracell TDMA is advantageous over CDMA in terms of ergodic symmetric throughput and that random user locations increase throughput. In the case of multi-cell processing, it is shown that intracell TDMA is suboptimal in terms of ergodic symmetric capacity, which is in sharp contrast to results obtained under the Wyner model wherein intracell TDMA is proved to be optimal.
Jiaming Xu 0002, Jun Zhang 0004, Jeffrey G. Andrews
GLOBECOM2
2010 Adaptive Spatial Intercell Interference Cancellation in Multicell Wireless Networks
abstract
Downlink spatial intercell interference cancellation (ICIC) is considered for mitigating other-cell interference using multiple transmit antennas. A principle question we explore is whether it is better to do ICIC or simply standard single-cell beamforming. We explore this question analytically and show that beamforming is preferred for all users when the edge SNR (signal-to-noise ratio) is low (10 dB), for example in an urban setting. At medium SNR, a proposed adaptive strategy, where multiple base stations jointly select transmission strategies based on the user location, outperforms both while requiring a lower feedback rate than the pure ICIC approach. The employed metric is sum rate, which is normally a dubious metric for cellular systems, but surprisingly we show that even with this reward function the adaptive strategy also improves fairness. When the channel information is provided by limited feedback, the impact of the induced quantization error is also investigated. The analysis provides insights on the feedback design, and it is shown that ICIC with well-designed feedback strategies still provides significant throughput gain.
Jun Zhang 0004, Jeffrey G. Andrews
IEEE J. Sel. Areas Commun.1
2009 Block Diagonalization in the MIMO Broadcast Channel with Delayed CSIT
abstract
This paper investigates the impact of delayed channel state information at the transmitter (CSIT) on the MIMO broadcast channel with block diagonalization (BD) preceding. First, an upper bound for the achievable throughput is provided, which shows that BD is more robust to imperfect CSIT than zero-forcing precoding as it has fewer inter-user interfering streams. Due to residual inter-user interference, the throughput of BD still saturates at high SNR, which motivates switching between single-user and multi-user precoding. An accurate closed-form approximation is derived for the achievable throughput of the BD system, which provides guidance on the preferred transmission technique for a given scenario.
Jun Zhang 0004, Jeffrey G. Andrews, Robert W. Heath Jr.
GLOBECOM1
2009 Achievable throughput of multi-mode multiuser MIMO with imperfect CSI constraints
abstract
In the multiple-input multiple-output (MIMO) broadcast channel with imperfect channel state information (CSI), neither the capacity nor the optimal transmission technique have been fully discovered. In this paper, we derive achievable ergodic rates for a multi-antenna fading broadcast channel when CSI at the transmitter (CSIT) is delayed and quantized. It is shown that not all possible users should be supported with spatial division multiplexing due to the residual inter-user interference caused by imperfect CSIT. Based on the derived achievable rates, we propose a multi-mode transmission strategy to maximize the throughput, which adaptively adjusts the number of active users based on the channel statistics information.
Jun Zhang 0004, Marios Kountouris, Jeffrey G. Andrews, Robert W. Heath Jr.
ISIT1
2009 Networked MIMO with clustered linear precoding
abstract
A clustered base transceiver station (BTS) coordination strategy is proposed for a large cellular MIMO network, which includes full intra-cluster coordination-to enhance the sum rate-and limited inter-cluster coordination-to reduce interference for the cluster edge users. Multi-cell block diagonalization is used to coordinate the transmissions across multiple BTSs in the same cluster. To satisfy per-BTS power constraints, three combined precoder and power allocation algorithms are proposed with different performance and complexity tradeoffs. For inter-cluster coordination, the coordination area is chosen to balance fairness for edge users and the achievable sum rate. It is shown that a small cluster size (about 7 cells) is sufficient to obtain most of the sum rate benefits from clustered coordination while greatly relieving channel feedback requirement. Simulations show that the proposed coordination strategy efficiently reduces interference and provides a considerable sum rate gain for cellular MIMO networks.
Jun Zhang 0004, Runhua Chen, Jeffrey G. Andrews, Arunabha Ghosh, Robert W. Heath Jr.
IEEE Trans. Wirel. Commun.1
2008 Distributed Antenna Systems with Randomness
abstract
In a cellular distributed antenna system (DAS), distributed antenna elements (AEs) are connected to the base station via an offline dedicated link, e.g. fiber optics or line-of-sight RF. Distributed antennas have been recently shown to provide considerable gains in coverage and capacity, at much lower cost than decreasing cell size. Previous studies have neglected the key sources of randomness in such systems, notably (i) random channel effects (fading and shadowing) and (ii) the random quantity and locations of both the mobile users and the AEs. Typically, path loss has been the focus, and the AEs are assumed to be regularly spaced, both of which are significant idealizations. First, we develop an analytical framework that allows random channels to be accommodated. We use this approach to show that selection transmission (using a single AE) is preferable to maximum ratio transmission (which uses all the AEs) in a multicell environment. Interestingly, the opposite is true in an isolated cell. Second, since AEs are placed opportunistically (on tall structures with backhaul access) rather than regularly, we develop a stochastic geometry-inspired approach to determine the outage probability as a function of the number of randomly placed AEs, which we model as a point process. With selection transmission, the outage probability is shown to decrease exponentially with the number of AEs and users. In the most general setup - with multiple distributed antennas and users, and both AE selection and user selection - we show that randomly deployed AEs provide nearly the same performance as regularly spaced AEs.
Jun Zhang 0004, Jeffrey G. Andrews
IEEE Trans. Wirel. Commun.1
2007 QuickCN: A Combined Approach for Efficient Keyword Search over Databases
Jun Zhang 0004, Zhaohui Peng, Shan Wang 0001
DASFAA1
2007 Cellular Communication with Randomly Placed Distributed Antennas
abstract
A cellular distributed antenna system with randomly located distributed antenna elements (AEs) and mobile users is considered. The AEs are connected to the base station via an offline dedicated link (such as fiber optics or line-of- sight RF). A user is served by selecting the AE with the best channel to it (often the closest one). The outage probability is derived for both AE selection and user selection individually for an isolated cell, and shown to decrease exponentially with the number of AEs and users. Due to selection diversity, fading and shadowing typically are desirable effects that decrease the likelihood of outage. In a more general setup - with multiple distributed antennas and both AE selection and user selection - the outage probability decreases exponentially with the number of users but not with the number of AEs. Due to diminishing returns, there is no need to deploy more AEs than a certain extent, which is determined by the transmit power of AEs. With a sufficient number of AEs, randomly deployed AEs can provide nearly the same performance as regularly deployed AEs, and both have a common outage probability floor which is determined by user density.
Jun Zhang 0004, Jeffrey G. Andrews
GLOBECOM1
2007 CLASCN: Candidate Network Selection for Efficient Top- k Keyword Queries over Databases
Jun Zhang 0004, Zhaohui Peng, Shan Wang 0001, Huijing Nie
J. Comput. Sci. Technol.1
2006 Multiple-Source Multiple-Relay Cooperation System
abstract
In this paper, we study the multiple-source multiple-relay cooperation system. In the transmission protocol we proposed, the relay nodes not only decode and forward the symbols from the source nodes, but also encode the incoming symbols according to the cooperative code, resulting in higher efficiency and flexibility compared with previous repetition-coded cooperation system. We analyze the impact of the noisy interuser channel between the source nodes and the relay nodes on the system performance, and then propose an adaptive cooperative coding scheme to compensate for it. The analysis and the simulation results show that the proposed system achieves full diversity.
Jun Zhang 0004, Tat-Ming Lok
ICC1
2006 Si-SEEKER: Ontology-Based Semantic Search over Databases
Jun Zhang 0004, Zhaohui Peng, Shan Wang 0001, Huijing Nie
KSEM1
2006 NUITS: A Novel User Interface for Efficient Keyword Search over Databases
Shan Wang 0001, Zhaohui Peng, Jun Zhang 0004, Lu Qin 0001, Jeffrey Xu Yu, Bolin Ding
VLDB3
2006 TreeCluster: Clustering Results of Keyword Search over Databases
Zhaohui Peng, Jun Zhang 0004, Shan Wang 0001, Lu Qin 0001
WAIM2
2006 Performance comparison of conventional and cooperative multihop transmission
abstract
In this paper, we analyze the decode-and-forward cooperative multihop transmission. Different from conventional multihop transmission, each node achieves the spatial diversity by combining all the independently fading symbols from previous nodes on the route line. We consider the possibility for the relay node to forward error-detected symbols, and propose an adaptive transmission protocol to compensate for it. The bit error rate is derived for the cooperation system, and the comparison with the conventional multihop transmission shows that this system can achieve full diversity order. Practical cooperative range and the impact of relay nodes distribution on the system performance are discussed for the practical implementation
Jun Zhang 0004, Tat-Ming Lok
WCNC1
2006 PreCN: Preprocessing Candidate Networks for Efficient Keyword Search over Databases
Jun Zhang 0004, Zhaohui Peng, Shan Wang 0001, Huijing Nie
WISE1
2005 Traffic Measurement and Analysis of TUNET
abstract
Traffic measurement and analysis, as one of the important methods of understanding and characterizing network, can provide significant support for network management. After a brief introduction of a novel NP (network processor)-based architecture of traffic measurement, the paper presents the detailed analysis results of the traffic collected from the gigabit link connecting Tsinghua University campus network (in short, TUNET) to its upstream ISP, China Education and Research NETwork (in short, CERNET). Then, the paper comprehensively analyzes the traffic from multi-dimension viewpoints, including temporal distribution, packet length distribution, port-based distribution, protocol-based distribution, and TopN statistics. Such analysis not only provides support for the study of user behavior, but also enriches traffic measurement technology
Jun Zhang 0004, Jiahai Yang 0001, Changqing An, Jilong Wang 0001
CW1