VLDB 2026 Research / reviewers in the wild / expert
Qianqian Yang 0002
dblp:10/9659-2
· DBLP profile ↗
77ranked-venue papers
9as first author
64since 2021 · last 2026
0000-0003-4747-9410ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 49 · 5 first-author · 42 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 4 since 2021Theory of computation · 3 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AFDM-guided Deep Joint Source-Channel Coding for Satellite Communication
Junyu Pan, Kaiyi Chi, Qianqian Yang 0002, Zhiguo Shi 0001 |
ICC | 3 |
| 2026 | AVA: Towards Agentic Video Analytics with Vision Language Models
Yuxuan Yan, Shiqi Jiang 0002, Ting Cao 0003, Yifan Yang 0004, Qianqian Yang 0002, Yuanchao Shu, Yuqing Yang 0001, Lili Qiu |
NSDI | 5 |
| 2026 | Federated Split Learning for Large Language Models With RSMAabstractABSTRACT This study proposes a federated split learning framework for large language models (FedsLLM) integrated with rate‐splitting multiple access (RSMA), aimed at enhancing the efficiency and privacy of LLM training in wireless communication systems. By leveraging low‐rank adaptation (LoRA) to distribute computational loads and a fluid antenna system to dynamically optimize channel capacity, the framework effectively reduces training latency through joint optimization of learning accuracy and communication resources. Experimental results demonstrate that the proposed framework significantly outperforms traditional time‐division multiple access including time division multiple access, frequency division multiple access (FDMA), enhanced bandwidth FDMA, and fairness‐enhanced FDMA across multiple metrics: at a transmit power of 20 dBm, RSMA reduces task completion time by 8.3%; under 20 MHz bandwidth, it achieves a 25% performance improvement; and even with a data volume of 900 Kbits, it maintains a 12% advantage. The adopted alternating optimization algorithm converges rapidly, reaching 95% of the optimal value within only 5 iterations, substantially outperforming the fixed‐point method. Overall, FedsLLM‐RSMA effectively addresses privacy, computational and communication bottlenecks in distributed LLM training. Compared to TDMA, it reduces total training latency by 28% and improves communication efficiency by 35%, while achieving higher model accuracy and faster convergence. This work provides a viable pathway for efficient and scalable deployment of LLMs in 6G networks. Jianxin Dai, Feibo Jiang, Zhaohui Yang 0001, Qianqian Yang 0002, Zhaoyang Zhang 0001, Linqing Gui |
IET Commun. | 5 |
| 2026 | Online Energy Efficient Multimodal Probabilistic Semantic CommunicationabstractIn this paper, we investigate an uplink multi-modal probabilistic semantic communication (PSCom) system based on probability graph in the satellite scenario. The system consists of both: semantic computation and traditional communication. Firstly, at the user end, the transmitted data is compressed based on the probabilistic graph. Then, the compressed data is transmitted to the satellite, which uses the same probabilistic graph to recover the received data. In the considered model, this paper addresses an optimization problem for multi-modal multi-user semantic communication across multiple time slots. An optimization problems formulated aiming to minimize the total energy consumption of the PSCom system, with satisfying the transmission time, transmission power, transmission bandwidth, local computation frequency, and transmission data requirements. To solve this problem, the Lyapunov drift-plus-penalty function based on online optimization is first used to transform the multi-slot problem into a stochastic single-slot problem, thereby converting the optimization problem into a trade-off between system energy consumption and queue length. Subsequently, an alternating algorithm is proposed to iteratively optimize, semantic compression rate, local computation frequency, transmission bandwidth, transmission power, and time allocation variables. Finally, simulation experiments demonstrates the effectiveness of the proposed algorithm. Jianxin Dai, Zhouxiang Zhao, Zhaohui Yang 0001, Jianglin Ye, Qianqian Yang 0002, Chongwen Huang, Zhaoyang Zhang 0001 |
IEEE Internet Things J. | 6 |
| 2026 | Bidirectional Motion-Enhanced Semantic Communication for Wireless Video TransmissionabstractWith the increasing proliferation of Ultra-High-Definition (UHD) videos, the demand for efficient video transmission schemes to alleviate network congestion is growing. In this paper, we propose a bi-directional motion enhanced semantic communication (SemCom) system for efficient and robust video transmission. In particular, we introduce a bi-directional motion estimation module to capture inter-frame differences caused by camera movements, where the obtained forward and backward motion vectors are combined with the residual information to generate motion-compensated frames. We also introduce a predicted feature module to discard semantically redundant features, prioritizing crucial semantic-related content. Leveraging information from previously reconstructed frames, the frame prediction module refines predicted frames with the assistance of the motion compensation module. To enhance the system’s robustness to channel noise, we propose a noise attention module that assigns varying importance weights to the extracted features under different channel conditions. Experimental results show that our proposed method outperforms existing deep learning (DL)-based approaches in terms of transmission efficiency, achieving about 33.3% reduction in the number of transmitted symbols while improving the peak signal-to-noise ratio (PSNR) and multi-scale structural similarity index measure (MS-SSIM) performance by an average of 0.56 dB and 0.0024 over an additive white Gaussian noise channel for different schemes. When employing the same compression ratio, our method achieves an average gain of 0.637 dB in PSNR and 0.0038 in MS-SSIM over the slow Rayleigh fading channel. Qianqian Yang 0002, Shibo He, Zhiguo Shi 0001 |
IEEE Internet Things J. | 2 |
| 2026 | Can Knowledge Improve Security? A Coding-Enhanced Jamming Approach for Semantic CommunicationabstractAs semantic communication (SemCom) attracts growing attention as a novel communication paradigm, ensuring the security of transmitted semantic information over open wireless channels has become a critical issue. However, traditional encryption methods often introduce significant additional communication overhead to maintain stability, and conventional learning-based secure SemCom methods typically rely on a channel capacity advantage for the legitimate receiver, which is challenging to guarantee in real-world scenarios. In this paper, we propose a coding-enhanced jamming method that eliminates the need to transmit a secret key by utilizing shared knowledge–potentially part of the training set of the SemCom system–between the legitimate receiver and the transmitter. Specifically, we leverage the shared private knowledge base to generate a set of private digital codebooks in advance using neural network (NN)-based encoders. For each transmission, we encode the transmitted data into digital sequence Y1and associate Y1with a sequence randomly picked from the private codebook, denoted as Y2, through superposition coding. Here, Y1serves as the outer code and Y2as the inner code. By optimizing the power allocation between the inner and outer codes, the legitimate receiver can reconstruct the transmitted data using successive decoding with the index of Y2shared, while the eavesdropper’s decoding performance is severely degraded, potentially to the point of random guessing. Experimental results demonstrate that our method achieves security comparable to state-of-the-art approaches while significantly improving the reconstruction performance of the legitimate receiver by more than 1 dB across varying channel signal-to-noise ratios (SNRs) and compression ratios. Weixuan 'Vincent' Chen, Qianqian Yang 0002, Shuo Shao 0001, Zhiguo Shi 0001, Jiming Chen 0001, Xuemin Shen |
IEEE J. Sel. Areas Commun. | 2 |
| 2026 | DeepGuard: Defending Deep Joint Source-Channel Coding Against Eavesdropping at Physical-LayerabstractDeep joint source-channel coding (DeepJSCC) has emerged as a promising paradigm for efficient and robust information transmission. However, its intrinsic characteristics also pose new security challenges, notably an increased vulnerability to eavesdropping attacks. Existing studies on defending against eavesdropping attacks in DeepJSCC, while demonstrating certain effectiveness, often incur considerable computational overhead or introduce performance trade-offs that may adversely affect legitimate users. In this paper, we present DeepGuard, to the best of our knowledge, the first physical-layer defense framework for DeepJSCC against eavesdropping attacks, validated through over-the-air experiments using software-defined radios (SDRs). Considering that existing eavesdropping attacks against DeepJSCC are limited to simulation under ideal channels, we take a step further by identifying and implementing four representative types of attacks under various configurations in orthogonal frequency-division multiplexing systems. These attacks are evaluated over-the-air under diverse scenarios, allowing us to comprehensively characterize the real-world threat landscape. To mitigate these threats, DeepGuard introduces a novel preamble perturbation mechanism that modifies the preamble shared only between legitimate transceivers. To realize it, we first conduct a theoretical analysis of the perturbation’s impact on the signals intercepted by the eavesdropper. Building upon this, we develop an end-to-end perturbation optimization algorithm that significantly degrades eavesdropping performance while preserving reliable communication for legitimate users. We prototype DeepGuard using SDRs and conduct extensive over-the-air experiments in practical scenarios. Extensive experiments demonstrate that DeepGuard effectively mitigates eavesdropping threats while preserving reliable communication for legitimate users. In particular, DeepGuard can reduce the eavesdropper’s reconstruction performance by as much as 29 dB in PSNR and decrease classification accuracy by up to 91% compared with the performance achieved by the legitimate user. Kaiyi Chi, Yinghui He, Qianqian Yang 0002, Yuanchao Shu, Zhiqin Wang, Jun Luo 0001, Jiming Chen 0001 |
IEEE J. Sel. Areas Commun. | 3 |
| 2026 | Bridging the Modality Gap: Enhancing Channel Prediction With Semantically Aligned LLMs and Knowledge DistillationabstractAccurate channel prediction is essential in massive multiple-input multiple-output (m-MIMO) systems to improve precoding effectiveness and reduce the overhead of channel state information (CSI) feedback. However, existing methods often suffer from accumulated prediction errors and poor generalization to dynamic wireless environments, making it challenging to maintain high prediction accuracy. Large language models (LLMs) have demonstrated remarkable modeling and generalization capabilities in tasks such as time series prediction, making them a promising solution. Nevertheless, a significant modality gap exists between the linguistic knowledge embedded in pretrained LLMs and the intrinsic characteristics of CSI, posing substantial challenges for their direct application to channel prediction. Moreover, the large parameter size of LLMs hinders their practical deployment in real-world communication systems with stringent latency constraints. To address these challenges, we propose a novel channel prediction framework based on semantically aligned large models, referred to as CSI-ALM, which bridges the modality gap between natural language and channel information. Specifically, we design a cross-modal fusion module that aligns CSI representations with the language feature space using a pretrained corpus. Additionally, we maximize the cosine similarity between word embeddings and CSI embeddings to construct semantic cues, effectively leveraging the latent knowledge in LLMs. To reduce complexity and enable practical implementation, we further introduce a lightweight version of the proposed approach, called CSI-ALM-Light. This variant is derived via a knowledge distillation strategy based on attention matrices, which extracts essential features from the teacher model, CSI-ALM, and transfers them to a compact, efficient student model, CSI-ALM-Light. Extensive experimental results demonstrate that CSI-ALM consistently outperforms state-of-the-art deep learning methods across various communication scenarios, achieving substantial performance gains. Moreover, under limited training data conditions—where all models are trained using only 10% of the original training dataset—CSI-ALM-Light, with only 0.34M parameters, attains performance comparable to CSI-ALM and significantly outperforms conventional deep learning approaches. These validate the effectiveness of the proposed approach for accurate and efficient channel prediction in m-MIMO systems. Zhaoyang Li 0005, Qianqian Yang 0002, Zehui Xiong, Zhiguo Shi 0001, Tony Q. S. Quek |
IEEE J. Sel. Areas Commun. | 2 |
| 2026 | Spatial Context-Aware Dynamic Fusion With Mixture-of-Experts for Wireless LocalizationabstractMultimodal learning emerges as a promising solution for high-precision localization, a cornerstone of 6G integrated sensing and communications (ISAC), by integrating measurements from different data sources. Yet its real-world deployment remains challenging because(i)the quality and relevance of different modalities fluctuate with frequency, noise, and antenna heterogeneity and(ii)spatial and fingerprint ambiguities under non-line-of-sight (NLOS) propagation obscure the mapping between channel measurements and positions. To overcome these challenges, we propose a spatial-context-aware dynamicfusion architecture built on the mixture-of-experts (SCADF-MoE) backbone. We first construct a million-scale comprehensive ray-tracing dataset measuring synchronized angle, distance, gain, and channel across diverse carrier frequencies, antenna geometries, and noise levels. A three-stage pre-processing pipeline then clusters neighboring points into short trajectories, enriching data samples with spatial context information. The resulting sequences are fed into SCADF-MoE: first, multimodal soft MoE blocks with learnable routing matrices dynamically fuse heterogeneous inputs according to their modality relevance in different environmental contexts; second, a modality-task MoE formulates position estimation as a multi-objective problem, simultaneously predicting coordinates of neighboring points to leverage their shared spatial correlations. Additionally, we introduce a regularization loss that enforces expert diversity and mitigates gradient conflicts during multi-task optimization. Simulations across three environments (dense-urban, suburban, canyon) and three heterogeneity dimensions (frequency, noise, antenna) demonstrate that SCADF-MoE achieves consistent sub-meter accuracy in all conditions, reducing overall MSE by 63%, and cuts unseen-NLOS error by 55% compared to state-of-the-art methods. To the best of our knowledge, this is the first work that leverages large-scale multimodal MoEs for high-precision ISAC localization. Chenwei Wu 0006, Chongwen Huang, Yongliang Shen 0001, Zhaohui Yang 0001, Qianqian Yang 0002, Zhaoyang Zhang 0001, Sami Muhaidat, Chau Yuen |
IEEE J. Sel. Areas Commun. | 6 |
| 2026 | Implicit Neural Compression of Point CloudsabstractPoint clouds have gained prominence across numerous applications due to their ability to accurately represent 3D objects and scenes. However, efficiently compressing unstructured, high-precision point cloud data remains a significant challenge. In this paper, we propose NeRC ${}^{\textbf {3}}$ , a novel point cloud compression framework that leverages implicit neural representations (INRs) to encode both geometry and attributes of dense point clouds. Our approach employs two coordinate-based neural networks: one maps spatial coordinates to voxel occupancy, while the other maps occupied voxels to their attributes, thereby implicitly representing the geometry and attributes of a voxelized point cloud. The encoder quantizes and compresses network parameters alongside auxiliary information required for reconstruction, while the decoder reconstructs the original point cloud by inputting voxel coordinates into the neural networks. Furthermore, we extend our method to dynamic point cloud compression through techniques that reduce temporal redundancy, including a 4D spatio-temporal representation termed 4D-NeRC ${}^{\textbf {3}}$ . Experimental results validate the effectiveness of our approach: For static point clouds, NeRC ${}^{\textbf {3}}$ outperforms octree-based G-PCC standard and existing INR-based methods. For dynamic point clouds, 4D-NeRC ${}^{\textbf {3}}$ achieves superior geometry compression performance compared to the latest G-PCC and V-PCC standards, while matching state-of-the-art learning-based methods. It also demonstrates competitive performance in joint geometry and attribute compression. Hongning Ruan, Yulin Shao, Qianqian Yang 0002, Liang Zhao 0004, Zhaoyang Zhang 0001, Dusit Niyato |
IEEE Trans. Image Process. | 3 |
| 2026 | A Superposition Code-Based Semantic Communication Approach With Quantifiable and Controllable SecurityabstractThis paper addresses the challenge of achieving security in semantic communication (SemCom) over a wiretap channel, where a legitimate receiver coexists with an eavesdropper experiencing a poorer channel condition. Despite previous efforts to secure SemCom against eavesdroppers, guarantee of approximately zero information leakage remains an open issue. In this work, we propose a secure SemCom approach based on superposition code, aiming to provide quantifiable and controllable security for digital SemCom systems. The proposed method employs a double-layered constellation map, where semantic information is associated with satellite constellation points and cloud center constellation points are randomly selected. By carefully allocating power between these two layers of constellation, we ensure that the symbol error probability (SEP) of the eavesdropper when decoding satellite constellation points is nearly equivalent to random guessing, while maintaining a low SEP for the legitimate receiver to successfully decode the semantic information. Simulation results demonstrate that the peak signal-to-noise ratio (PSNR) and mean squared error (MSE) of the eavesdropper's reconstructed data, under the proposed method, can range from decoding Gaussian-distributed random noise to approaching the variance of the data. This validates the effectiveness of our method in nearly achieving the experimental upper bound of security for digital SemCom systems when both eavesdroppers and legitimate users utilize identical decoding schemes. Furthermore, the proposed method consistently outperforms benchmark techniques, showcasing superior data security and robustness against eavesdropping. The implementation code is publicly available at:https://github.com/1weixuanchen/A-Superposition-Code-Based-Semantic-Communication. Weixuan 'Vincent' Chen, Shuo Shao 0001, Qianqian Yang 0002, Zhaoyang Zhang 0001, Ping Zhang 0003 |
IEEE Trans. Mob. Comput. | 3 |
| 2026 | Lightweight Semantic Communication-Compliant Shortest Path Selection in Large-Scale LEO Satellite NetworksabstractEnhanced by inter-satellite links and satellite direct-to-device capabilities, satellite networks can offer low-latency communication globally. However, limited spectrum resources and the capacity bounds of the Shannon's information theory pose fundamental challenges for supporting bandwidth-intensive multimedia services. Semantic communication (SemCom) offers a promising solution by transmitting compressed semantic representations instead of raw data, thereby alleviating bandwidth pressure. However, it also introduces SemCom-related constraints that render conventional schemes such as contact graph routing inapplicable. To overcome this challenge, we investigate SemCom-compliant path selection and formulate it as a non-NP hard mixed-integer linear programming problem. To address the problem, we develop a graph-based scheme that exploits the special structure of the solution space, the sparsity of SemCom-capable satellites, and the property of Dijkstra's algorithm, thus achieving optimal solutions with polynomial-time complexity. Simulation results on the Starlink constellation confirm that the proposed scheme facilitates SemCom with negligible computational overhead and significant bandwidth reduction. While the bandwidth reduction comes at the cost of increased delay and path hops, these effects are shown to be mitigatable through higher SemCom deployment in a satellite network or by enabling semantic processing at the user side. Binquan Guo, Zehui Xiong, Zhou Zhang 0004, Qianqian Yang 0002, Dusit Niyato, Mohsen Guizani, Zhu Han 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2026 | Adaptive Semantic Compression and Transmission With Joint Resource Allocation Optimization for Multi-User Image ClassificationabstractTask-oriented semantic communication, leveraging learning-based joint source-channel coding (JSCC), has emerged as a key paradigm for low-latency, high-precision edge-assisted Internet of Things systems. However, the direct mapping of source data to continuous channel symbols in JSCC poses a great challenge in compatibility with existing digital systems. To address this, we propose a digital semantic communication scheme, i.e., an AdaptiveSemanticCompression with jointResourceAllocation andModulation (Adaptive-SCRAM) optimization scheme for multi-user image classification. This scheme, with the semantics quantized by a compressed codebook, enables the discrete semantic transmission with adaptive modulation, while achieving high accuracy and low latency with transmission resources optimized in multi-user classification task. Specifically, we first design a vector quantized-variational autoencoder-based digital JSCC framework with regional quantization, by jointly maximizing the semantic entropy and minimizing the codebook training loss with various SNRs and modulation orders considered in Rayleigh fading. Then based on the well trained end-to-end architecture, we mathematically fit the classification accuracy with respect to the effects of both compressed codebook size and received SNR under different modulation orders, providing an effective premise for the task performance optimization. Finally, we consider to maximize the overall multi-user classification accuracy under the transmission delay constraint, by optimizing the compression, modulation, power and bandwidth allocation for each user. To address the highly non-convex issue, we develop a dual-layer optimization algorithm. The outer-layer problem, which optimizes the compressed codebook size and modulation order, is solved by a cross-entropy-based learning algorithm. While for the inner-layer problem, a successive convex approximation method is used to optimize the power and bandwidth allocation. Simulation results show that our JSCC framework significantly reduces the semantic codebook size without compromising the classification accuracy, which is applicable to practical digital transmission systems. More importantly, compared to most existing comparable optimization schemes for image classification, our Adaptive-SCRAM optimization scheme with adaptive compression, modulation, and resource allocation can achieve much higher classification accuracy for multi-user tasks, while guaranteeing the transmission efficiency. Qian Wang 0030, Jiaqi Ye, Li Ping Qian 0001, Wei Jiang 0020, Qianqian Yang 0002, Ying-Chang Liang, Pooi Yuen Kam |
IEEE Trans. Mob. Comput. | 5 |
| 2025 | Adaptive Model Partitioning for Distributed LLM Inference Across Heterogeneous DevicesabstractLarge language models (LLMs) deliver strong performance across diverse tasks but their growing computational demands make single-device inference increasingly infeasible. Distributed inference on heterogeneous accelerators offers a scalable solution, yet existing partitioning strategies often depend on oversimplified latency models that fail to capture runtime variability caused by device diversity and fluctuating input lengths. We present an adaptive framework that jointly predicts submodel latency and optimally partitions LLMs to minimize end-to-end inference time in heterogeneous clusters. A multilayer perceptron (MLP) regression model integrates multidimensional features—such as submodel masks, device characteristics, token counts, and computational demand—and employs lightweight adapter layers for rapid fine-tuning to unseen environments. Using these latency predictions, we formulate pipeline-aware model partitioning as a dynamic programming problem, reducing search complexity from exponential to polynomial. Experiments on GPT2 show that our approach achieves high prediction accuracy (91.8% within 20% error) and improves adaptability, raising accuracy from 20.3% to 76.0% with only 50 fine-tuning samples. Overall, the proposed method reduces inference latency from 234.70 ms to 130.45 ms, significantly outperforming baseline loadbalancing methods. Junda Wang, Zhaoyang Li 0005, Qianqian Yang 0002 |
EUC | 3 |
| 2025 | CSI-ALM: Enhancing Channel State Information Prediction with Semantically Aligned Large Language Models
Zhaoyang Li 0005, Qianqian Yang 0002, Zhiguo Shi 0001, Zehui Xiong, Tony Q. S. Quek |
GLOBECOM | 2 |
| 2025 | Enabling Training-Free Semantic Communication Systems with Generative Diffusion ModelsabstractSemantic communication (SemCom) has recently emerged as a promising paradigm for next-generation wireless systems. Empowered by advanced artificial intelligence (AI) technologies, SemCom has achieved significant improvements in transmission quality and efficiency. However, existing SemCom systems either rely on training over large datasets and specific channel conditions or suffer from performance degradation under channel noise when operating in a training-free manner. To address these issues, we explore the use of generative diffusion models (GDMs) as training-free SemCom systems. Specifically, we design a semantic encoding and decoding method based on the inversion and sampling process of the denoising diffusion implicit model (DDIM), which introduces a two-stage forward diffusion process, split between the transmitter and receiver to enhance robustness against channel noise. Moreover, we optimize sampling steps to compensate for the increased noise level caused by channel noise. We also conduct a brief analysis to provide insights about this design. Simulations on the Kodak dataset validate that the proposed system outperforms the existing baseline SemCom systems across various metrics. Shunpu Tang, Qianqian Yang 0002, Ruichen Zhang 0001, Jihong Park, Dusit Niyato |
GLOBECOM | 3 |
| 2025 | Confidant: Customizing Transformer-based LLMs via Collaborative Training on Mobile DevicesabstractLarge language models (LLMs) have emerged as a cornerstone for advancing AI technologies. It revolutionizes the way we interact with devices, websites, and information, and paves the way for the development of highly intuitive and capable virtual assistants. Training of today's LLMs happens in cloud data centers due to the requirement of enormous data and a significant amount of computing power. Despite extensive research in mobile edge computing, fine-tuning pre-trained LLMs using resource-constrained devices like commodity smartphones remains highly under-explored. In this paper, we propose Confidant, a practical collaborative training framework that allows modern LLMs to be fine-tuned across multiple off-the-shelf mobile devices. To this end, Confidant partitions an LLM into several sub-models, allowing each of them to fit in the memory of a mobile device. Multiple mobile devices then collaborate to train the LLM by employing a novel pipeline parallel training approach. In specific, Confidant encompasses a memory-aware dynamic model partitioning and intra-device multi-processor scheduler to minimize the training time across heterogeneous platforms. To ensure resilient distributed training, a hybrid fault tolerance mechanism is devised to proactively manage potential device and network failures. We fully implemented Confidant in C++/Python, and built a cross-framework adapter, enabling collaborative training on a variety of mobile platforms. Experimental results show that Confidant excels in achieving computation-, memory-efficient, and robust customization of LLMs - it manages to train state-of-the-art billion-sized LLMs including BERT, GPT-2, Phi2, and LLaMA3, and fine-tunes Phi2-2.7B on Alpaca in just 40.1 hours using three consumer-grade mobile devices. Yuhao Chen 0005, Yuxuan Yan, Shuowei Ge, Yuyang Qin, Qianqian Yang 0002, Shibo He, Zhiguo Shi 0001, Jiming Chen 0001, Yuanchao Shu |
MobiCom | 6 |
| 2025 | Demo: Customizing Transformer-based LLMs via Collaborative Training on Mobile DevicesabstractDespite large language models (LLMs) being an essential part of our lives, training of LLMs still needs to be done in cloud data centers due to the large requirements of data and computing power, leaving fine-tuning pre-trained LLMs on resource-constrained mobile devices remains highly under-explored. In this demo, we present Confidant, a practical collaborative training system that allows modern LLMs to be fine-tuned across multiple off-the-shelf mobile devices. Confidant partitions an LLM into several sub-models, deploying each of them to a mobile device. Multiple mobile devices then collaborate to train the LLM by employing a novel pipeline parallel training approach. Specifically, Confidant encompasses a memory-aware dynamic model partitioning and intra-device multi-processor scheduler to minimize the training time across heterogeneous platforms. A hybrid fault tolerance mechanism is also devised to proactively manage potential device and network failures. By building a cross-framework adapter and fully implementing Confidant on smartphones and laptops, we present the demo of collaborative training on a variety of mobile platforms. Yuhao Chen 0005, Yuxuan Yan, Shuowei Ge, Qianqian Yang 0002, Shibo He, Zhiguo Shi 0001, Jiming Chen 0001, Yuanchao Shu |
MobiCom | 6 |
| 2025 | Cooperative Multi-Modal Semantic Communication Scheme for Semantic Segmentation in Autonomous Driving SystemsabstractIn recent years, multi-modal semantic segmentation in autonomous driving has garnered significant attention due to its effectiveness under challenging lighting conditions. However, current segmentation approaches primarily focus on segmentation techniques without addressing the critical communication challenges inherent in internet of vehicles (IoV). Unlike traditional communications that transmit source data, semantic communications transmit only task-relevant semantic information, significantly reducing data traffic while ensuring the accuracy of task execution. This paper introduces a novel cooperative multi-modal semantic communication framework designed to enhance semantic segmentation in autonomous driving systems. By compressing redundant information and transmitting only essential semantic features, the proposed scheme enables continuous data transmission with drastically reduced data volume. Moreover, this scheme not only improves communication efficiency but also ensures reliable segmentation performance across diverse data modalities. Experimental results validate the effectiveness of proposed scheme, demonstrating its ability to achieve high compression ratios, robust segmentation performance, and re-silience to channel noise under varying lighting conditions. Yunqi Feng 0001, Hesheng Shen, Xiufang Shi, Qianqian Yang 0002 |
WCNC | 4 |
| 2025 | Maximum Likelihood Estimation of Wiener Phase Noise Variance in MPSK Modulated SystemsabstractPhase noise is one of the fundamental impairments in radar, communications, and even the integration of sensing and communications, which is necessary to be suppressed to guarantee the system performance for high-order modulations. In order to obtain precise phase estimation or effectively track phase noise, many estimation algorithms rooted in digital signal processing operate under the premise that the variance of the phase noise is known. However, in practical applications, the receiver side can hardly get the premise knowledge of the phase noise variance. Thus, accurate estimation of the phase noise variance is significantly important for not only carrier recovery, but also performance monitoring. This paper proposes a maximum likelihood (ML)-based Wiener phase noise variance estimation scheme, based on the amplitude and phase-form of the noisy received signal model for$M$-ary phase-shift keying ($M$PSK) modulated systems. Specifically, by making full use of the explicit statistics of the received phase after raising to the Mth power, the closed-form expressions for ML estimation of the incremental phase noise and the Wiener phase noise variance are derived. The estimated mean square error is both theoretically and numerically analyzed to validate the unbiased ML estimator. Numerical results are given to verify the estimation accuracy in terms of varing signal-to-noise ratio and memory length. The proposed ML estimator is demonstrated to have precise estimation performance with low computational complexity. Qian Wang 0030, Xinwei Du, Li Ping Qian 0001, Qianqian Yang 0002, Pooi Yuen Kam |
WCNC | 5 |
| 2025 | Object-Attribute-Relation Representation-Based Video Semantic CommunicationabstractWith the rapid growth of multimedia data volume, there is an increasing need for efficient video transmission in applications such as virtual reality and future video streaming services. Semantic communication is emerging as a vital technique for ensuring efficient and reliable transmission in low-bandwidth, high-noise settings. However, most current approaches focus on joint source-channel coding (JSCC) that depends on end-to-end training. These methods often lack an interpretable semantic representation and struggle with adaptability to various downstream tasks. In this paper, we introduce the use of object-attribute-relation (OAR) as a semantic framework for videos to facilitate low bit-rate coding and enhance the JSCC process for more effective video transmission. We utilize OAR sequences for both low bit-rate representation and generative video reconstruction. Additionally, we incorporate OAR into the image JSCC model to prioritize communication resources for areas more critical to downstream tasks. Our experiments on traffic surveillance video datasets assess the effectiveness of our approach in terms of video transmission performance. The empirical findings demonstrate that our OAR-based video coding method not only outperforms H.265 coding at lower bit-rates but also synergizes with JSCC to deliver robust and efficient video transmission. Qiyuan Du, Yiping Duan, Qianqian Yang 0002, Xiaoming Tao 0001, Mérouane Debbah |
IEEE J. Sel. Areas Commun. | 3 |
| 2025 | Data-Free Cloud-Edge Distillation for Safe and Efficient Intelligent CommunicationsabstractEfficiency and security are the core challenges in the intelligent communication field. Lightweight neural networks have accelerated the information interpretation and communication efficiency, thereby fostering the rapid development of the Internet of Things. Enhancing the recognition capability of lightweight neural networks remains challenging. Knowledge distillation, a technique that transfers knowledge from a complex model to a smaller one, is often used to improve the recognition performance of lightweight networks. However, practical issues such as transmission constraints and user privacy make the original data required for knowledge distillation difficult to access directly. To tackle this issue, this paper proposes a Data-Free Cloud-Edge Knowledge Distillation (DF-CEKD) model, which uses a complex network in the cloud to provide training guidance for lightweight networks deployed on mobile devices. Specifically, DF-CEKD employs a novel Deep Inversion Diffusion Generation (DIDG) module to provide proxy data as input for the distillation process, thereby transferring the feature learning capability from the cloud network to the edge network. Meanwhile, a Multi-Layer Feature Joint Supervision Distillation (MLF-JSD) module is designed to further enhance the feature selection guidance provided by the teacher network in the cloud for training the lightweight student network. The simulation results demonstrate that the proposed DF-CEKD reduces the number of parameters to 1/20 and the floating-point operations to 1/12 when distilling from WRN40-2 to WRN16-1, resulting in only a 0.27% decrease in accuracy. Xiufang Li, Yiping Duan, Xiaoming Tao 0001, Qigong Sun, Qiyuan Du, Qianqian Yang 0002, Tony Q. S. Quek |
IEEE J. Sel. Areas Commun. | 6 |
| 2025 | Capacity Optimizing Resource Allocation in Joint Source-Channel Coding Systems With QoS ConstraintsabstractBenefited from the advances of deep learning (DL) techniques, deep joint source-channel coding (JSCC) has shown its great potential to improve the performance of wireless transmission. However, most of the existing works focus on the DL-based transceiver design of the JSCC model, while ignoring the resource allocation problem in wireless systems. In this paper, we consider a downlink resource allocation problem, where a base station (BS) jointly optimizes the compression ratio (CR) and power allocation as well as resource block (RB) assignment of each user according to the latency and performance constraints to maximize the number of users that successfully receive their requested content with desired quality. To solve this problem, we first decompose it into two subproblems without loss of optimality. The first subproblem is to minimize the required transmission power for each user under given RB allocation. We derive the closed-form expression of the optimal transmit power by searching the maximum feasible compression ratio. The second one aims at maximizing the number of supported users through optimal user-RB pairing, which is solved by utilizing bisection search as well as interior-point algorithm. To reduce the computational complexity, we propose a heuristic greedy algorithm to obtain a simplified problem. Then the Bregman alternating direction method of multipliers (BADMM) based algorithm is adopted to decompose the simplified problem into several subproblems that can be computed in parallel. Simulation results validate the effectiveness of the proposed resource allocation methods in terms of the number of satisfied users with given resources. It is also shown that the BADMM-based algorithm can significantly reduce the computational complexity and retain high performance. Kaiyi Chi, Qianqian Yang 0002, Zhaohui Yang 0001, Yiping Duan, Zhaoyang Zhang 0001 |
IEEE Trans. Commun. | 2 |
| 2024 | Secure Semantic Communication for Image Transmission in the Presence of EavesdroppersabstractSemantic communication (SemCom) has emerged as a key technology for the forthcoming sixth-generation (6G) network, attributed to its enhanced communication efficiency and robustness against channel noise. However, the open nature of wireless channels makes them vulnerable to eavesdropping, which poses a serious threat to privacy. To address this issue, we propose a novel secure semantic communication (SemCom) approach for image transmission, which integrates steganography technology to conceal private information within non-private images (host images). Specifically, we propose an invertible neural network (INN)-based signal steganography approach that embeds channel input signals of a private image into those of a host image before transmission. This ensures that the original private image can be reconstructed from the received signals at the legitimate receiver, while the eavesdropper can only decode the information of the host image. Simulation results demonstrate that the proposed approach maintains comparable reconstruction quality of both host and private images at the legitimate receiver, compared to scenarios without any secure mechanisms. Moreover, the results indicate that the eavesdropper is only able to reconstruct host images, showcasing the enhanced security provided by our approach. Shunpu Tang, Chen Liu 0034, Qianqian Yang 0002, Shibo He, Dusit Niyato |
GLOBECOM | 3 |
| 2024 | Semantic Communication for Efficient Point Cloud TransmissionabstractAs three-dimensional acquisition technologies like LiDAR cameras advance, the need for efficient transmission of 3D point clouds is becoming increasingly important. In this paper, we present a novel semantic communication (SemCom) approach for efficient 3D point cloud transmission. Different from existing methods that rely on downsampling and feature extraction for compression, our approach utilizes a parallel structure to separately extract both global and local information from point clouds. This system is composed of five key components: local semantic encoder, global semantic encoder, channel encoder, channel decoder, and semantic decoder. Our numerical results indicate that this approach surpasses both the traditional Octree compression methodology and alternative deep learning-based strategies in terms of reconstruction quality. Moreover, our system is capable of achieving high-quality point cloud reconstruction under adverse channel conditions, specifically maintaining a reconstruction quality of over 37dB even with severe channel noise. Shangzhuo Xie, Qianqian Yang 0002, Yuyi Sun, Tianxiao Han, Zhaohui Yang 0001, Zhiguo Shi 0001 |
GLOBECOM | 2 |
| 2024 | Robust Continuous-Time Beam Tracking with Liquid Neural NetworkabstractMillimeter-wave (mmWave) technology is increasingly recognized as a pivotal technology of the sixth-generation communication networks due to the large amounts of available spectrum at high frequencies. However, the huge overhead associated with beam training imposes a significant challenge in mmWave communications, particularly in urban environments with high background noise. To reduce this high overhead, we propose a novel solution for robust continuous-time beam tracking with liquid neural network, which dynamically adjust the narrow mmWave beams to ensure real-time beam alignment with mobile users. Through extensive simulations, we validate the effectiveness of our proposed method and demonstrate its superiority over existing state-of-the-art deep-learning-based approaches. Specifically, our scheme achieves at most 46.9% higher normalized spectral efficiency than the baselines when the user is moving at 5 m/s, demonstrating the potential of liquid neural networks to enhance mmWave mobile communication performance. Fenghao Zhu, Xinquan Wang, Chongwen Huang, Richeng Jin, Qianqian Yang 0002, Ahmed Al Hammadi, Zhaoyang Zhang 0001, Chau Yuen, Mérouane Debbah |
GLOBECOM | 5 |
| 2024 | Latency-minimizing Semantic Communication with Dynamic Model PartitioningabstractSemantic communication is an emerging communication approach that aims to enhance efficient transmission by conveying the essential semantic meaning of the information while eliminating redundancy. In the current deep learning (DL)-based semantic communication systems, the encoder and decoder at the sender and receiver persist without modification after deployment, irrespective of variations in device computing power and channel bandwidth. This lack of adaptability may result in a decline in performance. To overcome this issue, we introduce an adaptive semantic communication approach aimed at minimizing end-to-end latency by leveraging a dynamic model partitioning mechanism. This mechanism dynamically splits the overall model into the encoder and decoder components, with the partitioning points adapting to changing communication and computing resources. Furthermore, we present a training method referred to as scheduled random partition point training to ensure that changes in the partitioning points do not adversely impact the performance of downstream tasks. Our experimental results affirm the effectiveness of these methods in terms of reducing latency and improving task performance. Yuxuan Yan, Yuhao Chen 0005, Qianqian Yang 0002, Zhiguo Shi 0001 |
ICC | 3 |
| 2024 | Pilot-Free Semantic Communication Over Multi-User Mimo Fading ChannelsabstractWireless communication systems operating in fading channels often demand pilots for channel estimation and data recovery, leading to substantial transmission overhead. In this paper, we propose a novel pilot-free semantic communication system designed for transmitting images over multi-user MIMO (MU-MIMO) fading channels. Specifically, our method involves extracting multi-scale semantic features from the source image at the transmitter, effectively embedding pilot-like information. At the receiver, we extract channel features from these semantic features at each scale, enabling the reconstruction of the source image without requiring explicit channel estimation and signal detection. To enhance the image reconstruction process, we introduce a novel module, called Resnet Transformer, which combines multi-head self-attention (MHSA) with Resnet block. Our experimental results demonstrate that this pilot-free system outperforms existing pilot-aided semantic communication methods in terms of perceptual quality and transmission efficiency. Weixuan 'Vincent' Chen, Qianqian Yang 0002, Zhaohui Yang 0001, Yiping Duan, Zhaoyang Zhang 0001 |
ICIP | 2 |
| 2024 | Evolving Semantic Communication with Generative ModellingabstractLearning-based semantic communication (SemCom) has emerged as a promising solution for the upcoming 6G networks. In this paper, we explore an evolving SemCom system for image transmission, which can continuously adapt and enhance its transmission efficiency by exploiting knowledge accumulated during previous transmissions. Specifically, we propose a novel channel-aware semantic encoder that utilizes a pretrained generative model to extract channel-correlated latent variables consisting of several semantic vectors from the input images, which can be directly transmitted over a noisy channel without further channel coding. Moreover, we introduce a dynamic code construction mechanism that dynamically updates the codebook with transmitted semantic vectors to eliminate the need to transmit similar codes in subsequent transmissions, thus further reducing the communication overhead. Simulation results highlight the evolving performance of the proposed system in terms of transmission efficiency, achieving superior perceptual quality with an average bandwidth compression ratio (BCR) of $1 / 192$ for a sequence of 100 test images compared to DeepJSCC and InverseJSCC. Code used in this paper is available at https://github.com/recusant7/GAN_SeCom. Shunpu Tang, Qianqian Yang 0002, Deniz Gündüz, Zhaoyang Zhang 0001 |
PIMRC | 2 |
| 2024 | A Joint Communication and Learning Design for Secure Federated Learning with Differential PrivacyabstractIn this paper, the problem of resource allocation for non-orthogonal multiple access (NOMA) enabled secure federated learning (FL) is investigated. In the considered model, a set of users participate in the FL training through transmitting their trained FL model parameters to the base stations (BSs) via NOMA techniques. To prevent data leakage, each user uses the differential privacy (DP) technique through adding Gaussian noise to its FL model parameters. The problem of minimizing overall privacy leakage of all FL participaring users is formulated as an optimization problem through jointly optimizing the connections between users and BSs, transmit power of the users, and the DP noise power. To solve the formulated non-convex optimization problem, a genetic algorithm is proposed to search for feasible solutions in which user connection matrix is taken as gene and the objective function value is taken as the fitness of solution. Simulation results show that the proposed genetic algorithm reduces privacy leakage by up to 73% compared to the conventional alternating optimization algorithm. Licheng Lin, Zhaohui Yang 0001, Qianqian Yang 0002, Mingzhe Chen |
VTC Fall | 3 |
| 2024 | Knowledge-Aided Semantic Communication Leveraging Probabilistic Graphical ModelingabstractIn this paper, we propose a semantic communication approach based on probabilistic graphical model (PGM). The proposed approach involves constructing a PGM from a training dataset, which is then shared as common knowledge between the transmitter and receiver. We evaluate the importance of various semantic features and present a PGM-based compression algorithm designed to eliminate predictable portions of semantic information. Furthermore, we introduce a technique to reconstruct the discarded semantic information at the receiver end, generating approximate results based on the PGM. Simulation results indicate a significant improvement in transmission efficiency over existing methods, while maintaining the quality of the transmitted images. Haowen Wan, Qianqian Yang 0002, Jiancheng Tang, Zhiguo Shi 0001 |
VTC Fall | 2 |
| 2024 | Multi -Sources Information Fusion Learning for Multi-Points NLOS LocalizationabstractAccurate localization of mobile terminals is crucial for integrated sensing and communication systems. Existing fingerprint localization methods, which deduce coordinates from channel information in pre-defined rectangular areas, struggle with the heterogeneous fingerprint distribution inherent in non-line-of-sight (NLOS) scenarios. To address the problem, we introduce a novel multi-source information fusion learning framework referred to as the Autosync Multi-Domain NLOS Localization (AMDNLoc). Specifically, AMDNLoc employs a two-stage matched filter fused with a target tracking algorithm and iterative centroid-based clustering to automatically and irregularly segment NLOS regions, ensuring uniform fingerprint distribution within channel state information across frequency, power, and time-delay domains. Additionally, the framework utilizes a segment-specific linear classifier array, coupled with deep residual network-based feature extraction and fusion, to establish the correlation function between fingerprint features and coordinates within these regions. Simulation results demonstrate that AMDNLoc significantly enhances localization accuracy by over 40% compared with traditional convolutional neural networks on the wireless artificial intelligence research dataset. Fenghao Zhu, Mengbing Liu, Chongwen Huang, Qianqian Yang 0002, Ahmed Alhammadi, Zhaoyang Zhang 0001, Mérouane Debbah |
VTC Spring | 5 |
| 2024 | Spectral Efficiency Maximization for Probabilistic Semantic Communication with Rate SplittingabstractIn this paper, the problem of joint transmission and computation resource allocation for probabilistic semantic communication (PSC) network with rate splitting multiple access (RSMA) is investigated. In the considered model, the base station (BS) needs to transmit a large amount of data, which is represented by substantial knowledge graphs, to multiple users. Due to limited communication resource, the BS needs to utilize semantic communication techniques to compress the large-sized data. In this paper, the semantic communication is enabled by shared probability graphs between the BS and users. The process of semantic compression requires computation power at the BS, which has an impact on limited power budget. Therefore, it is necessary to balance the power between transmission and computation. Based on the probability graph, the semantic rate related to semantic compression ratio is first theoretically formulated. Then, the problem is formulated as an optimization problem with the aim of maximizing the sum semantic rate of all users under total power, semantic compression ratio, and rate allocation constraints. To tackle this problem, an iterative algorithm is accordingly proposed to obtain a suboptimal solution. Numerical results validate the effectiveness of the proposed scheme. Zhouxiang Zhao, Zhaohui Yang 0001, Mingzhe Chen, Xu Gan, Chongwen Huang, Yao Sun 0002, Qianqian Yang 0002, Wei Xu 0001, Zhaoyang Zhang 0001 |
VTC Spring | 7 |
| 2024 | A Joint Communication and Computation Design for Distributed RIS-Assisted Probabilistic Semantic Communication in IIoTabstractThe advent of Industry 4.0 has positioned the industrial Internet of Things (IIoT) as a cornerstone of future industry. In this article, the problem of spectral-efficient communication and computation resource allocation for distributed reconfigurable intelligent surfaces (RISs) assisted probabilistic semantic communication (PSC) in IIoT is investigated. In the considered model, multiple RISs are deployed to serve multiple users, while PSC adopts compute-then-transmit protocol to reduce the size of the transmission data. To support the high-rate transmission, the semantic compression ratio, transmit power allocation, and distributed RISs deployment must be jointly considered. This joint communication and computation problem is formulated as an optimization problem whose goal is to maximize the sum semantic-aware transmission rate of the system under the total transmit power, phase shift, RIS-user association, and semantic compression ratio constraints. To solve this problem, a many-to-many matching scheme is proposed to solve the RIS-user association subproblem, the semantic compression ratio subproblem is addressed following the greedy policy, while the phase shift of RIS can be optimized using the tensor-based beamforming. Numerical results verify the superiority of the proposed algorithm. Zhouxiang Zhao, Zhaohui Yang 0001, Chongwen Huang, Li Wei 0007, Qianqian Yang 0002, Caijun Zhong, Wei Xu 0001, Zhaoyang Zhang 0001 |
IEEE Internet Things J. | 5 |
| 2024 | Contrastive Learning-Based Semantic CommunicationsabstractRecently, there has been a growing interest in learning-based semantic communication because it can prioritize the preservation of meaningful semantic information over the accuracy of the transmitted symbols, resulting in improved communication efficiency. However, existing learning-based approaches still face limitations in defining semantic level loss and often struggle to find a good trade-off between preserving semantic information and preserving intricate details. In addition, the existing semantic communication approaches cannot effectively train semantic encoders and decoders without the support of downstream models. To address these limitations, this paper proposes a contrastive learning (CL)-based semantic communication system. First, inspired by practical observations, we introduce the concept of semantic contrastive loss and propose a semantic contrastive coding (SemCC) approach that treats data corruption during transmission as a form of data augmentation within the CL framework. Moreover, we propose a semantic re-encoding (SemRE) operation, which uses a duplicate of the semantic encoder deployed at the receiver to guide the entire training process when the downstream model is inaccessible. Further, we design the training procedure for SemCC and SemRE approaches, respectively, to balance the semantic information and intricate details. Finally, simulations are performed to demonstrate the superiority of the proposed approaches over competing approaches. In particular, our approaches achieve a significant accuracy improvement of up to 53% on the CIFAR-10 dataset with a bandwidth compression ratio of 1/24, and also obtain comparable image reconstruction quality as the bandwidth compression ratio is improved. Shunpu Tang, Qianqian Yang 0002, Lisheng Fan, Xianfu Lei, Arumugam Nallanathan, George K. Karagiannidis |
IEEE Trans. Commun. | 2 |
| 2024 | FTPipeHD: A Fault-Tolerant Pipeline-Parallel Distributed Training Approach for Heterogeneous Edge DevicesabstractWith the increasing proliferation of Internet-of-Things (IoT) devices, there is a growing trend towards distributing the power of deep learning (DL) among edge devices rather than centralizing it at the cloud. To deploy deep and complex models at edge devices with limited resources, model partitioning of deep neural network (DNN) models has been widely studied. However, most of the existing literature only considers distributing the inference model while still training the model at the cloud. In this paper, we propose FTPipeHD, a novel DNN training approach that trains DNN models across distributed heterogeneous devices with the fault-tolerance mechanism. To accelerate the training with the time-varying computing power of each device, we optimize the partition points dynamically according to real-time computing capacities. We also propose a novel weight redistribution approach that replicates the weights to both the neighboring nodes and the central node periodically, which combats the failure of multiple devices during training while incurring limited communication costs. Our numerical results demonstrate that FTPipeHD is 6.8 times faster in training than the state-of-the-art method when the computing capacity of the best device is 10 times greater than the worst one. It is also shown that the proposed method is able to accelerate the training even with the existence of device failures. Yuhao Chen 0005, Qianqian Yang 0002, Shibo He, Zhiguo Shi 0001, Jiming Chen 0001, Mohsen Guizani |
IEEE Trans. Mob. Comput. | 2 |
| 2024 | AccEPT: An Acceleration Scheme for Speeding up Edge Pipeline-Parallel TrainingabstractIt is usually infeasible to fit and train an entire large deep neural network (DNN) model using a single edge device due to the limited resources. To facilitate intelligent applications across edge devices, researchers have proposed partitioning a large model into several sub-models, and deploying each of them to a different edge device to collaboratively train a DNN model. However, the communication overhead caused by the large amount of data transmitted from one device to another during training, as well as the sub-optimal partition point due to the inaccurate latency prediction of computation at each edge device can significantly slow down training. In this paper, we propose AccEPT, an acceleration scheme for accelerating the edge collaborative pipeline-parallel training. In particular, we propose a light-weight adaptive latency predictor to accurately estimate the computation latency of each layer at different devices, which also adapts to unseen devices through continuous learning. Therefore, the proposed latency predictor leads to better model partitioning which balances the computation loads across participating devices. Moreover, we propose a bit-level computation-efficient data compression scheme to compress the data to be transmitted between devices during training. Our numerical results demonstrate that our proposed acceleration approach is able to significantly speed up edge pipeline parallel training up to 3 times faster in the considered experimental settings Yuhao Chen 0005, Yuxuan Yan, Qianqian Yang 0002, Yuanchao Shu, Shibo He, Zhiguo Shi 0001, Jiming Chen 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2024 | Soft Actor-Critic-Based Multi-User Multi-TTI MIMO Precoding in Multi-Modal Real-Time Broadband CommunicationsabstractThe next-generation wireless network is envisioned to support real-time broadband communication (RTBC) to provision services for immersive applications. Such applications (e.g., virtual reality, VR) usually need to simultaneously transmit multi-modal (e.g., visual, audio and haptic) data streams that have different traffic characteristics and transmission requirements, within multiple transmission time intervals (TTIs). In this paper, we formulate an optimization problem of multi-user multiple-input multiple-output (MIMO) precoding within multiple TTIs for multi-modal data transmission. As it is hard to find an optimal solution, we first resort to a novel soft actor-critic (SAC)-based learning approach. Specifically, a lightweight reinforcement learning architecture is employed to learn the adaptive priority weight of each user within multiple TTIs by taking into account its remaining multi-modal data amount and dynamic interaction state. The learned priority weights are then input to an iterative weighted minimum mean-square error (WMMSE) algorithm to adjust the precoder matrix and user transmission rates. With a scalable state design, the proposed algorithm can be tailored to different numbers of potential or active users. We also provide another practical solution to the formulated multi-TTI precoding problem, which transforms the problem into a single-TTI optimization problem by adding the quality-of-service (QoS) constraints into the traditional WMMSE problem and then solves it using the alternating direction method of multipliers (ADMM). Simulation results demonstrate the robustness and efficiency of the proposed algorithms, which show that the SAC-based precoding algorithm can achieve a 50.0% increment in system capacity compared to traditional WMMSE and a significant reduction in time complexity compared to the QoS-constrained WMMSE algorithm. Yingzhi Huang, Kaiyi Chi, Qianqian Yang 0002, Zhaohui Yang 0001, Zhaoyang Zhang 0001 |
IEEE Trans. Wirel. Commun. | 3 |
| 2023 | Deep Joint Source-Channel Coding for Wireless Image Transmission with Entropy-Aware Adaptive Rate ControlabstractAdaptive rate control for deep joint source and channel coding (JSCC) is considered as an effective approach to transmit sufficient information in scenarios with limited communication resources. We propose a deep JSCC scheme for wireless image transmission with entropy-aware adaptive rate control, using a single deep neural network to support multiple rates and automatically adjust the rate based on the feature maps of the input image and their entropy, as well as the channel conditions. In particular, we maximize the entropy of the feature maps to increase the average information carried by each transmitted symbol during the training. We further decide which feature maps should be activated based on their entropy, which improves the efficiency of the transmitted symbols. We also propose a pruning module to remove less important pixels in the activated feature maps in order to further improve transmission efficiency. The experimental results demonstrate that our proposed scheme learns an effective rate control strategy that reduces the required channel bandwidth while preserving the quality of the reconstructed images. Weixuan 'Vincent' Chen, Yuhao Chen 0005, Qianqian Yang 0002, Chongwen Huang, Qian Wang 0030, Zhaoyang Zhang 0001 |
GLOBECOM | 3 |
| 2023 | The Model Inversion Eavesdropping Attack in Semantic Communication SystemsabstractIn recent years, semantic communication has been a popular research topic for its superiority in communication efficiency. As semantic communication relies on deep learning to extract meaning from raw messages, it is vulnerable to attacks targeting deep learning models. In this paper, we introduce the model inversion eavesdropping attack (MIEA) to reveal the risk of privacy leaks in the semantic communication system. In MIEA, the attacker first eavesdrops the signal being transmitted by the semantic communication system and then performs model inversion attack to reconstruct the raw message, where both the white-box and black-box settings are considered. Evaluation results show that MIEA can successfully reconstruct the raw message with good quality under different channel conditions. We then propose a defense method based on random permutation and substitution to defend against MIEA in order to achieve secure semantic communication. Our experimental results demonstrate the effectiveness of the proposed defense method in preventing MIEA. Yuhao Chen 0005, Qianqian Yang 0002, Zhiguo Shi 0001, Jiming Chen 0001 |
GLOBECOM | 2 |
| 2023 | MIMO Precoding Design with QoS and Per-Antenna Power ConstraintsabstractPrecoding design for the downlink of multiuser multiple-input multiple-output (MU-MIMO) systems is a fundamental problem. In this paper, we aim to maximize the weighted sum rate (WSR) while considering both quality-of-service (QoS) constraints of each user and per-antenna power constraints (PAPCs) in the downlink MU-MIMO system. To solve the problem, we reformulate the original problem to an equivalent problem by using the well-known weighted minimal mean square error (WMMSE) framework, which can be tackled by iteratively solving three subproblems. Since the precoding matrices are coupled among the QoS constraints and PAPCs, we adopt alternating direction method of multipliers (ADMM) to obtain a distributed solution. Simulation results validate the effectiveness of the proposed algorithm. Kaiyi Chi, Yingzhi Huang, Qianqian Yang 0002, Zhaohui Yang 0001, Zhaoyang Zhang 0001 |
GLOBECOM | 3 |
| 2023 | Semantic-aware Transmission for Robust Point Cloud ClassificationabstractAs three-dimensional (3D) data acquisition devices become increasingly prevalent, the demand for 3D point cloud transmission is growing. In this study, we introduce a semanticaware communication system for robust point cloud classification that capitalizes on the advantages of pre-trained Point-BERT models. Our proposed method comprises four main components: the semantic encoder, channel encoder, channel decoder, and semantic decoder. By employing a two-stage training strategy, our system facilitates efficient and adaptable learning tailored to the specific classification tasks. The results show that the proposed system achieves classification accuracy of over 89% when SNR is higher than 10 dB and still maintains accuracy above 66.6% even at SNR of 4 dB. Compared to the existing method, our approach performs at 0.8% to 48% better across different SNR values, demonstrating robustness to channel noise. Our system also achieves a balance between accuracy and speed, being computationally efficient while maintaining high classification performance under noisy channel conditions. This adaptable and resilient approach holds considerable promise for a wide array of 3D scene understanding applications, effectively addressing the challenges posed by channel noise. Tianxiao Han, Kaiyi Chi, Qianqian Yang 0002, Zhiguo Shi 0001 |
GLOBECOM | 3 |
| 2023 | Soft Actor-Critic-Based Multi-TTI Precoding for Multi-Modal RTBC Over MIMO SystemsabstractThe 5.5th Generation is envisioned to support the Real-time Broadband Communication (RTBC) scenarios, which needs to satisfy the enhanced Mobile Broadband(eMBB) and Ultra-Reliable Low-Latency Communication (uRLLC) services requirements simultaneously. As an essential technique to improve the capacity of systems, pre coding in RTBC faces the challenge of finding the optimal solution over long-term transmission with multimodal streams for the system. To meet these requirements, we propose a soft actor-critic (SAC) based multiple transmission time interval (TTl) intelligent precoding algorithm that optimizes the multi-user precoding scheme by learning the priority weight of each user in the iterative weighted minimum mean-square error (MMSE) algorithm. Considering the remaining multimodal data and dynamic activation state of real-time interaction users, we design a novel and lightweight reinforcement learning architecture scalable to different numbers of potential or active users. Simulation results demonstrate the robustness and superiority of our precoding algorithm, which achieves 50% performance improvement in system capacity compared to the weighted MMSE algorithm. Yingzhi Huang, Kaiyi Chi, Qianqian Yang 0002, Zhaohui Yang 0001, Zhaoyang Zhang 0001 |
GLOBECOM | 3 |
| 2023 | Generative Model based Highly Efficient Semantic Communication Approach for Image TransmissionabstractDeep learning (DL) based semantic communication methods have been explored to transmit images efficiently in recent years. In this paper, we propose a generative model based semantic communication to further improve the efficiency of image transmission and protect private information. In particular, the transmitter extracts the interpretable latent representation from the original image by a generative model exploiting the GAN inversion method. We also employ a privacy filter and a knowledge base to erase private information and replace it with natural features in the knowledge base. The simulation results indicate that our proposed method achieves comparable quality of received images while significantly reducing communication costs compared to the existing methods. Tianxiao Han, Jiancheng Tang, Qianqian Yang 0002, Yiping Duan, Zhaoyang Zhang 0001, Zhiguo Shi 0001 |
ICASSP | 3 |
| 2023 | Resource Allocation for Capacity Optimization in Joint Source-Channel Coding SystemsabstractBenefited from the advances of deep learning (DL) techniques, deep joint source-channel coding (JSCC) has shown its great potential to improve the performance of wireless transmission. However, most of the existing works focus on the DL-based transceiver design of the JSCC model, while ignoring the resource allocation problem in wireless systems. In this paper, we consider a downlink resource allocation problem, where a base station (BS) jointly optimizes the compression ratio (CR) and power allocation as well as resource block (RB) assignment of each user according to the latency and performance constraints to maximize the number of users that successfully receive their requested content with desired quality. To solve this problem, we first decompose it into two subproblems without loss of optimality. The first subproblem is to minimize the required transmission power for each user under given RB allocation. We derive the closed-form expression of the optimal transmit power by searching the maximum feasible compression ratio. The second one aims at maximizing the number of supported users through optimal user-RB pairing, which we solve by utilizing bisection search as well as Karmarkar's algorithm. Simulation results validate the effectiveness of the proposed resource allocation method in terms of the number of satisfied users with given resources. Kaiyi Chi, Qianqian Yang 0002, Zhaohui Yang 0001, Yiping Duan, Zhaoyang Zhang 0001 |
ICC | 2 |
| 2023 | Efficient Pruning Method for Learned Lossy Image Compression Models Based on Side InformationabstractIn recent years, deep learning-based lossy image compression have achieved great success. However, the problem of their huge overhead in terms of computational and parametric costs has still not been adequately addressed. Inspired by the classical image compression methods, deep learning based models are usually combined with an entropy model to maintain the compression performance. Existing methods also introduce side information to serve as a prior on the parameters of the entropy model, which have achieved better rate-distortion performance. Based on the role of side information in learned image compression models, we propose an efficient pruning method for such models. In particular, the proposed pruning approach automatically searches for the optimal decoder architecture based on the extent to which each hidden layer in the decoder utilizes side information. The experiment results demonstrate the effectiveness of the proposed method and show that it outperforms all existing related studies in terms of compression performance. Weixuan 'Vincent' Chen, Qianqian Yang 0002 |
ICIP | 2 |
| 2023 | Task-Oriented Communication with Reliability-Driven Retransmission RequestabstractThe advanced Deep Learning (DL) techniques have enabled the development of semantic communication systems by their remarkable information processing and end - to-end optimization capabilities. However, the lack of performance guarantee of these DL-based methods also brings concerns on the reliability of semantic communication systems. To address this issue, we propose a semantic communication scheme with reliability-driven retransmission requests in order to improve the transmission efficiency and guarantee the reliability of the inference result at the same time. In particular, the transmitter first sends a basic amount of the information, and the receiver infers with this information, and assesses the reliability of the results. If the derived reliability is below a given threshold, a retransmission request is sent to the sender for more information to be transmitted. We exploit the entropy of the output logits to quantify the reliability of the classification results. More specifically, lower entropy corresponds to higher confidence, indicating higher reliability. Experimental results validate the effectiveness of the proposed scheme in terms of improving transmission efficiency and computational efficiency while maintaining the reliability of the inference results. Ziheng Ding, Qianqian Yang 0002, Zhaoyang Zhang 0001 |
ICNP | 2 |
| 2023 | Coded Parallelism for Distributed Deep LearningabstractWith the rapid development of deep learning, the parameters of modern neural network models, especially in the field of Natural Language Processing (NLP) are extremely huge. When the parameters of the model are larger even than the storage memory of a single device, it is necessary to split the original big learning model into different parts with each part assigned to one device, thus realizing joint model training over different devices (i.e., distributed training). In this paper, we aim to introduce the advanced coding scheme into the distributed parallel framework, which leads to the perfect combination of coding and the underlying calculation of neural networks. The proposed scheme is not only able to avoid the impact of poor computing power or low bandwidth and even dropped devices (stragglers) on system performance but also reduce the communication load between different devices, thereby greatly improving the performance of distributed parallel systems. Songting Ji, Zhaoyang Zhang 0001, Zhaohui Yang 0001, Richeng Jin, Qianqian Yang 0002 |
ISIT | 5 |
| 2023 | Contrastive Learning based Semantic Communication for Wireless Image TransmissionabstractRecently, semantic communication has been widely applied in wireless image transmission systems as it can prioritize the preservation of meaningful semantic information in images over the accuracy of transmitted symbols, leading to improved communication efficiency. However, existing semantic communication approaches still face limitations in achieving considerable inference performance in downstream AI tasks like image recognition, or balancing the inference performance with the quality of the reconstructed image at the receiver. Therefore, this paper proposes a contrastive learning (CL)-based semantic communication approach to overcome these limitations. Specifically, we regard the image corruption during transmission as a form of data augmentation in CL and leverage CL to reduce the semantic distance between the original and the corrupted reconstruction while maintaining the semantic distance among irrelevant images for better discrimination in downstream tasks. Moreover, we design a two-stage training procedure and the corresponding loss functions for jointly optimizing the semantic encoder and decoder to achieve a good trade-off between the performance of image recognition in the downstream task and reconstructed quality. Simulations are finally conducted to demonstrate the superiority of the proposed method over the competitive approaches. In particular, the proposed method can achieve up to 56% accuracy gain on the CIFAR10 dataset when the bandwidth compression ratio is 1/48. Shunpu Tang, Qianqian Yang 0002, Lisheng Fan, Xianfu Lei, Yansha Deng, Arumugam Nallanathan |
VTC Fall | 2 |
| 2023 | Deep Learning Enabled Semantic Communication Systems for Video TransmissionabstractSemantic communication has emerged as a promising approach for improving efficient transmission in the next generation of wireless networks. Inspired by the success of semantic communication in different areas, we aim to provide a new semantic communication scheme from the semantic level. In this paper, we propose a novel DL-based semantic communication system for video transmission, which compacts semantic-related information to improve transmission efficiency. In particular, we utilize the Bi-optical flow to estimate residual information of inter-frame details. We also propose a feature choice module and a feature fusion module to drop semantically redundant features while paying more attention to the important semantic-related content. We employ a frame prediction module to reconstruct semantic features of the prediction frame from the received signal at the receiver. To enhance the system’s robustness, we propose a noise attention module that assigns different importance weights to the extracted features. Simulation results indicate that our proposed method outperforms existing approaches in terms of transmission efficiency, achieving about 33.3% reduction in the number of transmitted symbols while improving the peak signal-to-noise ratio (PSNR) performance by an average of 0.56dB. Qianqian Yang 0002, Shibo He, Jiming Chen 0001 |
VTC Fall | 2 |
| 2023 | Semantic Communication with Probability Graph: A Joint Communication and Computation DesignabstractIn this paper, we present a probability graph-based semantic information compression system for scenarios where the base station (BS) and the user share common background knowledge. We employ probability graphs to represent the shared knowledge between the communicating parties. During the transmission of specific text data, the BS first extracts semantic information from the text, which is represented by a knowledge graph. Subsequently, the BS omits certain relational information based on the shared probability graph to reduce the data size. Upon receiving the compressed semantic data, the user can automatically restore missing information using the shared probability graph and predefined rules. This approach brings additional computational resource consumption while effectively reducing communication resource consumption. Considering the limitations of wireless resources, we address the problem of joint communication and computation resource allocation design, aiming at minimizing the total communication and computation energy consumption of the network while adhering to latency, transmit power, and semantic constraints. Simulation results demonstrate the effectiveness of the proposed system. Zhouxiang Zhao, Zhaohui Yang 0001, Quoc-Viet Pham, Qianqian Yang 0002, Zhaoyang Zhang 0001 |
VTC Fall | 4 |
| 2023 | Semantic-Preserved Communication System for Highly Efficient Speech TransmissionabstractDeep learning (DL) based semantic communication methods have been explored for the efficient transmission of images, text, and speech in recent years. In contrast to traditional wireless communication methods that focus on the transmission of abstract symbols, semantic communication approaches attempt to achieve better transmission efficiency by only sending the semantic-related information of the source data. In this paper, we consider semantic-oriented speech transmission which transmits only the semantic-relevant information over the channel for the speech recognition task, and a compact additional set of semantic-irrelevant information for the speech reconstruction task. We propose a novel end-to-end DL-based transceiver which extracts and encodes the semantic information from the input speech spectrums at the transmitter and outputs the corresponding transcriptions from the decoded semantic information at the receiver. In particular, we employ a soft alignment module and a redundancy removal module to extract only the text-related semantic features while dropping semantically redundant content, greatly reducing the amount of semantic redundancy compared to existing methods. We also propose a semantic correction module to further correct the predicted transcription with semantic knowledge by leveraging a pretrained language model. For the speech to speech transmission, we further include a CTC alignment module that extracts a small number of additional semantic-irrelevant but speech-related information, such as duration, pitch, power and speaker identification of the speech for the better reconstruction of the original speech signals at the receiver. We also introduce a two-stage training scheme which speeds up the training of the proposed DL model. The simulation results confirm that our proposed method outperforms current methods in terms of the accuracy of the predicted text for the speech to text transmission and the quality of the recovered speech signals for the speech to speech transmission, and significantly improves transmission efficiency. More specifically, the proposed method only sends 16% of the amount of the transmitted symbols required by the existing methods while achieving about a 10% reduction in WER for the speech to text transmission. For the speech to speech transmission, it results in an even more remarkable improvement in terms of transmission efficiency with only 0.2% of the amount of the transmitted symbols required by the existing method while preserving the comparable quality of the reconstructed speech signals. Tianxiao Han, Qianqian Yang 0002, Zhiguo Shi 0001, Shibo He, Zhaoyang Zhang 0001 |
IEEE J. Sel. Areas Commun. | 2 |
| 2023 | AIGAN: Attention-encoding Integrated Generative Adversarial Network for the reconstruction of low-dose CT and low-dose PET images
Yu Fu 0008, Shunjie Dong, Meng Niu, Le Xue, Hanning Guo, Yanyan Huang, Yuanfan Xu, Tianbai Yu, Kuangyu Shi, Qianqian Yang 0002, Yiyu Shi 0001, Cheng Zhuo |
Medical Image Anal. | 10 |
| 2023 | Partial Unbalanced Feature Transport for Cross-Modality Cardiac Image SegmentationabstractDeep learning based approaches have achieved great success on the automatic cardiac image segmentation task. However, the achieved segmentation performance remains limited due to the significant difference across image domains, which is referred to as domain shift. Unsupervised domain adaptation (UDA), as a promising method to mitigate this effect, trains a model to reduce the domain discrepancy between the source (with labels) and the target (without labels) domains in a common latent feature space. In this work, we propose a novel framework, named Partial Unbalanced Feature Transport (PUFT), for cross-modality cardiac image segmentation. Our model facilities UDA leveraging two Continuous Normalizing Flow-based Variational Auto-Encoders (CNF-VAE) and a Partial Unbalanced Optimal Transport (PUOT) strategy. Instead of directly using VAE for UDA in previous works where the latent features from both domains are approximated by a parameterized variational form, we introduce continuous normalizing flows (CNF) into the extended VAE to estimate the probabilistic posterior and alleviate the inference bias. To remove the remaining domain shift, PUOT exploits the label information in the source domain to constrain the OT plan and extracts structural information of both domains, which are often neglected in classical OT for UDA. We evaluate our proposed model on two cardiac datasets and an abdominal dataset. The experimental results demonstrate that PUFT achieves superior performance compared with state-of-the-art segmentation methods for most structural segmentation. Shunjie Dong, Zixuan Pan, Yu Fu 0008, Dongwei Xu, Kuangyu Shi, Qianqian Yang 0002, Yiyu Shi 0001, Cheng Zhuo |
IEEE Trans. Medical Imaging | 6 |
| 2022 | Blind Channel Estimation for MIMO Systems via Variational InferenceabstractIn this paper, we investigate the blind channel estimation problem for MIMO systems under Rayleigh fading channel. Conventional MIMO communication techniques require transmitting a considerable amount of training symbols as pilots in each data block to obtain the channel state information (CSI) such that the transmitted signals can be successfully recovered. However, the pilot overhead and contamination become a bottleneck for the practical application of MIMO systems with the increase of the number of antennas. To overcome this obstacle, we propose a blind channel estimation framework, where we introduce an auxiliary posterior distribution of CSI and the transmitted signals given the received signals to derive a lower bound to the intractable likelihood function of the received signal. Meanwhile, we generate this auxiliary distribution by a neural network based variational inference framework, which is trained by maximizing the lower bound. The optimal auxiliary distribution which approaches real prior distribution is then leveraged to obtain the maximum a posterior (MAP) estimation of channel matrix and transmitted data. The simulation results demonstrate that the performance of the proposed blind channel estimation method closely approaches that of the conventional pilot-aided methods in terms of the channel estimation error and symbol error rate (SER) of the detected signals even without the help of pilots. Jiancheng Tang, Qianqian Yang 0002, Zhaoyang Zhang 0001 |
ICC | 2 |
| 2022 | Self-Attention DDPG for Multi-Beam Combining in mmWave MIMO SystemsabstractIn this paper, we aim at an efficient multi-beam combining design with only requiring receive power measurements for a millimeter-wave (mmWave) multi-input multi-output (MIMO) communication system. A spectrum efficiency maximization problem is formulated with both beam selection and power constraints. To solve this problem, a reinforcement learning (RL)-based multi-beam combining algorithm is proposed. In particular, a self-attention deep deterministic policy gradient (DDPG) scheme is used to adaptively learn the serving beam sets and the corresponding combining weights without any channel state information (CSI). Moreover, the transformer is integrated into the DDPG to precisely capture the signal directions and relevant strengths. Experimental results show the effectiveness of the proposed learning structure in terms of system achievable rate, convergence, and network robustness. Yingzhi Huang, Zhaoyang Zhang 0001, Zhaohui Yang 0001, Qianqian Yang 0002 |
PIMRC | 4 |
| 2022 | Performance Optimization of Energy Efficient Semantic Communications over Wireless NetworksabstractIn this paper, the problem of wireless resource allocation and semantic information extraction for energy efficient semantic communications over wireless networks is investigated. In the considered model, each user first extracts the semantic information from its large-scale data, and then transmits the small-sized semantic information to the base station (BS) which recovers the original data. Due to the limited energy budget of wireless users, both local computational energy and transmission energy must be considered. This joint computation and communication problem is formulated as an optimization problem whose goal is to minimize the total energy consumption of the network under a latency constraint. To solve this problem, an iterative algorithm is proposed where the optimal solution for joint bandwidth allocation, power control, and computation frequency optimization problem can be obtained. Numerical results show the effectiveness of the proposed algorithm. Zhaohui Yang 0001, Mingzhe Chen, Zhaoyang Zhang 0001, Chongwen Huang, Qianqian Yang 0002 |
VTC Fall | 5 |
| 2022 | Semantic Communication Approach for Multi-Task Image TransmissionabstractThis paper presents a deep learning-based image features extraction and compression for multi-tasks, which can be applied to various intelligent tasks. We explore the multilevel features of the source by designing different information extraction networks, which contain text semantics, image segmentation, and pixel information. We propose a coarse-to-fine architecture to excavate the plentiful semantic information received from the encoder. The coarse module recovers the multi-granularity image according to the receiving symbols, and the fine module fuse different quality image to improve the reconstruction performance. In particular, we use multi-attention networks to extract and recover the image features at pixel levels. To overcome the artifact blocks phenomenon during the image reconstruction process that lacks necessary information, we design a dual features block that can mitigate the problem. Meanwhile, the system can accomplish different tasks by changing the last layers of the model. Qianqian Yang 0002, Shibo He, Zhiguo Shi 0001 |
VTC Fall | 2 |
| 2022 | DeU-Net 2.0: Enhanced deformable U-Net for 3D cardiac cine MRI segmentation
Shunjie Dong, Zixuan Pan, Yu Fu 0008, Qianqian Yang 0002, Yuanxue Gao, Tianbai Yu, Yiyu Shi 0001, Cheng Zhuo |
Medical Image Anal. | 4 |
| 2022 | Asynchronous Federated Learning Over Wireless Communication NetworksabstractThe conventional federated learning (FL) framework usually assumes synchronous reception and fusion of all the local models at the central aggregator and synchronous updating and training of the global model at all the agents as well. However, in a wireless network, due to limited radio resource, inevitable transmission failures and heterogeneous computing capacity, it is very hard to realize strict synchronization among all the involved user equipments (UEs). In this paper, we propose a novel asynchronous FL framework, which well adapts to the heterogeneity of users, communication environments and learning tasks, by considering both the possible delays in training and uploading the local models and the resultant staleness among the received models that has heavy impact on the global model fusion. A novel centralized fusion algorithm is designed to determine the fusion weight during the global update, which aims to make full use of the fresh information contained in the uploaded local models while avoiding the biased convergence by enforcing the impact of each UE’s local dataset to be proportional to its sample share. Numerical experiments validate that the proposed asynchronous FL framework can achieve fast and smooth convergence and enhance the training efficiency significantly. Zhaoyang Zhang 0001, Yuqing Tian, Qianqian Yang 0002, Hangguan Shan, Wei Wang 0021, Tony Q. S. Quek |
IEEE Trans. Wirel. Commun. | 4 |
| 2021 | Channel Prediction Based on A Novel Physics-Inspired Generative Learning StructureabstractIn this paper, we try to solve the problem of wireless channel prediction in a fixed area based only on position information of user's equipment. It is the first time that such a problem is proposed and discussed. Different from recent channel prediction methods which need a sequence of measured channel state information (CSI) as known factor, we view this task as a generative problem. A large amount of CSI data measured in the historical communication process can be made use of directly. For solving this problem in a data-driven way, a novel physics-inspired learning structure (C-GRBF) is proposed which fits the physics process of channel impulse response formulating perfectly. Scattering environment information is learned as parameters of the network and the principle of electromagnetic wave propagation is implicit represented by the structure of the network. In the meantime, the reason why conventional universal learning structures fail in solving this problem is analyzed. Experimental results show great performance in prediction accuracy, convergence speed and network robustness of the proposed learning structure. Zhuoran Xiao, Zhaoyang Zhang 0001, Chongwen Huang, Qianqian Yang 0002, Xiaoming Chen 0001 |
VTC Fall | 4 |
| 2021 | Communication-Efficient Federated Learning With Binary Neural NetworksabstractFederated learning (FL) is a privacy-preserving machine learning setting that enables many devices to jointly train a shared global model without the need to reveal their data to a central server. However, FL involves a frequent exchange of the parameters between all the clients and the server that coordinates the training. This introduces extensive communication overhead, which can be a major bottleneck in FL with limited communication links. In this paper, we consider training the binary neural networks (BNNs) in the FL setting instead of the typical real-valued neural networks to fulfill the stringent delay and efficiency requirement in wireless edge networks. We introduce a novel FL framework of training BNNs, where the clients only upload the binary parameters to the server. We also propose a novel parameter updating scheme based on the Maximum Likelihood (ML) estimation that preserves the performance of the BNN even without the availability of aggregated real-valued auxiliary parameters that are usually needed during the training of the BNN. Moreover, for the first time in the literature, we theoretically derive the conditions under which the training of BNN is converging. Numerical results show that the proposed FL framework significantly reduces the communication cost compared to the conventional neural networks with typical real-valued parameters, and the performance loss incurred by the binarization can be further compensated by a hybrid method. Yuzhi Yang, Zhaoyang Zhang 0001, Qianqian Yang 0002 |
IEEE J. Sel. Areas Commun. | 3 |
| 2021 | RCoNet: Deformable Mutual Information Maximization and High-Order Uncertainty-Aware Learning for Robust COVID-19 DetectionabstractThe novel 2019 Coronavirus (COVID-19) infection has spread worldwide and is currently a major healthcare challenge around the world. Chest computed tomography (CT) and X-ray images have been well recognized to be two effective techniques for clinical COVID-19 disease diagnoses. Due to faster imaging time and considerably lower cost than CT, detecting COVID-19 in chest X-ray (CXR) images is preferred for efficient diagnosis, assessment, and treatment. However, considering the similarity between COVID-19 and pneumonia, CXR samples with deep features distributed near category boundaries are easily misclassified by the hyperplanes learned from limited training data. Moreover, most existing approaches for COVID-19 detection focus on the accuracy of prediction and overlook uncertainty estimation, which is particularly important when dealing with noisy datasets. To alleviate these concerns, we propose a novel deep network named RCoNetksfor robust COVID-19 detection which employs Deformable Mutual Information Maximization (DeIM), Mixed High-order Moment Feature (MHMF), and Multiexpert Uncertainty-aware Learning (MUL). With DeIM, the mutual information (MI) between input data and the corresponding latent representations can be well estimated and maximized to capture compact and disentangled representational characteristics. Meanwhile, MHMF can fully explore the benefits of using high-order statistics and extract discriminative features of complex distributions in medical imaging. Finally, MUL creates multiple parallel dropout networks for each CXR image to evaluate uncertainty and thus prevent performance degradation caused by the noise in the data. The experimental results show that RCoNetksachieves the state-of-the-art performance on an open-source COVIDx dataset of 15 134 original CXR images across several metrics. Crucially, our method is shown to be more effective than existing methods with the presence of noise in the data. Shunjie Dong, Qianqian Yang 0002, Yu Fu 0008, Cheng Zhuo |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | Distributed Deep Convolutional Compression for Massive MIMO CSI FeedbackabstractMassive multiple-input multiple-output (MIMO) systems require downlink channel state information (CSI) at the base station (BS) to achieve spatial diversity and multiplexing gains. In a frequency division duplex (FDD) multiuser massive MIMO network, each user needs to compress and feedback its downlink CSI to the BS. The CSI overhead scales with the number of antennas, users and subcarriers, and becomes a major bottleneck for the overall spectral efficiency. In this paper, we propose a deep learning (DL)-based CSI compression scheme, calledDeepCMC, composed of convolutional layers followed by quantization and entropy coding blocks. In comparison with previous DL-based CSI reduction structures, DeepCMC proposes a novel fully-convolutional neural network (NN) architecture, with residual layers at the decoder, and incorporates quantization and entropy coding blocks into its design. DeepCMC is trained to minimize a weighted rate-distortion cost, which enables a trade-off between the CSI quality and its feedback overhead. Simulation results demonstrate that DeepCMC outperforms the state of the art CSI compression schemes in terms of the reconstruction quality of CSI for the same compression rate. We also propose a distributed version of DeepCMC for a multi-user MIMO scenario to encode and reconstruct the CSI from multiple users in a distributed manner. Distributed DeepCMC not only utilizes the inherent CSI structures of a single MIMO user for compression, but also benefits from the correlations among the channel matrices of nearby users to further improve the performance in comparison with DeepCMC. We also propose a reduced-complexity training method for distributed DeepCMC, allowing to scale it to multiple users, and suggest a cluster-based distributed DeepCMC approach for practical implementation. Mahdi Boloursaz Mashhadi, Qianqian Yang 0002, Deniz Gündüz |
IEEE Trans. Wirel. Commun. | 2 |
| 2020 | CNN-Based Analog CSI Feedback in FDD MIMO-OFDM SystemsabstractMassive multiple-input multiple-output (MIMO) systems require downlink channel state information (CSI) at the base station (BS) to better utilize the available spatial diversity and multiplexing gains. However, in a frequency division duplex (FDD) massive MIMO system, CSI feedback overhead degrades the overall spectral efficiency. Deep Learning (DL)-based CSI feedback compression schemes have received a lot of attention recently as they provide significant improvements in compression efficiency; however, they still require reliable feedback links to convey the compressed CSI information to the BS. Instead, we propose here a Convolutional neural network (CNN)-based analog feedback scheme, called AnalogDeepCMC, which directly maps the downlink CSI to uplink channel input. Corresponding noisy channel outputs are used by another CNN to reconstruct the downlink channel estimate. The proposed analog scheme not only outperforms existing digital CSI feedback schemes in terms of the achievable downlink rate, but also simplifies the feedback transmission as it does not require explicit quantization, coding, and modulation, and provides a low-latency alternative particularly in rapidly changing MIMO channels, where the CSI needs to be estimated and fed back periodically. Mahdi Boloursaz Mashhadi, Qianqian Yang 0002, Deniz Gündüz |
ICASSP | 2 |
| 2020 | Centralized Caching and Delivery of Correlated Contents Over Gaussian Broadcast ChannelsabstractContent delivery in a multi-user cache-aided broadcast network is studied, where a server holding a database of correlated contents communicates with the users over a Gaussian broadcast channel (BC). The minimum transmission power required to satisfy all possible demand combinations is studied, when the users are equipped with caches of equal size. Two centralized caching schemes are proposed, both of which not only utilize the user's local caches, but also exploit the correlation among the contents in the database. The first scheme implements uncoded cache placement and delivers coded contents to users using superposition coding. The second scheme, which is proposed for small cache sizes, places coded contents in users' caches and jointly encodes the cached contents of users and the messages targeted at them. The performance of the proposed schemes, which provide upper bounds on the required transmit power for a given cache capacity, is characterized. The scheme based on coded placement improves upon the first one for small cache sizes, and under certain conditions meets the uncoded placement lower bound. A lower bound on the required transmit power is also presented assuming uncoded cache placement. Our results indicate that exploiting the correlations among the contents in a cache-aided Gaussian BC can provide significant energy savings. Qianqian Yang 0002, Parisa Hassanzadeh, Deniz Gündüz, Elza Erkip |
IEEE Trans. Commun. | 1 |
| 2019 | Audience-Retention-Rate-Aware Caching and Coded Video Delivery With Asynchronous DemandsabstractMost of the current literature on coded caching focus on a static scenario in which a fixed number of users synchronously place their requests from a content library, and the performance is measured in terms of the latency in satisfying all of these requests. In practice, however, users start watching an online video content asynchronously over time, and often abort watching a video before it is completed. The latter behavior is captured by the notion of audience retention rate, which measures the portion of a video content watched on average. In order to bring coded caching one step closer to practice, asynchronous user demands are considered in this paper by allowing user demands to arrive randomly over time, and both the popularity of video files, and the audience retention rates are taken into account. A decentralized partial coded delivery (PCD) scheme is proposed, and two cache allocation schemes are employed; namely homogeneous cache allocation (HoCA) and heterogeneous cache allocation (HeCA), which allocate users' caches among different chunks of the video files in the library. Numerical results validate that the proposed PCD scheme, either with HoCA or HeCA, outperforms conventional uncoded caching as well as the state-of-the-art decentralized caching schemes, which consider only the file popularities, and are designed for synchronous demand arrivals. An information-theoretical lower bound on the average delivery rate is also presented. Qianqian Yang 0002, Mohammad Mohammadi Amiri, Deniz Gündüz |
IEEE Trans. Commun. | 1 |
| 2018 | Centralized Coded Caching of Correlated ContentsabstractCoded caching and delivery is studied taking into account the correlations among the contents in the library. Correlations are modeled as common parts shared by multiple contents; that is, each file in the database is composed of a group of subfiles, where each subfile is shared by a different subset of files. The number of files that include a certain subfile is defined as the level of commonness of this subfile. First, a correlation-aware uncoded caching scheme is proposed, and it is shown that the optimal placement for this scheme gives priority to the subfiles with the highest levels of commonness. Then a correlation- aware coded caching scheme is presented, and the cache capacity allocated to subfiles with different levels of commonness is optimized in order to minimize the delivery rate. The proposed correlation-aware coded caching scheme is shown to remarkably outperform state-of-the-art correlation-ignorant solutions, indicating the benefits of exploiting content correlations in coded caching and delivery in networks. Qianqian Yang 0002, Deniz Gündüz |
ICC | 1 |
| 2018 | Centralized caching and delivery of correlated contents over a Gaussian broadcast channelabstractContent delivery in a multi-user cache-aided broadcast network is studied, where a server holding a database of correlated contents communicates with the users over a Gaussian broadcast channel (BC). The minimum transmission power required to satisfy all possible demand combinations is studied, when the users are equipped with caches of equal size. A lower bound on the required transmit power is derived, assuming uncoded cache placement, as a function of the cache capacity. A centralized joint cache and channel coding scheme is proposed, which not only utilizes the user's local caches, but also exploits the correlation among the contents in the database. This scheme provides an upper bound on the minimum required transmit power for a given cache capacity. Our results indicate that exploiting the correlations among the contents in a cache-aided Gaussian BC can provide significant energy savings. Qianqian Yang 0002, Parisa Hassanzadeh, Deniz Gündüz, Elza Erkip |
WiOpt | 1 |
| 2018 | Coded Caching and Content Delivery With Heterogeneous Distortion RequirementsabstractCache-aided coded content delivery is studied for devices with diverse quality-of-service requirements, specified by a different average distortion target. The network consists of a server holding a database of independent contents, and users equipped with local caches of different capacities. User caches are filled by the server during a low traffic period without the knowledge of particular user demands. As opposed to the current literature, which assumes that the users request files in their entirety, it is assumed that the users in the system have distinct distortion requirements; and therefore, each user requests a single file from the database to be served at a different distortion level. Our goal in this paper is to characterize the minimum delivery rate the server needs to transmit over an error-free shared link to satisfy all possible demand combinations at the requested distortion levels, considering both centralized and decentralized cache placement. For centralized cache placement, the optimal delivery rate is characterized for the two-file two-user scenario for any pair of target distortion requirements, when the underlying source distribution is successively refinable. For the two-user scenario with more than two successively refinable files, the optimal scheme is characterized when the cache capacities of the users are the same and the number of files is a multiple of 3. For the general source distribution, not necessarily successively refinable, and with arbitrary number of users and files, a layered caching and delivery scheme is proposed, assuming that scalable source coding is employed at the server. This allows dividing the problem into two subproblems: the lossless caching of each layer with heterogeneous cache sizes, and cache allocation among layers. A delivery rate minimization problem is formulated and solved numerically for each layer; while two different schemes are proposed to allocate user caches among layers, namely, proportional cache allocation and ordered cache allocation. A decentralized lossy coded caching scheme is also proposed, and its delivery rate performance is studied. Simulation results validate the effectiveness of the proposed schemes in both settings. Qianqian Yang 0002, Deniz Gündüz |
IEEE Trans. Inf. Theory | 1 |
| 2017 | The multi-layer information bottleneck problemabstractThe muti-layer information bottleneck (IB) problem, where information is propagated (or successively refined) from layer to layer, is considered. Based on information forwarded by the preceding layer, each stage of the network is required to preserve a certain level of relevance with regards to a specific hidden variable, quantified by the mutual information. The hidden variables and the source can be arbitrarily correlated. The optimal trade-off between rates of relevance and compression (or complexity) is obtained through a singleletter characterization, referred to as the rate-relevance region. Conditions of successive refinabilty are given. Binary source with BSC hidden variables and binary source with BSC/BEC mixed hidden variables are both proved to be successively refinable. We further extend our result to Guassian models. A counterexample of successive refinability is also provided. Qianqian Yang 0002, Pablo Piantanida, Deniz Gündüz |
ITW | 1 |
| 2017 | Decentralized Caching and Coded Delivery With Distinct Cache CapacitiesabstractDecentralized proactive caching and coded delivery is studied in a content delivery network, where each user is equipped with a cache memory, not necessarily of equal capacity. Cache memories are filled in advance during the off-peak traffic period in a decentralized manner, i.e., without the knowledge of the number of active users, their identities, or their particular demands. User demands are revealed during the peak traffic period, and are served simultaneously through an error-free shared link. The goal is to find the minimum delivery rate during the peak traffic period that is sufficient to satisfy all possible demand combinations. A group-based decentralized caching and coded delivery scheme is proposed, and it is shown to improve upon the state of the art in terms of the minimum required delivery rate when there are more users in the system than files. Numerical results indicate that the improvement is more significant as the cache capacities of the users become more skewed. A new lower bound on the delivery rate is also presented, which provides a tighter bound than the classical cut-set bound. Mohammad Mohammadi Amiri, Qianqian Yang 0002, Deniz Gündüz |
IEEE Trans. Commun. | 2 |
| 2016 | Centralized coded caching for heterogeneous lossy requestsabstractCentralized coded caching of popular contents is studied for users with heterogeneous distortion requirements, corresponding to diverse processing and display capabilities of mobile devices. Users' distortion requirements are assumed to be fixed and known, while their particular demands are revealed only after the placement phase. Modeling each file in the database as an independent and identically distributed Gaussian vector, the minimum delivery rate that can satisfy any demand combination within the corresponding distortion target is studied. The optimal delivery rate is characterized for the special case of two users and two files for any pair of distortion requirements. For the general setting with multiple users and files, a layered caching and delivery scheme, which exploits the successive refinability of Gaussian sources, is proposed. This scheme caches each content in multiple layers, and it is optimized by solving two subproblems: lossless caching of each layer with heterogeneous cache capacities, and allocation of available caches among layers. The delivery rate minimization problem for each layer is solved numerically, while two schemes, called the proportional cache allocation (PCA) and ordered cache allocation (OCA), are proposed for cache allocation. These schemes are compared with each other and the cut-set bound through numerical simulations. Qianqian Yang 0002, Deniz Gündüz |
ISIT | 1 |
| 2016 | Coded caching for a large number of usersabstractWe consider the coded caching problem with a central server containing N files, each of length F bits, and K users, each equipped with a cache of capacity MF bits. We assume that coded contents can be proactively placed into users' caches at no cost during the placement phase. During the delivery phase, each user requests exactly one file from the database, and all the requests are served simultaneously by the server over an error-free common link. The goal is to utilize the local cache memories at the users to reduce the delivery rate from the server during the peak period. Here, we focus on a system which has more users than files, i.e., K > N. We first consider the centralized caching problem, in which the number and identity of active users are known in advance, and propose a group-based coded caching scheme for M = N/K, which improves upon the best achievable scheme in the literature. The proposed centralized caching scheme is then exploited in a decentralized setting, in which neither the number nor the identity of the active users are known during the placement phase. It is shown that the proposed coded caching scheme improves upon the best known decentralized delivery rate as well. Mohammad Mohammadi Amiri, Qianqian Yang 0002, Deniz Gündüz |
ITW | 2 |
| 2014 | Towards optimal barrier coverage in wireless sensor and actor networksabstractBarrier coverage in sensor networks has attracted much attention in recent years. Existing results revealed that sensor mobility can remarkably improve the coverage performance of sensor networks. Considering the high manufacture cost of mobile sensors, in this paper we propose to tradeoff the barrier coverage performance and deployment budget by employing a wireless sensor and actor network (WSAN), wherein an actor is used to move static sensors around in order to enhance the barrier coverage performance. We first formulate the barrier coverage problem in WSAN and propose a new coverage metric to evaluate the barrier coverage performance. Then we design an efficient actor movement scheme, S-AMS, for the case where the number of monitoring points can be divided by the number of available sensors. By exploiting the actor's mobility and clustering procedure, S-AMS is able to significantly improve barrier coverage. Based on the insight from S-AMS, we design G-AMS for the general case. We show that S-AMS achieves asymptotically optimal solution for the special case and G-AMS obtains close-to-optimal solution for the general case. Extensive simulations are conducted to demonstrate the performance of our proposed schemes. Qianqian Yang 0002, Shibo He, Jiming Chen 0001 |
GLOBECOM | 2 |
| 2013 | Energy-efficient area coverage in bistatic radar sensor networksabstractIn this paper we study area coverage in bistatic radar sensor networks (BRSN), which is composed of a collection of transmitters and receivers. Coverage in BRSN is much more difficult than that in traditional sensor networks as the sensing area of a bistatic radar depends on the positions of its component transmitter and receiver, and is in general of an elliptical shape. We first investigate the geometrical relationship between the c-coverage area of a bistatic radar and the distance between its component transmitter and receiver, based on which we reduce the number of candidate bistatic radars from all transmitter-receiver pairs. Then we reduce the problem dimension by transforming the area coverage problem to point coverage problem by employing intersection point concept. Finally we propose an efficient algorithm to solve the Point Coverage Problem, which thus solves the area coverage problem. We perform extensive simulations to validate our analysis and the performance of the proposed algorithm. Qianqian Yang 0002, Shibo He, Jiming Chen 0001 |
GLOBECOM | 1 |
| 2012 | Energy-efficient probabilistic full coverage in wireless sensor networksabstractIt is a common class of applications with wireless sensor network to provide full coverage to the region of interest (ROI), such as environment monitoring, military detection and agricultural observation. Existing literatures on full coverage are mostly based on the binary sensing model to simplify the problem. However, the results are far from the reality since binary sensing model as a coarse approximation is too conservative. The probabilistic sensing model has been proposed as a more realistic model to characterize the sensing region. In this paper, we introduce the concept of ε-full coverage based on probabilistic model, i.e., every point in ROI has at least a probability ε of being covered by sensors. We explore the mathematic relationship between the probabilities of two adjacent points being covered and transform ε-full coverage problem into point coverage problem. Then, we design ε-full coverage optimization (FCO) to select a subset of sensors to provide ε-full coverage dynamically so that the lifetime of network is prolonged. This algorithm outperforms the state-of-the-art solution significantly, which we have validated by simulations. Qianqian Yang 0002, Shibo He, Junkun Li, Jiming Chen 0001, Youxian Sun |
GLOBECOM | 1 |