EDBT 2026 Demo / reviewers in the wild / expert
Yuxuan Shi 0001
dblp:240/7215-1
· DBLP profile ↗
14ranked-venue papers
3as first author
14since 2021 · last 2026
0000-0003-4923-6930ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 13 · 2 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FDD CSI Feedback under Finite Downlink Training: A Rate-Distortion Perspective
Shuao Chen, Junyuan Gao, Yuxuan Shi 0001, Yongpeng Wu 0001, Giuseppe Caire, H. Vincent Poor, Wenjun Zhang 0001 |
ICC | 3 |
| 2026 | WVSC: Wireless Video Semantic Communication with Multi-Frame CompensationabstractExisting wireless video transmission schemes directly conduct video coding in pixel level, while neglecting the inner semantics contained in videos. In this paper, we propose a wireless video semantic communication framework, abbreviated as WVSC, which integrates the idea of semantic communication into wireless video transmission scenarios. WVSC first encodes original video frames as semantic frames and then conducts video coding based on such compact representations, enabling the video coding in semantic level rather than pixel level. Moreover, to further reduce the communication overhead, a reference semantic frame is introduced to substitute motion vectors of each frame in common video coding methods. At the receiver, multi-frame compensation (MFC) is proposed to produce compensated current semantic frame with a multi-frame fusion attention module. With both the reference frame transmission and MFC, the bandwidth efficiency improves with satisfying video transmission performance. Experimental results verify the performance gain of WVSC over other DL-based methods e.g. DVSC about 1 dB and traditional schemes about 2 dB in terms of PSNR. Bingyan Xie, Yongpeng Wu 0001, Yuxuan Shi 0001, Biqian Feng, Wenjun Zhang 0001, Jihong Park, Tony Q. S. Quek |
WCNC | 3 |
| 2026 | WirelessGPT: A Generative Foundation Model for Multi-Task Integrated Sensing and CommunicationabstractThis paper presents WirelessGPT, a generative foundation model designed for multi-task learning in integrated sensing and communication (ISAC) systems. Built upon large-scale heterogeneous wireless datasets including Traciverse, Sensiverse, and DeepMIMO, WirelessGPT learns universal spatio-temporal-frequency representations through self-supervised pretraining with masked channel token prediction. The proposed architecture introduces a multi-scale patch embedding module to capture both local and global channel features, and a triple-axis attention encoder to jointly model temporal, spatial, and frequency-domain dependencies. After pretraining, the model can be efficiently fine-tuned via lightweight adapters for diverse downstream tasks such as channel estimation, channel prediction, human activity recognition, environment reconstruction, and object tracking. Experimental results show that WirelessGPT achieves superior accuracy and generalization under limited labeled data and dynamic ISAC conditions, outperforming traditional and task-specific models in low-SNR and high-mobility scenarios while maintaining efficient inference suitable for edge deployment. By unifying communication and sensing functionalities within a single generative backbone, WirelessGPT establishes a scalable paradigm for AI-native 6G systems, enabling shared representations that support heterogeneous wireless tasks. Tingting Yang 0001, Ping Zhang 0003, Mengfan Zheng, Yuxuan Shi 0001, Liwen Jing 0001, Jianbo Huang, Nan Li 0011 |
IEEE J. Sel. Areas Commun. | 4 |
| 2026 | Wireless Video Semantic Communication With Decoupled Diffusion Multi-Frame CompensationabstractExisting wireless video transmission schemes directly conduct video coding in pixel level, while neglecting the inner semantics contained in videos. In this paper, we propose a wireless video semantic communication framework with decoupled diffusion multi-frame compensation (DDMFC), abbreviated as WVSC-D, which integrates the idea of semantic communication into wireless video transmission scenarios. WVSC-D first encodes original video frames as semantic frames and then conducts video coding based on such compact representations, enabling the video coding in semantic level rather than pixel level. Moreover, to further reduce the communication overhead, a reference semantic frame is introduced to substitute motion vectors of each frame in common video coding methods. At the receiver, DDMFC is proposed to generate compensated current semantic frame by a two-stage conditional diffusion process. With both the reference frame transmission and DDMFC frame compensation, the bandwidth efficiency improves with satisfying video transmission performance. Experimental results verify the performance gain of WVSC-D over other DL-based methods e.g. DVSC about 1.8 dB in terms of PSNR. Bingyan Xie, Yongpeng Wu 0001, Yuxuan Shi 0001, Biqian Feng, Wenjun Zhang 0001, Jihong Park, Tony Q. S. Quek |
IEEE Trans. Commun. | 3 |
| 2026 | Signal Compression for Wireless Communication and Sensing: A General Approach Utilizing Pretrained Wireless Foundation ModelsabstractArtificial intelligence is expected to play a central role in enabling future 6 G networks. Developing foundation models that support a wide range of downstream tasks is critical for advancing 6 G standardization. This paper proposes a general framework for compressing wireless channel state information (CSI) using pretrained wireless foundation models. The foundation model is pre-trained using self-supervised learning with a masked reconstruction objective, achieving a normalized mean square error on the order of$10^{-3}$during pretraining. The model is evaluated on a range of wireless communication and sensing tasks, including classification tasks where compressed CSI is directly used for prediction, and regression tasks that require full CSI reconstruction. For regression tasks such as massive MIMO CSI feedback, the pre-compressed output from the foundation model is used as an auxiliary input to the downstream compressor, effectively enhancing the reconstruction quality. Compared to the Type-I codebook with comparable number of feedback bits, our method improves SGCS by 16.22% and reduces NMSE by 93.24%. Additionally, it achieves comparable SGCS performance to the Type-II codebook while using only 29% of the feedback bits. Comparison with existing research further confirms the contribution of the foundation model's compressed output in improving CSI compression performance. For classification tasks such as WiFi-based human activity recognition and human identification, the compressed representations produced by the foundation model can be directly utilized without additional fine-tuning. These representations achieve over 97% accuracy, outperforming conventional AI-based methods even under higher compression ratios. These findings demonstrate that leveraging a pretrained wireless foundation model consistently enhances performance across both classification and regression tasks, underscoring its versatility and potential in wireless CSI processing. Liwen Jing 0001, Tingting Yang 0001, Han Zhang 0025, Yuxuan Shi 0001, Chi Zhang 0111, Bowen Zhang 0005 |
IEEE Trans. Mob. Comput. | 4 |
| 2025 | Joint Lossy Compression for a Vector Gaussian Source under Individual Distortion Criteria
Shuao Chen, Junyuan Gao, Yuxuan Shi 0001, Yongpeng Wu 0001, Giuseppe Caire, H. Vincent Poor, Wenjun Zhang 0001 |
GLOBECOM | 3 |
| 2025 | RWZC: A Model-Driven Approach for Learning-Based Robust Wyner-Ziv CodingabstractIn this paper, a novel learning-based Wyner-Ziv coding framework is considered under a distributed image transmission scenario, where the correlated source is only available at the receiver. Unlike other learnable frameworks, our approach demonstrates robustness to non-stationary source correlation, where the overlapping information between image pairs varies. Specifically, we first model the affine relationship between correlated images and leverage this model for learnable mask generation and rate-adaptive joint source-channel coding. Moreover, we also provide a warping-prediction network to remove the distortion from channel interference and affine transform. Intuitively, the observed performance improvement is largely due to focusing on the simple geometric relationship, rather than the complex joint distribution between the sources. Numerical results show that our framework achieves a 1.5 dB gain in PSNR and a 0.2 improvement in MS-SSIM, along with a significant superiority in perceptual metric, compared to state-of-the-art methods when applied to real-world samples with non-stationary correlations. Yuxuan Shi 0001, Shuo Shao 0001, Yongpeng Wu 0001, Wenjun Zhang 0001, Mérouane Debbah |
IEEE J. Sel. Areas Commun. | 1 |
| 2025 | SCSC: A Novel Standards-Compatible Semantic Communication Framework for Image TransmissionabstractJoint source-channel coding (JSCC) is a promising paradigm for next-generation communication systems, particularly in challenging transmission environments. In this paper, we propose a novel standard-compatible JSCC framework for the transmission of images over multiple-input multiple-output (MIMO) channels. Different from the existing end-to-end AI-based DeepJSCC schemes, our framework consists of learnable modules that enable communication using conventional separate source and channel codes (SSCC), which makes it amenable for easy deployment on legacy systems. Specifically, the learnable modules involve a preprocessing-empowered network (PPEN) for preserving essential semantic information, and a precoder & combiner-enhanced network (PCEN) for efficient transmission over a resource-constrained MIMO channel. We treat existing compression and channel coding modules as non-trainable blocks. Since the parameters of these modules are non-differentiable, we employ a proxy network that mimics their operations when training the learnable modules. Numerical results demonstrate that our scheme can save more than 29% of the channel bandwidth, and requires lower complexity compared to the constrained baselines. We also show its generalization capability to unseen datasets and tasks through extensive experiments. Xue Han 0003, Yongpeng Wu 0001, Zhen Gao 0001, Biqian Feng, Yuxuan Shi 0001, Deniz Gündüz, Wenjun Zhang 0001 |
IEEE Trans. Commun. | 5 |
| 2025 | Indirect Lossy Source Coding With Observed Source Reconstruction: Nonasymptotic Bounds and Second-Order AsymptoticsabstractThis paper considers the joint compression of a pair of correlated sources, where the encoder is allowed to access only one of the sources. The objective is to recover both sources under separate distortion constraints for each source while minimizing the rate. This problem generalizes the indirect lossy source coding problem by also requiring the recovery of the observed source. In this paper, we aim to study the nonasymptotic and second-order asymptotic properties of this problem. Specifically, we begin by deriving nonasymptotic achievability and converse bounds valid for general sources and distortion measures. The source dispersion (Gaussian approximation) is then determined through asymptotic analysis of the nonasymptotic bounds. We further examine the case of erased fair coin flips (EFCF) and provide its specific nonasymptotic achievability and converse bounds. Numerical results under the EFCF case demonstrate that our second-order asymptotic approximation closely approximates the optimum rate at appropriately large blocklengths. Huiyuan Yang, Yuxuan Shi 0001, Shuo Shao 0001, Xiaojun Yuan 0002 |
IEEE Trans. Commun. | 2 |
| 2025 | Semantic-Aided Parallel Image Transmission Compatible With Practical SystemabstractIn this paper, we propose a novel semantic-aided image communication framework for supporting the compatibility with practical separation-based coding architectures. Particularly, the deep learning (DL)-based joint source-channel coding (JSCC) is integrated into the classical separate source-channel coding (SSCC) to transmit the images via the combination of semantic stream and image stream from DL networks and SSCC respectively, which we name as parallel-stream transmission. The positive coding gain stems from the sophisticated design of the JSCC encoder, which leverages the residual information neglected by the SSCC to enhance the learnable image features. Furthermore, a conditional rate adaptation mechanism is introduced to adjust the transmission rate of semantic stream according to residual, rendering the framework more flexible and efficient to bandwidth allocation. We also design a dynamic stream aggregation strategy at the receiver, which provides the composite framework with more robustness to signal-to-noise ratio (SNR) fluctuations in wireless systems compared to a single conventional codec. Finally, the proposed framework is verified to surpass the performance of both traditional and DL-based competitors in a large range of scenarios and meanwhile, maintains lightweight in terms of the transmission and computational complexity of semantic stream, which exhibits the potential to be applied in real systems. Mingkai Xu, Yongpeng Wu 0001, Yuxuan Shi 0001, Xiang-Gen Xia 0001, Mérouane Debbah, Wenjun Zhang 0001, Ping Zhang 0003 |
IEEE Trans. Wirel. Commun. | 3 |
| 2024 | Robust Image Semantic Coding With Learnable CSI Fusion Masking Over MIMO Fading ChannelsabstractThough achieving marvelous progress in various scenarios, existing semantic communication frameworks mainly consider single-input single-output Gaussian channels or Rayleigh fading channels, neglecting the widely-used multiple-input multiple-output (MIMO) channels, which hinders the application into practical systems. One common solution to combat MIMO fading is to utilize feedback MIMO channel state information (CSI). In this paper, we incorporate MIMO CSI into system designs from a new perspective and propose the learnable CSI fusion semantic communication (LCFSC) framework, where CSI is treated as side information by the semantic extractor to enhance the semantic coding. To avoid feature fusion due to abrupt combination of CSI with features, we present a non-invasive CSI fusion multi-head attention module inside the Swin Transformer. With the learned attention masking map determined by both source and channel states, more robust attention distribution could be generated. Furthermore, the percentage of mask elements could be flexibly adjusted by the learnable mask ratio, which is produced based on the conditional variational interference in an unsupervised manner. In this way, CSI-aware semantic coding is achieved through learnable CSI fusion masking. Experiment results testify the superiority of LCFSC over traditional schemes and state-of-the-art Swin Transformer-based semantic communication frameworks in MIMO fading channels. Bingyan Xie, Yongpeng Wu 0001, Yuxuan Shi 0001, Wenjun Zhang 0001, Shuguang Cui, Mérouane Debbah |
IEEE Trans. Wirel. Commun. | 3 |
| 2023 | Communication-Efficient Framework for Distributed Image Semantic Wireless TransmissionabstractMultinode communication, which refers to the interaction among multiple devices, has attracted lots of attention in many Internet of Things (IoT) scenarios. However, its huge amounts of data flows and inflexibility for task extension have triggered the urgent requirement of communication-efficient distributed data transmission frameworks. In this article, inspired by the great superiorities on bandwidth reduction and task adaptation of semantic communications, we propose a federated learning (FL)-based semantic communication (FLSC) framework for multitask distributed image transmission with IoT devices. FL enables the design of independent semantic communication link of each user while further improves the semantic extraction and task performance through global aggregation. Each link in FLSC is composed of a hierarchical vision transformer (HVT)-based extractor and a task-adaptive translator for coarse-to-fine semantic extraction and meaning translation according to specific tasks. In order to extend the FLSC into more realistic conditions, we design a channel state information-based multiple-input–multiple-output transmission module to combat channel fading and noise. Simulation results show that the coarse semantic information can deal with a range of image-level tasks. Moreover, especially in low signal-to-noise ratio (SNR) and channel bandwidth ratio regimes, FLSC evidently outperforms the traditional scheme, e.g., about 10 peak SNR gain in the 3-dB channel condition. Bingyan Xie, Yongpeng Wu 0001, Yuxuan Shi 0001, Derrick Wing Kwan Ng, Wenjun Zhang 0001 |
IEEE Internet Things J. | 3 |
| 2023 | Excess Distortion Exponent Analysis for Semantic-Aware MIMO Communication SystemsabstractIn this paper, the analysis of excess distortion exponent for joint source-channel coding (JSCC) in semantic-aware communication systems is presented. By introducing an unobservable semantic source, we extend the classical results by Csiszar to semantic-aware communication systems. Both upper and lower bounds of the exponent for the discrete memoryless source-channel pair are established. Moreover, an extended achievable bound of the excess distortion exponent for MIMO systems is derived. Further analysis explores how the block fading and numbers of antennas influence the exponent of semantic-aware MIMO systems. Our results offer some theoretical bounds of error decay performance and can be used to guide future semantic communications with joint source-channel coding scheme. Yuxuan Shi 0001, Shuo Shao 0001, Yongpeng Wu 0001, Wenjun Zhang 0001, Xiang-Gen Xia 0001, Chengshan Xiao |
IEEE Trans. Wirel. Commun. | 1 |
| 2022 | Error exponent for concatenated codes in DNA data storage under substitution errors
Yuxuan Shi 0001, Shuo Shao 0001, Yongpeng Wu 0001 |
Sci. China Inf. Sci. | 1 |