EDBT 2026 Demo / reviewers in the wild / expert
Zhuoran Xiao
dblp:304/8993
· DBLP profile ↗
14ranked-venue papers
7as first author
14since 2021 · last 2026
0000-0003-0365-5398ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 9 · 4 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Uniair: A Unified AI Framework for Multi-Task Joint Optimization Over the Air Interface
Yijia Feng, Chenhui Ye, Tianyu Jiao, Yunbo Hu, Zhuoran Xiao, Tao Tao 0004 |
WCNC | 6 |
| 2026 | Towards Native Intelligence: 6G-LLM Trained with Reinforcement Learning from NDT Feedback
Zhuoran Xiao, Tao Tao 0004, Chenhui Ye, Yunbo Hu, Yijia Feng, Tianyu Jiao, Liyu Cai |
WCNC | 1 |
| 2026 | StableMIL: Entropy-Stabilized Attention-Based Multiple Instance Learning for Morphologically Variable Whole Slide ImagesabstractAggregating features of tens of thousands of patches into Whole Slide Images (WSIs) representations via aggregators is a crucial step in computational pathology. However, existing aggregation strategies overlook the morphological variability of tissue regions in WSIs stemming from differences in clinical procedures and tumor characteristics, leading to two critical limitations: 1) attention collapse in long sequences caused by significant variation in patch numbers across WSIs (ranging from thousands to tens of thousands per WSI); 2) attention misallocation due to under-trained positional embeddings resulting from the non-uniform spatial coordinates introduced by irregular patch distributions. Consequently, current attention-based methods struggle to generalize across this morphological variability, resulting in inconsistent aggregation performance and compromised model reliability in clinical settings. To address these issues, we propose a Entropy-Stabilized Attention-based Multiple Instance Learning (StableMIL) framework, which incorporates an entropy-stabilized attention mechanism to ensure consistent aggregation across WSIs with varying patch numbers and a Randomly Projected 2D rotary position embedding to enhance spatial representation robustness across irregular patch distributions. Extensive theoretical and experimental analyses on nine WSI datasets spanning diverse cancer types, across both classification and survival prediction tasks, demonstrate that StableMIL effectively overcomes the challenges of handling long instance sequences and out-of-distribution spatial coordinates. Our framework consistently outperforms representative baselines, particularly in survival prediction, with stable improvements observed across all evaluated cancer types and morphological scenarios, highlighting its potential for real-world clinical applications. Our source code is available at https://github.com/theeeqi/stableMIL. Yinuo Lu, Mingxin Qi, Yao Fu 0010, Zhuoran Xiao, Wei Shao 0005, Jie Tian 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2025 | Transmission With Machine Language Tokens: A Paradigm for Task-Oriented Agent CommunicationabstractThe rapid advancement in large foundation models is propelling the paradigm shifts across various industries. One significant change is that agents, instead of traditional machines or humans, will be the primary participants in the future production process, which consequently requires a novel AI-native communication system tailored for agent communications. Integrating the ability of large language models (LLMs) with task-oriented semantic communication is a potential approach. However, the output of existing LLM is human language, which is highly constrained and sub-optimal for agent-type communication. In this paper, we innovatively propose a task-oriented agent communication system. Specifically, we leverage the original LLM to learn a specialized machine language represented by token embeddings. Simultaneously, a multi-modal LLM is trained to comprehend the application task and to extract essential implicit information from multi-modal inputs, subsequently expressing it using machine language tokens. This representation is significantly more efficient for transmission over the air interface. Furthermore, to reduce transmission overhead, we introduce a joint token and channel coding (JTCC) scheme that compresses the token sequence by exploiting its sparsity while enhancing robustness against channel noise. Extensive experiments demonstrate that our approach reduces transmission overhead for downstream tasks while enhancing accuracy relative to the SOTA methods. Zhuoran Xiao, Chenhui Ye, Yijia Feng, Yunbo Hu, Tianyu Jiao, Liyu Cai, Guangyi Liu 0001 |
GLOBECOM | 1 |
| 2025 | Addressing the Curse of Scenario and Task Generalization in AI-6G: A Multi-Modal ParadigmabstractExisting works on machine learning (ML)-empowered wireless communication primarily focus on monolithic scenarios and single tasks. However, with the blooming growth of communication task classes coupled with various task requirements in future 6G systems, this working pattern is obviously unsustainable. Therefore, identifying a groundbreaking paradigm that enables a universal model to solve multiple tasks in the physical layer within diverse scenarios is crucial for future system evolution. This paper aims to fundamentally address the curse of ML model generalization across diverse scenarios and tasks by unleashing multi-modal feature integration capabilities in future systems. Given the universality of electromagnetic propagation theory, the communication process is determined by the scattering environment, which can be more comprehensively characterized by cross-modal perception, thus providing sufficient information for all communication tasks across varied environments. This fact motivates us to propose a transformative two-stage multi-modal pre-training and downstream task adaptation paradigm. In the pre-training stage, we introduce a multi-modal two-tower model and a corresponding contrastive learning method to integrate the explicit description of the scattering environment and implicit channel state information (CSI) into a universal representation, which encapsulates rich high-level knowledge and can be leveraged for all downstream tasks in different scenarios. Additionally, we present two specially designed model structures to enhance the interaction of communication modalities. In the second stage, based on the frozen pre-trained model, we propose a direct method and a pluggable method for flexible and low-cost task adaptation. Experimental results demonstrate that our proposed approach significantly outperforms benchmarks in both task performance and tuning parameter size for exemplary sub-tasks in unseen scenarios. Tianyu Jiao, Zhuoran Xiao, Yin Xu 0001, Chenhui Ye, Zhiyong Chen 0002, Liyu Cai, Dazhi He, Yunfeng Guan 0001, Guangyi Liu 0001, Wenjun Zhang 0001 |
IEEE Trans. Wirel. Commun. | 2 |
| 2024 | From Data-Driven Learning to Physics-Inspired Inferring: A Novel Mobile MIMO Channel Prediction Scheme Based on Neural ODEabstractIn this paper, we propose an innovative learning-based channel prediction scheme so as to achieve higher prediction accuracy and reduce the requirements of huge amounts and strict sequential format of channel data. Inspired by the idea of the neural ordinary differential equation (Neural ODE), we first prove that the channel prediction problem can be modeled as an ODE problem with a known initial value by analyzing the physical process of electromagnetic wave propagation within a mobile environment. Then, we design a novel physics-inspired spatial channel gradient network (SCGnet), which represents the derivative process of channel varying as a special neural network and can obtain the gradients at any relative displacement needed for the ODE solving. With the SCGnet, the static channel at any location served by the base station is accurately inferred through consecutive propagation and integration. Finally, we design an efficient recurrent positioning algorithm based on some prior knowledge of user mobility to obtain the velocity vector and propose an approximate Doppler compensation method to make up the instantaneous angular-delay domain channel. Only discrete historical channel data is needed for the training, whereas only a few fresh channel measurements are needed for the prediction, which ensures the scheme’s practicability. Comprehensive evaluations show that the proposed scheme is most efficient in representing, learning, and predicting mobile wireless channels. Zhuoran Xiao, Zhaoyang Zhang 0001, Zhaohui Yang 0001, Chongwen Huang, Xiaoming Chen 0001 |
IEEE Trans. Wirel. Commun. | 1 |
| 2023 | CSI of Each Subcarrier is a Fingerprint: Multi-Carrier Cumulative Learning Based Positioning in Massive MIMO SystemsabstractViewing channel state information (CSI) as a fingerprint to infer user position is a promising technology. Recently, with the help of deep learning (DL) methods, the performance of CSI-based positioning has been further improved. However, the performance of DL methods is highly dependent on the number of training samples. Thus, achieving high performance with limited training samples has become an important topic. Analyzing the prior properties of the task and introducing the prior into the neural network is an effective means. In this paper, by analyzing the physical correlation between CSI and position, we point out that the CSI of each subcarrier is a valid fingerprint of user position in a massive multiple-input multiple-output (MIMO) system. With this property, we propose a positioning scheme based on multi-carrier cumulative learning neural network (MCCNet). Instead of directly learning the mapping from the entire MIMO-orthogonal frequency division multiplexing (OFDM) CSI to position, MCCNet first learns to map from the CSI of each subcarrier to position and then accumulates the extracted features from each subcarrier for the final position inference. This a priori design makes the feature extraction from the entire CSI downsize to the CSI of each subcarrier, reducing the learning burden. Simulation experiments on line-of-sight and non-line-of-sight scenarios both show that compared to the existing methods, MCCNet can reduce the averaged position error by at least 53% with the same training samples or achieve the same performance with only a quarter of training samples. Zhaoyang Zhang 0001, Zhuoran Xiao, Chuanzhi Zhang, Zhaohui Yang 0001 |
PIMRC | 3 |
| 2023 | Deep Learning-Based Multi-User Positioning in Wireless FDMA Cellular NetworksabstractIn Cooperative Intelligent Transportation Systems (C-ITS) and Connected Automated Vehicles (CAV), accessing multiple users and providing high-precision positioning are both vital. This paper aims to design an efficient deep learning approach to extend current Channel State Information (CSI)-based positioning to Frequency Division Multiple Access (FDMA) mode. In FDMA mode, different users are allocated with different subcarriers, making the user CSI have diverse frequency domain characteristics. The diverse frequency domain characteristics bring huge interference to the neural network for stable position inference, and efficient designs are required to handle this challenge. This paper proposes a novel approach named multi-frequency fusion learning for CSI-based positioning. By first using a shareable method to extract position-related features from CSI on each subcarrier independently and then fusing the obtained features, the designed neural network obtains excellent frequency domain flexibility to cope with the diverse frequency address challenge in FDMA mode. Meanwhile, we provide the feasibility analysis of this learning approach in massive Multiple-Input Multiple-Output (MIMO) systems to ensure its stable application. Based on the architecture of multi-frequency fusion learning, we propose two specific positioning schemes with differentiated designs. One is a Multi-Frequency Ensemble Network (MFENet), which extracts and fuses frequency-independent features to ensure the network is utterly unharmed by the complicated frequency domain characteristics. The other is a Multi-Frequency Cumulative Network (MFCNet), which uses sufficient feature accumulation to achieve high precision positioning. The key performance indices and applications on vehicles are comprehensively compared with popular deep-learning methods. Experiment results show the effectiveness and superiority of the proposed schemes. Zhaoyang Zhang 0001, Zhuoran Xiao, Zhaohui Yang 0001, Richeng Jin |
IEEE J. Sel. Areas Commun. | 3 |
| 2023 | Viewing Channel as Sequence Rather Than Image: A 2-D Seq2Seq Approach for Efficient MIMO-OFDM CSI FeedbackabstractIn this paper, we aim to design an effective learning-based channel state information (CSI) feedback scheme for the multiple-input multiple-output (MIMO) orthogonal frequency division multiplexing (OFDM) systems from a physics-inspired perspective. We first argue that the CSI matrix of a MIMO-OFDM system is physically closer to a two-dimensional (2-D) sequence rather than an image due to its apparent unsmoothness, non-scalability, and translational variance within both the spatial and frequency domains. On this basis, we introduce a 2-D long short-term memory (LSTM) neural network to represent the CSI and propose a 2-D sequence-to-sequence (Seq2Seq) model for CSI compression and reconstruction. Specifically, one two-layer 2-D LSTM is used for CSI feature extraction, and the other is used for CSI representation and reconstruction. The proposed scheme can not only fully utilize the unique 2-D characteristics of CSI but also preserve the index information and unsmooth features of the CSI matrix compared with current convolutional neural network (CNN) based schemes. We show that the computational complexity of the proposed scheme is linear in the number of transmit antennas and subcarriers. Its key performances, like reconstruction accuracy, convergence speed, generalization ability after short-term training, and robustness to lossy feedback, are comprehensively compared with existing popular convolutional networks. Experimental results show that our scheme can bring up to nearly 7 dB gain in reconstruction accuracy under the same overhead and reduce feedback overhead by up to 75% under the same accuracy compared with the conventional CNN-based approaches. Zhaoyang Zhang 0001, Zhuoran Xiao, Zhaohui Yang 0001, Kai-Kit Wong |
IEEE Trans. Wirel. Commun. | 3 |
| 2022 | Mobile MIMO Channel Prediction with ODE-RNN: a Physics-Inspired Adaptive ApproachabstractObtaining accurate channel state information (CSI) is crucial and challenging for multiple-input multiple-output (MIMO) wireless communication systems. The conventional channel estimation method cannot guarantee the accuracy of mobile CSI while requiring high signaling overhead. Through exploring the intrinsic correlation among a set of historical CSI instances randomly obtained in a certain communication environment, channel prediction can significantly increase CSI accuracy and save signaling overhead. In this paper, we propose a novel channel prediction method based on ordinary differential equation (ODE)-recurrent neural network (RNN) for accurate and flexible mobile MIMO channel prediction. Different from existing works using sequential network structures for exploring the numerical correlation between observed data, our proposed method tries to represent the implicit physics process of path responses changing by a specially designed continuous learning network with ODE structure. Due to the targeted design of the learning network, our proposed method fits the mathematics feature of CSI data better and enjoy higher network interpretability. Experimental results show that the proposed learning approach outperforms existing methods, especially for long time interval of the CSI sequence and large channel measurement error. Zhuoran Xiao, Zhaoyang Zhang 0001, Zhaohui Yang 0001, Richeng Jin |
PIMRC | 1 |
| 2022 | Viewing the MIMO Channel as Sequence Rather than Image: A Seq2Seq Approach for Efficient CSI FeedbackabstractIn a massive multiple-input multiple-output (MIMO) system, channel state information (CSI) is essential for the base station (BS) to achieve high performance gain. The user equipment (UE) needs to estimate CSI and then feeds it back to the BS in the frequency division duplexing (FDD) mode. Effective compression of CSI will significantly reduce the cost of channel feedback, and many deep learning (DL) based channel compression schemes have been proposed to achieve this goal. In this paper, rather than viewing CSI as images as in most existing works, we propose a new perspective of viewing CSI as an information sequence and analogize CSI feedback to a machine translation task. Further, we propose a novel sequence to sequence (Seq2Seq) model for CSI feedback composed of only recurrent neural networks and a small-scale fully connected layer, avoiding the convolution and pooling structure commonly used in current DL-based works. The advantage of this scheme is that it fully integrates the physical characteristics of the MIMO channel in the spatial domain into the structure of the neural network and avoids the loss of some prominent unsmooth or discontinuous features caused by the inappropriate convolution and pooling operation. Simulation results show that the proposed Seq2Seq model outperforms other DL-based CSI compression techniques under various communication scenarios. Zhaoyang Zhang 0001, Zhuoran Xiao |
WCNC | 3 |
| 2022 | C-GRBFnet: A Physics-Inspired Generative Deep Neural Network for Channel Representation and PredictionabstractIn this paper, we aim to efficiently and accurately predict the static channel impulse response (CIR) with only the user’s position information and a set of channel instances obtained within a certain wireless communication environment. Such a problem is by no means trivial since it needs to reconstruct the high-dimensional information (here the CIR everywhere) from the extremely low-dimensional data (here the location coordinates), which often results in overfitting and large prediction error. To this end, we resort to a novel physics-inspired generative approach. Specifically, we first use a forward deep neural network to infer the positions of all possible images of the source reflected by the surrounding scatterers within that environment, and then use the well-known Gaussian Radial Basis Function network (GRBF) to approximate the amplitudes of all possible propagation paths. We further incorporate the most recently developed sinusoidal representation network (SIREN) into the proposed network to implicitly represent the highly dynamic phases of all possible paths, which usually cannot be well predicted by the conventional neural networks with non-periodic activators. The resultant framework of Cosine-Gaussian Radial Basis Function network (C-GRBFnet) is also extended to the MIMO channel case. Key performance measures including prediction accuracy, convergence speed, network scale and robustness to channel estimation error are comprehensively evaluated and compared with existing popular networks, which show that our proposed network is much more efficient in representing, learning and predicting wireless channels in a given communication environment. Zhuoran Xiao, Zhaoyang Zhang 0001, Chongwen Huang, Xiaoming Chen 0001, Caijun Zhong, Mérouane Debbah |
IEEE J. Sel. Areas Commun. | 1 |
| 2021 | GPAE-LSTMnet: A Novel Learning Structure for Mobile MIMO Channel PredictionabstractMobile channel estimation is very challenging as usually it requires more pilots and channel observations to obtain the channel state information (CSI) and the resultant estimation accuracy may decrease with the number of antennas and sub-carriers. Through exploring the long-and-short-term intrinsic spatial and temporal correlation among a set of historic channel instances randomly obtained within a certain communication environment, channel prediction can help increase the CSI accuracy w.r.t. to that obtained from only the pilots, and thus save signaling overhead and computational cost. In this paper, we propose a novel generative Periodic-Activator-enabled Auto Encoder-LSTM network (GPAE-LSTMnet) for accurate channel prediction of mobile MIMO channels, which first compresses the high dimensional channel matrix with high-frequency features to a low dimensional space with relatively low-frequency feature space that has high data smoothness and is suitable for time-series sequence prediction. After that, a LSTM network is used to predict the channel in the low dimensional space, which ensures high accuracy and low computational cost. Experimental results show that our proposed learning structure outperforms existing methods especially when the dimension of CSI to be predicted is relatively high, the time interval of the CSI sequence is relatively long and the number of network parameters is highly limited. Zhuoran Xiao, Zhaoyang Zhang 0001, Chongwen Huang, Caijun Zhong, Xiaoming Chen 0001 |
PIMRC | 1 |
| 2021 | Channel Prediction Based on A Novel Physics-Inspired Generative Learning StructureabstractIn this paper, we try to solve the problem of wireless channel prediction in a fixed area based only on position information of user's equipment. It is the first time that such a problem is proposed and discussed. Different from recent channel prediction methods which need a sequence of measured channel state information (CSI) as known factor, we view this task as a generative problem. A large amount of CSI data measured in the historical communication process can be made use of directly. For solving this problem in a data-driven way, a novel physics-inspired learning structure (C-GRBF) is proposed which fits the physics process of channel impulse response formulating perfectly. Scattering environment information is learned as parameters of the network and the principle of electromagnetic wave propagation is implicit represented by the structure of the network. In the meantime, the reason why conventional universal learning structures fail in solving this problem is analyzed. Experimental results show great performance in prediction accuracy, convergence speed and network robustness of the proposed learning structure. Zhuoran Xiao, Zhaoyang Zhang 0001, Chongwen Huang, Qianqian Yang 0002, Xiaoming Chen 0001 |
VTC Fall | 1 |