VLDB 2026 Research / reviewers in the wild / expert
Bokai Xu
dblp:315/0762
· DBLP profile ↗
12ranked-venue papers
3as first author
12since 2021 · last 2026
0009-0007-3444-0708ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 9 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scaling Law for Large Wireless ModelsabstractEmerging from recent advances in foundation models, Large Wireless Models (LWMs) represent a new paradigm of general-purpose intelligence for wireless communications that transcends task-specific engineering. The success of foundation models is critically underpinned by scaling laws, which provide a predictable roadmap for how performance scales with resources. However, established scaling laws from language and vision, charting performance as a power-law of model and dataset sizes, are ill-suited for the wireless domain, as their core formulations cannot model the structured nature of the physical channel. To address this, we propose a novel wireless scaling law that extends the classical formulation by modeling two wireless-native factors: channel heterogeneity and discretization granularity. These two factors reshape scaling behavior via nested linear and power-law relationships, recasting the scaling law's parameters (notably the scaling exponent and irreducible loss) from universal constants into dynamic variables dictated by the physical environment. Our physics-aware formulation reveals two key insights: first, that compute-optimal scaling is not dictated by a fixed model-data ratio but is instead a dynamic function of heterogeneity and granularity, and second, that this dependency is particularly sensitive to granularity, allowing significant performance to be unlocked from existing data simply by refining its resolution. Crucially, this establishes a reliable roadmap for designing powerful yet resource-efficient LWMs, translating theoretical insights into actionable engineering principles. Extensive experiments validate our wireless scaling law, showing a 32.31% prediction accuracy improvement over classical laws in diverse wireless scenarios where they fail. Jiayi Zhang 0001, Bokai Xu, Yiyang Zhu, Enyu Shi |
AAAI | 4 |
| 2026 | Enhancing Physical Layer Security for SIM-aided Cell-free mMIMO Systems
Jiayi Zhang 0001, Enyu Shi, Jiakang Zheng, Bokai Xu, Bo Ai 0001 |
ICC | 5 |
| 2026 | MaLAM4Com: Multi-Agent Cooperative Large AI Models for Wireless CommunicationsabstractLarge artificial intelligence (AI) models for wireless communications have demonstrated remarkable success across a range of wireless downstream tasks. However, their high computational overhead, low training efficiency, and limited privacy protection pose significant challenges for deployment on resource-constrained terminal devices. To address this issue, we propose a novel distributed framework that utilizes a three-layer cooperative paradigm to effectively achieve cooperation among agents, namely Multi-agent cooperative Large AI Models for Wireless Communications: MaLAM4Com. However, two key challenges in MaLAM4Com are how to effectively extract knowledge from shared information and how to alleviate the significant complexity arising from high-dimensional information sharing. To address these bottlenecks, we introduce federated distillation and Lyapunov cooperation to achieve robust knowledge transfer and consistent dynamic evolution, enabling the agents to capture the intrinsic structure of wireless channels. Subsequently, we innovatively utilize low-dimensional embeddings to facilitate information sharing among agents, significantly reducing cooperation complexity by up to 94% while enhancing privacy protection. This breaks traditional cooperative paradigms that rely on wireless channels. Moreover, we further introduce dataset distillation to enhance training efficiency by synthesizing elite data instead of directly utilizing raw datasets. Numerical results demonstrate that MaLAM4Com significantly outperforms existing baselines, with gains exceeding 45% under low sampling ratios. Remarkably, low-dimensional embeddings have also shown significant advantages in downstream tasks, reducing inference complexity by over 96%. Jiayi Zhang 0001, Yiyang Zhu, Enyu Shi, Bokai Xu, Dusit Niyato, Shi Jin 0002, Bo Ai 0001 |
IEEE J. Sel. Areas Commun. | 5 |
| 2026 | Performance Analysis and Optimization Design of Uplink RSMA-Enabled Cell-Free Massive MIMO Systems With Hardware ImpairmentsabstractCell-free (CF) massive multiple-input multiple-output (MIMO) has emerged as a promising technique to deliver uniform signal coverage and high data rates. However, employing low-precision hardware in user equipment introduces susceptibility to hardware impairments (HI), resulting in significantly degraded channel state information (CSI) accuracy. Fortunately, rate-splitting multiple access (RSMA) has been proposed as a robust solution to mitigate the adverse effects of imperfect CSI by performing message splitting at the transmitter and successive interference cancellation (SIC) at the receiver. In this paper, we incorporate RSMA into CF massive MIMO systems to tackle the problem posed by imperfect CSI. Taking into account inevitable pilot contamination, we first derive a novel and closed-form expression for the spectral efficiency (SE) to analytically characterize the performance of RSMA-enabled CF massive MIMO systems under spatially correlated Rician fading channels. Subsequently, we focus on optimizing the decoding order, power allocation, and fronthaul weights to maximize the system’s sum SE. To address this mixed-integer nonlinear programming (MINLP) problem, we initially propose an alternating optimization (AO)-based optimization method that decomposes the original intractable problem into three manageable subproblems, which are iteratively handled until convergence. Considering the significant computational complexity associated with the AO-based approach, we further propose a proximal policy optimization (PPO)-based method to establish an effective and low-complexity optimization framework. Simulation results unveil the detrimental impact of HI on both CSI accuracy and the overall sum SE performance. In particular, the presence of HI introduces residual interference that limits the performance gains achievable through additional RSMA layers, especially in strong line-of-sight scenarios, highlighting the trade-off between these gains and the SIC-related costs in terms of computational complexity and decoding latency. Xilai Feng, Jiakang Zheng, Jiayi Zhang 0001, Bokai Xu, Derrick Wing Kwan Ng, Bo Ai 0001, Victor C. M. Leung |
IEEE Trans. Wirel. Commun. | 4 |
| 2026 | Double-Layer Over-the-Air Synchronization Scheme for Cell-Free Massive MIMO SystemsabstractThe distributed deployment of communication infrastructure is a promising evolutionary trend in the next-generation wireless communication systems, as exemplified by the novel cell-free massive multiple-input multiple-output (CF mMIMO) technology. In user-centric CF mMIMO systems, synchronization among access points (APs) is a critical challenge that significantly impacts the effectiveness of coherent joint processing gains. In this paper, we investigate a CF mMIMO system featuring distributed AP deployments and low-resolution analog-to-digital converters (ADCs). To guarantee precise phase synchronization, we first propose two double-layer AP clustering approaches for rapid synchronization using the Leader-Follower paradigm: one based on the K-means algorithm and the other utilizing classical graph theory with geographical distance metrics in AP deployment. Specifically, in the first layer, a designated Leader AP keeps synchronization with its serving secondary Follower-1 APs, while in the second layer, each Follower-1 AP communicates with its neighboring Follower-2 APs. Next, we propose novel phase synchronization and carrier frequency synchronization strategies among APs based on an over-the-air synchronization signal transmission mechanism, which enables mutual calibration without transmitting any measurements to the central processing unit via fronthaul links. Furthermore, we consider the effect of quantization accuracy of radio frequency hardware on synchronization performance, thereby facilitating the adoption of low-cost components. Finally, simulation results demonstrate that synchronization precision can be significantly improved, reaching values on the order of$10^{-5}$. Additionally, even with moderately coarse ADC quantization, near-optimal performance can be achieved in practical scenarios. Jiayi Zhang 0001, Jiakang Zheng, Bokai Xu, Arumugam Nallanathan, Bo Ai 0001 |
IEEE Trans. Wirel. Commun. | 4 |
| 2026 | Low-Complexity Distributed Combining Design for Near-Field Cell-Free XL-MIMO SystemsabstractIn this paper, we investigate the low-complexity distributed combining scheme design for near-field cell-free extremely large-scale multiple-input-multiple-output (CF XL-MIMO) systems. Firstly, we construct the uplink spectral efficiency (SE) performance analysis framework for CF XL-MIMO systems over centralized and distributed processing schemes. Notably, we derive the centralized minimum mean-square error (CMMSE) and local minimum mean-square error (LMMSE) combining schemes over arbitrary channel estimators. Then, focusing on the CMMSE and LMMSE combining schemes, we propose five low-complexity distributed combining schemes based on the matrix approximation methodology or the symmetric successive over relaxation (SSOR) algorithm. More specifically, we propose two matrix approximation methodology-aided combining schemes: Global Statistics & Local Instantaneous information-based MMSE (GSLI-MMSE) and Statistics matrix Inversion-based LMMSE (SI-LMMSE). These two schemes are derived by approximating the global instantaneous information in the CMMSE combining and the local instantaneous information in the LMMSE combining with the global and local statistics information by asymptotic analysis and matrix expectation approximation, respectively. Moreover, by applying the low-complexity SSOR algorithm to iteratively solve the matrix inversion in the LMMSE combining, we derive three distributed SSOR-based LMMSE combining schemes, distinguished from the applied information and initial values. Zhe Wang 0018, Jiayi Zhang 0001, Bokai Xu, Dusit Niyato, Bo Ai 0001, Shiwen Mao, Zhu Han 0001 |
IEEE Trans. Wirel. Commun. | 3 |
| 2026 | Asynchronous Distributed Beamforming for Beyond-Diagonal RIS-Aided Movable Antenna SystemsabstractMovable antenna (MA) technology has recently attracted significant research attention as a promising solution for enhancing wireless network performance. However, conventional MAs can only effectively serve users in close proximity, resulting in restricted coverage. To overcome this limitation, in this paper, we explore a beyond-diagonal reconfigurable intelligent surface (BD-RIS)-aided MA system. First, we propose a penalty-based block coordinate descent optimization algorithm tailored to the new constraints imposed by BD-RIS-aided MA systems. Specifically, our method decouples the inherently non-convex and coupled antenna distance constraints by introducing auxiliary optimization variables. Subsequently, the resulting problem is efficiently addressed via alternating optimization, with closed-form updates for the auxiliary variables. Furthermore, recognizing the challenges posed by large-scale BD-RIS deployments, which have the potential for serving a substantial number of users, traditional centralized optimization frameworks encounter considerable difficulties, including high computational complexity, excessive communication overheads, as well as limited scalability with increasing system size. To address these limitations, we propose an efficient asynchronous alternating direction method of multipliers (AS-ADMM) scheme aimed at maximizing the sum rate. Our numerical results demonstrate that the BD-RIS-aided MA system achieves superior performance compared to both conventional fixed position antenna and BD-RIS-aided systems. Furthermore, the proposed AS-ADMM framework can achieve a trade-off between performance and computational overhead, highlighting its potential for practical implementation in large-scale wireless communication networks. Bokai Xu, Jiayi Zhang 0001, Zhe Wang 0018, Bo Ai 0001, Derrick Wing Kwan Ng |
IEEE Trans. Wirel. Commun. | 1 |
| 2025 | VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality DocumentsabstractRetrieval-augmented generation (RAG) is an effective technique that enables large language models (LLMs) to utilize external knowledge sources for generation. However, current RAG systems are solely based on text, rendering it impossible to utilize vision information like layout and images that play crucial roles in real-world multi-modality documents. In this paper, we introduce VisRAG, which tackles this issue by establishing a vision-language model (VLM)-based RAG pipeline. In this pipeline, instead of first parsing the document to obtain text, the document is directly embedded using a VLM as an image and then retrieved to enhance the generation of a VLM. Compared to traditional text-based RAG, VisRAG maximizes the retention and utilization of the data information in the original documents, eliminating the information loss introduced during the parsing process. We collect both open-source and synthetic data to train the retriever in VisRAG and explore a variety of generation methods. Experiments demonstrate that VisRAG outperforms traditional RAG in both the retrieval and generation stages, achieving a 20–40% end-to-end performance gain over traditional text-based RAG pipeline. Further analysis reveals that VisRAG is efficient in utilizing training data and demonstrates strong generalization capability, positioning it as a promising solution for RAG on multi-modality documents. Our code and data are available at https://github.com/openbmb/visrag. Shi Yu 0001, Chaoyue Tang, Bokai Xu, Junbo Cui, Junhao Ran, Yukun Yan, Zhenghao Liu 0001, Shuo Wang 0013, Xu Han 0007, Zhiyuan Liu 0001, Maosong Sun 0001 |
ICLR | 3 |
| 2025 | GCN-Based Low-Complexity Downlink Beamforming for Cell-Free Massive MIMO Systems With Partially Coherent Joint TransmissionabstractTo enhance the capacity and reliability of next-generation wireless communication systems, the novel cell-free massive multiple-input multiple-output (mMIMO) has emerged as a pivotal technology in satisfying the stringent quality of service requirements of massive network-connected devices. In this paper, we propose a partially coherent joint transmission (PCJT) approach that draws insights from both coherent and non-coherent joint transmission (NCJT) strategies. Specifically, we design the downlink transmit beamformers to maximize the weighted sum rate (WSR) and compare the performance in three distinct joint transmission modes, ranging from coherent and partially coherent, to non-coherent joint transmission. Specifically, a non-convex optimization problem is formulated that incorporates multiple data stream transmission and transmit power constraints. Given the intractability of the problem, the weighted minimum mean square error (WMMSE) approach is introduced to transform it into an equivalent form, which facilitates the development of a low-complexity and low-interaction reduced WMMSE (R-WMMSE) beamforming algorithm design to acquire an effective solution. For further reducing communication overhead and improving convergence rates, we propose a novel graph convolution network-based unfolding technique for R-WMMSE algorithm. It significantly reduces the number of iterations required while achieving similar performance to the original WMMSE algorithm, thus alleviating the signaling overhead burdens in distributive implementation. Simulation results demonstrate the significant performance gains achieved by the proposed algorithm in terms of superior WSR and rapid convergence performance. Furthermore, it is evident that the performance of PCJT can promote the performance achieved by NCJT, positioning it as an alternative between the existing two joint transmission strategies. Jiayi Zhang 0001, Bokai Xu, Derrick Wing Kwan Ng, Arumugam Nallanathan, Bo Ai 0001 |
IEEE Trans. Wirel. Commun. | 3 |
| 2025 | Deep Unfolding Beamforming and Power Control Designs for Multi-Port Matching NetworksabstractThe key technologies of sixth generation (6G), such as ultra-massive multiple-input multiple-output (MIMO), enable intricate interactions between antennas and wireless propagation environments. As a result, it becomes necessary to develop joint models that encompass both antennas and wireless propagation channels. To achieve this, we utilize the multi-port communication theory, which considers impedance matching among the source, transmission medium, and load to facilitate efficient power transfer. Specifically, we first investigate the impact of insertion loss, mutual coupling, and other factors on the performance of multi-port matching networks. Next, to further improve system performance, we explore two important deep unfolding designs for the multi-port matching networks: beamforming and power control, respectively. For the hybrid beamforming, we develop a deep unfolding framework, i.e., projected gradient descent (PGD)-Net based on unfolding projected gradient descent. For the power control, we design a deep unfolding network, graph neural network (GNN) aided alternating optimization (AO)-Net, which considers the interaction between different ports in optimizing power allocation. Numerical results verify the necessity of considering insertion loss in the dynamic metasurface antenna (DMA) performance analysis. Besides, the proposed PGD-Net based hybrid beamforming approaches approximate the conventional model-based algorithm with very low complexity. Moreover, our proposed power control scheme has a fast run time compared to the traditional weighted minimum mean squared error (WMMSE) method. Bokai Xu, Jiayi Zhang 0001, Qingfeng Lin, Huahua Xiao, Yik-Chung Wu, Bo Ai 0001 |
IEEE Trans. Wirel. Commun. | 1 |
| 2023 | Enhancing Chat Language Models by Scaling High-quality Instructional ConversationsabstractFine-tuning on instruction data has been widely validated as an effective practice for implementing chat language models like ChatGPT.Scaling the diversity and quality of such data, although straightforward, stands a great chance of leading to improved performance.This paper aims to push the upper bound of opensource models further.We first provide a systematically designed, diverse, informative, large-scale dataset of instructional conversations, UltraChat, which does not involve human queries.Our objective is to capture the breadth of interactions between a human user and an AI assistant and employs a comprehensive framework to generate multi-turn conversation iteratively.UltraChat contains 1.5 million high-quality multi-turn dialogues and covers a wide range of topics and instructions.Our statistical analysis of UltraChat reveals its superiority in various key metrics, including scale, average length, diversity, coherence, etc., solidifying its position as a leading opensource dataset.Building upon UltraChat, we fine-tune a LLaMA model to create a powerful conversational model, UltraLM.Our evaluations indicate that UltraLM consistently outperforms other open-source models, including WizardLM and Vicuna, the previously recognized state-of-the-art open-source models. Ning Ding 0002, Yulin Chen 0001, Bokai Xu, Yujia Qin, Shengding Hu, Zhiyuan Liu 0001, Maosong Sun 0001, Bowen Zhou 0002 |
EMNLP | 3 |
| 2023 | Low-Complexity Precoding for Extremely Large-Scale MIMO Over Non-Stationary ChannelsabstractExtremely large-scale multiple-input-multiple-output (XL-MIMO) is a promising technology for the future sixth-generation (6G) networks to achieve higher performance. In practice, various linear precoding schemes, such as zero-forcing (ZF) and regularized zero-forcing (RZF) precoding, are capable of achieving both large spectral efficiency (SE) and low bit error rate (BER) in traditional massive MIMO (mMIMO) systems. However, these methods are not efficient in extremely large-scale regimes due to the inherent spatial non-stationarity and high computational complexity. To address this problem, we investigate a low-complexity precoding algorithm, e.g., randomized Kaczmarz (rKA), taking into account the spatial non-stationary properties in XL-MIMO systems. Furthermore, we propose a novel mode of randomization, i.e., sampling without replacement rKA (SwoR-rKA), which enjoys a faster convergence speed than the rKA algorithm. Besides, the closed-form expression of SE considering the interference between subarrays in downlink XL-MIMO systems is derived. Numerical results show that the complexity given by both rKA and SwoR-rKA algorithms has 51.3% reduction than the traditional RZF algorithm with similar SE performance. More importantly, our algorithms can effectively reduce the BER when the transmitter has imperfect channel estimation. Bokai Xu, Zhe Wang 0018, Huahua Xiao, Jiayi Zhang 0001, Bo Ai 0001, Derrick Wing Kwan Ng |
ICC | 1 |