VLDB 2026 Research / reviewers in the wild / expert
Rongwei Lu
dblp:355/8929
· DBLP profile ↗
11ranked-venue papers
4as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 5 since 2021Computer networks · 5 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enabling Trustworthy Recommendations in the Federated IoT: A Secure and Verifiable Tensor Learning ApproachabstractHigh-dimensional tensors are a cornerstone of context-aware recommendation systems. However, deploying tensor-based models within a federated learning framework faces formidable hurdles in computational efficiency, privacy preservation, and system robustness. These challenges are exacerbated by the untrusted server, which could infer sensitive user data or corrupt the global model. To systematically address these obstacles, we design and implement CSV-FTL, a Cloud-Edge Synergistic Secure and Verifiable Federated Tensor Learning framework. Our framework’s core innovations are fourfold: (1) a natively designed federated tensor model optimized for high efficiency in distributed environments; (2) a lossless double-masking mechanism that achieves strong privacy without compromising model accuracy; (3) a reconstruction algorithm for user dropout tolerance; and (4) a homomorphic hashing scheme enabling users to verify the integrity of server-side aggregation. Consequently, CSV-FTL strikes a superior balance between recommendation accuracy, computational cost, and security guarantees. To the best of our knowledge, this is the first work to seamlessly integrate a native federated tensor model with lossless privacy, dropout tolerance, and verifiable aggregation within a single, unified framework. Both rigorous theoretical analysis and extensive experiments on four real-world recommendation-related datasets validate the provable security and superior performance of our framework over state-of-the-art methods. Rongwei Lu, Xue-Feng Duan, Guoqiang Deng, Yong Ding 0005 |
IEEE Internet Things J. | 1 |
| 2025 | Beyond A Single AI Cluster: A Survey of Decentralized LLM TrainingabstractThe emergence of large language models (LLMs) has revolutionized AI development, yet their resource demands beyond a single cluster or even datacenter, limiting accessibility to well-resourced organizations.Decentralized training has emerged as a promising paradigm to leverage dispersed resources across clusters, datacenters and even regions, offering the potential to democratize LLM development for broader communities.As the first comprehensive exploration of this emerging field, we present decentralized LLM training as a resource-driven paradigm and categorize existing efforts into community-driven and organizational approaches.We further clarify this through: (1) a comparison with related paradigms, (2) characterization of decentralized resources, and (3) a taxonomy of recent advancements.We also provide up-to-date case studies and outline future directions to advance research in decentralized LLM training. Haotian Dong, Jingyan Jiang, Rongwei Lu, Jiajun Luo, Jiajun Song, Zhi Wang 0001 |
EMNLP | 3 |
| 2025 | DICE: Staleness-Centric Optimizations for Parallel Diffusion MoE Inference
Jiajun Luo, Lizhuo Luo, Jianru Xu, Jiajun Song, Rongwei Lu, Zhi Wang 0001 |
ICCV | 5 |
| 2025 | γ-FedHT: Stepsize-Aware Hard-Threshold Gradient Compression in Federated Learning
Rongwei Lu, Yifei Zhu 0001, Bin Chen 0011, Zhi Wang 0001 |
INFOCOM | 1 |
| 2025 | Accelerating Parallel Diffusion Model Serving with Residual CompressionabstractDiffusion models produce realistic images and videos but require substantial computational resources, necessitating multi-accelerator parallelism for real-time deployment. However, parallel inference introduces significant communication overhead from exchanging large activations between devices, limiting efficiency and scalability. We present CompactFusion, a compression framework that significantly reduces communication while preserving generation quality. Our key observation is that diffusion activations exhibit strong temporal redundancy—adjacent steps produce highly similar activations, saturating bandwidth with near-duplicate data carrying little new information. To address this inefficiency, we seek a more compact representation that encodes only the essential information. CompactFusion achieves this via Residual Compression that transmits only compressed residuals (step-wise activation differences). Based on empirical analysis and theoretical justification, we show that it effectively removes redundant data, enabling substantial data reduction while maintaining high fidelity. We also integrate lightweight error feedback to prevent error accumulation. CompactFusion establishes a new paradigm for parallel diffusion inference, delivering lower latency and significantly higher generation quality than prior methods. On 4$\times$L20, it achieves $3.0\times$ speedup while greatly improving fidelity. It also uniquely supports communication-heavy strategies like sequence parallelism on slow networks, achieving $6.7\times$ speedup over prior overlap-based method. CompactFusion applies broadly across diffusion models and parallel settings, and integrates easily without requiring pipeline rework. Portable implementation demonstrated on xDiT is publicly available at https://github.com/Cobalt-27/CompactFusion Jiajun Luo, Yicheng Xiao, Jianru Xu, Yangxiu You, Rongwei Lu, Jingyan Jiang, Zhi Wang 0001 |
NeurIPS | 5 |
| 2025 | LLM4Band: Enhancing Reinforcement Learning with Large Language Models for Accurate Bandwidth EstimationabstractReal-time communication (RTC) applications rely on accurate bandwidth estimation to ensure high-quality communication and user experience. Traditional heuristic and reinforcement learning (RL)-based methods often face challenges with the dynamic nature of real-time networks, leading to issues with generalization. Inspired by the success of Large Language Models (LLMs)---which, with billions of parameters pre-trained on massive datasets, have demonstrated exceptional capabilities in semantic representation, adaptability, and transfer learning---we propose LLM4Band, a novel framework that integrates LLMs with offline reinforcement learning to tackle bandwidth estimation in RTC scenarios. By leveraging the powerful feature extraction capabilities of LLMs and combining them with an offline RL algorithm, LLM4Band incorporates a Balanced Replay Buffer and an LLM-based policy network to significantly enhance robustness and adaptability. Extensive experiments demonstrate that LLM4Band surpasses state-of-the-art methods, achieving a 12.35% improvement in estimation accuracy and a 21% enhancement in communication quality. Rongwei Lu, Cédric Westphal, Dongbiao He, Jingyan Jiang |
NOSSDAV | 2 |
| 2025 | Data-Aware Gradient Compression for FL in Communication-Constrained Mobile ComputingabstractFederated Learning (FL) in mobile environments faces significant communication bottlenecks. Gradient compression has proven as an effective solution to this issue, offering substantial benefits in environments with limited bandwidth and metered data. Yet, it encounters severe performance drops in non-IID environments due to a one-size-fits-all compression approach, which does not account for the varying data volumes across workers. Assigning varying compression ratios to workers with distinct data distributions and volumes is therefore a promising solution. This work derives the convergence rate of distributed SGD with non-uniform compression, which reveals the intricate relationship between model convergence and the compression ratios applied to individual workers. Accordingly, we frame the relative compression ratio assignment as an$n$-variable chi-squared nonlinear optimization problem, constrained by a limited communication budget. We propose DAGC-R, which assigns conservative compression to workers handling larger data volumes. Recognizing the computational limitations of mobile devices, we propose the DAGC-A, which is computationally less demanding and enhances the robustness of compression in non-IID scenarios. Our experiments confirm that the DAGC-R and DAGC-A can speed up the training speed by up to 25.43% and 16.65% compared to the uniform compression respectively, when dealing with highly imbalanced data volume distribution and restricted communication. Rongwei Lu, Yinan Mao, Bin Chen 0011, Laizhong Cui, Zhi Wang 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2024 | Retraining-free Model Quantization via One-Shot Weight-Coupling LearningabstractQuantization is of significance for compressing the over-parameterized deep neural models and deploying them on resource-limited devices. Fixed-precision quantization suf-fers from performance drop due to the limited numerical representation ability. Conversely, mixed-precision quan-tization (MPQ) is advocated to compress the model ef-fectively by allocating heterogeneous bit-width for layers. MPQ is typically organized into a searching-retraining two-stage process. Previous works only focus on determining the optimal bit-width configuration in the first stage effi-ciently, while ignoring the considerable time costs in the second stage and thus hindering deployment efficiency sig-nificantly. In this paper, we devise a one-shot training-searching paradigm for mixed-precision model compression. Specifically, in the first stage, all potential bit-width configurations are coupled and thus optimized simultane-ously within a set of shared weights. However, our ob-servations reveal a previously unseen and severe bit-width interference phenomenon among highly coupled weights during optimization, leading to considerable performance degradation under a high compression ratio. To tackle this problem, we first design a bit-width scheduler to dy-namically freeze the most turbulent bit-width of layers during training, to ensure the rest bit-widths converged prop-erly. Then, taking inspiration from information theory, we present an information distortion mitigation technique to align the behaviour of the bad-performing bit-widths to the well-performing ones. In the second stage, an inference-only greedy search scheme is devised to evaluate the good-ness of configurations without introducing any additional training costs. Extensive experiments on three representative models and three datasets demonstrate the effective-ness of the proposed method. Code can be available on https://github.com/1hunters/retraining-free-quantization. Shuzhao Xie, Rongwei Lu, Xinzhu Ma, Zhi Wang 0001, Wenwu Zhu 0001 |
CVPR | 5 |
| 2024 | MesonGS: Post-training Compression of 3D Gaussians via Efficient Attribute Transformation
Shuzhao Xie, Weixiang Zhang, Yunpeng Bai, Rongwei Lu, Shijia Ge, Zhi Wang 0001 |
ECCV (33) | 5 |
| 2024 | A Joint Approach to Local Updating and Gradient Compression for Efficient Asynchronous Federated Learning
Jiajun Song, Jiajun Luo, Rongwei Lu, Shuzhao Xie, Bin Chen 0011, Zhi Wang 0001 |
Euro-Par (3) | 3 |
| 2023 | DAGC: Data-Aware Adaptive Gradient CompressionabstractGradient compression algorithms are widely used to alleviate the communication bottleneck in distributed ML. However, existing gradient compression algorithms suffer from accuracy degradation in Non-IID scenarios, because a uniform compression scheme is used to compress gradients at workers with different data distributions and volumes, since workers with larger volumes of data are forced to adapt to the same aggressive compression ratios as others. Assigning different compression ratios to workers with different data distributions and volumes is thus a promising solution. In this study, we first derive a function from capturing the correlation between the number of training iterations for a model to converge to the same accuracy, and the compression ratios at different workers; This function particularly shows that workers with larger data volumes should be assigned with higher compression ratios1to guarantee better accuracy. Then, we formulate the assignment of compression ratios to the workers as an n-variables chi-square nonlinear optimization problem under fixed and limited total communication constrain. We propose an adaptive gradient compression strategy called DAGC, which assigns each worker a different compression ratio according to their data volumes. Our experiments confirm that DAGC can achieve better performance facing highly imbalanced data volume distribution and restricted communication. Rongwei Lu, Jiajun Song, Bin Chen 0011, Laizhong Cui, Zhi Wang 0001 |
INFOCOM | 1 |