VLDB 2026 Research / reviewers in the wild / expert
Xiaozhe Ren
dblp:248/7679
· DBLP profile ↗
17ranked-venue papers
1as first author
17since 2021 · last 2026
0000-0002-0432-5510ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 1 first-author · 12 since 2021Computer networks · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BitDP: Ultra-low-bit Communication for Data Parallelism in LLM TrainingabstractTraining large language models (LLMs) with billions of parameters on trillion-token datasets requires distributed data parallelism at increasingly large scales, where gradient synchronization becomes a communication bottleneck, especially in bandwidth-constrained environments. Although gradient quantization presents a promising solution, it faces two key challenges: maintaining training stability and accuracy for transformer architectures and adapting to modern distributed communication systems. In this paper, we propose BitDP, an ultra-low-bit gradient quantization system that reduces communication costs by up to 32× while preserving model accuracy with less than 1% performance degradation. Our approach achieves numerical stability for large transformer models and seamlessly integrates with existing infrastructures. We evaluate BitDP's effectiveness across various LLM sizes, architectures and optimizers. The results demonstrate significant training efficiency improvements while maintaining convergence quality, establishing BitDP as a scalable and reliable solution for real-world LLM training at industrial scales. Xiaozhe Ren, Qiong Luo 0001 |
AAAI | 1 |
| 2026 | An Efficient Edge-Cloud Collaboration System With Foundational Models for Open-Set IoT ApplicationsabstractArtificial intelligence (AI) models have been widely deployed on edge devices, enabling various IoT applications. However, lightweight on-device AI models on resource-limited edge devices hinder their adaptability to dynamic environments and tasks. Despite the superior generalization capabilities of recently developed Foundation Models (FMs), utilizing their extensive knowledge on the resource-constrained edge platforms remains unexplored. In this work, we introduce DeepEdgeFM, an edge-cloud collaborative system with FMs that enables open-set learning, simultaneously achieving generalizability and efficiency for IoT applications. DeepEdgeFM employs a spatiotemporalaware semantic customization approach that leverages spatial, temporal, and domain-specific knowledge from FMs to continuously customize edge models using unlabeled sensor data in emerging IoT environments. Meanwhile, DeepEdgeFM utilizes a dynamic model switching strategy to selectively query the knowledge of FMs based on sensor-data uncertainty and real-time network fluctuations. We implement DeepEdgeFM on five FMs and multi-modal large language models (MLLMs), covering four types of sensor data modalities. We evaluate DeepEdgeFM on two edge platforms, five public datasets, and two self-collected datasets covering both indoor and outdoor real-world environments. The results show that DeepEdgeFM outperforms state-ofthe- art baselines, achieving up to an 18.6% accuracy gain and a 38.6 Bufang Yang, Wenrui Lu, Lixing He, Neiwen Ling, Zhenyu Yan 0002, Guoliang Xing, Xian Shuai, Xiaozhe Ren, Xin Jiang 0002 |
IEEE Trans. Mob. Comput. | 8 |
| 2025 | DAPE V2: Process Attention Score as Feature Map for Length ExtrapolationabstractThe attention mechanism is a fundamental component of the Transformer model, contributing to interactions among distinct tokens. In general, the attention scores are determined simply by the key-query products. However, this work’s occasional trial (combining DAPE and NoPE) of including additional MLPs on attention scores without position encoding indicates that the classical key-query multiplication may limit the performance of Transformers. In this work, we conceptualize attention as a feature map and apply the convolution operator (for neighboring attention scores across different heads) to mimic the processing methods in computer vision. Specifically, the main contribution of this paper is identifying and interpreting the Transformer length extrapolation problem as a result of the limited expressiveness of the naive query and key dot product, and we successfully translate the length extrapolation issue into a well-understood feature map processing problem, which is called Convolutional Data-Adaptive Position Encoding (CDAPE).The novel insight, which can be adapted to various attention-related models, reveals that the current Transformer architecture has the potential for further evolution. Extensive experiments demonstrate that treating attention as a feature map and applying convolution as a processing method significantly enhances Transformer performance. Chuanyang Zheng, Yihang Gao, Jiankai Sun, Jingyao Li 0001, Minbin Huang, Xiaozhe Ren, Michael Kwok-Po Ng, Zhenguo Li, Yu Li 0006 |
ACL (1) | 8 |
| 2025 | Self-Adjust SoftmaxabstractChuanyang Zheng, Yihang Gao, Guoxuan Chen, Han Shi, Jing Xiong, Xiaozhe Ren, Chao Huang, Zhenguo Li, Yu Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Chuanyang Zheng, Yihang Gao, Guoxuan Chen, Xiaozhe Ren, Zhenguo Li, Yu Li 0006 |
EMNLP | 6 |
| 2025 | SepLLM: Accelerate Large Language Models by Compressing One Segment into One SeparatorabstractLarge Language Models (LLMs) have exhibited exceptional performance across a spectrum of natural language processing tasks. However, their substantial sizes pose considerable challenges, particularly in computational demands and inference speed, due to their quadratic complexity. In this work, we have identified a key pattern: certain seemingly meaningless separator tokens (i.e., punctuations) contribute disproportionately to attention scores compared to semantically meaningful tokens. This observation suggests that information of the segments between these separator tokens can be effectively condensed into the separator tokens themselves without significant information loss. Guided by this insight, we introduce SepLLM, a plug-and-play framework that accelerates inference by compressing these segments and eliminating redundant tokens. Additionally, we implement efficient kernels for training acceleration. Experimental results across training-free, training-from-scratch, and post-training settings demonstrate SepLLM's effectiveness. Notably, using the Llama-3-8B backbone, SepLLM achieves over 50% reduction in KV cache on the GSM8K-CoT benchmark while maintaining comparable performance. Furthermore, in streaming settings, SepLLM effectively processes sequences of up to 4 million tokens or more while maintaining consistent language modeling capabilities. Guoxuan Chen, Yihang Gao, Xiaozhe Ren, Zhenguo Li, Weiyang Liu |
ICML | 5 |
| 2025 | DeepDiver: Adaptive Web-Search Intensity Scaling via Reinforcement LearningabstractInformation seeking demands iterative evidence gathering and reflective reasoning, yet large language models (LLMs) still struggle with it in open-web question answering. Existing prompting and supervised fine-tuning (SFT) methods remain fixed by prompt rules or training corpora, and are usually benchmarked only on well-structured wiki sources, limiting real-world adaptability. We introduce $\textbf{WebPuzzle}$, a 24k-sample training and 275-sample test benchmark that evaluates information seeking on the live internet, across both wiki and open-domain queries. Leveraging 7k WebPuzzle instances, we develop $\textbf{DeepDiver}$, a reinforcement-learning (RL) framework that cultivates $\textbf{Search Intensity Scaling (SIS)}$—an emergent ability to escalate search frequency and depth instead of settling on overconfident, under-evidenced answers. With SIS, Qwen2.5-7B-Instruct and Pangu-7B-Reasoner attain performance on real-web tasks comparable to the 671B-parameter DeepSeek-R1. We detail DeepDiver’s curriculum from cold-start SFT to a well designed RL procedure, and show that its seeking policy generalized from closed-ended queries to open-ended generation such as long-form writing. Our results advance adaptive information seeking in LLMs and provide a rigorous benchmark for future work. Haochen Tan, Chuqiao Kuang, Hanting Chen, Xiaozhe Ren, Yasheng Wang, Lifeng Shang |
NeurIPS | 6 |
| 2025 | TaskSense: A Translation-like Approach for Tasking Heterogeneous Sensor Systems with LLMsabstractAn increasing number of environments, such as smart homes and factories, are being equipped with multiple sensor systems to enable diverse intelligent applications. However, most existing sensor coordination systems require manually predefined rules, limiting their ability to handle flexible and complex tasks. While recent approaches leverage large language models (LLMs) to interact with external APIs, they struggle to fully understand the capabilities and data dependencies of practical sensor systems. This paper introduces TaskSense, a novel system that coordinates multiple sensor systems in response to users' complex queries. TaskSense introduces a sensor language that automatically translates the capabilities and data dependencies of sensor systems into vocabularies and grammar rules that can be understood by LLMs. It then interprets user intentions into executable task plans for sensor systems using this sensor language in combination with LLMs. Meanwhile, TaskSense checks the solvability of user queries and verifies the correctness of task plan dependencies. To further enhance robustness, TaskSense incorporates a dynamic plan execution mechanism that adjusts plans based on real-time feedback from sensor data availability, data quality and execution results. TaskSense is deployed on real-world smart home systems, utilizing six popular LLMs. The system is evaluated across 4 scenarios involving 9 types of sensor systems, over 60 APIs, 170 tasks and 5 types of data modalities. Results show that TaskSense achieves up to 2× higher planning accuracy and a 75% increase in answer accuracy using the similar amount of tokens compared with baseline approaches. Kaiwei Liu 0001, Bufang Yang, Lilin Xu, Yunqi Guo, Guoliang Xing, Xian Shuai, Xiaozhe Ren, Xin Jiang 0002, Zhenyu Yan 0002 |
SenSys | 7 |
| 2024 | PIXART-Σ: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation
Junsong Chen, Chongjian Ge, Enze Xie, Yue Wu 0002, Lewei Yao, Xiaozhe Ren, Zhongdao Wang, Ping Luo 0002, Huchuan Lu, Zhenguo Li |
ECCV (32) | 6 |
| 2024 | ScheMoE: An Extensible Mixture-of-Experts Distributed Training System with Tasks SchedulingabstractIn recent years, large-scale models can be easily scaled to trillions of parameters with sparsely activated mixture-of-experts (MoE), which significantly improves the model quality while only requiring a sub-linear increase in computational costs. However, MoE layers require the input data to be dynamically routed to a particular GPU for computing during distributed training. The highly dynamic property of data routing and high communication costs in MoE make the training system low scaling efficiency on GPU clusters. In this work, we propose an extensible and efficient MoE training system, ScheMoE, which is equipped with several features. 1) ScheMoE provides a generic scheduling framework that allows the communication and computation tasks in training MoE models to be scheduled in an optimal way. 2) ScheMoE integrates our proposed novel all-to-all collective which better utilizes intra- and inter-connect bandwidths. 3) ScheMoE supports easy extensions of customized all-to-all collectives and data compression approaches while enjoying our scheduling algorithm. Extensive experiments are conducted on a 32-GPU cluster and the results show that ScheMoE outperforms existing state-of-the-art MoE systems, Tutel and Faster-MoE, by 9%-30%. Shaohuai Shi, Xinglin Pan, Qiang Wang 0022, Chengjian Liu, Xiaozhe Ren, Zhongzhe Hu, Bo Li 0001, Xiaowen Chu 0001 |
EuroSys | 5 |
| 2024 | DAPE: Data-Adaptive Positional Encoding for Length ExtrapolationabstractPositional encoding plays a crucial role in transformers, significantly impact- ing model performance and length generalization. Prior research has introduced absolute positional encoding (APE) and relative positional encoding (RPE) to distinguish token positions in given sequences. However, both APE and RPE remain fixed after model training regardless of input data, limiting their adaptability and flexibility. Hence, we expect that the desired positional encoding should be data-adaptive and can be dynamically adjusted with the given attention. In this paper, we propose a Data-Adaptive Positional Encoding (DAPE) method, which dynamically and semantically adjusts based on input context and learned fixed priors. Experimental validation on real-world datasets (Arxiv, Books3, and CHE) demonstrates that DAPE enhances model performances in terms of trained length and length generalization, where the improvements are statistically significant. The model visualization suggests that our model can keep both local and anti-local information. Finally, we successfully train the model on sequence length 128 and achieve better performance at evaluation sequence length 8192, compared with other static positional encoding methods, revealing the benefit of the adaptive positional encoding method. Chuanyang Zheng, Yihang Gao, Minbin Huang, Jingyao Li 0001, Xiaozhe Ren, Michael Kwok-Po Ng, Zhenguo Li, Yu Li 0006 |
NeurIPS | 7 |
| 2024 | Poster Abstract: Tasking Heterogeneous Sensor Systems with LLMsabstractDespite the extensive use of sensors enabling intelligent applications, the complementary potential of co-existing sensor systems is often not fully utilized, limiting more advanced applications. This paper introduces a novel solution using Large Language Models (LLMs) to coordinate sensor systems for handling complex user queries. It defines a sensor language for sensor systems, including vocabulary set and grammar rules, analogous to natural language components, enabling LLMs to translate user intentions into sensor coordination plans. Preliminary results show that our approach significantly outperforms the existing solution at plan generation, execution and response generation stages. Kaiwei Liu 0001, Bufang Yang, Lilin Xu, Yunqi Guo, Neiwen Ling, Guoliang Xing, Xian Shuai, Xiaozhe Ren, Xin Jiang 0002, Zhenyu Yan 0002 |
SenSys | 9 |
| 2023 | CAME: Confidence-guided Adaptive Memory Efficient OptimizationabstractAdaptive gradient methods, such as Adam and LAMB, have demonstrated excellent performance in the training of large language models.Nevertheless, the need for adaptivity requires maintaining second-moment estimates of the per-parameter gradients, which entails a high cost of extra memory overheads.To solve this problem, several memory-efficient optimizers (e.g., Adafactor) have been proposed to obtain a drastic reduction in auxiliary memory usage, but with a performance penalty.In this paper, we first study a confidence-guided strategy to reduce the instability of existing memory efficient optimizers.Based on this strategy, we propose CAME to simultaneously achieve two goals: fast convergence as in traditional adaptive methods, and low memory usage as in memory-efficient methods.Extensive experiments demonstrate the training stability and superior performance of CAME across various NLP tasks such as BERT and GPT-2 training.Notably, for BERT pre-training on the large batch size of 32,768, our proposed optimizer attains faster convergence and higher accuracy compared with the Adam optimizer.The implementation of CAME is publicly available 1 . Xiaozhe Ren, Zangwei Zheng, Zhuo Jiang, Xin Jiang 0002, Yang You 0001 |
ACL (1) | 2 |
| 2023 | A Study on Transformer Configuration and Training ObjectiveabstractTransformer-based models have delivered impressive results on many tasks, particularly vision and language tasks. In many model training situations, conventional configurations are often adopted. For example, we usually set the base model with hidden size (i.e. model width) to be 768 and the number of transformer layers (i.e. model depth) to be 12. In this paper, we revisit these conventional configurations by studying the the relationship between transformer configuration and training objective. We show that the optimal transformer configuration is closely related to the training objective. Specifically, compared with the simple classification objective, the masked autoencoder is effective in alleviating the over-smoothing issue in deep transformer training. Based on this finding, we propose “Bamboo”, a notion of using deeper and narrower transformer configurations, for masked autoencoder training. On ImageNet, with such a simple change in configuration, the re-designed Base-level transformer achieves 84.2% top-1 accuracy and outperforms SoTA models like MAE by $0.9%$. On language tasks, re-designed model outperforms BERT with the default setting by 1.1 points on average, on GLUE benchmark with 8 datasets. Fuzhao Xue, Jianghai Chen, Aixin Sun, Xiaozhe Ren, Zangwei Zheng, Xiao-Xin He, Yongming Chen, Xin Jiang 0002, Yang You 0001 |
ICML | 4 |
| 2023 | Response Length Perception and Sequence Scheduling: An LLM-Empowered LLM Inference PipelineabstractLarge language models (LLMs) have revolutionized the field of AI, demonstrating unprecedented capacity across various tasks. However, the inference process for LLMs comes with significant computational costs. In this paper, we propose an efficient LLM inference pipeline that harnesses the power of LLMs. Our approach begins by tapping into the potential of LLMs to accurately perceive and predict the response length with minimal overhead. By leveraging this information, we introduce an efficient sequence scheduling technique that groups queries with similar response lengths into micro-batches. We evaluate our approach on real-world instruction datasets using the LLaMA-based model, and our results demonstrate an impressive 86% improvement in inference throughput without compromising effectiveness. Notably, our method is orthogonal to other inference acceleration techniques, making it a valuable addition to many existing toolkits (e.g., FlashAttention, Quantization) for LLM inference. Zangwei Zheng, Xiaozhe Ren, Fuzhao Xue, Xin Jiang 0002, Yang You 0001 |
NeurIPS | 2 |
| 2023 | EdgeFM: Leveraging Foundation Model for Open-set Learning on the EdgeabstractDeep Learning (DL) models have been widely deployed on IoT devices with the help of advancements in DL algorithms and chips. However, the limited resources of edge devices make these on-device DL models hard to be generalizable to diverse environments and tasks. Although the recently emerged foundation models (FMs) show impressive generalization power, how to effectively leverage the rich knowledge of FMs on resource-limited edge devices is still not explored. In this paper, we propose EdgeFM, a novel edge-cloud cooperative system with open-set recognition capability. EdgeFM selectively uploads unlabeled data to query the FM on the cloud and customizes the specific knowledge and architectures for edge models. Meanwhile, EdgeFM conducts dynamic model switching at run-time taking into account both data uncertainty and dynamic network variations, which ensures the accuracy always close to the original FM. We implement EdgeFM using two FMs on two edge platforms. We evaluate EdgeFM on three public datasets and two self-collected datasets. Results show that EdgeFM can reduce the end-to-end latency up to 3.2x and achieve 34.3% accuracy increase compared with the baseline. Bufang Yang, Lixing He, Neiwen Ling, Zhenyu Yan 0002, Guoliang Xing, Xian Shuai, Xiaozhe Ren, Xin Jiang 0002 |
SenSys | 7 |
| 2022 | AutoBERT-Zero: Evolving BERT Backbone from ScratchabstractTransformer-based pre-trained language models like BERT and its variants have recently achieved promising performance in various natural language processing (NLP) tasks. However, the conventional paradigm constructs the backbone by purely stacking the manually designed global self-attention layers, introducing inductive bias and thus leads to sub-optimal. In this work, we make the first attempt to automatically discover novel pre-trained language model (PLM) backbone on a flexible search space containing the most fundamental operations from scratch. Specifically, we propose a well-designed search space which (i) contains primitive math operations in the intra-layer level to explore novel attention structures, and (ii) leverages convolution blocks to be the supplementary for attentions in the inter-layer level to better learn local dependency. To enhance the efficiency for finding promising architectures, we propose an Operation-Priority Neural Architecture Search (OP-NAS) algorithm, which optimizes both the search algorithm and evaluation of candidate models. Specifically, we propose Operation-Priority (OP) evolution strategy to facilitate model search via balancing exploration and exploitation. Furthermore, we design a Bi-branch Weight-Sharing (BIWS) training strategy for fast model evaluation. Extensive experiments show that the searched architecture (named AutoBERT-Zero) significantly outperforms BERT and its variants of different model capacities in various downstream tasks, proving the architecture's transfer and scaling abilities. Remarkably, AutoBERT-Zero-base outperforms RoBERTa-base (using much more data) and BERT-large (with much larger model size) by 2.4 and 1.4 higher score on GLUE test set. Jiahui Gao 0002, Hang Xu 0004, Xiaozhe Ren, Philip L. H. Yu, Xiaodan Liang, Xin Jiang 0002, Zhenguo Li |
AAAI | 4 |
| 2021 | SparseBERT: Rethinking the Importance Analysis in Self-attentionabstractTransformer-based models are popularly used in natural language processing (NLP). Its core component, self-attention, has aroused widespread interest. To understand the self-attention mechanism, a direct method is to visualize the attention map of a pre-trained model. Based on the patterns observed, a series of efficient Transformers with different sparse attention masks have been proposed. From a theoretical perspective, universal approximability of Transformer-based models is also recently proved. However, the above understanding and analysis of self-attention is based on a pre-trained model. To rethink the importance analysis in self-attention, we study the significance of different positions in attention matrix during pre-training. A surprising result is that diagonal elements in the attention map are the least important compared with other attention positions. We provide a proof showing that these diagonal elements can indeed be removed without deteriorating model performance. Furthermore, we propose a Differentiable Attention Mask (DAM) algorithm, which further guides the design of the SparseBERT. Extensive experiments verify our interesting findings and illustrate the effect of the proposed algorithm. Jiahui Gao 0002, Xiaozhe Ren, Hang Xu 0004, Xiaodan Liang, Zhenguo Li, James T. Kwok |
ICML | 3 |