VLDB 2026 Research / reviewers in the wild / expert
Wenxiang Lin
dblp:273/2988
· DBLP profile ↗
11ranked-venue papers
4as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Computer networks · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap
Wenxiang Lin, Xinglin Pan, Lin Zhang 0059, Shaohuai Shi, Xuan Wang 0002, Xiaowen Chu 0001 |
INFOCOM | 1 |
| 2026 | ZipCCL: Efficient Lossless Data Compression of Communication Collectives for Accelerating LLM TrainingabstractCommunication has emerged as a critical bottleneck in the distributed training of large language models (LLMs). While numerous approaches have been proposed to reduce communication overhead, the potential of lossless compression has remained largely underexplored since compression and decompression typically consume larger overheads than the benefits of reduced communication traffic. We observe that the communication data, including activations, gradients and parameters, during training often follows a near-Gaussian distribution, which is a key feature for data compression. Thus, we introduce ZipCCL, a lossless compressed communication library of collectives for LLM training. ZipCCL is equipped with our novel techniques: (1) theoretically grounded exponent coding that exploits the Gaussian distribution of LLM tensors to accelerate compression without expensive online statistics, (2) GPU-optimized compression and decompression kernels that carefully design memory access patterns and pipeline using communication-aware data layout, and (3) adaptive communication strategies that dynamically switch collective operations based on workload patterns and system characteristics. Evaluated on a 64-GPU cluster using both mixture-of-experts and dense transformer models, ZipCCL reduces communication time by up to 1.35X and achieves end-to-end training speedups of up to 1.18X without any impact on model quality. Wenxiang Lin, Xinglin Pan, Ruibo Fan, Shaohuai Shi, Xiaowen Chu 0001 |
SIGCOMM | 1 |
| 2026 | CoLOR-DP: Conjugate Low-Rank Differential Privacy for Structure-Aware LoRA Fine-Tuning
Kai Zhang 0074, Wenxiang Lin, Pei-Wei Tsai, Xin Yuan 0004, Minhui Xue 0001 |
WWW | 3 |
| 2026 | K-TCDP: A Temporal Correlated DP Mechanism for LoRA Supervised Fine-Tuning
Kai Zhang 0074, Wenxiang Lin, Pei-Wei Tsai, Xin Yuan 0004, Minhui Xue 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts ModelsabstractRecent large language models (LLMs) have tended to leverage sparsity to reduce computations, employing the sparsely activated mixture-of-experts (MoE) technique. MoE introduces four modules, including token routing, token communication, expert computation, and expert parallelism, that impact model quality and training efficiency. To enable ver- satile usage of MoE models, we introduce FSMoE, a flexible training system optimizing task scheduling with three novel techniques: 1) Unified abstraction and online profiling of MoE modules for task scheduling across various MoE implementations. 2) Co-scheduling intra-node and inter-node communications with computations to minimize communication overheads. 3) To support near-optimal task scheduling, we design an adaptive gradient partitioning method for gradient aggregation and a schedule to adaptively pipeline communications and computations. We conduct extensive experiments with configured MoE layers and real-world MoE models on two GPU clusters. Experimental results show that 1) our FSMoE supports four popular types of MoE routing functions and is more efficient than existing implementations (with up to a 1.42× speedup), and 2) FSMoE outperforms the state-of-the-art MoE training systems (DeepSpeed-MoE and Tutel) by 1.18×-1.22× on 1458 MoE layers and 1.19×-3.01× on real-world MoE models based on GPT-2 and Mixtral using a popular routing function. In this work, we present a flexible training system named FSMoE to optimize task scheduling. To achieve this goal: 1) we design unified abstraction and online profiling of MoE modules across various MoE implementations, 2) we co-schedule intra-node and inter-node communications with computations to minimize communication overhead, and 3) we design an adaptive gradient partitioning method for gradient aggregation and a schedule to adaptively pipeline communications and computations. Experimental results on two clusters up to 48 GPUs show that our FSMoE outperforms the state-of-the-art MoE training systems (DeepSpeed-MoE and Tutel) with speedups of 1.18x-1.22x on 1458 customized MoE layers and 1.19x-3.01x on real-world MoE models based on GPT-2 and Mixtral. Xinglin Pan, Wenxiang Lin, Lin Zhang 0059, Shaohuai Shi, Zhenheng Tang, Rui Wang 0172, Bo Li 0001, Xiaowen Chu 0001 |
ASPLOS (1) | 2 |
| 2025 | ScheInfer: Efficient Inference of Large Language Models with Task Scheduling on Moderate GPUs
Wenxiang Lin, Xinglin Pan, Shaohuai Shi, Xuan Wang 0002, Xiaowen Chu 0001 |
Euro-Par (3) | 1 |
| 2025 | Mast: Efficient Training of Mixture-of-Experts Transformers with Task Pipelining and OrderingabstractThe utilization of the sparsely activated mixture-of-experts (MoE) technique has enabled the expansion of modern large language models (LLMs) to trillion-level sizes while maintaining a sub-linear increase in computations. This involves equipping an MoE layer with multiple experts, where only one or two experts are activated for each input data. However, the dynamic activation of MoE experts introduces extensive communications, limiting the scaling efficiency of distributed systems. In this work, we propose Mast to efficiently train MoE models by pipelining and re-ordering communication and computation tasks to effectively hide communication costs. Specifically, we first propose to overlap tasks in both attention layers and MoE layers. Then we theoretically analyze the task overlaps between communications and computations, identifying the inefficiencies of existing schedules. We then develop an optimization formulation to determine a near-optimal order for task pipelining with the objective of minimizing iteration time. We conduct extensive experiments on two 32-GPU clusters employing 432 configured MoE layers and three real-world MoE models based on BERT, GPT-2 and Mistral. The experimental results demonstrate that Mast outperforms state-of-the-art MoE training systems (DeepSpeed-MoE, Tutel, PipeMoE and CoCoNet) with an average speedup 1.13 ×-1.43 × on the MoE models. Wenxiang Lin, Xinglin Pan, Shaohuai Shi, Xuan Wang 0002, Bo Li 0001, Xiaowen Chu 0001 |
ICDCS | 1 |
| 2025 | Mitigating Contention in Stream Multiprocessors for Pipelined Mixture of Experts: An SM-Aware Scheduling ApproachabstractSparsely activated Mixture-of-Experts (MoEs) models have become prominent in Large Language Models (LLMs) due to their ability to expand model capacity without proportional increases in computation. MoE layers feature multiple experts, with only a few activated per sample, enhancing model performance across various domains such as natural language generation and translation. The dynamic activation of MoE experts introduces extensive communications in distributed training. However, this dynamic activation creates communication challenges in distributed training. While previous work attempted to pipeline computation and communication through input chunking, we found that these tasks compete for Stream Multiprocessors (SMs) on GPUs, making the scheduling ineffective. In this paper, we update the optimization problem to minimize training time while accounting for SM contentions. We develop performance models for computation and communication tasks to identify MoE layer bottlenecks. By delaying GEMM launching and splitting GEMM operations, we enable communication to preempt SMs, enhancing overall efficiency and more stability. Xinglin Pan, Rui Wang 0172, Wenxiang Lin, Shaohuai Shi, Xiaowen Chu 0001 |
ICDCS | 3 |
| 2025 | SP-MoE: Expediting Mixture-of-Experts Training with Optimized Pipelining PlanningabstractSparsely activated Mixture-of-Experts (MoE) has emerged as a key technique to expand the size of Transformer-based large language models (LLMs) while maintaining low computational costs. However, MoE layers require to route the input data to distributed devices, incurring significant communication latency. Existing studies have primarily focused on alleviating this problem by overlapping computation and communication tasks within a single MoE layer, which fails to achieve sufficient overlap and results in limited performance gains. In this work, we introduce an orthogonal partitioning dimension from existing task-parallel methods by leveraging the autoregressive nature of causal Transformer-based LLMs, i.e. partitioning tasks along the sequence dimension. This provides more flexible and efficient overlaps among tasks from both non-MoE and MoE layers. To this end, we propose an efficient MoE training approach, SP-MoE, with two innovative designs. 1) It incorporates non-MoE layers into the overlapping with not only the current MoE layer but also the preceding MoE layer, thereby facilitating more efficient training; 2) It identifies the optimal combination of pipeline degrees for non-MoE and MoE layers and devises the best scheduling plans for load-imbalanced non-MoE and uniform MoE layers to achieve the goal of minimizing the total training latency. Extensive experiments conducted on two GPU clusters demonstrate that SP-MoE can effectively identify the optimal combination of pipeline degrees and achieve 16.1% - 34.3% reduction in training latency compared to three state-of-the-art MoE systems. Ne Wang, Wenxiang Lin, Lin Zhang 0059, Shaohuai Shi, Ruiting Zhou, Bo Li 0001 |
INFOCOM | 2 |
| 2024 | Parm: Efficient Training of Large Sparsely-Activated Models with Dedicated SchedulesabstractSparsely-activated Mixture-of-Expert (MoE) layers have found practical applications in enlarging the model size of large-scale foundation models, with only a sub-linear increase in computation demands. Despite the wide adoption of hybrid parallel paradigms like model parallelism, expert parallelism, and expert-sharding parallelism (i.e., MP+EP+ESP) to support MoE model training on GPU clusters, the training efficiency is hindered by communication costs introduced by these parallel paradigms. To address this limitation, we propose Parm, a system that accelerates MP+EP+ESP training by designing two dedicated schedules for placing communication tasks. The proposed schedules eliminate redundant computations and communications and enable overlaps between intra-node and inter-node communications, ultimately reducing the overall training time. As the two schedules are not mutually exclusive, we provide comprehensive theoretical analyses and derive an automatic and accurate solution to determine which schedule should be applied in different scenarios. Experimental results on an 8-GPU server and a 32-GPU cluster demonstrate that Parm outperforms the state-of-the-art MoE training system, DeepSpeed-MoE, achieving 1.13× to 5.77× speedup on 1296 manually configured MoE layers and approximately 3× improvement on two real-world MoE models based on BERT and GPT-2. Xinglin Pan, Wenxiang Lin, Shaohuai Shi, Xiaowen Chu 0001, Weinong Sun, Bo Li 0001 |
INFOCOM | 2 |
| 2022 | AFINet: Attentive Feature Integration Networks for image classification
Xinglin Pan, Yu Pan 0005, Liangjian Wen, Wenxiang Lin, Hongguang Fu, Zenglin Xu |
Neural Networks | 5 |