EDBT 2026 Demo / reviewers in the wild / expert
Xinglin Pan
dblp:273/3352
· DBLP profile ↗
22ranked-venue papers
5as first author
21since 2021 · last 2026
0000-0002-1172-9935ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Systems, architecture and hardware · 7 · 2 first-author · 7 since 2021Computer networks · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless CompressionabstractLossless model compression holds tremendous promise for alleviating the memory and bandwidth bottlenecks in bitexact Large Language Model (LLM) serving. However, existing approaches often result in substantial inference slowdowns due to fundamental design mismatches with GPU architectures: at the kernel level, variable-length bitstreams produced by traditional entropy codecs break SIMT parallelism; at the system level, decoupled pipelines lead to redundant memory traffic. We present ZipServ, a lossless compression framework co-designed for efficient LLM inference. ZipServ introduces Tensor-Core-Aware Triple Bitmap Encoding (TCA-TBE), a novel fixed-length format that enables constant-time, parallel decoding, together with a fused decompression-GEMM (ZipGEMM) kernel that decompresses weights on-the-fly directly into Tensor Core registers. This "load-compressed, compute-decompressed" design eliminates intermediate buffers and maximizes compute intensity. Experiments show that ZipServ reduces the model size by up to 30%, achieves up to 2.21× kernel-level speedup over NVIDIA’s cuBLAS, and expedites end-to-end inference by an average of 1.22× over vLLM. ZipServ is the first lossless compression system that provides both storage savings and substantial acceleration for LLM inference on GPUs. Ruibo Fan, Xiangrui Yu, Xinglin Pan, Weile Luo, Qiang Wang 0022, Wei Wang 0030, Xiaowen Chu 0001 |
ASPLOS (2) | 3 |
| 2026 | HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap
Wenxiang Lin, Xinglin Pan, Lin Zhang 0059, Shaohuai Shi, Xuan Wang 0002, Xiaowen Chu 0001 |
INFOCOM | 2 |
| 2026 | Compass: Dissecting Communication and Computation Operators for Efficient LLM TrainingabstractOverlapping communication and computation operators is a common practice to hide communication overheads, accelerating large language models (LLMs) training on GPU clusters. Existing systems achieve this through either intra-operator fusion (IntraFusion), which packs operators into a single large kernel, or inter-operator decomposition (InterDecom), which splits a tensor into multiple parts for pipelined execution. However, current IntraFusion methods underutilize network topology, causing suboptimal bandwidth usage on multi-GPU systems, while InterDecom struggles to determine the optimal number of decomposed parts for peak performance. To address these issues, we introduce Compass, which employs systematic optimization and comprehensive modeling. First, we design a novel IntraFusion algorithm leveraging double-ring communications to maximize bandwidth utilization in hybrid NVLink-PCIe systems, achieving 1.5x-2.5x speedups. Second, we develop a decomposition model that mathematically derives the optimal tensor decomposition degree for InterDecom, improving performance by up to 1.3x. Finally, we develop a unified performance framework that accurately determines the best strategy for different scenarios. We validate Compass through extensive evaluation across 288 configurations and end-to-end experiments on real-world applications. The results demonstrate that Compass consistently selects the optimal strategy, achieving up to a 1.42x end-to-end speedup compared to the Megatron-LM baseline. Guangyu Xiang, Lin Zhang 0059, Haoxuan Yu, Xinglin Pan, Shaohuai Shi, Xiaowen Chu 0001 |
INFOCOM | 4 |
| 2026 | ZipCCL: Efficient Lossless Data Compression of Communication Collectives for Accelerating LLM TrainingabstractCommunication has emerged as a critical bottleneck in the distributed training of large language models (LLMs). While numerous approaches have been proposed to reduce communication overhead, the potential of lossless compression has remained largely underexplored since compression and decompression typically consume larger overheads than the benefits of reduced communication traffic. We observe that the communication data, including activations, gradients and parameters, during training often follows a near-Gaussian distribution, which is a key feature for data compression. Thus, we introduce ZipCCL, a lossless compressed communication library of collectives for LLM training. ZipCCL is equipped with our novel techniques: (1) theoretically grounded exponent coding that exploits the Gaussian distribution of LLM tensors to accelerate compression without expensive online statistics, (2) GPU-optimized compression and decompression kernels that carefully design memory access patterns and pipeline using communication-aware data layout, and (3) adaptive communication strategies that dynamically switch collective operations based on workload patterns and system characteristics. Evaluated on a 64-GPU cluster using both mixture-of-experts and dense transformer models, ZipCCL reduces communication time by up to 1.35X and achieves end-to-end training speedups of up to 1.18X without any impact on model quality. Wenxiang Lin, Xinglin Pan, Ruibo Fan, Shaohuai Shi, Xiaowen Chu 0001 |
SIGCOMM | 2 |
| 2025 | FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts ModelsabstractRecent large language models (LLMs) have tended to leverage sparsity to reduce computations, employing the sparsely activated mixture-of-experts (MoE) technique. MoE introduces four modules, including token routing, token communication, expert computation, and expert parallelism, that impact model quality and training efficiency. To enable ver- satile usage of MoE models, we introduce FSMoE, a flexible training system optimizing task scheduling with three novel techniques: 1) Unified abstraction and online profiling of MoE modules for task scheduling across various MoE implementations. 2) Co-scheduling intra-node and inter-node communications with computations to minimize communication overheads. 3) To support near-optimal task scheduling, we design an adaptive gradient partitioning method for gradient aggregation and a schedule to adaptively pipeline communications and computations. We conduct extensive experiments with configured MoE layers and real-world MoE models on two GPU clusters. Experimental results show that 1) our FSMoE supports four popular types of MoE routing functions and is more efficient than existing implementations (with up to a 1.42× speedup), and 2) FSMoE outperforms the state-of-the-art MoE training systems (DeepSpeed-MoE and Tutel) by 1.18×-1.22× on 1458 MoE layers and 1.19×-3.01× on real-world MoE models based on GPT-2 and Mixtral using a popular routing function. In this work, we present a flexible training system named FSMoE to optimize task scheduling. To achieve this goal: 1) we design unified abstraction and online profiling of MoE modules across various MoE implementations, 2) we co-schedule intra-node and inter-node communications with computations to minimize communication overhead, and 3) we design an adaptive gradient partitioning method for gradient aggregation and a schedule to adaptively pipeline communications and computations. Experimental results on two clusters up to 48 GPUs show that our FSMoE outperforms the state-of-the-art MoE training systems (DeepSpeed-MoE and Tutel) with speedups of 1.18x-1.22x on 1458 customized MoE layers and 1.19x-3.01x on real-world MoE models based on GPT-2 and Mixtral. Xinglin Pan, Wenxiang Lin, Lin Zhang 0059, Shaohuai Shi, Zhenheng Tang, Rui Wang 0172, Bo Li 0001, Xiaowen Chu 0001 |
ASPLOS (1) | 1 |
| 2025 | ScheInfer: Efficient Inference of Large Language Models with Task Scheduling on Moderate GPUs
Wenxiang Lin, Xinglin Pan, Shaohuai Shi, Xuan Wang 0002, Xiaowen Chu 0001 |
Euro-Par (3) | 2 |
| 2025 | Mast: Efficient Training of Mixture-of-Experts Transformers with Task Pipelining and OrderingabstractThe utilization of the sparsely activated mixture-of-experts (MoE) technique has enabled the expansion of modern large language models (LLMs) to trillion-level sizes while maintaining a sub-linear increase in computations. This involves equipping an MoE layer with multiple experts, where only one or two experts are activated for each input data. However, the dynamic activation of MoE experts introduces extensive communications, limiting the scaling efficiency of distributed systems. In this work, we propose Mast to efficiently train MoE models by pipelining and re-ordering communication and computation tasks to effectively hide communication costs. Specifically, we first propose to overlap tasks in both attention layers and MoE layers. Then we theoretically analyze the task overlaps between communications and computations, identifying the inefficiencies of existing schedules. We then develop an optimization formulation to determine a near-optimal order for task pipelining with the objective of minimizing iteration time. We conduct extensive experiments on two 32-GPU clusters employing 432 configured MoE layers and three real-world MoE models based on BERT, GPT-2 and Mistral. The experimental results demonstrate that Mast outperforms state-of-the-art MoE training systems (DeepSpeed-MoE, Tutel, PipeMoE and CoCoNet) with an average speedup 1.13 ×-1.43 × on the MoE models. Wenxiang Lin, Xinglin Pan, Shaohuai Shi, Xuan Wang 0002, Bo Li 0001, Xiaowen Chu 0001 |
ICDCS | 2 |
| 2025 | Mitigating Contention in Stream Multiprocessors for Pipelined Mixture of Experts: An SM-Aware Scheduling ApproachabstractSparsely activated Mixture-of-Experts (MoEs) models have become prominent in Large Language Models (LLMs) due to their ability to expand model capacity without proportional increases in computation. MoE layers feature multiple experts, with only a few activated per sample, enhancing model performance across various domains such as natural language generation and translation. The dynamic activation of MoE experts introduces extensive communications in distributed training. However, this dynamic activation creates communication challenges in distributed training. While previous work attempted to pipeline computation and communication through input chunking, we found that these tasks compete for Stream Multiprocessors (SMs) on GPUs, making the scheduling ineffective. In this paper, we update the optimization problem to minimize training time while accounting for SM contentions. We develop performance models for computation and communication tasks to identify MoE layer bottlenecks. By delaying GEMM launching and splitting GEMM operations, we enable communication to preempt SMs, enhancing overall efficiency and more stability. Xinglin Pan, Rui Wang 0172, Wenxiang Lin, Shaohuai Shi, Xiaowen Chu 0001 |
ICDCS | 1 |
| 2025 | CORE: CORrelation-Guided Feature Enhancement for Few-Shot Image ClassificationabstractFew-shot classification aims to adapt classifiers trained on base classes to novel classes with a few shots. However, the limited amount of training data is often inadequate to represent the intraclass variations in novel classes. This can result in biased estimation of the feature distribution, which in turn results in inaccurate decision boundaries, especially when the support data are outliers. To address this issue, we propose a feature enhancement method called CORrelation-guided feature Enrichment that generates improved features for novel classes using weak supervision from the base classes. The proposed CORrelation-guided feature Enhancement (CORE) method utilizes an autoencoder (AE) architecture but incorporates classification information into its latent space. This design allows the CORE to generate more discriminative features while discarding irrelevant content information. After being trained on base classes, CORE's generative ability can be transferred to novel classes that are similar to those in the base classes. By using these generative features, we can reduce the estimation bias of the class distribution, which makes few-shot learning (FSL) less sensitive to the selection of support data. Our method is generic and flexible and can be used with any feature extractor and classifier. It can be easily integrated into existing FSL approaches. Experiments with different backbones and classifiers show that our proposed method consistently outperforms existing methods on various widely used benchmarks. Xinglin Pan, Jingquan Wang, Wenjie Pei, Qing Liao 0001, Zenglin Xu |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Multi-task Domain Adaptation for Language Grounding with 3D Objects
Penglei Sun, Yaoxian Song, Xinglin Pan, Peijie Dong, Xiaofei Yang 0002, Qiang Wang 0022, Zhixu Li, Tiefeng Li, Xiaowen Chu 0001 |
ECCV (34) | 3 |
| 2024 | ScheMoE: An Extensible Mixture-of-Experts Distributed Training System with Tasks SchedulingabstractIn recent years, large-scale models can be easily scaled to trillions of parameters with sparsely activated mixture-of-experts (MoE), which significantly improves the model quality while only requiring a sub-linear increase in computational costs. However, MoE layers require the input data to be dynamically routed to a particular GPU for computing during distributed training. The highly dynamic property of data routing and high communication costs in MoE make the training system low scaling efficiency on GPU clusters. In this work, we propose an extensible and efficient MoE training system, ScheMoE, which is equipped with several features. 1) ScheMoE provides a generic scheduling framework that allows the communication and computation tasks in training MoE models to be scheduled in an optimal way. 2) ScheMoE integrates our proposed novel all-to-all collective which better utilizes intra- and inter-connect bandwidths. 3) ScheMoE supports easy extensions of customized all-to-all collectives and data compression approaches while enjoying our scheduling algorithm. Extensive experiments are conducted on a 32-GPU cluster and the results show that ScheMoE outperforms existing state-of-the-art MoE systems, Tutel and Faster-MoE, by 9%-30%. Shaohuai Shi, Xinglin Pan, Qiang Wang 0022, Chengjian Liu, Xiaozhe Ren, Zhongzhe Hu, Bo Li 0001, Xiaowen Chu 0001 |
EuroSys | 2 |
| 2024 | Pruner-Zero: Evolving Symbolic Pruning Metric From Scratch for Large Language ModelsabstractDespite the remarkable capabilities, Large Language Models (LLMs) face deployment challenges due to their extensive size. Pruning methods drop a subset of weights to accelerate, but many of them require retraining, which is prohibitively expensive and computationally demanding. Recently, post-training pruning approaches introduced novel metrics, enabling the pruning of LLMs without retraining. However, these metrics require the involvement of human experts and tedious trial and error. To efficiently identify superior pruning metrics, we develop an automatic framework for searching symbolic pruning metrics using genetic programming. In particular, we devise an elaborate search space encompassing the existing pruning metrics to discover the potential symbolic pruning metric. We propose an opposing operation simplification strategy to increase the diversity of the population. In this way, Pruner-Zero allows auto-generation of symbolic pruning metrics. Based on the searched results, we explore the correlation between pruning metrics and performance after pruning and summarize some principles. Extensive experiments on LLaMA and LLaMA-2 on language modeling and zero-shot tasks demonstrate that our Pruner-Zero obtains superior performance than SOTA post-training pruning methods. Code at: https://github.com/pprp/Pruner-Zero. Peijie Dong, Lujun Li 0001, Zhenheng Tang, Xiang Liu 0001, Xinglin Pan, Qiang Wang 0022, Xiaowen Chu 0001 |
ICML | 5 |
| 2024 | Parm: Efficient Training of Large Sparsely-Activated Models with Dedicated SchedulesabstractSparsely-activated Mixture-of-Expert (MoE) layers have found practical applications in enlarging the model size of large-scale foundation models, with only a sub-linear increase in computation demands. Despite the wide adoption of hybrid parallel paradigms like model parallelism, expert parallelism, and expert-sharding parallelism (i.e., MP+EP+ESP) to support MoE model training on GPU clusters, the training efficiency is hindered by communication costs introduced by these parallel paradigms. To address this limitation, we propose Parm, a system that accelerates MP+EP+ESP training by designing two dedicated schedules for placing communication tasks. The proposed schedules eliminate redundant computations and communications and enable overlaps between intra-node and inter-node communications, ultimately reducing the overall training time. As the two schedules are not mutually exclusive, we provide comprehensive theoretical analyses and derive an automatic and accurate solution to determine which schedule should be applied in different scenarios. Experimental results on an 8-GPU server and a 32-GPU cluster demonstrate that Parm outperforms the state-of-the-art MoE training system, DeepSpeed-MoE, achieving 1.13× to 5.77× speedup on 1296 manually configured MoE layers and approximately 3× improvement on two real-world MoE models based on BERT and GPT-2. Xinglin Pan, Wenxiang Lin, Shaohuai Shi, Xiaowen Chu 0001, Weinong Sun, Bo Li 0001 |
INFOCOM | 1 |
| 2024 | Should We Really Edit Language Models? On the Evaluation of Edited Language ModelsabstractModel editing has become an increasingly popular alternative for efficiently updating knowledge within language models.
Current methods mainly focus on reliability, generalization, and locality, with many methods excelling across these criteria.
Some recent works disclose the pitfalls of these editing methods such as knowledge distortion or conflict. However, the general abilities of post-edited language models remain unexplored.
In this paper, we perform a comprehensive evaluation on various editing methods and different language models, and have following findings.
(1) Existing editing methods lead to inevitable performance deterioration on general benchmarks, indicating that existing editing methods maintain the general abilities of the model within only a few dozen edits.
When the number of edits is slightly large, the intrinsic knowledge structure of the model is disrupted or even completely damaged.
(2) Instruction-tuned models are more robust to editing, showing less performance drop on general knowledge after editing.
(3) Language model with large scale is more resistant to editing compared to small model.
(4) The safety of the edited model, is significantly weakened, even for those safety-aligned models.
Our findings indicate that current editing methods are only suitable for small-scale knowledge updates within language models, which motivates further research on more practical and reliable editing methods. Xiang Liu 0001, Zhenheng Tang, Peijie Dong, Xinglin Pan, Xiaowen Chu 0001 |
NeurIPS | 6 |
| 2023 | Generalized Category Discovery with Clustering Assignment Consistency
Xiangli Yang, Xinglin Pan, Irwin King, Zenglin Xu |
ICONIP (5) | 2 |
| 2023 | PipeMoE: Accelerating Mixture-of-Experts through Adaptive PipeliningabstractLarge models have attracted much attention in the AI area. The sparsely activated mixture-of-experts (MoE) technique pushes the model size to a trillion-level with a sub-linear increase of computations as an MoE layer can be equipped with many separate experts, but only one or two experts need to be trained for each input data. However, the feature of dynamically activating experts of MoE introduces extensive communications in distributed training. In this work, we propose PipeMoE to adaptively pipeline the communications and computations in MoE to maximally hide the communication time. Specifically, we first identify the root reason why a higher pipeline degree does not always achieve better performance in training MoE models. Then we formulate an optimization problem that aims to minimize the training iteration time. To solve this problem, we build performance models for computation and communication tasks in MoE and develop an optimal solution to determine the pipeline degree such that the iteration time is minimal. We conduct extensive experiments with 174 typical MoE layers and two real-world NLP models on a 64-GPU cluster. Experimental results show that our PipeMoE almost always chooses the best pipeline degree and outperforms state-of-the-art MoE training systems by 5%-77% in training time. Shaohuai Shi, Xinglin Pan, Xiaowen Chu 0001, Bo Li 0001 |
INFOCOM | 2 |
| 2023 | Multi-behavior recommendation based on intent learning
Xinglin Pan, Mingxin Gan |
Multim. Syst. | 1 |
| 2023 | RegNet: Self-Regulated Network for Image ClassificationabstractThe ResNet and its variants have achieved remarkable successes in various computer vision tasks. Despite its success in making gradient flow through building blocks, the information communication of intermediate layers of blocks is ignored. To address this issue, in this brief, we propose to introduce a regulator module as a memory mechanism to extract complementary features of the intermediate layers, which are further fed to the ResNet. In particular, the regulator module is composed of convolutional recurrent neural networks (RNNs) [e.g., convolutional long short-term memories (LSTMs) or convolutional gated recurrent units (GRUs)], which are shown to be good at extracting spatio-temporal information. We named the new regulated network as regulated residual network (RegNet). The regulator module can be easily implemented and appended to any ResNet architecture. Experimental results on three image classification datasets have demonstrated the promising performance of the proposed architecture compared with the standard ResNet, squeeze-and-excitation ResNet, and other state-of-the-art architectures. Yu Pan 0005, Xinglin Pan, Steven C. H. Hoi, Zhang Yi 0001, Zenglin Xu |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | ENIGMA: Low-Latency and Privacy-Preserving Edge Inference on Heterogeneous Neural Network AcceleratorsabstractTime-efficient artificial intelligence (AI) service has recently witnessed increasing interest from academia and industry due to the urgent needs in massive smart applications such as self-driving cars, virtual reality, high-resolution video streaming, etc. Existing solutions to reduce AI latency, like edge computing and heterogeneous neural-network accelerators (NNAs), face high risk of privacy leakage. To achieve both low-latency and privacy-preserving purposes on edge servers (e.g., NNAs), this paper proposes ENIGMA that can exploit the trusted execution environment (TEE) and heterogeneous NNAs of edge servers for edge inference. The low-latency is supported by a new ahead-of-time analysis framework for analyzing the linearity of multilayer neural networks, which automatically slices forward-graph and assigns sub-graphs to TEE or NNA. To avoid privacy leakage issue, we then introduce a pre-forwarded cipher generation (PFCG) scheme for computing linear sub-forward-graphs on NNA. The input data is encrypted to ciphertext that can be computed directly by linear sub-graphs, and the output can be decrypted to obtain the correct output. To enable non-linear computation of sub-graphs on TEE, we use ring-cache and automatic vectorization optimization to address the memory limitation of TEE. Qualitative analysis and quantitative experiments on GPU, NPU and TPU demonstrate that ENIGMA is not only compatible with heterogeneous NNAs, but also can avoid leakages of private features with latency as low as 50-milliseconds. Qiushi Li 0002, Ju Ren 0001, Xinglin Pan, Yue-Zhi Zhou, Yaoxue Zhang |
ICDCS | 3 |
| 2022 | Alleviating the Sample Selection Bias in Few-shot Learning by Removing Projection to the CentroidabstractFew-shot learning (FSL) targets at generalization of vision models towards unseen tasks without sufficient annotations. Despite the emergence of a number of few-shot learning methods, the sample selection bias problem, i.e., the sensitivity to the limited amount of support data, has not been well understood. In this paper, we find that this problem usually occurs when the positions of support samples are in the vicinity of task centroid—the mean of all class centroids in the task. This motivates us to propose an extremely simple feature transformation to alleviate this problem, dubbed Task Centroid Projection Removing (TCPR). TCPR is applied directly to all image features in a given task, aiming at removing the dimension of features along the direction of the task centroid. While the exact task centoid cannot be accurately obtained from limited data, we estimate it using base features that are each similar to one of the support features. Our method effectively prevents features from being too close to the task centroid. Extensive experiments over ten datasets from different domains show that TCPR can reliably improve classification accuracy across various feature extractors, training algorithms and datasets. The code has been made available at https://github.com/KikimorMay/FSL-TCBR. Xu Luo 0003, Xinglin Pan, Wenjie Pei, Zenglin Xu |
NeurIPS | 3 |
| 2022 | AFINet: Attentive Feature Integration Networks for image classification
Xinglin Pan, Yu Pan 0005, Liangjian Wen, Wenxiang Lin, Hongguang Fu, Zenglin Xu |
Neural Networks | 1 |
| 2020 | InvisibleFL: Federated Learning over Non-Informative Intermediate Updates against Multimedia Privacy LeakagesabstractIn cloud and edge networks, federated learning involves training statistical models over decentralized data, where servers aggregate models through intermediate updates trained from clients. By utilizing private and local data it improves quality of personalized services and reduces user's concern for privacy. However, federated learning still leaks multimedia features through trained intermediate updates and thereby is not privacy-preserving for multimedia. Existing techniques applied from secure community attempt to avoid multimedia features leakages for federated learning but yet cannot address issues of privacy. In this paper, we propose a privacy-preserving solution that avoids multimedia privacy leakages in federated learning. Firstly, we devise a novel encryption scheme called Non-Informative Transformation (NIT) for federated aggregation to eliminates residual multimedia features in intermediate updates. Based on the scheme, we then propose Just-Learn-over-Ciphertext (JLoC) mechanism for federated learning, which includes three stages in each model iteration. The Encrypt stage encrypts intermediate updates and makes it non-informative distribution at clients. The Aggregate stage performs model aggregation without decryption at servers. Specifically, this stage just computes over ciphertext, and its output of aggregation also keeps non-informative. The Decrypt stage converts non-informative outputs of aggregation to available parameters for the next iteration at clients. Moreover, we implement a prototype and conduct experiments to evaluate its privacy and performance on real devices. The experimental results demonstrate that our methods can defend against potential attacks for multimedia privacy leakages without accuracy loss in commercial off-the-shelf products. Qiushi Li 0002, Wenwu Zhu 0001, Chao Wu 0002, Xinglin Pan, Fan Yang 0134, Yue-Zhi Zhou, Yaoxue Zhang |
ACM Multimedia | 4 |