Shenggui Li

dblp:290/1198 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0003-2037-2496ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TetriServe: Efficiently Serving Mixed DiT Workloads
Runyu Lu, Shiqi He, Wenxuan Tan, Shenggui Li, Jeff J. Ma, Ang Chen 0001, Mosharaf Chowdhury
ASPLOS (2)4
2026 SPPO: Making Million-Token LLM Training Practical on Modest GPU Clusters
abstract
In recent years, Large Language Models (LLMs) have exhibited remarkable capabilities, driving advancements in real-world applications. However, training LLMs on increasingly long input sequences imposes significant challenges due to high GPU memory and computational demands. Existing solutions face two key limitations: (1) memory reduction techniques, such as activation recomputation and CPU offloading, compromise training efficiency; (2) distributed parallelism strategies require excessive GPU resources, limiting the scalability of input sequence length.
Qiaoling chen, Shenggui Li, Wei Gao 0064, Peng Sun 0006, Yonggang Wen 0001, Tianwei Zhang 0004
ICS2
2024 GliDe with a CaPE: A Low-Hassle Method to Accelerate Speculative Decoding
abstract
Speculative decoding is a relatively new decoding framework that leverages small and efficient draft models to reduce the latency of LLMs. In this study, we introduce GliDe and CaPE, two low-hassle modifications to vanilla speculative decoding to further improve the decoding speed of a frozen LLM. Specifically, GliDe is a modified draft model architecture that reuses the cached keys and values from the target LLM, while CaPE is a proposal expansion method that uses the draft model’s confidence scores to help select additional candidate tokens for verification. Extensive experiments on different benchmarks demonstrate that our proposed GliDe draft model significantly reduces the expected decoding latency. Additional evaluation using walltime reveals that GliDe can accelerate Vicuna models up to 2.17x and further extend the improvement to 2.61x with CaPE. We will release our code, data, and the trained draft models.
Cunxiao Du, Jing Jiang 0001, Yuanchen Xu, Jiawei Wu 0003, Sicheng Yu, Yongqi Li 0001, Shenggui Li, Liqiang Nie, Zhaopeng Tu
ICML7
2023 Sequence Parallelism: Long Sequence Training from System Perspective
abstract
Transformer achieves promising results on various tasks.However, self-attention suffers from quadratic memory requirements with respect to the sequence length.Existing work focuses on reducing time and space complexity from an algorithm perspective.In this work, we propose sequence parallelism, a memory-efficient parallelism to solve this issue from system perspective instead.Our approach is compatible with most existing parallelisms (e.g., data, pipeline, and tensor parallelism), which means our sequence parallelism makes 4D parallelism possible.More importantly, we no longer require a single device to hold the whole sequence.Besides, using efficient attention with linear complexity, our sequence parallelism enables us to train transformer with infinite long sequence.Specifically, we split the input sequence into multiple chunks and feed each chunk into its corresponding device (i.e., GPU).To compute the attention output, we integrated ring-style communication with self-attention calculation and proposed Ring Self-Attention (RSA).Experiments show that sequence parallelism performs well when scaling with batch size and sequence length.Compared with tensor parallelism, our approach achieved 13.7× and 3.0× maximum batch size and sequence length respectively when scaling up to 64 NVIDIA P100 GPUs.With efficient attention, sequence can handle sequence with over 114K tokens, which is over 27× longer than existing efficient attention works holding the whole sequence on a single device.
Shenggui Li, Fuzhao Xue, Chaitanya Baranwal, Yang You 0001
ACL (1)1
2023 Colossal-AI: A Unified Deep Learning System For Large-Scale Parallel Training
abstract
The success of Transformer models has pushed the deep learning model scale to billions of parameters, but the memory limitation of a single GPU has led to an urgent need for training on multi-GPU clusters. However, the best practice for choosing the optimal parallel strategy is still lacking, as it requires domain expertise in both deep learning and parallel computing. The Colossal-AI system addressed the above challenge by introducing a unified interface to scale your sequential code of model training to distributed environments. It supports parallel training methods such as data, pipeline, tensor, and sequence parallelism and is integrated with heterogeneous training and zero redundancy optimizer. Compared to the baseline system, Colossal-AI can achieve up to 2.76 times training speedup on large-scale models.
Shenggui Li, Hongxin Liu, Zhengda Bian, Jiarui Fang, Haichen Huang, Yang You 0001
ICPP1
2023 Parallel Training of Pre-Trained Models via Chunk-Based Dynamic Memory Management
abstract
The pre-trained model (PTM) is revolutionizing Artificial Intelligence (AI) technology. However, the hardware requirement of PTM training is prohibitively high, making it a game for a small proportion of people. Therefore, we proposed PatrickStar system to lower the hardware requirements of PTMs and make them accessible to everyone. PatrickStar uses the CPU-GPU heterogeneous memory space to store the model data. Different from existing works, we organize the model data in memory chunks and dynamically distribute them in the heterogeneous memory. Guided by the runtime memory statistics collected in a warm-up iteration, chunks are orchestrated efficiently in heterogeneous memory and generate lower CPU-GPU data transmission volume and higher bandwidth utilization. Symbiosis with the Zero Redundancy Optimizer, PatrickStar scales to multiple GPUs on multiple nodes. The system can train tasks on bigger models and larger batch sizes, which cannot be accomplished by existing works. Experimental results show that PatrickStar extends model scales 2.27 and 2.5 times of DeepSpeed, and exhibits significantly higher execution speed. PatricStar also successfully runs the 175B GPT3 training task on a 32 GPU cluster. Our code is available athttps://github.com/Tencent/PatrickStar.
Jiarui Fang, Zilin Zhu, Shenggui Li, Hui Su, Yang Yu 0038, Jie Zhou 0016, Yang You 0001
IEEE Trans. Parallel Distributed Syst.3
2022 Critique of "MemXCT: Memory-Centric X-Ray CT Reconstruction With Massive Parallelization" by SCC Team From Nanyang Technological University
abstract
In this technical report, we focus on reproducing the results reported in the paper “MemXCT: Memory-Centric X-ray CT Reconstruction with Massive Parallelization” [1]. MemXCT is a scalable approach to X-ray Computed Tomography reconstruction which removes redundant computation. We reproduced the single CPU/GPU performance as well as strong scaling experiments. We set up our configurations on Microsoft Azure CycleCloud and have two clusters. One cluster has 4 nodes with 60 CPUs on each node and the other cluster has 4 nodes with 4 NVIDIA V100 GPUs on each node. Both clusters come with InfiniBand. The original author conducted his experiments on Theta and Blue Waters supercomputers. We were able to reproduce part of the results in the original paper, however, failed to produce similar performance on other experiments. This report was submitted as part of the reproducibility challenge in SC20 Student Cluster Competition. Digital artifacts from these experiments are available at: 10.5281/zenodo.5598108.
Shenggui Li, Bu-Sung Lee
IEEE Trans. Parallel Distributed Syst.1
2021 Online evolutionary batch size orchestration for scheduling deep learning workloads in GPU clusters
abstract
Efficient GPU resource scheduling is essential to maximize resource utilization and save training costs for the increasing amount of deep learning workloads in shared GPU clusters. Existing GPU schedulers largely rely on static policies to leverage the performance characteristics of deep learning jobs. However, they can hardly reach optimal efficiency due to the lack of elasticity. To address the problem, we propose ONES, an ONline Evolutionary Scheduler for elastic batch size orchestration. ONES automatically manages the elasticity of each job based on the training batch size, so as to maximize GPU utilization and improve scheduling efficiency. It determines the batch size for each job through an online evolutionary search that can continuously optimize the scheduling decisions. We evaluate the effectiveness of ONES with 64 GPUs on TACC's Longhorn supercomputers. The results show that ONES can outperform the prior deep learning schedulers with a significantly shorter average job completion time.
Zhengda Bian, Shenggui Li, Wei Wang 0225, Yang You 0001
SC2