EDBT 2026 Demo / reviewers in the wild / expert
Xuegui Zheng
dblp:357/2100
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2026
0009-0005-4635-7309ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in ProductionabstractWe present MegaScale-MoE, a production system tailored for the efficient training of large-scale mixture-of-experts (MoE) models. MoE emerges as a promising architecture to scale large language models (LLMs) to unprecedented sizes, thereby enhancing model performance. However, existing MoE training systems experience a degradation in training efficiency, exacerbated by the escalating scale of MoE models and the continuous evolution of hardware. Chao Jin 0007, Ziheng Jiang, Zhihao Bai, Juncai Liu, Xiang Li 0067, Ningxin Zheng, Qi Huang 0001, Wen Heng, Yiyuan Ma, Wenlei Bao, Size Zheng 0001, Xuegui Zheng, Yanghua Peng, Haibin Lin, Xuanzhe Liu, Xin Jin 0008, Xin Liu 0086 |
EuroSys | 15 |
| 2026 | UniEP: Unified Expert-Parallel MegaKernel MoE for LLM TrainingabstractAs LLM training grows increasingly resource-intensive and expert parallelism (EP) becomes essential for scaling MoE models, EP optimizations are widely adopted in production frameworks like Megatron-LM. Existing solutions often rely on ad-hoc, complex kernels that lack adaptability across diverse optimization configurations and frequently neglect numerical stability, failing to meet the strict precision requirements of large-scale training. Size Zheng 0001, Xuegui Zheng, Li-Wen Chang, Jidong Zhai |
HPDC | 2 |
| 2026 | Tetris: Efficient Long-context LLM Serving with Chunkwise Dynamic Sequence Parallelism
Xuegui Zheng, Yijin Guan, Size Zheng 0001, Li-Wen Chang, Shufan Liu, Xin Liu 0086, Guangyu Sun 0003 |
ISCA | 3 |
| 2024 | Tetris: Accelerating Sparse Convolution by Exploiting Memory Reuse on GPUabstractConvolutional neural networks (CNNs) have achieved remarkable success in various application fields. Although model compression techniques mitigate the ever-increasing resource demands of large CNN models, the compressed models usually exhibit irregular memory access and unstructured sparsity, which are difficult for dominant operators such as sparse convolution to achieve expected performance speedup on popular inference platforms such as GPU. In this paper, we propose Tetris, an efficient sparse convolution approach optimized for GPU. Tetris first fully exploits the input reuse opportunity of sparse convolution to reduce the memory accesses to global memory. It then adopts a stride packed filter (SPF) format and a bank-sensing reorganization scheme to eliminate the irregular memory accesses caused by unstructured sparsity. It also leverages a filter group reorder technique to address load imbalance among threads, and a parameter tuning method to determine the optimal parameters of the sparse convolution implementation. The experiment results show that Tetris outperforms dense/sparse convolution libraries and cutting-edge implementations with promising performance speedup. Xuegui Zheng, Hailong Yang 0002, Zhongzhi Luan, Depei Qian 0001 |
PPoPP | 2 |
| 2024 | Towards optimized tensor code generation for deep learning on sunway many-core processor
Mingzhen Li 0001, Changxi Liu, Jianjin Liao, Xuegui Zheng, Hailong Yang 0002, Rujun Sun, Lin Gan 0001, Guangwen Yang 0002, Zhongzhi Luan, Depei Qian 0001 |
Frontiers Comput. Sci. | 4 |