Kinman Lei

dblp:380/6574 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2026
0009-0006-7633-9010ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-author · 6 since 2021
YearPublicationVenuePosition
2026 SYCL++: A Unified Programming Framework for Heterogeneous Supercomputers at Scale
Zitao Shen, Yuyang Jin 0001, Kinman Lei, Zixuan Ma, Zhenchuan Chen, Di Wei, Fei Wang 0096, Ying Liu 0055, Lin Gan 0001, Jidong Zhai
HPDC3
2026 RoMeo: Mitigating Dual-dimensional Outliers with Rotated Mixed Precision Quantization
abstract
Mixed precision quantization has been adopted to accelerate large language models (LLMs) serving by leveraging high-throughput low-precision compute units in GPUs while preserving outliers in higher precision to maintain model accuracy. However, existing methods focus on mitigating single-dimensional channel-wise outliers, leading to model accuracy degradation when scaled to 4-bit precision.
Qihao Zhang, Mingliang Tang, Mingshu Zhai, Kinman Lei, Jidong Zhai
PPoPP4
2025 FlashTensor: Optimizing Tensor Programs by Leveraging Fine-grained Tensor Property
abstract
Deep neural networks (DNNs) have shown significant effectiveness in natural language processing and video applications. However, DNN models, especially for long-context tasks, introduce extremely large intermediate tensors, producing substantial memory overhead. Although considerable efforts have been made to optimize DNNs, insufficient awareness of tensor properties has hindered effective memory optimization and can lead to inefficient computations in a long-context scenario.
Runxin Zhong, Yuyang Jin 0001, Chen Zhang 0001, Kinman Lei, Shuangyu Li, Jidong Zhai
PPoPP4
2025 UltraAttn: Efficiently Parallelizing Attention through Hierarchical Context-Tiling
abstract
Long-context comprehension is critical for large language models. Context parallelism and irregular block-sparse attention are keyss to accelerating long-context training and inference. Existing context parallelism suffers from poor scalability due to the striped-like partition pattern, which causes high communication traffic, and the ring-based communication pattern, which limits kernel granularity, reduces device utilization, and incurs redundant communication.
Zan Zong, Yuyang Jin 0001, Kinman Lei, Jiaao He, Qigang Yang, Jidong Zhai
SC4
2025 mTuner: Accelerating Parameter-Efficient Fine-Tuning on Multi-GPU Servers with Elastic Tensor
Kezhao Huang, Siqi Zhu, Mingshu Zhai, Liyan Zheng 0001, Kinman Lei, Jiaao He, Yuyang Jin 0001, Jidong Zhai
USENIX ATC5
2024 PUZZLE: Efficiently Aligning Large Language Models through Light-Weight Context Switch
Kinman Lei, Yuyang Jin 0001, Mingshu Zhai, Kezhao Huang, Haoxing Ye, Jidong Zhai
USENIX ATC1