EDBT 2026 Demo / reviewers in the wild / expert
Kinman Lei
dblp:380/6574
· DBLP profile ↗
6ranked-venue papers
1as first author
6since 2021 · last 2026
0009-0006-7633-9010ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SYCL++: A Unified Programming Framework for Heterogeneous Supercomputers at Scale
Zitao Shen, Yuyang Jin 0001, Kinman Lei, Zixuan Ma, Zhenchuan Chen, Di Wei, Fei Wang 0096, Ying Liu 0055, Lin Gan 0001, Jidong Zhai |
HPDC | 3 |
| 2026 | RoMeo: Mitigating Dual-dimensional Outliers with Rotated Mixed Precision QuantizationabstractMixed precision quantization has been adopted to accelerate large language models (LLMs) serving by leveraging high-throughput low-precision compute units in GPUs while preserving outliers in higher precision to maintain model accuracy. However, existing methods focus on mitigating single-dimensional channel-wise outliers, leading to model accuracy degradation when scaled to 4-bit precision. Qihao Zhang, Mingliang Tang, Mingshu Zhai, Kinman Lei, Jidong Zhai |
PPoPP | 4 |
| 2025 | FlashTensor: Optimizing Tensor Programs by Leveraging Fine-grained Tensor PropertyabstractDeep neural networks (DNNs) have shown significant effectiveness in natural language processing and video applications. However, DNN models, especially for long-context tasks, introduce extremely large intermediate tensors, producing substantial memory overhead. Although considerable efforts have been made to optimize DNNs, insufficient awareness of tensor properties has hindered effective memory optimization and can lead to inefficient computations in a long-context scenario. Runxin Zhong, Yuyang Jin 0001, Chen Zhang 0001, Kinman Lei, Shuangyu Li, Jidong Zhai |
PPoPP | 4 |
| 2025 | UltraAttn: Efficiently Parallelizing Attention through Hierarchical Context-TilingabstractLong-context comprehension is critical for large language models. Context parallelism and irregular block-sparse attention are keyss to accelerating long-context training and inference. Existing context parallelism suffers from poor scalability due to the striped-like partition pattern, which causes high communication traffic, and the ring-based communication pattern, which limits kernel granularity, reduces device utilization, and incurs redundant communication. Zan Zong, Yuyang Jin 0001, Kinman Lei, Jiaao He, Qigang Yang, Jidong Zhai |
SC | 4 |
| 2025 | mTuner: Accelerating Parameter-Efficient Fine-Tuning on Multi-GPU Servers with Elastic Tensor
Kezhao Huang, Siqi Zhu, Mingshu Zhai, Liyan Zheng 0001, Kinman Lei, Jiaao He, Yuyang Jin 0001, Jidong Zhai |
USENIX ATC | 5 |
| 2024 | PUZZLE: Efficiently Aligning Large Language Models through Light-Weight Context Switch
Kinman Lei, Yuyang Jin 0001, Mingshu Zhai, Kezhao Huang, Haoxing Ye, Jidong Zhai |
USENIX ATC | 1 |