Weile Luo

dblp:371/1187 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2026
0009-0007-2875-0056ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
abstract
Lossless model compression holds tremendous promise for alleviating the memory and bandwidth bottlenecks in bitexact Large Language Model (LLM) serving. However, existing approaches often result in substantial inference slowdowns due to fundamental design mismatches with GPU architectures: at the kernel level, variable-length bitstreams produced by traditional entropy codecs break SIMT parallelism; at the system level, decoupled pipelines lead to redundant memory traffic. We present ZipServ, a lossless compression framework co-designed for efficient LLM inference. ZipServ introduces Tensor-Core-Aware Triple Bitmap Encoding (TCA-TBE), a novel fixed-length format that enables constant-time, parallel decoding, together with a fused decompression-GEMM (ZipGEMM) kernel that decompresses weights on-the-fly directly into Tensor Core registers. This "load-compressed, compute-decompressed" design eliminates intermediate buffers and maximizes compute intensity. Experiments show that ZipServ reduces the model size by up to 30%, achieves up to 2.21× kernel-level speedup over NVIDIA’s cuBLAS, and expedites end-to-end inference by an average of 1.22× over vLLM. ZipServ is the first lossless compression system that provides both storage savings and substantial acceleration for LLM inference on GPUs.
Ruibo Fan, Xiangrui Yu, Xinglin Pan, Weile Luo, Qiang Wang 0022, Wei Wang 0030, Xiaowen Chu 0001
ASPLOS (2)5
2026 DynSpAttn: Efficient Attention via Dual-Side Dynamic Sparsity on Sparse Tensor Cores
abstract
The high computational complexity of the self-attention mechanism constitutes a primary performance bottleneck in LLM inference. Existing sparse attention mechanisms commonly adopt coarse-grained block sparsity to align with FlashAttention’s tiling and rely on dense Tensor Cores, leaving the potential of emerging hardware Sparse Tensor Cores (SpTCs) and semi-structured sparsity largely untapped. We present DynSpAttn, a dynamic sparse attention mechanism co-designed with NVIDIA Sparse Tensor Cores. DynSpAttn introduces a dual-side 2:4 structured sparsity strategy that prunes both the Query and Score matrices, thereby transforming the dominant matrix multiplications in attention into sparse matrix multiplications (SpMMs) executable on SpTCs. To realize this transformation, DynSpAttn incorporates lightweight in-register pruners and a shuffle-free operand remapping scheme within a fully fused, I/O-aware CUDA kernel. Evaluations on RTX 4090 and L20 GPUs show that DynSpAttn achieves up to 1.70 × kernel-level preformance improvement over FlashAttention and 1.58 × end-to-end inference speedup. These results demonstrate that co-designing semi-structured sparsity with hardware support across the full attention pipeline provides a practical and efficient solution for LLM inference.
Xiangrui Yu, Ruibo Fan, Weile Luo, Gu Gong, Xiaowen Chu 0001
ICS4
2026 ROME: Maximizing GPU Efficiency for All-Pairs Shortest Path via Taming Fine-Grained Irregularities
abstract
All-Pairs Shortest Path (APSP), a fundamental problem in graph analytics, can be solved efficiently by reducing the computational workload through vertex reordering. However, it fails on GPUs due to fine-grained granularity, shape, and dependency irregularities, which cause severe hardware underutilization. We introduce ROME, a system that tames these irregularities by spatially restructuring computation into regularized workloads and temporally overlapping them with an asynchronous pipeline. ROME achieves 14.7-244.5× speedup over the state-of-the-art multicore CPU solution and 11.2-338.0× speedup over the state-of-the-art GPU solution. Notably, our results achieve mostly above 20% and up to 34.7% of peak min-plus OPs across all tested graphs.
Weile Luo, Yuhan Chen 0008, Xiangrui Yu, Qiang Wang 0022, Ruibo Fan, Hongyuan Liu 0002, Xiaowen Chu 0001
PPoPP1
2024 Benchmarking and Dissecting the Nvidia Hopper GPU Architecture
abstract
Graphics processing units (GPUs) are continually evolving to cater to the computational demands of contemporary general-purpose workloads, particularly those driven by artificial intelligence (AI) utilizing deep learning techniques. A substantial body of studies have been dedicated to dissecting the microarchitectural metrics characterizing diverse GPU generations, which helps researchers understand the hardware details and leverage them to optimize the GPU programs. However, the latest Hopper GPUs present a set of novel attributes, including new tensor cores supporting FP8, DPX, and distributed shared memory. Their details still remain mysterious in terms of performance and operational characteristics. In this research, we propose an extensive benchmarking study focused on the Hopper GPU. The objective is to unveil its microarchitectural intricacies through an examination of the new instruction-set architecture (ISA) of Nvidia GPUs and the utilization of new CUDA APIs. Our approach involves two main aspects. Firstly, we conduct conventional latency and throughput comparison benchmarks across the three most recent GPU architectures, namely Hopper, Ada, and Ampere. Secondly, we delve into a comprehensive discussion and benchmarking of the latest Hopper features, encompassing the Hopper DPX dynamic programming (DP) instruction set, distributed shared memory, and the availability of FP8 tensor cores. The microbenchmarking results we present offer a deeper understanding of the novel GPU AI function units and programming features introduced by the Hopper architecture. This newfound understanding is expected to greatly facilitate software optimization and modeling efforts for GPU architectures. To the best of our knowledge, this study makes the first attempt to demystify the tensor core performance and programming instruction sets unique to Hopper GPUs.
Weile Luo, Ruibo Fan, Dayou Du, Qiang Wang 0022, Xiaowen Chu 0001
IPDPS1
2024 DSO: A GPU Energy Efficiency Optimizer by Fusing Dynamic and Static Information
abstract
Increased reliance on graphics processing units (GPUs) for high-intensity computing tasks raises challenges regarding energy consumption. To address this issue, dynamic voltage and frequency scaling (DVFS) has emerged as a promising technique for conserving energy while maintaining the quality of service (QoS) of GPU applications. However, existing solutions using DVFS are hindered by inefficiency or inaccuracy as they depend either on dynamic or static information respectively, which prevents them from being adopted to practical power management schemes. To this end, we propose a novel energy efficiency optimizer, called DSO, to explore a light weight solution that leverages both dynamic and static information to model and optimize the GPU energy efficiency. DSO firstly proposes a novel theoretical energy efficiency model which reflects the DVFS roofline phenomenon and considers the tradeoff between performance and energy. Then it applies machine learning techniques to predict the parameters of the above model with both GPU kernel runtime metrics and static code features. Experiments on modern DVFS-enabled GPUs indicate that DSO can enhance energy efficiency by 19% whilst maintaining performance within a 5% loss margin.
Laiyi Li, Weile Luo, Bingqiang Wang
IWQoS3