EDBT 2026 Demo / reviewers in the wild / expert
Xiangrui Yu
dblp:352/1552
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2026
0009-0005-2478-1512ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless CompressionabstractLossless model compression holds tremendous promise for alleviating the memory and bandwidth bottlenecks in bitexact Large Language Model (LLM) serving. However, existing approaches often result in substantial inference slowdowns due to fundamental design mismatches with GPU architectures: at the kernel level, variable-length bitstreams produced by traditional entropy codecs break SIMT parallelism; at the system level, decoupled pipelines lead to redundant memory traffic. We present ZipServ, a lossless compression framework co-designed for efficient LLM inference. ZipServ introduces Tensor-Core-Aware Triple Bitmap Encoding (TCA-TBE), a novel fixed-length format that enables constant-time, parallel decoding, together with a fused decompression-GEMM (ZipGEMM) kernel that decompresses weights on-the-fly directly into Tensor Core registers. This "load-compressed, compute-decompressed" design eliminates intermediate buffers and maximizes compute intensity. Experiments show that ZipServ reduces the model size by up to 30%, achieves up to 2.21× kernel-level speedup over NVIDIA’s cuBLAS, and expedites end-to-end inference by an average of 1.22× over vLLM. ZipServ is the first lossless compression system that provides both storage savings and substantial acceleration for LLM inference on GPUs. Ruibo Fan, Xiangrui Yu, Xinglin Pan, Weile Luo, Qiang Wang 0022, Wei Wang 0030, Xiaowen Chu 0001 |
ASPLOS (2) | 2 |
| 2026 | DynSpAttn: Efficient Attention via Dual-Side Dynamic Sparsity on Sparse Tensor CoresabstractThe high computational complexity of the self-attention mechanism constitutes a primary performance bottleneck in LLM inference. Existing sparse attention mechanisms commonly adopt coarse-grained block sparsity to align with FlashAttention’s tiling and rely on dense Tensor Cores, leaving the potential of emerging hardware Sparse Tensor Cores (SpTCs) and semi-structured sparsity largely untapped. We present DynSpAttn, a dynamic sparse attention mechanism co-designed with NVIDIA Sparse Tensor Cores. DynSpAttn introduces a dual-side 2:4 structured sparsity strategy that prunes both the Query and Score matrices, thereby transforming the dominant matrix multiplications in attention into sparse matrix multiplications (SpMMs) executable on SpTCs. To realize this transformation, DynSpAttn incorporates lightweight in-register pruners and a shuffle-free operand remapping scheme within a fully fused, I/O-aware CUDA kernel. Evaluations on RTX 4090 and L20 GPUs show that DynSpAttn achieves up to 1.70 × kernel-level preformance improvement over FlashAttention and 1.58 × end-to-end inference speedup. These results demonstrate that co-designing semi-structured sparsity with hardware support across the full attention pipeline provides a practical and efficient solution for LLM inference. Xiangrui Yu, Ruibo Fan, Weile Luo, Gu Gong, Xiaowen Chu 0001 |
ICS | 1 |
| 2026 | ROME: Maximizing GPU Efficiency for All-Pairs Shortest Path via Taming Fine-Grained IrregularitiesabstractAll-Pairs Shortest Path (APSP), a fundamental problem in graph analytics, can be solved efficiently by reducing the computational workload through vertex reordering. However, it fails on GPUs due to fine-grained granularity, shape, and dependency irregularities, which cause severe hardware underutilization. We introduce ROME, a system that tames these irregularities by spatially restructuring computation into regularized workloads and temporally overlapping them with an asynchronous pipeline. ROME achieves 14.7-244.5× speedup over the state-of-the-art multicore CPU solution and 11.2-338.0× speedup over the state-of-the-art GPU solution. Notably, our results achieve mostly above 20% and up to 34.7% of peak min-plus OPs across all tested graphs. Weile Luo, Yuhan Chen 0008, Xiangrui Yu, Qiang Wang 0022, Ruibo Fan, Hongyuan Liu 0002, Xiaowen Chu 0001 |
PPoPP | 3 |
| 2025 | SpInfer: Leveraging Low-Level Sparsity for Efficient Large Language Model Inference on GPUsabstractLarge Language Models (LLMs) have demonstrated remarkable capabilities, but their immense scale poses significant challenges in terms of both memory and computational costs. While unstructured pruning offers promising solutions by introducing sparsity to reduce resource requirements, realizing its benefits in LLM inference remains elusive. This is primarily due to the storage overhead of indexing non-zero elements and the inefficiency of sparse matrix multiplication (SpMM) kernels at low sparsity levels (around 50%). In this paper, we present SpInfer, a high-performance framework tailored for sparsified LLM inference on GPUs. SpInfer introduces Tensor-Core-Aware Bitmap Encoding (TCA-BME), a novel sparse format that minimizes indexing overhead by leveraging efficient bitmap-based indexing, optimized for GPU Tensor Core architectures. Furthermore, SpInfer integrates an optimized SpMM kernel with Shared Memory Bitmap Decoding (SMBD) and asynchronous pipeline design to enhance computational efficiency. Experimental results show that SpInfer significantly outperforms state-of-the-art SpMM implementations (up to 2.14× and 2.27× over Flash-LLM and SparTA, respectively) across a range of sparsity levels (30% to 70%), with substantial improvements in both memory efficiency and end-to-end inference speed (up to 1.58×). SpInfer outperforms highly optimized cuBLAS at sparsity levels as low as 30%, marking the first effective translation of unstructured pruning's theoretical advantages into practical performance gains for LLM inference. Ruibo Fan, Xiangrui Yu, Peijie Dong, Gu Gong, Qiang Wang 0022, Wei Wang 0030, Xiaowen Chu 0001 |
EuroSys | 2 |
| 2023 | Balancing Computation and Communication in Distributed Sparse Matrix-Vector MultiplicationabstractSparse Matrix-Vector Multiplication (SpMV) is a fundamental operation in a number of scientific and engineering problems. When the sparse matrices processed are large enough, distributed memory systems should be used to accelerate SpMV. At present, the optimization techniques for distributed SpMV mainly focus on reordering through graph or hypergraph partitioning. However, although the reordering could reduce the amount of communications in general, there are still load balancing challenges in computations and communications on distributed platforms that are not well addressed. In this paper, we propose two strategies to optimize SpMV on distributed clusters: (1) resizing the number of row blocks on the nodes for balancing the amount of computations, and (2) adjusting the column number of the diagonal blocks for balancing tasks and reducing communications among compute nodes. The experimental results show that compared with the classic distributed SpMV implementation and its variant reordered with graph partitioning, our algorithm achieves on average 77.20x and 5.18x (up to 460.52x and 27.50x) speedups, respectively. Also, our method bring on average 19.56x (up to 48.49x) speedup over a recently proposed hybrid distributed SpMV algorithm. In addition, our algorithm achieves obviously better scalability over these existing distributed SpMV methods. Hongli Mi, Xiangrui Yu, Xiaosong Yu, Shuangyuan Wu, Weifeng Liu 0002 |
CCGrid | 2 |