EDBT 2026 Demo / reviewers in the wild / expert
Zhengyang Lu 0003
dblp:254/3011-3
· DBLP profile ↗
5ranked-venue papers
2as first author
4since 2021 · last 2023
0000-0002-1540-0678ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | HASpMV: Heterogeneity-Aware Sparse Matrix-Vector Multiplication on Modern Asymmetric Multicore ProcessorsabstractSparse matrix-vector multiplication (SpMV) is a fundamental routine in computational science and engineering. Its optimization methods on various homogeneous parallel processors, such as CPUs and GPUs, received much attention. Recently, asymmetric multicore processors (AMPs) have heterogeneous performance and efficient cores (e.g., P- and E-cores from Intel and Apple, or Big.LITTLE cores from ARM), or cores with different cache structures (e.g., cores with/without 3D V-Cache from AMD) are becoming one of the mainstream in desktop and workstation computers. However, there lacks heterogeneity-aware research on accelerating SpMV on AMPs.We in this paper propose a parallel algorithm called heterogeneity-aware SpMV (HASpMV) for improving the performance of SpMV on the latest 12th- and 13th-Gen AMPs from Intel and Ryzen 9 AMPs from AMD. We first micro-benchmark bandwidth and multi-/single-core SpMV to collect performance characteristics and to motivate our algorithm design, and then develop several optimization techniques to assign workloads between the two types of cores for achieving significantly better cache locality and load balancing. The experimental results show that compared to the latest version of the Intel oneMKL library and the open-source works CSR5 and merge-SpMV, HASpMV achieves an average speedup of 2.61x, 2.31x, and 3.73x (up to 5.23x, 4.46x, and 8.23x) on the i9-12900KF processor. On the i9-13900KF processor, HASpMV achieves an average speedup of 3.17x, 1.52x, and 2.23x (up to 9.46x, 5.31x, and 4.49x). Additionally, when comparing AMD Ryzen 9 7950X3D and 7950X AMPs, HASpMV brings an average speedup of 1.43x, 1.3x, and 1.29x (up to 6.28x, 7.8x, and 10.8x) over AMD Optimizing CPU Libraries (AOCL), CSR5, and merge-SpMV, respectively. Helin Cheng, Zhengyang Lu 0003, Yuechen Lu, Weifeng Liu 0002 |
CLUSTER | 3 |
| 2023 | TileSpTRSV: a tiled algorithm for parallel sparse triangular solve on GPUs
Zhengyang Lu 0003, Weifeng Liu 0002 |
CCF Trans. High Perform. Comput. | 1 |
| 2022 | TileSpGEMM: a tiled algorithm for parallel sparse general matrix-matrix multiplication on GPUsabstractSparse general matrix-matrix multiplication (SpGEMM) is one of the most fundamental building blocks in sparse linear solvers, graph processing frameworks and machine learning applications. The existing parallel approaches for shared memory SpGEMM mostly use the row-row style with possibly good parallelism. However, because of the irregularity in sparsity structures, the existing row-row methods often suffer from three problems: (1) load imbalance, (2) high global space complexity and unsatisfactory data locality, and (3) sparse accumulator selection. Yuyao Niu, Zhengyang Lu 0003, Haonan Ji, Shuhui Song, Zhou Jin 0001, Weifeng Liu 0002 |
PPoPP | 2 |
| 2021 | TileSpMV: A Tiled Algorithm for Sparse Matrix-Vector Multiplication on GPUsabstractWith the extensive use of GPUs in modern supercomputers, accelerating sparse matrix-vector multiplication (SpMV) on GPUs received much attention in the last couple of decades. A number of techniques, such as increasing utilization of wide vector units, reducing load imbalance and selecting the best formats, have been developed. However, the 2D spatial sparsity structure has not been well exploited in the existing work for SpMV on GPUs. In this paper, we propose an efficient tiled algorithm called TileSpMV for optimizing SpMV on GPUs through exploiting 2D spatial structure of sparse matrices. We first implement seven warp-level SpMV methods for calculating sparse tiles stored in a variety of formats, and then design a selection method to find the best format and SpMV implementation for each tile. We also adaptively extract nonzeros in the very sparse tiles into a separate matrix to maximize the overall performance. The experimental results show that our method is faster than state-of-the-art SpMV methods such as Merge-SpMV, CSR5 and BSR in most matrices of the full SuiteSparse Matrix Collection and delivers up to 2.61x, 3.96x and 426.59x speedups, respectively. Yuyao Niu, Zhengyang Lu 0003, Meichen Dong, Zhou Jin 0001, Weifeng Liu 0002, Guangming Tan |
IPDPS | 2 |
| 2020 | Efficient Block Algorithms for Parallel Sparse Triangular SolveabstractThe sparse triangular solve (SpTRSV) kernel is an important building block for a number of linear algebra routines such as sparse direct and iterative solvers. The major challenge of accelerating SpTRSV lies in the difficulties of finding higher parallelism. Existing work mainly focuses on reducing dependencies and synchronizations in the level-set methods. However, the 2D block layout of the input matrix has been largely ignored in designing more efficient SpTRSV algorithms. Zhengyang Lu 0003, Yuyao Niu, Weifeng Liu 0002 |
ICPP | 1 |