Zhengyang Lu 0003

dblp:254/3011-3 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
4since 2021 · last 2023
0000-0002-1540-0678ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2023 HASpMV: Heterogeneity-Aware Sparse Matrix-Vector Multiplication on Modern Asymmetric Multicore Processors
abstract
Sparse matrix-vector multiplication (SpMV) is a fundamental routine in computational science and engineering. Its optimization methods on various homogeneous parallel processors, such as CPUs and GPUs, received much attention. Recently, asymmetric multicore processors (AMPs) have heterogeneous performance and efficient cores (e.g., P- and E-cores from Intel and Apple, or Big.LITTLE cores from ARM), or cores with different cache structures (e.g., cores with/without 3D V-Cache from AMD) are becoming one of the mainstream in desktop and workstation computers. However, there lacks heterogeneity-aware research on accelerating SpMV on AMPs.We in this paper propose a parallel algorithm called heterogeneity-aware SpMV (HASpMV) for improving the performance of SpMV on the latest 12th- and 13th-Gen AMPs from Intel and Ryzen 9 AMPs from AMD. We first micro-benchmark bandwidth and multi-/single-core SpMV to collect performance characteristics and to motivate our algorithm design, and then develop several optimization techniques to assign workloads between the two types of cores for achieving significantly better cache locality and load balancing. The experimental results show that compared to the latest version of the Intel oneMKL library and the open-source works CSR5 and merge-SpMV, HASpMV achieves an average speedup of 2.61x, 2.31x, and 3.73x (up to 5.23x, 4.46x, and 8.23x) on the i9-12900KF processor. On the i9-13900KF processor, HASpMV achieves an average speedup of 3.17x, 1.52x, and 2.23x (up to 9.46x, 5.31x, and 4.49x). Additionally, when comparing AMD Ryzen 9 7950X3D and 7950X AMPs, HASpMV brings an average speedup of 1.43x, 1.3x, and 1.29x (up to 6.28x, 7.8x, and 10.8x) over AMD Optimizing CPU Libraries (AOCL), CSR5, and merge-SpMV, respectively.
Helin Cheng, Zhengyang Lu 0003, Yuechen Lu, Weifeng Liu 0002
CLUSTER3
2023 TileSpTRSV: a tiled algorithm for parallel sparse triangular solve on GPUs
Zhengyang Lu 0003, Weifeng Liu 0002
CCF Trans. High Perform. Comput.1
2022 TileSpGEMM: a tiled algorithm for parallel sparse general matrix-matrix multiplication on GPUs
abstract
Sparse general matrix-matrix multiplication (SpGEMM) is one of the most fundamental building blocks in sparse linear solvers, graph processing frameworks and machine learning applications. The existing parallel approaches for shared memory SpGEMM mostly use the row-row style with possibly good parallelism. However, because of the irregularity in sparsity structures, the existing row-row methods often suffer from three problems: (1) load imbalance, (2) high global space complexity and unsatisfactory data locality, and (3) sparse accumulator selection.
Yuyao Niu, Zhengyang Lu 0003, Haonan Ji, Shuhui Song, Zhou Jin 0001, Weifeng Liu 0002
PPoPP2
2021 TileSpMV: A Tiled Algorithm for Sparse Matrix-Vector Multiplication on GPUs
abstract
With the extensive use of GPUs in modern supercomputers, accelerating sparse matrix-vector multiplication (SpMV) on GPUs received much attention in the last couple of decades. A number of techniques, such as increasing utilization of wide vector units, reducing load imbalance and selecting the best formats, have been developed. However, the 2D spatial sparsity structure has not been well exploited in the existing work for SpMV on GPUs. In this paper, we propose an efficient tiled algorithm called TileSpMV for optimizing SpMV on GPUs through exploiting 2D spatial structure of sparse matrices. We first implement seven warp-level SpMV methods for calculating sparse tiles stored in a variety of formats, and then design a selection method to find the best format and SpMV implementation for each tile. We also adaptively extract nonzeros in the very sparse tiles into a separate matrix to maximize the overall performance. The experimental results show that our method is faster than state-of-the-art SpMV methods such as Merge-SpMV, CSR5 and BSR in most matrices of the full SuiteSparse Matrix Collection and delivers up to 2.61x, 3.96x and 426.59x speedups, respectively.
Yuyao Niu, Zhengyang Lu 0003, Meichen Dong, Zhou Jin 0001, Weifeng Liu 0002, Guangming Tan
IPDPS2
2020 Efficient Block Algorithms for Parallel Sparse Triangular Solve
abstract
The sparse triangular solve (SpTRSV) kernel is an important building block for a number of linear algebra routines such as sparse direct and iterative solvers. The major challenge of accelerating SpTRSV lies in the difficulties of finding higher parallelism. Existing work mainly focuses on reducing dependencies and synchronizations in the level-set methods. However, the 2D block layout of the input matrix has been largely ignored in designing more efficient SpTRSV algorithms.
Zhengyang Lu 0003, Yuyao Niu, Weifeng Liu 0002
ICPP1