Ruifeng Zhang 0008

dblp:07/2301-8 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2026
0009-0002-3765-1555ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 SparseX: Synergizing GPU Libraries for Sparse Matrix Multiplication on Heterogeneous Processors
abstract
Sparse Matrix-Matrix Multiplication (SpMM) on GPU is critical to applications ranging from scientific simulations to Graph Neural Networks (GNNs) and Deep Neural Networks (DNNs). Modern GPUs offer diverse processing units, such as CUDA cores, Tensor Cores, and Sparse Tensor Cores. Many SpMM libraries have been built to harness those different types of processors. Although impressive performance has been reported by each, including cuSparse, Sputnik, CLASP, and Jigsaw, a systematic study in this work shows that no single library is a clear winner across all matrices and scenarios. Based on the empirical observations, this work proposes the first solution to synergize the various libraries to best harness the heterogeneous processors and the array of cutting-edge libraries. The solution is an extensible framework, namely SparseX, that can automatically select the best matrix multiplication library (and types of processors) on the fly for a given sparse matrix on a GPU through an agile accurate predictive model. Experiments show that SparseX can speed up sparse matrix multiplications on thousands of real-world matrices significantly over the SOTA GPU libraries, achieving significant speedups (e.g., as much as 95.34x over cuSparse). Its extensible design makes it easy to be extended to cover new libraries and hardware architectures.
Ruifeng Zhang 0008, Ang Li 0006, Xipeng Shen
CGO1
2026 N:M sparsity-oriented graph reordering for accelerating GNNs on GPU sparse tensor cores
abstract
Recent GPUs incorporate Sparse Tensor Cores (SPTC) to accelerate computations on matrices that satisfy structured N : M sparsity; software stacks further extend support to generalized V : N : M patterns. Although graphs in graph neural networks (GNNs) are typically sparse, their sparsity is irregular and rarely conforms to these patterns. This paper introduces a lossless graph reordering algorithm that reshapes irregular graph data into the required sparse patterns, allowing GNN workloads to exploit SPTC. The transformation preserves model accuracy and maintains the symmetry of graph adjacency matrices, ensuring compatibility with symmetry-based graph algorithms. On the SuiteSparse collection, our method removes 93.77%–98.51% of vector-level N : M violations and increases the fraction of conforming graphs from 5%–7% to 79.3%–91.3%. On NVIDIA A100 GPUs, it accelerates SpMM by up to 50.7 × (geometric-mean speedups 3.44 × –7.25 × ) over cuSPARSE, and speeds up key GNN graph operations on real graphs by as much as 6.7 × (2.6 × on average).
Ruifeng Zhang 0008, Jou-An Chen, Hsin-Hsuan Sung, Ang Li 0006, Xipeng Shen
J. Parallel Distributed Comput.1
2025 Accelerating GNNs on GPU Sparse Tensor Cores through N: M Sparsity-Oriented Graph Reordering
abstract
Recent GPUs have introduced Sparse Tensor Cores (SPTC) to accelerate computations on sparse matrices meeting the N:M sparse patterns. Software tools expand the support to more general V:N:M patterns. Graphs in Graph Neural Networks (GNNs) are typically sparse, but the sparsity is often irregular, not conforming to the required V:N:M sparse patterns. This paper proposes a novel graph reordering algorithm to transform irregular graph data into the required sparse patterns for GNNs to benefit from SPTC. The optimization is lossless, maintaining the accuracy of GNN. It at the same time keeps the symmetry of the adjacency matrices of the graphs so that the same matrices can remain compatible with many symmetry-based graph algorithms. The optimization successfully removes 98-100% violations of the N:M sparse patterns at the vector level and increases the portion of conforming graphs in the SuiteSparse collection from 5-9% to 88.7-93.5%. On A100 GPUs, the optimization accelerates Sparse Matrix Matrix (SpMM) by up to 43X (a geomean speedup of 2.3X - 7.5X) over cuSPARSE and speeds up the key graph operations in GNNs on real graphs by as much as 8.6X (3.5X on average).
Jou-An Chen, Hsin-Hsuan Sung, Ruifeng Zhang 0008, Ang Li 0006, Xipeng Shen
PPoPP3