Shaoshuai Zhang

dblp:215/7561 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0002-9525-1659ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 4 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Deterministic and Efficient Low Rank Approximation on GPUs
Lu Shi 0006, Weijie Shen, Ruiyi Zhan, Dajun Huang, Weiwei Xu 0006, Shaoshuai Zhang
HPDC6
2026 Non-Delayed Cholesky Factorization
abstract
While General Matrix-Matrix Multiplication (GEMM) often nearly approaches the peak performance of modern GPUs, other fundamental dense linear algebra algorithms, such as Cholesky factorization, exhibit a significant performance gap. We identify the root cause of this inefficiency as the conventional ‘delayed update’ paradigm, a multi-kernel approach that relies on coarse-grained synchronization and incurs substantial data movement overhead, particularly for small-to-medium matrices.
Shaoshuai Zhang, Weifeng Liu 0002
ICS3
2026 Towards Singular Value Decomposition for Rank-Deficient Matrices: An Efficient and Accurate Algorithm on GPU Architectures
abstract
Singular Value Decomposition (SVD) is a fundamental tool in numerous scientific and engineering domains. Many high-performance libraries, such as LAPACK, MAGMA, and cuSOLVER, provide general, truncated, and randomized SVD routines. However, when the input is a low-rank matrix whose rank is not explicitly known, existing routines usually treat it as full-rank, which leads to suboptimal performance. In this paper, we propose an efficient SVD algorithm specifically for rank-deficient matrices based on a recently proposed rank-revealing QR factorization, termed QB factorization. To further enhance numerical stability and efficiency, we introduce a Householder QB factorization and a mixed-precision SVD algorithm, accompanied by a rigorous error analysis demonstrating correctness and stability. Experimental results show that our method achieves up to 6978.71x speedup over the general (full) SVD routine in cuSOLVER and is 9.99x faster than randomized SVD in FP32 precision. Moreover, our method exhibits higher numerical accuracy than cuSOLVER full SVD, achieving substantially smaller backward errors while maintaining stable and reliable singular values. Beyond synthetic benchmarks, we also demonstrate its effectiveness in an image compression application with higher efficiency.
Lu Shi 0006, Weiwei Xu 0006, Shaoshuai Zhang
PPoPP3
2026 OSLA: A High Performance One-Sided Linear Algebra Library on GPU Architectures
Lu Shi 0006, Ruiyi Zhan, Gaoyuan Zou, Geyong Min, Hancong Duan, Shaoshuai Zhang
IEEE Trans. Parallel Distributed Syst.7
2025 Improving Tridiagonalization Performance on GPU Architectures
abstract
Tridiagonalization, which is a key step in symmetric eigenvalue decomposition (EVD), aims to convert a symmetric matrix to a tridiagonal form. In Nvidia's cuSOLVER library, the FP64 precision tridiagonalization process only reach 2.1 TFLOPs out of 67 TFLOPs on H100 GPU, and it consumes a significant portion of the elapsed time in the entire EVD process, accounting for over 97%. Thus, improving the tridiagonalization performance is crucial on accelerating EVD. In this paper, we analyze the reasons behind the suboptimal performance of tridiagonalization on GPU architectures, and we propose a new double blocking band reduction algorithm along with an implementation of GPU-based bulge chasing to improve the tridiagonalization performance. Through experimental evaluation, the proposed FP64 precision tridiagonalization method yields up to 19.6 TFLOPs which is 9.3x and 5.2x faster compared cuSOVLER and MAGMA, respectively.
Zhekai Duan, Zitian Zhao, Saiqi Zheng, Qiao Li 0001, Xu Jiang 0004, Shaoshuai Zhang
PPoPP8
2025 On the Scalability and Efficiency of Intra-Process Communication in Ros 2
abstract
The Robot Operating System 2 (ROS 2) has become a widely adopted middleware framework for building modular and distributed robotic systems. Its intra-process communication mechanism is designed to reduce latency by avoiding serialization and memory copying, which is often treated as a negligible or constant-cost operation in both system design and performance analysis. However, this assumption oversimplifies the underlying behavior and may lead to inaccurate performance models and misleading conclusions, especially in latency-sensitive applications. In this paper, we present a comprehensive analysis of intraprocess communication in ROS 2, revealing that its performance is highly sensitive to message configuration, workload structure, and message usage strategies. We identify a scalability risk caused by misaligned communication configurations and propose a guideline to ensure efficient and predictable intra-process communication across varying execution patterns. In addition, we uncover a performance bottleneck in the default ROS 2 implementation stemming from repeated message creation. To address this, we propose a novel message pooling mechanism that reuses message objects to exploit temporal locality and eliminate redundant allocations. Our design is fully compatible with existing ROS 2 APIs and requires no modifications to application-level code. Experimental evaluations using synthetic benchmarks and real-world case studies demonstrate substantial improvements in communication latency, validating the practicality of our design.
Xiantong Luo, Xu Jiang 0004, Nan Guan, Yue Tang 0001, Shaoshuai Zhang
RTSS6
2025 Rethinking Back Transformation in 2-stage Eigenvalue Decomposition on Heterogeneous Architectures
abstract
The 2-stage eigenvalue decomposition (EVD) method outperforms conventional 1-stage method on GPUs and heterogeneous architectures, especially when eigenvectors are not required. However, its performance advantage diminishes when performing back transformation to obtain eigenvectors. To address this, we propose two key solutions: 1) replacing BLAS3 operations with BLAS2 operations during the bulge-chasing back transformation for better performance, and 2) reordering the back transformation workflow from a backward pattern to a new parallelism-driven pattern to hide divide-and-conquer latency, at the cost of one additional GEMM computation. Experimentally, the proposed back transformation algorithm demonstrates significant performance improvements, outperforming the SOTA implementation in MAGMA by an average factor of 3.58x. For complete FP64 precision symmetric EVD with eigenvectors, the proposed algorithm, incorporating both solutions, surpasses the SOTA implementations in MAGMA and cuSOLVER by average factors of 2.62x and 2.21x, respectively.
Dajun Huang, Gaoyuan Zou, Lu Shi 0006, Xu Jiang 0004, Xi Wu 0004, Hancong Duan, Shaoshuai Zhang
SC8
2025 High Performance Householder QR Factorization on Emerging GPU Architectures Using Tensor Cores
abstract
Since 2017, NVIDIA GPUs have been equipped with specialized units known as Tensor Cores, which demonstrate remarkable efficiency in processing matrix multiplications (GEMMs). Beyond GEMMs, researchers have explored the potential applications of Tensor Cores in matrix factorization, such as QR factorization. However, the inside GEMMs in QR factorization are typically tall and skinny. Compared to compute-bound square GEMMs, these tall and skinny GEMMs are memory bound, leading to suboptimal performance on Tensor Cores. To solve this problem, we indicate the recursive QR factorization can convert the tall and skinny GEMMs to relatively square and large GEMMs, resulting in better performance on Tensor Cores. Besides, we extend the FP16 Tensor-Cores-based QR factorization to accommodate FP32 and FP64 on FP16 and INT8 Tensor Cores, respectively. Additionally, to address the issue of orthogonality loss in the preceding Tensor Cores-based QR factorization, we transition from the Gram-Schmidt to the Householder algorithm while preserving high performance. According to our experimental evaluation conducted on NVIDIA's A100 and GeForce RTX 3090 GPU, the precision levels of FP64, FP32, and FP16 are up to 6.22x, 8.67x, and 4.03x faster, respectively, than the current state-of-the-art implementations.
Yuhan Leng, Gaoyuan Zou, Panruo Wu, Shaoshuai Zhang
IEEE Trans. Parallel Distributed Syst.5
2023 Fast Symmetric Eigenvalue Decomposition via WY Representation on Tensor Core
abstract
Symmetric eigenvalue decomposition (EVD) is a fundamental analytic and numerical tool used in many scientific areas. The state-of-the-art algorithm in terms of performance is typically the two-stage tridiagonalization method. The first stage in the two-stage tridiagonalization is called successive band reduction (SBR), which reduces a symmetric matrix to a band form, and its computational cost usually dominates. When Tensor Core (specialized matrix computational accelerator) is used to accelerate the expensive EVD, the conventional ZY-representation-based method results in suboptimal performance due to unfavorable shapes of the matrix computations. In this paper, we propose a new method that uses WY representation instead of ZY representation (see Section 3.2 for details), which can provide a better combination of locality and parallelism so as to perform better on Tensor Cores. Experimentally, the proposed method can bring up to 3.7x speedup in SBR and 2.3x in the entire EVD compared to state-of-the-art implementations.
Shaoshuai Zhang, Ruchi Shah, Hiroyuki Ootomo, Rio Yokota, Panruo Wu
PPoPP1
2021 Recursion Brings Speedup to Out-of-Core TensorCore-based Linear Algebra Algorithms: A Case Study of Classic Gram-Schmidt QR Factorization
abstract
Out-of-core processing aims to handle large amount of data when the memory is limited. There exists several out-of-core applications including disk-memory and CPU-GPU processing. Ideally, these out-of-core applications can be expected to be close to the peak performance of the in-core computations, if the data movement between different memory hierarchies can be overlapped by the in-core computations effectively. However, with the emergence of matrix accelerators such as TensorCore GPU, the imbalance between the speed of computations and data movement is further exacerbated, such that even high computation intensity kernels can be dominated by data movement cost. In such cases, the algorithms need to be redesigned to reduce communication volume and overlap the data movement by pipelines. In this paper, we select classic Gram-Schmidt QR factorization as an example to illustrate our recursive strategy, which shows smaller amount of data movement and higher overlapping ratio than the conventional blocking QR factorization algorithm. The results suggest this technique can potentially be applied to broader matrix computations kernels.
Shaoshuai Zhang, Panruo Wu
ICPP1
2020 High Accuracy Matrix Computations on Neural Engines: A Study of QR Factorization and its Applications
abstract
Fueled by the surge of ever expanding successful applications of deep neural networks and the great computational power demanded, modern computer processors and accelerators are beginning to offer half precision floating point arithmetic support, and special units (neural engines) such as NVIDIA TensorCore on GPU and Google Tensor Processing Unit (TPU) to accelerate the training and prediction of deep neural networks. It remains unclear how neural engines can be profitably used in applications other than neural networks. In this paper we present an endeavor of accelerating and stabilizing a fundamental matrix factorization on neural engines---the QR factorization---which may open doors to much wider relevance to scientific, engineering, and data science. We show that traditional Householder QR algorithms and implementations do not have the necessary data locality, parallelism, accuracy, and robustness on neural engines which are characterized by extreme speed and low precision/range.
Shaoshuai Zhang, Elaheh Baharlouei, Panruo Wu
HPDC1
2020 TensorSVM: accelerating kernel machines with tensor engine
abstract
This paper explores the use of Tensor Engines to accelerate nonlinear and linear SVM training. Support Vector Machine(SVM) is a classical machine learning model for classification and regression and remains to be the state-of-the-art model for some tasks such as text classification and bioinformatics. However large scale SVM training is still challenging because of its high computational complexity. This is especially severe for non-linear SVM with kernel tricks. On the other hand, the surging importance of neural networks fuels the emergence of specialized processors called Tensor Units (TensorCore in GPU and Tensor Processing Unit of Google) which are characterized by extreme efficiency and very limited precision and range. This paper proposes a TensorCore GPU based SVM algorithm and software system that is faster and more scalable than state-of-the-art SVM solvers. It includes a fast, accurate low-rank Gram matrix approximation that effectively utilizes the TensorCore in GPU and a primal-dual interior-point method to solve the quadratic program with a fast and predictable convergence rate. The random projection based Gram matrix approximation can be substantially accelerated by TensorCore on GPU.
Shaoshuai Zhang, Ruchi Shah, Panruo Wu
ICS1
2019 xSVM: Scalable Distributed Kernel Support Vector Machine Training
abstract
Kernel Support Vector Machine (SVM) is a popular machine learning model for classification and regression. A significant challenge of large scale Kernel SVM is the size of the Gram matrix (n × n), which cannot be stored or processed efficiently when training data-set is large (e.g. n in the millions). This paper proposes a novel SVM training algorithm and its parallelization strategy that can efficiently train on data-sets with millions of samples on thousands of processors. It consists of an accurate, fast, and scalable low rank matrix approximation based on random projection, and a primal-dual interior point method to solve the approximated optimization problem. We demonstrate that xSVM is fast, scalable, and accurate on large scale data-sets and computing nodes. Compared to state-of-the-art distributed Kernel L1-SVM system xSVM is consistently several times faster, with comparable accuracy to the exact model trained by LIBSVM.
Ruchi Shah, Shaoshuai Zhang, Panruo Wu
IEEE BigData2