EDBT 2026 Demo / reviewers in the wild / expert
Shaoshuai Zhang
dblp:215/7561
· DBLP profile ↗
13ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0002-9525-1659ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 4 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Deterministic and Efficient Low Rank Approximation on GPUs
Lu Shi 0006, Weijie Shen, Ruiyi Zhan, Dajun Huang, Weiwei Xu 0006, Shaoshuai Zhang |
HPDC | 6 |
| 2026 | Non-Delayed Cholesky FactorizationabstractWhile General Matrix-Matrix Multiplication (GEMM) often nearly approaches the peak performance of modern GPUs, other fundamental dense linear algebra algorithms, such as Cholesky factorization, exhibit a significant performance gap. We identify the root cause of this inefficiency as the conventional ‘delayed update’ paradigm, a multi-kernel approach that relies on coarse-grained synchronization and incurs substantial data movement overhead, particularly for small-to-medium matrices. Shaoshuai Zhang, Weifeng Liu 0002 |
ICS | 3 |
| 2026 | Towards Singular Value Decomposition for Rank-Deficient Matrices: An Efficient and Accurate Algorithm on GPU ArchitecturesabstractSingular Value Decomposition (SVD) is a fundamental tool in numerous scientific and engineering domains. Many high-performance libraries, such as LAPACK, MAGMA, and cuSOLVER, provide general, truncated, and randomized SVD routines. However, when the input is a low-rank matrix whose rank is not explicitly known, existing routines usually treat it as full-rank, which leads to suboptimal performance. In this paper, we propose an efficient SVD algorithm specifically for rank-deficient matrices based on a recently proposed rank-revealing QR factorization, termed QB factorization. To further enhance numerical stability and efficiency, we introduce a Householder QB factorization and a mixed-precision SVD algorithm, accompanied by a rigorous error analysis demonstrating correctness and stability. Experimental results show that our method achieves up to 6978.71x speedup over the general (full) SVD routine in cuSOLVER and is 9.99x faster than randomized SVD in FP32 precision. Moreover, our method exhibits higher numerical accuracy than cuSOLVER full SVD, achieving substantially smaller backward errors while maintaining stable and reliable singular values. Beyond synthetic benchmarks, we also demonstrate its effectiveness in an image compression application with higher efficiency. Lu Shi 0006, Weiwei Xu 0006, Shaoshuai Zhang |
PPoPP | 3 |
| 2026 | OSLA: A High Performance One-Sided Linear Algebra Library on GPU Architectures
Lu Shi 0006, Ruiyi Zhan, Gaoyuan Zou, Geyong Min, Hancong Duan, Shaoshuai Zhang |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2025 | Improving Tridiagonalization Performance on GPU ArchitecturesabstractTridiagonalization, which is a key step in symmetric eigenvalue decomposition (EVD), aims to convert a symmetric matrix to a tridiagonal form. In Nvidia's cuSOLVER library, the FP64 precision tridiagonalization process only reach 2.1 TFLOPs out of 67 TFLOPs on H100 GPU, and it consumes a significant portion of the elapsed time in the entire EVD process, accounting for over 97%. Thus, improving the tridiagonalization performance is crucial on accelerating EVD. In this paper, we analyze the reasons behind the suboptimal performance of tridiagonalization on GPU architectures, and we propose a new double blocking band reduction algorithm along with an implementation of GPU-based bulge chasing to improve the tridiagonalization performance. Through experimental evaluation, the proposed FP64 precision tridiagonalization method yields up to 19.6 TFLOPs which is 9.3x and 5.2x faster compared cuSOVLER and MAGMA, respectively. Zhekai Duan, Zitian Zhao, Saiqi Zheng, Qiao Li 0001, Xu Jiang 0004, Shaoshuai Zhang |
PPoPP | 8 |
| 2025 | On the Scalability and Efficiency of Intra-Process Communication in Ros 2abstractThe Robot Operating System 2 (ROS 2) has become a widely adopted middleware framework for building modular and distributed robotic systems. Its intra-process communication mechanism is designed to reduce latency by avoiding serialization and memory copying, which is often treated as a negligible or constant-cost operation in both system design and performance analysis. However, this assumption oversimplifies the underlying behavior and may lead to inaccurate performance models and misleading conclusions, especially in latency-sensitive applications. In this paper, we present a comprehensive analysis of intraprocess communication in ROS 2, revealing that its performance is highly sensitive to message configuration, workload structure, and message usage strategies. We identify a scalability risk caused by misaligned communication configurations and propose a guideline to ensure efficient and predictable intra-process communication across varying execution patterns. In addition, we uncover a performance bottleneck in the default ROS 2 implementation stemming from repeated message creation. To address this, we propose a novel message pooling mechanism that reuses message objects to exploit temporal locality and eliminate redundant allocations. Our design is fully compatible with existing ROS 2 APIs and requires no modifications to application-level code. Experimental evaluations using synthetic benchmarks and real-world case studies demonstrate substantial improvements in communication latency, validating the practicality of our design. Xiantong Luo, Xu Jiang 0004, Nan Guan, Yue Tang 0001, Shaoshuai Zhang |
RTSS | 6 |
| 2025 | Rethinking Back Transformation in 2-stage Eigenvalue Decomposition on Heterogeneous ArchitecturesabstractThe 2-stage eigenvalue decomposition (EVD) method outperforms conventional 1-stage method on GPUs and heterogeneous architectures, especially when eigenvectors are not required. However, its performance advantage diminishes when performing back transformation to obtain eigenvectors. To address this, we propose two key solutions: 1) replacing BLAS3 operations with BLAS2 operations during the bulge-chasing back transformation for better performance, and 2) reordering the back transformation workflow from a backward pattern to a new parallelism-driven pattern to hide divide-and-conquer latency, at the cost of one additional GEMM computation. Experimentally, the proposed back transformation algorithm demonstrates significant performance improvements, outperforming the SOTA implementation in MAGMA by an average factor of 3.58x. For complete FP64 precision symmetric EVD with eigenvectors, the proposed algorithm, incorporating both solutions, surpasses the SOTA implementations in MAGMA and cuSOLVER by average factors of 2.62x and 2.21x, respectively. Dajun Huang, Gaoyuan Zou, Lu Shi 0006, Xu Jiang 0004, Xi Wu 0004, Hancong Duan, Shaoshuai Zhang |
SC | 8 |
| 2025 | High Performance Householder QR Factorization on Emerging GPU Architectures Using Tensor CoresabstractSince 2017, NVIDIA GPUs have been equipped with specialized units known as Tensor Cores, which demonstrate remarkable efficiency in processing matrix multiplications (GEMMs). Beyond GEMMs, researchers have explored the potential applications of Tensor Cores in matrix factorization, such as QR factorization. However, the inside GEMMs in QR factorization are typically tall and skinny. Compared to compute-bound square GEMMs, these tall and skinny GEMMs are memory bound, leading to suboptimal performance on Tensor Cores. To solve this problem, we indicate the recursive QR factorization can convert the tall and skinny GEMMs to relatively square and large GEMMs, resulting in better performance on Tensor Cores. Besides, we extend the FP16 Tensor-Cores-based QR factorization to accommodate FP32 and FP64 on FP16 and INT8 Tensor Cores, respectively. Additionally, to address the issue of orthogonality loss in the preceding Tensor Cores-based QR factorization, we transition from the Gram-Schmidt to the Householder algorithm while preserving high performance. According to our experimental evaluation conducted on NVIDIA's A100 and GeForce RTX 3090 GPU, the precision levels of FP64, FP32, and FP16 are up to 6.22x, 8.67x, and 4.03x faster, respectively, than the current state-of-the-art implementations. Yuhan Leng, Gaoyuan Zou, Panruo Wu, Shaoshuai Zhang |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2023 | Fast Symmetric Eigenvalue Decomposition via WY Representation on Tensor CoreabstractSymmetric eigenvalue decomposition (EVD) is a fundamental analytic and numerical tool used in many scientific areas. The state-of-the-art algorithm in terms of performance is typically the two-stage tridiagonalization method. The first stage in the two-stage tridiagonalization is called successive band reduction (SBR), which reduces a symmetric matrix to a band form, and its computational cost usually dominates. When Tensor Core (specialized matrix computational accelerator) is used to accelerate the expensive EVD, the conventional ZY-representation-based method results in suboptimal performance due to unfavorable shapes of the matrix computations. In this paper, we propose a new method that uses WY representation instead of ZY representation (see Section 3.2 for details), which can provide a better combination of locality and parallelism so as to perform better on Tensor Cores. Experimentally, the proposed method can bring up to 3.7x speedup in SBR and 2.3x in the entire EVD compared to state-of-the-art implementations. Shaoshuai Zhang, Ruchi Shah, Hiroyuki Ootomo, Rio Yokota, Panruo Wu |
PPoPP | 1 |
| 2021 | Recursion Brings Speedup to Out-of-Core TensorCore-based Linear Algebra Algorithms: A Case Study of Classic Gram-Schmidt QR FactorizationabstractOut-of-core processing aims to handle large amount of data when the memory is limited. There exists several out-of-core applications including disk-memory and CPU-GPU processing. Ideally, these out-of-core applications can be expected to be close to the peak performance of the in-core computations, if the data movement between different memory hierarchies can be overlapped by the in-core computations effectively. However, with the emergence of matrix accelerators such as TensorCore GPU, the imbalance between the speed of computations and data movement is further exacerbated, such that even high computation intensity kernels can be dominated by data movement cost. In such cases, the algorithms need to be redesigned to reduce communication volume and overlap the data movement by pipelines. In this paper, we select classic Gram-Schmidt QR factorization as an example to illustrate our recursive strategy, which shows smaller amount of data movement and higher overlapping ratio than the conventional blocking QR factorization algorithm. The results suggest this technique can potentially be applied to broader matrix computations kernels. Shaoshuai Zhang, Panruo Wu |
ICPP | 1 |
| 2020 | High Accuracy Matrix Computations on Neural Engines: A Study of QR Factorization and its ApplicationsabstractFueled by the surge of ever expanding successful applications of deep neural networks and the great computational power demanded, modern computer processors and accelerators are beginning to offer half precision floating point arithmetic support, and special units (neural engines) such as NVIDIA TensorCore on GPU and Google Tensor Processing Unit (TPU) to accelerate the training and prediction of deep neural networks. It remains unclear how neural engines can be profitably used in applications other than neural networks. In this paper we present an endeavor of accelerating and stabilizing a fundamental matrix factorization on neural engines---the QR factorization---which may open doors to much wider relevance to scientific, engineering, and data science. We show that traditional Householder QR algorithms and implementations do not have the necessary data locality, parallelism, accuracy, and robustness on neural engines which are characterized by extreme speed and low precision/range. Shaoshuai Zhang, Elaheh Baharlouei, Panruo Wu |
HPDC | 1 |
| 2020 | TensorSVM: accelerating kernel machines with tensor engineabstractThis paper explores the use of Tensor Engines to accelerate nonlinear and linear SVM training. Support Vector Machine(SVM) is a classical machine learning model for classification and regression and remains to be the state-of-the-art model for some tasks such as text classification and bioinformatics. However large scale SVM training is still challenging because of its high computational complexity. This is especially severe for non-linear SVM with kernel tricks. On the other hand, the surging importance of neural networks fuels the emergence of specialized processors called Tensor Units (TensorCore in GPU and Tensor Processing Unit of Google) which are characterized by extreme efficiency and very limited precision and range. This paper proposes a TensorCore GPU based SVM algorithm and software system that is faster and more scalable than state-of-the-art SVM solvers. It includes a fast, accurate low-rank Gram matrix approximation that effectively utilizes the TensorCore in GPU and a primal-dual interior-point method to solve the quadratic program with a fast and predictable convergence rate. The random projection based Gram matrix approximation can be substantially accelerated by TensorCore on GPU. Shaoshuai Zhang, Ruchi Shah, Panruo Wu |
ICS | 1 |
| 2019 | xSVM: Scalable Distributed Kernel Support Vector Machine TrainingabstractKernel Support Vector Machine (SVM) is a popular machine learning model for classification and regression. A significant challenge of large scale Kernel SVM is the size of the Gram matrix (n × n), which cannot be stored or processed efficiently when training data-set is large (e.g. n in the millions). This paper proposes a novel SVM training algorithm and its parallelization strategy that can efficiently train on data-sets with millions of samples on thousands of processors. It consists of an accurate, fast, and scalable low rank matrix approximation based on random projection, and a primal-dual interior point method to solve the approximated optimization problem. We demonstrate that xSVM is fast, scalable, and accurate on large scale data-sets and computing nodes. Compared to state-of-the-art distributed Kernel L1-SVM system xSVM is consistently several times faster, with comparable accuracy to the exact model trained by LIBSVM. Ruchi Shah, Shaoshuai Zhang, Panruo Wu |
IEEE BigData | 2 |