Lu Shi 0006

dblp:42/11188-6 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2026
0000-0003-2481-0318ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Deterministic and Efficient Low Rank Approximation on GPUs
Lu Shi 0006, Weijie Shen, Ruiyi Zhan, Dajun Huang, Weiwei Xu 0006, Shaoshuai Zhang
HPDC1
2026 Towards Singular Value Decomposition for Rank-Deficient Matrices: An Efficient and Accurate Algorithm on GPU Architectures
abstract
Singular Value Decomposition (SVD) is a fundamental tool in numerous scientific and engineering domains. Many high-performance libraries, such as LAPACK, MAGMA, and cuSOLVER, provide general, truncated, and randomized SVD routines. However, when the input is a low-rank matrix whose rank is not explicitly known, existing routines usually treat it as full-rank, which leads to suboptimal performance. In this paper, we propose an efficient SVD algorithm specifically for rank-deficient matrices based on a recently proposed rank-revealing QR factorization, termed QB factorization. To further enhance numerical stability and efficiency, we introduce a Householder QB factorization and a mixed-precision SVD algorithm, accompanied by a rigorous error analysis demonstrating correctness and stability. Experimental results show that our method achieves up to 6978.71x speedup over the general (full) SVD routine in cuSOLVER and is 9.99x faster than randomized SVD in FP32 precision. Moreover, our method exhibits higher numerical accuracy than cuSOLVER full SVD, achieving substantially smaller backward errors while maintaining stable and reliable singular values. Beyond synthetic benchmarks, we also demonstrate its effectiveness in an image compression application with higher efficiency.
Lu Shi 0006, Weiwei Xu 0006, Shaoshuai Zhang
PPoPP1
2026 OSLA: A High Performance One-Sided Linear Algebra Library on GPU Architectures
Lu Shi 0006, Ruiyi Zhan, Gaoyuan Zou, Geyong Min, Hancong Duan, Shaoshuai Zhang
IEEE Trans. Parallel Distributed Syst.1
2025 Rethinking Back Transformation in 2-stage Eigenvalue Decomposition on Heterogeneous Architectures
abstract
The 2-stage eigenvalue decomposition (EVD) method outperforms conventional 1-stage method on GPUs and heterogeneous architectures, especially when eigenvectors are not required. However, its performance advantage diminishes when performing back transformation to obtain eigenvectors. To address this, we propose two key solutions: 1) replacing BLAS3 operations with BLAS2 operations during the bulge-chasing back transformation for better performance, and 2) reordering the back transformation workflow from a backward pattern to a new parallelism-driven pattern to hide divide-and-conquer latency, at the cost of one additional GEMM computation. Experimentally, the proposed back transformation algorithm demonstrates significant performance improvements, outperforming the SOTA implementation in MAGMA by an average factor of 3.58x. For complete FP64 precision symmetric EVD with eigenvectors, the proposed algorithm, incorporating both solutions, surpasses the SOTA implementations in MAGMA and cuSOLVER by average factors of 2.62x and 2.21x, respectively.
Dajun Huang, Gaoyuan Zou, Lu Shi 0006, Xu Jiang 0004, Xi Wu 0004, Hancong Duan, Shaoshuai Zhang
SC4