Saiqi Zheng

dblp:398/7065 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
0009-0000-4569-1552ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
High-performance computing · 50% GPUs and heterogeneous computing · 50%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing
GPU computing
0.912025
Improving Tridiagonalization Performance on GPU Architectures · PPoPP 2025
GPUs and heterogeneous computing
GPU kernel optimization
0.912025
Improving Tridiagonalization Performance on GPU Architectures · PPoPP 2025
High-performance computing
numerical linear algebra
0.912025
Improving Tridiagonalization Performance on GPU Architectures · PPoPP 2025
High-performance computing › numerical linear algebra
tridiagonalization
0.912025
Improving Tridiagonalization Performance on GPU Architectures · PPoPP 2025

Methods — techniques the papers use, named apart from their topics

double blocking band reduction · 0.9bulge chasing · 0.9
YearPublicationVenuePosition
2025 Improving Tridiagonalization Performance on GPU Architectures
abstract
Tridiagonalization, which is a key step in symmetric eigenvalue decomposition (EVD), aims to convert a symmetric matrix to a tridiagonal form. In Nvidia's cuSOLVER library, the FP64 precision tridiagonalization process only reach 2.1 TFLOPs out of 67 TFLOPs on H100 GPU, and it consumes a significant portion of the elapsed time in the entire EVD process, accounting for over 97%. Thus, improving the tridiagonalization performance is crucial on accelerating EVD. In this paper, we analyze the reasons behind the suboptimal performance of tridiagonalization on GPU architectures, and we propose a new double blocking band reduction algorithm along with an implementation of GPU-based bulge chasing to improve the tridiagonalization performance. Through experimental evaluation, the proposed FP64 precision tridiagonalization method yields up to 19.6 TFLOPs which is 9.3x and 5.2x faster compared cuSOVLER and MAGMA, respectively.
Zhekai Duan, Zitian Zhao, Saiqi Zheng, Qiao Li 0001, Xu Jiang 0004, Shaoshuai Zhang
PPoPP5