Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Yuhan Leng

dblp:396/8063 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
High-performance computing · 54% GPUs and heterogeneous computing · 46%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing
GPU computing
0.912025
High Performance Householder QR Factorization on Emerging GPU Architectures Using Tensor Cores · IEEE Trans. Parallel Distributed Syst. 2025
High-performance computing
numerical linear algebra
0.912025
High Performance Householder QR Factorization on Emerging GPU Architectures Using Tensor Cores · IEEE Trans. Parallel Distributed Syst. 2025
High-performance computing › numerical linear algebra › matrix factorization
QR factorization
0.912025
High Performance Householder QR Factorization on Emerging GPU Architectures Using Tensor Cores · IEEE Trans. Parallel Distributed Syst. 2025
GPUs and heterogeneous computing › GPU computing
tensor cores
0.912025
High Performance Householder QR Factorization on Emerging GPU Architectures Using Tensor Cores · IEEE Trans. Parallel Distributed Syst. 2025
High-performance computing › numerical linear algebra
matrix multiplication
0.312025
High Performance Householder QR Factorization on Emerging GPU Architectures Using Tensor Cores · IEEE Trans. Parallel Distributed Syst. 2025

Methods — techniques the papers use, named apart from their topics

recursive QR · 0.9householder algorithm · 0.9gram-schmidt · 0.9
YearPublicationVenuePosition
2025 High Performance Householder QR Factorization on Emerging GPU Architectures Using Tensor Cores
abstract
Since 2017, NVIDIA GPUs have been equipped with specialized units known as Tensor Cores, which demonstrate remarkable efficiency in processing matrix multiplications (GEMMs). Beyond GEMMs, researchers have explored the potential applications of Tensor Cores in matrix factorization, such as QR factorization. However, the inside GEMMs in QR factorization are typically tall and skinny. Compared to compute-bound square GEMMs, these tall and skinny GEMMs are memory bound, leading to suboptimal performance on Tensor Cores. To solve this problem, we indicate the recursive QR factorization can convert the tall and skinny GEMMs to relatively square and large GEMMs, resulting in better performance on Tensor Cores. Besides, we extend the FP16 Tensor-Cores-based QR factorization to accommodate FP32 and FP64 on FP16 and INT8 Tensor Cores, respectively. Additionally, to address the issue of orthogonality loss in the preceding Tensor Cores-based QR factorization, we transition from the Gram-Schmidt to the Householder algorithm while preserving high performance. According to our experimental evaluation conducted on NVIDIA's A100 and GeForce RTX 3090 GPU, the precision levels of FP64, FP32, and FP16 are up to 6.22x, 8.67x, and 4.03x faster, respectively, than the current state-of-the-art implementations.
Yuhan Leng, Gaoyuan Zou, Panruo Wu, Shaoshuai Zhang
IEEE Trans. Parallel Distributed Syst.1