EDBT 2026 Demo / reviewers in the wild / expert
Gaoyuan Zou
dblp:396/8208
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2026
0009-0001-0550-2662ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
High-performance computing · 51% GPUs and heterogeneous computing · 39% Parallel and multicore computing · 10% |
Topics — the 12 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
GPUs and heterogeneous computing
GPU computing |
1.9 | 2 | 2026 | OSLA: A High Performance One-Sided Linear Algebra Library on GPU Architectures · IEEE Trans. Parallel Distributed Syst. 2026 High Performance Householder QR Factorization on Emerging GPU Architectures Using Tensor Cores · IEEE Trans. Parallel Distributed Syst. 2025 |
High-performance computing
linear algebra library |
1.0 | 1 | 2026 | OSLA: A High Performance One-Sided Linear Algebra Library on GPU Architectures · IEEE Trans. Parallel Distributed Syst. 2026 |
Parallel and multicore computing › parallel computing › parallel communication
one-sided communication |
1.0 | 1 | 2026 | OSLA: A High Performance One-Sided Linear Algebra Library on GPU Architectures · IEEE Trans. Parallel Distributed Syst. 2026 |
High-performance computing › numerical linear algebra
parallel linear algebra |
1.0 | 1 | 2026 | OSLA: A High Performance One-Sided Linear Algebra Library on GPU Architectures · IEEE Trans. Parallel Distributed Syst. 2026 |
High-performance computing › numerical linear algebra
dense linear algebra |
0.9 | 1 | 2025 | Rethinking Back Transformation in 2-stage Eigenvalue Decomposition on Heterogeneous Architectures · SC 2025 |
GPUs and heterogeneous computing › GPU-accelerated scientific computing
GPU-accelerated numerical linear algebra |
0.9 | 1 | 2025 | Rethinking Back Transformation in 2-stage Eigenvalue Decomposition on Heterogeneous Architectures · SC 2025 |
High-performance computing
numerical linear algebra |
0.9 | 1 | 2025 | High Performance Householder QR Factorization on Emerging GPU Architectures Using Tensor Cores · IEEE Trans. Parallel Distributed Syst. 2025 |
High-performance computing › numerical linear algebra › matrix factorization
QR factorization |
0.9 | 1 | 2025 | High Performance Householder QR Factorization on Emerging GPU Architectures Using Tensor Cores · IEEE Trans. Parallel Distributed Syst. 2025 |
GPUs and heterogeneous computing › GPU computing
tensor cores |
0.9 | 1 | 2025 | High Performance Householder QR Factorization on Emerging GPU Architectures Using Tensor Cores · IEEE Trans. Parallel Distributed Syst. 2025 |
GPUs and heterogeneous computing
heterogeneous architecture |
0.3 | 1 | 2025 | Rethinking Back Transformation in 2-stage Eigenvalue Decomposition on Heterogeneous Architectures · SC 2025 |
High-performance computing › numerical linear algebra
matrix multiplication |
0.3 | 1 | 2025 | High Performance Householder QR Factorization on Emerging GPU Architectures Using Tensor Cores · IEEE Trans. Parallel Distributed Syst. 2025 |
High-performance computing
performance optimization at scale |
0.3 | 1 | 2025 | Rethinking Back Transformation in 2-stage Eigenvalue Decomposition on Heterogeneous Architectures · SC 2025 |
Methods — techniques the papers use, named apart from their topics
one-sided communication · 1.0recursive QR · 0.9householder algorithm · 0.9gram-schmidt · 0.9divide-and-conquer · 0.9bulge chasing · 0.9BLAS3 · 0.9BLAS2 · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OSLA: A High Performance One-Sided Linear Algebra Library on GPU Architectures
Lu Shi 0006, Ruiyi Zhan, Gaoyuan Zou, Geyong Min, Hancong Duan, Shaoshuai Zhang |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2025 | Rethinking Back Transformation in 2-stage Eigenvalue Decomposition on Heterogeneous ArchitecturesabstractThe 2-stage eigenvalue decomposition (EVD) method outperforms conventional 1-stage method on GPUs and heterogeneous architectures, especially when eigenvectors are not required. However, its performance advantage diminishes when performing back transformation to obtain eigenvectors. To address this, we propose two key solutions: 1) replacing BLAS3 operations with BLAS2 operations during the bulge-chasing back transformation for better performance, and 2) reordering the back transformation workflow from a backward pattern to a new parallelism-driven pattern to hide divide-and-conquer latency, at the cost of one additional GEMM computation. Experimentally, the proposed back transformation algorithm demonstrates significant performance improvements, outperforming the SOTA implementation in MAGMA by an average factor of 3.58x. For complete FP64 precision symmetric EVD with eigenvectors, the proposed algorithm, incorporating both solutions, surpasses the SOTA implementations in MAGMA and cuSOLVER by average factors of 2.62x and 2.21x, respectively. Dajun Huang, Gaoyuan Zou, Lu Shi 0006, Xu Jiang 0004, Xi Wu 0004, Hancong Duan, Shaoshuai Zhang |
SC | 3 |
| 2025 | High Performance Householder QR Factorization on Emerging GPU Architectures Using Tensor CoresabstractSince 2017, NVIDIA GPUs have been equipped with specialized units known as Tensor Cores, which demonstrate remarkable efficiency in processing matrix multiplications (GEMMs). Beyond GEMMs, researchers have explored the potential applications of Tensor Cores in matrix factorization, such as QR factorization. However, the inside GEMMs in QR factorization are typically tall and skinny. Compared to compute-bound square GEMMs, these tall and skinny GEMMs are memory bound, leading to suboptimal performance on Tensor Cores. To solve this problem, we indicate the recursive QR factorization can convert the tall and skinny GEMMs to relatively square and large GEMMs, resulting in better performance on Tensor Cores. Besides, we extend the FP16 Tensor-Cores-based QR factorization to accommodate FP32 and FP64 on FP16 and INT8 Tensor Cores, respectively. Additionally, to address the issue of orthogonality loss in the preceding Tensor Cores-based QR factorization, we transition from the Gram-Schmidt to the Householder algorithm while preserving high performance. According to our experimental evaluation conducted on NVIDIA's A100 and GeForce RTX 3090 GPU, the precision levels of FP64, FP32, and FP16 are up to 6.22x, 8.67x, and 4.03x faster, respectively, than the current state-of-the-art implementations. Yuhan Leng, Gaoyuan Zou, Panruo Wu, Shaoshuai Zhang |
IEEE Trans. Parallel Distributed Syst. | 2 |