VLDB 2026 Research / reviewers in the wild / expert
Renqiu Ouyang
dblp:351/5290
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2025
0000-0002-8712-4561ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
GPUs and heterogeneous computing · 46% High-performance computing · 42% Memory systems · 12% |
Topics — the 6 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
High-performance computing › tensor computation
sparse tensor contraction |
1.4 | 2 | 2024 | BCB-SpTC: An Efficient Sparse High-Dimensional Tensor Contraction Employing Tensor Core Acceleration · IEEE Trans. Parallel Distributed Syst. 2024 A Novel Parallel Algorithm for Sparse Tensor Matrix Chain Multiplication via TCU-Acceleration · IEEE Trans. Parallel Distributed Syst. 2023 |
GPUs and heterogeneous computing › GPU computing
tensor cores |
1.4 | 2 | 2024 | BCB-SpTC: An Efficient Sparse High-Dimensional Tensor Contraction Employing Tensor Core Acceleration · IEEE Trans. Parallel Distributed Syst. 2024 A Novel Parallel Algorithm for Sparse Tensor Matrix Chain Multiplication via TCU-Acceleration · IEEE Trans. Parallel Distributed Syst. 2023 |
GPUs and heterogeneous computing
GPU computing |
0.8 | 1 | 2024 | BCB-SpTC: An Efficient Sparse High-Dimensional Tensor Contraction Employing Tensor Core Acceleration · IEEE Trans. Parallel Distributed Syst. 2024 |
Memory systems › memory bandwidth
memory bandwidth optimization |
0.8 | 1 | 2024 | BCB-SpTC: An Efficient Sparse High-Dimensional Tensor Contraction Employing Tensor Core Acceleration · IEEE Trans. Parallel Distributed Syst. 2024 |
High-performance computing › sparse linear algebra
sparse matrix storage format |
0.8 | 1 | 2024 | BCB-SpTC: An Efficient Sparse High-Dimensional Tensor Contraction Employing Tensor Core Acceleration · IEEE Trans. Parallel Distributed Syst. 2024 |
High-performance computing › tensor computation
tensor decomposition |
0.2 | 1 | 2023 | A Novel Parallel Algorithm for Sparse Tensor Matrix Chain Multiplication via TCU-Acceleration · IEEE Trans. Parallel Distributed Syst. 2023 |
Methods — techniques the papers use, named apart from their topics
vectorization · 0.8blocking · 0.8bitmap encoding · 0.8multi-dimensional tiling · 0.7hybrid tensor format · 0.7TCU-based parallel algorithm · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SASTC: Spatial-Aware Sparse Tensor Completion for Large-Scale Traffic Data Recovery
Renqiu Ouyang, Haotian Wang 0006, Yikun Hu 0001, Wangdong Yang, Kenli Li 0001 |
IEEE Internet Things J. | 1 |
| 2024 | BCB-SpTC: An Efficient Sparse High-Dimensional Tensor Contraction Employing Tensor Core AccelerationabstractSparse tensor contraction (SpTC) is an important operator in tensor networks, which tends to generate a large amount of sparse high-dimensional data, placing higher demands on the computational performance and storage bandwidth of the processor. Using GPUs with powerful arithmetic characteristics is a reliable choice for accelerating SpTC, however, the high dimensionality and sparsity of tensor makes GPU-accelerated SpTC operators suffer from the difficulties of low computational intensity and high memory consumption. The recent introduction of Tensor Core Units (TCUs) on GPUs brings even more powerful arithmetic, which exacerbates the memory wall problem. To cope with the challenges, this paper proposes a new BCB format that linearizes the indices of multidimensional blocks to reduce block index accesses and uses a bitmap to store the distribution of non-zero elements in a block to reduce the storage overhead. A parallel blocking algorithm of BCB-SpTC is designed to divide the binary linear indices into free and contracted indexes to improve the pairing overhead of computational tasks. Then based on the characteristic computation method of TCUs, the proprietary filling method of TCUs is designed to overcome the inefficiency of parallel computation of sparse data on TCUs. Finally, experimental results on the A100 dataset show that BCB-SpTC improves the acceleration ratio by$1.1\times$to$21.3\times$over the existing SpTC GPU method. Haotian Wang 0006, Wangdong Yang, Renqiu Ouyang, Keqin Li 0001, Kenli Li 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2023 | IAP-SpTV: An input-aware adaptive pipeline SpTV via GCN on CPU-GPU
Haotian Wang 0006, Wangdong Yang, Renqiu Ouyang, Kenli Li 0001, Keqin Li 0001 |
J. Parallel Distributed Comput. | 4 |
| 2023 | A Novel Parallel Algorithm for Sparse Tensor Matrix Chain Multiplication via TCU-AccelerationabstractAnalysis of multi-dimensional data, especially tensor decomposition, which extracts latent information, is becoming considerably popular. Although multi-dimensional sparse data is typically processed on multi-core processors, developing highly optimized GPU-basedSparseTensorMatrixChainMultiplication (SpTMCM) is challenging. The purpose of this paper is to investigate a novel approach named SpTMCM and to explore the discovery of SpTMCM coupled with the emerging computing core, Tensor Core Unit (TCU). In contrast to prior work, the proposed novel approach enables a uniform storage format and optimization approach for SpTMCM. We design a hybrid tensor format based on multi-dimensional tiling that divides the tensor depending on the tile threshold to address the inefficient memory accesses caused by the irregular nonzero distribution of the sparse tensor. Further, we develop a TCU-based tensor parallel algorithm with our novel approach to increase the memory bandwidth. Compared to state-of-the-art works, our method achieves$1.16\sim 24.12\times$speedup for SpMTTKRP and$5.07\sim 7.15\times$speedup for SpTTMChain across NVIDIA A100 GPU on a range of real-world sparse tensors. Haotian Wang 0006, Wangdong Yang, Renqiu Ouyang, Kenli Li 0001, Keqin Li 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |