Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Renqiu Ouyang

dblp:351/5290 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2025
0000-0002-8712-4561ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
GPUs and heterogeneous computing · 46% High-performance computing · 42% Memory systems · 12%

Topics — the 6 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing › tensor computation
sparse tensor contraction
1.422024
BCB-SpTC: An Efficient Sparse High-Dimensional Tensor Contraction Employing Tensor Core Acceleration · IEEE Trans. Parallel Distributed Syst. 2024
A Novel Parallel Algorithm for Sparse Tensor Matrix Chain Multiplication via TCU-Acceleration · IEEE Trans. Parallel Distributed Syst. 2023
GPUs and heterogeneous computing › GPU computing
tensor cores
1.422024
BCB-SpTC: An Efficient Sparse High-Dimensional Tensor Contraction Employing Tensor Core Acceleration · IEEE Trans. Parallel Distributed Syst. 2024
A Novel Parallel Algorithm for Sparse Tensor Matrix Chain Multiplication via TCU-Acceleration · IEEE Trans. Parallel Distributed Syst. 2023
GPUs and heterogeneous computing
GPU computing
0.812024
BCB-SpTC: An Efficient Sparse High-Dimensional Tensor Contraction Employing Tensor Core Acceleration · IEEE Trans. Parallel Distributed Syst. 2024
Memory systems › memory bandwidth
memory bandwidth optimization
0.812024
BCB-SpTC: An Efficient Sparse High-Dimensional Tensor Contraction Employing Tensor Core Acceleration · IEEE Trans. Parallel Distributed Syst. 2024
High-performance computing › sparse linear algebra
sparse matrix storage format
0.812024
BCB-SpTC: An Efficient Sparse High-Dimensional Tensor Contraction Employing Tensor Core Acceleration · IEEE Trans. Parallel Distributed Syst. 2024
High-performance computing › tensor computation
tensor decomposition
0.212023
A Novel Parallel Algorithm for Sparse Tensor Matrix Chain Multiplication via TCU-Acceleration · IEEE Trans. Parallel Distributed Syst. 2023

Methods — techniques the papers use, named apart from their topics

vectorization · 0.8blocking · 0.8bitmap encoding · 0.8multi-dimensional tiling · 0.7hybrid tensor format · 0.7TCU-based parallel algorithm · 0.7
YearPublicationVenuePosition
2025 SASTC: Spatial-Aware Sparse Tensor Completion for Large-Scale Traffic Data Recovery
Renqiu Ouyang, Haotian Wang 0006, Yikun Hu 0001, Wangdong Yang, Kenli Li 0001
IEEE Internet Things J.1
2024 BCB-SpTC: An Efficient Sparse High-Dimensional Tensor Contraction Employing Tensor Core Acceleration
abstract
Sparse tensor contraction (SpTC) is an important operator in tensor networks, which tends to generate a large amount of sparse high-dimensional data, placing higher demands on the computational performance and storage bandwidth of the processor. Using GPUs with powerful arithmetic characteristics is a reliable choice for accelerating SpTC, however, the high dimensionality and sparsity of tensor makes GPU-accelerated SpTC operators suffer from the difficulties of low computational intensity and high memory consumption. The recent introduction of Tensor Core Units (TCUs) on GPUs brings even more powerful arithmetic, which exacerbates the memory wall problem. To cope with the challenges, this paper proposes a new BCB format that linearizes the indices of multidimensional blocks to reduce block index accesses and uses a bitmap to store the distribution of non-zero elements in a block to reduce the storage overhead. A parallel blocking algorithm of BCB-SpTC is designed to divide the binary linear indices into free and contracted indexes to improve the pairing overhead of computational tasks. Then based on the characteristic computation method of TCUs, the proprietary filling method of TCUs is designed to overcome the inefficiency of parallel computation of sparse data on TCUs. Finally, experimental results on the A100 dataset show that BCB-SpTC improves the acceleration ratio by$1.1\times$to$21.3\times$over the existing SpTC GPU method.
Haotian Wang 0006, Wangdong Yang, Renqiu Ouyang, Keqin Li 0001, Kenli Li 0001
IEEE Trans. Parallel Distributed Syst.4
2023 IAP-SpTV: An input-aware adaptive pipeline SpTV via GCN on CPU-GPU
Haotian Wang 0006, Wangdong Yang, Renqiu Ouyang, Kenli Li 0001, Keqin Li 0001
J. Parallel Distributed Comput.4
2023 A Novel Parallel Algorithm for Sparse Tensor Matrix Chain Multiplication via TCU-Acceleration
abstract
Analysis of multi-dimensional data, especially tensor decomposition, which extracts latent information, is becoming considerably popular. Although multi-dimensional sparse data is typically processed on multi-core processors, developing highly optimized GPU-basedSparseTensorMatrixChainMultiplication (SpTMCM) is challenging. The purpose of this paper is to investigate a novel approach named SpTMCM and to explore the discovery of SpTMCM coupled with the emerging computing core, Tensor Core Unit (TCU). In contrast to prior work, the proposed novel approach enables a uniform storage format and optimization approach for SpTMCM. We design a hybrid tensor format based on multi-dimensional tiling that divides the tensor depending on the tile threshold to address the inefficient memory accesses caused by the irregular nonzero distribution of the sparse tensor. Further, we develop a TCU-based tensor parallel algorithm with our novel approach to increase the memory bandwidth. Compared to state-of-the-art works, our method achieves$1.16\sim 24.12\times$speedup for SpMTTKRP and$5.07\sim 7.15\times$speedup for SpTTMChain across NVIDIA A100 GPU on a range of real-world sparse tensors.
Haotian Wang 0006, Wangdong Yang, Renqiu Ouyang, Kenli Li 0001, Keqin Li 0001
IEEE Trans. Parallel Distributed Syst.4