Yijie Nie

dblp:401/8065 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
GPUs and heterogeneous computing · 61% Processor architecture and microarchitecture · 30% Hardware accelerators and domain-specific architectures · 9%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture
dataflow architecture
1.012026
Uni-STC: Unified Sparse Tensor Core · HPCA 2026
GPUs and heterogeneous computing › GPU computing › tensor cores
sparse tensor core
1.012026
Uni-STC: Unified Sparse Tensor Core · HPCA 2026
GPUs and heterogeneous computing › GPU computing
tensor cores
1.012026
Uni-STC: Unified Sparse Tensor Core · HPCA 2026
Hardware accelerators and domain-specific architectures
sparse computation
0.312026
Uni-STC: Unified Sparse Tensor Core · HPCA 2026

Methods — techniques the papers use, named apart from their topics

unified sparse format · 1.0fine-grained task partitioning · 1.0
YearPublicationVenuePosition
2026 Uni-STC: Unified Sparse Tensor Core
abstract
Modern processors are increasingly adopting tensor cores as key computational units. Compared to existing designs for dense and structured sparsity, recent dual-side sparse tensor cores have evolved to support general sparsity. However, existing methods still face limitations on generality (incomplete sparse kernel support prevents broad applicability) and performance (outer-product/row-row schemes yield unsatisfactory hardware utilisation, data reuse, and energy efficiency). In this paper, we propose Uni-STC, a unified sparse tensor core that delivers high-performance dataflows for four key sparse kernels: sparse matrix-vector multiplication (SpMV), sparse matrixsparse vector multiplication (SpMSpV), sparse matrix-multiple vector multiplication (SpMM), and sparse general matrix-matrix multiplication (SpGEMM). To efficiently support these diverse sparse workloads, we first introduce BBC, a unified sparse format co-designed with Uni-STC's dataflow. We then design UniSTC's architecture supporting (1) fine-grained task partitioning to improve resource utilisation, (2) parallel sparse-tile processing to enhance data reuse, and (3) a dynamic network to reduce intermediate data movement and energy consumption. Evaluated across 2893 SuiteSparse and 302 DLMC matrices, Uni-STC demonstrates significant improvements, outperforming the state-of-the-art RM-STC with a$2.21 \times$geomean speedup and$2.96 \times$higher energy efficiency.
Haocheng Lian, Meichen Dong, Yijie Nie, Junzhong Shen, Chun Huang 0006, Bingcai Sui, Weifeng Liu 0002
HPCA5
2024 Leda: Leveraging Tiling Dataflow to Accelerate SpMM on HBM-Equipped FPGAs for GNNs
abstract
Graph neural networks (GNNs) play a pivotal role in extracting insightful representations from graph-structured data, driving advancements across diverse domains. Central to GNNs is the sparse matrix-dense matrix multiplication (SpMM) kernel. However, challenges arise in accelerating SpMM due to the high sparsity and randomly distributed non-zeros in graph matrices. Recently, the high concurrency capability of high bandwidth memory (HBM) has provided a new opportunity for SpMM acceleration. Nonetheless, accelerating SpMM on HBM FPGAs is still non-trivial due to load imbalance and the random memory access patterns.
Enxin Yi, Jiarui Bai, Yijie Nie, Dan Niu, Zhou Jin 0001, Weifeng Liu 0002
ICCAD3