Jiabei Long

dblp:404/6511 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
0009-0004-4719-7303ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Hardware accelerators and domain-specific architectures · 75% Reconfigurable computing and FPGAs · 25%
Artificial intelligence
1 paper
Graph learning · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures › accelerator architecture
accelerator microarchitecture
0.912025
GCNTrain+: A Versatile and Efficient Accelerator for Graph Convolutional Neural Network Training · ACM Trans. Archit. Code Optim. 2025
Reconfigurable computing and FPGAs
dynamic reconfiguration
0.912025
GCNTrain+: A Versatile and Efficient Accelerator for Graph Convolutional Neural Network Training · ACM Trans. Archit. Code Optim. 2025
Hardware accelerators and domain-specific architectures › graph processing accelerator
GCN training accelerator
0.912025
GCNTrain+: A Versatile and Efficient Accelerator for Graph Convolutional Neural Network Training · ACM Trans. Archit. Code Optim. 2025
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.912025
GCNTrain+: A Versatile and Efficient Accelerator for Graph Convolutional Neural Network Training · ACM Trans. Archit. Code Optim. 2025
Machine learning › Graph learning › graph neural network
graph convolutional network
0.312025
GCNTrain+: A Versatile and Efficient Accelerator for Graph Convolutional Neural Network Training · ACM Trans. Archit. Code Optim. 2025
Machine learning › Graph learning › graph neural network training
mini-batch training
0.312025
GCNTrain+: A Versatile and Efficient Accelerator for Graph Convolutional Neural Network Training · ACM Trans. Archit. Code Optim. 2025

Methods — techniques the papers use, named apart from their topics

sparse-dense matrix multiplication · 1.7dynamic clustering · 1.7
YearPublicationVenuePosition
2025 GCNTrain+: A Versatile and Efficient Accelerator for Graph Convolutional Neural Network Training
abstract
Recently, graph convolutional networks (GCNs) have gained wide attention due to their ability to capture node relationships in graphs. One problem appears when full-batch GCN is trained on large graph datasets, where the computational and memory requirements are unacceptable. To address this issue, mini-batch GCN training is introduced to improve the scalability of GCN training for large datasets by sampling and training only a subset of the graph in each batch. Although several acceleration techniques have been designed for boosting the efficiency of full-batch GCN, they lack attention to mini-batch GCN, which differs from full-batch GCN in terms of the sampled dynamic graph structures. Based on our previous work, GCNTrain [ 28 ], which was originally excogitated for accelerating full-batch GCN training, we devise GCNTrain+—a universal accelerator to tackle the performance bottlenecks associated with both full-batch and mini-batch GCN training. GCNTrain+ is equipped with two engines to optimize computation and memory access in GCN training, respectively. To reduce the computation overhead, we propose to dynamically reconfigure the computation order based on the varying data dimensions involved in each training batch. Moreover, we build a unified computation engine to perform the sparse-dense matrix multiplications and sparse-sparse matrix multiplications discovered in GCN training uniformly. To alleviate the memory burden, we devise a two-phased dynamic clustering mechanism to capture data locality as well as customized hardware to reduce the clustering overhead. We evaluate GCNTrain+ on seven datasets, and the result shows that GCNTrain+ achieves 136.0×, 52.6×, 2.2×, and 1.5× speedup over CPU, GPU, GCNAX, and GCNTrain in full-batch GCN training. Additionally, GCNTrain+ outperforms them with speedups of 131.6×, 67.1×, 4.4×, and 1.5× in mini-batch GCN training.
Zhuoran Song, Jiabei Long, Li Jiang 0002, Naifeng Jing, Xiaoyao Liang
ACM Trans. Archit. Code Optim.2