He Bai 0005

dblp:73/5171-5 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2026
0000-0001-5418-0375ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021
YearPublicationVenuePosition
2026 Full-Core Fluid-Structure-Interaction Simulation of Nuclear Reactor on CPU+GPU Hybrid Clusters
abstract
Nuclear reactor FSI simulation faces two key challenges: "Mapping wall" bottleneck in data transfer across non-matching mesh coupling interfaces; Low hardware utilization from multi-physics solvers’ heterogeneous core tasks (compute- vs. memory-intensive). Therefore, an innovative FSI framework integrating two strategies is proposed: Scalable radial basis function mapping—restructuring the global problem into massive independent subproblems via task partitioning, preallocation, and multi-granularity load balancing to eliminate communication overhead; Dependency-aware multi-stream optimization—deeply overlapping heterogeneous solver tasks to maximize hardware utilization. It first achieves parameter transfer across ∼90,000 non-matching coupling interfaces in China Experimental Fast Reactor, with 86.36% strong scaling and 94.01% weak scaling. The combined optimizations yield ∼60% performance gain, increase strong scaling by over 20 percentage points, and achieve high weak scaling of ∼97%. Moreover, the FSI results align well with publicly available data, verifying its correctness.
Xue Miao, Jue Wang 0013, Qida Lin, Shufei Zhang, Rongqiang Cao, Chunbao Zhou, Ningming Nie, He Bai 0005, Yangang Wang 0002
HPDC8
2026 HIP-DFPT: Scalable Optimization of Irregular Workloads in Quantum Perturbation on GPU Clusters
Meng Wan, Jue Wang 0013, Shunde Li, Honghui Shang, He Bai 0005, Peng Shi 0006, Yuchen Pang, Ying Liu 0055, Jinrong Jiang, Yangang Wang 0002, Xuebin Chi
IEEE Trans. Parallel Distributed Syst.6
2025 MISA-AKMC : Achieve Kinetic Monte Carlo Simulation of 20 Quadrillion Atoms on GPU Clusters
abstract
The Atomic Kinetic Monte Carlo (AKMC) method provides insights into the macroscopic behavior of materials through atomistic-level simulations and finds broad applications in materials science innovation. Improving simulation scale and performance remains a consistent focus in the development of parallel AKMC software. We port the AKMC software to GPU clusters. To alleviate the memory pressure in large-scale complex system simulations, we redesign the data layout and propose the Lattice Data Compression and Vacancy Data Decompression algorithms. Additionally, We propose a multi-level pipeline scheme combined with an on-demand communication forwarding and merging strategy to reduce data transfer and communication overhead. Compared to state-of-the-art KMC software, MISA-AKMC achieves a 10.41-fold improvement in computational throughput and a 52.07-fold expansion in simulation scale. We implement the first true micrometer-scale AKMC simulation involving 20 quadrillion atoms on GPU clusters. MISA-AKMC achieves 96.03% parallel efficiency in weak scaling and 85.29% in strong scaling on 16,000 GPUs.
Shunde Li, Ningming Nie, Jue Wang 0013, He Bai 0005, Genshen Chu, Xinfu He, Yangang Wang 0002, Changjun Hu, Xuebin Chi
SC5
2023 Efficient Algorithm Design of Optimizing SpMV on GPU
abstract
Sparse matrix-vector multiplication (SpMV) is a fundamental building block for various numerical computing applications. However, most existing GPU-SpMV approaches may suffer from either long preprocessing overhead, load imbalance, format conversion, bad memory access patterns. In this paper, we proposed two new SpMV algorithms:flat andline-enhance, as well as their implementations, for GPU systems to overcome the above shortcomings. Our algorithms work directly on the CSR sparse matrix format. To achieve high performance: 1) for load balance, theflat algorithm uses non-zero splitting andline-enhance uses a mix of row and non-zero splitting; 2) memory access patterns are designed for both algorithms for data loading, storing and reduction steps; and 3) an adaptive approach is proposed to select appropriate algorithm and parameters based on matrix characteristics.
Genshen Chu, Yuanjie He, Lingyu Dong, Zhezhao Ding, He Bai 0005, Changjun Hu
HPDC6