Chunbao Zhou

dblp:83/5685 · DBLP profile ↗
← Back
12ranked-venue papers
0as first author
10since 2021 · last 2026
0000-0003-3812-5215ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 8 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Full-Core Fluid-Structure-Interaction Simulation of Nuclear Reactor on CPU+GPU Hybrid Clusters
abstract
Nuclear reactor FSI simulation faces two key challenges: "Mapping wall" bottleneck in data transfer across non-matching mesh coupling interfaces; Low hardware utilization from multi-physics solvers’ heterogeneous core tasks (compute- vs. memory-intensive). Therefore, an innovative FSI framework integrating two strategies is proposed: Scalable radial basis function mapping—restructuring the global problem into massive independent subproblems via task partitioning, preallocation, and multi-granularity load balancing to eliminate communication overhead; Dependency-aware multi-stream optimization—deeply overlapping heterogeneous solver tasks to maximize hardware utilization. It first achieves parameter transfer across ∼90,000 non-matching coupling interfaces in China Experimental Fast Reactor, with 86.36% strong scaling and 94.01% weak scaling. The combined optimizations yield ∼60% performance gain, increase strong scaling by over 20 percentage points, and achieve high weak scaling of ∼97%. Moreover, the FSI results align well with publicly available data, verifying its correctness.
Xue Miao, Jue Wang 0013, Qida Lin, Shufei Zhang, Rongqiang Cao, Chunbao Zhou, Ningming Nie, He Bai 0005, Yangang Wang 0002
HPDC6
2026 Hierarchical Reinforcement Learning Agent for Distributed Event-Driven Multi-asset LOBs Simulation
Zhenglie Sun, Tian Gou, Chunbao Zhou, Yangang Wang 0002
KSEM (2)3
2026 DESMA: A Distributed Discrete-Event Simulator for Synchronous Multi-asset Limit Order Books
Zhenglie Sun, Tian Gou, Chunbao Zhou, Yangang Wang 0002
KSEM (4)3
2026 TAC: Cache-Based System for Accelerating Billion-Scale GNN Training on Multi-GPU Platform
abstract
Graph neural networks (GNNs) have been proven to have increasingly widespread applications in the real world. In the mainstream mini-batch training mode, multiple cache-based GNN training acceleration systems have been proposed because of the possibility of selecting the same vertex multiple times during the sampling process. However, on ultra-large scale graphs, especially those exhibiting power-law characteristics, these systems are difficult to fully utilize the distribution characteristics of cached data, which limits training performance. To this end, we propose TAC, a GNN training acceleration system that fully exploits the distribution characteristics of cached data to optimize both data transmission and computational efficiency. Specifically, we have designed a data affinity optimization algorithm that significantly enhances the locality of cache access. Secondly, an adaptive sparse matrix operator for sparsity perception is proposed, which dynamically selects the optimal computing mode based on the location of data. Finally, we have constructed a fine-grained training pipeline that maximizes system parallelism by hiding the sampling and computation. The experimental results show that TAC significantly outperforms existing state-of-the-art cache acceleration systems on multiple benchmark datasets, demonstrating higher training efficiency.
Jue Wang 0013, Xingguo Shi, Junyu Gu, Peng Di, Sian Li, Chunbao Zhou, Lian Zhao, Yangang Wang 0002, Xuebin Chi
PPoPP10
2025 ParGNN: A Scalable Graph Neural Network Training Framework on multi-GPUs
abstract
Full-batch Graph Neural Network (GNN) training is indispensable for interdisciplinary applications. Although fullbatch training has advantages in convergence accuracy and speed, it still faces challenges such as severe load imbalance and high communication traffic overhead. In order to address these challenges, we propose ParGNN, an efficient full-batch training system for GNNs, which adopts a profiler-guided adaptive load balancing method along with graph over-partition to alleviate load imbalance. Based on the over-partition results, we present a subgraph pipeline algorithm to overlap communication and computation while maintaining the accuracy of GNN training. Extensive experiments demonstrate that ParGNN can not only obtain the highest accuracy but also reach the preset accuracy in the shortest time. In the end-to-end experiments performed on the four datasets, ParGNN outperforms the two state-of-theart full-batch GNN systems, PipeGCN and DGL, achieving the highest speedup of $2.7 \times$ and $21.8 \times$ times respectively.
Junyu Gu, Shunde Li, Rongqiang Cao, Jue Wang 0013, Shigang Li 0002, Chunbao Zhou, Yangang Wang 0002, Xuebin Chi
DAC9
2025 Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor Cores
abstract
General-purpose Sparse Matrix-Matrix Multiplication (SpMM) is a fundamental kernel in scientific computing and deep learning. The emergence of new matrix computation units such as Tensor Cores (TCs) brings more opportunities for SpMM acceleration. However, in order to fully unleash the power of hardware performance, systematic optimization is required. In this paper, we propose Acc-SpMM, a high-performance SpMM library on TCs, with multiple optimizations, including data-affinity-based reordering, memory efficient compressed format, high-throughput pipeline, and adaptive sparsity-aware load balancing. In contrast to the state-of-the-art SpMM kernels on various NVIDIA GPU architectures with a diverse range of benchmark matrices, Acc-SpMM achieves significant performance improvements, on average 2.52x (up to 5.11x) speedup on RTX 4090, on average 1.91x (up to 4.68x) speedup on A800, and on average 1.58x (up to 3.60x) speedup on H100 over cuSPARSE.
Haisha Zhao, San Li, Chunbao Zhou, Jue Wang 0013, Zhikuang Xin, Shunde Li, Yangang Wang 0002, Xuebin Chi
PPoPP4
2024 POSTER: ParGNN: Efficient Training for Large-Scale Graph Neural Network on GPU Clusters
abstract
Full-batch graph neural network (GNN) training is essential for interdisciplinary applications. Large-scale graph data is usually divided into subgraphs and distributed across multiple compute units to train GNN. The state-of-the-art load balancing method based on direct graph partition is too rough to effectively achieve true load balancing on GPU clusters. We propose ParGNN, which employs a profiler-guided load balance workflow in conjunction with graph repartition to alleviate load imbalance and minimize communication traffic. Experiments have verified that ParGNN has the capability to scale to larger clusters.
Shunde Li, Junyu Gu, Jue Wang 0013, Tiechui Yao, Yumeng Shi, Shigang Li 0002, Weiting Xi, Shushen Li, Chunbao Zhou, Yangang Wang 0002, Xuebin Chi
PPoPP10
2023 An Auto-Parallel Method for Deep Learning Models Based on Genetic Algorithm
abstract
As the size of datasets and neural network models increases, automatic parallelization methods for models have become a research hotspot in recent years. The existing auto-parallel methods based on machine learning or graph algorithms still have issues with search efficiency and applicability. This paper proposes an automatic parallel method based on a dual-population genetic algorithm, TGA, which transforms model partitioning and placement into an integer linear programming problem and constructs a cost model to evaluate the solution. The solution space is built using the neural network’s dataflow graph and device cluster’s topology, and the dual-population genetic algorithm is used to search for the optimal model parallel strategy. Experiments with various models show that the proposed method can improve single-step execution time by up to 42% compared to the Baechi method and up to 37.7% compared to the Hierarchical method.
Chengchuang Huang, Yijie Ni, Chunbao Zhou, Jue Wang 0013, Mingyao Zhou, Meiting Xue, Yunquan Zhang
ICPADS4
2023 A Scalable Hybrid Total FETI Method for Massively Parallel FEM Simulations
abstract
The Hybrid Total Finite Element Tearing and Interconnecting (HTFETI) method plays an important role in solving large-scale and complex engineering problems. This method needs to handle numerous matrix-vector multiplications. Directly calling the vendor-optimized library for general matrix-vector multiplication (gemv) on GPU leads to low performance, since it does not consider optimizations for different matrix sizes in HTFETI, i.e. different row and column sizes. In addition, state-of-the-art graph partitioning methods cannot guarantee load balancing for HTFETI, since the matrix size is determined by the length of the subdomain boundary. To solve the problems above, we first port gemv to the multi-stream pipeline scheme and develop a new batched kernel function on GPU, which brings 15%~30% throughput improvement and 37% average GFLOPs improvement, respectively. We also propose a multi-grained load-balancing scheme based on graph repartitioning and work-stealing, and the load imbalance ratio is down to 1.05~1.09 from 1.5. We have successfully applied the scalable HTFETI method to simulate the whole core assembly of China Experimental Fast Reactor (CEFR) for steady-state analysis, and the efficiencies of weak scalability and strong scalability reach 78% and 72% on 12,288 GPUs, respectively. As far as we know, this is the first time that HTFETI has been used in large-scale and high-fidelity whole core assembly simulation.
Kehao Lin, Chunbao Zhou, Ningming Nie, Jue Wang 0013, Shigang Li 0002, Yangde Feng, Yangang Wang 0002, Kehan Yao, Tiechui Yao, Jian Wan 0001
PPoPP2
2023 Large-Scale Simulation of Structural Dynamics Computing on GPU Clusters
abstract
Structural dynamics simulation plays an important role in research on reactor design and complex engineering. The Hybrid Total Finite Element Tearing and Interconnecting (HTFETI) method combined with Newmark method is an efficient way to solve large-scale structural dynamics problems. However, the sparse direct solver and the load imbalance caused by inconsistent density models are two critical issues limiting the performance and the scalability of structural dynamics computing. For the former, we propose an efficient variable-size batched method to accelerate SpMV on GPUs. For the latter, we establish an online performance prediction model, based on which we then design a novel inter-cluster subdomain fine-tuning algorithm to balance the workload of HTFETI parallel computing. We are the first to achieve the high-fidelity structural dynamics simulation of China Experimental Fast Reactor core assembly with up to 53.4 billion grids. The weak and strong scalability efficiencies reach 91.77% and 86.13% on 12,800 GPUs, respectively.
Yumeng Shi, Ningming Nie, Jue Wang 0013, Kehao Lin, Chunbao Zhou, Shigang Li 0002, Kehan Yao, Shunde Li, Yangde Feng, Yangang Wang 0002
SC5
2016 Extreme-scale phase field simulations of coarsening dynamics on the sunway taihulight supercomputer
abstract
Many important properties of materials such as strength, ductility, hardness and conductivity are determined by the microstructures of the material. During the formation of these microstructures, grain coarsening plays an important role. The Cahn-Hilliard equation has been applied extensively to simulate the coarsening kinetics of a two-phase microstructure. It is well accepted that the limited capabilities in conducting large scale, long time simulations constitute bottlenecks in predicting microstructure evolution based on the phase field approach. We present here a scalable time integration algorithm with large stepsizes and its efficient implementation on the Sunway TaihuLight supercomputer. The highly nonlinear and severely stiff Cahn-Hilliard equations with degenerate mobility for microstructure evolution are solved at extreme scale, demonstrating that the latest advent of high performance computing platform and the new advances in algorithm design are now offering us the possibility to simulate the coarsening dynamics accurately at unprecedented spatial and time scales.
Jian Zhang 0070, Chunbao Zhou, Yangang Wang 0002, Lili Ju, Qiang Du 0001, Xuebin Chi, Dexun Chen
SC2
2015 Optimizing the Bayesian Inference of Phylogeny on Graphic Processors
abstract
Searching for the evolutionary relationships between groups of organism has become a routine procedure in molecular biology. MrBayes is a popular model based phylogenetic inference tool using Bayesian statistics. Unfortunately, the computational cost is very high, resulting in undesirably long execution time. In this paper, we present what we believe the fastest solution of the MrBayes MC3 algorithm running on off-the-shelf graphic processors. The performance benefits are offered by the multi-granularity parallelism model, coarse-grained GPU kernel system, efficient thread arrangement strategy and GPU code level optimizations. MrBayes goMC3 (proposed herein) provides a significant performance improvement over the sequential MrBayes MC3 by a speedup of up to 48× when using single Tesla C2075 GPU card, whereas a speedup factor of 77× can be achieved when using dual GPUs. In comparison to the state-of-the-art version of other publicly available GPU implementations of MrBayes MC3, the cumulative optimizations adopted in goMC3 resulted in a speedup of up 2.5× over oMC3 (v1.0), 1.75× over tgMC3 (v1.0) and 1.46× over nMC3(v2.1.1) for realistic empirical biological datasets. Besides, experimental results indicated that goMC3 outstrips these GPU implementations on the analysis of simulated datasets composed of ultra-large-scale sequences. As a consequence, the reported performance improvement of goMC3 is significant and appears to scale well with increasing dataset sizes.
Cheng Ling, Chunbao Zhou, Arong Luo, Guoguang Zhao, Tsuyoshi Hamada, Xiaoyan Zhu 0001
CCGRID2