Zhikuang Xin

dblp:323/9202 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0003-3459-5105ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Systems, architecture and hardware · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021
YearPublicationVenuePosition
2026 RDMA-Aware gRPC for Scalable Distributed Reinforcement Learning on HPC Clusters
Zhikuang Xin, Zhenghong Wu, Benxi Tian, Jue Wang 0013, Haikuo Zhang, Rongqiang Cao, Yangang Wang 0002
KSEM (2)1
2026 SEEDTrans: Interpretable Day-Ahead Photovoltaic Power Forecasting with Multi-level Series Decomposition Transformer
Zhikuang Xin, Meng Wan, Benxi Tian, Jue Wang 0013, Peng Shi 0006, Haikuo Zhang, Rongqiang Cao, Xue Miao, Zhenbing Zhao, Yangang Wang 0002
KSEM (2)1
2025 Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor Cores
abstract
General-purpose Sparse Matrix-Matrix Multiplication (SpMM) is a fundamental kernel in scientific computing and deep learning. The emergence of new matrix computation units such as Tensor Cores (TCs) brings more opportunities for SpMM acceleration. However, in order to fully unleash the power of hardware performance, systematic optimization is required. In this paper, we propose Acc-SpMM, a high-performance SpMM library on TCs, with multiple optimizations, including data-affinity-based reordering, memory efficient compressed format, high-throughput pipeline, and adaptive sparsity-aware load balancing. In contrast to the state-of-the-art SpMM kernels on various NVIDIA GPU architectures with a diverse range of benchmark matrices, Acc-SpMM achieves significant performance improvements, on average 2.52x (up to 5.11x) speedup on RTX 4090, on average 1.91x (up to 4.68x) speedup on A800, and on average 1.58x (up to 3.60x) speedup on H100 over cuSPARSE.
Haisha Zhao, San Li, Chunbao Zhou, Jue Wang 0013, Zhikuang Xin, Shunde Li, Yangang Wang 0002, Xuebin Chi
PPoPP6
2024 Reinforcement Learning for Scientific Application: A Survey
Zhikuang Xin, Zhenghong Wu, Jue Wang 0013, Yangang Wang 0002
KSEM (5)1
2023 ANT-MOC: Scalable Neutral Particle Transport Using 3D Method of Characteristics on Multi-GPU Systems
abstract
The Method Of Characteristic (MOC) to solve the Neutron Transport Equation (NTE) is the core of full-core simulation for reactors. High resolution is enabled by discretizing the NTE through massive tracks to traverse the 3D reactor geometry. However, the 3D full-core simulation is prohibitively expensive because of the high memory consumption and the severe load imbalance. To deal with these challenges, we develop ANT-MOC1. Specifically, we build a performance model for memory footprint, computation and communication, based on which a track management strategy is proposed to overcome the resolution bottlenecks caused by limited GPU memory. Furthermore, we implement a novel multi-level load mapping strategy to ensure load balancing among nodes, GPUs, and CUs. ANT-MOC enables a 3D full-core reactor simulation with 100 billion tracks on 16,000 GPUs, with 70.69% and 89.38% parallel efficiency for strong scalability and weak scalability, respectively.
Shunde Li, Zongguo Wang, Lingkun Bu, Jue Wang 0013, Zhikuang Xin, Shigang Li 0002, Yangang Wang 0002, Yangde Feng, Peng Shi 0006, Xuebin Chi
SC5
2022 VenusAI: An artificial intelligence platform for scientific discovery on supercomputers
Tiechui Yao, Jue Wang 0013, Meng Wan, Zhikuang Xin, Yangang Wang 0002, Rongqiang Cao, Shigang Li 0002, Xuebin Chi
J. Syst. Archit.4