Honghui Shang

dblp:198/9226 · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
19since 2021 · last 2026
0000-0003-4957-4251ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 20 · 5 first-author · 19 since 2021
YearPublicationVenuePosition
2026 A Fully GPU-Accelerated Framework for High-Performance Configuration Interaction Selection with Neural Network Quantum States
abstract
AI-driven methods have demonstrated considerable success in tackling the central challenge of accurately solving the Schrödinger equation for complex many-body systems. Among neural network quantum state (NNQS) approaches, the NNQS-SCI (Selected Configuration Interaction) method stands out as a state-of-the-art technique, recognized for its high accuracy and scalability. However, its application to larger systems is severely constrained by a hybrid CPU-GPU architecture. Specifically, centralized CPU-based global de-duplication creates a severe scalability barrier due to communication bottlenecks, while host-resident coupled-configuration generation induces prohibitive computational overheads. We introduce QiankunNet-cuSCI, a fully GPU-accelerated SCI framework designed to overcome these bottlenecks. It first integrates a distributed, load-balanced global de-duplication algorithm to minimize redundancy and communication overhead at scale. To address compute limitations, it employs specialized, fine-grained CUDA kernels for exact coupled configuration generation. Finally, to break the single-GPU memory barrier exposed by this full acceleration, it incorporates a GPU memory-centric runtime featuring GPU-side pooling, streaming mini-batches, and overlapped offloading. This design enables much larger configuration spaces and shifts the bottleneck from host-side limitations back to on-device inference. Our evaluation demonstrates that our work fundamentally expands the scale of solvable problems. On an NVIDIA A100 cluster with 64 GPUs, our work achieves up to 2.32 × end-to-end speedup over the highly-optimized NNQS-SCI baseline while preserving the same chemical accuracy. Furthermore, it demonstrates excellent distributed performance, maintaining over 90% parallel efficiency in strong scaling tests.
Daran Sun, Bowen Kan, Haoquan Long, Hairui Zhao 0002, Haoxu Li, Ankang Feng, Wenjing Huang 0002, Yida Gu, Honghui Shang, Yunquan Zhang, Dingwen Tao, Ninghui Sun, Guangming Tan
HPDC12
2026 HIP-DFPT: Scalable Optimization of Irregular Workloads in Quantum Perturbation on GPU Clusters
Meng Wan, Jue Wang 0013, Shunde Li, Honghui Shang, He Bai 0005, Peng Shi 0006, Yuchen Pang, Ying Liu 0055, Jinrong Jiang, Yangang Wang 0002, Xuebin Chi
IEEE Trans. Parallel Distributed Syst.5
2025 TensorMD: Molecular Dynamics Simulation with Ab Initio Accuracy of 50 Billion Atoms
abstract
Molecular dynamics simulation emerges as an important area that HPC+AI helps to investigate the physical properties, with machine-learning interatomic potentials (MLIPs) being used. General-purpose machine-learning (ML) tools have been leveraged in MLIPs, but they are not perfectly matched with each other, since many optimization opportunities in MLIPs have been missed by ML tools. This inefficiency arises from the fact that HPC+AI applications work with far more computational complexity compared with pure AI scenarios. This paper has developed an MLIP, named TensorMD, independently from any ML tool. TensorMD has been evaluated on two supercomputers and scaled to 51.8 billion atoms, i.e., ~ 3× compared with state-of-the-art.
Yucheng Ouyang, Ying Liu 0055, Honghui Shang, Zhenchuan Chen, Jiahao Shan, Huimin Cui, Xiaobing Feng 0002, Xingyu Gao 0003, Haifeng Song 0003, Xin Chen 0023, Rongfen Lin
PPoPP3
2025 NNQS-SCI: Tackling Trillion-Dimensional Hilbert Space with Adaptive Neural Network Quantum States
abstract
Neural Network Quantum States (NNQS) offer a powerful variational Monte Carlo (VMC) approach for quantum many-body problems, balancing polynomial scaling with high expressive power. However, scaling NNQS to large chemical systems faces challenges in preserving accuracy with exact energy and managing vast configurations efficiently. In this work, we introduce NNQS-SCI, a high-performance Selected Configuration Interaction (SCI) based NNQS method designed to overcome these limitations. NNQS-SCI employs highly parallelized Slater-Condon rules for fast local energy evaluations, avoiding accuracy loss, while its adaptive SCI engine dynamically manages billions of configurations without space explosion or arbitrary cutoffs that plague other NNQS-CI approaches. Optimized for extreme scalability via multi-level parallelism and memory compression, NNQS-SCI successfully simulates systems up to 152 spin orbitals, tackling Hilbert space dimensions exceeding 1014 and demonstrating significant advances in scale and efficiency. NNQS-SCI thus provides a robust and scalable path towards high-accuracy quantum chemistry on high-performance computing platforms.
Bowen Kan, Yumeng Zhou, Daiyou Xie, Yunquan Zhang, Honghui Shang
SC6
2025 TENSORMD: Accelerating Molecular Dynamics with a High-Performance Machine Learning Interatomic Potential
abstract
AI has been integrated into HPC across various scientific fields, significantly enhancing performance. In molecular dynamics simulations, HPC+AI facilitates the investigation of atomic-scale physical properties using machine-learning interatomic potentials (MLIPs). However, general-purpose ML tools (e.g., TensorFlow) used in MLIPs are not optimally matched, leading to missed optimization opportunities due to the higher computational complexity and greater diversity of HPC+AI applications compared to pure AI scenarios. To address this, we introduce TensorMD, an MLIP independent of existing ML tools, enabling flexible optimizations that standard ML frameworks cannot support. TensorMD outperforms a state-of-the-art MLIP—winner of the 2020 Gordon Bell Prize and built on an ML tool—by 1.88 × on NVIDIA A100 GPU. Additionally, TensorMD was evaluated on two supercomputers with different architectures, achieving significantly reduced time-to-solution and supporting molecular dynamics simulations at scales beyond 50 billion atoms.
Yucheng Ouyang, Ying Liu 0055, Xin Chen 0023, Honghui Shang, Zhenchuan Chen, Rongfen Lin, Xingyu Gao 0003, Jiahao Shan, Haifeng Song 0003, Huimin Cui, Xiaobing Feng 0002, Jingling Xue
SC5
2025 Fast and Scalable Neural Network Quantum States Method for Molecular Potential Energy Surfaces
abstract
The Neural Network Quantum States (NNQS) method is highly promising for accurately solving the Schrödinger equation, yet it encounters challenges such as computational demands and slow rates of convergence. To address the high computational requirements, we introduce optimizations including a cross-sample KV cache sharing technique to enhance sampling efficiency, Quantum Bitwise and BloomHash methods for more efficient local energy computation, and mixed-precision training strategies to boost computational efficiency. To overcome the issue of slow convergence, we propose a parallel training algorithm for NNQS under second quantization to accelerate the training of base models for molecular potential surfaces. Our approach achieves up to 27-fold acceleration specifically in local energy calculations in systems with 154 spin orbitals and demonstrates strong and weak scaling efficiencies of 98% and 97%, respectively, on the H$_{2}$O$_{2}$potential surface training set. The parallelized implementation of transformer-based NNQS is highly portable on various high-performance computing architectures, offering new perspectives on quantum chemistry simulations.
Yangjun Wu, Wanlu Cao, Honghui Shang
IEEE Trans. Parallel Distributed Syst.4
2025 Large-Scale Neural Network Quantum States Calculation for Quantum Chemistry on a New Sunway Supercomputer
Yangjun Wu, Li Shen 0001, Hong Qian, Honghui Shang
IEEE Trans. Parallel Distributed Syst.5
2024 Scalable and Differentiable Simulator for Quantum Computational Chemistry
abstract
We develop a high-performance simulator for variational quantum eigensolver (VQE), the major innovations include: (1) A differentiable matrix product state (MPS) based VQE simulator that seamlessly integrates MPS into the automatic differentiation framework, which overcomes the exponential memory growth of state-vector simulator and can efficiently calculate gradients with a cost independent of the number of parameters; (2) A dynamic scheme to distribute the gradient calculations to achieve good load balance; (3) A parallel adaptive VQE which integrates our differentiable MPS simulator to further enhance the simulation performance; (4) Study of real chemical systems with convergence to chemical accuracy using our simulator, achieving nearly linearly strong and weak scaling for chemical systems with up to 100 qubits. Our simulator provides an ideal test ground for VQE and paves the way of benchmarking large-scale VQE experiments on near-term quantum computers.
Zhiqian Xu 0005, Honghui Shang, Xiongzhi Zeng, Yunquan Zhang, Chu Guo
IPDPS2
2024 Pushing the Limit of Quantum Mechanical Simulation to the Raman Spectra of a Biological System with 100 Million Atoms
abstract
Raman spectroscopy offers invaluable insights into the chemical composition and structural characteristics of various materials, making it a powerful tool for structural analysis. However, accurate quantum mechanical simulations of Raman spectra for large systems, such as biological materials, have been limited due to immense computational costs and technical challenges. In this study, we developed efficient algorithms and optimized implementations on heterogeneous computing architectures to enable fast and highly scalable ab initio simulations of Raman spectra for large-scale biological systems with up to 100 million atoms. Our simulations have achieved nearly linear strong and weak scaling on two cutting-edge high-performance computing systems, with peak FP64 performances reaching 400 PFLOPS on 96,000 nodes of new Sunway supercomputer and 85 PFLOPS on 6,000 node of ORISE supercomputer. These advances provide promising prospects for extending quantum mechanical simulations to biological systems.
Honghui Shang, Ying Liu 0055, Zhikun Wu, Zhenchuan Chen, Jinfeng Liu 0004, Meiyue Shao, Yingzhou Li, Bowen Kan, Huimin Cui, Xiaobing Feng 0002, Yunquan Zhang, Donald G. Truhlar, Hong An, Xiao He 0004, Jinlong Yang 0003
SC1
2023 NNQS-Transformer: an Efficient and Scalable Neural Network Quantum States Approach for Ab initio Quantum Chemistry
abstract
Neural network quantum state (NNQS) has emerged as a promising candidate for quantum many-body problems, but its practical applications are often hindered by the high cost of sampling and local energy calculation. We develop a high-performance NNQS method for ab initio electronic structure calculations. The major innovations include: (1) A transformer based architecture as the quantum wave function ansatz; (2) A data-centric parallelization scheme for the variational Monte Carlo (VMC) algorithm which preserves data locality and well adapts for different computing architectures; (3) A parallel batch sampling strategy which reduces the sampling cost and achieves good load balance; (4) A parallel local energy evaluation scheme which is both memory and computationally efficient; (5) Study of real chemical systems demonstrates both the superior accuracy of our method compared to state-of-the-art and the strong and weak scalability for large molecular systems with up to 120 spin orbitals.
Yangjun Wu, Chu Guo, Honghui Shang
SC5
2023 Portable and Scalable All-Electron Quantum Perturbation Simulations on Exascale Supercomputers
abstract
Quantum perturbation theory is pivotal in determining the critical physical properties of materials. The first-principles computations of these properties have yielded profound and quantitative insights in diverse domains of chemistry and physics. In this work, we propose a portable and scalable OpenCL implementation for quantum perturbation theory, which can be generalized across various high-performance computing (HPC) systems. Optimal portability is realized through the utilization of a cross-platform unified interface and a collection of performance-portable heterogeneous optimizations. Exceptional scalability is attained by addressing major constraints on memory and communication, employing a locality-enhancing task mapping strategy and a packed hierarchical collective communication scheme. Experiments on two advanced supercomputers demonstrate that our implementation exhibits remarkably performance on various material systems, scaling the system to 200,000 atoms with all-electron precision. This research enables all-electron quantum perturbation simulations on substantially larger molecular scales, with a potentially significant impact on progress in material sciences.
Zhikun Wu, Yangjun Wu, Ying Liu 0055, Honghui Shang, Yingxiang Gao, Zhongcheng Zhang, Yingchi Long, Xiaobing Feng 0002, Huimin Cui
SC4
2023 Redesigning OpenKMC for Multi-Component Trillion-Atom Simulations on the New Sunway Supercomputer
abstract
The atomic kinetic Monte Carlo method plays an important role in material simulations by connecting the microscale mechanism with macroscale evolution. However, the long-time simulation of multi-component materials is highly challenging because it demands significant computing resources. With the advent of exascale computing, ultra-high computing power can enable kinetic Monte Carlo (KMC) simulations. In this paper, we deeply optimize OpenKMC for the new-generation Sunway supercomputer. This includes optimizing the memory access for the SW39000 architecture, eliminating various redundant computations at growing scales, and proposing a communication strategy for heterogeneous platforms. In addition, we expanded OpenKMC's simulation for multi-component alloys. Finally, the acceleration framework can produces a$37\times$performance enhancement on the Sunway platform. Furthermore, when powered by 10 million cores, our program can perform trillion-atom simulations of complex multi-component alloys with 85% parallel efficiency.
Lei Xu 0023, Honghui Shang, Xin Chen 0023, Yunquan Zhang, Xingyu Gao 0003, Haifeng Song 0003
IEEE Trans. Parallel Distributed Syst.2
2022 Large-Scale Simulation of Quantum Computational Chemistry on a New Sunway Supercomputer
abstract
Quantum computational chemistry (QCC) is the use of quantum computers to solve problems in computational quantum chemistry. We develop a high performance variational quantum eigensolver (VQE) simulator for simulating quantum computational chemistry problems on a new Sunway supercomputer. The major innovations include: (1) a Matrix Product State (MPS) based VQE simulator to reduce the amount of memory needed and increase the simulation efficiency; (2) a combination of the Density Matrix Embedding Theory with the MPS-based VQE simulator to further extend the simulation range; (3) A three-level parallelization scheme to scale up to 20 million cores; (4) Usage of the Julia script language as the main programming language, which both makes the programming easier and enables cutting edge performance as native C or Fortran; (5) Study of real chemistry systems based on the VQE simulator, achieving nearly linearly strong and weak scaling. Our simulation demonstrates the power of VQE for large quantum chemistry systems, thus paves the way for large-scale VQE experiments on near-term quantum computers.
Honghui Shang, Li Shen 0001, Zhiqian Xu 0005, Chu Guo, Jie Liu 0069, Rongfen Lin, Yuling Yang, Zhuoya Wang, Yunquan Zhang
SC1
2022 Increasing the Efficiency of Massively Parallel Sparse Matrix-Matrix Multiplication in First-Principles Calculation on the New-Generation Sunway Supercomputer
abstract
The first-principles approach based on density-functional theory (DFT)/density-functional perturbation theory (DFPT) is widely used in calculations of the systems’ ground state energy, response properties (e.g., polarizability, phonon dispersions) and is playing an increasingly important role in chemistry, physics and materials science. For the large-scale calculations, the computation of the density matrix/response density matrix in DFT/DFPT has become the main performance bottleneck. One of the solutions is using the linear scaling method to get the density matrix and response density matrix. Here a massively parallel medium sparse matrix-matrix multiplication algorithm is designed for first-principle calculations and implemented on the new-generation Sunway supercomputer. Experiments show that the proposed method has obvious performance advantages compared to the original parallel version under moderate sparsity. The computing cores scale to 3,900,000 with strong scalability of 77.3$\%$.
Xin Chen 0023, Yingxiang Gao, Honghui Shang, Zhiqian Xu 0005, Xin Liu 0081, Dexun Chen
IEEE Trans. Parallel Distributed Syst.3
2022 Scaling Poisson Solvers on Many Cores via MMEwald
abstract
The Poisson solver for the calculation of the electrostatic potential is an essential primitive in quantum mechanics calculations. In this article, we adopt the Ewald method and propose a highly-optimized and scalable framework for Poisson solver, MMEwald, on the new generation Sunway supercomputer, capable of utilizing the collection of 390-core accelerators it uses. The MMEwald is based on a grid adapted cut-plane approach to partition the points into batches and distribute the batch to the processors. Furthermore, we propose a set of architecture-specific optimizations to efficiently utilize the memory bandwidth and computation capacity of the supercomputer. Experimental results demonstrate the efficiency of the MMEwald in providing strong and weak scaling performance.
Mingchuan Wu, Yangjun Wu, Honghui Shang, Ying Liu 0055, Huimin Cui, Xiaohui Duan, Yunquan Zhang, Xiaobing Feng 0002
IEEE Trans. Parallel Distributed Syst.3
2021 SW_Qsim: a minimize-memory quantum simulator with high-performance on a new Sunway supercomputer
abstract
Classical simulation of quantum computation plays a critical role in numerical studies of quantum algorithms and the validation of quantum devices. Here, we introduce SW_Qsim, a tensor-network-based quantum simulator, which is designed with a two-level parallel structure for efficient implementation on the many-core New Sunway Supercomputer. We propose a minimize-memory contraction path algorithm for rectangular quantum grids to reduce the memory overhead, and provide the memory-limited simulation capacity of SW26010pro. Moreover, tensor operations are carefully optimized on the SW processor to achieve high performance. We design a fault tolerance mechanism to improve the extreme-scale parallel stability. We benchmark SW_Qsim's simulation of RQCs up to 400-qubits, achieving near-linear strong and weak scaling with up to 28.75 million cores, far beyond the previous state of the art. Our work sheds light on the development of efficient quantum algorithms for use in the physical, chemical, and engineering science fields.
Xin Liu 0081, Pengpeng Zhao 0006, Yuling Yang, Honghui Shang, Weizhe Sun, Enming Dong, Dexun Chen
SC6
2021 TensorKMC: kinetic Monte Carlo simulation of 50 trillion atoms driven by deep learning on a new generation of Sunway supercomputer
abstract
The atomic kinetic Monte Carlo method plays an important role in multi-scale physical simulations because it bridges the micro and macro worlds. However, its accuracy is limited by empirical potentials. We therefore propose herein a triple-encoding algorithm and vacancy-cache mechanism to efficiently integrate ab initio neural network potentials (NNPs) with AKMC and implement them in our TensorKMC codes. We port our program to SW26010-pro and innovate a fast feature operator and a big fusion operator for the NNPs for fully utilizing the powerful heterogeneous computing units of the new-generation Sunway supercomputer. We further optimize memory usage. With these improvements, TensorKMC can simulate up to 54 trillions of atoms and achieve excellent strong and weak scaling performance up to 27,456,000 cores.
Honghui Shang, Xin Chen 0023, Xingyu Gao 0003, Rongfen Lin, Lei Xu 0023, Leilei Zhu, Fei Wang 0096, Yunquan Zhang, Haifeng Song 0003
SC1
2021 Accelerating all-electron ab initio simulation of raman spectra for biological systems
abstract
Raman spectroscopy provides chemical and compositional information that can serve as a structural fingerprint for various materials. Therefore, simulations of Raman spectra, including both quantum perturbation analyses and ground-state calculations are of significant interest. However, highly accurate full quantum mechanical (QM) simulations of Raman spectra have previously been confined to small systems. For large systems such as biological materials, the computational cost of full QM simulations is extremely high, and their extension to such systems remains challenging. In the work described here, by employing robust new algorithms and advances in implementation for the many-core architectures, we are able to perform fast, accurate, and massively parallel full ab initio simulations of the Raman spectra of biological systems with excellent strong and weak scaling, thereby providing a starting point for applying QM approaches to structural studies of such systems.
Honghui Shang, Yunquan Zhang, Ying Liu 0055, Mingchuan Wu, Yangjun Wu, Di Wei, Huimin Cui, Xin Liu 0081, Fei Wang 0096, Yuxi Ye, Yingxiang Gao, Shuang Ni, Xin Chen 0023, Dexun Chen
SC1
2021 Extreme-scale ab initio quantum raman spectra simulations on the leadership HPC system in China
abstract
Raman spectroscopy provides chemical and compositional information that can serve as a structural fingerprint for various materials. Therefore, simulations of Raman spectra, including both quantum perturbation analyses and ground-state calculations, are of significant interest. However, highly accurate full quantum mechanical (QM) simulations of Raman spectra have previously been confined to small systems. For large systems such as biological materials, full QM simulations have an extremely high computational cost and remain challenging. In this work, robust new algorithms and advanced implementations on many-core architectures are employed to enable fast, accurate, and massively parallel full ab initio simulations of the Raman spectra of realistic biological systems containing up to 3006 atoms, with excellent strong and weak scaling. Up to a performance of 468.5 PFLOP/s in double-precision and 813.7 PLOPS/s in mixed-half precision is achieved on the new-generation Sunway high-performance computing system, suggesting the potential for new applications of the QM approach to biological systems.
Honghui Shang, Yunquan Zhang, You Fu, Yingxiang Gao, Yangjun Wu, Xiaohui Duan, Rongfen Lin, Xin Liu 0081, Ying Liu 0055, Dexun Chen
SC1
2019 OpenKMC: a KMC design for hundred-billion-atom simulation using millions of cores on Sunway Taihulight
abstract
With more attention attached to nuclear energy, the formation mechanism of the solute clusters precipitation within complex alloys becomes intriguing research in the embrittlement of nuclear reactor pressure vessel (RPV) steels. Such phenomenon can be simulated with atomic kinetic Monte Carlo (AKMC) software, which evaluates the interactions of solute atoms with point defects in metal alloys. In this paper, we propose OpenKMC to accelerate large-scale KMC simulations on Sunway many-core architecture. To overcome the constraints caused by complex many-core architecture, we employ six levels of optimization in OpenKMC: (1) a new efficient potential computation model; (2) a group reaction strategy for fast event selection; (3) a software cache strategy; (4) combined communication optimizations; (5) a Transcription-Translation-Transmission algorithm for many-core optimization; (6) vectorization acceleration. Experiments illustrate that our OpenKMC has high accuracy and good scalability of applying hundred-billion-atom simulation over 5.2 million cores with a performance of over 80.1% parallel efficiency.
Kun Li 0016, Honghui Shang, Yunquan Zhang, Shigang Li 0002, Baodong Wu, Dexun Chen, Zhiqiang Wei 0004
SC2