Wanlu Cao

dblp:378/9768 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Emerging computing paradigms · 61% Parallel and multicore computing · 30% High-performance computing · 9%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Emerging computing paradigms › quantum computing
neural network quantum states
0.912025
Fast and Scalable Neural Network Quantum States Method for Molecular Potential Energy Surfaces · IEEE Trans. Parallel Distributed Syst. 2025
Parallel and multicore computing › parallel computing › parallel machine learning
parallel training
0.912025
Fast and Scalable Neural Network Quantum States Method for Molecular Potential Energy Surfaces · IEEE Trans. Parallel Distributed Syst. 2025
Emerging computing paradigms
quantum computing
0.912025
Fast and Scalable Neural Network Quantum States Method for Molecular Potential Energy Surfaces · IEEE Trans. Parallel Distributed Syst. 2025
High-performance computing › scientific computing systems
quantum chemistry simulation
0.312025
Fast and Scalable Neural Network Quantum States Method for Molecular Potential Energy Surfaces · IEEE Trans. Parallel Distributed Syst. 2025

Methods — techniques the papers use, named apart from their topics

transformer · 0.9mixed precision training · 0.9KV-cache sharing · 0.9
YearPublicationVenuePosition
2025 Fast and Scalable Neural Network Quantum States Method for Molecular Potential Energy Surfaces
abstract
The Neural Network Quantum States (NNQS) method is highly promising for accurately solving the Schrödinger equation, yet it encounters challenges such as computational demands and slow rates of convergence. To address the high computational requirements, we introduce optimizations including a cross-sample KV cache sharing technique to enhance sampling efficiency, Quantum Bitwise and BloomHash methods for more efficient local energy computation, and mixed-precision training strategies to boost computational efficiency. To overcome the issue of slow convergence, we propose a parallel training algorithm for NNQS under second quantization to accelerate the training of base models for molecular potential surfaces. Our approach achieves up to 27-fold acceleration specifically in local energy calculations in systems with 154 spin orbitals and demonstrates strong and weak scaling efficiencies of 98% and 97%, respectively, on the H$_{2}$O$_{2}$potential surface training set. The parallelized implementation of transformer-based NNQS is highly portable on various high-performance computing architectures, offering new perspectives on quantum chemistry simulations.
Yangjun Wu, Wanlu Cao, Honghui Shang
IEEE Trans. Parallel Distributed Syst.2