Zhaoqi Sun

dblp:376/6389 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2026
0009-0008-5947-7192ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
High-performance computing · 83% Parallel and multicore computing · 17%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing › scientific computing systems
cryo-EM 3D reconstruction
0.912025
Leveraging the Hardware Resources to Accelerate cryo-EM Reconstruction of RELION on the New Sunway Supercomputer · ACM Trans. Archit. Code Optim. 2025
High-performance computing
performance optimization at scale
0.912025
Leveraging the Hardware Resources to Accelerate cryo-EM Reconstruction of RELION on the New Sunway Supercomputer · ACM Trans. Archit. Code Optim. 2025
High-performance computing
scientific computing systems
0.912025
Leveraging the Hardware Resources to Accelerate cryo-EM Reconstruction of RELION on the New Sunway Supercomputer · ACM Trans. Archit. Code Optim. 2025
Parallel and multicore computing › parallelization strategies
multi-level parallelism
0.312025
Leveraging the Hardware Resources to Accelerate cryo-EM Reconstruction of RELION on the New Sunway Supercomputer · ACM Trans. Archit. Code Optim. 2025
Parallel and multicore computing
parallel programming models
0.312025
Leveraging the Hardware Resources to Accelerate cryo-EM Reconstruction of RELION on the New Sunway Supercomputer · ACM Trans. Archit. Code Optim. 2025

Methods — techniques the papers use, named apart from their topics

pipelining · 0.9operator optimization · 0.9multi-level parallelization · 0.9lock-free writing · 0.9
YearPublicationVenuePosition
2026 Redesign and accelerate the TIP4P water model on the new-generation sunway supercomputer
Puyu Xiong, Zhaoqi Sun, Mengyuan Hua, Jianqing Lan
J. Parallel Distributed Comput.4
2025 Leveraging the Hardware Resources to Accelerate cryo-EM Reconstruction of RELION on the New Sunway Supercomputer
abstract
The fast development of biomolecular structure determination has enabled the fine-grained study of objects in the micro-world, such as proteins and RNAs. The world is benefited. However, as the computational algorithms are constantly developed, the enrichment of features increases the algorithmic complexity and brings more computationally unfriendly modules. It calls for efficient solutions to leverage the rich and various hardware resources from the world’s most state-of-the-art supercomputing systems, and to fully accelerate the performance of the applications. In this article, we present our efforts on porting and optimizing the 3D reconstruction of RELION, one of the most popular cryo-EM software for biomolecular structure determinations, by leveraging different resources of the latest generation of Sunway heterogeneous supercomputer. Several novel approaches are proposed to resolve different challenges faced by the complex algorithm, including a multi-level parallel scheme and operator optimizations to smartly map and scale RELION, efficient strategies to largely address the memory bottlenecks and improve data locality, lock-free writing solutions to minimize write-write conflicts, and pipelining approaches to obtain excellent computation and communication overlap. Combining all proposed optimizations, the computation time is greatly reduced to under 2 hours, achieving 11.9× and 8.9× speedups on two different datasets. The overall design scales to 131,072 cores, increasing parallel efficiency from 33% to 61% and from 46% to 70%, respectively. To the best of our knowledge, this is the first work that fully optimized and scaled the 3D reconstruction of RELION using the latest Sunway system.
Jingle Xu, Jiayu Fu, Lin Gan 0001, Yaojian Chen, Zhaoqi Sun, Zhenchun Huang, Guangwen Yang 0002
ACM Trans. Archit. Code Optim.5
2024 A Low Overhead Heterogeneous Parallel Optimization Method Based on 3-D Elastic Wave Numerical Simulation
abstract
Applying the staggered grid finite difference method (SGFDM) for simulating acoustic responses in large-scale, complex, three-dimensional models poses substantial challenges in geophysics, especially in high-resolution stratigraphic model, due to the high computational burden. To address this problem, we have proposed a low overhead parallel optimization method (LOPOM) suitable for heterogeneous architectures, using high-resolution borehole models as examples. The LOPOM enhances computational intensity and optimizes memory bandwidth utilization. This is achieved through the technique of data reuse along the discontinuities of the model and by minimizing such discontinuities within the halo region. LOPOM was implemented and tested on the CUDA platform and the Sunway supercomputer. In each instance, LOPOM demonstrated an optimal acceleration ratio, thus proving its robust performance. Furthermore, the method’s effectiveness was confirmed through its application to fracture-vuggy formation models and digital core models. Finally, the numerical simulation results and the actual logging data are combined to illustrate the application value of LOPOM.
Zhuwen Wang, Zhaoqi Sun, Wubing Wan, Lin Gan 0008, Ruiyi Han, Yibo Wang 0002
IEEE Trans. Geosci. Remote. Sens.5