Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Huihai An

dblp:398/7052 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2026
0009-0009-4134-8277ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Parallel and multicore computing · 49% High-performance computing · 35% Cloud and datacenter computing · 12%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Environmental and earth informatics · 100%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing
parallel programming models
2.022026
Accelerating Molecular Dynamics Simulations on ARM Multi-Core Processors · IEEE Trans. Parallel Distributed Syst. 2026
SWGOMP: Extending OpenMP for Efficient Offloading on Sunway Heterogeneous Architecture · IEEE Trans. Parallel Distributed Syst. 2026
Cloud and datacenter computing
computation offloading
1.012026
SWGOMP: Extending OpenMP for Efficient Offloading on Sunway Heterogeneous Architecture · IEEE Trans. Parallel Distributed Syst. 2026
High-performance computing › scientific computing systems
molecular dynamics simulation
1.012026
Accelerating Molecular Dynamics Simulations on ARM Multi-Core Processors · IEEE Trans. Parallel Distributed Syst. 2026
Parallel and multicore computing › thread-level parallelism
multithreaded parallelization
1.012026
Accelerating Molecular Dynamics Simulations on ARM Multi-Core Processors · IEEE Trans. Parallel Distributed Syst. 2026
Parallel and multicore computing › parallel programming models › directive-based programming
OpenMP
1.012026
SWGOMP: Extending OpenMP for Efficient Offloading on Sunway Heterogeneous Architecture · IEEE Trans. Parallel Distributed Syst. 2026
High-performance computing
scientific computing systems
1.012026
Accelerating Molecular Dynamics Simulations on ARM Multi-Core Processors · IEEE Trans. Parallel Distributed Syst. 2026
Environmental and earth informatics › geoscience
earth system modeling
0.912025
An AI-Enhanced 1km-Resolution Seamless Global Weather and Climate Model to Achieve Year-Scale Simulation Speed using 34 Million Cores · PPoPP 2025
High-performance computing › large-scale simulation
climate and weather simulation
0.912025
An AI-Enhanced 1km-Resolution Seamless Global Weather and Climate Model to Achieve Year-Scale Simulation Speed using 34 Million Cores · PPoPP 2025
GPUs and heterogeneous computing
heterogeneous architecture
0.312026
SWGOMP: Extending OpenMP for Efficient Offloading on Sunway Heterogeneous Architecture · IEEE Trans. Parallel Distributed Syst. 2026

Methods — techniques the papers use, named apart from their topics

mixed-precision optimization · 1.7OpenMP parallelization · 1.7vectorization · 1.0neighbor list algorithm · 1.0compiler directive extension · 1.0
YearPublicationVenuePosition
2026 SWGOMP: Extending OpenMP for Efficient Offloading on Sunway Heterogeneous Architecture
Qixin Chang, Xiaohui Duan, Huihai An, Yi Zhang 0127, Haohuan Fu, Bin Yang 0043, Yilun Han, Dongqiang Huang, Xiting Ju, Haopeng Huang, Wei Xue 0003, Lin Gan 0008, Maoxue Yu, Jian Li 0069, Zhao Jing, Hailong Liu 0007, Lixin Wu, Ren Hu
IEEE Trans. Parallel Distributed Syst.3
2026 Accelerating Molecular Dynamics Simulations on ARM Multi-Core Processors
abstract
LAMMPS is a widely used molecular dynamics (MD) software package in materials science, computational chemistry, and biophysics, supporting parallel computing from a single CPU core to large supercomputers. The Kunpeng processor features both high memory bandwidth and core density and is therefore an interesting candidate for accelerating compute-intensive workloads. In this paper, we target the Kunpeng multi-core architecture and focus on optimizing LAMMPS for modern ARM-based platforms by using the Lennard-Jones (L-J) and Tersoff potentials as representative case studies. We investigate both common and specific optimization challenges, and present a comprehensive performance analysis addressing four key aspects: neighbor list algorithm design, force computation optimization, efficient vectorization, and multi-thread parallelization. Experimental results show that the optimized potentials achieve speedups of approximately$2 \times$and$5 \times$, reaching$4.55 \times$and$7.04\times$the performance of the original Intel version for L-J and Tersoff, respectively. Both potentials outperform Intel's acceleration library, with a peak performance up to$2.9\times$-$3.5\times$. In terms of parallel efficiency, we evaluate scalability both within a single CPU (small-scale) and across multiple nodes (large-scale). Strong and weak scaling tests within a single CPU show that when the expansion factor is 32 times, parallel efficiency remains above$90\%$. Large-scale weak scaling across multiple nodes achieves up to$86\%$efficiency when the expansion factor is 32. Using 32 nodes (18,432 processes), our implementation enables billion-atom simulations with L-J and Tersoff potentials. This work achieves breakthrough performance and provides critical support for large-scale molecular dynamics in engineering applications.
Huihai An, Zhihua Sa, Ping Gao 0005, Xiaohui Duan, Bertil Schmidt, Yizhen Chen, Lin Gan 0001, Guangwen Yang 0002
IEEE Trans. Parallel Distributed Syst.2
2025 An AI-Enhanced 1km-Resolution Seamless Global Weather and Climate Model to Achieve Year-Scale Simulation Speed using 34 Million Cores
abstract
Global Storm Resolving Models (GSRMs) is crucial for understanding extreme weather events under the climate change background. In this study, we optimize Global-Regional Integrated Forecast System (GRIST), which is a unified weather-climate modeling system designed for research and operation, for the next-generation Sunway supercomputer, incorporating AI-enhanced physics suite, OpenMP-based parallelization, and mixed-precision optimizations to enhance both efficiency and performance portability, as well as the unified modeling capability. Our experiments successfully capture significant events during the "23.7" extreme rainfall over northern China influenced by super Typhoon Doksuri, at 1km resolution. Notably, our work scales to 34 million cores, enabling simulation speeds at 491 SDPD (3km) and 181 SDPD (1km).
Xiaohui Duan, Yi Zhang 0127, Haohuan Fu, Bin Yang 0043, Yilun Han, Dongqiang Huang, Huihai An, Xiting Ju, Haopeng Huang, Wei Xue 0003, Jianye Hou, Maoxue Yu, Jian Li 0069, Zhao Jing, Hailong Liu 0007, Lixin Wu
PPoPP12