EDBT 2026 Demo / reviewers in the wild / expert
Huihai An
dblp:398/7052
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2026
0009-0009-4134-8277ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Parallel and multicore computing · 49% High-performance computing · 35% Cloud and datacenter computing · 12% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Environmental and earth informatics · 100% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing
parallel programming models |
2.0 | 2 | 2026 | Accelerating Molecular Dynamics Simulations on ARM Multi-Core Processors · IEEE Trans. Parallel Distributed Syst. 2026 SWGOMP: Extending OpenMP for Efficient Offloading on Sunway Heterogeneous Architecture · IEEE Trans. Parallel Distributed Syst. 2026 |
Cloud and datacenter computing
computation offloading |
1.0 | 1 | 2026 | SWGOMP: Extending OpenMP for Efficient Offloading on Sunway Heterogeneous Architecture · IEEE Trans. Parallel Distributed Syst. 2026 |
High-performance computing › scientific computing systems
molecular dynamics simulation |
1.0 | 1 | 2026 | Accelerating Molecular Dynamics Simulations on ARM Multi-Core Processors · IEEE Trans. Parallel Distributed Syst. 2026 |
Parallel and multicore computing › thread-level parallelism
multithreaded parallelization |
1.0 | 1 | 2026 | Accelerating Molecular Dynamics Simulations on ARM Multi-Core Processors · IEEE Trans. Parallel Distributed Syst. 2026 |
Parallel and multicore computing › parallel programming models › directive-based programming
OpenMP |
1.0 | 1 | 2026 | SWGOMP: Extending OpenMP for Efficient Offloading on Sunway Heterogeneous Architecture · IEEE Trans. Parallel Distributed Syst. 2026 |
High-performance computing
scientific computing systems |
1.0 | 1 | 2026 | Accelerating Molecular Dynamics Simulations on ARM Multi-Core Processors · IEEE Trans. Parallel Distributed Syst. 2026 |
Environmental and earth informatics › geoscience
earth system modeling |
0.9 | 1 | 2025 | An AI-Enhanced 1km-Resolution Seamless Global Weather and Climate Model to Achieve Year-Scale Simulation Speed using 34 Million Cores · PPoPP 2025 |
High-performance computing › large-scale simulation
climate and weather simulation |
0.9 | 1 | 2025 | An AI-Enhanced 1km-Resolution Seamless Global Weather and Climate Model to Achieve Year-Scale Simulation Speed using 34 Million Cores · PPoPP 2025 |
GPUs and heterogeneous computing
heterogeneous architecture |
0.3 | 1 | 2026 | SWGOMP: Extending OpenMP for Efficient Offloading on Sunway Heterogeneous Architecture · IEEE Trans. Parallel Distributed Syst. 2026 |
Methods — techniques the papers use, named apart from their topics
mixed-precision optimization · 1.7OpenMP parallelization · 1.7vectorization · 1.0neighbor list algorithm · 1.0compiler directive extension · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SWGOMP: Extending OpenMP for Efficient Offloading on Sunway Heterogeneous Architecture
Qixin Chang, Xiaohui Duan, Huihai An, Yi Zhang 0127, Haohuan Fu, Bin Yang 0043, Yilun Han, Dongqiang Huang, Xiting Ju, Haopeng Huang, Wei Xue 0003, Lin Gan 0008, Maoxue Yu, Jian Li 0069, Zhao Jing, Hailong Liu 0007, Lixin Wu, Ren Hu |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2026 | Accelerating Molecular Dynamics Simulations on ARM Multi-Core ProcessorsabstractLAMMPS is a widely used molecular dynamics (MD) software package in materials science, computational chemistry, and biophysics, supporting parallel computing from a single CPU core to large supercomputers. The Kunpeng processor features both high memory bandwidth and core density and is therefore an interesting candidate for accelerating compute-intensive workloads. In this paper, we target the Kunpeng multi-core architecture and focus on optimizing LAMMPS for modern ARM-based platforms by using the Lennard-Jones (L-J) and Tersoff potentials as representative case studies. We investigate both common and specific optimization challenges, and present a comprehensive performance analysis addressing four key aspects: neighbor list algorithm design, force computation optimization, efficient vectorization, and multi-thread parallelization. Experimental results show that the optimized potentials achieve speedups of approximately$2 \times$and$5 \times$, reaching$4.55 \times$and$7.04\times$the performance of the original Intel version for L-J and Tersoff, respectively. Both potentials outperform Intel's acceleration library, with a peak performance up to$2.9\times$-$3.5\times$. In terms of parallel efficiency, we evaluate scalability both within a single CPU (small-scale) and across multiple nodes (large-scale). Strong and weak scaling tests within a single CPU show that when the expansion factor is 32 times, parallel efficiency remains above$90\%$. Large-scale weak scaling across multiple nodes achieves up to$86\%$efficiency when the expansion factor is 32. Using 32 nodes (18,432 processes), our implementation enables billion-atom simulations with L-J and Tersoff potentials. This work achieves breakthrough performance and provides critical support for large-scale molecular dynamics in engineering applications. Huihai An, Zhihua Sa, Ping Gao 0005, Xiaohui Duan, Bertil Schmidt, Yizhen Chen, Lin Gan 0001, Guangwen Yang 0002 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2025 | An AI-Enhanced 1km-Resolution Seamless Global Weather and Climate Model to Achieve Year-Scale Simulation Speed using 34 Million CoresabstractGlobal Storm Resolving Models (GSRMs) is crucial for understanding extreme weather events under the climate change background. In this study, we optimize Global-Regional Integrated Forecast System (GRIST), which is a unified weather-climate modeling system designed for research and operation, for the next-generation Sunway supercomputer, incorporating AI-enhanced physics suite, OpenMP-based parallelization, and mixed-precision optimizations to enhance both efficiency and performance portability, as well as the unified modeling capability. Our experiments successfully capture significant events during the "23.7" extreme rainfall over northern China influenced by super Typhoon Doksuri, at 1km resolution. Notably, our work scales to 34 million cores, enabling simulation speeds at 491 SDPD (3km) and 181 SDPD (1km). Xiaohui Duan, Yi Zhang 0127, Haohuan Fu, Bin Yang 0043, Yilun Han, Dongqiang Huang, Huihai An, Xiting Ju, Haopeng Huang, Wei Xue 0003, Jianye Hou, Maoxue Yu, Jian Li 0069, Zhao Jing, Hailong Liu 0007, Lixin Wu |
PPoPP | 12 |