Yuhu Chen

dblp:295/0442 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2025
0000-0002-0228-3126ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
High-performance computing · 92% GPUs and heterogeneous computing · 8%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing › large-scale simulation
exascale simulation
0.912025
Kilometer-Scale AI-Powered and Performance-Portable Earth System Model (AP3ESM) to Achieve Year-Scale Simulation Speed on Heterogeneous Supercomputers · SC 2025
High-performance computing › performance engineering
performance portability
0.912025
Kilometer-Scale AI-Powered and Performance-Portable Earth System Model (AP3ESM) to Achieve Year-Scale Simulation Speed on Heterogeneous Supercomputers · SC 2025
High-performance computing › scientific computing systems
climate modeling
0.612022
Enabling Large-Scale Simulation of CAM on the Sunway TaihuLight Supercomputer · IEEE Trans. Computers 2022
High-performance computing
performance optimization at scale
0.612022
Enabling Large-Scale Simulation of CAM on the Sunway TaihuLight Supercomputer · IEEE Trans. Computers 2022
GPUs and heterogeneous computing
heterogeneous supercomputing
0.312025
Kilometer-Scale AI-Powered and Performance-Portable Earth System Model (AP3ESM) to Achieve Year-Scale Simulation Speed on Heterogeneous Supercomputers · SC 2025
High-performance computing › supercomputing
sunway taihulight
0.212022
Enabling Large-Scale Simulation of CAM on the Sunway TaihuLight Supercomputer · IEEE Trans. Computers 2022
High-performance computing
supercomputing
0.212022
Enabling Large-Scale Simulation of CAM on the Sunway TaihuLight Supercomputer · IEEE Trans. Computers 2022

Methods — techniques the papers use, named apart from their topics

mixed-precision computation · 0.9kokkos · 0.9OpenMP · 0.9AI-enhanced parameterization · 0.9vectorization · 0.6load balancing · 0.6domain decomposition · 0.6athread · 0.6OpenACC · 0.6
YearPublicationVenuePosition
2025 Kilometer-Scale AI-Powered and Performance-Portable Earth System Model (AP3ESM) to Achieve Year-Scale Simulation Speed on Heterogeneous Supercomputers
abstract
Kilometer-scale Earth system models (ESMs) necessitate exascale supercomputers to facilitate realistic simulations of weather phenomena and climate variability over a time span ranging from days to decades. We present AP3ESM, an ultra‑high‑resolution, AI‑Powered, Performance‑Portable ESM coupling atmosphere, land surface, ocean, and sea ice components. By leveraging the performance portability features of Kokkos and OpenMP, the AP3ESM operates efficiently on two heterogeneous systems while incurring minimal development overhead. Advanced optimization techniques, such as adaptive parallel algorithms, AI-enhanced physical parameterizations, and mixed-precision computations, have been implemented to further boost the computational efficiency. Breaking the 1-km resolution barrier, AP3ESM delivers 0.85 and 1.98 simulated-years-per-day (SYPD) for the standalone atmosphere and ocean components on 34.1 million Sunway cores and 16085 GPUs, respectively; the holistic AP3ESM achieves 0.54 SYPD on 37.2 million Sunway cores. Notably, the forecast experiment successfully captures Super Typhoon Doksuri in 2023 and its associated extreme rainfall across China.
Maoxue Yu, Yuhu Chen, Jiaying Song, Xiaohui Duan, Junwei Wei, Jiangfeng Yu, Hailong Liu 0007, Jinrong Jiang, Yi Zhang 0127, Pengfei Lin 0004, Weipeng Zheng, Jingwei Xie, Jiakang Zhang, Zilu Liu, Xiaoyu Jin, Jilin Wei, Qixin Chang, Qingxia Lin, Yanzhi Zhou, Wei Xue 0003, Haohuan Fu, Yue Yu 0001, Xuebin Chi, Lixin Wu
SC3
2024 swCUDA: Auto parallel code translation framework from CUDA to ATHREAD for new generation sunway supercomputer
abstract
Abstract Since specific hardware characteristics and low-level programming model are adapted to both NVIDIA GPU and new generation Sunway architecture, automatically translating mature CUDA kernels to Sunway ATHREAD kernels are realistic but challenging work. To address this issue, swCUDA, an auto parallel code translation framework is proposed. To that end, we create scale affine translation to transform CUDA thread hierarchy to Sunway index, directive based memory hierarchy and data redirection optimization to assign optimal memory usage and data stride strategy, directive based grouping-calculation-asynchronous-reduction (GCAR) algorithm to provide general solution for random access issue. swCUDA utilizes code generator ANTLR as compiler frontend to parse CUDA kernel and integrate novel algorithms in the node of abstracted syntax tree (AST) depending on directives. Automatically translation is performed on the entire Polybench suite and NBody simulation benchmark. We get an average 40x speedup compared with baseline on the Sunway architecture, average speedup of 15x compared to x86 CPU and average 27 percentage higher than NVIDIA GPU. Further, swCUDA is implemented to translate major kernels of the real world application Gromacs. The translated version achieves up to 17x speedup.
Maoxue Yu, Guanghao Ma, Zhuoya Wang, Yuhu Chen, Yucheng Wang 0002, Dongning Jia
CCF Trans. High Perform. Comput.5
2022 Enabling Large-Scale Simulation of CAM on the Sunway TaihuLight Supercomputer
abstract
The Community Atmosphere Model (CAM) has been ported, redesigned, and scaled to the full system of the Sunway TaihuLight, and provides peta-scale climate modeling performance. Based on a novel domain decomposition method, we have fully optimized the complete model code by using both OpenACC refactoring and more aggressive and finer-grained Athread approaches. The Athread approach enables us to achieve exceptional memory control and usage, efficient vectorization, and sophisticated utilization of the thread-level communication mechanism. We have also further refined the load-balance behaviors towards ultra-large-scale numerical simulation. By combining all these novelties, we achieved a simulation speed of 7.2 and 25.6 simulation-year-per-day (SYPD) for global 25-km and 100-km resolution, respectively (1.2- to 2.2-fold improvements over previous efforts), and a sustainable double-precision performance of 3.3 PFlops for a 750-m global simulation when using 10075000 cores.
Xiaohui Duan, Lin Gan 0001, Wubing Wan, Yuhu Chen, Jinzhe Yang, Wei Xue 0003, Haohuan Fu, Guangwen Yang 0002
IEEE Trans. Computers5