Maoxue Yu

dblp:384/1136 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0002-7253-4947ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 1 first-author · 7 since 2021
YearPublicationVenuePosition
2026 swKokkos: An Athread Backend for Enhanced Kokkos with the Sunway Heterogeneous Architecture
Junlin Wei, Jinrong Jiang, Chen Li 0068, Yehong Zhang, Yue Yu 0001, Lian Zhao, Zhenjia Li, Feng Zhang 0048, Yidi Bai, Maoxue Yu, Hailong Liu 0007, Xuebin Chi
EuroSys13
2026 SWGOMP: Extending OpenMP for Efficient Offloading on Sunway Heterogeneous Architecture
Qixin Chang, Xiaohui Duan, Huihai An, Yi Zhang 0127, Haohuan Fu, Bin Yang 0043, Yilun Han, Dongqiang Huang, Xiting Ju, Haopeng Huang, Wei Xue 0003, Lin Gan 0008, Maoxue Yu, Jian Li 0069, Zhao Jing, Hailong Liu 0007, Lixin Wu, Ren Hu
IEEE Trans. Parallel Distributed Syst.20
2025 An AI-Enhanced 1km-Resolution Seamless Global Weather and Climate Model to Achieve Year-Scale Simulation Speed using 34 Million Cores
abstract
Global Storm Resolving Models (GSRMs) is crucial for understanding extreme weather events under the climate change background. In this study, we optimize Global-Regional Integrated Forecast System (GRIST), which is a unified weather-climate modeling system designed for research and operation, for the next-generation Sunway supercomputer, incorporating AI-enhanced physics suite, OpenMP-based parallelization, and mixed-precision optimizations to enhance both efficiency and performance portability, as well as the unified modeling capability. Our experiments successfully capture significant events during the "23.7" extreme rainfall over northern China influenced by super Typhoon Doksuri, at 1km resolution. Notably, our work scales to 34 million cores, enabling simulation speeds at 491 SDPD (3km) and 181 SDPD (1km).
Xiaohui Duan, Yi Zhang 0127, Haohuan Fu, Bin Yang 0043, Yilun Han, Dongqiang Huang, Huihai An, Xiting Ju, Haopeng Huang, Wei Xue 0003, Jianye Hou, Maoxue Yu, Jian Li 0069, Zhao Jing, Hailong Liu 0007, Lixin Wu
PPoPP20
2025 Kilometer-Scale AI-Powered and Performance-Portable Earth System Model (AP3ESM) to Achieve Year-Scale Simulation Speed on Heterogeneous Supercomputers
abstract
Kilometer-scale Earth system models (ESMs) necessitate exascale supercomputers to facilitate realistic simulations of weather phenomena and climate variability over a time span ranging from days to decades. We present AP3ESM, an ultra‑high‑resolution, AI‑Powered, Performance‑Portable ESM coupling atmosphere, land surface, ocean, and sea ice components. By leveraging the performance portability features of Kokkos and OpenMP, the AP3ESM operates efficiently on two heterogeneous systems while incurring minimal development overhead. Advanced optimization techniques, such as adaptive parallel algorithms, AI-enhanced physical parameterizations, and mixed-precision computations, have been implemented to further boost the computational efficiency. Breaking the 1-km resolution barrier, AP3ESM delivers 0.85 and 1.98 simulated-years-per-day (SYPD) for the standalone atmosphere and ocean components on 34.1 million Sunway cores and 16085 GPUs, respectively; the holistic AP3ESM achieves 0.54 SYPD on 37.2 million Sunway cores. Notably, the forecast experiment successfully captures Super Typhoon Doksuri in 2023 and its associated extreme rainfall across China.
Maoxue Yu, Yuhu Chen, Jiaying Song, Xiaohui Duan, Junwei Wei, Jiangfeng Yu, Hailong Liu 0007, Jinrong Jiang, Yi Zhang 0127, Pengfei Lin 0004, Weipeng Zheng, Jingwei Xie, Jiakang Zhang, Zilu Liu, Xiaoyu Jin, Jilin Wei, Qixin Chang, Qingxia Lin, Yanzhi Zhou, Wei Xue 0003, Haohuan Fu, Yue Yu 0001, Xuebin Chi, Lixin Wu
SC2
2024 A Performance-Portable Kilometer-Scale Global Ocean Model on ORISE and New Sunway Heterogeneous Supercomputers
abstract
Ocean general circulation models (OGCMs) are indispensable for studying the multi-scale oceanic processes and climate change. High-resolution ocean simulations require immense computational power and thus become a challenge in climate science. We present LICOMK++, a performance-portable OGCM using Kokkos, to facilitate global kilometer-scale ocean simulations. The breakthroughs include: (1) we enhance cuttingedge Kokkos with the Sunway architecture, enabling LICOMK++ to become the first performance-portable OGCM on diversified architectures, i.e., Sunway processors, CUDA/HIP-based GPUs, and ARM CPUs. (2) LICOMK++ overcomes the one simulated-years-per-day (SYPD) performance challenge for global realistic OGCM at $1-\mathrm{km}$ resolution. It records $\mathbf{1. 0 5}$ and 1.70 SYPD with a parallel efficiency of 54.8% and 55.6% scaling on almost the entire new Sunway supercomputer and two-thirds of the ORISE supercomputer. (3) LICOMK++ is the first global 1-km-resolution realistic OGCM to generate scientific results. It successfully reproduces mesoscale and submesoscale structures that have considerable climate effects.
Junlin Wei, Jiangfeng Yu, Jinrong Jiang, Hailong Liu 0007, Pengfei Lin 0004, Maoxue Yu, Lian Zhao, Weipeng Zheng, Jingwei Xie, Yanzhi Zhou, Tao Zhang 0096, Feng Zhang 0048, Yehong Zhang, Yue Yu 0001, Yidi Bai, Chen Li 0068, Zipeng Yu, Xuebin Chi
SC7
2024 Heterogeneous many-core optimization for Monte Carlo path-tracing on new generation Sunway HPC system
abstract
Abstract We present swRender, a new parallel rendering pipeline based on the new Sunway many-core architecture (SW26010P) for the Monte Carlo path-tracing algorithm. Previous parallel rendering schemes are unsuitable for our task due to issues such as vast differences in hardware architectures and bottlenecks in I/O communication efficiency. To that end, we create a new two-level parallel tile rendering framework to fully utilize the Sunway computing resources, a practical tile-grouping load-balancing method to maintain the framework’s stability, and a novel many-core acceleration optimization to improve the rendering performance at the pixel level. Our method achieves (1) an average speedup of 16x in multiple benchmarks when compared to the baseline path-tracing model on the Sunway architecture, and (2) an average speedup of 2x when compared to state-of-the-art CPU, co-processor, and GPU-based parallel rendering approaches. Moreover, we scale swRender to run on 15 million cores and obtain high scalable parallel efficiency of 92%.
Xinjie Wang 0003, Guanghao Ma, Jiaying Song, Mingyao Geng, Wenhui Hu, Xi Duan, Xiaogang Jin 0001, Dexun Chen, Maoxue Yu
CCF Trans. High Perform. Comput.12
2024 swCUDA: Auto parallel code translation framework from CUDA to ATHREAD for new generation sunway supercomputer
abstract
Abstract Since specific hardware characteristics and low-level programming model are adapted to both NVIDIA GPU and new generation Sunway architecture, automatically translating mature CUDA kernels to Sunway ATHREAD kernels are realistic but challenging work. To address this issue, swCUDA, an auto parallel code translation framework is proposed. To that end, we create scale affine translation to transform CUDA thread hierarchy to Sunway index, directive based memory hierarchy and data redirection optimization to assign optimal memory usage and data stride strategy, directive based grouping-calculation-asynchronous-reduction (GCAR) algorithm to provide general solution for random access issue. swCUDA utilizes code generator ANTLR as compiler frontend to parse CUDA kernel and integrate novel algorithms in the node of abstracted syntax tree (AST) depending on directives. Automatically translation is performed on the entire Polybench suite and NBody simulation benchmark. We get an average 40x speedup compared with baseline on the Sunway architecture, average speedup of 15x compared to x86 CPU and average 27 percentage higher than NVIDIA GPU. Further, swCUDA is implemented to translate major kernels of the real world application Gromacs. The translated version achieves up to 17x speedup.
Maoxue Yu, Guanghao Ma, Zhuoya Wang, Yuhu Chen, Yucheng Wang 0002, Dongning Jia
CCF Trans. High Perform. Comput.1