EDBT 2026 Demo / reviewers in the wild / expert
He Zhang 0005
dblp:24/2058-5
· DBLP profile ↗
8ranked-venue papers
0as first author
3since 2021 · last 2023
0000-0001-7222-9072ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Heterogeneous acceleration algorithms for shallow cumulus convection scheme over GPU clusters
Fei Li 0042, Jinrong Jiang, He Zhang 0005, Xuebin Chi |
Future Gener. Comput. Syst. | 4 |
| 2023 | AGCM-3DLF: Accelerating Atmospheric General Circulation Model via 3-D Parallelization and Leap-FormatabstractThe atmospheric general circulation model (AGCM) has been an important research tool in the study of climate change for decades. As the demand for high-resolution simulation is becoming urgent, the scalability and simulation efficiency is faced with great challenges, especially for the latitude-longitude mesh-based models. In this paper, we propose a highly scalable 3-D atmospheric general circulation model based on leap-format, namely AGCM-3DLF. First, it utilizes a 3-D decomposition method allowing for parallelism release in all three physical dimensions. Then the leap-format difference computation scheme is adopted to maintain computational stability in grid updating and avoid additional filtering at the high latitudes. A novel shifting window communication algorithm is designed for parallelization of the unified model. Furthermore, a series of optimizations are conducted to improve the effectiveness of large-scale simulations. Experiment results in different platforms demonstrate good efficiency and scalability of the model. AGCM-3DLF scales up to the entire CAS-Xiandao1 supercomputer (196,608 CPU cores), attaining the speed of 11.1 simulation-year-per-day (SYPD) at a high resolution of 25KM. In addition, simulations conducted on the Sunway TaihuLight supercomputer exhibit a 1.06 million cores scalability with 36.1% parallel efficiency. He Zhang 0005, Yunquan Zhang, Baodong Wu, Kun Li 0016, Shigang Li 0002, Pengqi Lu, Junmin Xiao |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2022 | CC-RRTMG_SW++: Further optimizing a shortwave radiative transfer scheme on GPU
Fei Li 0042, Xiaohui Ji, Jinrong Jiang, Xiaoyong Tang, He Zhang 0005 |
J. Supercomput. | 7 |
| 2020 | A Highly Efficient Dynamical Core of Atmospheric General Circulation Model based on Leap-FormatabstractThe finite-difference dynamical core based on the equal-interval latitude-longitude mesh has been widely used for numerical simulations of the Atmospheric General Circulation Model (AGCM). Previous work utilizes different filtering schemes to alleviate the instability problem incurred by the unequal physical spacing at different latitudes, but they all incur high communication and computation overhead and become a scaling bottleneck. This paper proposes a new leap-format finite-difference computing scheme. It generalizes the usual finite-difference format with adaptive wider intervals and is able to maintain the computational stability in the grid updating. Therefore, the costly filtering scheme is eliminated. The new scheme is parallelized with a shifting communication method and implemented with fine communication optimizations based on a 3D decomposition. With the proposed leap-format computation scheme, the communication overhead of the AGCM is significantly reduced and good load balance is exhibited. The simulation results verify the correctness of the new leap-format scheme. The new scheme achieves the speed of 16.6 simulation-year-per-day (SYPD) and up to 3.3x speedup over the latest implementation. He Zhang 0005, Baodong Wu, Shigang Li 0002, Pengqi Lu, Yunquan Zhang, Yongjun Xu 0001 |
IPDPS | 3 |
| 2018 | AGCM3D: A Highly Scalable Finite-Difference Dynamical Core of Atmospheric General Circulation Model Based on 3D DecompositionabstractIt is commonly recognized that the dynamical core of the atmospheric model based on latitude-longitude mesh has poor parallel scalability, since it has to perform the costly polar or high-latitude filtering to dump out the unwanted modes. To parallelize the algorithm, only two dimensions can be partitioned even for a 3-dimensional mesh because of the costly filtering, which hinders the scalability of the algorithm. In this paper, we develop a highly scalable finite-difference dynamical core based on the latitude-longitude mesh using a 3D decomposition method, named as AGCM3D. Different from the traditional methods, our method releases the parallelism in all three dimensions, namely latitude, longitude, and level. To replace the costly Fast Fourier Transform (FFT) filtering, we propose a novel adaptive Gaussian filtering scheme, whose filtering strength increases as the latitude increases. Compared with the parallel FFT filtering, the parallel adaptive Gaussian filtering is far more efficient. In addition, we use the techniques of communication avoiding and message aggregation to further reduce the communication overhead. Experiments are conducted on Tianhe-2 supercomputer, and the resolution of the model is set as 0.5°x0.5°(50 km). Results show that our implementation scales up to 32,768 CPU cores in strong scaling and achieves the maximal simulation speed of 15.6 simulation-year-per-day (SYPD). Baodong Wu, Shigang Li 0002, Yunquan Zhang, He Zhang 0005, Junmin Xiao |
ICPADS | 5 |
| 2018 | Communication-Avoiding for Dynamical Core of Atmospheric General Circulation ModelabstractDynamical core is one of the most time-consuming parts in the global atmospheric general circulation model, which is widely used for the numerical simulation of the dynamic evolution process of global atmosphere. Due to its complicated calculation procedures and the non-uniformity of latitude-longitude mesh, the parallelization suffers from high communication overhead. In this paper, we deduce the operator form of the calculating flow in the dynamical core. Furthermore, it is abstracted out that the stencil and collection alternate action is the basic operation in the dynamic core. Based on the operator form of the calculation flow, we propose the corresponding optimization strategy for each operator. In the end, we develop a communication-avoiding algorithm to reduce communication overhead in the dynamic core. Our experiments show that the communication-avoiding algorithm reduces the total runtime by 54% at most for a 50 km resolution model running 10 years. Especially for communication reduction, the new algorithm achieves 1.4x speedup on average for the collective communication and 3.9x speedup on average for the communication involved in the stencil computation. Junmin Xiao, Shigang Li 0002, Baodong Wu, He Zhang 0005, Kun Li 0016, Erlin Yao, Yunquan Zhang, Guangming Tan |
ICPP | 4 |
| 2018 | An efficient parallel algorithm for the coupling of global climate models and regional climate models on a large-scale multi-core cluster
Jinrong Jiang, Junqiang Zhang, Juanxiong He, He Zhang 0005, Xuebin Chi, Tianxiang Yue |
J. Supercomput. | 5 |
| 2017 | A scalable parallel algorithm for atmospheric general circulation models on a multi-core cluster
Jinrong Jiang, He Zhang 0005, Lizhe Wang 0001, Rajiv Ranjan 0001, Albert Y. Zomaya |
Future Gener. Comput. Syst. | 3 |