Jinrong Jiang

dblp:45/7449 · DBLP profile ↗
← Back
22ranked-venue papers
0as first author
15since 2021 · last 2026
0000-0003-4463-8666ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 17 · 12 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 swKokkos: An Athread Backend for Enhanced Kokkos with the Sunway Heterogeneous Architecture
Junlin Wei, Jinrong Jiang, Chen Li 0068, Yehong Zhang, Yue Yu 0001, Lian Zhao, Zhenjia Li, Feng Zhang 0048, Yidi Bai, Maoxue Yu, Hailong Liu 0007, Xuebin Chi
EuroSys2
2026 HIP-DFPT: Scalable Optimization of Irregular Workloads in Quantum Perturbation on GPU Clusters
Meng Wan, Jue Wang 0013, Shunde Li, Honghui Shang, He Bai 0005, Peng Shi 0006, Yuchen Pang, Ying Liu 0055, Jinrong Jiang, Yangang Wang 0002, Xuebin Chi
IEEE Trans. Parallel Distributed Syst.10
2025 Kilometer-Scale AI-Powered and Performance-Portable Earth System Model (AP3ESM) to Achieve Year-Scale Simulation Speed on Heterogeneous Supercomputers
abstract
Kilometer-scale Earth system models (ESMs) necessitate exascale supercomputers to facilitate realistic simulations of weather phenomena and climate variability over a time span ranging from days to decades. We present AP3ESM, an ultra‑high‑resolution, AI‑Powered, Performance‑Portable ESM coupling atmosphere, land surface, ocean, and sea ice components. By leveraging the performance portability features of Kokkos and OpenMP, the AP3ESM operates efficiently on two heterogeneous systems while incurring minimal development overhead. Advanced optimization techniques, such as adaptive parallel algorithms, AI-enhanced physical parameterizations, and mixed-precision computations, have been implemented to further boost the computational efficiency. Breaking the 1-km resolution barrier, AP3ESM delivers 0.85 and 1.98 simulated-years-per-day (SYPD) for the standalone atmosphere and ocean components on 34.1 million Sunway cores and 16085 GPUs, respectively; the holistic AP3ESM achieves 0.54 SYPD on 37.2 million Sunway cores. Notably, the forecast experiment successfully captures Super Typhoon Doksuri in 2023 and its associated extreme rainfall across China.
Maoxue Yu, Yuhu Chen, Jiaying Song, Xiaohui Duan, Junwei Wei, Jiangfeng Yu, Hailong Liu 0007, Jinrong Jiang, Yi Zhang 0127, Pengfei Lin 0004, Weipeng Zheng, Jingwei Xie, Jiakang Zhang, Zilu Liu, Xiaoyu Jin, Jilin Wei, Qixin Chang, Qingxia Lin, Yanzhi Zhou, Wei Xue 0003, Haohuan Fu, Yue Yu 0001, Xuebin Chi, Lixin Wu
SC11
2025 A parallel algorithm for an Ocean General Circulation Model based on a unified dynamics framework
Xuebin Chi, Jinrong Jiang, Run Guo, Lian Zhao, Chen Li 0068, Yidi Bai, Junlin Wei, Guangqing Zhou
CCF Trans. High Perform. Comput.3
2025 HIP-RRTMG_SW: Accelerating a shortwave radiative transfer scheme under the heterogeneous-compute interface for portability (HIP) framework
Jinrong Jiang
J. Parallel Distributed Comput.4
2025 A Coupled Transformer-CNN Network: Advancing Sea Surface Temperature Forecast Accuracy
abstract
Sea surface temperature (SST) is critically important for understanding ocean dynamics and supporting various marine activities, making accurate short-term SST forecasting highly significant. However, accurately modeling the multi-scale variability of SST remains challenging for existing deep learning (DL) models. This study introduces the Coupled Transformer-CNN Network (CoTCN), a hybrid architecture designed to leverage the multi-scale variability of SST. The CoTCN combines the strengths of Transformers and convolutional neural networks (CNNs), significantly enhancing SST forecasts’ spatial continuity and predictive accuracy. Compared to five state-of-the-art DL models based on Transformer or CNN that include ConvLSTM, ConvGRU, AFNO, PredRNN, and SwinLSTM, CoTCN demonstrates superior performance in global and local areas of SST forecasting. At 1-day lead time, CoTCN reduces the global average root mean square error (RMSE) by over 15%, with forecast errors ranging from 0.20°C to 0.53°C across 1–10 day lead times. Moreover, the CoTCN effectively mitigates the checkerboard artifacts inherent to the Vision Transformer architecture. These findings highlight the effectiveness of CoTCN in capturing SST’s multi-scale features and underscore the promising potential of hybrid architectures for future DL models.
Tao Zhang 0096, Pengfei Lin 0004, Hailong Liu 0007, Weipeng Zheng, Jinrong Jiang, Lian Zhao
IEEE Trans. Geosci. Remote. Sens.9
2024 Generative Evolution Attacks Portfolio Selection
abstract
It is agreed that portfolio selection is of great importance for the financial market. Numerous outstanding exact and heuristic algorithms have been proposed in the past decades. However, their development always demands meticulous human ideas and could be time-consuming. Moreover, most of them tend to suffer from performance degradation when exposed to new portfolio selection models and different investment environments. Learning-enabled approaches have recently yielded impressive results, but these methods still grapple with challenges in model design and training. In this paper, we explore the mutual facilitation of large language models (LLMs) and huristic approaches in portfolio selection, and propose a novel LLM-based multi-objective evolutionary algorithm (MOEA) named IlmPC-NSGA-II. In this algorithm, the LLM with carefully-designed well-structured prompts serves as a straightforward yet effective engine for generating new solutions, non-dominated sorting and crowding distance calculation are adopted to enable the LLM and the evolutionary process to mutually guide toward the optimal region of the solution space. Experimental results on various scales of constrained multi-objective portfolio selection models and four benchmark problems demonstrate that our proposed approach can achieve a more competitive performance compared to widely-used MOEAs and the LLM-only method.
Chen Li 0068, Jinrong Jiang, Lian Zhao, Yidi Bai, Zhonghua Lu, Xuebin Chi
CEC3
2024 A Performance-Portable Kilometer-Scale Global Ocean Model on ORISE and New Sunway Heterogeneous Supercomputers
abstract
Ocean general circulation models (OGCMs) are indispensable for studying the multi-scale oceanic processes and climate change. High-resolution ocean simulations require immense computational power and thus become a challenge in climate science. We present LICOMK++, a performance-portable OGCM using Kokkos, to facilitate global kilometer-scale ocean simulations. The breakthroughs include: (1) we enhance cuttingedge Kokkos with the Sunway architecture, enabling LICOMK++ to become the first performance-portable OGCM on diversified architectures, i.e., Sunway processors, CUDA/HIP-based GPUs, and ARM CPUs. (2) LICOMK++ overcomes the one simulated-years-per-day (SYPD) performance challenge for global realistic OGCM at $1-\mathrm{km}$ resolution. It records $\mathbf{1. 0 5}$ and 1.70 SYPD with a parallel efficiency of 54.8% and 55.6% scaling on almost the entire new Sunway supercomputer and two-thirds of the ORISE supercomputer. (3) LICOMK++ is the first global 1-km-resolution realistic OGCM to generate scientific results. It successfully reproduces mesoscale and submesoscale structures that have considerable climate effects.
Junlin Wei, Jiangfeng Yu, Jinrong Jiang, Hailong Liu 0007, Pengfei Lin 0004, Maoxue Yu, Lian Zhao, Weipeng Zheng, Jingwei Xie, Yanzhi Zhou, Tao Zhang 0096, Feng Zhang 0048, Yehong Zhang, Yue Yu 0001, Yidi Bai, Chen Li 0068, Zipeng Yu, Xuebin Chi
SC4
2024 Accelerating LASG/IAP climate system ocean model version 3 for performance portability using Kokkos
Junlin Wei, Pengfei Lin 0004, Jinrong Jiang, Hailong Liu 0007, Lian Zhao, Yehong Zhang, Feng Zhang 0048, Youyun Li, Yue Yu 0001, Xuebin Chi
Future Gener. Comput. Syst.3
2023 Spatiotemporal networks for ENSO forecasting with LICOM3 and remote sensing data
Xuanying Zhang, Lianjing Wei, Jinrong Jiang, Pengfei Lin 0004, Hailong Liu 0007
Eng. Appl. Artif. Intell.4
2023 Heterogeneous acceleration algorithms for shallow cumulus convection scheme over GPU clusters
Fei Li 0042, Jinrong Jiang, He Zhang 0005, Xuebin Chi
Future Gener. Comput. Syst.3
2023 A GPU-enabled acceleration algorithm for the CAM5 cloud microphysics scheme
Xuanying Zhang, Jinrong Jiang
J. Supercomput.6
2023 LICOM3-CUDA: a GPU version of LASG/IAP climate system ocean model version 3 based on CUDA
Junlin Wei, Jinrong Jiang, Hailong Liu 0007, Feng Zhang 0048, Pengfei Lin 0004, Yongqiang Yu, Xuebin Chi, Lian Zhao, Mengrong Ding, Zipeng Yu, Weipeng Zheng
J. Supercomput.2
2022 CC-RRTMG_SW++: Further optimizing a shortwave radiative transfer scheme on GPU
Fei Li 0042, Xiaohui Ji, Jinrong Jiang, Xiaoyong Tang, He Zhang 0005
J. Supercomput.5
2021 GPUs-RRTMG_LW: high-efficient and scalable computing for a longwave radiative transfer model on multiple GPUs
Mingxin Guo, Jinrong Jiang
J. Supercomput.4
2019 DLENSO: A Deep Learning ENSO Forecasting Model
Dandan He, Pengfei Lin 0004, Hailong Liu 0007, Jinrong Jiang
PRICAI (2)5
2018 An efficient parallel algorithm for the coupling of global climate models and regional climate models on a large-scale multi-core cluster
Jinrong Jiang, Junqiang Zhang, Juanxiong He, He Zhang 0005, Xuebin Chi, Tianxiang Yue
J. Supercomput.2
2017 A scalable parallel algorithm for atmospheric general circulation models on a multi-core cluster
Jinrong Jiang, He Zhang 0005, Lizhe Wang 0001, Rajiv Ranjan 0001, Albert Y. Zomaya
Future Gener. Comput. Syst.2
2016 Using Supercomputer to Speed up Neural Network Training
abstract
Recent works in deep learning have shown that large models can dramatically improve performance. In this paper, we accelerated the deep network training using many GPUs. We have developed a framework based on Caffe called Caffe-HPC that can utilize computing clusters with multiple GPUs to train large models. Caffe[6] provides multimedia scientists and practitioners with a clean and modifiable framework for state-of-the-art deep learning algorithms and a collection of reference models. And Caffe-HPC retains all the features of the original Caffe, the model trained on original Caffe can be continue to trained on Caffe-HPC. It provides a convenient solution for people who are using Caffe and want to speed up the training. Using an Asynchronous Stochastic Gradient Descent optimizer, We made a good acceleration on training a CNN model on ILSVRC[5] 2012 dataset. And we have compared the convergence of different SGD algorithms. We believe our work will makes it possible to train larger networks on larger training sets in a reasonable amount of time.
Jinrong Jiang, Xuebin Chi
ICPADS2
2016 A distributed load balancing algorithm for climate big data processing over a multi-core CPU cluster
abstract
Summary Load imbalance is a common problem to be tackled urgently in large scale data‐driven simulation systems or data intensive computing. According to the coupler, the Chinese Academy of Sciences‐Earth System Model (CAS‐ESM) implements one‐way nesting of the Institute of Atmospheric Physics of Chinese Academy of Sciences Atmospheric General Circulation Model version 4.0 (IAP AGCM4.0) and Weather Research and Forecasting model (WRF). The METGRID (meteorological grid) and REAL program modules in the WRF are used to process meteorological data. In the CAS‐ESM, the load of the METGRID module is seriously unbalanced on many CPU cores. The load imbalance has a serious impact on the processing speed of meteorological data, so this study designs an optimization algorithm to solve the problem. Numerical experiments show that compared to before optimization, the optimization algorithm can solve the load imbalance of the METGRID, and the computation speed of the METGRID and REAL modules after optimization on 64 CPU cores is about 7.2 times faster than before. Meanwhile, the whole computation speed of the CAS‐ESM can improve by 217.53%. In addition, results indicate that they also can reach to a similar speedup on different numbers of CPU cores. Copyright © 2016 John Wiley & Sons, Ltd.
Jinrong Jiang, Huang Ye, Juanxiong He
Concurr. Comput. Pract. Exp.2
2015 Forecast Verification and Visualization based on Gaussian Mixture Model Co-estimation
abstract
Abstract Precipitation forecast verification is essential to the quality of a forecast. The Gaussian mixture model (GMM) can be used to approximate the precipitation of several rain bands and provide a concise view of the data, which is especially useful for comparing forecast and observation data. The robustness of such comparison mainly depends on the consistency of and the correspondence between the extracted rain bands in the forecast and observation data. We propose a novel co‐estimation approach based on GMM in which forecast and observation data are analysed simultaneously. This approach naturally increases the consistency of and correspondence between the extracted rain bands by exploiting the similarity between both forecast and observation data. Moreover, a novel visualization and exploration framework is implemented to help the meteorologists gain insight from the forecast. The proposed approach was applied to the forecast and observation data provided by the China Meteorological Administration. The results are evaluated by meteorologists and novel insight has been gained.
Yunhai Wang, Chaoran Fan, Jian Zhang 0070, Tao Niu, Song Zhang 0004, Jinrong Jiang
Comput. Graph. Forum6
2009 A Parallel Refined Block Arnoldi Algorithm for Large Unsymmetric Matrices
abstract
This paper proposed a parallel refined block Arnoldi method for computing a few eigenvalues with largest or smallest real parts. The method accelerated by Chebyshev iteration is also investigated. We report some numerical results and compare the parallel refined block methods with single vector counterparts. The results show that the proposed method is more efficient than single vector counterparts.
Xuebin Chi, Jinrong Jiang, Jun Liu 0059, Zhonghua Lu
HPCC3