Lin Gan 0008

dblp:120/9592-8 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
9since 2021 · last 2026
0000-0003-3486-6016ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 SWGOMP: Extending OpenMP for Efficient Offloading on Sunway Heterogeneous Architecture
Qixin Chang, Xiaohui Duan, Huihai An, Yi Zhang 0127, Haohuan Fu, Bin Yang 0043, Yilun Han, Dongqiang Huang, Xiting Ju, Haopeng Huang, Wei Xue 0003, Lin Gan 0008, Maoxue Yu, Jian Li 0069, Zhao Jing, Hailong Liu 0007, Lixin Wu, Ren Hu
IEEE Trans. Parallel Distributed Syst.19
2026 Exploiting the Performance Potential of Extreme-Scale Earthquake Simulation: Achieving 86.7 PFLOPS With Over 39 Million Cores
abstract
Leveraging the latest Sunway supercomputer, we developed a fully optimized earthquake simulation model that accurately captures topographic effects for realistic seismic analysis. Optimizing for the SW26010Pro architecture with DMA/RMA communication mechanisms, data compression schemes, and vectorization, we achieved a speedup exceeding 160×. Our pipeline-based computation and communication overlapping scheme, combined with performance prediction models further minimized computational costs. These optimizations enabled the largest-scale curvilinear grid finite-difference method (CGFDM) earthquake simulations to date, covering 197 trillion grid points and achieving 86.7 PFLOPS on 39 million cores with a weak scaling efficiency of 97.9%. These advancements enabled the successful simulation of the 2008 Wenchuan earthquake, providing high-resolution seismic insights and robust assessments for regional hazard mitigation and disaster preparedness.
Lin Gan 0008, Wubing Wan, Zekun Yin, Zhong He, Ping Gao 0005, Xiaohui Duan, Wei Xue 0003, Haohuan Fu, Guangwen Yang 0002
IEEE Trans. Parallel Distributed Syst.1
2026 SMEStencil: Optimizing High-Order Stencils on ARM Multicore Using SME Unit
abstract
Matrix-accelerated stencil computation is a hot research topic, yet its application to 3 dimensional (3D) high-order stencils and HPC remains underexplored. With the emergence of Scalable Matrix Extension(SME) on ARMv9-A CPU, we analyze SME-based accelerating strategies and tailor an optimal approach for 3D high-order stencils. We introduce algorithmic optimizations based on Scalable Vector Extension(SVE) and SME unit to address strided memory accesses, alignment conflicts, and redundant accesses. We propose memory optimizations to boost on-package memory efficiency, and a novel multi-thread parallelism paradigm to overcome data-sharing challenges caused by the absence of shared data caches. SMEStencil sustains consistently high hardware utilization across diverse stencil shapes and dimensions. Our DMA-based inter-NUMA communication further mitigates NUMA effects and MPI limitations in hybrid parallelism. Combining all the innovations, SMEStencil outperforms state-of-the-art libraries on Nividia A100 GPGPU by up to 2.1× . Moreover, the performance improvements enabled by our optimizations translate directly to real-world HPC applications and enable Reverse Time Migration(RTM) real-world applications to yield 1.8x speedup versus highly-optimized Nvidia A100 GPGPU version.
Tianqi Mao 0003, Lin Gan 0008, Wubing Wan, Jiayu Fu, Lanke He, Zekun Yin, Wei Xue 0003, Guangwen Yang 0002
IEEE Trans. Parallel Distributed Syst.3
2025 Accelerating Half-Precision Seismic Simulation on Neural Processing Unit
abstract
Due to the superiority of handling irregular regions of interest, the curvilinear grid finite difference method (CGFDM) has become wildly used in seismic simulation for earthquake hazard evaluation and understanding of earthquake physics. This paper proposes a novel approach that optimizes a CGFDM solver on the Ascend, a cutting-edge Neural Processing(NPU) Unit using half-precision storage and mixed-precision arithmetic. The approach increases the data throughput and computing efficiency, enabling more effective seismic modeling. Furthermore, we propose an efficient matrix unit enabled 3D difference algorithm that employs matrix unit on NPU to accelerate the computation. By fully exploiting the capability of matrix unit and wide SIMD lane, our solver on Ascend achieves a speedup of 4.19 × over the performance of parallel solver on two AMD CPUs and has successfully simulated real-world Wenchuan earthquake. For the best of our knowledge, we are the first to conduct seismic simulations on NPU.
Wubing Wan, Lin Gan 0008, Ping Gao 0005, Haohuan Fu, Wei Xue 0003, Guangwen Yang 0002
IEEE Trans. Parallel Distributed Syst.5
2024 C³DA: A Universal Domain Adaptation Method for Scene Classification From Remote Sensing Imagery
abstract
Various remote sensing applications have widely used domain adaptation (DA) methods. Since it does not need to add human interpretation in the target domain, it can be used in cross-region, multi-temporal, and multi-sensor application scenarios. In order to further optimize the design of the loss function and better address the challenges of DA in remote sensing, in this paper, we propose a new universal DA method named C3DA for scene recognition of remote sensing images. It has a comprehensive C3criterion for recognizing the "unknown" classes by innovatively fusing confidence, consistency, and certainty of samples to make our network training more efficient. We evaluate the performance of our proposed method based on six transfer tasks on three remote sensing datasets. The evaluation results show that our proposed method achieves an average H-score of 58.44%, significantly higher than other SOTA universal DA methods with an average improvement of 2.32~29.43%. Compared to the baseline ResNet-50, it achieves up to 19.92% improvement, demonstrating that the proposed method outperforms in the universal DA scenario. In the future, we also plan to expand the application of this method to more scenarios.
Jiaxu Guo, Yushan Lai, Jinxiao Zhang, Juepeng Zheng, Haohuan Fu, Lin Gan 0008, Liang Hu 0001, Gaochao Xu, Xilong Che
IEEE Geosci. Remote. Sens. Lett.6
2024 A Low Overhead Heterogeneous Parallel Optimization Method Based on 3-D Elastic Wave Numerical Simulation
abstract
Applying the staggered grid finite difference method (SGFDM) for simulating acoustic responses in large-scale, complex, three-dimensional models poses substantial challenges in geophysics, especially in high-resolution stratigraphic model, due to the high computational burden. To address this problem, we have proposed a low overhead parallel optimization method (LOPOM) suitable for heterogeneous architectures, using high-resolution borehole models as examples. The LOPOM enhances computational intensity and optimizes memory bandwidth utilization. This is achieved through the technique of data reuse along the discontinuities of the model and by minimizing such discontinuities within the halo region. LOPOM was implemented and tested on the CUDA platform and the Sunway supercomputer. In each instance, LOPOM demonstrated an optimal acceleration ratio, thus proving its robust performance. Furthermore, the method’s effectiveness was confirmed through its application to fracture-vuggy formation models and digital core models. Finally, the numerical simulation results and the actual logging data are combined to illustrate the application value of LOPOM.
Zhuwen Wang, Zhaoqi Sun, Wubing Wan, Lin Gan 0008, Ruiyi Han, Yibo Wang 0002
IEEE Trans. Geosci. Remote. Sens.7
2024 Acceleration of Multi-Body Molecular Dynamics With Customized Parallel Dataflow
abstract
FPGAs are drawing increasing attention in resolving molecular dynamics (MD) problems, and have already been applied in problems such as two-body potentials, force fields composed of these potentials, etc. Competitive performance is obtained compared with traditional counterparts such as CPUs and GPUs. However, as far as we know, FPGA solutions for more complex and real-world MD problems, such as multi-body potentials, are seldom to be seen. This work explores the prospects of state-of-the-art FPGAs in accelerating multi-body potential. An FPGA-based accelerator with customized parallel dataflow that features multi-body potential computation, motion update, and internode communication is designed. Major contributions include: (1) parallelization applied at different levels of the accelerator; (2) an optimized dataflow mixing atom-level pipeline and cell-level pipeline to achieve high throughput; (3) a mixed-precision method using different precision at different stages of simulations; and (4) a communication-efficient method for internode communication. Experiments show that, our single-node accelerator is over 2.7× faster than an 8-core CPU design, performing 20.501 ns/day on a 55,296-atom system for theTersoffsimulation. Regarding power efficiency, our accelerator is 28.9× higher than I7-11700 and 4.8× higher than RTX 3090 when running the same test case.
Quan Deng 0001, Qiang Liu 0011, Xiaohui Duan, Lin Gan 0008, Jinzhe Yang, Wenlai Zhao, Zhenxiang Zhang, Guiming Wu, Wayne Luk, Haohuan Fu, Guangwen Yang 0002
IEEE Trans. Parallel Distributed Syst.5
2023 GEO-WMS: an improved approach to geoscientific workflow management system on HPC
Jiaxu Guo, Yidan Xu, Haohuan Fu, Wei Xue 0003, Lin Gan 0008, Mengxuan Tan, Tingye Wu, Yutong Shen, Xianwei Wu, Liang Hu 0001, Xilong Che
CCF Trans. High Perform. Comput.5
2023 Redesign and Accelerate the AIREBO Bond-Order Potential on the New Sunway Supercomputer
abstract
Molecular dynamics (MD) is one of the most crucial computer simulation methods for understanding real-world processes at the atomic level. Reactive potentials based on the bond order concept have the ability to model dynamic bond breaking and formation with close to quantum mechanical (QM) precision without actually requiring expensive QM calculations. In this article, we focus on the adaptive intermolecular reactive empirical bond-order (AIREBO) potential in LAMMPS for the simulation of carbon and hydrocarbon systems on the new Sunway supercomputer. To achieve scalable performance, we propose a parallel two-level building scheme and periodic buffering strategy for the tailored data design to explore data locality and data reuse. Furthermore, we design two optimized nearest-neighbor access algorithms: the redistribution of accumulated coefficients algorithm and the double-end search connectivity algorithm. Finally, we implement parallel force computation with an AoS data layout and hardware/software co-cache. In addition, we have designed a low-overhead atomic operation-based load balancing method and vectorization. The overall performance of AIREBO achieves a speedup of nearly$20\times$on a single core group (CG), and more than$5\times$and$4\times$over an Intel Xeon E5 2680 v3 core and an Intel Xeon Gold 6138 core, respectively. Compared with the Intel accelerator package in LAMMPS, our performance further achieves$3.0\times$of an Intel Xeon E5 2680 v3 core and is better than that of an Intel Xeon Gold 6138 core. We complete the validation of the results in no more than 20.5 hours on a single node with 2,000,000 running steps (i.e., 1 ns). Our experiments show that the simulation of 2,139,095,040 atoms on 798,720 ((1MPE+64CPEs) × 12,288 processes) cores exhibits a parallel efficiency of 88% under weak scaling.
Ping Gao 0005, Xiaohui Duan, Bertil Schmidt, Wubing Wan, Jiaxu Guo, Wusheng Zhang, Lin Gan 0008, Haohuan Fu, Wei Xue 0003, Guangwen Yang 0002
IEEE Trans. Parallel Distributed Syst.7