EDBT 2026 Demo / reviewers in the wild / expert
Wubing Wan
dblp:316/3086
· DBLP profile ↗
7ranked-venue papers
1as first author
7since 2021 · last 2026
0009-0001-5174-1832ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploiting the Performance Potential of Extreme-Scale Earthquake Simulation: Achieving 86.7 PFLOPS With Over 39 Million CoresabstractLeveraging the latest Sunway supercomputer, we developed a fully optimized earthquake simulation model that accurately captures topographic effects for realistic seismic analysis. Optimizing for the SW26010Pro architecture with DMA/RMA communication mechanisms, data compression schemes, and vectorization, we achieved a speedup exceeding 160×. Our pipeline-based computation and communication overlapping scheme, combined with performance prediction models further minimized computational costs. These optimizations enabled the largest-scale curvilinear grid finite-difference method (CGFDM) earthquake simulations to date, covering 197 trillion grid points and achieving 86.7 PFLOPS on 39 million cores with a weak scaling efficiency of 97.9%. These advancements enabled the successful simulation of the 2008 Wenchuan earthquake, providing high-resolution seismic insights and robust assessments for regional hazard mitigation and disaster preparedness. Lin Gan 0008, Wubing Wan, Zekun Yin, Zhong He, Ping Gao 0005, Xiaohui Duan, Wei Xue 0003, Haohuan Fu, Guangwen Yang 0002 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2026 | SMEStencil: Optimizing High-Order Stencils on ARM Multicore Using SME UnitabstractMatrix-accelerated stencil computation is a hot research topic, yet its application to 3 dimensional (3D) high-order stencils and HPC remains underexplored. With the emergence of Scalable Matrix Extension(SME) on ARMv9-A CPU, we analyze SME-based accelerating strategies and tailor an optimal approach for 3D high-order stencils. We introduce algorithmic optimizations based on Scalable Vector Extension(SVE) and SME unit to address strided memory accesses, alignment conflicts, and redundant accesses. We propose memory optimizations to boost on-package memory efficiency, and a novel multi-thread parallelism paradigm to overcome data-sharing challenges caused by the absence of shared data caches. SMEStencil sustains consistently high hardware utilization across diverse stencil shapes and dimensions. Our DMA-based inter-NUMA communication further mitigates NUMA effects and MPI limitations in hybrid parallelism. Combining all the innovations, SMEStencil outperforms state-of-the-art libraries on Nividia A100 GPGPU by up to 2.1× . Moreover, the performance improvements enabled by our optimizations translate directly to real-world HPC applications and enable Reverse Time Migration(RTM) real-world applications to yield 1.8x speedup versus highly-optimized Nvidia A100 GPGPU version. Tianqi Mao 0003, Lin Gan 0008, Wubing Wan, Jiayu Fu, Lanke He, Zekun Yin, Wei Xue 0003, Guangwen Yang 0002 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2025 | Accelerating Half-Precision Seismic Simulation on Neural Processing UnitabstractDue to the superiority of handling irregular regions of interest, the curvilinear grid finite difference method (CGFDM) has become wildly used in seismic simulation for earthquake hazard evaluation and understanding of earthquake physics. This paper proposes a novel approach that optimizes a CGFDM solver on the Ascend, a cutting-edge Neural Processing(NPU) Unit using half-precision storage and mixed-precision arithmetic. The approach increases the data throughput and computing efficiency, enabling more effective seismic modeling. Furthermore, we propose an efficient matrix unit enabled 3D difference algorithm that employs matrix unit on NPU to accelerate the computation. By fully exploiting the capability of matrix unit and wide SIMD lane, our solver on Ascend achieves a speedup of 4.19 × over the performance of parallel solver on two AMD CPUs and has successfully simulated real-world Wenchuan earthquake. For the best of our knowledge, we are the first to conduct seismic simulations on NPU. Wubing Wan, Lin Gan 0008, Ping Gao 0005, Haohuan Fu, Wei Xue 0003, Guangwen Yang 0002 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2024 | A Low Overhead Heterogeneous Parallel Optimization Method Based on 3-D Elastic Wave Numerical SimulationabstractApplying the staggered grid finite difference method (SGFDM) for simulating acoustic responses in large-scale, complex, three-dimensional models poses substantial challenges in geophysics, especially in high-resolution stratigraphic model, due to the high computational burden. To address this problem, we have proposed a low overhead parallel optimization method (LOPOM) suitable for heterogeneous architectures, using high-resolution borehole models as examples. The LOPOM enhances computational intensity and optimizes memory bandwidth utilization. This is achieved through the technique of data reuse along the discontinuities of the model and by minimizing such discontinuities within the halo region. LOPOM was implemented and tested on the CUDA platform and the Sunway supercomputer. In each instance, LOPOM demonstrated an optimal acceleration ratio, thus proving its robust performance. Furthermore, the method’s effectiveness was confirmed through its application to fracture-vuggy formation models and digital core models. Finally, the numerical simulation results and the actual logging data are combined to illustrate the application value of LOPOM. Zhuwen Wang, Zhaoqi Sun, Wubing Wan, Lin Gan 0008, Ruiyi Han, Yibo Wang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | 69.7-PFlops Extreme Scale Earthquake Simulation with Crossing Multi-faults and Topography on SunwayabstractA high-scalable and fully optimized earthquake model is presented based on the latest Sunway supercomputer. Contributions include: 1) the curvilinear grid finite-difference method (CGFDM) and flexible model applying perfectly matched layer (PML) and enabling more accurate and realistic terrain descriptions; 2) a hybrid and non-uniform domain decomposition scheme that efficiently maps the model across different levels of the computing system; and 3) sophisticated optimizations that largely alleviate or even eliminate bottlenecks in memory, communication, etc., obtaining a speedup of over 140×. Combining all innovations, the design fully exploits the hardware potential of all aspects and enables us to perform the largest CGFDM-based earthquake simulation ever reported (69.7 PFlops using over 39 million cores). Based on our design, the Turkey earthquakes (February 6, 2023), and the Ridgecrest earthquake (July 4, 2019), are successfully simulated with a maximum resolution of 12-m. Precise hazard evaluations for the hazardous reduction of earthquake-stricken areas are also conducted. Wubing Wan, Lin Gan 0001, Zekun Yin, Haodong Tian, Mengyuan Hua, Shengye Xiang, Zhongqiu He, Ping Gao 0005, Xiaohui Duan, Wei Xue 0003, Haohuan Fu, Guangwen Yang 0002, Yaojian Chen, Xin Liu 0081, Wei Zhang 0321 |
SC | 1 |
| 2023 | Redesign and Accelerate the AIREBO Bond-Order Potential on the New Sunway SupercomputerabstractMolecular dynamics (MD) is one of the most crucial computer simulation methods for understanding real-world processes at the atomic level. Reactive potentials based on the bond order concept have the ability to model dynamic bond breaking and formation with close to quantum mechanical (QM) precision without actually requiring expensive QM calculations. In this article, we focus on the adaptive intermolecular reactive empirical bond-order (AIREBO) potential in LAMMPS for the simulation of carbon and hydrocarbon systems on the new Sunway supercomputer. To achieve scalable performance, we propose a parallel two-level building scheme and periodic buffering strategy for the tailored data design to explore data locality and data reuse. Furthermore, we design two optimized nearest-neighbor access algorithms: the redistribution of accumulated coefficients algorithm and the double-end search connectivity algorithm. Finally, we implement parallel force computation with an AoS data layout and hardware/software co-cache. In addition, we have designed a low-overhead atomic operation-based load balancing method and vectorization. The overall performance of AIREBO achieves a speedup of nearly$20\times$on a single core group (CG), and more than$5\times$and$4\times$over an Intel Xeon E5 2680 v3 core and an Intel Xeon Gold 6138 core, respectively. Compared with the Intel accelerator package in LAMMPS, our performance further achieves$3.0\times$of an Intel Xeon E5 2680 v3 core and is better than that of an Intel Xeon Gold 6138 core. We complete the validation of the results in no more than 20.5 hours on a single node with 2,000,000 running steps (i.e., 1 ns). Our experiments show that the simulation of 2,139,095,040 atoms on 798,720 ((1MPE+64CPEs) × 12,288 processes) cores exhibits a parallel efficiency of 88% under weak scaling. Ping Gao 0005, Xiaohui Duan, Bertil Schmidt, Wubing Wan, Jiaxu Guo, Wusheng Zhang, Lin Gan 0008, Haohuan Fu, Wei Xue 0003, Guangwen Yang 0002 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2022 | Enabling Large-Scale Simulation of CAM on the Sunway TaihuLight SupercomputerabstractThe Community Atmosphere Model (CAM) has been ported, redesigned, and scaled to the full system of the Sunway TaihuLight, and provides peta-scale climate modeling performance. Based on a novel domain decomposition method, we have fully optimized the complete model code by using both OpenACC refactoring and more aggressive and finer-grained Athread approaches. The Athread approach enables us to achieve exceptional memory control and usage, efficient vectorization, and sophisticated utilization of the thread-level communication mechanism. We have also further refined the load-balance behaviors towards ultra-large-scale numerical simulation. By combining all these novelties, we achieved a simulation speed of 7.2 and 25.6 simulation-year-per-day (SYPD) for global 25-km and 100-km resolution, respectively (1.2- to 2.2-fold improvements over previous efforts), and a sustainable double-precision performance of 3.3 PFlops for a 750-m global simulation when using 10075000 cores. Xiaohui Duan, Lin Gan 0001, Wubing Wan, Yuhu Chen, Jinzhe Yang, Wei Xue 0003, Haohuan Fu, Guangwen Yang 0002 |
IEEE Trans. Computers | 4 |