Shan Liang 0005

dblp:29/6005-5 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2026
0000-0001-6781-3565ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 WindStencil: Unleashing GPU Potential for High-Order Stencil Computation in High-Performance Inviscid CFD Simulations
Xiazhen Liu, Runfeng Jin, Jian Zhang 0070, Wu Yuan 0002, Shan Liang 0005, Zhonghua Lu
ICS8
2026 From Optimal Solutions to Reusable Rules: A Knowledge-Driven Stencil Scheduling Framework
Jian Zhang 0070, Xiazhen Liu, Wu Yuan 0002, Shan Liang 0005
KSEM (6)6
2025 An Integrated Topology-Aware Mapping Scheme for Large-Scale CFD Applications on the ORISE Supercomputer
Lin Gao 0004, Wu Yuan 0002, Xiazhen Liu, Shan Liang 0005, Jian Zhang 0070, Zhonghua Lu
ICA3PP (6)4
2025 FlashMP: Fast Discrete Transform-Based Solver for Preconditioning Maxwell's Equations on GPUs
abstract
Efficiently solving large-scale linear systems is a critical challenge in electromagnetic simulations, particularly when using the Crank-Nicolson Finite-Difference Time-Domain method. Existing iterative solvers are commonly employed to handle the resulting sparse systems but suffer from slow convergence due to the ill-conditioned nature of the double-curl operator. Approximate preconditioners, like SOR and Incomplete LU decomposition (ILU) provide insufficient convergence, while direct solvers are impractical due to excessive memory requirements. To address this, we propose FlashMP, a novel preconditioning system that designs a subdomain exact solver based on discrete transforms. FlashMP provides an efficient GPU implementation that achieves multi-GPU scalability through domain decomposition. Evaluations on AMD MI60 GPU clusters (up to$\mathbf{1 0 0 0 ~ G P U s}$) show that FlashMP reduces iteration counts by up to$16 \times$and achieves speedups of$2.5 \times$to$4.9 \times$compared to baseline implementations in state-of-the-art libraries Hypre. Weak scalability tests show parallel efficiencies up to 84.1 %.
Yaqian Gao, Runfeng Jin, Yidong Chen 0014, Wu Yuan 0002, Wenpeng Ma, Shan Liang 0005, Jian Zhang 0070, Zhonghua Lu
ICCD10