EDBT 2026 Demo / reviewers in the wild / expert
Shan Liang 0005
dblp:29/6005-5
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2026
0000-0001-6781-3565ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | WindStencil: Unleashing GPU Potential for High-Order Stencil Computation in High-Performance Inviscid CFD Simulations
Xiazhen Liu, Runfeng Jin, Jian Zhang 0070, Wu Yuan 0002, Shan Liang 0005, Zhonghua Lu |
ICS | 8 |
| 2026 | From Optimal Solutions to Reusable Rules: A Knowledge-Driven Stencil Scheduling Framework
Jian Zhang 0070, Xiazhen Liu, Wu Yuan 0002, Shan Liang 0005 |
KSEM (6) | 6 |
| 2025 | An Integrated Topology-Aware Mapping Scheme for Large-Scale CFD Applications on the ORISE Supercomputer
Lin Gao 0004, Wu Yuan 0002, Xiazhen Liu, Shan Liang 0005, Jian Zhang 0070, Zhonghua Lu |
ICA3PP (6) | 4 |
| 2025 | FlashMP: Fast Discrete Transform-Based Solver for Preconditioning Maxwell's Equations on GPUsabstractEfficiently solving large-scale linear systems is a critical challenge in electromagnetic simulations, particularly when using the Crank-Nicolson Finite-Difference Time-Domain method. Existing iterative solvers are commonly employed to handle the resulting sparse systems but suffer from slow convergence due to the ill-conditioned nature of the double-curl operator. Approximate preconditioners, like SOR and Incomplete LU decomposition (ILU) provide insufficient convergence, while direct solvers are impractical due to excessive memory requirements. To address this, we propose FlashMP, a novel preconditioning system that designs a subdomain exact solver based on discrete transforms. FlashMP provides an efficient GPU implementation that achieves multi-GPU scalability through domain decomposition. Evaluations on AMD MI60 GPU clusters (up to$\mathbf{1 0 0 0 ~ G P U s}$) show that FlashMP reduces iteration counts by up to$16 \times$and achieves speedups of$2.5 \times$to$4.9 \times$compared to baseline implementations in state-of-the-art libraries Hypre. Weak scalability tests show parallel efficiencies up to 84.1 %. Yaqian Gao, Runfeng Jin, Yidong Chen 0014, Wu Yuan 0002, Wenpeng Ma, Shan Liang 0005, Jian Zhang 0070, Zhonghua Lu |
ICCD | 10 |