EDBT 2026 Demo / reviewers in the wild / expert
Hao Luo 0015
dblp:14/3727-15
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2025
0009-0008-1038-944XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | StructILU: Dependency-Preserving Incomplete LU with Hierarchical Parallelism for Structured Grid PDEs on GPUsabstractThe Incomplete LU (ILU) computation is a crucial component for solving large-scale sparse linear systems arising from partial differential equations (PDEs), many of which are discretized on structured grids.However, due to inherent loop-carried data dependencies in ILU computation, implementing it on GPUs with massive computing units poses significant challenges.Existing methods either experience Hao Luo 0015, Qianchao Zhu, Xiaochen Hao, Chunxi Lei, Chengdi Ma, Yun Liang 0001, Chao Yang 0002 |
ICS | 1 |
| 2025 | Leonid: Exploring Automated Kernel Fusion in Performance-Portable Programming Models for Scientific ComputationabstractWith advances in hardware performance, architectural divergence and the growing gap between computational power and memory bandwidth have become increasingly pronounced.Existing performance-portable models address hardware divergence but lack automated kernel fusion to optimize memory-bound scientific applications.To address this issue, we propose Leonid, a performance-portable programming model designed to support automated kernel fusion.Leonid integrates separate modules for unified global and scratchpad memory management, and for unified parallel and serial execution patterns, both specifically tailored for automated kernel fusion, alongside an integrated automated kernel fusion module.These components ensure the compatibility across CPUs, GPUs, and Sunway platforms for automated kernel fusion.Performance evaluations demonstrate that Leonid achieves up to 1.52× speedup (averaging 1.19×) over manually implemented code, outperforms Kokkos and RAJA, and matches the efficiency of manually fused code in bandwidthlimited algorithms and applications.For bandwidth-limited and fusion-eligible code, Leonid offers a significant advantage over other performance-portable models that lack automated kernel fusion capabilities. Hao Luo 0015, Chao Yang 0002 |
ICS | 2 |
| 2025 | Telos: A Dataflow Accelerator for Sparse Triangular Solver of Partial Differential EquationsabstractPartial Differential Equations (PDEs) serve as the backbone of numerous scientific problems.Their solutions often rely on numerical methods, which transform these equations into large, sparse systems of linear equations.These systems, solved with iterative methods, exhibit structured sparsity patterns when derived from stencil-based numerical schemes.In preconditioned solvers, the sparse triangular solve procedure, SpTRSV, usually dominates the entire execution due to its loop-carried dependencies.Optimizing SpTRSV requires extracting parallelism from dependent computations.However, prior works have struggled to achieve both high parallelism and data locality, leading to suboptimal performance.We propose Telos, a dataflow accelerator for SpTRSV that exploits structured sparsity patterns in PDE solving.The dataflow execution leverages stencil patterns, efficiently utilizing pipeline parallelism to resolve data dependencies with minimal overhead.We tackle the challenge of complex data dependencies by proposing a plane-parallel pipelining technique that maps computations onto processing elements (PEs) while preserving data locality.A cross-plane communication aggregation technique is developed to streamline data transfers into a systolic manner.Our accelerator features effective pipelining of dependent computations and overlapping of computations with memory accesses.Experiments demonstrate that Telos delivers average speedups of 61×, 8×, and 11× over CPUs, GPUs, state-of-the-art accelerator, respectively. Xiaochen Hao, Hao Luo 0015, Chao Yang 0002, Yun Liang 0001 |
ISCA | 2 |
| 2021 | Enabling and scaling the HPCG benchmark on the newest generation Sunway supercomputer with 42 million heterogeneous coresabstractWe study and evaluate performance optimization techniques for the HPCG benchmark on the newest generation Sunway supercomputer. Specifically, a two-level blocking scheme is proposed to expose adequate parallelism in the symmetric Gauss-Seidel kernel while keeping a fast convergence rate, a fine-grained kernel fusion technique is developed to alleviate the bandwidth load on local storage with small capacity, and a low overhead thread collaboration method is presented to efficiently move data between threads and hide its cost with data transfer operations. Test results show that the optimized HPCG code is able to exploit 73.0% of the theoretical memory bandwidth, and scale to over 42 million heterogeneous cores with 95.5% weak-scaling efficiency and 5.91 Pflop/s performance. We also study how the performance can be improved if the specific rules of HPCG are not fully obeyed, and design dependency preserving parallelization and vectorization methods, further boosting performance to 27.6 Pflop/s. Qianchao Zhu, Hao Luo 0015, Chao Yang 0002, Mingshuo Ding, Wanwang Yin, Xinhui Yuan |
SC | 2 |