EDBT 2026 Demo / reviewers in the wild / expert
Jiayu Fu
dblp:145/9723
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HierCut: Enabling 16-bit Format Mixed Precision for Molecular Dynamics through Hierarchical CutoffabstractMixed-precision methods offer the potential to achieve better performance while maintaining accuracy comparable to that of high-precision formats. However, the adoption of mixed precision—particularly with 16-bit formats—in scientific computing remains limited due to precision truncation. Lin Gan 0001, Xiaohui Duan, Zhengrui Li, Jiayu Fu, Guangzhao Li, Guangwen Yang 0002 |
PPoPP | 5 |
| 2026 | SMEStencil: Optimizing High-Order Stencils on ARM Multicore Using SME UnitabstractMatrix-accelerated stencil computation is a hot research topic, yet its application to 3 dimensional (3D) high-order stencils and HPC remains underexplored. With the emergence of Scalable Matrix Extension(SME) on ARMv9-A CPU, we analyze SME-based accelerating strategies and tailor an optimal approach for 3D high-order stencils. We introduce algorithmic optimizations based on Scalable Vector Extension(SVE) and SME unit to address strided memory accesses, alignment conflicts, and redundant accesses. We propose memory optimizations to boost on-package memory efficiency, and a novel multi-thread parallelism paradigm to overcome data-sharing challenges caused by the absence of shared data caches. SMEStencil sustains consistently high hardware utilization across diverse stencil shapes and dimensions. Our DMA-based inter-NUMA communication further mitigates NUMA effects and MPI limitations in hybrid parallelism. Combining all the innovations, SMEStencil outperforms state-of-the-art libraries on Nividia A100 GPGPU by up to 2.1× . Moreover, the performance improvements enabled by our optimizations translate directly to real-world HPC applications and enable Reverse Time Migration(RTM) real-world applications to yield 1.8x speedup versus highly-optimized Nvidia A100 GPGPU version. Tianqi Mao 0003, Lin Gan 0008, Wubing Wan, Jiayu Fu, Lanke He, Zekun Yin, Wei Xue 0003, Guangwen Yang 0002 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2025 | T2-RELION: Task Parallelism, Tensor Core Accelerated RELION for Cryo-EM 3D ReconstructionabstractCryo-electron microscopy (cryo-EM) is a key technique for structural biology, but its computational efficiency, particularly during 3D reconstruction, remains a bottleneck. We introduce T2-RELION, a highly optimized version of RELION for cryo-EM 3D reconstruction on CPU-GPU platforms. RELION is a widely used open-source package in the cryo-EM community. We identify and resolve key inefficiencies in RELION’s parallelization strategy and memory management by proposing task parallelism and a three-phase GPU memory management strategy. Furthermore, we leverage Tensor Cores to accelerate the hot-spot kernel for difference calculation, employing an advanced pipelining strategy to hide latency and enable thread-block-level data reuse. On a quad-A100 GPU machine, performance evaluations demonstrate that T2-RELION outperforms RELION 4.0. For the hot-spot kernel, our optimizations achieve 1.90-23.7 times speedup. For the whole application using CNG and Trpv1 datasets, we observe 3.86 times and 2.68 times speedups, respectively. Jiayu Fu, Jingle Xu, Lin Gan 0001, Tianqi Mao 0003, Zirong Shen, Xiaohui Duan, Wei Xue 0003, Guangwen Yang 0002 |
SC | 1 |
| 2025 | Leveraging the Hardware Resources to Accelerate cryo-EM Reconstruction of RELION on the New Sunway SupercomputerabstractThe fast development of biomolecular structure determination has enabled the fine-grained study of objects in the micro-world, such as proteins and RNAs. The world is benefited. However, as the computational algorithms are constantly developed, the enrichment of features increases the algorithmic complexity and brings more computationally unfriendly modules. It calls for efficient solutions to leverage the rich and various hardware resources from the world’s most state-of-the-art supercomputing systems, and to fully accelerate the performance of the applications. In this article, we present our efforts on porting and optimizing the 3D reconstruction of RELION, one of the most popular cryo-EM software for biomolecular structure determinations, by leveraging different resources of the latest generation of Sunway heterogeneous supercomputer. Several novel approaches are proposed to resolve different challenges faced by the complex algorithm, including a multi-level parallel scheme and operator optimizations to smartly map and scale RELION, efficient strategies to largely address the memory bottlenecks and improve data locality, lock-free writing solutions to minimize write-write conflicts, and pipelining approaches to obtain excellent computation and communication overlap. Combining all proposed optimizations, the computation time is greatly reduced to under 2 hours, achieving 11.9× and 8.9× speedups on two different datasets. The overall design scales to 131,072 cores, increasing parallel efficiency from 33% to 61% and from 46% to 70%, respectively. To the best of our knowledge, this is the first work that fully optimized and scaled the 3D reconstruction of RELION using the latest Sunway system. Jingle Xu, Jiayu Fu, Lin Gan 0001, Yaojian Chen, Zhaoqi Sun, Zhenchun Huang, Guangwen Yang 0002 |
ACM Trans. Archit. Code Optim. | 2 |
| 2025 | Ocean Surface Currents Measurement From GEO-LEO Bistatic Along-Track Interferometric SAR: Methods and OptimizationabstractSpaceborne Synthetic Aperture Radar (SAR) Along-Track Interferometry (ATI) serves as a primary approach to measure the Total Surface Current Vector (TSCV) of the ocean. However, single spaceborne SAR systems are constrained by limited observation perspectives, leading to challenges in measuring two-dimensional (2D) TSCV. To address the 2D TSCV inversion issue, a bistatic SAR that utilizes Geosynchronous SAR (GEO SAR) as the radiation source and Low Earth Orbit SARs (LEO SARs) as passive receivers is proposed. This bistatic SAR configuration outperforms existing bistatic SAR ATI systems in terms of cost efficiency as well as inter-satellite synchronization simplicity. For the GEO-LEO SAR system, we develop new interferometric signal models and processing methods to enable one-dimensional (1D) and 2D TSCV estimation. Secondly, to enhance measurement performance, a system configuration optimization algorithm that considered both imaging and ATI performances is applied. Optimization results demonstrate that a larger GEO SAR elevation angle improves estimation accuracy. Finally, based on the currently operating L-band GEO SAR system LSAR4-01, optimization and full-link ATI simulations are conducted, verifying the effectiveness of the optimization algorithm and highlighting the potential design of LEO SAR orbits for high accuracy ATI observations. Simulations achieve an imaging resolution of <100 m and a 2D TSCV inversion bias of 0.1 m/s. The influence of errors on simulation accuracy is also discussed in depth. Yuanhao Li 0001, Jiayu Fu, Zhiyang Chen 0001, Cheng Hu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |