Hongkun Yu 0002

dblp:147/2253-2 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
1since 2021 · last 2025
0000-0002-2591-6328ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
GPUs and heterogeneous computing · 44% High-performance computing · 44% Parallel and multicore computing · 13%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing › scientific computing systems
cryo-EM 3D reconstruction
0.412019
GPU-based 3D cryo-EM reconstruction with key-value streams: poster · PPoPP 2019
GPUs and heterogeneous computing
GPU computing
0.412019
GPU-based 3D cryo-EM reconstruction with key-value streams: poster · PPoPP 2019
Parallel and multicore computing › parallel algorithms
parallel algorithm design
0.112019
GPU-based 3D cryo-EM reconstruction with key-value streams: poster · PPoPP 2019

Methods — techniques the papers use, named apart from their topics

key-value streams · 0.4
YearPublicationVenuePosition
2025 Auto-Stencil: Performance-Driven Stencil Optimization with Hardware Feedback for LLMs
abstract
Stencil computation is an important computing pattern from numerous scientific simulations, and optimizing stencil for modern GPU architectures demands specialized expertise in both parallel programming and hardware-specific optimizations. However, traditional domain-specific languages based auto-tuning tools offer limited flexibility; generalized auto-parallelization tools offer limited performance; and large language models produce code that often fails to compile or underperforms. This paper presents Auto-Stencil, a novel framework that bridges this gap by integrating LLMs with hardware-aware reinforcement learning. Our approach combines a comprehensive stencil optimization dataset with a dual-objective training methodology that systematically aligns model outputs with both functional correctness and performance requirements. By incorporating execution feedback through a performance-driven reward model, Auto-Stencil generates highly optimized CUDA implementations that not only pass unit tests but also deliver exceptional performance across diverse stencil patterns. Experimental results demonstrate the superiority of the framework over state-of-the-art alternatives: 100% compilation accuracy, 100% optimization rate, and an average of 171 × speedup across test cases compared to a single CPU core. The results show Auto-Stencil is particularly suitable for large-scale HPC workflows where automation of complex, architecture-specific optimizations can significantly reduce development effort while maintaining excellent performance.
Quan Deng 0001, Lin Gan 0001, Hongkun Yu 0002, Wenlai Zhao, Guangwen Yang 0002
ICPP3
2019 Large-scale Parallel Design for Cryo-EM Structure Determination on Heterogeneous Many-core Architectures
abstract
Cryo-EM structure determination is the most important research area in structural biology. With the development of cryo-electron microscopy, the resolution has been enhanced significantly, which leads to the huge computation to reconstruct the biomolecule in recent years. In this paper, we present a large-scale parallel design for Cryo-EM structure determination on heterogeneous many-core architectures. A novel task parallel strategy is proposed to reduce the redundant computation and improve the scalability on large-scale systems. Further, We distribute the reconstruction model to each node and rearrange the data layout to achieve high parallel efficiency and reduce the memory footprint. The proposed comprehensive parallel design shows highly parallel efficiency and scalability on large-scale heterogeneous architectures, which could significantly accelerate the whole period of Cryo-EM structure determination process.
Hongkun Yu 0002, Ruixin Sun, Wenlai Zhao, Haohuan Fu, Guangwen Yang 0002
BIBM2
2019 Parallelizing cryo-EM 3D reconstruction on GPU cluster with a partitioned and streamed model
abstract
As a vital approach to determine the structure of biomacromolecules, high-resolution cryo-electron microscopy (cryo-EM) 3D reconstruction is extremely compute-intensive, and has gradually migrated to GPU accelerators in recent years. With certain kernels already achieving high speedup and efficiency on GPUs, the reconstruction part, which inherently requires accesses of a large 3D model in different orientations, brings tough challenges to GPU architectures and has no effective GPU-based options. To fill the above gap, in this paper, we propose Stream3D, a novel GPU-based parallel design for cryo-EM 3D reconstruction. Our major idea is to reorganize the related problem space as streams of key-value pairs, so that we can achieve both the flexibility and efficiency to compute and accumulate the contribution to the final 3D model from all different 2D image inputs. In addition, we design a hybrid communication mechanism to reduce intra-node communications and enable the solving process on a larger scale. With the addition of our GPU-based reconstruction design, we are able to improve the performance of the reconstruction part itself by 9.50 times, and the performance of the entire processing part (the reconstruction part and the other parts with mature GPU options) by 2.83 times. Moreover, Stream3D enables using the approach at a large scale, with 65.32-fold speedup when using up to 80 GPUs.
Shizhen Xu, Haohuan Fu, Hongkun Yu 0002, Wenlai Zhao, Guangwen Yang 0002
ICS4
2019 GPU-based 3D cryo-EM reconstruction with key-value streams: poster
abstract
The 3D reconstruction of cryo-electron microscopy (cryo-EM) structural determination process is highly compute-intensive. It inherently requires accesses of a large 3D model in different and variable orientations, brings tough challenges to GPU architecture and has no effective solutions currently. To fill this gap, we propose a novel GPU-based parallel design for cryo-EM 3D reconstruction. The major idea is to reorganize the related problem space as streams of key-value pairs, so that we can achieve both the flexibility and efficiency to compute and accumulate the contribution to the final 3D model from all different 2D image inputs. In addition, we design a hybrid communication mechanism to reduce intra-node communications and enable the solving process on a larger scale.
Shizhen Xu, Hongkun Yu 0002, Haohuan Fu, Guangwen Yang 0002
PPoPP3