Wenlai Zhao

dblp:153/0665 · DBLP profile ↗
← Back
21ranked-venue papers
4as first author
4since 2021 · last 2025
0000-0003-1036-4732ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 18 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Auto-Stencil: Performance-Driven Stencil Optimization with Hardware Feedback for LLMs
abstract
Stencil computation is an important computing pattern from numerous scientific simulations, and optimizing stencil for modern GPU architectures demands specialized expertise in both parallel programming and hardware-specific optimizations. However, traditional domain-specific languages based auto-tuning tools offer limited flexibility; generalized auto-parallelization tools offer limited performance; and large language models produce code that often fails to compile or underperforms. This paper presents Auto-Stencil, a novel framework that bridges this gap by integrating LLMs with hardware-aware reinforcement learning. Our approach combines a comprehensive stencil optimization dataset with a dual-objective training methodology that systematically aligns model outputs with both functional correctness and performance requirements. By incorporating execution feedback through a performance-driven reward model, Auto-Stencil generates highly optimized CUDA implementations that not only pass unit tests but also deliver exceptional performance across diverse stencil patterns. Experimental results demonstrate the superiority of the framework over state-of-the-art alternatives: 100% compilation accuracy, 100% optimization rate, and an average of 171 × speedup across test cases compared to a single CPU core. The results show Auto-Stencil is particularly suitable for large-scale HPC workflows where automation of complex, architecture-specific optimizations can significantly reduce development effort while maintaining excellent performance.
Quan Deng 0001, Lin Gan 0001, Hongkun Yu 0002, Wenlai Zhao, Guangwen Yang 0002
ICPP4
2024 A Joint Time-Frequency Domain Transformer for multivariate time series forecasting
Yushu Chen, Shengzhuo Liu, Jinzhe Yang, Wenlai Zhao, Guangwen Yang 0002
Neural Networks5
2024 Acceleration of Multi-Body Molecular Dynamics With Customized Parallel Dataflow
abstract
FPGAs are drawing increasing attention in resolving molecular dynamics (MD) problems, and have already been applied in problems such as two-body potentials, force fields composed of these potentials, etc. Competitive performance is obtained compared with traditional counterparts such as CPUs and GPUs. However, as far as we know, FPGA solutions for more complex and real-world MD problems, such as multi-body potentials, are seldom to be seen. This work explores the prospects of state-of-the-art FPGAs in accelerating multi-body potential. An FPGA-based accelerator with customized parallel dataflow that features multi-body potential computation, motion update, and internode communication is designed. Major contributions include: (1) parallelization applied at different levels of the accelerator; (2) an optimized dataflow mixing atom-level pipeline and cell-level pipeline to achieve high throughput; (3) a mixed-precision method using different precision at different stages of simulations; and (4) a communication-efficient method for internode communication. Experiments show that, our single-node accelerator is over 2.7× faster than an 8-core CPU design, performing 20.501 ns/day on a 55,296-atom system for theTersoffsimulation. Regarding power efficiency, our accelerator is 28.9× higher than I7-11700 and 4.8× higher than RTX 3090 when running the same test case.
Quan Deng 0001, Qiang Liu 0011, Xiaohui Duan, Lin Gan 0008, Jinzhe Yang, Wenlai Zhao, Zhenxiang Zhang, Guiming Wu, Wayne Luk, Haohuan Fu, Guangwen Yang 0002
IEEE Trans. Parallel Distributed Syst.7
2023 Automatic Deep Learning Operator Fusion on Sunway SW26010 Many-Core Processor
abstract
Deep learning networks (DNNs) have been growing rapidly in recent years, with increasing demands on computing power. Therefore, accelerating the execution of DNN models has become a research hotspot. Operator fusion is a critical optimization strategy to enhance DNN performance in Deep Learning (DL) frameworks, such as TensorFlow, Pytorch, TVM and Halide. However, these frameworks are designed for general optimization and cannot fully harness the specific features of emerging hardware. Moreover, they primarily implement operator fusion at the operator level, missing out on many fusion opportunities and heavily relying on extensive manual optimizations for fused operators. Targeting the Sunway SW26010 Many-Core processor, the basic building block of Sunway TaihuLight supercomputer, we introduce swAutoFuser, an end-to-end automatic operator fusion and code generation framework. swAutoFuser proposes a set of low-level primitives to leverage hardware features and employs an autofuser to achieve primitive level fusion, which breaks operator boundaries and enables more fusion opportunities. In addition, swAutoFuser can automatically generate high-performance fused operator implementations based on a static cost model, significantly reducing the overhead of manually optimizing fused operators. Our experiments demonstrate that swAutoFuser can improve operator performance by 10% to 56%.
Wenxiang Zhang, Wenzhao Wu, Yanjie Zhen, Wenlai Zhao, Guangwen Yang 0002
ICPADS5
2020 High performance reconfigurable computing for numerical simulation and deep learning
Lin Gan 0001, Jinzhe Yang, Wenlai Zhao, Wayne Luk, Guangwen Yang 0002
CCF Trans. High Perform. Comput.4
2020 Large-Scale Automatic K-Means Clustering for Heterogeneous Many-Core Supercomputer
abstract
This article presents an automatic k-means clustering solution targeting the Sunway TaihuLight supercomputer. We first introduce a multilevel parallel partition approach that not only partitions by dataflow and centroid, but also by dimension, which unlocks the potential of the hierarchical parallelism in the heterogeneous many-core processor and the system architecture of the supercomputer. The parallel design is able to process large-scale clustering problems with up to 196,608 dimensions and over 160,000 targeting centroids, while maintaining high performance and high scalability. Furthermore, we propose an automatic hyper-parameter determination process for k-means clustering, by automatically generating and executing the clustering tasks with a set of candidate hyper-parameter, and then determining the optimal hyper-parameter using a proposed evaluation method. The proposed autoclustering solution can not only achieve high performance and scalability for problems with massive high-dimensional data, but also support clustering without sufficient prior knowledge for the number of targeted clusters, which can potentially increase the scope of k-means algorithm to new application areas.
Wenlai Zhao, Pan Liu 0002, Vladimir Janjic, Xiaohan Yan, Shicai Wang, Haohuan Fu, Guangwen Yang 0002, John Thomson
IEEE Trans. Parallel Distributed Syst.2
2019 Large-scale Parallel Design for Cryo-EM Structure Determination on Heterogeneous Many-core Architectures
abstract
Cryo-EM structure determination is the most important research area in structural biology. With the development of cryo-electron microscopy, the resolution has been enhanced significantly, which leads to the huge computation to reconstruct the biomolecule in recent years. In this paper, we present a large-scale parallel design for Cryo-EM structure determination on heterogeneous many-core architectures. A novel task parallel strategy is proposed to reduce the redundant computation and improve the scalability on large-scale systems. Further, We distribute the reconstruction model to each node and rearrange the data layout to achieve high parallel efficiency and reduce the memory footprint. The proposed comprehensive parallel design shows highly parallel efficiency and scalability on large-scale heterogeneous architectures, which could significantly accelerate the whole period of Cryo-EM structure determination process.
Hongkun Yu 0002, Ruixin Sun, Wenlai Zhao, Haohuan Fu, Guangwen Yang 0002
BIBM5
2019 Severe Convective Weather Classification in Remote Sensing Images by Semantic Segmentation
ZhiLei Chai, Wenlai Zhao
ICANN (3)3
2019 swATOP: Automatically Optimizing Deep Learning Operators on SW26010 Many-Core Processor
abstract
Achieving an optimized mapping of Deep Learning (DL) operators to new hardware architectures is the key to building a scalable DL system. However, handcrafted optimization involves huge engineering efforts, due to the variety of DL operator implementations and complex programming skills. Targeting the innovative many-core processor SW26010 adopted by the 3rd fastest supercomputer Sunway TaihuLight, an end-to-end automated framework called swATOP is presented as a more practical solution for DL operator optimization. Arithmetic intensive DL operators are expressed into an auto-tuning-friendly form, which is based on tensorized primitives. By describing the algorithm of a DL operator using our domain specific language (DSL), swATOP is able to derive and produce an optimal implementation by separating hardware-dependent optimization and hardware-agnostic optimization. Hardware-dependent optimization is encapsulated in a set of tensorized primitives with sufficient utilization of the underlying hardware features. The hardware-agnostic optimization contains a scheduler, an intermediate representation (IR) optimizer, an auto-tuner, and a code generator. These modules cooperate to perform an automatic design space exploration, to apply a set of programming techniques, to discover a near-optimal solution, and to generate the executable code. Our experiments show that swATOP is able to bring significant performance improvement on DL operators in over 88% of cases, compared with the best-handcrafted optimization. Compared to a black-box autotuner, the tuning and code generation time can be reduced to minutes from days using swATOP.
Jiarui Fang, Wenlai Zhao, Jinzhe Yang, Long Wang 0014, Lin Gan 0001, Haohuan Fu, Guangwen Yang 0002
ICPP3
2019 Parallelizing cryo-EM 3D reconstruction on GPU cluster with a partitioned and streamed model
abstract
As a vital approach to determine the structure of biomacromolecules, high-resolution cryo-electron microscopy (cryo-EM) 3D reconstruction is extremely compute-intensive, and has gradually migrated to GPU accelerators in recent years. With certain kernels already achieving high speedup and efficiency on GPUs, the reconstruction part, which inherently requires accesses of a large 3D model in different orientations, brings tough challenges to GPU architectures and has no effective GPU-based options. To fill the above gap, in this paper, we propose Stream3D, a novel GPU-based parallel design for cryo-EM 3D reconstruction. Our major idea is to reorganize the related problem space as streams of key-value pairs, so that we can achieve both the flexibility and efficiency to compute and accumulate the contribution to the final 3D model from all different 2D image inputs. In addition, we design a hybrid communication mechanism to reduce intra-node communications and enable the solving process on a larger scale. With the addition of our GPU-based reconstruction design, we are able to improve the performance of the reconstruction part itself by 9.50 times, and the performance of the entire processing part (the reconstruction part and the other parts with mature GPU options) by 2.83 times. Moreover, Stream3D enables using the approach at a large scale, with 65.32-fold speedup when using up to 80 GPUs.
Shizhen Xu, Haohuan Fu, Hongkun Yu 0002, Wenlai Zhao, Guangwen Yang 0002
ICS5
2018 swCaffe: A Parallel Framework for Accelerating Deep Learning Applications on Sunway TaihuLight
abstract
This paper reports our efforts on swCaffe, a highly efficient parallel framework for accelerating deep neural networks (DNNs) training on Sunway TaihuLight, the current fastest supercomputer in the world that adopts a unique many-core heterogeneous architecture, with 40,960 SW26010 processors connected through a customized communication network.First, we point out some insightful principles to fully exploit the performance of the innovative many-core architecture.Second, we propose a set of optimization strategies for redesigning a variety of neural network layers based on Caffe.Third, we put forward a topology-aware parameter synchronization scheme to scale the synchronous Stochastic Gradient Descent (SGD) method to multiple processors efficiently.We evaluate our framework by training a variety of widely used neural networks with the ImageNet dataset.On a single node, swCaffe can achieve 23%˜119% overall performance compared with Caffe running on K40m GPU.As compared with the Caffe on CPU, swCaffe runs 3.04˜7.84xfaster on all the networks.Finally, we present the scalability of swCaffe for training of ResNet-50 and AlexNet on the scale of 1024 nodes.
Liandeng Li, Jiarui Fang, Haohuan Fu, Jinlei Jiang, Wenlai Zhao, Conghui He, Xin You 0001, Guangwen Yang 0002
CLUSTER5
2018 Large-scale hierarchical k-means for heterogeneous many-core supercomputers
Liandeng Li, Wenlai Zhao, Haohuan Fu, Guangwen Yang 0002, John Thomson
SC3
2018 Optimizing Convolutional Neural Networks on the Sunway TaihuLight Supercomputer
abstract
The Sunway TaihuLight supercomputer is powered by SW26010, a new 260-core processor designed with on-chip fusion of heterogeneous cores. In this article, we present our work on optimizing the training process of convolutional neural networks (CNNs) on the Sunway TaihuLight supercomputer. Specifically, a highly efficient library (swDNN) and a customized Caffe framework (swCaffe) are proposed. Architecture-oriented optimization methods targeting the many-core architecture of SW26010 are introduced and are able to achieve 48× speedup for the convolution routine in swDNN and 4× speedup for the complete training process of the VGG-16 network using swCaffe, compared to the unoptimized algorithm and framework. Compared to the cuDNN library and the Caffe framework based on the NVIDIA K40m GPU, the proposed swDNN library and swCaffe framework on SW26010 have nearly half the performance of K40m in single -precision and have 3.6× and 1.8× speedup over K40m in double precision, respectively.
Wenlai Zhao, Haohuan Fu, Jiarui Fang, Weijie Zheng 0001, Lin Gan 0001, Guangwen Yang 0002
ACM Trans. Archit. Code Optim.1
2017 swDNN: A Library for Accelerating Deep Learning Applications on Sunway TaihuLight
abstract
To explore the potential of training complex deep neural networks (DNNs) on other commercial chips rather than GPUs, we report our work on swDNN, which is a highly-efficient library for accelerating deep learning applications on the newly announced world-leading supercomputer, Sunway TaihuLight. Targeting SW26010 processor, we derive a performance model that guides us in the process of identifying the most suitable approach for mapping the convolutional neural networks (CNNs) onto the 260 cores within the chip. By performing a systematic optimization that explores major factors, such as organization of convolution loops, blocking techniques, register data communication schemes, as well as reordering strategies for the two pipelines of instructions, we manage to achieve a double-precision performance over 1.6 Tflops for the convolution kernel, achieving 54% of the theoretical peak. Compared with Tesla K40m with cuDNNv5, swDNN results in 1.91-9.75x performance speedup in an evaluation with over 100 parameter configurations.
Jiarui Fang, Haohuan Fu, Wenlai Zhao, Bingwei Chen, Weijie Zheng 0001, Guangwen Yang 0002
IPDPS3
2016 Unleashing the performance potential of CPU-GPU platforms for the 3D atmospheric Euler solver
abstract
As a traditional application on various supercomputers, atmospheric modeling has long been suffering from the low performance efficiency. In this paper, we pick the 3D Euler equation solver (the most essential dynamic component for a non-hydrostatic atmospheric model) as the target application, and explore the maximum performance efficiency that can be achieved on CPU-GPU hybrid architectures. Besides presenting the suitable hybrid domain decomposition methodology and taking proper usage of tuning techniques for both the CPU and GPU parts, we further propose a novel GPU tuning technique, namely the customizable data caching mechanism with thread warp rescheduling scheme, which is specifically designed for the Euler solver. Combining all the optimizing approaches together, remarkable performance boost has been achieved on mainstream GPU architectures including Tesla Fermi C2050, K20×, K40 and K80. Especially, on the latest Tesla K80, we demonstrate a 31.64× speedup over the performance of 12-core E5-2697 CPU. In addition, based on a hybrid CPU-GPU node with two 12-core E5-2697 CPUs and two Tesla K80 GPUs, a sustained double-precision performance of 1.04 Tflops (16% of the peak) is achieved, which is remarkably higher than the efficiency of similar optimizing tasks based on heterogeneous platforms (strictly less than 10%, as demonstrated in the related work). In addition, a nearly linear weak scaling efficiency is achieved which demonstrate the effectiveness of our domain decomposition method.
Haohuan Fu, Jingheng Xu, Lin Gan 0001, Chao Yang 0002, Wei Xue 0003, Wenlai Zhao, Guangwen Yang 0002
ASAP6
2016 Relation-oriented resource allocation for multi-accelerator systems
abstract
This paper presents a novel approach for allocating resources in systems with multiple accelerators. It has three main contributions. First, a new model based on Birkhoff's representation theory in capturing the ordering properties of resource allocation requests (RArs). Second, an effective technique for resource allocation based on this model, targeting systems with multiple accelerators. Third, the evaluation of the proposed approach for Maxeler MPC-X multi-accelerator systems, demonstrating time-efficiency and 30%-50% failure-rate decrease (FRD) on random input dataset.
Mark Stillwell, José Gabriel F. Coutinho, Wenlai Zhao, Shuang Liang 0012, Wayne Luk, Alexander L. Wolf, Yuchun Ma
ASAP5
2016 F-CNN: An FPGA-based framework for training Convolutional Neural Networks
abstract
This paper presents a novel reconfigurable framework for training Convolutional Neural Networks (CNNs). The proposed framework is based on reconfiguring a streaming datapath at runtime to cover the training cycle for the various layers in a CNN. The streaming datapath can support various parameterized modules which can be customized to produce implementations with different trade-offs in performance and resource usage. The modules follow the same input and output data layout, simplifying configuration scheduling. For different layers, instances of the modules contain different computation kernels in parallel, which can be customized with different layer configurations and data precision. The associated models on performance, resource and bandwidth can be used in deriving parameters for the datapath to guide the analysis of design trade-offs to meet application requirements or platform constraints. They enable estimation of the implementation specifications given different layer configurations, to maximize performance under the constraints on bandwidth and hardware resources. Experimental results indicate that the proposed module design targeting Maxeler technology can achieve a performance of 62.06 GFLOPS for 32-bit floating-point arithmetic, outperforming existing accelerators. Further evaluation based on training LeNet-5 shows that the proposed framework achieves about 4 times faster than CPU implementation of Caffe and about 7.5 times more energy efficient than the GPU implementation of Caffe.
Wenlai Zhao, Haohuan Fu, Wayne Luk, Yuchun Ma, Guangwen Yang 0002
ASAP1
2016 Generalized GPU Acceleration for Applications Employing Finite-Volume Methods
abstract
Scientific HPC applications are increasingly ported to GPUs to benefit from both the high throughput and the powerful computing capacity. Many of these applications, such as atmospheric modeling and hydraulic erosion simulation, are adopting the finite volume method (FVM) as the solver algorithm. However, the communication components inside these applications generally lead to a low flop-to-byte ratio and an inefficient utilization of GPU resources. This paper aims at optimizing FVM solver based on the structured mesh. Besides a high-level overview of the finite-volume method as well as its basic optimizations on modern GPU platforms, we further present two generalized tuning techniques including an explicit cache mechanism as well as an inner-thread rescheduling method that tries to achieve a suitable mapping between the algorithm feature and the platform architecture. To the end, we demonstrate the impact of our generalized optimization methods in two typical atmospheric dynamic kernels (Euler and SWE) based on four mainstream GPU platforms. According to the experimental results of Tesla K80, speedups of 24.4x for SWE and 31.5x for Euler could be achieved over a 12-core Intel E5-2697 CPU, which is a great promotion compared with its original speedup (18x and 15.47x) without applying these two methods.
Jingheng Xu, Haohuan Fu, Lin Gan 0001, Chao Yang 0002, Wei Xue 0003, Shizhen Xu, Wenlai Zhao, Bingwei Chen, Guangwen Yang 0002
CCGrid7
2015 Optimizing Residue Number Reverse Converters through Bitwise Arithmetic on FPGAs
abstract
As a promising number representation method to provide inspiring operational performance, the Residue Number System (RNS) has been widely applied in many key applications for data pocessing. However, a highly-efficient and general-purpose reverse converter, which is the key component in an RNS system, is still less to be seen, due to the costly and complex operators that require large amounts of computing resources and a long latency to accomplish. In this paper, we are targeting at reverse converters that are highly efficient and can support general moduli sets. We first propose optimizing methods based on the bit wise arithmetic to improve the performance of general reverse converters such as CRT and New CRT. The methods are capable of replacing expensive operations such as additions and multiplications with bit wise operations. We also optimize the performance of specific reverse converter through condition reduction and pre-calculation methods. Furthermore, we develop a user controlled FPGA design generator that can produce optimized reverse converter designs for a number of different moduli sets. Compared with the existing optimized converter designs, our proposed methods can further reduce the latency and resource consumption by 54.2% to 84.6% and 65% to 88.5% respectively.
Bangtian Liu, Haohuan Fu, Lin Gan 0001, Wenlai Zhao, Guangwen Yang 0002
FCCM4
2014 A Fully-Pipelined FPGA Design for Tree-Reweighted Message Passing Algorithm
Wenlai Zhao, Haohuan Fu, Guangwen Yang 0002
FCCM1
2014 Patra: Parallel tree-reweighted message passing architecture
abstract
Maximum a posteriori probability inference algorithms for Markov Random Field are widely used in many applications, such as computer vision and machine learning. Sequential tree-reweighted message passing (TRW-S) is an inference algorithm which shows good quality in finding optimal solutions. However, the performance of TRW-S in software cannot meet the requirements of many real-time applications, due to the sequential scheme and the high memory, bandwidth and computational costs. This paper proposes Patra, a novel parallel tree-reweighted message passing architecture, which involves a fully pipelined design targeting FPGA technology. We build a hybrid CPU/FPGA system to test the performance of Patra for stereo matching. Experimental results show that Patra provides about 100 times faster than a software implementation of TRW-S, and 12 times faster than a GPU-based message passing algorithm. Compared with an existing design in four FPGAs, we can achieve 2 times speedup in a single FPGA. Moreover, Patra can work at video rate in many cases, such as a rate of 167 frame/sec for a standard stereo matching test case, which makes it promising for many real-time applications.
Wenlai Zhao, Haohuan Fu, Guangwen Yang 0002, Wayne Luk
FPL1