Ruizhe Zhao

dblp:130/9004 · DBLP profile ↗
← Back
13ranked-venue papers
7as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 5 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorComputer networks · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2025 Linear Array Motion-Based Time-Varying Channel Estimation for Three-Dimensional Millimeter-Wave Systems
abstract
Using linear arrays for compressive sensing-based (CS) three-dimensional (3D) millimeter-wave (mmWave) channel estimation presents significant challenges due to the dimensional limitations of the array aperture, which prevents direct sparse 3D channel representation. Fortunately, the emergence of array motion provides a chance for that. In this paper, by exploring the motion of linear arrays in time-varying channels and identifying their similarities to planar arrays, we propose a method for sparse representation of 3D channels based on linear arrays. Utilizing the sparsity, a two-step sparse channel recovery algorithm is presented. Initially, the on-grid and off-grid approaches are employed to efficiently and accurately obtain the angles of arrival (AoAs). Subsequently, path gains are estimated based on these AoAs. Simulation results show that the proposed method can effectively estimate 3D time-varying mmWave channels using a linear array with only a small number of pilots.
Ruizhe Zhao, Zhi Zhang 0003, Jianhong Chu, Tianshu Su
WCNC1
2022 POLSCA: Polyhedral High-Level Synthesis with Compiler Transformations
abstract
Polyhedral optimization can parallelize nested affine loops for high-level synthesis (HLS), but polyhedral tools are HLS-agnostic and can worsen performance. Moreover, HLS tools require user directives which can produce unreadable polyhedral-transformed code. To address these two challenges, we present POLSCA, a compiler framework that improves polyhedral HLS workflow by automatic code transformation. POLSCA decomposes a design before polyhedral optimization to balance code complexity and parallelism, while revising memory interfaces of polyhedral-transformed code to make partitioning explicit for HLS tools; it enables designs to benefit more easily from polyhedral optimization. Experiments on Polybench/C show that POLSCA designs are 1.5 times faster on average compared with baseline designs generated directly from applying HLS on C code.
Ruizhe Zhao, Jianyi Cheng, Wayne Luk, George A. Constantinides
FPL1
2021 Polygeist: Raising C to Polyhedral MLIR
abstract
We present Polygeist, a new compilation flow that connects the MLIR compiler infrastructure to cutting edge polyhedral optimization tools. It consists of a C and C++ frontend capable of converting a broad range of existing codes into MLIR suitable for polyhedral transformation and a bi-directional conversion between MLIR and OpenScop exchange format. The Polygeist/MLIR intermediate representation featuring high-level (affine) loop constructs and n-D arrays embedded into a single static assignment (SSA) substrate enables an unprecedented combination of SSA-based and polyhedral optimizations. We illustrate this by proposing and implementing two extra transformations: statement splitting and reduction parallelization. Our evaluation demonstrates that Polygeist outperforms on average both an LLVM IR-level optimizer (Polly) and a source-to-source state-of-the-art polyhedral compiler (Pluto) when exercised on the Polybench/C benchmark suite in sequential (2.53x vs 1.41x, 2.34x) and parallel mode (9.47x vs 3.26x, 7.54x) thanks to the new representation and transformations.
William S. Moses, Lorenzo Chelini, Ruizhe Zhao, Oleksandr Zinenko
PACT3
2021 In-circuit tuning of deep learning designs
Zhiqiang Que, Daniel H. Noronha, Ruizhe Zhao, Xinyu Niu, Steve Wilton, Wayne Luk
J. Syst. Archit.3
2020 Reducing Underflow in Mixed Precision Training by Gradient Scaling
abstract
By leveraging the half-precision floating-point format (FP16) well supported by recent GPUs, mixed precision training (MPT) enables us to train larger models under the same or even smaller budget. However, due to the limited representation range of FP16, gradients can often experience severe underflow problems that hinder backpropagation and degrade model accuracy. MPT adopts loss scaling, which scales up the loss value just before backpropagation starts, to mitigate underflow by enlarging the magnitude of gradients. Unfortunately, scaling once is insufficient: gradients from distinct layers can each have different data distributions and require non-uniform scaling. Heuristics and hyperparameter tuning are needed to minimize these side-effects on loss scaling. We propose gradient scaling, a novel method that analytically calculates the appropriate scale for each gradient on-the-fly. It addresses underflow effectively without numerical problems like overflow and the need for tedious hyperparameter tuning. Experiments on a variety of networks and tasks show that gradient scaling can improve accuracy and reduce overall training effort compared with the state-of-the-art MPT.
Ruizhe Zhao, Brian K. Vogel, Wayne Luk
IJCAI1
2019 On-chip FPGA Debug Instrumentation for Machine Learning Applications
abstract
FPGAs provide a promising implementation option for many machine learning applications. Although simulations or software models can be used to explore the design space of these applications, often the final behaviour can not be evaluated until the design is mapped to the FPGA and integrated into the target system. This may be because long run-times are required, or because the environment can not be adequately described using a software model. Once unexpected behaviour is observed, on-chip debug is notoriously difficult; typically a design is instrumented with on-chip trace buffers that record the run-time behaviour for later interrogation. In this paper, we describe instrumentation that can accelerate the process of debugging machine learning applications implemented on an FPGA. Unlike previous work, our instrumentation is optimized to take advantage of characteristics of this application domain. Our instruments gather useful domain-specific information about the observed variables instead of recording the raw values of those elements. Results show that the proposed instruments provide at least 17.8x longer visibility in the most conservative of our experiments at a low area and latency cost.
Daniel H. Noronha, Ruizhe Zhao, Jeffrey B. Goeders, Wayne Luk, Steve Wilton
FPGA2
2019 Towards In-Circuit Tuning of Deep Learning Designs
abstract
This paper presents InTune, a novel approach for in-circuit tuning of deep learning designs targeting implementations in field-programmable gate array technology. This approach combines two promising techniques: domain-specific adaptation and in-circuit tuning. Domain-specific adaptation exploits domain-specific information in adapting pre-trained models to specific application domains, replacing standard convolution layers with efficient convolution blocks; the effects of such adaptation are then assessed by in-circuit tuning instruments to provide information to application builders for tuning the design. This approach is illustrated by its deployment in tuning deep neural networks, and its potential for a new generation of domain-specific tools with tight integration of synthesis and in-circuit tuning is explored.
Zhiqiang Que, Daniel H. Noronha, Ruizhe Zhao, Steve Wilton, Wayne Luk
ICCAD3
2018 Hardware Compilation of Deep Neural Networks: An Overview
abstract
Deploying a deep neural network model on a reconfigurable platform, such as an FPGA, is challenging due to the enormous design spaces of both network models and hardware design. A neural network model has various layer types, connection patterns and data representations, and the corresponding implementation can be customised with different architectural and modular parameters. Rather than manually exploring this design space, it is more effective to automate optimisation throughout an end-to-end compilation process. This paper provides an overview of recent literature proposing novel approaches to achieve this aim. We organise materials to mirror a typical compilation flow: front end, platform-independent optimisation and back end. Design templates for neural network accelerators are studied with a specific focus on their derivation methodologies. We also review previous work on network compilation and optimisation for other hardware platforms to gain inspiration regarding FPGA implementation. Finally, we propose some future directions for related research.
Ruizhe Zhao, Shuanglong Liu, Ho-Cheung Ng, Erwei Wang, James J. Davis 0001, Xinyu Niu, Huifeng Shi, George A. Constantinides, Peter Y. K. Cheung, Wayne Luk
ASAP1
2018 Automatic Optimising CNN with Depthwise Separable Convolution on FPGA: (Abstact Only)
abstract
Convolution layers in Convolutional Neural Networks (CNNs) are effective in vision feature extraction but quite inefficient in computational resource usage. Depthwise separable convolution layer has been proposed in recent publications to enhance the efficiency without reducing the effectiveness by separately computing the spatial and cross-channel correlations from input images and has proven successful in state-of-the-art networks such as MobileNets [1] and Xception [2]. Based on the facts that depthwise separable convolution is highly structured and uses limited resources, we argue that it can well fit reconfigurable platforms like FPGA. To benefit FPGA platforms with this new layer, in this paper, we present a novel framework that can automatically generate and optimise hardware designs for depthwise separable CNNs. Besides, in our framework, existing conventional CNNs can be systematically converted to ones whose standard convolution layers are selectively replaced with functionally identical depthwise separable convolution layers, by carefully balancing the trade-off among speed, accuracy, and resource usage through resource usage modelling and network fine-tuning. Results show that hardware designs generated by our framework can reach at most 231.7 frames per second regarding MobileNets, and for VGG-16 [3], we gain 3.43 times speed-up and 3.54% accuracy decrease on the ImageNet [4] dataset comparing the original model and a layer replaced one.
Ruizhe Zhao, Xinyu Niu, Wayne Luk
FPGA1
2018 Towards Efficient Convolutional Neural Network for Domain-Specific Applications on FPGA
abstract
FPGA becomes a popular technology for implementing Convolutional Neural Network (CNN) in recent years. Most CNN applications on FPGA are domain-specific, e.g., detecting objects from specific categories, in which commonly-used CNN models pre-trained on general datasets may not be efficient enough. This paper presents TuRF, an end-to-end CNN acceleration framework to efficiently deploy domain-specific applications on FPGA by transfer learning that adapts pre-trained models to specific domains, replacing standard convolution layers with efficient convolution blocks, and applying layer fusion to enhance hardware design performance. We evaluate TuRF by deploying a pre-trained VGG-16 model for a domain-specific image recognition task onto a Stratix V FPGA. Results show that designs generated by TuRF achieve better performance than prior methods for the original VGG-16 and ResNet-50 models, while for the optimised VGG-16 model TuRF designs are more accurate and easier to process. and layer fusion to enhance hardware design performance. We evaluate TuRF by end-to-end deploying a pre-trained VGG-16 model for a domain-specific image recognition task onto Stratix V FPGA. Results show that designs generated by our framework achieve better performance than prior works, regarding the original VGG-16 and ResNet-50, and the optimised VGG-16 model is more accurate and easier to process.
Ruizhe Zhao, Ho-Cheung Ng, Wayne Luk, Xinyu Niu
FPL1
2017 DeepPump: Multi-pumping deep Neural Networks
abstract
This paper presents DeepPump, a novel approach for generating and optimising hardware designs of deep Convolutional Neural Networks (CNNs) with multi-pumping on FPGA platforms. Multi-pumping [1] is a promising technique to save hardware resource usage by replacing M parallel units with one clocked at M times the global clock rate. DeepPump aims at automatically adopting multi-pumping when generating hardware designs for CNNs. It has three components: a parameterised CNN accelerator architecture that supports multi-pumping, a design model for trade-off analysis related to multi-pumping, and an optimisation flow for improving the architecture based on the design model.
Ruizhe Zhao, Tim Todman, Wayne Luk, Xinyu Niu
ASAP1
2017 Scale-Free Sparse Matrix-Vector Multiplication on Many-Core Architectures
abstract
Sparse matrix-vector multiplication (SpMV) is one of the most important kernels for many applications. In this paper, we study the implementation of SpMV for scale-free matrices on many-core architectures including graphic processing units and Xeon Phi coprocessors. We first propose a hardware oblivious implementation for heterogeneous many-core processors using OpenCL. Our OpenCL implementation uses a novel SpMV format called hybrid COO+CSR (HCC), which employs 2-D jagged partitioning to balance the workload among a large number of cores and improve the data locality. Moreover, the OpenCL implementation is designed to be parametric, which allows systematic performance tuning. We conduct experiments to evaluate the efficiency of our hardware oblivious implementation. Experiments show that it achieves comparable performance to the Intel MKL and state-of-the-art OpenCL-based ViennaCL library implementation. Although the OpenCL implementation provides functional portability for heterogeneous systems, it fails to take advantage of the low-level architectural features. To further improve the performance, we propose a hardware conscious implementation using the native parallel programming language. We use the Xeon Phi platform as a case study. In our hardware conscious implementation, we ensure that the HCC format efficiently utilizes the vector process units on Xeon Phi by employing low-level intrinsics, and improve the overall performance through locality-aware block mapping, and intrablock tiling. Experiments using a wide range of representative scale-free matrices demonstrate that compared with the OpenCL-based hardware oblivious implementation, the hardware conscious implementation achieves 2.2× speedup on average. Compared with MKL, the hardware conscious implementation achieves 3.1× speedup on Xeon Phi.
Yun Liang 0001, Wai Teng Tang, Ruizhe Zhao, Mian Lu, Huynh Phung Huynh, Rick Siow Mong Goh
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2015 Optimizing and auto-tuning scale-free sparse matrix-vector multiplication on Intel Xeon Phi
abstract
Recently, the Intel Xeon Phi coprocessor has received increasing attention in high performance computing due to its simple programming model and highly parallel architecture. In this paper, we implement sparse matrix vector multiplication (SpMV) for scale-free matrices on the Xeon Phi architecture and optimize its performance. Scale-free sparse matrices are widely used in various application domains, such as in the study of social networks, gene networks and web graphs. We propose a novel SpMV format called vectorized hybrid COO+CSR (VHCC). Our SpMV implementation employs 2D jagged partitioning, tiling and vectorized prefix sum computations to improve hardware resource utilization, and thus overall performance. As the achieved performance depends on the number of vertical panels, we also develop a performance tuning method to guide its selection. Experimental results demonstrate that our SpMV implementation achieves an average 3× speedup over Intel MKL for a wide range of scale-free matrices.
Wai Teng Tang, Ruizhe Zhao, Mian Lu, Yun Liang 0001, Huynh Phung Huyng, Xibai Li, Rick Siow Mong Goh
CGO2