VLDB 2026 Research / reviewers in the wild / expert
Chuliang Guo
dblp:261/7710
· DBLP profile ↗
11ranked-venue papers
8as first author
9since 2021 · last 2025
0000-0001-7403-0163ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 7 first-author · 8 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FexMo: Enabling Fuse Execution Mode for Multi-task CGRAs
Chenhao Xie 0001, Chuliang Guo, Liansheng Liu, Xiyuan Peng, Datong Liu, Yu Peng 0002 |
MICRO | 3 |
| 2025 | FPGA-based component-wise LSTM training accelerator for neural granger causality analysis
Chuliang Guo, Yu Fu 0008 |
Neurocomputing | 1 |
| 2025 | Highly Parallel CNN Accelerator for RepVGG-Like Network Training on FPGAsabstractIn this article, we propose a generic FPGA-based training accelerator tailored for RepVGG-like networks, which strikes a balance between maximizing training-time accuracy and minimizing inference-time latency. The proposed accelerator leverages fine-grain channel-level parallelism within computational units specially designed for multiple branches of the basic building block within the RepVGG-like network. Specifically, we employ a Conv block for forward Conv and backward deConv, along with a dilated Conv block, including a weight kernel partition scheme for efficient weight gradient calculation. Furthermore, we aggressively exploit a 2-stage coarse-grain task-level parallelism for low-latency CNN training: 1) parallelism among multiple branches of the basic building block of RepVGG and 2) parallelism between error back-propagation and weight gradient calculation in the backward path. Through experiments on the CIFAR-10 dataset using 16-bit fixed-point arithmetic, we demonstrate state-of-the-art batch 1 throughput of 150 GOPs for training and 183 GOPs for inference. Chuliang Guo, Binglei Lou, David Boland, Philip H. W. Leong |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | Single-Batch CNN Training using Block Minifloats on FPGAsabstractTraining convolutional neural networks remains a challenge on resource-limited edge devices due to its intensive computations, large storage requirements, and high bandwidth. Error back-propagation, gradient generation, and weight update usually require high precision to guarantee model accuracy, which places a further burden on computation and bandwidth. This paper presents the first parallel FPGA CNN training accelerator with block minifloat datatypes. We first propose a heuristic bit-width allocation technique to derive a unified 8-bit block minifloat format with a sign bit, 2 exponent bits, and 5 mantissa bits. In contrast to previous techniques, the same data format is used for weights, activations, errors, and gradients. Using this format, accuracy similar to 32-bit single precision floating point is achieved and thus simplifies the FPGA-based designs of computational units such as multiply-and-add. In addition, we propose a unified Conv block to deal with Conv and transposed Conv in the forward and backward paths respectively; and a dilated Conv block with a weight kernel partition scheme for gradient generation. Both Conv blocks support non-unit stride, this being crucial for the residual connections that appear in modern CNNs. For training of ResNet20 on the CIFAR-10 dataset with a batch size of 1, our accelerator on a Xilinx Ultrascale+ ZCU102 FPGA achieves state-of-the-art single-batch throughput of 144.64 and 192.68 GOPs with and without batch normalisation layers respectively. Chuliang Guo, Binglei Lou, Xueyuan Liu 0002, David Boland, Philip H. W. Leong |
FPGA | 1 |
| 2023 | BOOST: Block Minifloat-Based On-Device CNN Training Accelerator with Transfer LearningabstractAdapting CNNs to changing problems is challenging on resource-limited edge devices due to intensive computations, high precision requirements, large storage needs, and high bandwidth. This paper presents BOOST, a novel block minifloat (BM)-based parallel CNN training accelerator on memory- and computation-constrained FPGAs for transfer learning (TL). By updating a small number of layers online, BOOST enables adaptation to changing problems. Our approach utilizes a unified 8-bit BM datatype (bm(2,5) ), i.e., with a sign bit, 2 exponent bits, and 5 mantissa bits, and proposes unified Conv and dilated Conv blocks that support non-unit stride and enable task-level parallelism during back-propagation to minimize latency. For ResNet20 and VGG-like training on CIFAR-10 and SVHN datasets, BOOST achieves near 32-bit floating point accuracy, reducing latency by 21%-43% and BRAM usage by 63%-66% compared to back-propagation training without TL. Notably, BOOST outperforms the prior SOTA works to achieve perbatch throughput of 131 and 209 GOPs for ResNet20 and VGG-like respectively. Chuliang Guo, Binglei Lou, Xueyuan Liu 0002, David Boland, Philip H. W. Leong, Cheng Zhuo |
ICCAD | 1 |
| 2023 | GANDSE: Generative Adversarial Network-based Design Space Exploration for Neural Network Accelerator DesignabstractWith the popularity of deep learning, the hardware implementation platform of deep learning has received increasing interest. Unlike the general purpose devices, e.g., CPU or GPU, where the deep learning algorithms are executed at the software level, neural network hardware accelerators directly execute the algorithms to achieve higher energy efficiency and performance improvements. However, as the deep learning algorithms evolve frequently, the engineering effort and cost of designing the hardware accelerators are greatly increased. To improve the design quality while saving the cost, design automation for neural network accelerators was proposed, where design space exploration algorithms are used to automatically search the optimized accelerator design within a design space. Nevertheless, the increasing complexity of the neural network accelerators brings the increasing dimensions to the design space. As a result, the previous design space exploration algorithms are no longer effective enough to find an optimized design. In this work, we propose a neural network accelerator design automation framework named GANDSE, where we rethink the problem of design space exploration, and propose a novel approach based on the generative adversarial network (GAN) to support an optimized exploration for high-dimension large design space. The experiments show that GANDSE is able to find the more optimized designs in negligible time compared with approaches including multilayer perceptron and deep reinforcement learning. Lang Feng 0001, Chuliang Guo, Ke Tang 0006, Cheng Zhuo, Zhongfeng Wang 0001 |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2022 | Landslide susceptibility assessment based on multi GPUs: a deep learning approach
Chuliang Guo, Jinxia Wu, Shuaihe Zhao, Sansar Raj Meena, Feng Zhang 0012 |
CCF Trans. High Perform. Comput. | 1 |
| 2021 | Joint Sparsity with Mixed Granularity for Efficient GPU Implementation
Chuliang Guo, Xingang Yan, Yufei Chen 0007, He Li 0008, Xunzhao Yin, Cheng Zhuo |
DATE | 1 |
| 2021 | A Reconfigurable Multiplier for Signed Multiplications with Asymmetric Bit-WidthsabstractMultiplications have been commonly conducted in quantized CNNs, filters, and reconfigurable cores, and so on, which are widely deployed in mobile and embedded applications. Most multipliers are designed to perform multiplications with symmetric bit-widths, i.e., n - by n -bit multiplication. Such features would cause extra area overhead and performance loss when m - by n -bit multiplications ( m > n ) are deployed in the same hardware design, resulting in inefficient multiplication operations. It is highly desired and challenging to propose a reconfigurable multiplier design to accommodate operands with both symmetric and asymmetric bit-widths. In this work, we propose a reconfigurable approximate multiplier to support multiplications at various precisions, i.e., bit-widths. Unlike prior works of approximate adders assuming a uniform weight distribution with bit-wise independence, scenarios like a quantized CNN may have a centralized weight distribution and hence follow a Gaussian-like distribution with correlated adjacent bits. Thus, a new block-based approximate adder is also proposed as part of the multiplier to ensure energy-efficient operation with an awareness of the bit-wise correlation. Our experimental results show that the proposed approximate adder significantly reduces the error rate by 76% to 98% over a state-of-the-art approximate adder for Gaussian-like distribution scenarios. Evaluation results show that the proposed multiplier is 19% faster and 22% more power saving than a Xilinx multiplier IP at the same bit precision and achieves a 23.94-dB peak signal-to-noise ratio, which is comparable to the accurate one of 24.10 dB when deployed in a Gaussian filter for image processing tasks. Chuliang Guo, Li Zhang 0021, Grace Li Zhang, Bing Li 0005, Weikang Qian, Xunzhao Yin, Cheng Zhuo |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2020 | A Reconfigurable Approximate Multiplier for Quantized CNN ApplicationsabstractQuantized CNNs, featured with different bit-widths at different layers, have been widely deployed in mobile and embedded applications. The implementation of a quantized CNN may have multiple multipliers at different precisions with limited resource reuse or one multiplier at higher precision than needed causing area overhead. It is then highly desired to design a multiplier by accounting for the characteristics of quantized CNNs to ensure both flexibility and energy efficiency. In this work, we present a reconfigurable approximate multiplier to support multiplications at various precisions, i.e., bit-widths. Moreover, unlike prior works assuming uniform distribution with bit-wise independence, a quantized CNN may have centralized weight distribution and hence follow a Gaussian-like distribution with correlated adjacent bits. Thus, a new block-based approximate adder is also proposed as part of the multiplier to ensure energy efficient operation with awareness of bit-wise correlation. Our experimental results show that the proposed adder significantly reduces the error rate by 76-98% over a state-of-the-art approximate adder for such scenarios. Moreover, with the deployment of the proposed multiplier, which is 17% faster and 22% more power saving than a Xilinx multiplier IP at the same precision, a quantized CNN implemented in FPGA achieves 17% latency reduction and 15% power saving compared with a full precision case. Chuliang Guo, Li Zhang 0021, Weikang Qian, Cheng Zhuo |
ASP-DAC | 1 |
| 2020 | A Convolutional Neural Network Accelerator Architecture with Fine-Granular Mixed Precision ConfigurabilityabstractConvolutional neural networks (CNNs) have been widely deployed in deep learning applications, especially on power hungry GP-GPUs. Recent efforts in designing CNN accelerators are considered as a promising alternative to achieve higher energy efficiency. Unfortunately, with the growing complexity of CNN, the demanded computational and storage resources for accelerators keep increasing, hindering its wider applications in mobile devices. On the other hand, many quantization algorithms have been proposed for efficient CNN training, which brings many small or zero weights. This is a unique opportunity for accelerator designers to employ much fewer bits, e.g., 4 bits, in both arithmetic core and storage, thereby saving significant design cost. However, such a single precision strategy inevitably compromises the accuracy as some key operations may demand a higher precision. Thus, this paper proposes a low power CNN accelerator architecture that can simultaneously conduct computations with mixed precisions and assign the appropriate arithmetic cores to operation with different precision demands. This proposed architecture can achieve significant area and energy savings, without accuracy compromise. The experimental results show that the proposed architecture implemented on FPGA can reduces almost half of the weight storage and MAC area, and lower the dynamic power by 12.1% when compared with a state-of-the-art CNN accelerator design. Li Zhang 0021, Chuliang Guo, Xunzhao Yin, Cheng Zhuo |
ISCAS | 3 |