EDBT 2026 Demo / reviewers in the wild / expert
Ruiqi Chen 0001
dblp:184/7077-1
· DBLP profile ↗
28ranked-venue papers
9as first author
28since 2021 · last 2026
0000-0001-6837-5675ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 26 · 9 first-author · 26 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Equicore: Accelerating Clebsch-Gordan Tensor Product of Equivariant Neural Networks on FPGAabstractEquivariant neural networks (ENNs) are a powerful framework for modeling 3D geometric data in physical and biological systems. The Clebsch–Gordan tensor product (CGTP)—a core operation for preserving equivariance—remains the primary computational bottleneck in ENNs. Although Clebsch–Gordan (CG) coefficients exhibit pronounced structural sparsity, prior work has neither fully leveraged this property nor adopted hardware-friendly quantization, leading to limited efficiency. We present Equicore, a software–hardware co-design framework to accelerate CGTP in ENNs. Equicore introduces three key innovations: (1) a sparse-bypass strategy that exploits the CG structural sparsity together with a novel CG data format to pack the overlapping non-zeros, bypassing redundant data accesses and computations comparing to previous sparse solutions; (2) a merged-shift quantization strategy that enables full Int8 representation of irreps, weights, and CG coefficients using shift-only operations; and (3) a cascaded processing unit that tightly couples the FPGA hardware resources to achieve high operating frequency while supporting efficient sparse and quantized computation. Deployed on a AMD Virtex VCU128 platform, Equicore delivers up to 10.5× speedup and 17.4× energy-efficiency improvement over state-of-the-art GPU libraries and FPGA designs across diverse CGTP types in a benchmark of eleven ENN models. Shidi Tang, Chuanzhao Zhang, Ruiqi Chen 0001, Yuxuan Lv, Bruno da Silva 0001 |
DATE | 3 |
| 2026 | Breaking the BRAM Wall: Scalable Vina FPGA Acceleration via Distributed Grid Storage and Cross-Board Long-Ring PipelinesabstractAutoDock Vina (Vina), a gold standard for molecular docking, is hampered by computational expense. Previous FPGA hardware accelerators were fundamentally constrained by an on-chip memory bottleneck, where storing large, pre-computed energy grids consumes the majority of BRAM resources. This leads to imbalanced resource utilization, as the exhausted on-chip memory makes it impossible to further increase the degree of intra-node parallelism by instantiating more processing units. This paper introduces a novel, scalable multi-FPGA architecture that systematically removes this limitation.Our architecture’s innovation is a synergistic combination of three mechanisms. First, Distributed Grid Storage partitions the energy grid across all nodes to break the BRAM bottleneck. Second, a Cross-Board Long-Ring Pipeline creates a high-throughput dataflow for distributed energy calculations. Third, a dynamic intra-node scheduler unlocks massive fine-grained parallelism within each node. Together, these mechanisms create a powerful synergistic effect where adding nodes not only increases aggregate throughput but also enhances the performance of each individual node. This intrinsic amplification of per-node capacity is the direct driver of the system’s super-linear performance scaling.Implemented on a three-ZCU102 FPGA system and without sacrificing accuracy, our single-board normalized performance is 7.6× faster than the Vina-FPGA and 1.95× faster than the state-of-the-art Vina-FPGA-Cluster. Critically, the architecture demonstrates super-linear performance scaling: the three-board system achieves a 3.7× speedup over a single node, outperforming the ideal linear 3× speedup. Ankun Tian, Shidi Tang, Ruiqi Chen 0001 |
DATE | 3 |
| 2026 | IDSPfree: An FPGA-Based Intrusion Detection System with DSP-Free Design
Abdessamad Nassihi, Ahmed Sadaqa, Muhammad Iqbal Khan, Ruiqi Chen 0001, Bruno da Silva 0001 |
ISCAS | 5 |
| 2026 | Power-Efficient Spiking Conversion of Deep Unfolded Transformers
Ahmed Sadaqa, Brent De Weerdt, Ruiqi Chen 0001, Nikos Deligiannis, Bruno da Silva 0001 |
ISCAS | 3 |
| 2026 | BenDan: Benchmarking DPU performance on FPGAs
Ahmed Sadaqa, Yanxiang Zhu, Shidi Tang, Ruiqi Chen 0001, Bruno da Silva 0001 |
Integr. | 7 |
| 2026 | FP8ApproxLib: An FPGA-based approximate multiplier library for 8-bit floating point
Ruiqi Chen 0001, Yangxintong Lyu, Shidi Tang, Jindong Li 0001, Yanxiang Zhu, Bruno da Silva 0001 |
J. Syst. Archit. | 1 |
| 2026 | DIF-LUT Pro: An Automated Tool for Simple yet Scalable Approximation of Nonlinear Activation on FPGAabstractNonlinear activation plays an essential role in neural networks (NNs) for their generalization ability. However, implementing intricate mathematical operations on hardware platforms, including Field-Programmable Gate Arrays (FPGAs), presents significant challenges. Prior works based on piecewise functions or look-up table (LUT) have encountered difficulties in balancing precision requirements with fair hardware overhead and often necessitating complex manual interventions. To address these issues, this paper proposes DIF-LUT Pro, an automated tool for simple yet scalable approximation for various nonlinear activations on FPGA. Specifically, the proposed algorithm achieves self-adaptive hardware design oriented towards target precision, by piecewise linear matching to fit the function derivative roughly and range addressable LUT to offset the difference. Moreover, DIF-LUT Pro integrates the algorithm into an automated tool, allowing users to configure the customized interface and generate the corresponding hardware description language (HDL) code with a single click. Experimental results show that (1) DIF-LUT Pro features robust automation and fair generality, capable of generating equitable hardware designs under various user configurations across different FPGA platforms; (2) DIF-LUT Pro produces approximations that are simple yet effective, achieving competitive performance compared to previous expert-crafted designs. Furthermore, two detailed case studies demonstrate the efficient application of DIF-LUT Pro on NeRF and SEResnet, proving its practical value. Our source code is open-source and available at https://github.com/AdrianLiu00/DIF-LUT-Tool. Yang Liu 0376, Yu Li 0003, Ruiqi Chen 0001, Jun Yu 0010, Kun Wang 0005 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2026 | Diff-Acc: An Efficient FPGA Accelerator for Unconditional Diffusion ModelsabstractThe diffusion model has achieved remarkable success in the era of Artificial Intelligence Generated Content (AIGC) across various tasks, such as image, video, text, material modeling, and molecular design. However, the diffusion model is computational intensive due to the long iteration of the reverse denoising process, which hinders its further advancement. Therefore, there is an urgent need to accelerate the diffusion model, especially in edge scenarios that require real-time computation. While researchers have made efforts to accelerate the diffusion model at the algorithm level using either efficient sampling or model quantization, they still suffer from accuracy degradation. More importantly, they have overlooked the hardware-level acceleration challenge. This work aims to bridge the gap by introducing Diff-Acc , the first FPGA accelerator for unconditional diffusion models with a novel step-wise quantization method that requires minimal calibration data to achieve the state-of-the-art (SOTA) PTQ quantization accuracy. Additionally, we adopt several hardware-oriented optimizations to reduce the computational overhead. At the architecture level, we fully analyze the computation flow of diffusion models and propose a novel architecture with group-wise parallelism to tackle the long iteration challenge. Besides, we decouple the data dependencies and adopt proper computational transformations at the micro-architecture level. Experiments on two unconditional diffusion models (DDIM and DDPM) with two image datasets (CIFAR-10 and ImageNet) demonstrate that our quantization method achieves the substantial improvements in image quality (FID: 6.67, sFID: 11.24) under 8-bit PTQ quantization. Compared with both server-based (Tesla V100 and Intel Xeon) and edge-based (Raspberry Pi 4 and Jetson Nano) platforms, Diff-Acc implemented on the Zynq UltraScale+ XCZU9EG FPGA demonstrates an up-to 12.5× energy efficiency. Particularly versus edge-based platforms, Diff-Acc achieves up to 10.26× and 1.97× performance improvements over CPU and GPU, respectively. Shidi Tang, Ruiqi Chen 0001, Yuxuan Lv, Pengwei Zheng, He Li 0008 |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2026 | FANE: FPGA-Based FP8 Approximate Neural Network EngineabstractThe 8-bit floating-point (FP8) format has gained growing interest in neural networks (NNs) for its superior dynamic range over traditional INT8. However, multiply-accumulate (MAC) operations remain a major source of power consumption during inference of NNs, which makes DSP-free design important, especially for edge FPGAs with few or no DSPs. Therefore, this brief presents FPGA-based FP8 approximate neural network engine (FANE), an FPGA-based approximate NN engine for FP8. We first introduce a novel approximation method that replaces the multiplications by linear additions. This approximate method reduces power consumption while maintaining high accuracy, outperforming the latest FP8 approximate multiplier by 53.15%. Based on this design, we construct an FP8 MAC unit and integrate it into both a convolution engine and a matrix–vector multiplication (MVM) unit. Finally, we integrate our design into a large language model (LLM). The result shows 61.5% higher efficiency (TOPS/W) than the previous design, demonstrating the superiority of FANE in terms of performance and power efficiency. The code of FANE is available on ourhttps://github.com/hanbao04/FANE-FPGA-based-FP8-Approximate-Neural-Network-Engine.git Shidi Tang, Jingdong Li, Ahmed Sadaqa, Ruiqi Chen 0001, Bruno da Silva 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2025 | ATE-GCN: An FPGA-Based Graph Convolutional Network Accelerator with Asymmetrical Ternary QuantizationabstractTernary quantization can effectively simplify matrix multiplication, which is the primary computational operation in neural network models. It has shown success in FPGA-based accelerator designs for emerging models such as GAT and Transformer. However, existing ternary quantization methods can lead to substantial accuracy loss under certain weight distribution pat-terns, such as GCN. Furthermore, current FPGA-based ternary weight designs often focus on reducing resource consumption while neglecting full utilization of FPGA DSP blocks, limiting maximum performance. To address these challenges, we propose ATE-GCN, an FPGA-based asymmetrical ternary quantization GCN accelerator using a software-hardware co-optimization approach. First, we adopt an asymmetrical quantization strategy with specific interval divisions tailored to the bimodal distribution of GCN weights, reducing accuracy loss. Second, we design a unified processing element (PE) array on FPGA to support various matrix computation forms, optimizing FPGA resource usage while leveraging the benefits of cascade design and ternary quantization, significantly boosting performance. Finally, we implement the ATE-GCN prototype on the VCU118 FPGA board. The results show that ATE-GCN maintains an accuracy loss below 2%. Additionally, ATE-GCN achieves average performance improvements of$224.13\times$and$11.1\times$, with up to$898.82\times$and$69.9\times$energy consumption saving compared to CPU and GPU, respectively. Moreover, compared to state-of-the-art FPGA-based GCN accelerators, ATE-GCN improves DSP efficiency by 63% with an average latency reduction of 11%. Ruiqi Chen 0001, Shidi Tang, Yang Liu 0376, Yanxiang Zhu, Bruno da Silva 0001 |
DATE | 1 |
| 2025 | FPGA-Based Approximate Multiplier for FP8abstractThe 8-bit floating-point (FP8) data format has been increasingly adopted in neural network (NN) computations due to its superior dynamic range compared to traditional INT8. However, FP8-based multiplication, a core operation in NNs, still incurs significant power consumption. To address this issue, this paper presents an FPGA-based approximate multiplier design for FP8. Firstly, we conduct a bit-level analysis of the approximation method. Based on this analysis, we implement a fine-grained optimized design on mainstream FPGAs (AMD and Altera) using primitives and templates combined with physical layout constraints. Then, the accuracy and resource utilization of the FP8 approximate multiplier are evaluated and analyzed. The results indicate that, compared to previous FPGA-based 8-bit designs, our design achieves the minimal LUT consumption. Finally, we integrate the design into the inference phase of a representative NN model, demonstrating its excellent power efficiency. To the best of our knowledge, this is the first FPGA-based FP8 approximate multiplier design, which can serve as a benchmark for future designs and comparisons of FPGA-based low-precision floating-point approximate multipliers. The code of this work is available in our GitLab. Ruiqi Chen 0001, Yangxintong Lyu, Yanxiang Zhu, Shidi Tang, Bruno da Silva 0001 |
FCCM | 1 |
| 2025 | TrackGNN: A Highly Parallelized and Self-Adaptive GNN Accelerator for Track Reconstruction on FPGAsabstractReal-time track reconstruction in high energy physics imposes stringent latency constraints, hindering the deployment of graph neural networks (GNNs) on general-purpose platforms. We present TrackGNN11https//github.com/silvenachen/TrackGNN, an open-sourced GNN accelerator for track reconstruction. Using a dataflow architecture with multiple parallelism and a self-adaptive renaming mechanism, TrackGNN shows 27.6× speedup over CPUs, up to 101.1× over GPUs, and 5.7× over an FPGA overlay. Compared with FlowGNN, the renaming mechanism also reduces end-to-end latency by 1.12-1.16× with negligible resource overhead. Ruiqi Chen 0001, Bruno da Silva 0001, Giorgian Borca-Tasciuc, Dantong Yu, Cong Hao |
FCCM | 3 |
| 2025 | Hummingbird: A Smaller and Faster Large Language Model Accelerator on Embedded FPGAabstractDeploying large language models (LLMs) on embedded devices remains a significant research challenge due to the high computational and memory demands of LLMs and the limited hardware resources available in such environments. While embedded FPGAs have demonstrated performance and energy efficiency in traditional deep neural networks, their potential for LLM inference remains largely unexplored. Recent efforts to deploy LLMs on FPGAs have primarily relied on large, expensive cloud-grade hardware and have only shown promising results on relatively small LLMs, limiting their real-world applicability. In this work, we present Hummingbird, a novel FPGA accelerator designed specifically for LLM inference on embedded FPGAs. Hummingbird is smaller—targeting embedded FPGAs such as the KV260 and ZCU104 with 67% LUT, 39% DSP, and 42% power savings over existing research. Hummingbird is stronger—targeting LLaMA3-8B and supporting longer contexts, overcoming the typical 4GB memory constraint of embedded FPGAs through offloading strategies. Finally, Hummingbird is faster—achieving 4.8 tokens/s and 8.6 tokens/s for LLaMA3-8B on the KV260 and ZCU104 respectively, with 93-94% model bandwidth utilization, outperforming the prior 4.9 token/s for LLaMA2-7B with 84% bandwidth utilization baseline. We further demonstrate the viability of industrial applications by deploying Hummingbird on a cost-optimized Spartan UltraScale FPGA, paving the way for affordable LLM solutions at the edge. Jindong Li 0001, Ruiqi Chen 0001, Guobin Shen, Dongcheng Zhao, Qian Zhang 0080, Yi Zeng 0001 |
ICCAD | 3 |
| 2025 | Diff-DiT: Temporal Differential Accelerator for Low-bit Diffusion Transformers on FPGAabstractDiffusion Transformer (DiT) models have shown superior generative capabilities in image and video synthesis, yet their high computational cost during inference remains a critical bottleneck. Temporal differential computation offers a promising solution to low-bit quantization by exploiting the temporal similarity in activations. However, applying this technique to DiT’s Attention layers introduces substantial memory and computation overheads.In this paper, we present Diff-DiT, the first FPGA accelerator designed for low-bit DiT inference with differential computation. To overcome the unique challenges of DiT quantization and hardware acceleration, we propose: (1) an approximated differential attention (ADA) method that selectively approximates attention computations across time steps using a significance score, enabling low-bit on-chip execution while minimizing memory overhead; (2) an optimal cross-cast data accessing pattern with flexible data reuse to maximize computational intensity during matrix multiplications; and (3) a half-condition splitting (HCS) dataflow optimization and fine-grained pipelining to reduce the computation and memory access latency.Extensive experiments show that Diff-DiT outperforms NVIDIA V100 GPU by 1.39× in end-to-end throughput and 5.60× in energy efficiency. When compared with state-of-the-art diffusion model accelerators, Diff-DiT also achieves 2.81× and 2.77× improvements in throughput and energy efficiency, respectively. Code is available on GitHub1. Shidi Tang, Pengwei Zheng, Ruiqi Chen 0001, Yuxuan Lv, Bruno da Silva 0001 |
ICCAD | 3 |
| 2025 | Vina-FPGA2: a high-level parallelized hardware-accelerated molecular docking tool based on the inter-module pipelineabstractAutoDock Vina (Vina) is a widely adopted molecular docking tool, often regarded as a standard or used as a baseline in numerous studies. However, its computational process is highly time-consuming. The pioneering field-programmable gate array (FPGA)-based accelerator of Vina, known as Vina-FPGA, offers a high energy-efficiency approach to speed up the docking process. However, the computation modules in the Vina-FPGA design are not efficiently used. This is due to Vina exhibiting irregular behaviors in the form of nested loops with changing upper bounds and differing control flows. Fortunately, Vina employs the Monte Carlo iterative search method, which requires independent computations for different random initial inputs. This characteristic provides an opportunity to implement further parallel computation designs. To this end, this paper proposes Vina-FPGA2, an inter-module pipeline design for further accelerating Vina-FPGA. First, we use individual computational task (Task) independence by sequentially filling Tasks into computation modules. Then, we implement an inter-module pipeline parallel design by the Tag Checker module and architectural modifications, named Vina-FPGA2-Baseline. Next, to achieve resource-efficient hardware implementation, we describe it as an optimization problem and develop a reinforcement learning-based solver. Targeting the Xilinx UltraScale XCKU060 platform, this solver yields a more efficient implementation, named Vina-FPGA2-Enhanced. Finally, experiments show that Vina-FPGA2-Enhanced achieves an average 12.6× performance improvement over the central processing unit (CPU) and a 3.3× improvement over Vina-FPGA. Compared to Vina-GPU, Vina-FPGA2 achieves a 7.2× enhancement in energy efficiency. Shidi Tang, Ruiqi Chen 0001, Yanxiang Zhu |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2025 | EEVS: Redeploying Discarded Smartphones for Economic and Ecological Drug Molecules Virtual ScreeningabstractVirtual screening plays an indispensable role in the early stages of drug discovery, which utilizes high-throughput molecular docking to find potential drug candidates from vast databases. Virtual screening necessitates considerable computational resources to analyze tremendous compounds. However, the substantial demand for computational resources and the challenges in accessing high performance hardware hinders the development of drug discovery. This work introduces EEVS (Economic and Ecological Virtual Screening), an innovative framework that utilizes the computational capabilities of discarded smartphones for cost-effective and eco-friendly virtual screening. EEVS, with 16 discarded smartphones in this study, greatly reduces the construction cost of virtual screening, which is only 38.7%, 11.9%, and 26.9% of those of CPU, GPU, and FPGA implementations, respectively. Moreover, EEVS achieves a 4.05× improvement in screening speed while maintaining similar power and docking accuracy with CPU. When compared with GPU and FPGA, EEVS attains advantages of 4.93× in screening power and 1.08× in screening speed, respectively. Furthermore, we proposed the PCSA algorithm to further accelerate the screening speed of EEVS by a maximum of 33.6% while balancing various thermal dissipation requirements. To the best of our knowledge, this work is the first virtual screening framework that leverages discarded smartphones to accelerate drug discovery. Chuanzhao Zhang, Shidi Tang, Ruiqi Chen 0001, Yanxiang Zhu |
IEEE Trans. Sustain. Comput. | 4 |
| 2025 | FASE: An FPGA-Based Accelerator for Lightweight Sample Entropy With Monte Carlo SamplingabstractSample entropy (SampEn) is an algorithm within information entropy that enables effective analysis of biological signals. Due to the need for extensive similarity matching operations, the SampEn calculation process is time-consuming. Although a series of fast SampEn algorithms have been proposed, they remain time-intensive when processing large data volumes. Additionally, previous field-programmable gate array (FPGA)-based hardware accelerators designed for SampEn suffer from architectural design limitations, consuming substantial on-chip memory resources and operating at low frequencies. In this article, we propose FASE, an FPGA-based accelerator for lightweight sample entropy (LW-SampEn) with Monte Carlo (MC) sampling. The FASE design comprises two main parts: algorithm and hardware optimizations. On the algorithmic side, we introduce MC sampling into the merge-sort-based LW-SampEn algorithm, named MCLW-SampEn. MCLW-SampEn effectively reduces the computation load for large data volumes while maintaining algorithmic accuracy. For hardware, we first design efficient sorting and allocation modules to address boundary localization and load imbalance issues in previous accelerator designs. Then, we replicate the computation across the main phases to enable parallel processing. Finally, we deploy the design on the Pynq-Z2 board for validation. Experimental results show that the proposed MCLW-SampEn algorithm achieves an average speed up of$3\times $over the LW-SampEn algorithm, with accuracy losses kept within 0.5%. Compared to state-of-the-art (SOTA) designs, FASE achieves an average speed up of$12.8\times $while reducing power consumption by 89.3%. Ablation studies indicate that, for the same algorithm, FASE offers a$7.4\times $speedup over related FPGA designs. Yuanhang Li, Zhengyang Huang, Chao Chen 0042, Ruiqi Chen 0001, Bruno da Silva 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2024 | S-LGCN: Software-Hardware Co-Design for Accelerating LightGCNabstractGraph Convolutional Networks (GCNs) have garnered significant attention in recent years, finding applications across various domains, including recommendation systems, knowledge graphs, and biological prediction. One prominent GCN-based recommendation model, LightGCN, optimizes embeddings for final prediction through graph convolution operations, and has achieved outstanding performance in commodity recommendation and molecular property prediction. However, LightGCN suffers from suboptimal layer combination parameters and limited nonlinear modeling capabilities on the software side. On the hardware side, due to the irregularity of the aggregation phase of LightGCN, CPU and GPU executions are not efficient, and designing a accelerator will be constrained by transmission bandwidth and the efficiency of the sparse matrix multiplication kernel. In this paper, we optimize the layer combination parameters of LightGCN by Q-learning and add hardware-friendly activation function to enhance its nonlinear modeling capability. The optimized LightGCN not only performs well on the original dataset and some molecular prediction tasks, but also does not incur significant hardware overhead. Subsequently, we propose an efficient architecture to accelerate the inference of LightGCN to improve its adaptability to real-time tasks. Comparing S-LGCN to Intel(R) Xeon(R) Gold 5218R CPU and NVIDIA RTX3090 GPU, we observe that S-LGCN is 1576.4 × and 21.8 × faster with energy consumption reductions of 3211.6 × and 71.6 ×, respectively. Compared to FPGA-based accelerator, S-LGCN demonstrates 1.5-4.5 × lower latency and 2.03 × higher throughput. Ruiqi Chen 0001, Enhao Tang, Kun Wang 0005 |
DATE | 2 |
| 2024 | FPGA-Based Sparse Matrix Multiplication Accelerators: From State-of-the-Art to Future OpportunitiesabstractSparse matrix multiplication (SpMM) plays a critical role in high-performance computing applications, such as deep learning, image processing, and physical simulation. Field-Programmable Gate Arrays (FPGAs), with their configurable hardware resources, can be tailored to accelerate SpMMs. There has been considerable research on deploying sparse matrix multipliers across various FPGA platforms. However, the FPGA-based design of sparse matrix multipliers still presents numerous challenges. Therefore, it is necessary to summarize and organize the current work to provide a reference for further research. This article first introduces the computational method of SpMM and categorizes the different challenges of FPGA deployment. Following this, we introduce and analyze a variety of state-of-the-art FPGA-based accelerators tailored for SpMMs. In addition, a comparative analysis of these accelerators is performed, examining metrics including compression rate, throughput, and resource utilization. Finally, we propose potential research directions and challenges for further study of FPGA-based SpMM accelerators. Ruiqi Chen 0001, Bruno da Silva 0001 |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2024 | Graph-OPU: A Highly Flexible FPGA-Based Overlay Processor for Graph Neural NetworksabstractField-programmable gate arrays (FPGAs) are an ideal candidate for accelerating graph neural networks (GNNs). However, the FPGA redeployment process is time-consuming when updating or switching between diverse GNN models across different applications. Existing GNN processors eliminate the need for FPGA redeployment when switching between different GNN models. However, adapting matrix multiplication types by switching processing units decreases hardware utilization. In addition, the bandwidth of DDR limits further improvements in hardware performance. This article proposes a highly flexible FPGA-based overlay processor for GNN accelerations. Graph-OPU provides excellent flexibility and programmability for users, as the executable code of GNN models is automatically compiled and reloaded without requiring FPGA redeployment. First, we customize the compiler and instruction sets for the inference process of different GNN models. Second, we customize the datapath and optimize the data format in the microarchitecture to fully leverage the advantages of high bandwidth memory (HBM). Third, we design a unified matrix multiplication to handle both sparse-dense matrix multiplication (SpMM) and general matrix multiplication (GEMM), enhancing Graph-OPU performance. During Graph-OPU execution, the computational units are shared between SpMM and GEMM instead of being switched, which improves the hardware utilization. Finally, we implement a hardware prototype on the Xilinx Alveo U50 and test the mainstream GNN models using various datasets. Experimental results show that Graph-OPU achieves up to 1,654 \(\times\) and 63 \(\times\) speedup, as well as up to 5,305 \(\times\) and 422 \(\times\) energy efficiency boosts, compared to implementations on CPU and GPU, respectively. Graph-OPU outperforms state-of-the-art (SOTA) end-to-end overlay accelerators for GNN, reducing latency by an average of 1.36 \(\times\) and improving energy efficiency by 1.41 \(\times\) on average. Moreover, Graph-OPU exhibits an average 1.45 \(\times\) speed improvement in end-to-end latency over the SOTA GNN processor. Graph-OPU represents an in-depth study of an FPGA-based overlay processor for GNNs, offering high flexibility, speedup, and energy efficiency. Enhao Tang, Ruiqi Chen 0001, Hao Zhou 0008, Yuhanxiao Ma, Jun Yu 0010, Kun Wang 0005 |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2023 | Graph-OPU: An FPGA-Based Overlay Processor for Graph Neural NetworksabstractGraph Neural Networks (GNNs) have outstanding performance on graph-structured data and have been extensively accelerated by field-programmable gate array (FPGA) in various ways. However, existing accelerators significantly lack flexibility, especially in the following two aspects: 1) Many FPGA-based accelerators only support one GNN model. 2) The processes of re-synthesizing and bitstream re-generating are very time-consuming for new GNN models. To this end, we propose a highly integrated FPGA-based overlay processor for general GNN accelerations named Graph-OPU. Regarding the data structure and operation irregularity, we customize the instruction sets to support irregular operation patterns in the inference process of GNN models. Then, we customize our datapath and optimize the data format in the microarchitecture to take full advantage of high bandwidth memory (HBM). Moreover, we design the computation module to ensure a unified and fully-pipelined process of sparse matrix multiplication (SpMM) and general matrix multiplication (GEMM). Users can avoid the process of FPGA reconfiguration or RTL regeneration for the newly invented GNN models. We implement the hardware prototype on Xilinx Alveo U50 and test the mainstream GNN models with 9 datasets. Graph-OPU can achieve an average of 435× and 18× speedup, while 2013× and 109× better energy efficiency, compared with the Intel I7-12700KF processor and NVIDIA RTX3090 GPU, respectively. To the best of our knowledge, Graph-OPU is the first in-depth study on FPGA-based general processors for GNN acceleration with high speedup and energy efficiency. Ruiqi Chen 0001, Yuhanxiao Ma, Enhao Tang, Yanxiang Zhu, Jun Yu 0010, Kun Wang 0005 |
FPGA | 1 |
| 2023 | Graph-OPU: A Highly Integrated FPGA-Based Overlay Processor for Graph Neural NetworksabstractField-programmable gate array (FPGA) is an ideal candidate for accelerating graph neural networks (GNNs). However, FPGA reconfiguration is a time-consuming process when updating or switching between diverse GNN models across different applications. This paper proposes a highly integrated FPGA-based overlay processor for GNN accelerations. Graph-OPU provides excellent flexibility and software-like programmability for GNN end-users, as the executable code of GNN models are automatically compiled and reloaded without requiring FPGA reconfiguration. First, we customize the instruction sets for the inference qprocess of different GNN models. Second, we propose a microarchitecture ensuring a fully-pipelined process for GNN inference. Third, we design a unified matrix multiplication to process sparse-dense matrix multiplication and general matrix multiplication to increase the Graph-OPU performance. Finally, we implement a hardware prototype on the Xilinx Alveo U50 and test the mainstream GNN models using various datasets. Graph-OPU takes an average of only 2 minutes to switch between different GNN models, exhibiting average 128× speedup compared to related works. In addition, Graph-OPU outperforms state-of-the-art end-to-end overlay accelerators for GNN, reducing latency by an average of 1.36× and improving energy efficiency by an average of 1.41×. Moreover, Graph-OPU achieves up to 1654× and 63× speedup, as well as up to 5305× and 422× energy efficiency boosts, compared to implementations on CPU and GPU, respectively. To the best of our knowledge, Graph-OPU represents the first in-depth study of an FPGA-based overlay processor for GNNs, offering high flexibility, speedup, and energy efficiency. Ruiqi Chen 0001, Enhao Tang, Jun Yu 0010, Kun Wang 0005 |
FPL | 1 |
| 2023 | FPGA Accelerating Multi-Source Transfer Learning with GAT for Bioactivities of Ligands Targeting Orphan G Protein-Coupled ReceptorsabstractMachine learning has been used extensively in the bioactivity value (BAV) prediction of G Protein-Coupled Receptors (GPCR) targeting ligands. However, the performance of over 140 types of GPCR endogenous ligands, also called orphan GPCRs (oGPCRs), is still unsatisfactory due to the limited sample size. Also, current works are far from meeting the demand for fast inference time and energy efficiency. We propose the Multi-Source Transfer-Graph Attention Network (MSTL-GAT), as well as its FPGA-based accelerator. Firstly, we make use of the three ideal data sources for transfer learning, oGPCRs, experimentally validated GPCRs, and invalidated GPCRs similar to the former one. Secondly, we transform GPCRs from the SIMLEs format to graphics as the input of GAT to improve prediction accuracy. Moreover, we propose an FPGA-based accelerator tailored for the inference phase of MSTL-GAT. Finally, our experimental results show that MSTL-GAT remarkably improves the prediction of GPCRs ligand activity value compared with previous studies. On average, the two evaluation indexes we adopt, R2 and RMSE, improve by 34.76% and 13.16%, respectively. The proposed FPGA accelerator achieves 2.7× and 4.7× speedup, 29.7×, and 3.6× energy efficiency compared with works on GPU implementation and the state-of-the-art FPGA accelerator, respectively. Ruiqi Chen 0001, Jun Yu 0010, Kun Wang 0005 |
FPL | 1 |
| 2023 | g-BERT: Enabling Green BERT Deployment on FPGA via Hardware-Aware Hybrid PruningabstractTransformer-based models suffer from large num-ber of parameters and high inference latency, whose deployment are not green due to the potential environmental damage caused by high inference energy consumption. In addition, it is difficult to deploy such models on devices, especially on resource constrained devices such as FPGA. Various model pruning methods are proposed to shrink the model size and resource consumption, so as to fit the models on hardware. However, such methods often introduce floating point of operations (FLOPs) as an agent of hardware performance, which is not accurate. Furthermore, structural pruning methods are always in a single head-wise or layer-wise pattern, which fails to compress the models to the extreme. To resolve the above issues, we propose a green BERT deployment method on FPGA via hardware-aware and hybrid pruning, named g-BERT. Specifically, two hardware-aware metrics are introduced by High Level Synthesis (HLS) to evaluate the latency and power consumption of inference on FPGA, which can be optimized directly while pruning. Moreover, we simultaneously consider pruning of heads and full encoder layers. To efficiently find the optimal structure, g-BERT applies differentiable neural architecture search (NAS) with a special 0–1 loss function. Compared with the BERT-base, g-BERT achieves$2.1\times$speedup,$1.9\times$power consumption reduction and$1.8\times$model size reduction with comparable accuracy, on par with the state-of-the-art methods. Yueyin Bai, Hao Zhou 0008, Ruiqi Chen 0001, Kuangjie Zou, Jialin Cao, Jianli Chen, Jun Yu 0010, Kun Wang 0005 |
ICC | 3 |
| 2023 | Edge FPGA-based Onsite Neural Network TrainingabstractConjugate gradient (CG) is widely used in training sparse neural networks. However, CG, involving a large amount of sparse matrix and vector operations, cannot be efficiently implemented on resource-limited edge devices. In this paper, a high-performance and energy-efficient CG accelerator implemented on edge Field Programmable Gate Array is proposed for fast onsite neural networks training. According to the profiling, we propose a unified matrix multiplier that is compatible with the sparse and dense matrix. We also design a novel T-engine to handle transpose operation with the compressed sparse format. Experimental results show that our proposal outperforms the state-of-the-art FPGA work with a resource reduction of up to 41.3%. In addition, we achieve on average$10.2\times$and$2.0\times$speedup, while$10.1\times$and$3.5\times$better energy efficiency than implementations on CPU and GPU, respectively. Ruiqi Chen 0001, Yu Li 0003, Runzhou Zhang, Jun Yu 0010, Kun Wang 0005 |
ISCAS | 1 |
| 2023 | eSSpMV: An Embedded-FPGA-based Hardware Accelerator for Symmetric Sparse Matrix-Vector MultiplicationabstractSymmetric Sparse Matrix-Vector Multiplication (SSpMV) is a prevalent operation in numerous application domains (e.g., physical simulations, machine learning, and graph processing). Existing researches focus on the SSpMV implementation and its improvement on high-performance computing platforms but ignore the resource-limited edge platforms due to the main challenges: memory access overload and limited computing parallelism feasibility. To this end, this paper proposes an embedded-FPGA-based hardware accelerator for SSpMV, called eSSpMV. We first propose an optimized data format, named Symmetric Compressed Sparse Row (SCSR), to reduce memory consumption. Moreover, a fully-pipelined computation unit is proposed to be compatible with the optimized data format. Experimental results show that eSSpMV outperforms the state-of-the-art FPGA implementation for 2.9 x speedup, while still achieving a computing resource reduction of 39.3% and 32.3% for LUT and DSP, respectively. As for edge CPU and GPU implementations, eSSpMV achieves 9.3x speedup over CPU while acquiring 13.1 x better power latency product than GPU. Ruiqi Chen 0001, Yuhanxiao Ma, Jianli Chen, Jun Yu 0010, Kun Wang 0005 |
ISCAS | 1 |
| 2023 | Vina-FPGA: A Hardware-Accelerated Molecular Docking Tool With Fixed-Point Quantization and Low-Level ParallelismabstractMolecular docking (MD) is one of the core steps in the expensive and time-consuming process of drug design, which is basically an optimization problem based on scoring functions. AutoDock series MD software is widely accepted by academia and industry, among which AutoDock Vina (Vina) is the latest and most popular version due to its accuracy and relatively high speed. However, contrast to its prior version, i.e., AutoDock4, hardware acceleration approaches of Vina are rarely reported. In this article, we propose Vina-field-programmable gate array (FPGA), a hardware-accelerated Vina implementation with FPGA that exploits the low-level parallelism. First, the fixed-point quantization is analyzed and realized to accelerate the MD algorithm with a better energy efficiency in hardware. To boost the performance of the module-level computation, multiple in- module hardware pipelines have been designed and implemented. Besides, a strategy for fast accessing to block RAM (BRAM) is implemented by utilizing the layout of data, which brings four times memory access speed to the intermolecular and intramolecular energy computing modules. Under the same 140 ligand–receptor benchmarks, Vina-FPGA performs up to$6.9\times $(average$3.7\times$) faster than a state-of-the-art CPU does while consuming only 2.5% energy with similar docking accuracies. Compared to the GPU-accelerated implementation or Vina-GPU, the average energy consumption of Vina-FPGA is merely 45%. Qingde Lin, Ruiqi Chen 0001, Haimeng Qi, Mengru Lin, Yanxiang Zhu |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2022 | Biological Activity Prediction of GPCR-targeting Ligands on Heterogeneous FPGA-based AcceleratorsabstractIn the drug discovery process, the biological activity value (BAV) of G Protein-Coupled Receptors (GPCRs) targeting ligands is a large consideration. Past BAV prediction on CPU consumes tremendous time and power, yet there is rarely any related acceleration research. Therefore, this paper proposes a series of heterogeneous FPGA-based accelerators for well-performing algorithms to predict GPCRs ligands BAV. Communication delay is reduced by compressing the sparse matrix and directly coupling accelerators on the system BUS. Computation is accelerated by the remapping during the weight storage. Experimental results show that our FPGA accelerator implemented on Xilinx XCZU7EV performs 54.5× faster than CPU and 35.2× more energy-efficient than GPU. Ruiqi Chen 0001, Yuhanxiao Ma, Shaodong Zheng, Shizhen Huang, Chao Chen 0042, Jun Yu 0010, Kun Wang 0005 |
FCCM | 1 |