VLDB 2026 Research / reviewers in the wild / expert
Yanxiang Zhu
dblp:235/9434
· DBLP profile ↗
9ranked-venue papers
0as first author
9since 2021 · last 2026
0000-0002-5014-3419ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BenDan: Benchmarking DPU performance on FPGAs
Ahmed Sadaqa, Yanxiang Zhu, Shidi Tang, Ruiqi Chen 0001, Bruno da Silva 0001 |
Integr. | 5 |
| 2026 | FP8ApproxLib: An FPGA-based approximate multiplier library for 8-bit floating point
Ruiqi Chen 0001, Yangxintong Lyu, Shidi Tang, Jindong Li 0001, Yanxiang Zhu, Bruno da Silva 0001 |
J. Syst. Archit. | 6 |
| 2025 | ATE-GCN: An FPGA-Based Graph Convolutional Network Accelerator with Asymmetrical Ternary QuantizationabstractTernary quantization can effectively simplify matrix multiplication, which is the primary computational operation in neural network models. It has shown success in FPGA-based accelerator designs for emerging models such as GAT and Transformer. However, existing ternary quantization methods can lead to substantial accuracy loss under certain weight distribution pat-terns, such as GCN. Furthermore, current FPGA-based ternary weight designs often focus on reducing resource consumption while neglecting full utilization of FPGA DSP blocks, limiting maximum performance. To address these challenges, we propose ATE-GCN, an FPGA-based asymmetrical ternary quantization GCN accelerator using a software-hardware co-optimization approach. First, we adopt an asymmetrical quantization strategy with specific interval divisions tailored to the bimodal distribution of GCN weights, reducing accuracy loss. Second, we design a unified processing element (PE) array on FPGA to support various matrix computation forms, optimizing FPGA resource usage while leveraging the benefits of cascade design and ternary quantization, significantly boosting performance. Finally, we implement the ATE-GCN prototype on the VCU118 FPGA board. The results show that ATE-GCN maintains an accuracy loss below 2%. Additionally, ATE-GCN achieves average performance improvements of$224.13\times$and$11.1\times$, with up to$898.82\times$and$69.9\times$energy consumption saving compared to CPU and GPU, respectively. Moreover, compared to state-of-the-art FPGA-based GCN accelerators, ATE-GCN improves DSP efficiency by 63% with an average latency reduction of 11%. Ruiqi Chen 0001, Shidi Tang, Yang Liu 0376, Yanxiang Zhu, Bruno da Silva 0001 |
DATE | 5 |
| 2025 | FPGA-Based Approximate Multiplier for FP8abstractThe 8-bit floating-point (FP8) data format has been increasingly adopted in neural network (NN) computations due to its superior dynamic range compared to traditional INT8. However, FP8-based multiplication, a core operation in NNs, still incurs significant power consumption. To address this issue, this paper presents an FPGA-based approximate multiplier design for FP8. Firstly, we conduct a bit-level analysis of the approximation method. Based on this analysis, we implement a fine-grained optimized design on mainstream FPGAs (AMD and Altera) using primitives and templates combined with physical layout constraints. Then, the accuracy and resource utilization of the FP8 approximate multiplier are evaluated and analyzed. The results indicate that, compared to previous FPGA-based 8-bit designs, our design achieves the minimal LUT consumption. Finally, we integrate the design into the inference phase of a representative NN model, demonstrating its excellent power efficiency. To the best of our knowledge, this is the first FPGA-based FP8 approximate multiplier design, which can serve as a benchmark for future designs and comparisons of FPGA-based low-precision floating-point approximate multipliers. The code of this work is available in our GitLab. Ruiqi Chen 0001, Yangxintong Lyu, Yanxiang Zhu, Shidi Tang, Bruno da Silva 0001 |
FCCM | 5 |
| 2025 | Vina-FPGA2: a high-level parallelized hardware-accelerated molecular docking tool based on the inter-module pipelineabstractAutoDock Vina (Vina) is a widely adopted molecular docking tool, often regarded as a standard or used as a baseline in numerous studies. However, its computational process is highly time-consuming. The pioneering field-programmable gate array (FPGA)-based accelerator of Vina, known as Vina-FPGA, offers a high energy-efficiency approach to speed up the docking process. However, the computation modules in the Vina-FPGA design are not efficiently used. This is due to Vina exhibiting irregular behaviors in the form of nested loops with changing upper bounds and differing control flows. Fortunately, Vina employs the Monte Carlo iterative search method, which requires independent computations for different random initial inputs. This characteristic provides an opportunity to implement further parallel computation designs. To this end, this paper proposes Vina-FPGA2, an inter-module pipeline design for further accelerating Vina-FPGA. First, we use individual computational task (Task) independence by sequentially filling Tasks into computation modules. Then, we implement an inter-module pipeline parallel design by the Tag Checker module and architectural modifications, named Vina-FPGA2-Baseline. Next, to achieve resource-efficient hardware implementation, we describe it as an optimization problem and develop a reinforcement learning-based solver. Targeting the Xilinx UltraScale XCKU060 platform, this solver yields a more efficient implementation, named Vina-FPGA2-Enhanced. Finally, experiments show that Vina-FPGA2-Enhanced achieves an average 12.6× performance improvement over the central processing unit (CPU) and a 3.3× improvement over Vina-FPGA. Compared to Vina-GPU, Vina-FPGA2 achieves a 7.2× enhancement in energy efficiency. Shidi Tang, Ruiqi Chen 0001, Yanxiang Zhu |
Frontiers Inf. Technol. Electron. Eng. | 5 |
| 2025 | EEVS: Redeploying Discarded Smartphones for Economic and Ecological Drug Molecules Virtual ScreeningabstractVirtual screening plays an indispensable role in the early stages of drug discovery, which utilizes high-throughput molecular docking to find potential drug candidates from vast databases. Virtual screening necessitates considerable computational resources to analyze tremendous compounds. However, the substantial demand for computational resources and the challenges in accessing high performance hardware hinders the development of drug discovery. This work introduces EEVS (Economic and Ecological Virtual Screening), an innovative framework that utilizes the computational capabilities of discarded smartphones for cost-effective and eco-friendly virtual screening. EEVS, with 16 discarded smartphones in this study, greatly reduces the construction cost of virtual screening, which is only 38.7%, 11.9%, and 26.9% of those of CPU, GPU, and FPGA implementations, respectively. Moreover, EEVS achieves a 4.05× improvement in screening speed while maintaining similar power and docking accuracy with CPU. When compared with GPU and FPGA, EEVS attains advantages of 4.93× in screening power and 1.08× in screening speed, respectively. Furthermore, we proposed the PCSA algorithm to further accelerate the screening speed of EEVS by a maximum of 33.6% while balancing various thermal dissipation requirements. To the best of our knowledge, this work is the first virtual screening framework that leverages discarded smartphones to accelerate drug discovery. Chuanzhao Zhang, Shidi Tang, Ruiqi Chen 0001, Yanxiang Zhu |
IEEE Trans. Sustain. Comput. | 5 |
| 2023 | Graph-OPU: An FPGA-Based Overlay Processor for Graph Neural NetworksabstractGraph Neural Networks (GNNs) have outstanding performance on graph-structured data and have been extensively accelerated by field-programmable gate array (FPGA) in various ways. However, existing accelerators significantly lack flexibility, especially in the following two aspects: 1) Many FPGA-based accelerators only support one GNN model. 2) The processes of re-synthesizing and bitstream re-generating are very time-consuming for new GNN models. To this end, we propose a highly integrated FPGA-based overlay processor for general GNN accelerations named Graph-OPU. Regarding the data structure and operation irregularity, we customize the instruction sets to support irregular operation patterns in the inference process of GNN models. Then, we customize our datapath and optimize the data format in the microarchitecture to take full advantage of high bandwidth memory (HBM). Moreover, we design the computation module to ensure a unified and fully-pipelined process of sparse matrix multiplication (SpMM) and general matrix multiplication (GEMM). Users can avoid the process of FPGA reconfiguration or RTL regeneration for the newly invented GNN models. We implement the hardware prototype on Xilinx Alveo U50 and test the mainstream GNN models with 9 datasets. Graph-OPU can achieve an average of 435× and 18× speedup, while 2013× and 109× better energy efficiency, compared with the Intel I7-12700KF processor and NVIDIA RTX3090 GPU, respectively. To the best of our knowledge, Graph-OPU is the first in-depth study on FPGA-based general processors for GNN acceleration with high speedup and energy efficiency. Ruiqi Chen 0001, Yuhanxiao Ma, Enhao Tang, Yanxiang Zhu, Jun Yu 0010, Kun Wang 0005 |
FPGA | 6 |
| 2023 | Vina-FPGA: A Hardware-Accelerated Molecular Docking Tool With Fixed-Point Quantization and Low-Level ParallelismabstractMolecular docking (MD) is one of the core steps in the expensive and time-consuming process of drug design, which is basically an optimization problem based on scoring functions. AutoDock series MD software is widely accepted by academia and industry, among which AutoDock Vina (Vina) is the latest and most popular version due to its accuracy and relatively high speed. However, contrast to its prior version, i.e., AutoDock4, hardware acceleration approaches of Vina are rarely reported. In this article, we propose Vina-field-programmable gate array (FPGA), a hardware-accelerated Vina implementation with FPGA that exploits the low-level parallelism. First, the fixed-point quantization is analyzed and realized to accelerate the MD algorithm with a better energy efficiency in hardware. To boost the performance of the module-level computation, multiple in- module hardware pipelines have been designed and implemented. Besides, a strategy for fast accessing to block RAM (BRAM) is implemented by utilizing the layout of data, which brings four times memory access speed to the intermolecular and intramolecular energy computing modules. Under the same 140 ligand–receptor benchmarks, Vina-FPGA performs up to$6.9\times $(average$3.7\times$) faster than a state-of-the-art CPU does while consuming only 2.5% energy with similar docking accuracies. Compared to the GPU-accelerated implementation or Vina-GPU, the average energy consumption of Vina-FPGA is merely 45%. Qingde Lin, Ruiqi Chen 0001, Haimeng Qi, Mengru Lin, Yanxiang Zhu |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2022 | Disclosing incoherent sparse and low-rank patterns inside homologous GPCR tasks for better modelling of ligand bioactivities
Chuangchuang Lan, Xuelin Ye, Jiale Deng, Wanqing Huang, Xueni Yang, Yanxiang Zhu, Haifeng Hu 0004 |
Frontiers Comput. Sci. | 7 |