Shidi Tang

dblp:248/7983 · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
16since 2021 · last 2026
0000-0001-5493-7411ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 3 first-author · 13 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Equicore: Accelerating Clebsch-Gordan Tensor Product of Equivariant Neural Networks on FPGA
abstract
Equivariant neural networks (ENNs) are a powerful framework for modeling 3D geometric data in physical and biological systems. The Clebsch–Gordan tensor product (CGTP)—a core operation for preserving equivariance—remains the primary computational bottleneck in ENNs. Although Clebsch–Gordan (CG) coefficients exhibit pronounced structural sparsity, prior work has neither fully leveraged this property nor adopted hardware-friendly quantization, leading to limited efficiency. We present Equicore, a software–hardware co-design framework to accelerate CGTP in ENNs. Equicore introduces three key innovations: (1) a sparse-bypass strategy that exploits the CG structural sparsity together with a novel CG data format to pack the overlapping non-zeros, bypassing redundant data accesses and computations comparing to previous sparse solutions; (2) a merged-shift quantization strategy that enables full Int8 representation of irreps, weights, and CG coefficients using shift-only operations; and (3) a cascaded processing unit that tightly couples the FPGA hardware resources to achieve high operating frequency while supporting efficient sparse and quantized computation. Deployed on a AMD Virtex VCU128 platform, Equicore delivers up to 10.5× speedup and 17.4× energy-efficiency improvement over state-of-the-art GPU libraries and FPGA designs across diverse CGTP types in a benchmark of eleven ENN models.
Shidi Tang, Chuanzhao Zhang, Ruiqi Chen 0001, Yuxuan Lv, Bruno da Silva 0001
DATE1
2026 Breaking the BRAM Wall: Scalable Vina FPGA Acceleration via Distributed Grid Storage and Cross-Board Long-Ring Pipelines
abstract
AutoDock Vina (Vina), a gold standard for molecular docking, is hampered by computational expense. Previous FPGA hardware accelerators were fundamentally constrained by an on-chip memory bottleneck, where storing large, pre-computed energy grids consumes the majority of BRAM resources. This leads to imbalanced resource utilization, as the exhausted on-chip memory makes it impossible to further increase the degree of intra-node parallelism by instantiating more processing units. This paper introduces a novel, scalable multi-FPGA architecture that systematically removes this limitation.Our architecture’s innovation is a synergistic combination of three mechanisms. First, Distributed Grid Storage partitions the energy grid across all nodes to break the BRAM bottleneck. Second, a Cross-Board Long-Ring Pipeline creates a high-throughput dataflow for distributed energy calculations. Third, a dynamic intra-node scheduler unlocks massive fine-grained parallelism within each node. Together, these mechanisms create a powerful synergistic effect where adding nodes not only increases aggregate throughput but also enhances the performance of each individual node. This intrinsic amplification of per-node capacity is the direct driver of the system’s super-linear performance scaling.Implemented on a three-ZCU102 FPGA system and without sacrificing accuracy, our single-board normalized performance is 7.6× faster than the Vina-FPGA and 1.95× faster than the state-of-the-art Vina-FPGA-Cluster. Critically, the architecture demonstrates super-linear performance scaling: the three-board system achieves a 3.7× speedup over a single node, outperforming the ideal linear 3× speedup.
Ankun Tian, Shidi Tang, Ruiqi Chen 0001
DATE2
2026 BenDan: Benchmarking DPU performance on FPGAs
Ahmed Sadaqa, Yanxiang Zhu, Shidi Tang, Ruiqi Chen 0001, Bruno da Silva 0001
Integr.6
2026 Hessian-driven N:M sparsity and quantization co-optimization for edge device deployment
Minhua Ren, Zhihua Cai, Shidi Tang, Jianjun Li 0001
Integr.6
2026 FP8ApproxLib: An FPGA-based approximate multiplier library for 8-bit floating point
Ruiqi Chen 0001, Yangxintong Lyu, Shidi Tang, Jindong Li 0001, Yanxiang Zhu, Bruno da Silva 0001
J. Syst. Archit.4
2026 Diff-Acc: An Efficient FPGA Accelerator for Unconditional Diffusion Models
abstract
The diffusion model has achieved remarkable success in the era of Artificial Intelligence Generated Content (AIGC) across various tasks, such as image, video, text, material modeling, and molecular design. However, the diffusion model is computational intensive due to the long iteration of the reverse denoising process, which hinders its further advancement. Therefore, there is an urgent need to accelerate the diffusion model, especially in edge scenarios that require real-time computation. While researchers have made efforts to accelerate the diffusion model at the algorithm level using either efficient sampling or model quantization, they still suffer from accuracy degradation. More importantly, they have overlooked the hardware-level acceleration challenge. This work aims to bridge the gap by introducing Diff-Acc , the first FPGA accelerator for unconditional diffusion models with a novel step-wise quantization method that requires minimal calibration data to achieve the state-of-the-art (SOTA) PTQ quantization accuracy. Additionally, we adopt several hardware-oriented optimizations to reduce the computational overhead. At the architecture level, we fully analyze the computation flow of diffusion models and propose a novel architecture with group-wise parallelism to tackle the long iteration challenge. Besides, we decouple the data dependencies and adopt proper computational transformations at the micro-architecture level. Experiments on two unconditional diffusion models (DDIM and DDPM) with two image datasets (CIFAR-10 and ImageNet) demonstrate that our quantization method achieves the substantial improvements in image quality (FID: 6.67, sFID: 11.24) under 8-bit PTQ quantization. Compared with both server-based (Tesla V100 and Intel Xeon) and edge-based (Raspberry Pi 4 and Jetson Nano) platforms, Diff-Acc implemented on the Zynq UltraScale+ XCZU9EG FPGA demonstrates an up-to 12.5× energy efficiency. Particularly versus edge-based platforms, Diff-Acc achieves up to 10.26× and 1.97× performance improvements over CPU and GPU, respectively.
Shidi Tang, Ruiqi Chen 0001, Yuxuan Lv, Pengwei Zheng, He Li 0008
ACM Trans. Embed. Comput. Syst.1
2026 FANE: FPGA-Based FP8 Approximate Neural Network Engine
abstract
The 8-bit floating-point (FP8) format has gained growing interest in neural networks (NNs) for its superior dynamic range over traditional INT8. However, multiply-accumulate (MAC) operations remain a major source of power consumption during inference of NNs, which makes DSP-free design important, especially for edge FPGAs with few or no DSPs. Therefore, this brief presents FPGA-based FP8 approximate neural network engine (FANE), an FPGA-based approximate NN engine for FP8. We first introduce a novel approximation method that replaces the multiplications by linear additions. This approximate method reduces power consumption while maintaining high accuracy, outperforming the latest FP8 approximate multiplier by 53.15%. Based on this design, we construct an FP8 MAC unit and integrate it into both a convolution engine and a matrix–vector multiplication (MVM) unit. Finally, we integrate our design into a large language model (LLM). The result shows 61.5% higher efficiency (TOPS/W) than the previous design, demonstrating the superiority of FANE in terms of performance and power efficiency. The code of FANE is available on ourhttps://github.com/hanbao04/FANE-FPGA-based-FP8-Approximate-Neural-Network-Engine.git
Shidi Tang, Jingdong Li, Ahmed Sadaqa, Ruiqi Chen 0001, Bruno da Silva 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2025 ATE-GCN: An FPGA-Based Graph Convolutional Network Accelerator with Asymmetrical Ternary Quantization
abstract
Ternary quantization can effectively simplify matrix multiplication, which is the primary computational operation in neural network models. It has shown success in FPGA-based accelerator designs for emerging models such as GAT and Transformer. However, existing ternary quantization methods can lead to substantial accuracy loss under certain weight distribution pat-terns, such as GCN. Furthermore, current FPGA-based ternary weight designs often focus on reducing resource consumption while neglecting full utilization of FPGA DSP blocks, limiting maximum performance. To address these challenges, we propose ATE-GCN, an FPGA-based asymmetrical ternary quantization GCN accelerator using a software-hardware co-optimization approach. First, we adopt an asymmetrical quantization strategy with specific interval divisions tailored to the bimodal distribution of GCN weights, reducing accuracy loss. Second, we design a unified processing element (PE) array on FPGA to support various matrix computation forms, optimizing FPGA resource usage while leveraging the benefits of cascade design and ternary quantization, significantly boosting performance. Finally, we implement the ATE-GCN prototype on the VCU118 FPGA board. The results show that ATE-GCN maintains an accuracy loss below 2%. Additionally, ATE-GCN achieves average performance improvements of$224.13\times$and$11.1\times$, with up to$898.82\times$and$69.9\times$energy consumption saving compared to CPU and GPU, respectively. Moreover, compared to state-of-the-art FPGA-based GCN accelerators, ATE-GCN improves DSP efficiency by 63% with an average latency reduction of 11%.
Ruiqi Chen 0001, Shidi Tang, Yang Liu 0376, Yanxiang Zhu, Bruno da Silva 0001
DATE3
2025 FPGA-Based Approximate Multiplier for FP8
abstract
The 8-bit floating-point (FP8) data format has been increasingly adopted in neural network (NN) computations due to its superior dynamic range compared to traditional INT8. However, FP8-based multiplication, a core operation in NNs, still incurs significant power consumption. To address this issue, this paper presents an FPGA-based approximate multiplier design for FP8. Firstly, we conduct a bit-level analysis of the approximation method. Based on this analysis, we implement a fine-grained optimized design on mainstream FPGAs (AMD and Altera) using primitives and templates combined with physical layout constraints. Then, the accuracy and resource utilization of the FP8 approximate multiplier are evaluated and analyzed. The results indicate that, compared to previous FPGA-based 8-bit designs, our design achieves the minimal LUT consumption. Finally, we integrate the design into the inference phase of a representative NN model, demonstrating its excellent power efficiency. To the best of our knowledge, this is the first FPGA-based FP8 approximate multiplier design, which can serve as a benchmark for future designs and comparisons of FPGA-based low-precision floating-point approximate multipliers. The code of this work is available in our GitLab.
Ruiqi Chen 0001, Yangxintong Lyu, Yanxiang Zhu, Shidi Tang, Bruno da Silva 0001
FCCM6
2025 D-GCN: A Dynamic Pruning Accelerator for Deep Graph Convolutional Networks with Hybrid Dataflow
Hao Zhou 0008, Enhao Tang, Yang Liu 0376, Shidi Tang
ACM Great Lakes Symposium on VLSI5
2025 DiffDock-FPGA: A Hardware Accelerator for Molecular Docking with Customized Tensor Product Framework and Sparse-Aware Access Strategy
Chuanzhao Zhang, Shidi Tang
ACM Great Lakes Symposium on VLSI3
2025 Diff-DiT: Temporal Differential Accelerator for Low-bit Diffusion Transformers on FPGA
abstract
Diffusion Transformer (DiT) models have shown superior generative capabilities in image and video synthesis, yet their high computational cost during inference remains a critical bottleneck. Temporal differential computation offers a promising solution to low-bit quantization by exploiting the temporal similarity in activations. However, applying this technique to DiT’s Attention layers introduces substantial memory and computation overheads.In this paper, we present Diff-DiT, the first FPGA accelerator designed for low-bit DiT inference with differential computation. To overcome the unique challenges of DiT quantization and hardware acceleration, we propose: (1) an approximated differential attention (ADA) method that selectively approximates attention computations across time steps using a significance score, enabling low-bit on-chip execution while minimizing memory overhead; (2) an optimal cross-cast data accessing pattern with flexible data reuse to maximize computational intensity during matrix multiplications; and (3) a half-condition splitting (HCS) dataflow optimization and fine-grained pipelining to reduce the computation and memory access latency.Extensive experiments show that Diff-DiT outperforms NVIDIA V100 GPU by 1.39× in end-to-end throughput and 5.60× in energy efficiency. When compared with state-of-the-art diffusion model accelerators, Diff-DiT also achieves 2.81× and 2.77× improvements in throughput and energy efficiency, respectively. Code is available on GitHub1.
Shidi Tang, Pengwei Zheng, Ruiqi Chen 0001, Yuxuan Lv, Bruno da Silva 0001
ICCAD1
2025 Vina-FPGA2: a high-level parallelized hardware-accelerated molecular docking tool based on the inter-module pipeline
abstract
AutoDock Vina (Vina) is a widely adopted molecular docking tool, often regarded as a standard or used as a baseline in numerous studies. However, its computational process is highly time-consuming. The pioneering field-programmable gate array (FPGA)-based accelerator of Vina, known as Vina-FPGA, offers a high energy-efficiency approach to speed up the docking process. However, the computation modules in the Vina-FPGA design are not efficiently used. This is due to Vina exhibiting irregular behaviors in the form of nested loops with changing upper bounds and differing control flows. Fortunately, Vina employs the Monte Carlo iterative search method, which requires independent computations for different random initial inputs. This characteristic provides an opportunity to implement further parallel computation designs. To this end, this paper proposes Vina-FPGA2, an inter-module pipeline design for further accelerating Vina-FPGA. First, we use individual computational task (Task) independence by sequentially filling Tasks into computation modules. Then, we implement an inter-module pipeline parallel design by the Tag Checker module and architectural modifications, named Vina-FPGA2-Baseline. Next, to achieve resource-efficient hardware implementation, we describe it as an optimization problem and develop a reinforcement learning-based solver. Targeting the Xilinx UltraScale XCKU060 platform, this solver yields a more efficient implementation, named Vina-FPGA2-Enhanced. Finally, experiments show that Vina-FPGA2-Enhanced achieves an average 12.6× performance improvement over the central processing unit (CPU) and a 3.3× improvement over Vina-FPGA. Compared to Vina-GPU, Vina-FPGA2 achieves a 7.2× enhancement in energy efficiency.
Shidi Tang, Ruiqi Chen 0001, Yanxiang Zhu
Frontiers Inf. Technol. Electron. Eng.2
2025 EEVS: Redeploying Discarded Smartphones for Economic and Ecological Drug Molecules Virtual Screening
abstract
Virtual screening plays an indispensable role in the early stages of drug discovery, which utilizes high-throughput molecular docking to find potential drug candidates from vast databases. Virtual screening necessitates considerable computational resources to analyze tremendous compounds. However, the substantial demand for computational resources and the challenges in accessing high performance hardware hinders the development of drug discovery. This work introduces EEVS (Economic and Ecological Virtual Screening), an innovative framework that utilizes the computational capabilities of discarded smartphones for cost-effective and eco-friendly virtual screening. EEVS, with 16 discarded smartphones in this study, greatly reduces the construction cost of virtual screening, which is only 38.7%, 11.9%, and 26.9% of those of CPU, GPU, and FPGA implementations, respectively. Moreover, EEVS achieves a 4.05× improvement in screening speed while maintaining similar power and docking accuracy with CPU. When compared with GPU and FPGA, EEVS attains advantages of 4.93× in screening power and 1.08× in screening speed, respectively. Furthermore, we proposed the PCSA algorithm to further accelerate the screening speed of EEVS by a maximum of 33.6% while balancing various thermal dissipation requirements. To the best of our knowledge, this work is the first virtual screening framework that leverages discarded smartphones to accelerate drug discovery.
Chuanzhao Zhang, Shidi Tang, Ruiqi Chen 0001, Yanxiang Zhu
IEEE Trans. Sustain. Comput.3
2024 Vina-GPU 2.1: Towards Further Optimizing Docking Speed and Precision of AutoDock Vina and Its Derivatives
abstract
AutoDock Vina and its derivatives have established themselves as a prevailing pipeline for virtual screening in contemporary drug discovery. Our Vina-GPU method leverages the parallel computing power of GPUs to accelerate AutoDock Vina, and Vina-GPU 2.0 further enhances the speed of AutoDock Vina and its derivatives. Given the prevalence of large virtual screens in modern drug discovery, the improvement of speed and accuracy in virtual screening has become a longstanding challenge. In this study, we propose Vina-GPU 2.1, aimed at enhancing the docking speed and precision of AutoDock Vina and its derivatives through the integration of novel algorithms to facilitate improved docking and virtual screening outcomes. Building upon the foundations laid by Vina-GPU 2.0, we introduce a novel algorithm, namely Reduced Iteration and Low Complexity BFGS (RILC-BFGS), designed to expedite the most time-consuming operation. Additionally, we implement grid cache optimization to further enhance the docking speed. Furthermore, we employ optimal strategies to individually optimize the structures of ligands, receptors, and binding pockets, thereby enhancing the docking precision. To assess the performance of Vina-GPU 2.1, we conduct extensive virtual screening experiments on three prominent targets, utilizing two fundamental compound libraries and seven docking tools. Our results demonstrate that Vina-GPU 2.1 achieves an average 4.97-fold acceleration in docking speed and an average 342% improvement in EF1% compared to Vina-GPU 2.0.
Shidi Tang, Ji Ding 0002, Haitao Zhao 0004
IEEE ACM Trans. Comput. Biol. Bioinform.1
2023 Effectiveness Analysis of Multiple Initial States Simulated Annealing Algorithm, a Case Study on the Molecular Docking Tool AutoDock Vina
abstract
Simulated Annealing (SA) algorithm is not effective with large optimization problems for its slow convergence. Hence, several parallel Simulated Annealing (pSA) methods have been proposed, where the increase of searching threads can boost the speed of convergence. Although satisfactory solutions can be obtained by these methods, there is no rigorous mathematical analyses on their effectiveness. Thus, this article introduces a probabilistic model, on which a theorem about the effectiveness of multiple initial states parallel SA (MISPSA) has been proven. The theorem also demonstrates that the increasing parallelism in pSA algorithm with the reducing of search depth in each thread could obtain almost the same probability of finding the global optimal solution. We validated our theorem on AutoDock Vina, a widely used molecular docking tool with high accuracy and docking speed. AutoDock Vina uses a pSA strategy to find optimal molecular conformations. Under the premise that the total searching workload (i.e., thread number * iteration depth of each thread) remains unchanged, the docking accuracy from an aggressively parallelized SA searching method is almost the same or even better than those from the default exhaustiveness (parallelism degree) configuration of AutoDock Vina. Taking complex '1hnn' as an example,with the increase (125x) in the number of initial states (from 8 to 1000) and the decrease in the search depth for each thread (from 15540 to 124, or 1/125 of the original search depth), the mean energy is -7.80 and -7.94, while the mean RMSD is 3.4 and 3.14, respectively. The result also implies that a considerable speedup (in this case 125x in theory) can be obtained by a highly parallelized SA algorithm implementation.
Xingxing Zhou, Qingde Lin, Shidi Tang, Haifeng Hu 0004
IEEE ACM Trans. Comput. Biol. Bioinform.4