Yuxuan Lv

dblp:288/1167 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2026
0009-0002-9128-1479ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Equicore: Accelerating Clebsch-Gordan Tensor Product of Equivariant Neural Networks on FPGA
abstract
Equivariant neural networks (ENNs) are a powerful framework for modeling 3D geometric data in physical and biological systems. The Clebsch–Gordan tensor product (CGTP)—a core operation for preserving equivariance—remains the primary computational bottleneck in ENNs. Although Clebsch–Gordan (CG) coefficients exhibit pronounced structural sparsity, prior work has neither fully leveraged this property nor adopted hardware-friendly quantization, leading to limited efficiency. We present Equicore, a software–hardware co-design framework to accelerate CGTP in ENNs. Equicore introduces three key innovations: (1) a sparse-bypass strategy that exploits the CG structural sparsity together with a novel CG data format to pack the overlapping non-zeros, bypassing redundant data accesses and computations comparing to previous sparse solutions; (2) a merged-shift quantization strategy that enables full Int8 representation of irreps, weights, and CG coefficients using shift-only operations; and (3) a cascaded processing unit that tightly couples the FPGA hardware resources to achieve high operating frequency while supporting efficient sparse and quantized computation. Deployed on a AMD Virtex VCU128 platform, Equicore delivers up to 10.5× speedup and 17.4× energy-efficiency improvement over state-of-the-art GPU libraries and FPGA designs across diverse CGTP types in a benchmark of eleven ENN models.
Shidi Tang, Chuanzhao Zhang, Ruiqi Chen 0001, Yuxuan Lv, Bruno da Silva 0001
DATE4
2026 Diff-Acc: An Efficient FPGA Accelerator for Unconditional Diffusion Models
abstract
The diffusion model has achieved remarkable success in the era of Artificial Intelligence Generated Content (AIGC) across various tasks, such as image, video, text, material modeling, and molecular design. However, the diffusion model is computational intensive due to the long iteration of the reverse denoising process, which hinders its further advancement. Therefore, there is an urgent need to accelerate the diffusion model, especially in edge scenarios that require real-time computation. While researchers have made efforts to accelerate the diffusion model at the algorithm level using either efficient sampling or model quantization, they still suffer from accuracy degradation. More importantly, they have overlooked the hardware-level acceleration challenge. This work aims to bridge the gap by introducing Diff-Acc , the first FPGA accelerator for unconditional diffusion models with a novel step-wise quantization method that requires minimal calibration data to achieve the state-of-the-art (SOTA) PTQ quantization accuracy. Additionally, we adopt several hardware-oriented optimizations to reduce the computational overhead. At the architecture level, we fully analyze the computation flow of diffusion models and propose a novel architecture with group-wise parallelism to tackle the long iteration challenge. Besides, we decouple the data dependencies and adopt proper computational transformations at the micro-architecture level. Experiments on two unconditional diffusion models (DDIM and DDPM) with two image datasets (CIFAR-10 and ImageNet) demonstrate that our quantization method achieves the substantial improvements in image quality (FID: 6.67, sFID: 11.24) under 8-bit PTQ quantization. Compared with both server-based (Tesla V100 and Intel Xeon) and edge-based (Raspberry Pi 4 and Jetson Nano) platforms, Diff-Acc implemented on the Zynq UltraScale+ XCZU9EG FPGA demonstrates an up-to 12.5× energy efficiency. Particularly versus edge-based platforms, Diff-Acc achieves up to 10.26× and 1.97× performance improvements over CPU and GPU, respectively.
Shidi Tang, Ruiqi Chen 0001, Yuxuan Lv, Pengwei Zheng, He Li 0008
ACM Trans. Embed. Comput. Syst.4
2025 Diff-DiT: Temporal Differential Accelerator for Low-bit Diffusion Transformers on FPGA
abstract
Diffusion Transformer (DiT) models have shown superior generative capabilities in image and video synthesis, yet their high computational cost during inference remains a critical bottleneck. Temporal differential computation offers a promising solution to low-bit quantization by exploiting the temporal similarity in activations. However, applying this technique to DiT’s Attention layers introduces substantial memory and computation overheads.In this paper, we present Diff-DiT, the first FPGA accelerator designed for low-bit DiT inference with differential computation. To overcome the unique challenges of DiT quantization and hardware acceleration, we propose: (1) an approximated differential attention (ADA) method that selectively approximates attention computations across time steps using a significance score, enabling low-bit on-chip execution while minimizing memory overhead; (2) an optimal cross-cast data accessing pattern with flexible data reuse to maximize computational intensity during matrix multiplications; and (3) a half-condition splitting (HCS) dataflow optimization and fine-grained pipelining to reduce the computation and memory access latency.Extensive experiments show that Diff-DiT outperforms NVIDIA V100 GPU by 1.39× in end-to-end throughput and 5.60× in energy efficiency. When compared with state-of-the-art diffusion model accelerators, Diff-DiT also achieves 2.81× and 2.77× improvements in throughput and energy efficiency, respectively. Code is available on GitHub1.
Shidi Tang, Pengwei Zheng, Ruiqi Chen 0001, Yuxuan Lv, Bruno da Silva 0001
ICCAD4