EDBT 2026 Demo / reviewers in the wild / expert
Yang Liu 0376
dblp:51/3710-376
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2026
0009-0001-0911-0666ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 3 first-author · 10 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Physically-aware Framework for Joint MBFF Synthesis with OPTICS-based DebankingabstractMBFF banking is a standard technique for clock power reduction in modern IC design, yet two systemic flaws in prior works limit its practical gains: geometric abstractions that ignore placement congestion, and open-loop workflows where late legalization failures nullify power savings. We propose a self-correcting framework that co-optimizes banking, placement, and debanking via three innovations: (1) Mahalanobis-distance clustering for placement-feasible MBFF formation; (2) a legalization-driven feedback loop with cost-aware debanking to recover unplaceable MBFFs; and (3) an OPTICS-based debanking that splits problematic MBFFs at highest-cost boundaries. Evaluated on ICCAD 2024 CAD Contest Problem B and large-scale benchmarks, our framework outperforms the 1st, 2nd, and 3rd place winners by 2.3%, 9.3%, and 4.0% in final weighted score, respectively. Benchao Zhu, Yang Liu 0376, Jianli Chen, Keren Zhu 0001 |
ACM Great Lakes Symposium on VLSI | 3 |
| 2026 | DIF-LUT Pro: An Automated Tool for Simple yet Scalable Approximation of Nonlinear Activation on FPGAabstractNonlinear activation plays an essential role in neural networks (NNs) for their generalization ability. However, implementing intricate mathematical operations on hardware platforms, including Field-Programmable Gate Arrays (FPGAs), presents significant challenges. Prior works based on piecewise functions or look-up table (LUT) have encountered difficulties in balancing precision requirements with fair hardware overhead and often necessitating complex manual interventions. To address these issues, this paper proposes DIF-LUT Pro, an automated tool for simple yet scalable approximation for various nonlinear activations on FPGA. Specifically, the proposed algorithm achieves self-adaptive hardware design oriented towards target precision, by piecewise linear matching to fit the function derivative roughly and range addressable LUT to offset the difference. Moreover, DIF-LUT Pro integrates the algorithm into an automated tool, allowing users to configure the customized interface and generate the corresponding hardware description language (HDL) code with a single click. Experimental results show that (1) DIF-LUT Pro features robust automation and fair generality, capable of generating equitable hardware designs under various user configurations across different FPGA platforms; (2) DIF-LUT Pro produces approximations that are simple yet effective, achieving competitive performance compared to previous expert-crafted designs. Furthermore, two detailed case studies demonstrate the efficient application of DIF-LUT Pro on NeRF and SEResnet, proving its practical value. Our source code is open-source and available at https://github.com/AdrianLiu00/DIF-LUT-Tool. Yang Liu 0376, Yu Li 0003, Ruiqi Chen 0001, Jun Yu 0010, Kun Wang 0005 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2025 | Deploying Diffusion Models with Scheduling Space Search and Memory Overflow Prevention Based on Graph OptimizationabstractIn recent years, Neural Networks developed rapidly to deal with tasks in the field of Computer Vision and Natural Language Process, etc. With the development of AI Generated Content, U-Net based Diffusion Models (DM) take image synthesis to new heights. U-Net performs the noise prediction of DM, the latency of which accounts for the majority of the end-to-end latency. Although FPGA has been proven to be a high performance platform to deploy NN, a series of facts still pose challenges for efficient U-Net based DMs deployment based on FPGA. The input vector length and type of the special function vary between different layers. The absence of model periodicity increases the granularity and complexity of operator scheduling. Skip-connection and residual connection inside model cause meta-data retaining in the memory, which is not conductive to avoiding memory overflow and decreasing total off-chip memory access. Hao Zhou 0008, Yang Liu 0376, Enhao Tang, Guohao Dai 0001, Yongpan Liu, Kun Wang 0005 |
ASP-DAC | 2 |
| 2025 | ATE-GCN: An FPGA-Based Graph Convolutional Network Accelerator with Asymmetrical Ternary QuantizationabstractTernary quantization can effectively simplify matrix multiplication, which is the primary computational operation in neural network models. It has shown success in FPGA-based accelerator designs for emerging models such as GAT and Transformer. However, existing ternary quantization methods can lead to substantial accuracy loss under certain weight distribution pat-terns, such as GCN. Furthermore, current FPGA-based ternary weight designs often focus on reducing resource consumption while neglecting full utilization of FPGA DSP blocks, limiting maximum performance. To address these challenges, we propose ATE-GCN, an FPGA-based asymmetrical ternary quantization GCN accelerator using a software-hardware co-optimization approach. First, we adopt an asymmetrical quantization strategy with specific interval divisions tailored to the bimodal distribution of GCN weights, reducing accuracy loss. Second, we design a unified processing element (PE) array on FPGA to support various matrix computation forms, optimizing FPGA resource usage while leveraging the benefits of cascade design and ternary quantization, significantly boosting performance. Finally, we implement the ATE-GCN prototype on the VCU118 FPGA board. The results show that ATE-GCN maintains an accuracy loss below 2%. Additionally, ATE-GCN achieves average performance improvements of$224.13\times$and$11.1\times$, with up to$898.82\times$and$69.9\times$energy consumption saving compared to CPU and GPU, respectively. Moreover, compared to state-of-the-art FPGA-based GCN accelerators, ATE-GCN improves DSP efficiency by 63% with an average latency reduction of 11%. Ruiqi Chen 0001, Shidi Tang, Yang Liu 0376, Yanxiang Zhu, Bruno da Silva 0001 |
DATE | 4 |
| 2025 | D-GCN: A Dynamic Pruning Accelerator for Deep Graph Convolutional Networks with Hybrid Dataflow
Hao Zhou 0008, Enhao Tang, Yang Liu 0376, Shidi Tang |
ACM Great Lakes Symposium on VLSI | 4 |
| 2024 | CSTrans-OPU: An FPGA-based Overlay Processor with Full Compilation for Transformer Networks via Sparsity ExplorationabstractA few overlay processors for transformer networks emerge to achieve reconfigurable architectures and dynamic instructions. However, these processors consistently neglect exploring network sparsity, while existing sparse accelerators inefficiently utilize resources with separate computation parts. Furthermore, mainstream compilers for instruction generation are intricate and demand significant engineering efforts. In this work, we propose CSTrans-OPU, an FPGA-based overlay processor with full compilation for transformer networks via sparsity exploration. Specifically, we customize a multi-precision processing element (PE) array with DSP-packing for unified computation format with full resource utilization. Additionally, the introduced sorting and computation mode selection modules make it possible to explore the token sparsity. Moreover, equipped with a user-friendly compiler, CSTrans-OPU enables model parsing, operation fusion, model quantization, instruction generation and reordering directly from model files. Experimental results show that CSTrans-OPU achieves 6.92-20.06× speedup and 182.48× higher energy efficiency compared with CPU, and 1.47-3.85× latency reduction with 4.63-52.53× better energy efficiency compared with GPU. Furthermore, we observe up to 4.28× better latency and 4.94× higher energy efficiency compared with previously customized accelerators, and can be up to 1.93× faster and 4.39× more energy efficient than FPGA processors. To the best of our knowledge, our CSTrans-OPU is the first overlay processor for transformer networks considering sparsity. Yueyin Bai, Keqing Zhao, Yang Liu 0376, Hao Zhou 0008, Xiaoxing Wu, Jun Yu 0010, Kun Wang 0005 |
DAC | 3 |
| 2024 | SDAcc: A Stable Diffusion Accelerator on FPGA via Software-Hardware Co-DesignabstractStable Diffusion has become one of the mainstream image synthesis algorithms. The mainstream computing platform for Stable Diffusion is GPU. However, the deployment of Stable Diffusion on GPU still faces the problems of power consumption. With dedicated hardware design and optimization, FPGA based Stable Diffusion accelerator can achieve better performance of energy efficiency. In this paper, we propose SDAcc for realizing efficient inference of Stable Diffusion on FPGA. SDAcc is 4.40× faster than CPU. Compared to GPU and CPU, SDAcc achieves 1.27× and 19.66× energy efficiency improvement, respectively. Hao Zhou 0008, Yang Liu 0376, Enhao Tang, Kun Wang 0005 |
FCCM | 2 |
| 2024 | Fitop-Trans: Maximizing Transformer Pipeline Efficiency through Fixed-Length Token Pruning on FPGAabstractRecent years have witnessed Transformers emerge as a groundbreaking innovation in the Natural Language Processing (NLP) field. Unlike Recurrent Neural Network (RNN) models, Transformers process sequences in parallel, boosting accuracy for longer sequences. However, Transformers face challenges with extended processing time. This is particularly due to the requirement of padding inputs to match the longest sentence in a batch, thereby increasing computational demands. In this paper, we present Fitop-Trans, the first algorithm-hardware co-optimized framework using Fixed-Length Token Pruning strategy while deploying Transformers on FPGA. At the algorithmic level, we propose Fixed-Length Token Pruning. It is a novel pruning method which can maximize hardware efficiency in attention computation, aimed at eliminating unimportant tokens before the first layer. On the hardware side, a token selector is designed for Fixed-Length Token Pruning, which minimizes off-chip memory traffic. In addition, a partitionable Systolic Array (SA) is adopted, which is capable of handling varying input lengths and maximizing Digital Signal Processor (DSP) resource utilization. Furthermore, a scheduling module is designed to optimize hardware resource allocation and enhance pipeline attention throughput. Experimental results reveal that our hardware design on FPGA achieves a speedup of $580 \times$ and $6.39 \times$ in latency compared to Intel Xeon Gold CPU and NVIDIA GeForce RTX 3090. Kejia Shi, Manting Zhang, Keqing Zhao, Xiaoxing Wu, Yang Liu 0376, Jun Yu 0010, Kun Wang 0005 |
FPL | 5 |
| 2024 | TransLib: An Extensible Graph-Aware Library Framework for Automated Generation of Transformer Operators on FPGA
Yang Liu 0376, Zexu Zhang, Jun Yu 0010, Kun Wang 0005 |
ICCAD | 1 |
| 2023 | DIF-LUT: A Simple Yet Scalable Approximation for Non-Linear Activation Function on FPGAabstractNon-linear activation function plays an essential role in neural networks (NNs) for their generalization ability. However, deploying the intricate mathematical operations on hardware platforms like Field-Programmable Gate Array (FPGA) turns out a great challenge. Prior works based on piecewise functions or look-up table (LUT) either involve complex manual operations or neglect hardware overhead. To this end, this paper proposes a simple yet scalable and effective approximation called DIF-LUT, which is applicable to various non-linear functions. Specifically, the proposed method can achieve accurate approximation by piecewise linear matching to fit the function derivative roughly and range addressable LUT to offset the difference. Moreover, self-adaptive mechanisms are applied to automatically minimize hardware cost in terms of different accuracies. The experiments show that compared to state-of-the-art methods, DIF-LUT costs 43.68% fewer LUTs and 70.8% fewer flip-flops (FFs) without any digital signal processor (DSP), while achieving 2.7x approximation accuracy at 554.1MHz on Xilinx Zynq UltraScale+. Yang Liu 0376, Xiaoming He 0004, Jun Yu 0010, Kun Wang 0005 |
FPL | 1 |