EDBT 2026 Demo / reviewers in the wild / expert
Tiandong Zhao
dblp:256/0442
· DBLP profile ↗
9ranked-venue papers
3as first author
6since 2021 · last 2025
0009-0000-5348-1715ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 3 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | METAL: A Memory-Efficient Transformer Architecture for Long-Context Inference on FPGAabstractTransformer-based models have shown remarkable proficiency in extensive tasks for natural language processing, which are facing the ever-increasing need of processing long-context inputs. However, the memory footprint of the self-attention mechanism grows quadratically with the context length and becomes the bandwidth and memory bottleneck. Existing accelerators are mainly tailored for short sequences and struggle to handle attention in long-context scenarios. While some works attempt to mitigate this memory overhead with algorithmic optimizations, they suffer from limited hardware efficiency due to the sequential execution of backward-dependent iterations and additional computations. To this end, this paper proposes METAL, an algorithm-architecture co-optimized approach to support long-context inference with minimized memory overhead. First, we propose a hardware-friendly attention algorithm that eliminates the data dependency across inner loops, enabling full pipelining while keeping nonincreasing on-chip memory requirements, regardless of context length. Second, we develop a unified PE array for different dataflows across the transformer block to consistently support the entire inference with efficient data reuse. Moreover, an advanced non-linear operation module is designed to properly match the throughput of PE arrays with maximized resource sharing. Experimental results show that METAL on the Xilinx U200 FPGA outperforms other FPGA accelerators by$1.23-2.89 \times$in normalized throughput and achieves up to$\mathbf{5 1. 0 \%}$BRAM savings across different input sequence lengths. Zicheng He, Shaoqiang Lu, Tiandong Zhao, Jinlong Yan, Lei He 0001 |
ASAP | 3 |
| 2025 | MambaOPU: An FPGA Overlay Processor for State-space-duality-based Mamba ModelsabstractState-space models (SSMs), such as Mamba, have emerged as a promising alternative to Transformers. However, the recently developed Mamba2, based on state space duality (SSD), is highly memorybound and suffers from limited computation efficiency. This inefficiency arises from its irregular broadcast element-wise multiplications and structured sparse computations. In this work, we propose MambaOPU, an FPGA overlay processor, to accelerate SSD. First, to reduce memory overhead, we introduce a software-hardware co-optimized operator fusion framework. Specifically, operator merging combines adjacent broadcast multiplication and summation operations into a single descriptor, while operator backward shifting embeds segment multiplication into subsequent operations. Both techniques shorten the computation path and improve computation efficiency. Second, to enhance sparse computation efficiency, we skip zero-region computations using a tensor-reorder-and-group algorithm combined with a sparse-predefined data fetcher. Additionally, since Mamba integrates linear operations with SSD, we develop a reconfigurable systolic array to improve data reuse across different computation modes. Extensive experiment results demonstrate that MambaOPU achieves up to $1812 \times$ and $880.79 \times$ higher normalized throughput and up to $12908 \times$ and $24.27 \times$ higher energy efficiency over Intel Xeon Gold 6348 CPU and NVIDIA A100 GPU, respectively. Shaoqiang Lu, Xuliang Yu, Tiandong Zhao, Siyuan Miao, Xinsong Sheng, Ting-Jung Lin, Lei He 0001 |
DAC | 3 |
| 2025 | MCoreOPU: An FPGA-based Multi-Core Overlay Processor for Transformer-based ModelsabstractTransformer-based models have achieved extensive success with increasingly large numbers of parameters and computations, for which many multi-core accelerators have been developed. Nevertheless, they suffer from limited throughput due to either low operating frequency or high communication overhead between cores. This article proposes an FPGA-based multi-core overlay processor, named MCoreOPU, to optimize intra-core computation and inter-core communication. First, we boost the operating frequency of the processing element (PE) array to double the rest of the processor to improve the intra-core throughput. Second, we develop on-chip synchronization routers to reduce off-chip memory traffic, where only the partial sum and maximum are communicated between cores rather than entire vectors for layer normalization and softmax. Moreover, we pipeline synchronization to reduce synchronization latency and develop a bypass of the interconnect bus to reduce the off-chip memory access latency. Finally, we optimize the multi-core model allocation and scheduling to minimize the inter-core communications and maximize the intra-core computation efficiency. The MCoreOPU is implemented in 8-bit fixed-point precision with four cores and four DDRs on the Xilinx U200 FPGA, where the PE array runs at 600 MHz while the rest runs at 300 MHz. Experimental results show that the throughput per MAC of MCoreOPU for BERT, ViT, GPT-2, and LLaMA inference is 1.31 \(\times\) –7.18 \(\times\) higher than other FPGA-based accelerators. Compared with the A100 GPU, the throughput per equivalent MAC efficiency is improved by 22.52 \(\times\) –27.12 \(\times\) . Shaoqiang Lu, Tiandong Zhao, Ting-Jung Lin, Rumin Zhang, Lei He 0001 |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2024 | ChatOPU: An FPGA-based Overlay Processor for Large Language Models with Unstructured SparsityabstractLarge language models (LLMs) have achieved notable success on many applications with increasingly tremendous parameters and computations. While hardware accelerators on Transformer-based models have been extensively studied, recent work mainly assumes structured model pruning and does not work well for unstructured sparsity that could lead to more parameter and computation reduction. The reason behind is that it is difficult to exploit data reuse from the unstructured sparsity, leaving hardware underutilized. This paper proposes ChatOPU, an FPGA-based overlay processor for LLMs, to support unstructured model pruning with better data reuse. First, we propose a new diagonal dataflow on a systolic array to obtain efficient data reuse for both sparse and dense matrix multiplication. Second, we develop efficient encoding and decoding for the sparse parameters to save off-chip memory traffic. Moreover, we boost the off-chip bandwidth utilization with pinned on-chip KV cache allocation and coalesced access throughout the LLM inference. Experimental results show that ChatOPU on Xilinx U200 FPGA outperforms GPU and other FPGA-based accelerators on token/s by 2.29× and 1.63× on LLMs with unstructured sparsity across different input and output sequence lengths. Tiandong Zhao, Shaoqiang Lu, Lei He 0001 |
ICCAD | 1 |
| 2023 | Token Packing for Transformers with Variable-Length InputsabstractTransformer-based models has achieved remarkable success in extensive tasks for natural language processing. To face the variable-length sentences in human language, popular deep learning frameworks rely on zero padding for batch processing, which introduces significant computation and memory overhead. Existing works attempt to eliminate padding redundancy but results in low hardware efficiency due to the mismatch between the variable shape of operations and fixed shape of processing elements (PEs). This paper proposes a reconfigurable systolic array with token packing in three folds to boost hardware efficiency. First, matrix multiplications for different tokens can be packed along the array columns to improve spatial efficiency. Meanwhile, for temporal efficiency, we develop a coarse-grained pipeline for attention, where stages can run on different parts of the array at the same time. We further exploit the masking redundancy in the Transformer decoder with runtime reconfigurable inter-PE connection and buffer switching. Applied to GPT, our FPGA design has achieved 1.16× higher normalized throughput and 1.94× better runtime MAC utilization over the state-of-the-art GPU performance for variable-length input sequences from GLUE and SQuAD dataset. Tiandong Zhao, Siyuan Miao, Shaoqiang Lu, Jialin Cao, Xiao Shi 0001, Kun Wang 0005, Lei He 0001 |
FPL | 1 |
| 2021 | Heterogeneous Dual-Core Overlay Processor for Light-Weight CNNsabstractConvolutional neural networks (CNNs) have achieved extensive success on miscellaneous artificial intelligence applications such as image classification and object detection. A plethora of models emerge with different operators and architectures, gradually shifting attention from accuracy to efficiency in terms of speed and power, since VGG-like architecture from early stage has significant redundancy. Light-weight CNNs are proposed to reduce computation complexity and parameter amount. MobileNets, one typical example of light-weight CNNs, adopt depthwise separable convolution, while others such as SqueezeNet alter model topology to spare computation power. Tiandong Zhao, Yunxuan Yu, Kun Wang 0005, Lei He 0001 |
FCCM | 1 |
| 2020 | Light-OPU: An FPGA-based Overlay Processor for Lightweight Convolutional Neural NetworksabstractLightweight convolutional neural networks (LW-CNNs) such as MobileNet, ShuffleNet, SqueezeNet, etc., have emerged in the past few years for fast inference on embedded and mobile system. However, lightweight operations limit acceleration potential by GPU due to their memory bounded nature and their parallel mechanisms that are not friendly to SIMD. This calls for more specific accelerators. In this paper, we propose an FPGA-based overlay processor with a corresponding compilation flow for general LW-CNN accelerations, called Light-OPU. Software-hardware co-designed Light-OPU reformulates and decomposes lightweight operations for efficient acceleration. Moreover, our instruction architecture considers sharing of major computation engine between LW operations and conventional convolution operations. This improves the run-time resource efficiency and overall power efficiency. Finally, Light-OPU is software programmable, since loading of compiled codes and kernel weights completes switch of targeted network without FPGA reconfiguration. Our experiments on seven major LW-CNNs show that Light-OPU achieves 5.5x better latency and 3.0x higher power efficiency on average compared with edge GPU NVIDIA Jetson TX2. Furthermore, Light-OPU has 1.3x to 8.4x better power efficiency compared with previous customized FPGA accelerators. To the best of our knowledge, Light-OPU is the first in-depth study on FPGA-based general processor for LW-CNNs acceleration with high performance and power efficiency, which is evaluated using all major LW-CNNs including the newly released MobileNetV3. Yunxuan Yu, Tiandong Zhao, Kun Wang 0005, Lei He 0001 |
FPGA | 2 |
| 2020 | OPU: An FPGA-Based Overlay Processor for Convolutional Neural NetworksabstractField-programmable gate array (FPGA) provides rich parallel computing resources with high energy efficiency, making it ideal for deep convolutional neural network (CNN) acceleration. In recent years, automatic compilers have been developed to generate network-specific FPGA accelerators. However, with more cascading deep CNN algorithms adapted by various complicated tasks, reconfiguration of FPGA devices during runtime becomes unavoidable when network-specific accelerators are employed. Such reconfiguration can be difficult for edge devices. Moreover, network-specific accelerator means regeneration of RTL code and physical implementation whenever the network is updated. This is not easy for CNN end users. In this article, we propose a domain-specific FPGA overlay processor, named OPU to accelerate CNN networks. It offers software-like programmability for CNN end users, as CNN algorithms are automatically compiled into executable codes, which are loaded and executed by OPU without reconfiguration of FPGA for switch or update of CNN networks. Our OPU instructions have complicated functions with variable runtimes but a uniform length. The granularity of instruction is optimized to provide good performance and sufficient flexibility, while reducing complexity to develop microarchitecture and compiler. Experiments show that OPU can achieve an average of 91% runtime multiplication and accumulation unit (MAC) efficiency (RME) among nine different networks. Moreover, for VGG and YOLO networks, OPU outperforms automatically compiled network-specific accelerators in the literature. In addition, OPU shows 5.35× better power efficiency compared with Titan Xp. For a real-time cascaded CNN networks scenario, OPU is 2.9× faster compared with edge computing GPU Jetson Tx2, which has a similar amount of computing resources. Yunxuan Yu, Tiandong Zhao, Kun Wang 0005, Lei He 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2020 | Uni-OPU: An FPGA-Based Uniform Accelerator for Convolutional and Transposed Convolutional NetworksabstractIn this article, we design the first full software/ hardware stack, called Uni-OPU, for an efficient uniform hardware acceleration of different types of transposed convolutional (TCONV) networks and conventional convolutional (CONV) networks. Specifically, a software compiler is provided to transform the computation of various TCONV, i.e., zero-inserting-based TCONV (zero-TCONV), nearest-neighbor resizing-based TCONV (NN-TCONV), and CONV layers into the same pattern. The compiler conducts the following optimizations: 1) eliminating up to 98.4% of operations in TCONV by making use of the fixed pattern of TCONV upsampling; 2) decomposing and reformulating TCONV and CONV into streaming parallel vector multiplication with a uniform address generation scheme and data flow pattern; and 3) efficient scheduling and instruction compilation to map networks onto a hardware processor. An instruction-based hardware acceleration processor is developed to efficiently speedup our uniform computation pattern with throughput up to 2.35 TOPS for the TCONV layer, consuming only 2.89 W dynamic power. We evaluate Uni-OPU on a benchmark set composed of six TCONV networks from different application fields. Extensive experimental results indicate that Uni-OPU is able to gain 1.45× to 3.68× superior power efficiency compared with state-of-the-art zero-TCONV accelerators. High acceleration performance is also achieved on NN-TCONV networks, the acceleration of which have not been explored before. In summary, we observe 1.90× and 1.63× latency reduction, as well as 15.04× and 12.43× higher power efficiency on zero-TCONV and NN-TCONV networks compared with Titan Xp GPU on average. To the best of our knowledge, ours is the first in-depth study to completely unify the computation process of zero-TCONV, NN-TCONV, and CONV layers. Yunxuan Yu, Tiandong Zhao, Kun Wang 0005, Lei He 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |