Shaoqiang Lu

dblp:360/8136 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
11since 2021 · last 2026
0009-0002-8624-8666ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Mixture-of-Trees: Learning to Select and Weigh Reasoning Paths for Efficient LLM Inference
abstract
We introduce Mixture-of-Trees (MoT), a novel framework that integrates sparse expert activation with structured tree-based reasoning for efficient LLM inference. MoT employs a learned gating mechanism to selectively activate only the most relevant expert reasoning trees for each problem, where experts use models of varying capacities based on task complexity. The framework features three key innovations: (1) sparse expert activation through unified gating networks, (2) specialized expert trees that leverage domain-specific expertise while optimizing the quality-efficiency trade-off, and (3) collaborative debate mechanisms for conflicting solutions. Additionally, MoT includes a shared baseline tree with early stopping—activated experts perform lightweight validation and terminate early when confidence is high. Experiments across five benchmarks (GSM8K, MATH, AIME 2024, MMLU, HotpotQA) show that MoT achieves 2-7 percentage point accuracy improvements while reducing LLM calls by 37-40% compared to existing multi-path methods.
Yangbo Wei, Zhen Huang 0007, Shaoqiang Lu, Junhong Qian, Dongge Qin, Ting-Jung Lin, Wei W. Xing, Lei He 0001
AAAI3
2026 dLLM-OPU: An FPGA Overlay Processor for Accelerated Diffusion Large Language Models
abstract
Large Language Models (LLMs) are achieving unprecedented performance across diverse tasks, benefiting from autoregressive generation. However, this left-to-right decoding paradigm inherently limits contextual understanding quality. Diffusion-based LLMs (dLLMs) offer a promising alternative by iteratively refining sequences via denoising, enabling stronger bidirectional context modeling and improved generation quality. However, dLLMs face two main challenges: redundant computation and memory overhead in multi-step denoising, and excessive inference cost from over-denoising under fixed-step schedules. To address these issues, we propose dLLM-OPU, an FPGA overlay processor to accelerate dLLMs. Our solution features two key innovations: (1) a Region-Adaptive Caching for Dynamic Column Sparsity Framework that exploits temporal locality for selective recomputation without model retraining, and (2) a Token Entropy-based Early Stopping strategy that dynamically terminates the denoising process based on token-level convergence metrics. We implement these innovations through a specialized sparse processing element (PE) array that maximizes top-k sparsity utilization by minimizing idle cycles via row-column concatenation, complemented by an efficient cache management system that reduces memory access latency and a flexible entropybased decoding unit. Implemented on a U200 FPGA, dLLM-OPU achieves $2.2 \times-5.1 \times$ speedup and $7.6 \times-20.3 \times$ energy efficiency over RTX4090 in LLaDA.
Yangbo Wei, Shaoqiang Lu, Junhong Qian, Lei He 0001, Dongge Qin, Xiao Shi 0001
ASP-DAC2
2026 DFVG: A Heterogeneous Architecture for Speculative Decoding with Draft-on-FPGA and Verify-on-GPU
Shaoqiang Lu, Yangbo Wei, Junhong Qian, Dongge Qin, Shiji Gao, Yizhi Ding, Xiao Shi 0001, Lei He 0001
ASPLOS (2)1
2026 Harnessing Spatiotemporal Redundancy for Fast Diffusion Models on FPGA
Dongge Qin, Junhong Qian, Shaoqiang Lu, Yangbo Wei, Ruizhe Deng, Xiao Shi 0001, Longxing Shi, Lei He 0001
ISCAS3
2025 METAL: A Memory-Efficient Transformer Architecture for Long-Context Inference on FPGA
abstract
Transformer-based models have shown remarkable proficiency in extensive tasks for natural language processing, which are facing the ever-increasing need of processing long-context inputs. However, the memory footprint of the self-attention mechanism grows quadratically with the context length and becomes the bandwidth and memory bottleneck. Existing accelerators are mainly tailored for short sequences and struggle to handle attention in long-context scenarios. While some works attempt to mitigate this memory overhead with algorithmic optimizations, they suffer from limited hardware efficiency due to the sequential execution of backward-dependent iterations and additional computations. To this end, this paper proposes METAL, an algorithm-architecture co-optimized approach to support long-context inference with minimized memory overhead. First, we propose a hardware-friendly attention algorithm that eliminates the data dependency across inner loops, enabling full pipelining while keeping nonincreasing on-chip memory requirements, regardless of context length. Second, we develop a unified PE array for different dataflows across the transformer block to consistently support the entire inference with efficient data reuse. Moreover, an advanced non-linear operation module is designed to properly match the throughput of PE arrays with maximized resource sharing. Experimental results show that METAL on the Xilinx U200 FPGA outperforms other FPGA accelerators by$1.23-2.89 \times$in normalized throughput and achieves up to$\mathbf{5 1. 0 \%}$BRAM savings across different input sequence lengths.
Zicheng He, Shaoqiang Lu, Tiandong Zhao, Jinlong Yan, Lei He 0001
ASAP2
2025 MambaOPU: An FPGA Overlay Processor for State-space-duality-based Mamba Models
abstract
State-space models (SSMs), such as Mamba, have emerged as a promising alternative to Transformers. However, the recently developed Mamba2, based on state space duality (SSD), is highly memorybound and suffers from limited computation efficiency. This inefficiency arises from its irregular broadcast element-wise multiplications and structured sparse computations. In this work, we propose MambaOPU, an FPGA overlay processor, to accelerate SSD. First, to reduce memory overhead, we introduce a software-hardware co-optimized operator fusion framework. Specifically, operator merging combines adjacent broadcast multiplication and summation operations into a single descriptor, while operator backward shifting embeds segment multiplication into subsequent operations. Both techniques shorten the computation path and improve computation efficiency. Second, to enhance sparse computation efficiency, we skip zero-region computations using a tensor-reorder-and-group algorithm combined with a sparse-predefined data fetcher. Additionally, since Mamba integrates linear operations with SSD, we develop a reconfigurable systolic array to improve data reuse across different computation modes. Extensive experiment results demonstrate that MambaOPU achieves up to $1812 \times$ and $880.79 \times$ higher normalized throughput and up to $12908 \times$ and $24.27 \times$ higher energy efficiency over Intel Xeon Gold 6348 CPU and NVIDIA A100 GPU, respectively.
Shaoqiang Lu, Xuliang Yu, Tiandong Zhao, Siyuan Miao, Xinsong Sheng, Ting-Jung Lin, Lei He 0001
DAC1
2025 C2OPU: Hybrid Compute-in-Memory and Coarse-Grained Reconfigurable Architecture for Overlay Processing of Transformers
abstract
Transformer-based models have shown huge success in natural language processing (NLP) with increasing model size and attention mechanism. However, this makes Von Neumann architecture based accelerators memory-bound such that the accelerators cannot leverage all the advantages of Transformer-based models. Although computing-in-memory (CIM) processors have emerged to tackle this problem through in-situ computing, the mismatch of computing patterns and low computing precision of CIM make it still challenging to accelerate Transformers. In this paper, we propose C2OPU, a hybrid dual-core processor to accelerate Transformers with hardware and software co-optimization. The dual-core architecture uses CIM arrays to accelerate weight-stationary vector-matrix multiplications, which accounts for the main computation complexity of Transformers. Meanwhile, a coarse-grained reconfigurable architecture (CGRA) is used to address the issues of computing pattern mismatch and low precision of the CIM. In addition, we propose an accuracy-bound workload allocation strategy, which considers non-ideal characteristics in analog computing, to balance throughput and accuracy. Furthermore, C2OPU provides a compiler to automatically determine optimal system configurations when Transformer model changes. Experimental results show that C2OPU achieves an average speedup of 145.41×, 4.73×, 4.70x and 3.85×, and 1.37x compared to CPU, GPU, Science23, Nature23, and VLSI24, respectively, on ten different Transformer models.
Siyuan Miao, Lingkang Zhu, Shaoqiang Lu, Jinming Lyu, Lei He 0001
FCCM4
2025 MoE-OPU: An FPGA Overlay Processor Leveraging Expert Parallelism for MoE-based Large Language Models
abstract
The advent of Large Language Models (LLMs) like DeepSeek, empowered by the Mixture-of-Experts (MoE) architecture, has driven significant advancements across diverse applications. However, a critical challenge arises during inference: Only a small fraction of experts are activated, causing severe token allocation imbalances among experts. This inefficiency poses substantial storage and computational burdens on resource-constrained devices, exacerbated by the lack of optimization strategies that integrate expert usage-aware parameter pruning and parallel scheduling, ultimately leading to suboptimal resource utilization. To address these limitations, we propose MoE-OPU, an FPGA-based overlay processor that optimizes parallel MoE inference through three key innovations. First, we introduce N:M sparsity (1:4/2:4/4:8/6:8/8:8) in the MLP layers and mixed-precision quantization (BF16/FP8/INT4) guided by expert activation frequency, reducing the parameter size by up to 2.76× while maintaining model accuracy (only 1.53% average drop after fine-tuning). Second, a lightweight prediction network dynamically predicts next-layer "hot" experts by analyzing historical activation patterns and current hidden states, achieving an average prediction hit rate of 83.4%. Third, a reconfigurable multi-core architecture maximizes the utilization of HBM bandwidth via a systolic array that natively supports sparse and mixed-precision computations, coupled with parallel concatenation to balance compute and memory efficiency. Experimental results on a Xilinx V80 FPGA with the DeepSeek-V2-lite model demonstrate that MoE-OPU outperforms the NVIDIA A100 GPU, delivering a 6.78× higher token throughput. Compared to RTX 4090 and U200 FPGA, MoE-OPU achieves 13.37× and 7.85× improvements, respectively. These advancements highlight the potential of algorithm-hardware co-design for scalable deployment of MoE-based LLMs on edge devices.
Shaoqiang Lu, Yangbo Wei, Junhong Qian, Xiao Shi 0001, Lei He 0001
ICCAD1
2025 MCoreOPU: An FPGA-based Multi-Core Overlay Processor for Transformer-based Models
abstract
Transformer-based models have achieved extensive success with increasingly large numbers of parameters and computations, for which many multi-core accelerators have been developed. Nevertheless, they suffer from limited throughput due to either low operating frequency or high communication overhead between cores. This article proposes an FPGA-based multi-core overlay processor, named MCoreOPU, to optimize intra-core computation and inter-core communication. First, we boost the operating frequency of the processing element (PE) array to double the rest of the processor to improve the intra-core throughput. Second, we develop on-chip synchronization routers to reduce off-chip memory traffic, where only the partial sum and maximum are communicated between cores rather than entire vectors for layer normalization and softmax. Moreover, we pipeline synchronization to reduce synchronization latency and develop a bypass of the interconnect bus to reduce the off-chip memory access latency. Finally, we optimize the multi-core model allocation and scheduling to minimize the inter-core communications and maximize the intra-core computation efficiency. The MCoreOPU is implemented in 8-bit fixed-point precision with four cores and four DDRs on the Xilinx U200 FPGA, where the PE array runs at 600 MHz while the rest runs at 300 MHz. Experimental results show that the throughput per MAC of MCoreOPU for BERT, ViT, GPT-2, and LLaMA inference is 1.31 \(\times\) –7.18 \(\times\) higher than other FPGA-based accelerators. Compared with the A100 GPU, the throughput per equivalent MAC efficiency is improved by 22.52 \(\times\) –27.12 \(\times\) .
Shaoqiang Lu, Tiandong Zhao, Ting-Jung Lin, Rumin Zhang, Lei He 0001
ACM Trans. Reconfigurable Technol. Syst.1
2024 ChatOPU: An FPGA-based Overlay Processor for Large Language Models with Unstructured Sparsity
abstract
Large language models (LLMs) have achieved notable success on many applications with increasingly tremendous parameters and computations. While hardware accelerators on Transformer-based models have been extensively studied, recent work mainly assumes structured model pruning and does not work well for unstructured sparsity that could lead to more parameter and computation reduction. The reason behind is that it is difficult to exploit data reuse from the unstructured sparsity, leaving hardware underutilized. This paper proposes ChatOPU, an FPGA-based overlay processor for LLMs, to support unstructured model pruning with better data reuse. First, we propose a new diagonal dataflow on a systolic array to obtain efficient data reuse for both sparse and dense matrix multiplication. Second, we develop efficient encoding and decoding for the sparse parameters to save off-chip memory traffic. Moreover, we boost the off-chip bandwidth utilization with pinned on-chip KV cache allocation and coalesced access throughout the LLM inference. Experimental results show that ChatOPU on Xilinx U200 FPGA outperforms GPU and other FPGA-based accelerators on token/s by 2.29× and 1.63× on LLMs with unstructured sparsity across different input and output sequence lengths.
Tiandong Zhao, Shaoqiang Lu, Lei He 0001
ICCAD2
2023 Token Packing for Transformers with Variable-Length Inputs
abstract
Transformer-based models has achieved remarkable success in extensive tasks for natural language processing. To face the variable-length sentences in human language, popular deep learning frameworks rely on zero padding for batch processing, which introduces significant computation and memory overhead. Existing works attempt to eliminate padding redundancy but results in low hardware efficiency due to the mismatch between the variable shape of operations and fixed shape of processing elements (PEs). This paper proposes a reconfigurable systolic array with token packing in three folds to boost hardware efficiency. First, matrix multiplications for different tokens can be packed along the array columns to improve spatial efficiency. Meanwhile, for temporal efficiency, we develop a coarse-grained pipeline for attention, where stages can run on different parts of the array at the same time. We further exploit the masking redundancy in the Transformer decoder with runtime reconfigurable inter-PE connection and buffer switching. Applied to GPT, our FPGA design has achieved 1.16× higher normalized throughput and 1.94× better runtime MAC utilization over the state-of-the-art GPU performance for variable-length input sequences from GLUE and SQuAD dataset.
Tiandong Zhao, Siyuan Miao, Shaoqiang Lu, Jialin Cao, Xiao Shi 0001, Kun Wang 0005, Lei He 0001
FPL3