Siyuan Miao

dblp:360/8417 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2026
0009-0003-3813-141XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Multi-Level Interconnect Planning for Signal-Power-Thermal Integrity in 2.5D/3D Integration
abstract
Chiplets are a promising architecture for high-performance AI computing, but their package-level interconnects create a tightly coupled multiphysics problem involving signal delivery, power delivery, and heat dissipation. This challenge is compounded by the need to co-optimize the interposer and substrate, which have divergent design rules and performance sensitivities. To address these challenges, we propose MIP-SPT, a framework for multi-level interconnect planning. We introduce a hierarchical variable sched- uling strategy that decouples interposer and substrate variables, significantly reducing the search space. MIP-SPT then employs a multi-phase Bayesian optimization scheme to fully explore the streamlined design space. Crucially, our framework quantitatively models the effects of multiphysics coupling during planning to achieve rapid design closure. Experimental results show that our work reduces manufacturing cost by 22.4% compared to the baseline single-phase Bayesian optimization under equivalent design con- straints. In addition, it outperforms two existing works, lowering interconnect cost by 23.1% and 18.1%, respectively.
Siyuan Miao, Lingkang Zhu, Xiangqiao Meng, Wenkai Yang, Chengyu Zhu, Lei He 0001
ISPD1
2025 MambaOPU: An FPGA Overlay Processor for State-space-duality-based Mamba Models
abstract
State-space models (SSMs), such as Mamba, have emerged as a promising alternative to Transformers. However, the recently developed Mamba2, based on state space duality (SSD), is highly memorybound and suffers from limited computation efficiency. This inefficiency arises from its irregular broadcast element-wise multiplications and structured sparse computations. In this work, we propose MambaOPU, an FPGA overlay processor, to accelerate SSD. First, to reduce memory overhead, we introduce a software-hardware co-optimized operator fusion framework. Specifically, operator merging combines adjacent broadcast multiplication and summation operations into a single descriptor, while operator backward shifting embeds segment multiplication into subsequent operations. Both techniques shorten the computation path and improve computation efficiency. Second, to enhance sparse computation efficiency, we skip zero-region computations using a tensor-reorder-and-group algorithm combined with a sparse-predefined data fetcher. Additionally, since Mamba integrates linear operations with SSD, we develop a reconfigurable systolic array to improve data reuse across different computation modes. Extensive experiment results demonstrate that MambaOPU achieves up to $1812 \times$ and $880.79 \times$ higher normalized throughput and up to $12908 \times$ and $24.27 \times$ higher energy efficiency over Intel Xeon Gold 6348 CPU and NVIDIA A100 GPU, respectively.
Shaoqiang Lu, Xuliang Yu, Tiandong Zhao, Siyuan Miao, Xinsong Sheng, Ting-Jung Lin, Lei He 0001
DAC4
2025 C2OPU: Hybrid Compute-in-Memory and Coarse-Grained Reconfigurable Architecture for Overlay Processing of Transformers
abstract
Transformer-based models have shown huge success in natural language processing (NLP) with increasing model size and attention mechanism. However, this makes Von Neumann architecture based accelerators memory-bound such that the accelerators cannot leverage all the advantages of Transformer-based models. Although computing-in-memory (CIM) processors have emerged to tackle this problem through in-situ computing, the mismatch of computing patterns and low computing precision of CIM make it still challenging to accelerate Transformers. In this paper, we propose C2OPU, a hybrid dual-core processor to accelerate Transformers with hardware and software co-optimization. The dual-core architecture uses CIM arrays to accelerate weight-stationary vector-matrix multiplications, which accounts for the main computation complexity of Transformers. Meanwhile, a coarse-grained reconfigurable architecture (CGRA) is used to address the issues of computing pattern mismatch and low precision of the CIM. In addition, we propose an accuracy-bound workload allocation strategy, which considers non-ideal characteristics in analog computing, to balance throughput and accuracy. Furthermore, C2OPU provides a compiler to automatically determine optimal system configurations when Transformer model changes. Experimental results show that C2OPU achieves an average speedup of 145.41×, 4.73×, 4.70x and 3.85×, and 1.37x compared to CPU, GPU, Science23, Nature23, and VLSI24, respectively, on ten different Transformer models.
Siyuan Miao, Lingkang Zhu, Shaoqiang Lu, Jinming Lyu, Lei He 0001
FCCM1
2023 Token Packing for Transformers with Variable-Length Inputs
abstract
Transformer-based models has achieved remarkable success in extensive tasks for natural language processing. To face the variable-length sentences in human language, popular deep learning frameworks rely on zero padding for batch processing, which introduces significant computation and memory overhead. Existing works attempt to eliminate padding redundancy but results in low hardware efficiency due to the mismatch between the variable shape of operations and fixed shape of processing elements (PEs). This paper proposes a reconfigurable systolic array with token packing in three folds to boost hardware efficiency. First, matrix multiplications for different tokens can be packed along the array columns to improve spatial efficiency. Meanwhile, for temporal efficiency, we develop a coarse-grained pipeline for attention, where stages can run on different parts of the array at the same time. We further exploit the masking redundancy in the Transformer decoder with runtime reconfigurable inter-PE connection and buffer switching. Applied to GPT, our FPGA design has achieved 1.16× higher normalized throughput and 1.94× better runtime MAC utilization over the state-of-the-art GPU performance for variable-length input sequences from GLUE and SQuAD dataset.
Tiandong Zhao, Siyuan Miao, Shaoqiang Lu, Jialin Cao, Xiao Shi 0001, Kun Wang 0005, Lei He 0001
FPL2