Kejia Shi

dblp:364/0129 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2025
0009-0001-8471-0995ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2025 ToMamba: Towards Token-Efficient Mamba Architecture on FPGA
abstract
The State Space Model (SSM), particularly the Mamba implementation, has demonstrated impressive capabilities across various domains. It offers a significant reduction in computational complexity compared to Transformers while achieving higher algorithm accuracy. However, the ineffectiveness of spatially unfolding the SSM layer leads to increased latency as sentence length grows, especially when being deployed on FPGA. Previous token reduction methods introduced in Transformers fail to maintain high performance in Mamba. Moreover, the dispersed outliers, complex model structure and variety of non-linear operators obstruct its efficient implementation on FPGA. To address these challenges, we propose ToMamba, the first algorithm-architecture co-design to optimize Mamba implementation. At the algorithmic level, ToMamba incorporates a novel progressive token merging algorithm with minimal hardware consumption and a hardware-aware fine-grained quantization strategy. On the hardware side, a dualflow systolic array is designed to unify convolution and matrix multiplication, supporting both weight stationary and output stationary dataflow. A fine-grained pipeline design is adopted for SSM computation to maximize hardware efficiency and enhance throughput. Furthermore, efficient hardware architecture and approximation method for nonlinear function units are proposed. To enable merging after the Mamba layer, ToMamba also adopts a dedicated data mapping scheme. Comprehensive evaluations across multiple benchmarks demonstrate that the token reduction method of ToMamba achieves 10% sparsity with only 0.25% accuracy loss, improving up to 16.89% in accuracy compared to previous methods. ToMamba hardware implementation on U280 FPGA achieves up to 636.00×/11.01×/1.39× speedup compared to Intel Xeon Platinum 8369B CPU, NVIDIA Tesla A100 GPU and ASIC platforms and 1280×/44.32× energy efficiency improvement compared to CPU and GPU platforms.
Kejia Shi, Yuhang Du, Jianli Chen, Jun Yu 0010, Kun Wang 0005
ICCAD1
2024 FNM-Trans: Efficient FPGA-based Transformer Architecture with Full N: M Sparsity
abstract
Transformer models have become popular in various AI applications due to their exceptional performance. However, their impressive performance comes with significant computing and memory costs, hindering efficient deployment of Transformer-based applications. Many solutions focus on leveraging sparsity in weight matrix and attention computation. However, previous studies fail to exploit unified sparse pattern to accelerate all three modules of Transformer (QKV generation, attention computation and FFN). In this paper, we propose FNM-Trans, an adaptable and efficient algorithm-hardware co-design aimed at optimizing all three modules of the Transformer by fully harnessing N : M sparsity. At the algorithm level, we fully explore the interplay of dynamic pruning with static pruning under high N : M sparsity. At the hardware level, we develop a dedicated hardware architecture featuring a custom computing engine and a softmax module, tailored to support varying levels of N : M sparsity. Experiment results show that, our algorithm optimizes accuracy by 11.03% under 2:16 attention sparsity and 4:16 weight sparsity, compared to other methods. Additionally, FNM-Trans achieves speedups of 27.13× and 21.24× over Intel i9-9900X and NVIDIA RTX 2080 Ti, respectively, and outpaces current FPGA-based Transformers by 1.88× to 36.51×.
Manting Zhang, Jialin Cao, Kejia Shi, Keqing Zhao, Genhao Zhang, Jun Yu 0010, Kun Wang 0005
DAC3
2024 Fitop-Trans: Maximizing Transformer Pipeline Efficiency through Fixed-Length Token Pruning on FPGA
abstract
Recent years have witnessed Transformers emerge as a groundbreaking innovation in the Natural Language Processing (NLP) field. Unlike Recurrent Neural Network (RNN) models, Transformers process sequences in parallel, boosting accuracy for longer sequences. However, Transformers face challenges with extended processing time. This is particularly due to the requirement of padding inputs to match the longest sentence in a batch, thereby increasing computational demands. In this paper, we present Fitop-Trans, the first algorithm-hardware co-optimized framework using Fixed-Length Token Pruning strategy while deploying Transformers on FPGA. At the algorithmic level, we propose Fixed-Length Token Pruning. It is a novel pruning method which can maximize hardware efficiency in attention computation, aimed at eliminating unimportant tokens before the first layer. On the hardware side, a token selector is designed for Fixed-Length Token Pruning, which minimizes off-chip memory traffic. In addition, a partitionable Systolic Array (SA) is adopted, which is capable of handling varying input lengths and maximizing Digital Signal Processor (DSP) resource utilization. Furthermore, a scheduling module is designed to optimize hardware resource allocation and enhance pipeline attention throughput. Experimental results reveal that our hardware design on FPGA achieves a speedup of $580 \times$ and $6.39 \times$ in latency compared to Intel Xeon Gold CPU and NVIDIA GeForce RTX 3090.
Kejia Shi, Manting Zhang, Keqing Zhao, Xiaoxing Wu, Yang Liu 0376, Jun Yu 0010, Kun Wang 0005
FPL1
2023 PP-Transformer: Enable Efficient Deployment of Transformers Through Pattern Pruning
abstract
Transformer models have been widely adopted in the field of Natural Language Processing (NLP) and Computer Vision (CV). However, the excellent performance of Transformers comes at the cost of heavy memory footprints and gigantic computing complexity. To deploy Transformers on resource constrained platforms, e.g., FPGA, diverse weight pruning strategies have been proposed. However, pattern pruning, as an alternative pruning method, is not well explored in the context of Transformers. In this paper, we propose PP-Transformer, a framework specifically designed to efficiently deploy Transformer models on FPGA using pattern pruning. At the algorithm level, we leverage pattern pruning, a coarse-grained structured pruning strategy, to reduce parameter storage. Meanwhile, we have developed a dedicated hardware architecture, featuring a custom computing engine tailored to support pattern pruning algorithm. Experimental results demonstrate that our algorithm achieves up to$2.26\times$reduction in parameter storage with acceptable accuracy degradation. Additionally, our hardware implementation exhibits$839.72\times$and$5.72\times$speedup in comparison to CPU and GPU implementations.
Jialin Cao, Xuanda Lin, Manting Zhang, Kejia Shi, Jun Yu 0010, Kun Wang 0005
ICCAD4