EDBT 2026 Demo / reviewers in the wild / expert
Kairui Sun
dblp:398/9930
· DBLP profile ↗
4ranked-venue papers
3as first author
4since 2021 · last 2026
0009-0000-9880-4338ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | S2Mamba: An Efficient Mamba Accelerator With Word-Importance SSM SparsityabstractThe Mamba model, based on state space models (SSM), excels at long-sequence modeling with linear complexity, surpassing Transformers and emerging as a strong LLM candidate. However, various low-arithmetic-intensity operations in Mamba, along with the complex computational dependencies of the SSM, result in low efficiency on general-purpose computing platforms. Therefore, this paper introduces S$\rm ^{2}$Mamba, an efficient Mamba accelerator that leverages the sparsity of SSM. First, we develop a Mamba processing core (MPC) for low-arithmetic-intensity linear operations. By combining the linear-conv layer fusion scheme, the MPC facilitates fast depthwise convolution (DWC) computation, while continuous element-wise (EW) operations are designed to achieve efficient SSM computation. Second, we propose a word importance sparsity (WIS) algorithm that takes advantage of redundancy in natural language to filter 40.20% of unimportant words, leading to an average 37.25% reduction in the SSM computation. The saved computation is accelerated dynamically by the Dynamic Series Modules. Finally, we introduce a reconfigurable SiLU/Softplus unit (RSSU) for performing low-arithmetic-intensity nonlinear operations. The accelerator is implemented using a 28 nm CMOS process with an area of 17.56 mm$\rm ^{2}$. Extensive evaluations on representative benchmark tasks show that, S$\rm ^{2}$Mamba delivers 35.48–$97.11\times $speedup and 147.72–$332.47\times $energy efficiency gains over GPUs, while outperforming state-of-the-art Mamba accelerators with 1.52–$25.67\times $speedup and 1.10–$25.25\times $energy savings. Kairui Sun, Junhai Zhou, Zhongfeng Wang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2025 | UEDA: A Universal And Efficient Deformable Attention Accelerator For Various Vision TasksabstractDeformable attention (DA) provides an efficient and adaptive solution for capturing diverse object shapes, reducing computational complexity across multiple vision tasks. However, the various DA types lead to flexible matrix computation dimensions and complex attention pipelines. Moreover, the sampling operator's dynamic and irregular memory access significantly reduces data reuse and processing elements (PE) utilization, hindering DA from fully leveraging its low computational complexity. In this paper, we propose UEDA, a universal and efficient accelerator for DA. Specifically, a flexible 3D Folded Dimension Systolic Array (FDSA) is designed for efficient matrix multiplication computations with various dimensions, while achieving consistent high efficiency in supporting multiple networks. Secondly, a Reorganized Feature Map (RFM) sampling strategy is proposed to address parallel memory access conflicts, boosting the sampling module's processing rate by up to 4 times. Finally, an Inter-tile Cross Parallel (ITCP) dataflow is proposed to further hide the sampling module's latency, enhancing circuit throughput. The proposed UEDA is implemented on a Xilinx UltraScale+ FPGA. Experimental results show that UEDA achieves 11.7--14.8× speedup and 17.8--29.1x energy efficiency compared with GPU. Furthermore, we observe up to 2.19× better speedup and 27.8× higher energy efficiency compared to prior FPGA accelerators. Kairui Sun, Junhai Zhou, Zhongfeng Wang 0001 |
ASP-DAC | 1 |
| 2025 | LLM4GV: An LLM-Based Flexible Performance-Aware Framework for GEMM Verilog GenerationabstractAdvancements in AI have increased the demand for specialized AI accelerators, with design for general matrix multiplication (GEMM) module being crucial but time-consuming. While large language models (LLMs) show promise for automating GEMM design, challenges arise from GEMM's vast design space and performance requirements. Existing LLM-based frameworks for RTL code generation often lack flexibility and performance awareness. To overcome the challenges, we propose LLM4GV, a multi-agent LLM-based framework that integrates hardware optimization techniques (HOTs) and performance modeling, improving correctness and performance of the generated code over prior works. Dingyang Zou, Gaoche Zhang, Kairui Sun, Zhe Wen, Zhongfeng Wang 0001 |
DATE | 3 |
| 2025 | RETA-AD: A Reconfigurable and Efficient Transformer Accelerator for Autonomous DrivingabstractThe Transformer model is widely used in autonomous driving (AD) networks. However, diverse attention mechanisms and varying operators in AD networks lead to high computational complexity and substantial memory access requirements, presenting challenges to existing Transformer accelerators. To address these issues, we propose RETA-AD, a reconfigurable and efficient Transformer accelerator tailored for speeding up both normal attention (NA) and deformable attention (DA) computations within AD networks. First, a highly flexible 3-D Folded Dimension Systolic Array (FDSA) is developed, which is capable of processing small-width matrix multiplications (SWMMs) across various DA configurations, significantly improving hardware utilization and speed. Second, the computation of noncomputation-intensive (NCI) operators in AD networks is optimized, including a Reorganized Feature Map (RFM) sampling strategy to reduce the sampling time, and a pipeline reconfigurable (PR) Softmax module incorporating both coarse and fine-grained pipelines to support varying attention configurations with constantly high efficiency. Lastly, an Inter-tile Cross Parallel (ITCP) dataflow is designed to minimize on-chip storage requirements and hide the latency of NCI operations. The proposed RETA-AD is implemented on a Xilinx UltraScale+ FPGA development board. Experimental results show a 6.3–$8.9\times $speedup and a 13.7–$22.5\times $improvement in energy efficiency over the GPU A100 and the edge GPU Jetson AGX Xavier. Compared to previous FPGA-based Transformer accelerators, RETA-AD demonstrates a 1.1–$1.62\times $speedup and a 2.20–$2.62\times $improvement in energy efficiency. Kairui Sun, Zhongfeng Wang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |