VLDB 2026 Research / reviewers in the wild / expert
Kai Li 0024
dblp:181/2853-24
· DBLP profile ↗
13ranked-venue papers
1as first author
13since 2021 · last 2025
0000-0003-3251-931XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 1 first-author · 13 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | An MLA-LLM Hardware Acceleration with Tensor-Train Decomposition on Group Vector Systolic AcceleratorabstractLarge language models (LLMs) are both storageintensive and computation-intensive, posing significant challenges when deployed on resource-constrained hardware. As linear layers in LLMs are mainly resource consuming parts and KV cache affects the deployment of LLMs, this paper develops a posttraining process for converting multi-head attention (MHA) into multi-head latent attention (MLA) and a tensor-train decomposition (TTD) for LLMs with a further hardware implementation on FPGA. The post-training MLA process reduces the KV cache by 25 %. And TTD compression is applied to the linear layers in LLaMA2-7B models with compression ratios (CRs) for the whole network of$2.45 \times$. The compressed LLMs are further implemented on FPGA hardware within a highly efficient group vector systolic array (GVSA) architecture, which has DSP-shared parallel vector PEs for TTD inference, as well as optimized data communication in this accelerator. Experimental results show that the corresponding TTD based LLM accelerator implemented on FPGA achieves a peak speed of 64.20 tokens/s and$1.54 \times$reduction in first token delay for LLaMA2-7B models. Compared with NVIDIA A100 GPU, it achieves 46 % higher throughput. Sixiao Huang, Tintin Wang, Keyao Jiang, Kai Li 0024, Mingqiang Huang, Hao Yu 0001 |
ASAP | 5 |
| 2025 | A 20.98TOPS/W Energy-Efficient Binary BERT Model on Group Vector Systolic CIM AcceleratorabstractTransformer-based large language models (LLMs) impose significant bandwidth and compute challenges when deployed on edge devices. SRAM-based compute-in-memory (CIM) accelerators offer a promising solution to reduce data movement but are still limited by model size. This work develops a ternary weight splitting (TWS) binarization to obtain Brain-Floating-Point-16×INT1 (BF16×1-b) and INT8×INT1 (8-b×1-b) based transformers that exhibit competitive accuracy while significantly reducing model size compared to full precision counterparts. Then, a fully digital SRAM-based CIM accelerator is designed incorporating a bit-parallel SRAM macro within a highly efficient group vector systolic architecture, which can store one column of BERT-Tiny model with stationary systolic data reuse. The design in a 28nm technology only requires 2KB SRAM with an area of 2mm2. It achieves a throughput of 6.55TOPS and consumes a total power of 312.5mW and 221mW at 400MHz, resulting in a state-of-the-art area efficiency of 3.3TOPS/mm2and normalized energy efficiency of 20.98TOPS/W and 34.35TOPS/W for BF16×1-b and 8-b×1-b respectively on BERT-Tiny model, demonstrating a 10.25× improvement in area efficiency and a 2.23× improvement in energy efficiency compared to other state-of-the-art counterparts. Additionally, our proposed configuration compresses the model size by 32% with only a 0.5% accuracy loss on SST-2. Dingbang Liu, Qilong Chen, Jingyun Gu, Jiaqi Yang 0009, Kai Li 0024, Wei Mao 0002, Ngai Wong 0001, Chang Wen Chen, Hao Yu 0001 |
ISLPED | 6 |
| 2025 | EdgeLLM: A Highly Efficient CPU-FPGA Heterogeneous Edge Accelerator for Large Language ModelsabstractThe rapid advancements in artificial intelligence (AI), particularly the Large Language Models (LLMs), have profoundly affected our daily work and communication forms. However, it is still a challenge to deploy LLMs on resource-constrained edge devices (such as robots), due to the intensive computation requirements, heavy memory access, diverse operator types and difficulties in compilation. In this work, we proposed EdgeLLM to address the above issues. Firstly, focusing on the computation, we designed mix-precision processing element array together with group systolic architecture, that can efficiently support both FP$16\ast $FP16 for the MHA block (Multi-Head Attention) and FP$16\ast $INT4 for the FFN layer (Feed-Forward Network). Meanwhile specific optimization on log-scale structured weight sparsity, has been used to further increase the efficiency. Secondly, to address the compilation and deployment issue, we analyzed the whole operators within LLM models and developed a universal data parallelism scheme, by which all of the input and output features maintain the same data shape, enabling to process different operators without any data rearrangement. Then we proposed an end-to-end compiler to map the whole LLM model on CPU-FPGA heterogeneous system (AMD Xilinx VCU128 FPGA). The accelerator achieves$1.91\times $higher throughput and$7.55\times $higher energy efficiency than the commercial GPU (NVIDIA A100-SXM4-80G). When compared with state-of-the-art FPGA accelerator of FlightLLM, it shows 10-24% better performance in terms of HBM bandwidth utilization, energy efficiency and LLM throughput. Mingqiang Huang, Kai Li 0024, Haoxiang Peng, Yupeng Su, Hao Yu 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2025 | A 28-nm 135.19 TOPS/W Bootstrapped-SRAM Compute-in-Memory Accelerator With Layer-Wise Precision and SparsityabstractArtificial intelligence (AI) edge devices demand high energy efficiency as well as inference accuracy. SRAM-based compute-in-memory (CIM) accelerators have great potential for power reduction but still need to exploit higher throughput and better linearity performance. To meet edge-AI computing demands by CIM works, it is crucial to optimize algorithms and parameters for specific circuit systems to achieve hardware acceleration. This work firstly employs neural network search (NAS) method to find out the layer-wise optimized precisions and sparsities for convolutional neural networks (CNNs). Then, a 144-Kb charge-domain signed mixed-precision (2/4/8-bit) CIM accelerator employing bootstrapped SRAM cells with 9-transistors and 1-capacitor (9T1C) structure is proposed that incorporates a bit-level sparsity-aware analog-to-digital converter (ADC). This work not only achieves highly linear parallel accumulation operations to meet AI computing demands but also implements a hardware and software co-optimization system tailored to specific data characteristics. The design is verified on NAS-optimized networks VGG-16 and ResNet-18 using Cifar-10 dataset, which could achieve an equivalent accuracy at 4-bit of 68.68% while maintaining a high energy efficiency at 2-bit of 135.19TOPS/W by measurements. Wei Mao 0002, Dingbang Liu, Haoxiang Zhou, Fuyi Li, Kai Li 0024, Qiuping Wu, Jiaqi Yang 0009, Liuyang Zhang, Hao Yu 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2024 | LAMPS: A Layer-wised Mixed-Precision-and-Sparsity Accelerator for NAS-Optimized CNNs on FPGAabstractThe increasing model size and computation load of convolutional neural networks (CNN) pose a grand challenge to deploy CNN models on edge computing devices. To further improve performance without significant accuracy loss, this paper developed a neural architecture search (NAS) method to achieve a layer-wise mixed-precision-and-sparsity (LAMPS) CNN. However, this optimization cannot be fully utilized and directly mapped to existing AI accelerators due to the irregu- lar computation of sparse and multi-precision data. To tackle this challenge, this work proposed a LAMPS vector systolic accelerator and demonstrated state-of-the-art results. Experi- mental results show that the LAMPS accelerator on Xilinx ZCU102 achieves an average performance of 756.83 GOPS and 470.25 GOPS when accelerating the NAS-optimized VGG16 and Resnet18, respectively, leading to 1.3-6.0x speed-up over the state- of-the-art accelerators on FPGA. Shuxin Yang, Chenchen Ding, Mingqiang Huang, Kai Li 0024, Chenghao Li 0010, Zikun Wei, Sixiao Huang, Jingyao Dong, Liuyang Zhang, Hao Yu 0001 |
FCCM | 4 |
| 2023 | Multi-bit-width CNN Accelerator with Systolic-in-Systolic Dataflow and Single DSP Multiple Multiplication SchemeabstractMulti-bit-width neural network enlightens a promising method for high performance yet energy efficient edge computing due to its balance between software algorithm accuracy and hardware efficiency. To date, FPGA has been one of the core hardware platforms for deploying various neural networks. However, it is still difficult to fully make use of the dedicated digital signal processing (DSP) blocks in FPGA for accelerating the multi-bit-width network. In this work, we develop state-of-the-art multi-bit-width convolutional neural network accelerator with novel systolic-in-systolic type of dataflow and single DSP multiple multiplication (SDMM) INT2/4/8 execution scheme. Multi-level optimizations have also been adopted to further improve the performance, including group-vector systolic array for maximizing the circuit efficiency as well as minimizing the systolic delay, and differential neural architecture search (NAS) method for the high accuracy multi-bit-width network generation. The proposed accelerator has been practically deployed on Xilinx ZCU102 with accelerating NAS optimized VGG16 and Resnet18 networks as case studies. Average performance on accelerating the convolutional layer in VGG16 and Resnet18 is 1289GOPs and 1155GOPs, respectively. Throughput for running the full multi-bit-width VGG16 network is 870.73 GOPS at 250MHz, which has exceeded all of previous CNN accelerators on the same platform. Mingqiang Huang, Yucen Liu, Sixiao Huang, Kai Li 0024, Qiuping Wu, Hao Yu 0001 |
FPGA | 4 |
| 2022 | An Energy-Efficient Bit-Split-and-Combination Systolic Accelerator for NAS-Based Multi-Precision Convolution Neural NetworksabstractOptimized convolutional neural network (CNN) models and energy-efficient hardware design are of great importance in edge-computing applications. The neural architecture search (NAS) methods are employed for CNN model optimization with multi-precision networks. To satisfy the computation requirements, multi-precision convolution accelerators are highly desired. The existing high-precision-split (HPS) designs reduce the additional logics for reconfiguration while resulting in low throughput for low precisions. The low-precision-combination (LPC) designs improve the low-precision throughput with large hardware cost. In this work, a bit-split-and-combination (BSC) systolic accelerator is proposed to overcome the bottlenecks. Firstly, BSC-based multiply-accumulate (MAC) unit is designed to support multi-precision computation operations. Secondly, multi-precision systolic dataflow is developed with improved data-reuse and transmission efficiency. The proposed work is designed by Chisel and synthesized in 28-nm process. The BSC MAC unit achieves maximum 2.40× and 1.64× energy efficiency than HPS and LPC units, respectively. Compared with published accelerator designs Gemmini, Bit-fusion and Bit-serial, the proposed accelerator achieves up to 2.94 × area efficiency and 6.38 × energy-saving performance on the multi-precision VGG-16, ResNet-18 and LeNet-5 benchmarks. Liuyao Dai, Gengbin Huang, Junzhuo Zhou, Kai Li 0024, Wei Mao 0002, Hao Yu 0001 |
ASP-DAC | 6 |
| 2022 | A Precision-Scalable Energy-Efficient Bit-Split-and-Combination Vector Systolic Accelerator for NAS-Optimized DNNs on EdgeabstractOptimized model and energy-efficient hardware are both required for deep neural networks (DNNs) in edge-computing area. Neural architecture search (NAS) methods are employed for DNN model optimization with resulted multi-precision networks. Previous works have proposed low-precision-combination (LPC) and high-precision-split (HPS) methods for multi-precision networks, which are not energy-efficient for precision-scalable vector implementation. In this paper, a bit-split-and-combination (BSC) based vector systolic accelerator is developed for a precision-scalable energy-efficient convolution on edge. The maximum energy efficiency of the proposed BSC vector processing element (PE) is up to 1.95× higher in 2-bit, 4-bit and 8-bit operations when compared with LPC and HPS PEs. Further with NAS optimized multi-precision CNN networks, the averaged energy efficiency of the proposed vector systolic BSC PE array achieves up to 2.18× higher in 2-bit, 4-bit and 8-bit operations than that of LPC and HPS PE arrays. Kai Li 0024, Junzhuo Zhou, Junyi Luo, Zhengke Yang, Shuxin Yang, Wei Mao 0002, Mingqiang Huang, Hao Yu 0001 |
DATE | 1 |
| 2022 | A High Throughput Multi-bit-width 3D Systolic Accelerator for NAS Optimized Deep Neural Networks on FPGAabstractNeural architecture search (NAS) optimized multi-bit-width convolutional neural network (CNN) maintains the balance between network performance and efficiency, thus enlightening a promising method for accurate yet energy-efficient edge computing. In this work, we propose a high throughput three-dimensional (3D) systolic accelerator for NAS optimized CNNs, in which the input feature matrix, weight matrix and output feature matrix are delivering vertically, horizontally and perpendicularly through the systolic array respectively. With 3D systolic data flow, the processing time and logic resources consumption can be both reduced compared to the classical non-stationary systolic array. Besides, Booth-based multi-bit-width (INT2/4/8) multiply-add-accumulation (MAC) unit is developed within the 3D systolic accelerator. Deployed on FPGA platform Xilinx ZCU102, peek performance of the convolutional layer can reach as high as 2775 GOPS for INT2, 1650 GOPS for INT4, and 816 GOPS for INT8 respectively. The average performance on accelerating full NAS VGG16 network is 647 GOPS. Mingqiang Huang, Yucen Liu, Shuxin Yang, Kai Li 0024, Junyi Luo, Zhengke Yang, Qiufeng Li, Hao Yu 0001, Changhai Man |
FPGA | 5 |
| 2022 | A High Performance Multi-Bit-Width Booth Vector Systolic Accelerator for NAS Optimized Deep Learning Neural NetworksabstractMulti-bit-width convolutional neural network (CNN) maintains the balance between network accuracy and hardware efficiency, thus enlightening a promising method for accurate yet energy-efficient edge computing. In this work, we develop state-of-the-art multi-bit-width accelerator for NAS Optimized deep learning neural networks. To efficiently process the multi-bit-width network inferencing, multi-level optimizations have been proposed. Firstly, differential Neural Architecture Search (NAS) method is adopted for the high accuracy multi-bit-width network generation. Secondly, hybrid Booth based multi-bit-width multiply-add-accumulation (MAC) unit is developed for data processing. Thirdly, vector systolic array is proposed for effectively accelerating the matrix multiplications. With vector-style systolic dataflow, both the processing time and logic resources consumption can be reduced when compared with the classical systolic array. Finally, The proposed multi-bit-width CNN acceleration scheme has been practically deployed on FPGA platform of Xilinx ZCU102. Average performance on accelerating the full NAS optimized VGG16 network is 784.2 GOPS, and peek performance of the convolutional layer can reach as high as 871.26 GOPS for INT8, 1676.96 GOPS for INT4, and 2863.29 GOPS for INT2 respectively, which is among the best results in previous CNN accelerator benchmarks. Mingqiang Huang, Yucen Liu, Changhai Man, Kai Li 0024, Wei Mao 0002, Hao Yu 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2022 | An Energy-Efficient Mixed-Bitwidth Systolic Accelerator for NAS-Optimized Deep Neural NetworksabstractOptimized deep neural network (DNN) models and energy-efficient hardware designs are of great importance in edge-computing applications. The neural architecture search (NAS) methods are employed for DNN model optimization with mixed-bitwidth networks. To satisfy the computation requirements, mixed-bitwidth convolution accelerators are highly desired for low-power and high-throughput performance. There exist several methods to support mixed-bitwidth multiply-accumulate (MAC) operations in DNN accelerator designs. The low-bitwidth-combination (LBC) method improves the low-bitwidth throughput with a large hardware cost. The high-bitwidth-split (HBS) method minimizes the additional logic gates for configuration. However, the throughput performance in the low-bitwidth mode is poor. In this work, a bit-split-and-combination (BSC) systolic accelerator is proposed. The BSC-based MAC unit is designed to support mixed-bitwidth operations with the best overall performance. Besides, interprocessing element (PE) systolic and intra-PE paralleled dataflow not only improves throughput performance in mixed-bitwidth modes, but also saves power performance for data transmission. The proposed work is designed and synthesized in a 28-nm process. The BSC MAC unit achieves a maximum $2.08\times $ and $1.75\times $ energy efficiency improvement than the HBS and LBC unit, respectively. Compared with the state-of-the-art accelerators, the proposed work also achieves excellent energy-efficient performance with 20.02, 23.55, and 30.17 TOPS/W on mixed-bitwidth VGG-16, ResNet-18, and LeNet-5 benchmarks at 0.6 V, respectively. Wei Mao 0002, Liuyao Dai, Kai Li 0024, Laimin Du, Shaobo Luo, Mingqiang Huang, Hao Yu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2022 | A Configurable Floating-Point Multiple-Precision Processing Element for HPC and AI Converged ComputingabstractThere is an emerging need to design configurable accelerators for the high-performance computing (HPC) and artificial intelligence (AI) applications in different precisions. Thus, the floating-point (FP) processing element (PE), which is the key basic unit of the accelerators, is necessary to meet multiple-precision requirements with energy-efficient operations. However, the existing structures by using high-precision-split (HPS) and low-precision-combination (LPC) methods result in low utilization rate of the multiplication array and long multiterm processing period, respectively. In this article, a configurable FP multiple-precision PE design is proposed with the LPC structure. Half precision, single precision, and double precision are supported. The 100% multiplier utilization rate of the multiplication array for all precisions is achieved with improved speed in the comparison and summation process. The proposed design is realized in a 28-nm process with 1.429-GHz clock frequency. Compared with the existing multiple-precision FP methods, the proposed structure achieves 63% and 88% area-saving performance for FP16 and FP32 operations, respectively. The$4\times $and$20\times $maximum throughput rates are obtained when compared with fixed FP32 and FP64 operations. Compared with the previous multiple-precision PEs, the proposed one achieves the best energy-efficiency performance with 975.13 GFLOPS/W. Wei Mao 0002, Kai Li 0024, Liuyao Dai, Xinang Xie, He Li 0008, Longyang Lin, Hao Yu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2021 | A Reconfigurable Multiple-Precision Floating-Point Dot Product Unit for High-Performance ComputingabstractThere is an emerging need to optimize floating-point (FP) dot product units (DPU) for high-performance scientific computing as well as training deep learning models. Due to different precision requirements of applications, a reconfigurable multiple-precision DPU operation can largely reduce the cost of area and power. However, the existing methods could result in redundant bits for unit multipliers, but also leave idle hardware resources for the operations in different precisions. In this paper, a reconfigurable multiple-precision FP DPU design is proposed for high-performance computing (HPC) applications. The FP DPU can be reconfigured as follows. A bit-partitioning method is provided to minimize the redundant bits with a configurable mixed-precision multiplier for three-mode operations: 20 half-precision Dot Product (DP), 5 single-precision DP, and 1 double-precision DP operations. Any of the modes can be executed in two successive clock cycles without idle hardware resources. The proposed design is realized by using the UMC 55-nm process with simulation results. Compared with the existing multiple-precision FP methods, the proposed DPU achieves 88.9% and 35.8% area-saving performance for FP16 and FP32 operations, respectively. Moreover, when using benchmarked HPC applications where multiple precisions can be used, the proposed reconfigurable DPU can accelerate up to 4× and 20× maximum throughput rates when compared with fixed FP32 and FP64 operations, respectively. Wei Mao 0002, Kai Li 0024, Xinang Xie, Shirui Zhao, He Li 0008, Hao Yu 0001 |
DATE | 2 |