EDBT 2026 Demo / reviewers in the wild / expert
Xueming Li 0001
dblp:67/2097-1
· DBLP profile ↗
6ranked-venue papers
1as first author
6since 2021 · last 2026
0000-0002-9700-4272ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An efficient DSP packing framework for FPGA-based mixed-precision DCNN processor
Xueming Li 0001, Jinhui Pan, Hongmin Huang, Yuanmiao Lin, Xianghong Hu 0001, Shuting Cai, Xiaoming Xiong |
J. Syst. Archit. | 1 |
| 2026 | An FPGA-Efficient CNN Accelerator for Hybrid Model Compression With Scalable Bit-Serial and Bit-Parallel MACabstractMixed-precision quantization and unstructured pruning have emerged as two effective compression techniques, demonstrating great potential in reducing model size and computational cost in the deployment of convolutional neural networks (CNNs). However, their joint deployment still faces two major challenges: 1) The former introduces heterogeneous bit-widths in the bit-level, while the latter results in irregular sparsity in the value-level; their fundamentally incompatible data representations and computation patterns require two distinct types of hardware overhead to process them separately, which severely limits hardware execution efficiency. 2) Jointly applying both techniques often leads to notable accuracy degradation. In this paper, we propose a novel compression perspective that reinterprets zero-values generated by unstructured pruning as multiple consecutive 0-bits. We further introduce column-based bit-level sparsity, which provides a unified representation for weights after mixed-precision quantization and unstructured pruning, requiring only a single type of hardware overhead. Based on these techniques, we develop a hybrid compression framework that jointly optimizes model size, accuracy, and hardware implementation. Our method achieves weight/activation precision of 2.13b/4.06b on VGG16, delivering 7.40$\times $compression and 2.74$\times $speedup with 0.92% accuracy loss compared to the 8b baseline. Compared to state-of-the-art accelerators, our design achieves 1.12$\times $-6.23$\times $and 1.31$\times $-6.60$\times $improvements in energy efficiency and LUT efficiency when deploying VGG16 and ResNet50. Yuanmiao Lin, Xueming Li 0001, Hongmin Huang, Heng Mai, Ruidian Zhan, Xianghong Hu 0001, Shuting Cai, Xiaoming Xiong |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2026 | Efficient FPGA Acceleration for 4-bit CNNs via Quantization-Induced Structured Sparsity and LUT-Based MultiplicationabstractN:M structured sparsity is key to convolutional neural network (CNN) compression and acceleration, but two challenges remain. From the algorithm perspective, prior works have mainly focused on 8-bit quantized models, where N:M sparsity yields limited hardware efficiency. From the hardware perspective, the cost differences across N:M sparsity have not been analyzed. To address these issues, we present a unified algorithm–hardware co-design framework for 4-bit CNN acceleration. We show that 4-bit quantization induces over 80% zero weights and strongly structured sparsity, with over 95% of weight groups satisfying 4:8, 8:16, or 16:32 patterns. We propose a pruning-after-quantization (PAQ) algorithm that enforces strict N:M sparsity with minimal accuracy loss. We also analyze the hardware overhead of activation fetch units (AFUs) under different N:M sparsity patterns (4:8, 8:16, 16:32), revealing that the 4:8 AFU reduces look-up table (LUT) cost by up to 66.7% compared to 16:32. Finally, we introduce a 4-bit LUT-based sign-magnitude multiplier (LBSMM) requiring only 11 LUT6 resources, outperforming existing multipliers. Integrated on a Xilinx VCU118 field-programmable gate array (FPGA), our accelerator achieves$2.51\times $–$12.89\times $improvements in equivalent LUT efficiency over SOTA designs. The implementations of the PAQ algorithm and the RTL of LBSMM are available athttps://github.com/haden-01/PAQ-and-LBSMM.git Yuanmiao Lin, Xueming Li 0001, Hongmin Huang, Ruidian Zhan, Xianghong Hu 0001, Shuting Cai, Xiaoming Xiong |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2025 | An FPGA-based bit-level weight sparsity and mixed-bit accelerator for neural networks
Xianghong Hu 0001, Shansen Fu, Yuanmiao Lin, Xueming Li 0001, Chaoming Yang, Rongfeng Li 0001, Hongmin Huang, Shuting Cai, Xiaoming Xiong |
J. Syst. Archit. | 4 |
| 2025 | A Precision-Scalable Accelerator with Sign-Magnitude Representation and Dual Adder TreesabstractCurrently, there are two mainstream acceleration methods; one is mixed precision and the other is sparsity. Few accelerators support both mixed precision and sparsity, and most enable precision configurations across layers rather than within a single layer. Furthermore, most of accelerators adopt the traditional two’s complement (2C) data representation method, and we found that 2C brings many invalid ”1” when representing signed data, which brings more resources overhead for mixed precision and many invalid operations for bit-level sparsity. Therefore, we propose a high-efficiency accelerator featuring a precision-scalable Sign-Magnitude Processing Element (SM-PE), which adopts a data representation method of SM and can flexibly support various precision calculations (2, 4, 8 bits) and bit-level sparsity. In addition, a dynamic quantization algorithm named DoReFaLike and a bit-level column sparsity (BLCS) technique are proposed to improve the efficiency of SM-PEs. Under the same accuracy constraint, the sparsity rate of the SM scheme is 3.5× higher than that of the 2C format. The accelerator has been synthesized on a 55nm CMOS ASIC platform. When scaled to 28nm, experimental results show that the energy efficiency of the proposed accelerator reaches 15.50, 25.37, 101.54 TOPS/W with 8-bit, 4-bit, and 2-bit input activations, respectively, and weights represented in sparse 8-bit precision, operating at 400 MHz. Compared to state-of-the-art accelerators, the proposed design achieves a performance improvement of 1.1× to 3.9×. Xianghong Hu 0001, Chaoming Yang, Xueming Li 0001, Rongfeng Li 0001, Yuanmiao Lin, Shansen Fu, Hongmin Huang, Shuting Cai, Xiaoming Xiong |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2023 | High-performance Reconfigurable DNN Accelerator on a Bandwidth-limited Embedded SystemabstractDeep convolutional neural networks (DNNs) have been widely used in many applications, particularly in machine vision. It is challenging to accelerate DNNs on embedded systems because real-world machine vision applications should reserve a lot of external memory bandwidth for other tasks, such as video capture and display, while leaving little bandwidth for accelerating DNNs. In order to solve this issue, in this study, we propose a high-throughput accelerator, called reconfigurable tiny neural network accelerator (ReTiNNA), for the bandwidth-limited system and present a real-time object detection system for the high-resolution video image. We first present a dedicated computation engine that takes different data mapping methods for various filter types to improve data reuse and reduce hardware resources. We then propose an adaptive layer-wise tiling strategy that tiles the feature maps into strips to reduce the control complexity of data transmission dramatically and to improve the efficiency of data transmission. Finally, a design space exploration (DSE) approach is presented to explore design space more accurately in the case of insufficient bandwidth to improve the performance of the low-bandwidth accelerator. With a low bandwidth of 2.23 GB/s and a low hardware consumption of 90.261K LUTs and 448 DSPs, ReTiNNA can still achieve a high performance of 155.86 GOPS on VGG16 and 68.20 GOPS on ResNet50, which is better than other state-of-the-art designs implemented on FPGA devices. Furthermore, the real-time object detection system can achieve a high object detection speed of 19 fps for high-resolution video. Xianghong Hu 0001, Hongmin Huang, Xueming Li 0001, Xin Zheng 0001, Qinyuan Ren, Jingyu He, Xiaoming Xiong |
ACM Trans. Embed. Comput. Syst. | 3 |