Xianghong Hu 0001

dblp:243/1419-1 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0002-1237-4945ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2026 An efficient DSP packing framework for FPGA-based mixed-precision DCNN processor
Xueming Li 0001, Jinhui Pan, Hongmin Huang, Yuanmiao Lin, Xianghong Hu 0001, Shuting Cai, Xiaoming Xiong
J. Syst. Archit.5
2026 An FPGA-Efficient CNN Accelerator for Hybrid Model Compression With Scalable Bit-Serial and Bit-Parallel MAC
abstract
Mixed-precision quantization and unstructured pruning have emerged as two effective compression techniques, demonstrating great potential in reducing model size and computational cost in the deployment of convolutional neural networks (CNNs). However, their joint deployment still faces two major challenges: 1) The former introduces heterogeneous bit-widths in the bit-level, while the latter results in irregular sparsity in the value-level; their fundamentally incompatible data representations and computation patterns require two distinct types of hardware overhead to process them separately, which severely limits hardware execution efficiency. 2) Jointly applying both techniques often leads to notable accuracy degradation. In this paper, we propose a novel compression perspective that reinterprets zero-values generated by unstructured pruning as multiple consecutive 0-bits. We further introduce column-based bit-level sparsity, which provides a unified representation for weights after mixed-precision quantization and unstructured pruning, requiring only a single type of hardware overhead. Based on these techniques, we develop a hybrid compression framework that jointly optimizes model size, accuracy, and hardware implementation. Our method achieves weight/activation precision of 2.13b/4.06b on VGG16, delivering 7.40$\times $compression and 2.74$\times $speedup with 0.92% accuracy loss compared to the 8b baseline. Compared to state-of-the-art accelerators, our design achieves 1.12$\times $-6.23$\times $and 1.31$\times $-6.60$\times $improvements in energy efficiency and LUT efficiency when deploying VGG16 and ResNet50.
Yuanmiao Lin, Xueming Li 0001, Hongmin Huang, Heng Mai, Ruidian Zhan, Xianghong Hu 0001, Shuting Cai, Xiaoming Xiong
IEEE Trans. Circuits Syst. I Regul. Pap.7
2026 Efficient FPGA Acceleration for 4-bit CNNs via Quantization-Induced Structured Sparsity and LUT-Based Multiplication
abstract
N:M structured sparsity is key to convolutional neural network (CNN) compression and acceleration, but two challenges remain. From the algorithm perspective, prior works have mainly focused on 8-bit quantized models, where N:M sparsity yields limited hardware efficiency. From the hardware perspective, the cost differences across N:M sparsity have not been analyzed. To address these issues, we present a unified algorithm–hardware co-design framework for 4-bit CNN acceleration. We show that 4-bit quantization induces over 80% zero weights and strongly structured sparsity, with over 95% of weight groups satisfying 4:8, 8:16, or 16:32 patterns. We propose a pruning-after-quantization (PAQ) algorithm that enforces strict N:M sparsity with minimal accuracy loss. We also analyze the hardware overhead of activation fetch units (AFUs) under different N:M sparsity patterns (4:8, 8:16, 16:32), revealing that the 4:8 AFU reduces look-up table (LUT) cost by up to 66.7% compared to 16:32. Finally, we introduce a 4-bit LUT-based sign-magnitude multiplier (LBSMM) requiring only 11 LUT6 resources, outperforming existing multipliers. Integrated on a Xilinx VCU118 field-programmable gate array (FPGA), our accelerator achieves$2.51\times $–$12.89\times $improvements in equivalent LUT efficiency over SOTA designs. The implementations of the PAQ algorithm and the RTL of LBSMM are available athttps://github.com/haden-01/PAQ-and-LBSMM.git
Yuanmiao Lin, Xueming Li 0001, Hongmin Huang, Ruidian Zhan, Xianghong Hu 0001, Shuting Cai, Xiaoming Xiong
IEEE Trans. Very Large Scale Integr. Syst.7
2025 An FPGA-based bit-level weight sparsity and mixed-bit accelerator for neural networks
Xianghong Hu 0001, Shansen Fu, Yuanmiao Lin, Xueming Li 0001, Chaoming Yang, Rongfeng Li 0001, Hongmin Huang, Shuting Cai, Xiaoming Xiong
J. Syst. Archit.1
2025 A Precision-Scalable Accelerator with Sign-Magnitude Representation and Dual Adder Trees
abstract
Currently, there are two mainstream acceleration methods; one is mixed precision and the other is sparsity. Few accelerators support both mixed precision and sparsity, and most enable precision configurations across layers rather than within a single layer. Furthermore, most of accelerators adopt the traditional two’s complement (2C) data representation method, and we found that 2C brings many invalid ”1” when representing signed data, which brings more resources overhead for mixed precision and many invalid operations for bit-level sparsity. Therefore, we propose a high-efficiency accelerator featuring a precision-scalable Sign-Magnitude Processing Element (SM-PE), which adopts a data representation method of SM and can flexibly support various precision calculations (2, 4, 8 bits) and bit-level sparsity. In addition, a dynamic quantization algorithm named DoReFaLike and a bit-level column sparsity (BLCS) technique are proposed to improve the efficiency of SM-PEs. Under the same accuracy constraint, the sparsity rate of the SM scheme is 3.5× higher than that of the 2C format. The accelerator has been synthesized on a 55nm CMOS ASIC platform. When scaled to 28nm, experimental results show that the energy efficiency of the proposed accelerator reaches 15.50, 25.37, 101.54 TOPS/W with 8-bit, 4-bit, and 2-bit input activations, respectively, and weights represented in sparse 8-bit precision, operating at 400 MHz. Compared to state-of-the-art accelerators, the proposed design achieves a performance improvement of 1.1× to 3.9×.
Xianghong Hu 0001, Chaoming Yang, Xueming Li 0001, Rongfeng Li 0001, Yuanmiao Lin, Shansen Fu, Hongmin Huang, Shuting Cai, Xiaoming Xiong
ACM Trans. Embed. Comput. Syst.1
2023 High-performance Reconfigurable DNN Accelerator on a Bandwidth-limited Embedded System
abstract
Deep convolutional neural networks (DNNs) have been widely used in many applications, particularly in machine vision. It is challenging to accelerate DNNs on embedded systems because real-world machine vision applications should reserve a lot of external memory bandwidth for other tasks, such as video capture and display, while leaving little bandwidth for accelerating DNNs. In order to solve this issue, in this study, we propose a high-throughput accelerator, called reconfigurable tiny neural network accelerator (ReTiNNA), for the bandwidth-limited system and present a real-time object detection system for the high-resolution video image. We first present a dedicated computation engine that takes different data mapping methods for various filter types to improve data reuse and reduce hardware resources. We then propose an adaptive layer-wise tiling strategy that tiles the feature maps into strips to reduce the control complexity of data transmission dramatically and to improve the efficiency of data transmission. Finally, a design space exploration (DSE) approach is presented to explore design space more accurately in the case of insufficient bandwidth to improve the performance of the low-bandwidth accelerator. With a low bandwidth of 2.23 GB/s and a low hardware consumption of 90.261K LUTs and 448 DSPs, ReTiNNA can still achieve a high performance of 155.86 GOPS on VGG16 and 68.20 GOPS on ResNet50, which is better than other state-of-the-art designs implemented on FPGA devices. Furthermore, the real-time object detection system can achieve a high object detection speed of 19 fps for high-resolution video.
Xianghong Hu 0001, Hongmin Huang, Xueming Li 0001, Xin Zheng 0001, Qinyuan Ren, Jingyu He, Xiaoming Xiong
ACM Trans. Embed. Comput. Syst.1
2022 SDQ: Stochastic Differentiable Quantization with Mixed Precision
abstract
In order to deploy deep models in a computationally efficient manner, model quantization approaches have been frequently used. In addition, as new hardware that supports various-bit arithmetic operations, recent research on mixed precision quantization (MPQ) begins to fully leverage the capacity of representation by searching various bitwidths for different layers and modules in a network. However, previous studies mainly search the MPQ strategy in a costly scheme using reinforcement learning, neural architecture search, etc., or simply utilize partial prior knowledge for bitwidth distribution, which might be biased and sub-optimal. In this work, we present a novel Stochastic Differentiable Quantization (SDQ) method that can automatically learn the MPQ strategy in a more flexible and globally-optimized space with a smoother gradient approximation. Particularly, Differentiable Bitwidth Parameters (DBPs) are employed as the probability factors in stochastic quantization between adjacent bitwidth. After the optimal MPQ strategy is acquired, we further train our network with the entropy-aware bin regularization and knowledge distillation. We extensively evaluate our method on different networks, hardwares (GPUs and FPGA), and datasets. SDQ outperforms all other state-of-the-art mixed or single precision quantization with less bitwidth, and are even better than the original full-precision counterparts across various ResNet and MobileNet families, demonstrating the effectiveness and superiority of our method. Code will be publicly available.
Xijie Huang, Shichao Li 0002, Zechun Liu, Xianghong Hu 0001, Jeffry Wicaksana, Eric P. Xing, Kwang-Ting Cheng
ICML5
2021 Subgraph feature extraction based on multi-view dictionary learning for graph classification
Xin Zheng 0001, Shouzhi Liang, Bo Liu 0002, Xiaoming Xiong, Xianghong Hu 0001, Yuan Liu 0022
Knowl. Based Syst.5
2020 The Software/Hardware Co-Design and Implementation of SM2/3/4 Encryption/Decryption and Digital Signature System
abstract
The security of smart devices is facing great challenges. This article presents a new hybrid cipher framework suitable for such devices. Using software/hardware (SW/HW) co-design method, an efficient encryption/decryption and digital signature scheme based on SM2, SM3, and SM4 algorithms is implemented. First, the framework is partitioned into software and hardware parts based on the analysis result of a pure software solution and the hardware overhead under the conditions of design constraints. Second, the flow of the whole scheme and embedded algorithms are realized by SW/HW co-design. Finally, an improved implementation of SM2/3/4 algorithms is proposed to achieve higher efficiency. In this implementation, some SW/HW modules are parallelized to reduce the running time and enhance the performance. The proposed design is safe and can resist simple power analysis (SPA) attacks. Especially, the AHB bus interface IP and software scheduling approach are adopted in data transfer and manipulation. The design is taped-out on a silicon chip with SMIC 110-nm technology process. The chip uses about 199K logic gates and 1-mm2areas. The operating frequency of the design is 36 MHz and the chip consumes 23-mW power. The comparison with similar previous works shows our proposed design is more efficient with speed increasing of more than 10%.
Xin Zheng 0001, Chongyao Xu, Xianghong Hu 0001, Yun Zhang 0001, Xiaoming Xiong
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3