EDBT 2026 Demo / reviewers in the wild / expert
Hongmin Huang
dblp:291/7104
· DBLP profile ↗
10ranked-venue papers
0as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 6 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An efficient DSP packing framework for FPGA-based mixed-precision DCNN processor
Xueming Li 0001, Jinhui Pan, Hongmin Huang, Yuanmiao Lin, Xianghong Hu 0001, Shuting Cai, Xiaoming Xiong |
J. Syst. Archit. | 3 |
| 2026 | An FPGA-Efficient CNN Accelerator for Hybrid Model Compression With Scalable Bit-Serial and Bit-Parallel MACabstractMixed-precision quantization and unstructured pruning have emerged as two effective compression techniques, demonstrating great potential in reducing model size and computational cost in the deployment of convolutional neural networks (CNNs). However, their joint deployment still faces two major challenges: 1) The former introduces heterogeneous bit-widths in the bit-level, while the latter results in irregular sparsity in the value-level; their fundamentally incompatible data representations and computation patterns require two distinct types of hardware overhead to process them separately, which severely limits hardware execution efficiency. 2) Jointly applying both techniques often leads to notable accuracy degradation. In this paper, we propose a novel compression perspective that reinterprets zero-values generated by unstructured pruning as multiple consecutive 0-bits. We further introduce column-based bit-level sparsity, which provides a unified representation for weights after mixed-precision quantization and unstructured pruning, requiring only a single type of hardware overhead. Based on these techniques, we develop a hybrid compression framework that jointly optimizes model size, accuracy, and hardware implementation. Our method achieves weight/activation precision of 2.13b/4.06b on VGG16, delivering 7.40$\times $compression and 2.74$\times $speedup with 0.92% accuracy loss compared to the 8b baseline. Compared to state-of-the-art accelerators, our design achieves 1.12$\times $-6.23$\times $and 1.31$\times $-6.60$\times $improvements in energy efficiency and LUT efficiency when deploying VGG16 and ResNet50. Yuanmiao Lin, Xueming Li 0001, Hongmin Huang, Heng Mai, Ruidian Zhan, Xianghong Hu 0001, Shuting Cai, Xiaoming Xiong |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2026 | Efficient FPGA Acceleration for 4-bit CNNs via Quantization-Induced Structured Sparsity and LUT-Based MultiplicationabstractN:M structured sparsity is key to convolutional neural network (CNN) compression and acceleration, but two challenges remain. From the algorithm perspective, prior works have mainly focused on 8-bit quantized models, where N:M sparsity yields limited hardware efficiency. From the hardware perspective, the cost differences across N:M sparsity have not been analyzed. To address these issues, we present a unified algorithm–hardware co-design framework for 4-bit CNN acceleration. We show that 4-bit quantization induces over 80% zero weights and strongly structured sparsity, with over 95% of weight groups satisfying 4:8, 8:16, or 16:32 patterns. We propose a pruning-after-quantization (PAQ) algorithm that enforces strict N:M sparsity with minimal accuracy loss. We also analyze the hardware overhead of activation fetch units (AFUs) under different N:M sparsity patterns (4:8, 8:16, 16:32), revealing that the 4:8 AFU reduces look-up table (LUT) cost by up to 66.7% compared to 16:32. Finally, we introduce a 4-bit LUT-based sign-magnitude multiplier (LBSMM) requiring only 11 LUT6 resources, outperforming existing multipliers. Integrated on a Xilinx VCU118 field-programmable gate array (FPGA), our accelerator achieves$2.51\times $–$12.89\times $improvements in equivalent LUT efficiency over SOTA designs. The implementations of the PAQ algorithm and the RTL of LBSMM are available athttps://github.com/haden-01/PAQ-and-LBSMM.git Yuanmiao Lin, Xueming Li 0001, Hongmin Huang, Ruidian Zhan, Xianghong Hu 0001, Shuting Cai, Xiaoming Xiong |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2025 | An FPGA-based bit-level weight sparsity and mixed-bit accelerator for neural networks
Xianghong Hu 0001, Shansen Fu, Yuanmiao Lin, Xueming Li 0001, Chaoming Yang, Rongfeng Li 0001, Hongmin Huang, Shuting Cai, Xiaoming Xiong |
J. Syst. Archit. | 7 |
| 2025 | A Precision-Scalable Accelerator with Sign-Magnitude Representation and Dual Adder TreesabstractCurrently, there are two mainstream acceleration methods; one is mixed precision and the other is sparsity. Few accelerators support both mixed precision and sparsity, and most enable precision configurations across layers rather than within a single layer. Furthermore, most of accelerators adopt the traditional two’s complement (2C) data representation method, and we found that 2C brings many invalid ”1” when representing signed data, which brings more resources overhead for mixed precision and many invalid operations for bit-level sparsity. Therefore, we propose a high-efficiency accelerator featuring a precision-scalable Sign-Magnitude Processing Element (SM-PE), which adopts a data representation method of SM and can flexibly support various precision calculations (2, 4, 8 bits) and bit-level sparsity. In addition, a dynamic quantization algorithm named DoReFaLike and a bit-level column sparsity (BLCS) technique are proposed to improve the efficiency of SM-PEs. Under the same accuracy constraint, the sparsity rate of the SM scheme is 3.5× higher than that of the 2C format. The accelerator has been synthesized on a 55nm CMOS ASIC platform. When scaled to 28nm, experimental results show that the energy efficiency of the proposed accelerator reaches 15.50, 25.37, 101.54 TOPS/W with 8-bit, 4-bit, and 2-bit input activations, respectively, and weights represented in sparse 8-bit precision, operating at 400 MHz. Compared to state-of-the-art accelerators, the proposed design achieves a performance improvement of 1.1× to 3.9×. Xianghong Hu 0001, Chaoming Yang, Xueming Li 0001, Rongfeng Li 0001, Yuanmiao Lin, Shansen Fu, Hongmin Huang, Shuting Cai, Xiaoming Xiong |
ACM Trans. Embed. Comput. Syst. | 7 |
| 2023 | High-performance Reconfigurable DNN Accelerator on a Bandwidth-limited Embedded SystemabstractDeep convolutional neural networks (DNNs) have been widely used in many applications, particularly in machine vision. It is challenging to accelerate DNNs on embedded systems because real-world machine vision applications should reserve a lot of external memory bandwidth for other tasks, such as video capture and display, while leaving little bandwidth for accelerating DNNs. In order to solve this issue, in this study, we propose a high-throughput accelerator, called reconfigurable tiny neural network accelerator (ReTiNNA), for the bandwidth-limited system and present a real-time object detection system for the high-resolution video image. We first present a dedicated computation engine that takes different data mapping methods for various filter types to improve data reuse and reduce hardware resources. We then propose an adaptive layer-wise tiling strategy that tiles the feature maps into strips to reduce the control complexity of data transmission dramatically and to improve the efficiency of data transmission. Finally, a design space exploration (DSE) approach is presented to explore design space more accurately in the case of insufficient bandwidth to improve the performance of the low-bandwidth accelerator. With a low bandwidth of 2.23 GB/s and a low hardware consumption of 90.261K LUTs and 448 DSPs, ReTiNNA can still achieve a high performance of 155.86 GOPS on VGG16 and 68.20 GOPS on ResNet50, which is better than other state-of-the-art designs implemented on FPGA devices. Furthermore, the real-time object detection system can achieve a high object detection speed of 19 fps for high-resolution video. Xianghong Hu 0001, Hongmin Huang, Xueming Li 0001, Xin Zheng 0001, Qinyuan Ren, Jingyu He, Xiaoming Xiong |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2022 | Prioritized Contention Access Based MAC Protocol for In-Vivo Wireless NanoSensor NetworksabstractTerahertz based in-vivo Wireless NanoSensor Networks (iWNSNs) is a new type of NanoSensor Networks which take terahertz wave as its carrier and works in the human body. Multiple nano devices in the network are connected by wireless communication. The propagation characteristics of terahertz wave in-vivo are different from those in free space, and there are more serious molecular absorption noises and path losses. Besides, the nano devices are limited in battery energy, so the existing MAC (Medium Access Control) protocols cannot be directly utilized in the Terahertz based in-vivo Wireless NanoSensor Networks. To investigate the MAC protocol which is suitable for the Terahertz based on iWNSNs Networks, this paper proposes a Prioritized Contention Access Based MAC protocol (PCAB-MAC). Nanosensor nodes access the channel through competition, and a two-way handshake is established to ensure nodes can transmit data undisturbed. Considering that data has different priorities, the PCAB-MAC adopts a policy based on priority, and sets different backoff windows for data of different priorities to ensure priority transmission. The simulation results show that the PCAB-MAC protocol can ensure data transmission without conflict, and has excellent performance in delay and throughput. Juan Xu 0003, Hongmin Huang, Yakun Zhao, Lin Lin 0002 |
PIMRC | 2 |
| 2022 | Energy-balanced routing protocol based on data priority for lung terahertz nanosensor networksabstractConsidering problems that there are different urgency levels of health data and limited resources of nanonodes, an energy-balanced routing protocol based on data priority (DPER) is proposed. The priority of data and the lung terahertz time-varying channel characteristics are considered in this protocol creatively. The link selection function proposed in this paper can achieve the goal of allocating different data transmission rates for different priorities of data. The simulation results prove that the DPER protocol can provide different average delays for data with different priorities, and the degree of differentiation of data with different priorities is higher. In addition, the DPER protocol also has greater advantages in terms of packet transmission success rate and average throughput. Juan Xu 0003, Hongmin Huang, Jiali Kan |
VTC Spring | 2 |
| 2021 | A MAC Protocol Based on Energy Scheduling for In-Vivo Wireless NanoSensor NetworksabstractTerahertz based in-vivo Wireless NanoSensor Networks (iWNSNs) is a novel sensor network using electromagnetic waves in terahertz band as carrier and nanotechnology work in human body. Due to the different propagation characteristic of in-vivo terahertz channel and the limited resource of nano device, current MAC (Medium Access Control) protocol cannot be applied to Terahertz based iWNSNs. In this paper, a MAC protocol based on energy scheduling, called ES-MAC, is proposed. This protocol adopts nano energy harvesting system, designs a reward function based on the amount of transmitted data and the priority of data. Nano sensor nodes adopt Sarsa learning algorithm to make dynamic channel access decision according to their status. Simulation results show ES-MAC protocol can decrease average end-to-end delay and prolong lifetime of the network while providing differentiated services. Juan Xu 0003, Hongmin Huang, Yakun Zhao, Lin Lin 0002 |
WCNC | 2 |
| 2021 | Topological Optimization of Lung Wireless Nanosensor NetworkabstractThe early detection and prevention of lung disease is always a difficult point in the medical field, but it is of great significance. Some biological signals of the lung contain abundant information, especially the disease signals with great value. With the development of sensing technology, these specific biological signals are enough to be detected by sensors. However, if the information is to be transmitted from lung sensors to medical institutions in the macroworld, a suitable sensor network should be proposed. At the same time, the structure of lung is complex, and the channel in human body is special, so the construction of network is subject to many constraints. In order to construct an in vivo wireless nanosensor network (WNSN) that can meet the requirements of delay and energy, the model of lung WNSN is constructed in this article. On this basis, the performance of in vivo WNSNs (iWNSNs) with different equalization methods and different topologies is compared. Finally, a low-delay energy-efficient (LDEE) lung WNSN topology model is proposed. In this model, the total network delay and energy consumption of iWNSNs are taken as optimization objectives, the terahertz channel environment of lung and the properties of nanonodes are taken as constraint variables, and the topology model with the best performance is solved by MATLAB. The model can meet the necessary conditions of lung signal transmission, and has the characteristics of low network delay, high throughput, and long network lifetime. Juan Xu 0003, Hongmin Huang, Jiali Kan, Yongfa Hong |
IEEE Internet Things J. | 2 |