Heng Zhang 0025

dblp:55/826-25 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
9since 2021 · last 2026
0009-0005-3618-5849ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 3 first-author · 9 since 2021
YearPublicationVenuePosition
2026 Lightweight Multi-View EEG Seizure Detection with a Two-Stage Low-Power BFP FFT-NN Accelerator
Chi Ben, Youbin Luo, Zhenglin Gu, Kexin Tian, Heng Zhang 0025, Li Li 0003
ISCAS5
2026 FRRAP: A Fast-Response Reconfigurable AI Processor and Its Application in Four-Class Seizure Monitoring
Heng Zhang 0025, Qiang Tao, Chi Ben, Youbin Luo, Guoqiang He, Sirui Zhu, Jingyi Ma, Li Li 0003
ISCAS1
2026 Thermal-Aware 3D-IC Floorplan Based On TSV-Coordination Simulated Annealing
Xingjie Zou, Linfeng Wu, Heng Zhang 0025, Xinyu Wang 0027, Li Li 0003
ISCAS4
2026 A 1.1 μJ/Inference Binary Spiking Neural Network Accelerator for DVS Gesture Recognition
abstract
Dynamic vision sensors (DVS) are bioinspired sensors that can generate sparse data streams with low latency, low power consumption and high dynamic range. Spiking neural networks (SNNs), which are inspired by biological brains, are event-based models, and therefore they can be used to process the binary data streams produced by such sensors naturally. However, SNN accelerators usually require more memory and longer time for inference, due to the extra time dimension in SNNs. In this paper, an energy efficient binary spiking neural network (BSNN) accelerator for DVS gesture recognition is proposed with algorithm and hardware codesign. We integrate the binary neural network (BNN) training method into SNN to directly train a BSNN model, which significantly reduces memory consumption. A temporal pooling (TP) layer is further proposed to reduce the time steps in SNNs while maintaining competitive accuracy. The proposed BSNN accelerator can achieve high parallelism with high resource utilization, and the sparsity of input spikes is utilized to further reduce power consumption. The proposed BSNN model achieves an accuracy of 95.49% on IBM DVS Gesture dataset. The implementation results show that the BSNN accelerator can achieve 38.2k inference per second with$1.1~\mu $J/inference energy consumption and 216.9 TOPS/W energy efficiency.
Congyi Sun, Xusen Zeng, Qiang Tao, Heng Zhang 0025, Qinyu Chen, Li Li 0003
IEEE Trans. Circuits Syst. I Regul. Pap.4
2026 QSNNA: An Energy-Efficient Quaternary Spiking Neural Network Accelerator for Seizure Detection
Heng Zhang 0025, Linfeng Wu, Linxiang Wang, Youbin Luo, Haochuan Pan, Xinyu Wang 0027, Guoqiang He, Qinyu Chen, Li Li 0003
IEEE Trans. Very Large Scale Integr. Syst.1
2025 NAME: NoC-based Accelerators Mapping Exploration for High Performance DNN Inference
abstract
The rapid advancement of deep learning, with increasingly large deep neural networks (DNNs), has led to the use of multi-core parallel processing in accelerators, utilizing Network-on-Chip (NoC) for interconnection. However, while multi-core architectures improve computational performance, they also increase data movement overhead. This paper analyzes NoC traffic patterns during DNN processing and explores mapping optimization to reduce computational and communicational overheads. We propose NoC-based Accelerators Mapping Exploration (NAME), an automated mapper for generating high-performance task mappings for NoC-based accelerators. NAME balances computational and communication overheads by dividing DNNs into scalable groups and interleaving NoC link usage, reducing traffic bottlenecks. Experiments show that NAME reduces execution time by up to 82% and 52%, respectively, compared to fixed mapping and the state-of-the-art framework AOME (Autonomous Optimal Mapping Exploration).
Jinlun Ji, Hengyue Gao, Heng Zhang 0025, Yuqi Lu, Yulong Song, Wenjie Fan 0004, Li Li 0003
ISCAS3
2025 HengNet: An Ultra-lightweight Model with Two-level Reuse Algorithm for Seizure Detection and Prediction
abstract
Traditional models based on electroencephalographic (EEG) signals for seizure monitoring encounter difficulties in simultaneously optimizing accuracy, response latency, and computational load. These challenges hinder their deployment in edge computing environments, where real-time local inference is critical. To address these issues, we introduce a novel network architecture, designated as HengNet. This architecture integrates a Two-level Reuse Algorithm (TRA), which strategically reutilizes outputs from intermediate layers, considerably reducing the average computational load per inference—vital for scenarios requiring frequent inferences. When tested on the CHB-MIT dataset, this patient-specific model attains classification accuracies of 95.67% and 99.60% for seizure prediction and detection, respectively. Notably, it maintains an average computational load of merely 0.05 million multiply-accumulate operations (MACs) per inference and has a compact model size of 6.87 K parameters. These results represent a significant advancement compared with existing methods. Operating at a rate of 32 inferences per second, the computational load of the model for seizure prediction has been reduced by more than 19.4 times, and for seizure detection, by more than 6.4 times.
Heng Zhang 0025, Linxiang Wang, Wenjie Fan 0004, Zhenglin Gu, Youbin Luo, Xingjie Zou, Chang Gao 0002, Qinyu Chen, Li Li 0003
ISCAS1
2025 FAS-NoC: A Real-Time Fused Approximation Scheme Coordinating Communication and Computation for NoC-Based NN Accelerators
abstract
Network-on-Chip (NoC) is a scalable on-chip communication architecture widely used in neural network accelerators. However, data-intensive applications like machine learning place significant demands on the NoC’s communication and computation, and often have a degree of resilience to data noise, which allows to use approximation techniques to reduce execution time and energy consumption for both computation and communication, under the constraints of acceptable quality loss. Traditional approximate NoCs do not consider the data distribution characteristics of the neural networks, resulting in a lower approximate rate. Moreover, these schemes do not take the synergistic optimization of computation and communication, which limits reductions in execution time. In this paper, we propose a Fused Approximation Scheme of NoC (FAS-NoC) that incorporates the characteristics of data distribution in neural networks. FAS-NoC includes an approximate compression and recovery scheme based on data hierarchy, and uses a congestion-aware scheme to adjust the approximate rate of the node. Additionally, by leveraging the characteristics of recovered data after approximate communication, the scheme optimizes the design of computing units within the computing array. FAS-NoC collaboratively optimizes approximate communication and computation, organically integrating the two aspects. Compared with the state-of-the-art approximate framework ACDC (ACDC_ABDTR and ACDC_APPROX), the execution time of FAS-NoC is reduced by$48.94\%$and$47.72\%$, respectively. The experimental results show that the additional area overhead of the FAS-NoC only accounts for$0.81\%$of the original node, the additional power consumption overhead only accounts for$0.75\%$of the original node.
Chuanzhu Liu, Wenjie Fan 0004, Heng Zhang 0025, Chenyang Dai, Congyi Sun, Xinyu Wang 0027, Li Li 0003
IEEE Trans. Circuits Syst. I Regul. Pap.4
2024 TTNNM: Thermal- and Traffic-Aware Neural Network Mapping on 3D-NoC-based Accelerator
abstract
3D Network on Chips (3D-NoCs) have ample on-chip wiring resources and high bandwidth, yet face numerous hotspots and higher temperature gradients due to increased integration and power density. This could lead to device failure, impacting system stability. Our paper introduces a thermal- and traffic-aware mapping method for 3D-NoC-based neural network accelerators. Firstly, based on the average load of different neural network layer, we determine their mapping sequences and suitable dies. Secondly, to minimize delay and alleviate hotspot temperatures, we allocate groups to appropriate nodes. Compared with previous works, TTNNM reduces the average temperature by 3.0°C, 2.2°C, 2.4°C, temperature variance by 58.4%, 64.8%, 73.0%, maximum temperature by 9.3°C, 7.9°C, 12.0°C, and packet latency by 31.7%, 17.2%, 25.1%.
Wenjie Fan 0004, Heng Zhang 0025, Jinlun Ji, Tong Cheng, Shiping Li, Li Li 0003
ACM Great Lakes Symposium on VLSI3