EDBT 2026 Demo / reviewers in the wild / expert
Heng Zhang 0025
dblp:55/826-25
· DBLP profile ↗
9ranked-venue papers
3as first author
9since 2021 · last 2026
0009-0005-3618-5849ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 3 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Lightweight Multi-View EEG Seizure Detection with a Two-Stage Low-Power BFP FFT-NN Accelerator
Chi Ben, Youbin Luo, Zhenglin Gu, Kexin Tian, Heng Zhang 0025, Li Li 0003 |
ISCAS | 5 |
| 2026 | FRRAP: A Fast-Response Reconfigurable AI Processor and Its Application in Four-Class Seizure Monitoring
Heng Zhang 0025, Qiang Tao, Chi Ben, Youbin Luo, Guoqiang He, Sirui Zhu, Jingyi Ma, Li Li 0003 |
ISCAS | 1 |
| 2026 | Thermal-Aware 3D-IC Floorplan Based On TSV-Coordination Simulated Annealing
Xingjie Zou, Linfeng Wu, Heng Zhang 0025, Xinyu Wang 0027, Li Li 0003 |
ISCAS | 4 |
| 2026 | A 1.1 μJ/Inference Binary Spiking Neural Network Accelerator for DVS Gesture RecognitionabstractDynamic vision sensors (DVS) are bioinspired sensors that can generate sparse data streams with low latency, low power consumption and high dynamic range. Spiking neural networks (SNNs), which are inspired by biological brains, are event-based models, and therefore they can be used to process the binary data streams produced by such sensors naturally. However, SNN accelerators usually require more memory and longer time for inference, due to the extra time dimension in SNNs. In this paper, an energy efficient binary spiking neural network (BSNN) accelerator for DVS gesture recognition is proposed with algorithm and hardware codesign. We integrate the binary neural network (BNN) training method into SNN to directly train a BSNN model, which significantly reduces memory consumption. A temporal pooling (TP) layer is further proposed to reduce the time steps in SNNs while maintaining competitive accuracy. The proposed BSNN accelerator can achieve high parallelism with high resource utilization, and the sparsity of input spikes is utilized to further reduce power consumption. The proposed BSNN model achieves an accuracy of 95.49% on IBM DVS Gesture dataset. The implementation results show that the BSNN accelerator can achieve 38.2k inference per second with$1.1~\mu $J/inference energy consumption and 216.9 TOPS/W energy efficiency. Congyi Sun, Xusen Zeng, Qiang Tao, Heng Zhang 0025, Qinyu Chen, Li Li 0003 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2026 | QSNNA: An Energy-Efficient Quaternary Spiking Neural Network Accelerator for Seizure Detection
Heng Zhang 0025, Linfeng Wu, Linxiang Wang, Youbin Luo, Haochuan Pan, Xinyu Wang 0027, Guoqiang He, Qinyu Chen, Li Li 0003 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2025 | NAME: NoC-based Accelerators Mapping Exploration for High Performance DNN InferenceabstractThe rapid advancement of deep learning, with increasingly large deep neural networks (DNNs), has led to the use of multi-core parallel processing in accelerators, utilizing Network-on-Chip (NoC) for interconnection. However, while multi-core architectures improve computational performance, they also increase data movement overhead. This paper analyzes NoC traffic patterns during DNN processing and explores mapping optimization to reduce computational and communicational overheads. We propose NoC-based Accelerators Mapping Exploration (NAME), an automated mapper for generating high-performance task mappings for NoC-based accelerators. NAME balances computational and communication overheads by dividing DNNs into scalable groups and interleaving NoC link usage, reducing traffic bottlenecks. Experiments show that NAME reduces execution time by up to 82% and 52%, respectively, compared to fixed mapping and the state-of-the-art framework AOME (Autonomous Optimal Mapping Exploration). Jinlun Ji, Hengyue Gao, Heng Zhang 0025, Yuqi Lu, Yulong Song, Wenjie Fan 0004, Li Li 0003 |
ISCAS | 3 |
| 2025 | HengNet: An Ultra-lightweight Model with Two-level Reuse Algorithm for Seizure Detection and PredictionabstractTraditional models based on electroencephalographic (EEG) signals for seizure monitoring encounter difficulties in simultaneously optimizing accuracy, response latency, and computational load. These challenges hinder their deployment in edge computing environments, where real-time local inference is critical. To address these issues, we introduce a novel network architecture, designated as HengNet. This architecture integrates a Two-level Reuse Algorithm (TRA), which strategically reutilizes outputs from intermediate layers, considerably reducing the average computational load per inference—vital for scenarios requiring frequent inferences. When tested on the CHB-MIT dataset, this patient-specific model attains classification accuracies of 95.67% and 99.60% for seizure prediction and detection, respectively. Notably, it maintains an average computational load of merely 0.05 million multiply-accumulate operations (MACs) per inference and has a compact model size of 6.87 K parameters. These results represent a significant advancement compared with existing methods. Operating at a rate of 32 inferences per second, the computational load of the model for seizure prediction has been reduced by more than 19.4 times, and for seizure detection, by more than 6.4 times. Heng Zhang 0025, Linxiang Wang, Wenjie Fan 0004, Zhenglin Gu, Youbin Luo, Xingjie Zou, Chang Gao 0002, Qinyu Chen, Li Li 0003 |
ISCAS | 1 |
| 2025 | FAS-NoC: A Real-Time Fused Approximation Scheme Coordinating Communication and Computation for NoC-Based NN AcceleratorsabstractNetwork-on-Chip (NoC) is a scalable on-chip communication architecture widely used in neural network accelerators. However, data-intensive applications like machine learning place significant demands on the NoC’s communication and computation, and often have a degree of resilience to data noise, which allows to use approximation techniques to reduce execution time and energy consumption for both computation and communication, under the constraints of acceptable quality loss. Traditional approximate NoCs do not consider the data distribution characteristics of the neural networks, resulting in a lower approximate rate. Moreover, these schemes do not take the synergistic optimization of computation and communication, which limits reductions in execution time. In this paper, we propose a Fused Approximation Scheme of NoC (FAS-NoC) that incorporates the characteristics of data distribution in neural networks. FAS-NoC includes an approximate compression and recovery scheme based on data hierarchy, and uses a congestion-aware scheme to adjust the approximate rate of the node. Additionally, by leveraging the characteristics of recovered data after approximate communication, the scheme optimizes the design of computing units within the computing array. FAS-NoC collaboratively optimizes approximate communication and computation, organically integrating the two aspects. Compared with the state-of-the-art approximate framework ACDC (ACDC_ABDTR and ACDC_APPROX), the execution time of FAS-NoC is reduced by$48.94\%$and$47.72\%$, respectively. The experimental results show that the additional area overhead of the FAS-NoC only accounts for$0.81\%$of the original node, the additional power consumption overhead only accounts for$0.75\%$of the original node. Chuanzhu Liu, Wenjie Fan 0004, Heng Zhang 0025, Chenyang Dai, Congyi Sun, Xinyu Wang 0027, Li Li 0003 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2024 | TTNNM: Thermal- and Traffic-Aware Neural Network Mapping on 3D-NoC-based Acceleratorabstract3D Network on Chips (3D-NoCs) have ample on-chip wiring resources and high bandwidth, yet face numerous hotspots and higher temperature gradients due to increased integration and power density. This could lead to device failure, impacting system stability. Our paper introduces a thermal- and traffic-aware mapping method for 3D-NoC-based neural network accelerators. Firstly, based on the average load of different neural network layer, we determine their mapping sequences and suitable dies. Secondly, to minimize delay and alleviate hotspot temperatures, we allocate groups to appropriate nodes. Compared with previous works, TTNNM reduces the average temperature by 3.0°C, 2.2°C, 2.4°C, temperature variance by 58.4%, 64.8%, 73.0%, maximum temperature by 9.3°C, 7.9°C, 12.0°C, and packet latency by 31.7%, 17.2%, 25.1%. Wenjie Fan 0004, Heng Zhang 0025, Jinlun Ji, Tong Cheng, Shiping Li, Li Li 0003 |
ACM Great Lakes Symposium on VLSI | 3 |