EDBT 2026 Demo / reviewers in the wild / expert
Zhenhui Dai
dblp:249/3449
· DBLP profile ↗
9ranked-venue papers
1as first author
9since 2021 · last 2025
0000-0002-8711-6502ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Reconfigurable Digital Compute-In-Memory Heterogeneous Macro for Differential Frame Convolution and Spiking Neural NetworkabstractThe application of artificial neural network (ANN) in video processing encounters significant challenges, including large data volumes, numerous linear operations, and high power consumption. The fusion of convolutional neural network (CNN) and spiking neural network (SNN) provides a dual benefit of achieving high accuracy while maintaining low power consumption. However, ongoing challenges remain in minimizing multiply-accumulate (MAC) operations and optimizing data movement. In this work, we propose a reconfigurable digital compute-in-memory (RDCIM) heterogeneous macro without the sense amplifier, tailored for the diverse computational demands of CNN and SNN. To improve energy efficiency, differential frame convolution (DFC) is adopted to mitigate the computational overhead. In addition, computational resources are functionally reused to accommodate four data flow types, supporting both DFC and SNN operations. Implemented by TSMC 28nm technology, the proposed RDCIM heterogeneous macro achieves the peak energy efficiency of 29.13 TOPS/W for DFC and 0.56 pJ/SOP for SNN, operating at a frequency of 284 MHz. Li Lun, Zhenhui Dai, Yingying Cui, Xiaoxin Cui |
ISCAS | 3 |
| 2024 | An Energy-Efficient Differential Frame Convolutional Accelerator with on-Chip Fusion Storage Architecture and Pixel-Level Pipeline Data FlowabstractConvolutional neural networks require a huge amount of computation in video applications. For some specific tasks, such as surveillance, differential frame convolution reuses inter-frame data and significantly reduces multiplication and accumulation. However, there are still some challenges in improving energy efficiency of differential frame convolution on chips. Firstly, differential frame convolution brings additional on-chip storage for reusing inter-frame data. Secondly, in post-processing of differential frame convolution, there are more memory accessing and arithmetic logic operations. Therefore, sparse working mode is of vital importance for the post-processing. In response to these challenges, this work proposes an on-chip fusion storage architecture for energy-efficient differential frame convolution and a pixel-level pipeline data flow that supports the sparsity of features. The simulation of our accelerator implemented in 28nm CMOS can achieve energy efficiency by 3.09× compared with other state-of-the-art works Zhenhui Dai, Yi Zhong 0002, Kunyu Feng, Yuan Wang 0001, Dunshan Yu, Xiaoxin Cui |
ISCAS | 1 |
| 2024 | SPAT: FPGA-based Sparsity-Optimized Spiking Neural Network Training Accelerator with Temporal Parallel DataflowabstractSpiking neural networks (SNNs), as biologically inspired computational models, possess significant advantages in energy efficiency due to their event-driven operations. However, challenges remain in attaining high computational efficiency for SNN training. In this work, we propose a novel SNN training accelerator employing temporal parallelism and sparsity optimizations to achieve superior efficiency. A temporal parallel dataflow is designed to concurrently integrate spikes across multiple time steps, enhancing throughput and data reuse. To reduce latency and improve energy efficiency, we leverage the sparsity of SNNs and employ methods such as zero gating and zero skipping. Implemented on a field-programmable gate array (FPGA), the proposed training accelerator demonstrates 2.3-fold speedup and 15.7-fold energy reduction compared to NVIDIA A100 GPU on N-MNIST dataset. Li Lun, Mingqi Yin, Zhenhui Dai, Xiaole Cui, Xiaoxin Cui |
ISCAS | 6 |
| 2024 | A 16.41 TOPS/W CNN Accelerator with Event-Based Layer Fusion for Real-Time InferenceabstractThis paper proposes a convolutional neural network (CNN) accelerator architecture for real-time tasks in edge devices. An event-based layer fusion technique is adopted to eliminate on-chip storage requirements and off-chip data movement caused by features. Cross-layer pipeline is elaborated during layer fusion to obtain high throughput and low latency. An adaptive fully unrolling event-driven core is designed and a cyclic storage method is exploited to reduce the storage space for partial sum in the core. Modified LeNet is accelerated with the proposed architecture. The accelerator can reach an energy efficiency of 16.41 TOPS/W and a latency of 0.85μs under TSMC 28nm technology, and a frame rate of 369.4K FPS under FPGA. Li Lun, Zhenhui Dai, Xiaoxin Cui |
ISCAS | 3 |
| 2023 | Breath-Hold CBCT-Guided CBCT-to-CT Synthesis via Multimodal Unsupervised Representation Disentanglement LearningabstractAdaptive radiation therapy (ART) aims to deliver radiotherapy accurately and precisely in the presence of anatomical changes, in which the synthesis of computed tomography (CT) from cone-beam CT (CBCT) is an important step. However, because of serious motion artifacts, CBCT-to-CT synthesis remains a challenging task for breast-cancer ART. Existing synthesis methods usually ignore motion artifacts, thereby limiting their performance on chest CBCT images. In this paper, we decompose CBCT-to-CT synthesis into artifact reduction and intensity correction, and we introduce breath-hold CBCT images to guide them. To achieve superior synthesis performance, we propose a multimodal unsupervised representation disentanglement (MURD) learning framework that disentangles the content, style, and artifact representations from CBCT and CT images in the latent space. MURD can synthesize different forms of images using the recombination of disentangled representations. Also, we propose a multipath consistency loss to improve structural consistency in synthesis and a multidomain generator to improve synthesis performance. Experiments on our breast-cancer dataset show that MURD achieves impressive performance with a mean absolute error of 55.23±9.94 HU, a structural similarity index measurement of 0.721±0.042, and a peak signal-to-noise ratio of 28.26±1.93 dB in synthetic CT. The results show that compared to state-of-the-art unsupervised synthesis methods, our method produces better synthetic CT images in terms of both accuracy and visual quality. Chuanpu Li, Zhenhui Dai, Liming Zhong, Xuetao Wang, Wei Yang 0006 |
IEEE Trans. Medical Imaging | 3 |
| 2022 | MRI-guided Automated Delineation of Gross Tumor Volume for Nasopharyngeal Carcinoma using Deep LearningabstractIn this paper, we propose a novel deep learning-based automatic delineation method of nasopharynx gross tumor volume (GTVnx) by combing computed tomography (CT) and magnetic resonance imaging (MRI) modalities. The purpose of this study is to explore whether MRI can provide additional information to improve the accuracy of delineation on CT. The proposed model can adaptively leverage the high contrast information of MRI into the automated delineation of GTVnx on CT in nasopharyngeal carcinoma (NPC) radiotherapy. In this study, the dataset collected from 192 patients with NPC was used to verify the performance of the proposed method. The average Dice Similarity Coefficient, 95% Hausdorff Distance and Average Symmetric Surface Distance of the segmentation results predicted by the proposed model are 0.7181, 9.6637mm, and 2.8014mm, respectively, which outperformed that of the single-modal and the concatenation-based multi-modal segmentation models. Meiyan Yue, Zhenhui Dai, Jiahui He 0003, Yaoqin Xie, Nazar Zaki, Wenjian Qin |
CBMS | 2 |
| 2022 | An Event-driven Spiking Neural Network Accelerator with On-chip Sparse WeightabstractSpiking neural networks (SNNs) have widely drew attention of recent research. With brain-spired dynamics and spike-based communication, SNN is supposed to be a more energy-efficient neural network than existing artificial neural network (ANN). To make better use of the temporal sparsity of spikes and spatial sparsity of weights in SNN, this paper presents a sparse SNN accelerator. It adopts a novel self-adaptive spike compressing and decompressing (SASCD) mechanism for different input spike sparsity, as well as on-chip compressed weight storage and processing. We implement the octa-core design on field programmable gate array (FPGA). The results demonstrate a peak performance of 35.84 GSOPs/s, which is equivalent to 358.4 GSOPs/s in dense SNN accelerators for 90% weight sparsity. For the single-layer perceptron model in rate coding implemented on the hardware, SASCD reduces the time step intervals from 2.15 $\mu$ s to 0.55 $\mu$ s. Yisong Kuang, Xiaoxin Cui, Chenglong Zou, Yi Zhong 0002, Zhenhui Dai, Zilin Wang 0001, Kefei Liu 0002, Dunshan Yu, Yuan Wang 0001 |
ISCAS | 5 |
| 2022 | ESSA: Design of a Programmable Efficient Sparse Spiking Neural Network AcceleratorabstractSpiking neural networks (SNNs) have been witnessing the developing trends to reduce the model size and improve the hardware efficiency for area- and energy-based applications, which are processed by model pruning and data compressions. However, it is challenging to exploit the unstructured sparsity of SNNs for the dense neuromorphic processors. In this article, we present an efficient sparse SNN accelerator (ESSA), which leverages both the temporal sparsity of spike events and the spatial sparsity of weights in SNN inference. It provides both the compressed weights for sparse SNNs and the uncompressed weights for compact SNNs. The self-adaptive spike compression is proposed for sparse spike scenarios, leading to the improvement of throughput by$3.2\times $. ESSA executes a flexible fan-in–fan-out tradeoff by using combinable dendrites, which overcomes the fan-in limitation in neuromorphic systems. Furthermore, a low-latency intrachip spike multicast method is adopted to reduce the resource overhead. Implemented on the Xilinx Kintex Ultrascale field-programmable gate array (FPGA), ESSA achieves an equivalent performance of 253.1 GSOP/s and an energy efficiency of 32.1 GSOP/W for 75% weight sparsity at 140 MHz. The implementation of a four-layer fully connected SNN is expected to perform$2.6~\mu \text{s}$per time step and the energy consumption is$14.6~\mu \text{J}$. Our results demonstrate that ESSA outperforms several state-of-the-art application-specific integrated circuit (ASIC) or FPGA neuromorphic processors. Yisong Kuang, Xiaoxin Cui, Zilin Wang 0001, Chenglong Zou, Yi Zhong 0002, Kefei Liu 0002, Zhenhui Dai, Dunshan Yu, Yuan Wang 0001, Ru Huang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2021 | A 28-nm 0.34-pJ/SOP Spike-Based Neuromorphic Processor for Efficient Artificial Neural Network ImplementationsabstractNeuromorphic hardware platforms inspired by human brain have emerged as novel non von Neumann computing architectures. They were proved excellent platforms for spiking neural network (SNN) implementations. However, implementing artificial neural networks (ANNs) on existing neuromorphic hardware platforms is still a daunting task because of critical limitations on coding scheme, maximum of fan-in, and highest weight precision in them. In this paper, we introduce a neuromorphic processor developed for various neural networks implementations including ANNs and SNNs. We employ spatio-temporal coding scheme based on spike events. By combining low-precision dendrites, the chip can implement weight precision between 1 bit and 8 bits and scalable fan-in. The 3.66-mm2chip fabricated in 28-nm CMOS with a maximum fan-in of 72 K per neuron demonstrates unprecedented compatibility with ANN applications compared to previously-proposed neuromorphic chips. Yisong Kuang, Xiaoxin Cui, Yi Zhong 0002, Kefei Liu 0002, Chenglong Zou, Zhenhui Dai, Dunshan Yu, Yuan Wang 0001, Ru Huang 0001 |
ISCAS | 6 |