EDBT 2026 Demo / reviewers in the wild / expert
Haoming Chu
dblp:228/3454
· DBLP profile ↗
5ranked-venue papers
0as first author
4since 2021 · last 2026
0000-0002-4749-0954ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | WaferBRAIN: Whole-Brain Scale Neuromorphic Architecture Based on Wafer-Scale Integration
Yukun Feng, Liangyu Gan, Haoming Chu, Yufan He, Jiaxin Yin, Lirong Zheng 0001, Yuxiang Huan |
ISCA | 4 |
| 2026 | PicoSleepNet: An Ultra Lightweight Sleep Stage Classification by Spike Neural Network Using Single-Channel EEG SignalabstractThis study introduces PicoSleepNet, an ultra-lightweight sleep stage classification method that utilizes a spiking neural network (SNN) with single-channel electroencephalogram (EEG) signals. Traditional methods use multi-bit Nyquist sampling and dense computing, which result in high complexity and power consumption, hindering their deployment on wearable devices. To address these limitations, we propose an innovative pipeline combining single-bit sub-Nyquist level-crossing sampling (LCS) and sparse computing based on SNN. First, LCS adaptively encodes EEG signals into event-driven spike sequences, reducing data volume by 6.98× while preserving essential signal characteristics compared to Nyquist sampling. Second, a sparse recurrent spiking neural network (RSNN) architecture, optimized by the masked backpropagation and sparse regularization (Masked-BPSR) technique, improves performance and reduces computational costs. Third, quantization-aware training (QAT) ensures that the model maintains high accuracy with low-bitwidth quantization, significantly reducing computational power consumption and enabling hardware-friendly deployment. Compared with current state-of-the-art sleep staging approaches, PicoSleepNet achieves competitive performance on three public datasets (Sleep-EDF-20, Sleep-EDF-78, and ISRUC-Sleep) with accuracies of 83.5%, 77.9%, 79.4% and macro-F1 scores of 75.2%, 68.1%, 77.2%, respectively. Meanwhile, by leveraging the computational sparsity design of RSNN and the joint optimization of Masked-BPSR and QAT, PicoSleepNet achieves an ultra-lightweight model with only 14.0-25.8 K parameters (reduced by nearly 2×) and 681.4-842.0 K operations (reduced by 27×), reducing computational power consumption by 1480×. This approach demonstrates the feasibility of deploying ultra-lightweight sleep staging systems in wearable devices and neuromorphic hardware, paving the way for broader applications in real-time health monitoring. Shengnan Liu, Haoming Chu, Yukun Feng, Yulong Yan, Yuxiang Huan |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | A Low-Power Hybrid-Precision Neuromorphic Processor With INT8 Inference and INT16 Online Learning in 40-nm CMOSabstractIn this work, we present a neuromorphic processor for artificial intelligence of things (AIoT) applications featuring low-power consumption, a small footprint, STDP-based online learning, and the ability to adapt to multiple applications. Hybrid precision, i.e., INT8 for inference and INT16 for training, is suggested to achieve balanced accuracy and energy efficiency. A precision-configurable leaky integrate-and-fire(LIF) neuron unit and a unified memory architecture are designed to maximize datapath reuse. A dynamic pruning technique is proposed to exploit the temporal sparsity, yielding synaptic operations reduction by 3.68x in training and 1.63x in inference, respectively. The design is implemented and fabricated in a 40-nm CMOS process, with a core area of 0.87 mm2. It is measured to consume a minimal power of$680~\mu \text{W}$at 70 MHz under a 0.75 V power supply, corresponding to 9.9 pJ per synaptic operation. Evaluated with typical spatial, temporal, and spatiotemporal datasets (MNIST, MIT-BIH, and N-MNIST), the proposed design achieve energy efficiency comparable to the best-in-class solutions with handcrafted training and customized ASICs, while demonstrating improved versatility across multiple applications with balanced accuracy, power consumption, and model adaptability. Congyang Liu, Ziyi Yang 0014, Zikai Zhu, Haoming Chu, Yuxiang Huan, Lirong Zheng 0001, Zhuo Zou |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2021 | IECA: An In-Execution Configuration CNN Accelerator With 30.55 GOPS/mm² Area EfficiencyabstractIt remains challenging for a Convolutional Neural Network (CNN) accelerator to maintain high hardware utilization and low processing latency with restricted on-chip memory. This paper presents an In-Execution Configuration Accelerator (IECA) that realizes an efficient control scheme, exploring architectural data reuse, unified in-execution controlling, and pipelined latency hiding to minimize configuration overhead out of the computation scope. The proposed IECA achieves row-wise convolution with tiny distributed buffers and reduces the size of total on-chip memory by removing 40% of redundant memory storage with shared delay chains. By exploiting a reconfigurable Sequence Mapping Table (SMT) and Finite State Machine (FSM) control, the chip realizes cycle-accurate Processing Element (PE) control, automatic loop tiling and latency hiding without extra time slots for pre-configuration. Evaluated on AlexNet and VGG-16, the IECA retains over 97.3% PE utilization and over 95.6% memory access time hiding on average. The chip is designed and fabricated in a UMC 55-nm process running at a frequency of 250 MHz and achieves an area efficiency of 30.55 GOPS/mm2and 0.244 GOPS/KGE (kilo-gate-equivalent), which makes an over$2.0\times $and$2.1\times $improvement, respectively, compared with that of previous related works. Implementation of the IEC control scheme uses only a 0.55% area of the 2.75 mm2core. Boming Huang, Yuxiang Huan, Haoming Chu, Jiawei Xu 0002, Lizheng Liu, Lirong Zheng 0001, Zhuo Zou |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2018 | TMR Group Coding Method for Optimized SEU and MBU Tolerant Memory DesignabstractThis work proposes a fault tolerant memory design using the method of Triple Module Redundancy (TMR) group coding to tolerant the Single-Event Upset (SEU) and Multi-Bit Upset (MBU) influence on memory devices in space environment. The group coding method uses different models to partition and code each word line in memory with Hamming code to achieve best performance. TMR group coding method further increases the capability of self-correction for the errors occurred in parity bits. The evaluation results show that the suggested approach can obtain improved correctness for the memory output with optimized tradeoff between reliability and cost. At 5% error rate, the probability of correct output reaches 70.78% with small cost increment. To achieve 90% reliability, the accuracy improvement is 31.9% compared to TMR with 9% increased area. This solution proposed is evaluated on the memory rich micro-coded processor, but can be further extended to other memory-based processors that need high reliability for the SEU and MBU influence in aerospace applications. Yi Jin 0007, Yuxiang Huan, Haoming Chu, Zhuo Zou, Lirong Zheng 0001 |
ISCAS | 3 |