EDBT 2026 Demo / reviewers in the wild / expert
Hao Jiang 0024
dblp:38/6049-24
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0003-2337-2539ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A 1024-Ch 583-nW/Ch Spike-Sorting SoC With Sparsity-Aware Spike Detection Scratchpad and Ultra-Low-Leakage Dual-Voltage 5T-SRAM for 16K-Template ClusteringabstractThis paper presents an energy-efficient spike-sorting system-on-chip (SoC) designed for closed-loop brain-computer interfaces of massive probing channels. The design first incorporates a sparsity/similarity-aware spike detection scratchpad, leveraging a bit-wise differential encoder and zero-friendly read-out circuits, reducing the dynamic power consumption of spike detection by 77.7%. To mitigate static power dissipation, it also introduces an ultra-low-leakage dual-voltage 5T-SRAM array with level-shifter embedded sense amplifiers, achieving an 82.2% leakage power reduction of neural signal buffering by applying half$V_{DD}$on SRAM cells. Additionally, a memory hierarchy architecture combining on-chip SRAM and off-chip FeRAM, along with a firing-rate-based Osort for cluster template management, minimizes off-chip memory access to only 9.7% with a latency of$11.7\mu $s for 1024-channel spike sorting. A silicon prototype is fabricated in 28-nm CMOS technology, which achieves a power consumption of 583nW/channel and an area consumption of 0.0012mm2/channel. The chip supports real-time spike sorting with up to 16K templates,$21.3\times $greater than the state-of-the-art spike-sorting processor. Hao Jiang 0024, Zexing Chen, Jiajun Lu, Siqi He, Liangjian Lyu, Jiamin Xu, Shiwei Liu 0002, Yingping Chen, Chixiao Chen, Qi Liu 0010, Ming Liu 0022 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2025 | SDISC: A Spike-Driven Human-Machine Interface with In-Situ Computing for Real-Time Low-Power InteractionabstractFeature extraction and classification of bio-signals are crucial in human-machine interface (HMI), yet suffer from high delay and limited energy efficiency using conventional hardware. To mitigate this challenge, we propose an SDISC architecture, a neuromorphic HMI with the innovation from signal encoding, computing-in-memory (CIM) hardware, to algorithm-hardware co-optimization. The following strategies are implemented: (1) A spike-driven feature extractor, achieving > $10 \times$ sparser dataflow than frame-based method; (2) In-situ computing based on resistive random-access memory (RRAM), enabling energy-efficient (4.09 TOPS/W) spiking neural network (SNN) classifier; (3) A Spike-Activity-Distillation algorithm and an Aid-Loser-Only recovery scheme to alleviate the non-ideality of RRAM devices, ensuring SDISC maintains high accuracy ($\sim \mathbf{9 8. 0 \%}$) in long time inference ($\boldsymbol{\gt} \mathbf{1 5}$ days). We further develop an end-to-end SDISC system for real-time EMG-based robot control, achieving a low latency ($34 \mu \mathrm{~s}$) and low power ($39.72 \mu \mathrm{~W} /$ sample) interaction on edge. Fangduo Zhu, Jingsong Zhang, Xumeng Zhang, Siyuan Ouyang, Chenyang, Hao Jiang 0024, Qi Liu 0010 |
DAC | 7 |
| 2024 | HARDSEA: Hybrid Analog-ReRAM Clustering and Digital-SRAM In-Memory Computing Accelerator for Dynamic Sparse Self-Attention in TransformerabstractSelf-attention-based transformers have outperformed recurrent and convolutional neural networks (RNN/ CNNs) in many applications. Despite the effectiveness, calculating self-attention is prohibitively costly due to quadratic computation and memory requirements. To solve this challenge, this article proposes a hybrid analog-ReRAM and digital-SRAM in-memory computing accelerator (HARDSEA), a computing-in-memory (CIM) accelerator supporting self-attention in transformer applications. To trade off between energy efficiency and algorithm accuracy, HARDSEA features an algorithm-architecture-circuit codesign. A product-quantization-based scheme dynamically facilitates self-attention sparsity by predicting lightweight token relevance. A hybrid in-memory computing architecture employs both high-efficiency analog ReRAM-CIM and high-precision digital SRAM-CIM to implement the proposed new scheme. The ReRAM-CIM, whose precision is sensitive to circuit nonidealities, takes charge of token relevance prediction where only computing monotonicity is demanded. The SRAM-CIM, utilized for exact sparse attention computing, is reorganized as an on-memory-boundary computing scheme, thus adapting to irregular sparsity patterns. In addition, we propose a time-domain winner-take-all (WTA) circuit to replace the expensive ADCs in ReRAM-CIM macros. Experimental results show that HARDSEA prunes BERT and GPT-2 models to 12%–33% sparsity without accuracy loss, achieving$13.5\times $–$28.5\times $speedup and$291.6\times $–$1894.3\times $energy efficiency over GPU. Compared to state-of-the-art transformer accelerators, HARDSEA has$1.2\times $–$14.9\times $better energy efficiency at the same level of throughput. Shiwei Liu 0002, Chen Mu, Hao Jiang 0024, Yunzhengmao Wang, Jinshan Zhang 0006, Keji Zhou, Qi Liu 0010, Chixiao Chen |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2023 | TiPU: A Spatial-Locality-Aware Near-Memory Tile Processing Unit for 3D Point Cloud Neural NetworkabstractEnergy-efficient 3D point cloud neural network accelerators are desired for autonomous driving and AR/VR applications. This paper proposes TiPU, a spatial-locality-aware near-memory tile processing unit where the point clouds are partitioned into tiles to process spatial features locally. Intra-tile farthest point sampling and cross-tile neighbor search are employed to avoid unnecessary distance computing. To efficiently facilitate the tile operations, TiPU architecture consists of a tile-based unified distance computing unit, a near-CAM feature extractor, and a near-SRAM-computing MLP engine. The experimental results show that, compared to GPU implementation, TiPU achieves 15.7× processing speed and reduces 7308× energy consumption. Jiapei Zheng, Hao Jiang 0024, Xinkai Nie, Zhangcheng Huang 0001, Chixiao Chen, Qi Liu 0010 |
DAC | 2 |
| 2023 | A $2.53 \mu \mathrm{W}/\text{channel}$ Event-Driven Neural Spike Sorting Processor with Sparsity-Aware Computing-In-Memory MacrosabstractSpike sorting processors with high energy efficiency are widely used in large-scale neural signal processing tasks to monitor the activity of neurons in brains. This paper presents a low-power processor for high-accuracy spike sorting and on-chip incremental learning using an algorithm-hardware co-design approach. The processor introduces an event-driven mechanism with adaptive-threshold detection to conditionally activate the system in order to reduce power consumption. Sparsity-aware computing-in-memory (CIM) macros are also developed in our design to store templates and perform complicated computations efficiently. The prototype is designed using 28nm technology with an area of 0.018 mm2/channel and an overall power efficiency of$\mathbf{2.53} \mu \mathbf{W}/\mathbf{channel}$and 84nW/(channel.cluster) at the voltage of 0.72V. Moreover, the accuracy of the whole design can reach 94.5% in a 32-channel scenario. Hao Jiang 0024, Jiapei Zheng, Yunzhengmao Wang, Jinshan Zhang 0006, Haozhe Zhu, Liangjian Lyu, Yingping Chen, Chixiao Chen, Qi Liu 0010 |
ISCAS | 1 |
| 2023 | A Scalable Die-to-Die Interconnect with Replay and Repair Schemes for 2.5D/3D IntegrationabstractChiplet is a critical technology in the post-Moore era, and the die-to-die (D2D) interconnect is essential for communication between chiplets. Meanwhile, several edge-computing devices based on 2.5D/3D chiplet have recently emerged. However, a lightweight D2D interconnect for 2.5D/3D edge-computing systems is lacking. Given the differences between 2.5D/3D integration, a scalable D2D interconnect with replay and repair schemes is presented in this paper. A credit-based flow control scheme and a custom replay scheme are presented for high efficiency. An effective detection and repair scheme is proposed to enhance fault tolerance for the D2D interconnect. Compared with a previous D2D interconnect design, the proposed D2D interconnect delivers 1.07/1.09Gbps throughput ($\sim 2.4\times/\sim 3.9\times \text{for}\ \text{write}/\text{read}$) and significantly reduced energy/bit with only ∼1.7× increased hardware cost. Additionally, compared with a previous chip-to-chip interconnect design, the proposed D2D interconnect can be configured down to power consumption as low as 0.55pJ/bit and 38.40Gbps throughput, achieving ∼2.5× throughput and significantly reduced latency with a negligible increase in hardware cost. Bo Jiao 0003, Jinshan Zhang 0006, Shiwei Liu 0002, Hao Jiang 0024, Jun Tao 0001, Wenning Jiang, Qi Liu 0010, Lihua Zhang 0002, Haozhe Zhu, Chixiao Chen |
ISCAS | 5 |