EDBT 2026 Demo / reviewers in the wild / expert
Jinghai Wang
dblp:154/7532
· DBLP profile ↗
10ranked-venue papers
2as first author
10since 2021 · last 2026
0009-0006-5263-8120ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 2 first-author · 8 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SynapseHD: A unified training framework for bridging spiking neural networks and hyperdimensional computing
Lingfeng Zhou, Huiyao Wang, Jinghai Wang, Zhiyi Yu, Shanlin Xiao |
Neurocomputing | 4 |
| 2025 | SCSC: Leveraging Sparsity and Fault-Tolerance for Energy-Efficient Spiking Neural NetworksabstractSpiking neural networks (SNNs) are more energy-efficient for processing sparse spike signals and demonstrate better fault tolerance compared to artificial neural networks (ANNs). In neuromorphic chips, synaptic weight access and neuron computation operations constitute 75%-95% of the chip's energy consumption. Therefore, our primary strategy to achieve highly energy-efficient SNNs is to enhance network sparsity while leveraging SNNs' high fault tolerance to reduce both weight access and neuron computation energy. The coding module is an essential component of SNNs, responsible for encoding non-spiking inputs into spike trains. However, previous coding schemes often exhibit poor sparsity or fault tolerance performance. Thus, we propose a novel coding scheme for SNNs: spiking convolutional sparse coding (SCSC). SCSC utilizes convolutional kernels as dictionaries and achieves sparsity through neural layers. Additionally, dynamic firing thresholds in neural layers balance sparsity with network performance and fault tolerance. The experimental results indicate that SCSC can increase network sparsity by 10%-20% and achieves higher accuracy than baseline networks when dealing with disturbances. Furthermore, we utilize approximate DRAM to store synaptic weights and selectively deactivate specific neuronal computing modules. With only a 1% decrease in accuracy, SCSC can reduce synaptic weight access energy by 29% and neuronal computing energy by 49%. Wei Liu 0118, Jinghai Wang, Zhiyi Yu, Shanlin Xiao |
ASP-DAC | 4 |
| 2025 | Towards In-Situ Neuromorphic Computing Architecture for Event Stream Super-ResolutionabstractEvent-based cameras, with their unique event stream representation, effectively mitigate motion blur in highspeed, high-exposure scenarios but suffer from low spatial resolution. To address this, we propose a super-resolution hardware accelerator for event streams based on Spiking Neural Networks (SNNs). In terms of network architecture, we incorporate hardware-friendly algorithmic designs by simplifying neuron models and optimizing convolution operations. On the hardware side, the design adopts a hierarchical structure featuring a highly parallel computational array. Additionally, by proposing a Kernel-Channel-Timestamp-Row (KCTR) dataflow and dual-pipeline structure, the design achieves in-situ computing, eliminating intermediate storage within layers and significantly reducing inter-layer spike storage. Evaluations on the N-MNIST and ASL-DVS datasets demonstrate root mean square errors (RMSE) of $\mathbf{1. 2 9 6}$ and $\mathbf{0. 1 2 1}$ for reconstructed super-resolution event streams. In downstream applications, the classification accuracies reach 98.84% and 99.73%, respectively. The proposed accelerator, designed using a 28 nm CMOS process, improves reconstruction speed by 95.6% compared to a GPU, operates at 500 MHz, and consumes only 0.546 pJ per synaptic operation. Yihe Yu, Wei Liu 0118, Jinghai Wang, Zhiyi Yu, Shanlin Xiao |
DAC | 5 |
| 2025 | FDAIMC: A Fully-Differential Analog In-Memory-Computing for MAC in MRAM with Accuracy Calibration Under Process and Voltage VariationabstractAnalog in-memory-computing (AIMC) is adopted extensively in non-volatile memory for multibit multiply-and-accumulate (MAC) operation. However, the low-on/off-ratio feature of magnetic tunnel junction (MTJ) impedes a high-performance AIMC macro based on spin transfer torque magnetic random access memory (STT-MRAM). Secondly, because of the uncertainty feature of a mixed-signal system under process and voltage variation, a calibration support is indispensable. Moreover, the incompatibility between a nonlinear analog signal and a linear digital signal hinders accurate computation and calibration support. To overcome these challenges, this work proposes a STT-MRAM-AIMC macro featuring: 1) a 2-level-differential cell array and a linear computing scheme with a calibration support in analog domain; 2) an analog-digital-conversion (ADC) system, including a slew-rate-independent voltage-to-time converter (SRIVTC) scheme and a self-triggered time-to-MAC value converter (STTMC) scheme; 3) a compact layout design for high area efficiency. Finally, an average accuracy of 95.44% is obtained under the TT&0.9V corner. By using the calibration strategy, the average accuracy of 97.8% and 88.6% are obtained under FF&0.945V and SS&0.855V separately, with over 30% enhancement. Furthermore, a 1.64~21.18 times area FoM than state of the art is obtained. An energy efficiency of 87.2~312.4 TOPS/W is obtained. Weichong Chen, Ruida Hong, Jinghai Wang, Ningyuan Yin, Zhiyi Yu |
DATE | 4 |
| 2025 | Stair-LIF: Boosting the Representation of Spiking Neural Networks with Learnable Incremental Multi-Threshold NeuronsabstractSpiking neural networks (SNNs) have shown remarkable potential in processing spatio-temporal data by mimicking biological neuronal mechanisms and achieving low computational costs. However, previous SNNs often rely on neuron models with fixed and single threshold voltages and binary spikes across layers during training, which limits their capacity for accurate information representation and reduces their biological plausibility. Inspired by the diversity of neuronal behaviors in different brain regions, we propose a novel neuron called Stair-LIF, which introduces learnable incremental multi-threshold mechanisms to enhance neuronal representational capacity and utilizes multi-spike firing to improve the precision of information transmission. Furthermore, we propose a channel-wise parameterization method to expand representational capacity among Stair-LIF. Experimental results on static datasets (CIFAR-10, CIFAR-100) and neuromorphic dynamic datasets (CIFAR10-DVS and DVS128 Gesture) demonstrate that the Starir-LIF neuron achieves state-of-the-art performance. Jilong Luo, Yinsheng Chen, Jinghai Wang, Zhiyi Yu, Shanlin Xiao |
ICME | 4 |
| 2025 | ASNA-Flow: An Efficient Asynchronous Neuromorphic Accelerator for Real-Time Event-Based Optical FlowabstractOptical flow estimation constitutes a fundamental computational challenge in computer vision, with critical applications object trajectory prediction, depth reconstruction, and autonomous navigation systems. The emergence of neuromorphic vision systems, integrating event-driven cameras with spiking neural networks (SNNs), has recently gained attention as a promising paradigm for edge deployment of optical flow estimation due to their advantages in ultralow power and resource efficiency. However, current neuromorphic computing platforms lack specialized architectures optimized for this problem domain. Existing implementations either prioritize configurable architectures at the expense of energy efficiency or employ intricate hardware control mechanisms to manage the asynchronous and sparse computing patterns inherent in SNNs. To address these limitations, we present ASNA-Flow, an event-driven asynchronous neuromorphic accelerator featuring a pioneering algorithm–hardware co-design framework specifically tailored for event-based optical flow estimation. Our methodology encompasses three key innovations: 1) a hardware-aware algorithm optimization that maintains computational fidelity while enhancing implementation efficiency; 2) systematic data pattern analysis to inform architectural decisions; and 3) novel exploitation of optical flow’s spatial locality characteristics to enable efficient sparse computing. Implemented in TSMC 28-nm CMOS technology, ASNA-Flow achieves real-time performance of 104 frames per second (FPS) with ultralow power consumption of 7.9 mW, demonstrating superior energy efficiency of 0.3 pJ per synaptic operation (SOP). This work establishes the first dedicated neuromorphic computing solution that simultaneously addresses the temporal sparsity, event-driven processing, and energy constraints inherent in optical flow estimation tasks. Jinghai Wang, Jilong Luo, Lingfeng Zhou, Zhiyi Yu, Shanlin Xiao |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2025 | An End-to-End Bundled-Data Asynchronous Circuits Design Flow: From RTL to GDSabstractAsynchronous circuits with low power and robustness are revived in emerging applications such as the Internet of Things (IoT) and neuromorphic chips, thanks to clock-less and event-driven mechanisms. However, the lack of mature computer-aided design (CAD) tools for designing large-scale asynchronous circuits results in low design efficiency and high cost. This article proposes an end-to-end bundled-data (BD) asynchronous circuit design flow, which can facilitate building asynchronous circuits, even if the designer has little or no asynchronous circuit foundation. Three features that enable this are: 1) a lightweight circuit converter developed in Python can convert circuits from synchronous descriptions to corresponding asynchronous ones at register transfer level (RTL). Desynchronization flow helps designers maintain a “synchronization mentality” to construct asynchronous circuits; 2) a synchronization-like verification method is proposed for asynchronous circuits so that it can be functionally verified before synthesis. Avoids the risk of rework after logic defects are discovered during the synthesis and implementation, as asynchronous circuits often cannot be simulated until gate-level (GL) netlist generation; and 3) the whole implementation flow from RTL to graphic data system (GDS) is based on commercial electronic design automation (EDA) tools. Similar to the design flow of synchronous circuits, it helps designers implement asynchronous circuits with “synchronization habits.” Furthermore, to validate this methodology, two asynchronous processors were, respectively, implemented and evaluated in the TSMC 28-nm CMOS process. Compared to their synchronous counterparts, the general-purpose asynchronous RISC-V processor achieves 20.5% power savings. And the domain-specific asynchronous spiking neural network (SNN) accelerator achieves 58.46% power savings and$2.41\times $energy efficiency improvement at 70% input spike sparsity. Jinghai Wang, Shanlin Xiao, Jilong Luo, Lingfeng Zhou, Zhiyi Yu |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2024 | An Efficient Asynchronous Circuits Design Flow with Backward Delay Propagation ConstraintabstractAsynchronous circuits have recently become more popular in Internet of Things (IoT) and neural network chips because of their potential low power consumption. However, due to the lack of Electronic Design Automation (EDA) tools, the asynchronous circuits design efficiency remains low and faces challenges in large-scale applications. This paper proposes a new asynchronous circuits design flow using traditional EDA tools, and applies a new backward delay propagation constraint (BDPC) method. In this method, control paths and data paths are tightly coupled and analyzed together to improve the accuracy of static timing analysis. Compared to previous works, the proposed design flow and constraint method offer significant advantages in terms of accuracy and efficiency. To verify this flow, an asynchronous RISC-V processor was implemented on TSMC 65nm process. Compared to synchronous version, asynchronous processor achieves a power optimization of 17.4 % while main-taining the same speed and area. Lingfeng Zhou, Shanlin Xiao, Huiyao Wang, Jinghai Wang, Zeyang Xu, Zhiyi Yu |
DATE | 4 |
| 2024 | Better-Than-Worst-Case: A Frequency Adaptation Asynchronous RISC-V Core With Vector ExtensionabstractIn recent years, asynchronous circuits have become more popular in neural network chips and the Internet of Things (IoT) due to their potential advantages of low-power consumption and high performance. However, the existing design methods for asynchronous circuits are still constrained by critical paths, increasing power consumption and hindering the further improvement of performance. In this article, a fully digital design method for frequency adaptation asynchronous bundled-data (BD) circuits is proposed. The proposed method is straightforward, effective, widely applicable, and independent of asynchronous controllers. It allows to automatically work on different frequencies as required, which can improve performance and reduce power consumption, achieving better-than-worst-case. To verify the proposed method, an asynchronous RISC-V processor with vector acceleration extension is designed on both TSMC 65-nm process and field-programmable gate array (FPGA) platform. According to the postlayout simulation results, compared with its synchronous version, the asynchronous processor achieves a 10% speed improvement (from 227.3 to 250 MHz) with a 37% power reduction (from 135 to 85$\mu$W/MHz) under ideal conditions. Even under the worst conditions, the asynchronous processor achieves equivalent performance to the synchronous processor, while still reducing power consumption by 29% (from 133 to 95$\mu$W/MHz). On the FPGA platform, asynchronous processor also achieves higher speed while lower power consumption. Lingfeng Zhou, Shanlin Xiao, Huiyao Wang, Jinghai Wang, Zeyang Xu, Zhiyi Yu |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2024 | Toward Efficient Asynchronous Circuits Design Flow Using Backward Delay Propagation ConstraintabstractIn recent years, asynchronous circuits have gained attention in neural network chips and Internet of Things (IoT) due to their potential advantages of low power and high performance. However, design efficiency of asynchronous circuits remains low and faces challenges in large-scale applications because of the lack of electronic design automation (EDA) support. This article presents a new bundled-data (BD) asynchronous circuits’ design flow using traditional EDA tools, including a new backward delay propagation constraint (BDPC) method. In this method, control paths and data paths are analyzed together in a tightly coupled approach to improve the accuracy of static timing analysis (STA). Compared with other design flows, the proposed design flow and constraint method show significant advantages in aspects of STA accuracy, design efficiency, and design applicability, and solving the congestion issues of field-programmable gate array (FPGA) in a previous work. An asynchronous RISC-V processor was implemented to verify the method, with selective handshake technology to further reduce power. Compared with the synchronous processor, the asynchronous processor achieves a 17.4% power optimization on the TSMC 65-nm process and a 48.3% dynamic power savings on the FPGA while maintaining the same frequency and resource utilization. Lingfeng Zhou, Shanlin Xiao, Huiyao Wang, Jinghai Wang, Zeyang Xu, Zhiyi Yu |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |