VLDB 2026 Research / reviewers in the wild / expert
Chaoming Fang
dblp:309/0611
· DBLP profile ↗
11ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0001-8830-1294ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Heterogeneous Decision Spiking Transformer Accelerator with Locality-dependent KV Product Cache and Compute Pattern Reconfigurable Engine
Ziyang Shen, Zhipeng Liao, Sitan Shen, Chaoming Fang, Fengshi Tian, Jie Yang 0033, Mohamad Sawan |
ISCAS | 4 |
| 2025 | An Area-Efficient and Bit-Width Configurable Carry-Save Adder Tree for Spiking TransformersabstractSpiking transformers have been successfully applied to multiple applications with comparable accuracy with native transformers. Achieving a high energy efficiency with spiking transformers requires dedicated hardware design, especially a specialized matrix multiplication engine optimized for spike input. In this paper, we propose a carry-save adder (CSA) array with an improved energy and area efficiency for spiking transformer matrix multiplication computation. A two-staged CSA structure is proposed to support maximum logic reuse between 1b self-attention mode and 8b linear mode. Besides, a cubic-mesh architecture is proposed to organize CSA trees to reuse weight in different timesteps. Compared to a baseline accumulation design, the proposed architecture achieved a 2.7x area reduction and a 3.75x power reduction with the same throughput, showing that the optimized computation array has great potential to be applied in digital neuromorphic accelerators. Chaoming Fang, Ziyang Shen, Fengshi Tian, Jie Yang 0033, Mohamad Sawan |
ISCAS | 1 |
| 2025 | NeuroEye: A 54.59mW, 12200FPS Event-Driven Near-Sensor Eye-Tracking Processor with Pipelined Spatial-Temporal Spike-StreamingabstractThis paper presents a design of an eye tracking system based on neuromorphic computing to enhance user interaction in augmented reality (AR) and virtual reality (VR) environments. Traditional methods face challenges of high computational demands and power consumption. To address these issues, we propose a fully-spike eye-tracking system that utilizes dynamic vision sensors (DVS) for asynchronous pixel-level change detection, thereby reducing data redundancy and improving temporal resolution. We proposed a pipelined processor specifically tailored for handling DVS events and Spiking Neural Network (SNN) computations. Our spatial-temporal spike-streaming architecture enables cascaded computation across all layers, achieving high energy efficiency and high frame rate in eye-tracking tasks. Implemented in a 40nm CMOS process, NeuroEye demonstrates up to 12200 frame-per-second (FPS) and 4.47uJ/frame energy efficiency with 54.59mW power consumption in post-layout evaluations. Jiakun Zheng, Fengshi Tian, Jinbo Chen 0002, Chaoming Fang, Jie Yang 0033, Mohamad Sawan, Kwang-Ting Cheng, Chi-Ying Tsui |
ISCAS | 4 |
| 2025 | BoostViT: Booth-Serial Skipping and Tunable Scaling for Vision TransformersabstractVision Transformers (ViTs) have emerged as a dominant architecture in computer vision (CV), surpassing conventional neural network counterparts across diverse visual tasks. Despite their exceptional performance, ViTs incur substantial computational overhead characterized by high memory footprint, long inference latency, and elevated energy consumption. Current acceleration strategies for ViTs primarily focus on pruning operations or leveraging the inherent sparsity, requiring complex address control logic or position encoding. Alternatively, some software-based approaches attempt to pre-compute and separate dense and sparse matrix position encoding, while hardware solutions typically spend additional time and resources to obtain position encoding, decomposing matrix multiplications into structured forms. Through an analysis of ViTs’ parameters, we found that approximately 91.03% of the most significant bits (MSBs) are either 111s or 000s, and nearly 45% of 3 adjacent bits are identical. To leverage this characteristic of ViTs, we propose the Booth-Serial Skipping algorithm, which transforms the computation of consecutive 111 or 000 sequences into skip steps that require no additional computation time. Furthermore, the 4th to 6th bits of ViT weights can undergo aggressive scaling, enhancing the likelihood of Booth-skip operations with minimal impact on accuracy. The key innovation of this paper lies in exploiting the high proportion of naturally consecutive 0s or 1s in 8-bit weights during ViT inference and further expanding the skippable range through the Tunable Scaling strategy. At the hardware level, we develop a specialized accelerator to coordinate the proposed acceleration strategies. The processing element array in the accelerator is optimized for general matrix multiplication, it not only significantly improves the computation of multi-head self-attention but also enables resource reuse for linear transformations, ultimately optimizing end-to-end inference. Our design achieves$50.3\times $,$21.9\times $,$17.37\times $,$7.47\times $, and$1.49\times $an average end-to-end speedup on DeiT over CPU (Intel Xeon Gold 6152), EdgeGPU (NVIDIA Jetson Xavier NX), GPU (TITAN Xp), ViTCoD, and ViT-slice, respectively. Shiqi Zhao 0001, Chaoming Fang, Fengshi Tian, Jinbo Chen 0002, Changzeng Fu, Jie Yang 0033, Mohamad Sawan |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2024 | Exploring Effective Stimulus Encoding via Vision System Modeling for Visual ProsthesesabstractVisual prostheses are potential devices to restore vision for blind people, which highly depends on the quality of stimulation patterns of the implanted electrode array. However, existing processing frameworks prioritize the generation of stimulation while disregarding the potential impact of restoration effects and fail to assess the quality of the generated stimulation properly. In this paper, we propose for the first time an end-to-end visual prosthesis framework (StimuSEE) that generates stimulation patterns with proper quality verification using V1 neuron spike patterns as supervision. StimuSEE consists of a retinal network to predict the stimulation pattern, a phosphene model, and a primary vision system network (PVS-net) to simulate the signal processing from the retina to the visual cortex and predict the firing rate of V1 neurons. Experimental results show that the predicted stimulation shares similar patterns to the original scenes, whose different stimulus amplitudes contribute to a similar firing rate with normal cells. Numerically, the predicted firing rate and the recorded response of normal neurons achieve a Pearson correlation coefficient of 0.78. Chuanqing Wang, Di Wu 0057, Chaoming Fang, Jie Yang 0033, Mohamad Sawan |
ICLR | 3 |
| 2024 | Accelerating BPTT-Based SNN Training with Sparsity-Aware and Pipelined ArchitectureabstractOn-chip learning of Spiking Neural Networks (SNN) has been extensively researched to enhance adaptability and privacy protection, with Back-Propagation-Through-Time (BPTT) emerging as the top-performing method despite its resourceintensive nature. In this paper, we propose a dedicated training processor that accelerates the BPTT algorithm for SNNs. We analyze the bottlenecks and optimization opportunities in SNN- BPTT and introduce novel techniques such as recalculation of membrane potentials to reduce redundant data movement. Additionally, we implement a pipeline architecture with heterogeneous computing cores to maximize hardware utilization and parallelism. Exploiting three types of sparsity in BPTT allows us to skip unnecessary computations and memory access, further optimizing performance. The proposed processor, implemented using 40nm CMOS technology, achieves simulation results with an advanced training energy efficiency of 0.86pJ/OP. Chaoming Fang, Fengshi Tian, Jie Yang 0033, Mohamad Sawan |
ISCAS | 1 |
| 2024 | BOLS: A Bionic Sensor-direct On-chip Learning System with Direct-Feedback-Through-Time for Personalized Wearable Health MonitoringabstractPrecise bio-signal classification techniques for edge healthcare have been extensively researched, yet the scalability and efficiency of existing studies remain constrained by challenges in sensing, learning, and processing. Additionally, a deficiency in cross-level integration for the development of comprehensive healthcare systems has been observed. To tackle these issues and facilitate ultra-efficient personalized edge healthcare, this paper introduces the pioneering bionic sensor-direct on-chip learning and inference system with direct-feedback-through-time for user-specific cardiac arrhythmia detection, termed BOLS. This innovative system encompasses a compact sensor-direct feature extractor and a pipelined bionic processor, enabling end-to-end on-chip learning and inference. Employing cross-level co-design, our proposed bionic on-chip learning approach attains exceptional classification performance, boasting an accuracy of 98.6%, which ranks among the highest. The entire system has been implemented using 40nm CMOS process and subsequently verified. Remarkably, the proposed BOLS system consumes a mere 1.18mW for inference and 2.57mW for learning, resulting in an impressive power saving of over ×2000 compared to existing commercial training platforms. Fengshi Tian, Jiakun Zheng, Jingyu He, Jinbo Chen 0002, Chaoming Fang, Jie Yang 0033, Mohamad Sawan, Chi-Ying Tsui, Kwang-Ting Cheng |
ISCAS | 6 |
| 2023 | NBSSN: A Neuromorphic Binary Single-Spike Neural Network for Efficient Edge IntelligenceabstractNeuromorphic computing approaches such as Spiking Neural Networks (SNN) have been increasingly adopted in bio-signal processing and interpretation due to its intrinsic neurodynamic attribute. Nevertheless, reconciling performance and power efficiency in SNN implementation is still a bottleneck. Single-spike neural coding scheme, which is an extremely sparse coding scheme, provides a solution to bridge the gap. In this work, a neuromorphic architecture, using binary single spike neural signals, is proposed with both algorithm and hardware implementation. A sparsity-aware spatial-temporal back-propagation training method is proposed together with a single-spike coding scheme. Also, a novel neuromorphic accelerator is co-designed with algorithmic optimization and implemented in 40nm CMOS process. Experimental results show that the proposed processor reaches an accuracy of 94.61% on the MNIST dataset, 93.59% on the N-MNIST dataset, and 93.27% on the ECG dataset, respectively, while consumes$0.173\mu\mathrm{J}$per ECG classification task and 0.16mm2on-chip area. The overall power consumption is reduced by 91.68% compared to the state-of-the-art systems. Ziyang Shen, Fengshi Tian, Chaoming Fang, Xiaoyong Xue, Jie Yang 0033, Mohamad Sawan |
ISCAS | 4 |
| 2023 | SpikeSEE: An energy-efficient dynamic scenes processing framework for retinal prosthesesabstractIntelligent and low-power retinal prostheses are highly demanded in this era, where wearable and implantable devices are used for numerous healthcare applications. In this paper, we propose an energy-efficient dynamic scenes processing framework (SpikeSEE) that combines a spike representation encoding technique and a bio-inspired spiking recurrent neural network (SRNN) model to achieve intelligent processing and extreme low-power computation for retinal prostheses. The spike representation encoding technique could interpret dynamic scenes with sparse spike trains, decreasing the data volume. The SRNN model, inspired by the human retina's special structure and spike processing method, is adopted to predict the response of ganglion cells to dynamic scenes. Experimental results show that the Pearson correlation coefficient of the proposed SRNN model achieves 0.93, which outperforms the state-of-the-art processing framework for retinal prostheses. Thanks to the spike representation and SRNN processing, the model can extract visual features in a multiplication-free fashion. The framework achieves 8 times power reduction compared with the convolutional recurrent neural network (CRNN) processing-based framework. Our proposed SpikeSEE predicts the response of ganglion cells more accurately with lower energy consumption, which alleviates the precision and power issues of retinal prostheses and provides a potential solution for wearable or implantable prostheses. Chuanqing Wang, Chaoming Fang, Jie Yang 0033, Mohamad Sawan |
Neural Networks | 2 |
| 2022 | A Compact Online-Learning Spiking Neuromorphic Biosignal ProcessorabstractReal-time biosignal processing on wearable devices has attracted worldwide attention for its potential in healthcare applications. However, the requirement of low-area, low-power and high adaptability to different patients challenge conventional algorithms and hardware platforms. In this design, a compact online learning neuromorphic hardware architecture with ultralow power consumption designed explicitly for biosignal processing is proposed. A trace-based Spiking-Timing-Dependent-Plasticity (STDP) algorithm is applied to realize hardware-friendly online learning of a single-layer excitatory-inhibitory spiking neural network. Several techniques, including event-driven architecture and a fully optimized iterative computation approach, are adopted to minimize the hardware utilization and power consumption for the hardware implementation of online learning. Experiment results show that the proposed design reaches the accuracy of 87.36% and 83% for the Mixed National Institute of Standards and Technology database (MNIST) and ECG classification. The hardware architecture is implemented on a Zynq-7020 FPGA. Implementation results show that the Look-Up Table (LUT) and Flip Flops (FF) utilization reduced by 14.87 and 7.34 times, respectively, and the power consumption reduced by 21.69% compared to state of the art. Chaoming Fang, Ziyang Shen, Fengshi Tian, Jie Yang 0033, Mohamad Sawan |
ISCAS | 1 |
| 2022 | NIMBLE: A Neuromorphic Learning Scheme and Memristor Based Computing-in-Memory Engine for EMG Based Hand Gesture RecognitionabstractEMG based hand gesture recognition on convolutional neural networks (CNNs) has been widely learned, which gains high accuracy. However, CNN based systems are computationally complex and power consuming, thus hard to be deployed at edge. Biologically inspired, a new neuromorphic learning and computing approach for electromyogram (EMG) based hand gesture recognition tasks is proposed in this work. This approach designs an activate and inhibit joint processing spiking neural network (AIPS-SNN) which reaches an accuracy of 85.6% on Nina Pro dataset. Furthermore, the AIPS-SNN is deployed on the proposed memristor based computation in-memory (CIM) system, the power efficiency and area efficiency of which reach 10.146 TOPS/W and 35.399 GOPS/mm2, respectively. The experimental results indicate that the proposed neuromorphic CIM engine is promising for edge deployment. Fengshi Tian, Jinhao Liang, Jiahe Shi, Chaoming Fang, Hui Wu 0010, Xiaoyong Xue, Xiaoyang Zeng |
ISCAS | 6 |