EDBT 2026 Demo / reviewers in the wild / expert
Fengshi Tian
dblp:286/5428
· DBLP profile ↗
21ranked-venue papers
4as first author
21since 2021 · last 2026
0000-0003-1304-5593ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 4 first-author · 21 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Heterogeneous Decision Spiking Transformer Accelerator with Locality-dependent KV Product Cache and Compute Pattern Reconfigurable Engine
Ziyang Shen, Zhipeng Liao, Sitan Shen, Chaoming Fang, Fengshi Tian, Jie Yang 0033, Mohamad Sawan |
ISCAS | 5 |
| 2026 | BioSeek: A Design Generation Framework of Biosignal Processors with Large-Language Models for Edge Healthcare ApplicationsabstractDeep neural network (DNN)-based methodologies have shown impressive performance and robustness in the detection of abnormalities and decoding of multi-modal biosignals. While the use of DNNs provides promising classification and decoding capabilities, it also introduces significant design and cost challenges for the implementation of biomedical System on Chips (SoC). To address the increasing demand for advanced and efficient DNN-based healthcare solutions at the edge, we propose BioSeek, an agile design generation framework enhanced by cutting-edge large-language models (LLM). BioSeek offers a comprehensive solution to the design challenges associated with biosignal processors. The effectiveness of BioSeek is evaluated through the design generation of both application-specific and versatile biosignal processors, demonstrating performance that is competitive with existing solutions. Fengshi Tian, Jiakun Zheng, Hui Wu 0010, Zilu Liu, Jinbo Chen 0002, Shiqi Zhao 0001, Jie Yang 0033, Mohamad Sawan, Chi-Ying Tsui, Kwang-Ting Cheng |
ISCAS | 1 |
| 2025 | SynDCIM: A Performance-Aware Digital Computing-in-Memory Compiler with Multi-Spec-Oriented Subcircuit SynthesisabstractDigital Computing-in-Memory (DCIM) is an innovative technology that integrates multiply-accumulation (MAC) logic directly into memory arrays to enhance the performance of modern AI computing. However, the need for customized memory cells and logic components currently necessitates significant manual effort in DCIM design. Existing tools for facilitating DCIM macro designs struggle to optimize subcircuit synthesis to meet user-defined performance criteria, thereby limiting the potential system-level acceleration that DCIM can offer. To address these challenges and enable the agile design of DCIM macros with optimal architectures, we present SynDCIM - a performance-aware DCIM compiler that employs multi-spec-oriented subcircuit synthesis. SynDCIM features an automated performance-to-layout generation process that aligns with user-defined performance expectations. This is supported by a scalable subcircuit library and a multi-spec-oriented searching algorithm for effective subcircuit synthesis. The effectiveness of SynDCIM is demonstrated through extensive experiments and validated with a test chip fabricated in a 40nm CMOS process. Testing results reveal that designs generated by SynDCIM exhibit competitive performance when compared to state-of-the-art manually designed DCIM macros. Kunming Shao, Fengshi Tian, Jiakun Zheng, Jia Chen 0032, Jingyu He, Hui Wu 0010, Jinbo Chen 0002, Xihao Guan, Fengbin Tu, Jie Yang 0033, Mohamad Sawan, Kwang-Ting Cheng, Chi-Ying Tsui |
DATE | 2 |
| 2025 | An Area-Efficient and Bit-Width Configurable Carry-Save Adder Tree for Spiking TransformersabstractSpiking transformers have been successfully applied to multiple applications with comparable accuracy with native transformers. Achieving a high energy efficiency with spiking transformers requires dedicated hardware design, especially a specialized matrix multiplication engine optimized for spike input. In this paper, we propose a carry-save adder (CSA) array with an improved energy and area efficiency for spiking transformer matrix multiplication computation. A two-staged CSA structure is proposed to support maximum logic reuse between 1b self-attention mode and 8b linear mode. Besides, a cubic-mesh architecture is proposed to organize CSA trees to reuse weight in different timesteps. Compared to a baseline accumulation design, the proposed architecture achieved a 2.7x area reduction and a 3.75x power reduction with the same throughput, showing that the optimized computation array has great potential to be applied in digital neuromorphic accelerators. Chaoming Fang, Ziyang Shen, Fengshi Tian, Jie Yang 0033, Mohamad Sawan |
ISCAS | 3 |
| 2025 | Analysis and Prevention of Coupling-Dependent Data Flipping in Series-Series Resonant Wireless Power Transfer SystemsabstractLoad shift keying (LSK) is commonly used in wireless power transfer (WPT) systems for backscattering information from the receiver back to the transmitter. However, when the coupling coefficient (k) between the coupling coils falls below a threshold value (kDF), the demodulated LSK data can unexpectedly flip from '1' to '0' and '0' to '1'. This phenomenon is referred to as coupling-dependent data flipping (CDDF). This research investigates the factors that lead to CDDF in series-series resonant WPT systems, by taking into account of parasitic parameters of the coupled link and validates the analysis through SPICE simulations. To mitigate CDDF, we propose a carrier-frequency auto-tuning scheme that is also verified by simulation results. Sayan Sarkar, Fengshi Tian, Wing-Hung Ki, Chi-Ying Tsui, Yang Liu 0061 |
ISCAS | 3 |
| 2025 | A Flexible Precision Scaling Deep Neural Network Accelerator with Efficient Weight CombinationabstractDeploying mixed-precision neural networks on edge devices is friendly to hardware resources and power consumption. To support fully mixed-precision neural network inference, it is necessary to design flexible hardware accelerators for continuous varying precision operations. However, the previous works have issues on hardware utilization and overhead of reconfigurable logic. In this paper, we propose an efficient accelerator for 2 ∼ 8-bit precision scaling with serial activation input and parallel weight preloaded. First, we set two loading modes for the weight operands and decompose the weight into the corresponding bitwidths, which extends the weight precision support efficiently. Then, to improve hardware utilization of low-precision operations, we design the architecture that performs bit-serial MAC operation with systolic dataflow, and the partial sums are combined spatially. Furthermore, we designed an efficient carry save adder tree supporting both signed and unsigned number summation across rows. The experiment result shows that the proposed accelerator, synthesized with TSMC 28nm CMOS technology, achieves peak throughput of 4.09TOPS and peak energy efficiency of 68.94TOPS/W at 2/2-bit operations. Kunming Shao, Fengshi Tian, Kwang-Ting Cheng, Chi-Ying Tsui, Yi Zou 0001 |
ISCAS | 3 |
| 2025 | NeuroEye: A 54.59mW, 12200FPS Event-Driven Near-Sensor Eye-Tracking Processor with Pipelined Spatial-Temporal Spike-StreamingabstractThis paper presents a design of an eye tracking system based on neuromorphic computing to enhance user interaction in augmented reality (AR) and virtual reality (VR) environments. Traditional methods face challenges of high computational demands and power consumption. To address these issues, we propose a fully-spike eye-tracking system that utilizes dynamic vision sensors (DVS) for asynchronous pixel-level change detection, thereby reducing data redundancy and improving temporal resolution. We proposed a pipelined processor specifically tailored for handling DVS events and Spiking Neural Network (SNN) computations. Our spatial-temporal spike-streaming architecture enables cascaded computation across all layers, achieving high energy efficiency and high frame rate in eye-tracking tasks. Implemented in a 40nm CMOS process, NeuroEye demonstrates up to 12200 frame-per-second (FPS) and 4.47uJ/frame energy efficiency with 54.59mW power consumption in post-layout evaluations. Jiakun Zheng, Fengshi Tian, Jinbo Chen 0002, Chaoming Fang, Jie Yang 0033, Mohamad Sawan, Kwang-Ting Cheng, Chi-Ying Tsui |
ISCAS | 2 |
| 2025 | DIRC-RAG: Accelerating Edge RAG with Robust High-Density and High-Loading-Bandwidth Digital In-ReRAM ComputationabstractRetrieval-Augmented Generation (RAG) enhances large language models (LLMs) by integrating external knowledge retrieval but faces challenges on edge devices due to high storage, energy, and latency demands. Computing-in-Memory (CIM) offers a promising solution by storing document embeddings in CIM macros and enabling in-situ parallel retrievals but is constrained by either low memory density or limited computational accuracy. To address these challenges, we present DIRC-RAG, a novel edge RAG acceleration architecture leveraging Digital In-ReRAM Computation (DIRC). DIRC integrates a high-density multi-level ReRAM subarray with an SRAM cell, utilizing SRAM and differential sensing for robust ReRAM readout and digital multiply-accumulate (MAC) operations. By storing all document embeddings within the CIM macro, DIRC achieves ultra-low-power, single-cycle data loading, substantially reducing both energy consumption and latency compared to off-chip DRAM. A query-stationary (QS) dataflow is supported for RAG tasks, minimizing on-chip data movement and reducing SRAM buffer requirements. We introduce error optimization for the DIRC ReRAM-SRAM cell by extracting the bit-wise spatial error distribution of the ReRAM subarray and applying targeted bit-wise data remapping. An error detection circuit is also implemented to enhance readout resilience against device-and circuit-level variations.Simulation results demonstrate that DIRC-RAG under TSMC 40nm process achieves an on-chip non-volatile memory density of 5.18Mb/mm2and a throughput of 131 TOPS. It delivers a 4MB retrieval latency of 5.6μs/query and an energy consumption of 0.956μJ/query, while maintaining the retrieval precision. Kunming Shao, Zhipeng Liao, Jiangnan Yu, Xijie Huang, Jingyu He, Fengshi Tian, Yi Zou 0001, Kwang-Ting Cheng, Chi-Ying Tsui |
ISLPED | 8 |
| 2025 | BoostViT: Booth-Serial Skipping and Tunable Scaling for Vision TransformersabstractVision Transformers (ViTs) have emerged as a dominant architecture in computer vision (CV), surpassing conventional neural network counterparts across diverse visual tasks. Despite their exceptional performance, ViTs incur substantial computational overhead characterized by high memory footprint, long inference latency, and elevated energy consumption. Current acceleration strategies for ViTs primarily focus on pruning operations or leveraging the inherent sparsity, requiring complex address control logic or position encoding. Alternatively, some software-based approaches attempt to pre-compute and separate dense and sparse matrix position encoding, while hardware solutions typically spend additional time and resources to obtain position encoding, decomposing matrix multiplications into structured forms. Through an analysis of ViTs’ parameters, we found that approximately 91.03% of the most significant bits (MSBs) are either 111s or 000s, and nearly 45% of 3 adjacent bits are identical. To leverage this characteristic of ViTs, we propose the Booth-Serial Skipping algorithm, which transforms the computation of consecutive 111 or 000 sequences into skip steps that require no additional computation time. Furthermore, the 4th to 6th bits of ViT weights can undergo aggressive scaling, enhancing the likelihood of Booth-skip operations with minimal impact on accuracy. The key innovation of this paper lies in exploiting the high proportion of naturally consecutive 0s or 1s in 8-bit weights during ViT inference and further expanding the skippable range through the Tunable Scaling strategy. At the hardware level, we develop a specialized accelerator to coordinate the proposed acceleration strategies. The processing element array in the accelerator is optimized for general matrix multiplication, it not only significantly improves the computation of multi-head self-attention but also enables resource reuse for linear transformations, ultimately optimizing end-to-end inference. Our design achieves$50.3\times $,$21.9\times $,$17.37\times $,$7.47\times $, and$1.49\times $an average end-to-end speedup on DeiT over CPU (Intel Xeon Gold 6152), EdgeGPU (NVIDIA Jetson Xavier NX), GPU (TITAN Xp), ViTCoD, and ViT-slice, respectively. Shiqi Zhao 0001, Chaoming Fang, Fengshi Tian, Jinbo Chen 0002, Changzeng Fu, Jie Yang 0033, Mohamad Sawan |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2024 | PhotonNTT: Energy-Efficient Parallel Photonic Number Theoretic Transform AcceleratorabstractFully homomorphic encryption (FHE) presents a promising opportunity to remove privacy barriers in various scenarios including cloud computing and secure database search, by enabling computation on encrypted data. However, integrating FHE with real-world applications remains challenging due to its significant computational overhead. In the FHE scheme, Number Theoretic Transform (NTT) consumes the primary computing resources and has great potential for acceleration. For the first time, we present a photonic NTT accelerator, PhotonNTT, with high energy efficiency and parallelism to address the above challenge. Our approach involves formulating the NTT into matrix-vector multiplication (MVM) operations and mapping the data flow into parallel photonic MVM units. A dedicated data mapping scheme is proposed to introduce free spectral range (FSR) and distributed RAM design into the system, which enables a high bit-wise parallelism level. The system's reliability is validated through the Monte-Carlo BER analysis. The experimen-tal evaluation shows that the proposed architecture outperforms SOTA CiM-based NTT accelerators with an improvement of 50x in throughput and 63x improvement in energy efficiency. Yinyi Liu, Chengeng Li, Shixi Chen, Fengshi Tian, Wei Zhang 0012, Jiang Xu 0001 |
DATE | 9 |
| 2024 | ReSCIM: Variation-Resilient High Weight-Loading Bandwidth In-Memory Computation Based on Fine-Grained Hybrid Integration of Multi-Level ReRAM and SRAM CellsabstractSRAM-CIM is a promising approach to implement efficient accelerator architecture as it enables accurate, energy-efficient AI computing, supporting both analog and digital computation. However, it has low area efficiency. On the other hand, Resistive RAM (ReRAM) provides dense on-chip storage, especially with multi-level cells (MLC), but ReRAM-CIM may introduce inaccuracies due to device variation and only supports analog computation. To leverage the strengths of both technologies, a hybrid architecture that combines them at a fine granularity is desirable. Previous hybrid designs incorporate ReRAM resistors into SRAM to improve storage density. However, they face scalability limitations and restricted signal margins for multi-level RRAM readout, leading to degraded computation accuracy. In this work, we propose ReSCIM, a hybrid compute-in-memory (CIM) architecture that seamlessly integrates multi-level ReRAM into SRAM cells at a fine-grained level. By incorporating a compact ReRAM crossbar in each SRAM cell, a dense CIM marco using SRAM-based computation is achieved. We develop an energy-efficient differential sensing scheme that enables parallel weight loading from local ReRAM crossbars to SRAM cells. This scheme allows multi-bit ReRAM data readout using a single SRAM cell and offers resilience to device variations. Furthermore, We designed a ReSCIM accelerator architecture for efficient AI acceleration, fully utilizing the highly scalable storage and exceptional weight-loading bandwidth. We employ a folded weight-mapping approach for MLC ReRAM cells to guarantee accurate classification even under substantial ReRAM device variations. Experimental results show that ReSCIM accelerators based on both analog and digital-based CIM achieve 60% energy savings and 98% latency savings, and 59× higher area efficiency compared to state-of-the-art all-weights-on-chip AI accelerators on AlexNet. Jingyu He, Kunming Shao, Jiakun Zheng, Fengshi Tian, Kwang-Ting Cheng, Chi-Ying Tsui |
ICCAD | 5 |
| 2024 | A Low-Power Level-Crossing Analog-to-Spike Converter Intended for Neuromorphic Biomedical ApplicationsabstractThe increasing interests in building bio-signal recording and processing systems for personal healthcare applications have been hindered by the critical sampling energy consumption issues of conventional biomedical systems. To address these limits, we propose a comprehensive strategy centered around a low-power level-crossing analog-to-spike converter (LC-ASC). This strategy enables event-driven compressive sampling by leveraging signal sparsity, achieving lower average sampling rates than Nyquist sampling. Our strategy includes universal VerilogA LC-ASC models, evaluation tools, and a reconfigurable data interface for versatile digital processing. Specifically, we introduce an online open-source VerilogA LC-ASC model and compression performance calculation tools for evaluating its performance with different bio-signals. The implemented LC-ASC chip demonstrates very-low power consumption of 31.5125.3 nW validated through chip measurements. Additionally, the proposed reconfigurable data interface ensures seamless integration with synchronous and asynchronous digital processing modules without sacrificing system-level performance. These advancements pave the way for energy-efficient neuromorphic biomedical circuits and systems. Jinbo Chen 0002, Hui Wu 0010, Fengshi Tian, Qiming Hou, Jie Yang 0033, Mohamad Sawan |
ISCAS | 3 |
| 2024 | Accelerating BPTT-Based SNN Training with Sparsity-Aware and Pipelined ArchitectureabstractOn-chip learning of Spiking Neural Networks (SNN) has been extensively researched to enhance adaptability and privacy protection, with Back-Propagation-Through-Time (BPTT) emerging as the top-performing method despite its resourceintensive nature. In this paper, we propose a dedicated training processor that accelerates the BPTT algorithm for SNNs. We analyze the bottlenecks and optimization opportunities in SNN- BPTT and introduce novel techniques such as recalculation of membrane potentials to reduce redundant data movement. Additionally, we implement a pipeline architecture with heterogeneous computing cores to maximize hardware utilization and parallelism. Exploiting three types of sparsity in BPTT allows us to skip unnecessary computations and memory access, further optimizing performance. The proposed processor, implemented using 40nm CMOS technology, achieves simulation results with an advanced training energy efficiency of 0.86pJ/OP. Chaoming Fang, Fengshi Tian, Jie Yang 0033, Mohamad Sawan |
ISCAS | 2 |
| 2024 | BOLS: A Bionic Sensor-direct On-chip Learning System with Direct-Feedback-Through-Time for Personalized Wearable Health MonitoringabstractPrecise bio-signal classification techniques for edge healthcare have been extensively researched, yet the scalability and efficiency of existing studies remain constrained by challenges in sensing, learning, and processing. Additionally, a deficiency in cross-level integration for the development of comprehensive healthcare systems has been observed. To tackle these issues and facilitate ultra-efficient personalized edge healthcare, this paper introduces the pioneering bionic sensor-direct on-chip learning and inference system with direct-feedback-through-time for user-specific cardiac arrhythmia detection, termed BOLS. This innovative system encompasses a compact sensor-direct feature extractor and a pipelined bionic processor, enabling end-to-end on-chip learning and inference. Employing cross-level co-design, our proposed bionic on-chip learning approach attains exceptional classification performance, boasting an accuracy of 98.6%, which ranks among the highest. The entire system has been implemented using 40nm CMOS process and subsequently verified. Remarkably, the proposed BOLS system consumes a mere 1.18mW for inference and 2.57mW for learning, resulting in an impressive power saving of over ×2000 compared to existing commercial training platforms. Fengshi Tian, Jiakun Zheng, Jingyu He, Jinbo Chen 0002, Chaoming Fang, Jie Yang 0033, Mohamad Sawan, Chi-Ying Tsui, Kwang-Ting Cheng |
ISCAS | 1 |
| 2023 | Late Breaking Results: Weight Decay is ALL You Need for Neural Network SparsificationabstractThe heuristic iterative pruning strategy has been widely used for neural network sparsification. However, it is challenging to identify the right connections to remove at each pruning iteration with only a one-shot evaluation of weight magnitude, especially at the early pruning stage. The erroneously removed connections, unfortunately, can hardly be recovered. In this work, we propose a weight decay strategy as a substitute for pruning, which let the "insignificant" weights moderately decay instead of being directly clamped to zero. At the end of the training, the vast majority of redundant weights will naturally become close to zero, making it easier to identify which connections could be removed safely. Experimental results show that the proposed weight decay method can achieve an ultra-high sparsity of 99%. Compared to the current pruning strategy, the model size is further reduced by 34%, improving the compression rate from 69× to 106× at the same accuracy. Xizi Chen, Fengshi Tian, Chi-Ying Tsui |
DAC | 4 |
| 2023 | AutoDCIM: An Automated Digital CIM CompilerabstractDigital Computing-in-Memory (DCIM) is an emerging architecture that integrates digital logic into memory for efficient AI computing. However, current DCIM designs heavily rely on manual efforts. This increases DCIM design time and limits the optimization space, making it challenging to satisfy the user specifications of diverse AI applications. This paper presents AutoDCIM, the first automated DCIM compiler. Au-toDCIM takes the user specifications as inputs and generates a DCIM macro architecture with an optimized layout. AutoDCIM’s template-based generation balances handcrafted cell design and agile macro development. AutoDCIM’s layout exploration loop analyzes diverse DCIM array partitioning schemes to satisfy user specifications. The auto-generated DCIM macros present competitive efficiency results in comparison with state-of-the-art silicon-verified DCIM macros. Jia Chen 0032, Fengbin Tu, Kunming Shao, Fengshi Tian, Xiao Huo, Chi-Ying Tsui, Kwang-Ting Cheng |
DAC | 4 |
| 2023 | NBSSN: A Neuromorphic Binary Single-Spike Neural Network for Efficient Edge IntelligenceabstractNeuromorphic computing approaches such as Spiking Neural Networks (SNN) have been increasingly adopted in bio-signal processing and interpretation due to its intrinsic neurodynamic attribute. Nevertheless, reconciling performance and power efficiency in SNN implementation is still a bottleneck. Single-spike neural coding scheme, which is an extremely sparse coding scheme, provides a solution to bridge the gap. In this work, a neuromorphic architecture, using binary single spike neural signals, is proposed with both algorithm and hardware implementation. A sparsity-aware spatial-temporal back-propagation training method is proposed together with a single-spike coding scheme. Also, a novel neuromorphic accelerator is co-designed with algorithmic optimization and implemented in 40nm CMOS process. Experimental results show that the proposed processor reaches an accuracy of 94.61% on the MNIST dataset, 93.59% on the N-MNIST dataset, and 93.27% on the ECG dataset, respectively, while consumes$0.173\mu\mathrm{J}$per ECG classification task and 0.16mm2on-chip area. The overall power consumption is reduced by 91.68% compared to the state-of-the-art systems. Ziyang Shen, Fengshi Tian, Chaoming Fang, Xiaoyong Xue, Jie Yang 0033, Mohamad Sawan |
ISCAS | 2 |
| 2022 | An Event-Driven Compressive Neuromorphic System for Cardiac Arrhythmia DetectionabstractWearable electrocardiograph (ECG) recording and processing systems have been developed to detect cardiac arrhythmia to help prevent heart attacks. Conventional wearable systems, however, suffer from high energy consumption at both circuit and system levels. To overcome the design challenges, this paper proposes an event-driven compressive ECG recording and neuromorphic processing system for cardiac arrhythmia detection. The proposed system achieves low power consumption and high arrhythmia detection accuracy via system level co-design with spike-based information representation. Event-driven level-crossing ADC (LC-ADC) is exploited in the recording system, which utilizes the sparsity of ECG signal to enable compressive recording and save ADC energy during the silent signal period. Meanwhile, the proposed spiking convolutional neural network (SCNN) based neuromorphic arrhythmia detection method is inherently compatible with the spike-based output of LC-ADC, hence realizing accurate detection and low energy consumption at system level. Simulation results show that the proposed system with 5-bit LC-ADC achieves 88.6% reduction of sampled data points compared with Nyquist sampling in the MIT-BIH dataset, and 93.59% arrhythmia detection accuracy with SCNN, demonstrating the compression ability of LC-ADC and the effectiveness of system level co-design with SCNN. Jinbo Chen 0002, Fengshi Tian, Jie Yang 0033, Mohamad Sawan |
ISCAS | 2 |
| 2022 | A Compact Online-Learning Spiking Neuromorphic Biosignal ProcessorabstractReal-time biosignal processing on wearable devices has attracted worldwide attention for its potential in healthcare applications. However, the requirement of low-area, low-power and high adaptability to different patients challenge conventional algorithms and hardware platforms. In this design, a compact online learning neuromorphic hardware architecture with ultralow power consumption designed explicitly for biosignal processing is proposed. A trace-based Spiking-Timing-Dependent-Plasticity (STDP) algorithm is applied to realize hardware-friendly online learning of a single-layer excitatory-inhibitory spiking neural network. Several techniques, including event-driven architecture and a fully optimized iterative computation approach, are adopted to minimize the hardware utilization and power consumption for the hardware implementation of online learning. Experiment results show that the proposed design reaches the accuracy of 87.36% and 83% for the Mixed National Institute of Standards and Technology database (MNIST) and ECG classification. The hardware architecture is implemented on a Zynq-7020 FPGA. Implementation results show that the Look-Up Table (LUT) and Flip Flops (FF) utilization reduced by 14.87 and 7.34 times, respectively, and the power consumption reduced by 21.69% compared to state of the art. Chaoming Fang, Ziyang Shen, Fengshi Tian, Jie Yang 0033, Mohamad Sawan |
ISCAS | 3 |
| 2022 | NIMBLE: A Neuromorphic Learning Scheme and Memristor Based Computing-in-Memory Engine for EMG Based Hand Gesture RecognitionabstractEMG based hand gesture recognition on convolutional neural networks (CNNs) has been widely learned, which gains high accuracy. However, CNN based systems are computationally complex and power consuming, thus hard to be deployed at edge. Biologically inspired, a new neuromorphic learning and computing approach for electromyogram (EMG) based hand gesture recognition tasks is proposed in this work. This approach designs an activate and inhibit joint processing spiking neural network (AIPS-SNN) which reaches an accuracy of 85.6% on Nina Pro dataset. Furthermore, the AIPS-SNN is deployed on the proposed memristor based computation in-memory (CIM) system, the power efficiency and area efficiency of which reach 10.146 TOPS/W and 35.399 GOPS/mm2, respectively. The experimental results indicate that the proposed neuromorphic CIM engine is promising for edge deployment. Fengshi Tian, Jinhao Liang, Jiahe Shi, Chaoming Fang, Hui Wu 0010, Xiaoyong Xue, Xiaoyang Zeng |
ISCAS | 1 |
| 2021 | A New Neuromorphic Computing Approach for Epileptic Seizure PredictionabstractSeveral high specificity and sensitivity seizure prediction methods with convolutional neural networks (CNNs) are reported. However, CNNs are computationally expensive and power hungry. These inconveniences make CNN-based methods hard to be implemented on wearable devices. Motivated by the energy-efficient spiking neural networks (SNNs), a neuromorphic computing approach for seizure prediction is proposed in this work. This approach uses a designed gaussian random discrete encoder to generate spike sequences from the EEG samples and make predictions in a spiking convolutional neural network (Spiking-CNN) which combines the advantages of CNNs and SNNs. The experimental results show that the sensitivity, specificity and AUC can remain 95.1%, 99.2% and 0.912 respectively while the computation complexity is reduced by 98.58% compared to CNN, indicating that the proposed Spiking-CNN is hardware friendly and of high precision. Fengshi Tian, Jie Yang 0033, Shiqi Zhao 0001, Mohamad Sawan |
ISCAS | 1 |