Jie Yang 0033

dblp:12/1198-33 · DBLP profile ↗
← Back
37ranked-venue papers
4as first author
28since 2021 · last 2026
0000-0002-4148-0042ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 19 · 1 first-author · 18 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 TF-MV SwAV++: Subject-Invariant Time-Frequency Prototype Learning for Cross-Subject sEEG Seizure Detection
Yunsheng Liao, Jie Yang 0033, Mohamad Sawan
ISCAS2
2026 A Heterogeneous Decision Spiking Transformer Accelerator with Locality-dependent KV Product Cache and Compute Pattern Reconfigurable Engine
Ziyang Shen, Zhipeng Liao, Sitan Shen, Chaoming Fang, Fengshi Tian, Jie Yang 0033, Mohamad Sawan
ISCAS6
2026 BioSeek: A Design Generation Framework of Biosignal Processors with Large-Language Models for Edge Healthcare Applications
abstract
Deep neural network (DNN)-based methodologies have shown impressive performance and robustness in the detection of abnormalities and decoding of multi-modal biosignals. While the use of DNNs provides promising classification and decoding capabilities, it also introduces significant design and cost challenges for the implementation of biomedical System on Chips (SoC). To address the increasing demand for advanced and efficient DNN-based healthcare solutions at the edge, we propose BioSeek, an agile design generation framework enhanced by cutting-edge large-language models (LLM). BioSeek offers a comprehensive solution to the design challenges associated with biosignal processors. The effectiveness of BioSeek is evaluated through the design generation of both application-specific and versatile biosignal processors, demonstrating performance that is competitive with existing solutions.
Fengshi Tian, Jiakun Zheng, Hui Wu 0010, Zilu Liu, Jinbo Chen 0002, Shiqi Zhao 0001, Jie Yang 0033, Mohamad Sawan, Chi-Ying Tsui, Kwang-Ting Cheng
ISCAS7
2026 A plug-and-play hybrid pruning framework for spike-driven transformers via adaptive spiking statistical importance scoring
Hanfei Liu, Shiqi Zhao 0001, Changzeng Fu, Jie Yang 0033, Mohamad Sawan
Neurocomputing4
2026 HAQ-ViT: A hardware-aware post-training quantization for efficient vision transformer inference
Shiqi Zhao 0001, Haozu Sun, Changzeng Fu, Jie Yang 0033, Mohamad Sawan
Knowl. Based Syst.7
2026 Memory-Efficient Intrinsic Gating Adaptation for Enhanced On-Device Epilepsy Diagnosis
abstract
Recently, advances in neuroscience and the rise of artificial intelligence have significantly enhanced the capabilities of epilepsy diagnosis. While EEG-based diagnosis offer a promising avenue for detecting and predicting seizure activity, practical implementation in real-world scenarios remains hindered by the heterogeneity of epilepsy and the variability of patient-specific biomarkers over time. Conventional deep learning models, trained on historical EEG, often fail to adapt to such biomarker variations, leading to degraded performance. Moreover, the computational and memory constraints of edge devices further exacerbate the challenge of on-device learning. To address these challenges, we introduce a novel framework, Memory-Efficient Intrinsic Gating Adaptation (MEIGA), designed to enhance real-world epilepsy diagnosis on resource-constrained edge devices. Our approach pre-trains a model using historical EEG data and employs lightweight adapter networks for efficient on-device tuning across new sessions, addressing session-to-session variability. By leveraging Direct Feedback Alignment (DFA), MEIGA reduces memory usage and computational overhead while maintaining high classification accuracy. Extensive experiments on the CHB-MIT epilepsy dataset demonstrate that MEIGA outperforms the pretrained-only Vision Transformer baseline, raising seizure prediction accuracy from 47.88% to 86.77% with only 3,908 tunable parameters (5.05% of the backbone). For seizure detection, MEIGA improves accuracy from 85.06% to 96.29% by adapting 2,008 parameters (17.40% of the base architecture). Further experiments on the AES dataset demonstrate that MEIGA consistently delivers strong performance across subjects and scales effectively to larger networks.
Shanjin Li, Di Wu 0057, Shiqi Zhao 0001, Jie Yang 0033, Mohamad Sawan
IEEE J. Biomed. Health Informatics4
2026 Neuro-BERT: Rethinking Masked Autoencoding for Self-Supervised Neurological Pretraining
abstract
Deep learning associated with neurological signals is poised to drive major advancements in diverse fields such as medical diagnostics, neurorehabilitation, and brain-computer interfaces. The challenge in harnessing the full potential of these signals lies in the dependency on extensive, high-quality annotated data, which is often scarce and expensive to acquire, requiring specialized infrastructure and domain expertise. To address the appetite for data in deep learning, we present Neuro-BERT, a self-supervised pre-training framework of neurological signals based on masked autoencoding in the Fourier domain. The intuition behind our approach is simple: frequency and phase distribution of neurological signals can reveal intricate neurological activities. We propose a novel pre-training task dubbed Fourier Inversion Prediction (FIP), which randomly masks out a portion of the input signal and then predicts the missing information using the Fourier inversion theorem. Pre-trained models can be potentially used for various downstream tasks such as sleep stage classification and gesture recognition. Unlike contrastive-based methods, which strongly rely on carefully hand-crafted augmentations and siamese structure, our approach works reasonably well with a simple transformer encoder with no augmentation requirements. By evaluating our method on several benchmark datasets, we show that Neuro-BERT improves downstream neurological-related tasks by a large margin.
Di Wu 0057, Siyuan Li 0002, Jie Yang 0033, Mohamad Sawan
IEEE J. Biomed. Health Informatics3
2025 SynDCIM: A Performance-Aware Digital Computing-in-Memory Compiler with Multi-Spec-Oriented Subcircuit Synthesis
abstract
Digital Computing-in-Memory (DCIM) is an innovative technology that integrates multiply-accumulation (MAC) logic directly into memory arrays to enhance the performance of modern AI computing. However, the need for customized memory cells and logic components currently necessitates significant manual effort in DCIM design. Existing tools for facilitating DCIM macro designs struggle to optimize subcircuit synthesis to meet user-defined performance criteria, thereby limiting the potential system-level acceleration that DCIM can offer. To address these challenges and enable the agile design of DCIM macros with optimal architectures, we present SynDCIM - a performance-aware DCIM compiler that employs multi-spec-oriented subcircuit synthesis. SynDCIM features an automated performance-to-layout generation process that aligns with user-defined performance expectations. This is supported by a scalable subcircuit library and a multi-spec-oriented searching algorithm for effective subcircuit synthesis. The effectiveness of SynDCIM is demonstrated through extensive experiments and validated with a test chip fabricated in a 40nm CMOS process. Testing results reveal that designs generated by SynDCIM exhibit competitive performance when compared to state-of-the-art manually designed DCIM macros.
Kunming Shao, Fengshi Tian, Jiakun Zheng, Jia Chen 0032, Jingyu He, Hui Wu 0010, Jinbo Chen 0002, Xihao Guan, Fengbin Tu, Jie Yang 0033, Mohamad Sawan, Kwang-Ting Cheng, Chi-Ying Tsui
DATE12
2025 Towards Homogeneous Lexical Tone Decoding from Heterogeneous Intracranial Recordings
abstract
Recent advancements in brain-computer interfaces (BCIs) and deep learning have made decoding lexical tones from intracranial recordings possible, providing the potential to restore the communication ability of speech-impaired tonal language speakers. However, data heterogeneity induced by both physiological and instrumental factors poses a significant challenge for unified invasive brain tone decoding. Particularly, the existing heterogeneous decoding paradigm (training subject-specific models with individual data) suffers from the intrinsic limitation that fails to learn generalized neural representations and leverages data across subjects. To this end, we introduce Homogeneity-Heterogeneity Disentangled Learning for Neural Representations (H2DiLR), a framework that disentangles and learns the homogeneity and heterogeneity from intracranial recordings of multiple subjects. To verify the effectiveness of H2DiLR, we collected stereoelectroencephalography (sEEG) from multiple participants reading Mandarin materials containing 407 syllables (covering nearly all Mandarin characters). Extensive experiments demonstrate that H2DiLR, as a unified decoding paradigm, outperforms the naive heterogeneous decoding paradigm by a large margin. We also empirically show that H2DiLR indeed captures homogeneity and heterogeneity during neural representation learning.
Di Wu 0057, Siyuan Li 0002, Jie Yang 0033, Mohamad Sawan
ICLR6
2025 Neuromorphic Computing Chips: Challenges and Trends
abstract
Over the past decade, artificial intelligence (AI) has made unprecedented advancements in various fields. However, with the emergence of large-scale models in recent years, the energy consumption associated with AI computing has become a critical issue that urgently needs to be addressed. Furthermore, as Moore's Law reaches its limits, increasing computational power has become extremely challenging. The scarcity of energy and computational resources presents significant challenges to the further development and application of current AI technologies. These challenges provide an unprecedented opportunity for the introduction of neuromorphic computing chips which show great potential in addressing the problems currently faced by AI. We present in this article the evolution of brain-inspired neuromorphic computing chips, tracing their evolution from the early artificial retina to advanced designs incorporating millions of artificial neurons. Also, we explore future opportunities in key applications such as brain-computer interfaces, which can facilitate more efficient neural communication; embodied intelligence, which seeks to replicate human cognitive functions; and large-scale models, where these chips' energy-efficient processing capabilities can greatly enhance AI system performance and scalability.
Jie Yang 0033, Mohamad Sawan
ISCAS1
2025 An Area-Efficient and Bit-Width Configurable Carry-Save Adder Tree for Spiking Transformers
abstract
Spiking transformers have been successfully applied to multiple applications with comparable accuracy with native transformers. Achieving a high energy efficiency with spiking transformers requires dedicated hardware design, especially a specialized matrix multiplication engine optimized for spike input. In this paper, we propose a carry-save adder (CSA) array with an improved energy and area efficiency for spiking transformer matrix multiplication computation. A two-staged CSA structure is proposed to support maximum logic reuse between 1b self-attention mode and 8b linear mode. Besides, a cubic-mesh architecture is proposed to organize CSA trees to reuse weight in different timesteps. Compared to a baseline accumulation design, the proposed architecture achieved a 2.7x area reduction and a 3.75x power reduction with the same throughput, showing that the optimized computation array has great potential to be applied in digital neuromorphic accelerators.
Chaoming Fang, Ziyang Shen, Fengshi Tian, Jie Yang 0033, Mohamad Sawan
ISCAS4
2025 A 2.53 fJ/Conversion Low-Power Hybrid ADC with Level-Crossing Assisted Sparisty Adaptivity for Implantable Neural Interface
abstract
In the realm of implantable brain neural interfaces, a prominent challenge is the constraint of limited power, particularly in systems with a high channel count. Various generic and application-specific strategies have been proposed to enhance energy efficiency while preserving signal integrity, resulting in varying levels of effectiveness. We introduce in paper an innovative approach that employs an auxiliary bypass featuring a low-power, low-sampling-rate level-crossing analog-to-digital converter (LC-ADC) to exploit signal sparsity. This method facilitates adaptive power control of the core Successive Approximation Register analog-to-digital converter (SAR ADC), achieving a significant 65% reduction in power consumption compared to traditional high-precision SAR ADCs. The ADC is fabricated using a 40 nm technology node, providing a bandwidth range from 10 Hz to 1 MHz and attaining an effective number of bits (ENOB) reaching up to 10.56 bits, with a figure-of-merit (FoM) as low as 2.56 fJ/conversion. This advancement underscores the potential for enhanced energy efficiency in high-channel-count neural interfaces.
Yutao Mao, Jinbo Chen 0002, Hui Wu 0010, Jie Yang 0033, Xiaofei Kuang, Mohamad Sawan
ISCAS4
2025 Efficient Self-Adaptive Pseudo-Resistor with Rapid Settling and High Linearity for Neurorecording Front-End Circuits
abstract
In this paper, we present a novel self-adaptive pseudo-resistor (A-PR) designed to enhance the performance of neurorecording front-end circuits in terms of settling time, linearity, and tunability. We validate the effectiveness of the proposed A-PR through the implementation of a capacitively- coupled instrumentation amplifier (CCIA) recording front-end using TSMC 40-nm process technology. The results demonstrate that the A-PR enables continuous recording with minimal interruptions, enhancing the system’s robustness and enabling more reliable acquisition of neural signals. Notably, the A-PR achieves a significant reduction in settling time, reaching the millisecond level—1000 times faster than conventional pseudo-resistors—while also exhibiting wide linear characteristics and easy tunability.
Hui Wu 0010, Xing Liu 0014, Jinbo Chen 0002, Wenjun Zou, Qiming Hou, Yutao Mao, Xiaofei Kuang, Jie Yang 0033, Mohamad Sawan
ISCAS10
2025 NeuroEye: A 54.59mW, 12200FPS Event-Driven Near-Sensor Eye-Tracking Processor with Pipelined Spatial-Temporal Spike-Streaming
abstract
This paper presents a design of an eye tracking system based on neuromorphic computing to enhance user interaction in augmented reality (AR) and virtual reality (VR) environments. Traditional methods face challenges of high computational demands and power consumption. To address these issues, we propose a fully-spike eye-tracking system that utilizes dynamic vision sensors (DVS) for asynchronous pixel-level change detection, thereby reducing data redundancy and improving temporal resolution. We proposed a pipelined processor specifically tailored for handling DVS events and Spiking Neural Network (SNN) computations. Our spatial-temporal spike-streaming architecture enables cascaded computation across all layers, achieving high energy efficiency and high frame rate in eye-tracking tasks. Implemented in a 40nm CMOS process, NeuroEye demonstrates up to 12200 frame-per-second (FPS) and 4.47uJ/frame energy efficiency with 54.59mW power consumption in post-layout evaluations.
Jiakun Zheng, Fengshi Tian, Jinbo Chen 0002, Chaoming Fang, Jie Yang 0033, Mohamad Sawan, Kwang-Ting Cheng, Chi-Ying Tsui
ISCAS6
2025 BoostViT: Booth-Serial Skipping and Tunable Scaling for Vision Transformers
abstract
Vision Transformers (ViTs) have emerged as a dominant architecture in computer vision (CV), surpassing conventional neural network counterparts across diverse visual tasks. Despite their exceptional performance, ViTs incur substantial computational overhead characterized by high memory footprint, long inference latency, and elevated energy consumption. Current acceleration strategies for ViTs primarily focus on pruning operations or leveraging the inherent sparsity, requiring complex address control logic or position encoding. Alternatively, some software-based approaches attempt to pre-compute and separate dense and sparse matrix position encoding, while hardware solutions typically spend additional time and resources to obtain position encoding, decomposing matrix multiplications into structured forms. Through an analysis of ViTs’ parameters, we found that approximately 91.03% of the most significant bits (MSBs) are either 111s or 000s, and nearly 45% of 3 adjacent bits are identical. To leverage this characteristic of ViTs, we propose the Booth-Serial Skipping algorithm, which transforms the computation of consecutive 111 or 000 sequences into skip steps that require no additional computation time. Furthermore, the 4th to 6th bits of ViT weights can undergo aggressive scaling, enhancing the likelihood of Booth-skip operations with minimal impact on accuracy. The key innovation of this paper lies in exploiting the high proportion of naturally consecutive 0s or 1s in 8-bit weights during ViT inference and further expanding the skippable range through the Tunable Scaling strategy. At the hardware level, we develop a specialized accelerator to coordinate the proposed acceleration strategies. The processing element array in the accelerator is optimized for general matrix multiplication, it not only significantly improves the computation of multi-head self-attention but also enables resource reuse for linear transformations, ultimately optimizing end-to-end inference. Our design achieves$50.3\times $,$21.9\times $,$17.37\times $,$7.47\times $, and$1.49\times $an average end-to-end speedup on DeiT over CPU (Intel Xeon Gold 6152), EdgeGPU (NVIDIA Jetson Xavier NX), GPU (TITAN Xp), ViTCoD, and ViT-slice, respectively.
Shiqi Zhao 0001, Chaoming Fang, Fengshi Tian, Jinbo Chen 0002, Changzeng Fu, Jie Yang 0033, Mohamad Sawan
IEEE Trans. Circuits Syst. I Regul. Pap.8
2024 VSViG: Real-Time Video-Based Seizure Detection via Skeleton-Based Spatiotemporal ViG
Yankun Xu, Yun-Hsuan Chen, Jie Yang 0033, Wenjie Ming, Mohamad Sawan
ECCV (82)4
2024 Exploring Effective Stimulus Encoding via Vision System Modeling for Visual Prostheses
abstract
Visual prostheses are potential devices to restore vision for blind people, which highly depends on the quality of stimulation patterns of the implanted electrode array. However, existing processing frameworks prioritize the generation of stimulation while disregarding the potential impact of restoration effects and fail to assess the quality of the generated stimulation properly. In this paper, we propose for the first time an end-to-end visual prosthesis framework (StimuSEE) that generates stimulation patterns with proper quality verification using V1 neuron spike patterns as supervision. StimuSEE consists of a retinal network to predict the stimulation pattern, a phosphene model, and a primary vision system network (PVS-net) to simulate the signal processing from the retina to the visual cortex and predict the firing rate of V1 neurons. Experimental results show that the predicted stimulation shares similar patterns to the original scenes, whose different stimulus amplitudes contribute to a similar firing rate with normal cells. Numerically, the predicted firing rate and the recorded response of normal neurons achieve a Pearson correlation coefficient of 0.78.
Chuanqing Wang, Di Wu 0057, Chaoming Fang, Jie Yang 0033, Mohamad Sawan
ICLR4
2024 A Low-Power Level-Crossing Analog-to-Spike Converter Intended for Neuromorphic Biomedical Applications
abstract
The increasing interests in building bio-signal recording and processing systems for personal healthcare applications have been hindered by the critical sampling energy consumption issues of conventional biomedical systems. To address these limits, we propose a comprehensive strategy centered around a low-power level-crossing analog-to-spike converter (LC-ASC). This strategy enables event-driven compressive sampling by leveraging signal sparsity, achieving lower average sampling rates than Nyquist sampling. Our strategy includes universal VerilogA LC-ASC models, evaluation tools, and a reconfigurable data interface for versatile digital processing. Specifically, we introduce an online open-source VerilogA LC-ASC model and compression performance calculation tools for evaluating its performance with different bio-signals. The implemented LC-ASC chip demonstrates very-low power consumption of 31.5125.3 nW validated through chip measurements. Additionally, the proposed reconfigurable data interface ensures seamless integration with synchronous and asynchronous digital processing modules without sacrificing system-level performance. These advancements pave the way for energy-efficient neuromorphic biomedical circuits and systems.
Jinbo Chen 0002, Hui Wu 0010, Fengshi Tian, Qiming Hou, Jie Yang 0033, Mohamad Sawan
ISCAS6
2024 Accelerating BPTT-Based SNN Training with Sparsity-Aware and Pipelined Architecture
abstract
On-chip learning of Spiking Neural Networks (SNN) has been extensively researched to enhance adaptability and privacy protection, with Back-Propagation-Through-Time (BPTT) emerging as the top-performing method despite its resourceintensive nature. In this paper, we propose a dedicated training processor that accelerates the BPTT algorithm for SNNs. We analyze the bottlenecks and optimization opportunities in SNN- BPTT and introduce novel techniques such as recalculation of membrane potentials to reduce redundant data movement. Additionally, we implement a pipeline architecture with heterogeneous computing cores to maximize hardware utilization and parallelism. Exploiting three types of sparsity in BPTT allows us to skip unnecessary computations and memory access, further optimizing performance. The proposed processor, implemented using 40nm CMOS technology, achieves simulation results with an advanced training energy efficiency of 0.86pJ/OP.
Chaoming Fang, Fengshi Tian, Jie Yang 0033, Mohamad Sawan
ISCAS3
2024 BOLS: A Bionic Sensor-direct On-chip Learning System with Direct-Feedback-Through-Time for Personalized Wearable Health Monitoring
abstract
Precise bio-signal classification techniques for edge healthcare have been extensively researched, yet the scalability and efficiency of existing studies remain constrained by challenges in sensing, learning, and processing. Additionally, a deficiency in cross-level integration for the development of comprehensive healthcare systems has been observed. To tackle these issues and facilitate ultra-efficient personalized edge healthcare, this paper introduces the pioneering bionic sensor-direct on-chip learning and inference system with direct-feedback-through-time for user-specific cardiac arrhythmia detection, termed BOLS. This innovative system encompasses a compact sensor-direct feature extractor and a pipelined bionic processor, enabling end-to-end on-chip learning and inference. Employing cross-level co-design, our proposed bionic on-chip learning approach attains exceptional classification performance, boasting an accuracy of 98.6%, which ranks among the highest. The entire system has been implemented using 40nm CMOS process and subsequently verified. Remarkably, the proposed BOLS system consumes a mere 1.18mW for inference and 2.57mW for learning, resulting in an impressive power saving of over ×2000 compared to existing commercial training platforms.
Fengshi Tian, Jiakun Zheng, Jingyu He, Jinbo Chen 0002, Chaoming Fang, Jie Yang 0033, Mohamad Sawan, Chi-Ying Tsui, Kwang-Ting Cheng
ISCAS7
2024 Shorter latency of real-time epileptic seizure detection via probabilistic prediction
Yankun Xu, Jie Yang 0033, Wenjie Ming, Mohamad Sawan
Expert Syst. Appl.2
2023 NBSSN: A Neuromorphic Binary Single-Spike Neural Network for Efficient Edge Intelligence
abstract
Neuromorphic computing approaches such as Spiking Neural Networks (SNN) have been increasingly adopted in bio-signal processing and interpretation due to its intrinsic neurodynamic attribute. Nevertheless, reconciling performance and power efficiency in SNN implementation is still a bottleneck. Single-spike neural coding scheme, which is an extremely sparse coding scheme, provides a solution to bridge the gap. In this work, a neuromorphic architecture, using binary single spike neural signals, is proposed with both algorithm and hardware implementation. A sparsity-aware spatial-temporal back-propagation training method is proposed together with a single-spike coding scheme. Also, a novel neuromorphic accelerator is co-designed with algorithmic optimization and implemented in 40nm CMOS process. Experimental results show that the proposed processor reaches an accuracy of 94.61% on the MNIST dataset, 93.59% on the N-MNIST dataset, and 93.27% on the ECG dataset, respectively, while consumes$0.173\mu\mathrm{J}$per ECG classification task and 0.16mm2on-chip area. The overall power consumption is reduced by 91.68% compared to the state-of-the-art systems.
Ziyang Shen, Fengshi Tian, Chaoming Fang, Xiaoyong Xue, Jie Yang 0033, Mohamad Sawan
ISCAS6
2023 SpikeSEE: An energy-efficient dynamic scenes processing framework for retinal prostheses
abstract
Intelligent and low-power retinal prostheses are highly demanded in this era, where wearable and implantable devices are used for numerous healthcare applications. In this paper, we propose an energy-efficient dynamic scenes processing framework (SpikeSEE) that combines a spike representation encoding technique and a bio-inspired spiking recurrent neural network (SRNN) model to achieve intelligent processing and extreme low-power computation for retinal prostheses. The spike representation encoding technique could interpret dynamic scenes with sparse spike trains, decreasing the data volume. The SRNN model, inspired by the human retina's special structure and spike processing method, is adopted to predict the response of ganglion cells to dynamic scenes. Experimental results show that the Pearson correlation coefficient of the proposed SRNN model achieves 0.93, which outperforms the state-of-the-art processing framework for retinal prostheses. Thanks to the spike representation and SRNN processing, the model can extract visual features in a multiplication-free fashion. The framework achieves 8 times power reduction compared with the convolutional recurrent neural network (CRNN) processing-based framework. Our proposed SpikeSEE predicts the response of ganglion cells more accurately with lower energy consumption, which alleviates the precision and power issues of retinal prostheses and provides a potential solution for wearable or implantable prostheses.
Chuanqing Wang, Chaoming Fang, Jie Yang 0033, Mohamad Sawan
Neural Networks4
2022 An Event-Driven Compressive Neuromorphic System for Cardiac Arrhythmia Detection
abstract
Wearable electrocardiograph (ECG) recording and processing systems have been developed to detect cardiac arrhythmia to help prevent heart attacks. Conventional wearable systems, however, suffer from high energy consumption at both circuit and system levels. To overcome the design challenges, this paper proposes an event-driven compressive ECG recording and neuromorphic processing system for cardiac arrhythmia detection. The proposed system achieves low power consumption and high arrhythmia detection accuracy via system level co-design with spike-based information representation. Event-driven level-crossing ADC (LC-ADC) is exploited in the recording system, which utilizes the sparsity of ECG signal to enable compressive recording and save ADC energy during the silent signal period. Meanwhile, the proposed spiking convolutional neural network (SCNN) based neuromorphic arrhythmia detection method is inherently compatible with the spike-based output of LC-ADC, hence realizing accurate detection and low energy consumption at system level. Simulation results show that the proposed system with 5-bit LC-ADC achieves 88.6% reduction of sampled data points compared with Nyquist sampling in the MIT-BIH dataset, and 93.59% arrhythmia detection accuracy with SCNN, demonstrating the compression ability of LC-ADC and the effectiveness of system level co-design with SCNN.
Jinbo Chen 0002, Fengshi Tian, Jie Yang 0033, Mohamad Sawan
ISCAS3
2022 A Compact Online-Learning Spiking Neuromorphic Biosignal Processor
abstract
Real-time biosignal processing on wearable devices has attracted worldwide attention for its potential in healthcare applications. However, the requirement of low-area, low-power and high adaptability to different patients challenge conventional algorithms and hardware platforms. In this design, a compact online learning neuromorphic hardware architecture with ultralow power consumption designed explicitly for biosignal processing is proposed. A trace-based Spiking-Timing-Dependent-Plasticity (STDP) algorithm is applied to realize hardware-friendly online learning of a single-layer excitatory-inhibitory spiking neural network. Several techniques, including event-driven architecture and a fully optimized iterative computation approach, are adopted to minimize the hardware utilization and power consumption for the hardware implementation of online learning. Experiment results show that the proposed design reaches the accuracy of 87.36% and 83% for the Mixed National Institute of Standards and Technology database (MNIST) and ECG classification. The hardware architecture is implemented on a Zynq-7020 FPGA. Implementation results show that the Look-Up Table (LUT) and Flip Flops (FF) utilization reduced by 14.87 and 7.34 times, respectively, and the power consumption reduced by 21.69% compared to state of the art.
Chaoming Fang, Ziyang Shen, Fengshi Tian, Jie Yang 0033, Mohamad Sawan
ISCAS4
2022 Towards Task-aware Signal Compression for Efficient Continuous Health Monitoring
abstract
High-precision multi-channel bio-signals are the basis of reliable and accurate wearable and implantable continuous health monitoring systems. However, the limitations of transmission bandwidth and computation resources of these systems pose heavy constraints on either the communication or direct processing of the large volume of physiological signals. Although signal compression can be adopted to compress the signals, most existing compression methods are computationally expensive and completely overlook the actual monitoring task purpose, which causes the discard of task-relevant information. Moreover, a complex reconstruction process is needed for further signal analysis at the cost of a heavy computational burden for downstream devices. We propose in this paper a novel flexible health monitoring framework where the signal is compressed with a low computation and hardware cost in-sensor compression matrix, trained in a task-aware fashion to preserve task-relevant information. The resulting compressed signals can be transmitted with significantly lower bandwidth, analyzed directly without a dedicated reconstruction process, or reconstructed with high fidelity. We demonstrate the effectiveness of our proposed framework by showcasing a seizure monitoring system. Prediction accuracy, sensitivity, false prediction rate, and signal reconstruction quality are reported under different compression ratios. Extensive experiments show that the proposed framework is accurate, with an average seizure prediction accuracy of 91.44%.
Di Wu 0057, Jie Yang 0033, Mohamad Sawan
ISCAS2
2022 NeuroSEE: A Neuromorphic Energy-Efficient Processing Framework for Visual Prostheses
abstract
Visual prostheses with both comprehensive visual signal processing capability and energy efficiency are becoming increasingly demanded in the age of intelligent personal healthcare, particularly with the rise of wearable and implantable devices. To address this trend, we propose NeuroSEE, a neuromorphic energy-efficient processing framework that combines a spike representation encoding technique and a bio-inspired processing method. This framework first utilizes sparse spike trains to represent visual information, and then a bio-inspired spiking neural network (SNN) is adopted to process the spike trains. The SNN model makes use of an IF neuron with multiple spike-firing rates to decrease the energy consumption without compensating for prediction performance. The experimental results indicate that when predicting the response of the primary visual cortex, the framework achieves a state-of-the-art Pearson correlation coefficient performance. Spike-based recording and processing methods simplify the storage and transmission of redundant scene information and complex calculation processes. It could reduce power consumption by 15 times compared with the existing Convolutional neural network (CNN) processing framework. The proposed NeuroSEE framework predicts the response of the primary visual cortex in an energy efficient manner, making it a powerful tool for visual prostheses.
Chuanqing Wang, Jie Yang 0033, Mohamad Sawan
IEEE J. Biomed. Health Informatics2
2021 A New Neuromorphic Computing Approach for Epileptic Seizure Prediction
abstract
Several high specificity and sensitivity seizure prediction methods with convolutional neural networks (CNNs) are reported. However, CNNs are computationally expensive and power hungry. These inconveniences make CNN-based methods hard to be implemented on wearable devices. Motivated by the energy-efficient spiking neural networks (SNNs), a neuromorphic computing approach for seizure prediction is proposed in this work. This approach uses a designed gaussian random discrete encoder to generate spike sequences from the EEG samples and make predictions in a spiking convolutional neural network (Spiking-CNN) which combines the advantages of CNNs and SNNs. The experimental results show that the sensitivity, specificity and AUC can remain 95.1%, 99.2% and 0.912 respectively while the computation complexity is reduced by 98.58% compared to CNN, indicating that the proposed Spiking-CNN is hardware friendly and of high precision.
Fengshi Tian, Jie Yang 0033, Shiqi Zhao 0001, Mohamad Sawan
ISCAS2
2020 Binary Single-Dimensional Convolutional Neural Network for Seizure Prediction
abstract
Nowadays, several deep learning methods are proposed to tackle the challenge of epileptic seizure prediction. However, these methods still cannot be implemented as part of implantable or efficient wearable devices due to their large hardware and corresponding high-power consumption. They usually require complex feature extraction process, large memory for storing high precision parameters and complex arithmetic computation, which greatly increases required hardware resources. Moreover, available yield poor prediction performance, because they adopt network architecture directly from image recognition applications fails to accurately consider the characteristics of EEG signals. We propose in this paper a hardware-friendly network called Binary Single-dimensional Convolutional Neural Network (BSDCNN) intended for epileptic seizure prediction. BSDCNN utilizes 1D convolutional kernels to improve prediction performance. All parameters are binarized to reduce the required computation and storage, except the first layer. Overall area under curve, sensitivity, and false prediction rate reaches 0.915, 89.26%, 0.117/h and 0.970, 94.69%, 0.095/h on American Epilepsy Society Seizure Prediction Challenge (AES) dataset and the CHB-MIT one respectively. The proposed architecture outperforms recent works while offering 7.2 and 25.5 times reductions on the size of parameter and computation, respectively.
Shiqi Zhao 0001, Jie Yang 0033, Yankun Xu, Mohamad Sawan
ISCAS2
2019 Deep Single Image Enhancer
abstract
Surveillance cameras can be deployed in various environments where lighting conditions are constantly changing. However, due to the limited dynamic range of current image sensors, the captured images are only low dynamic range images that usually suffer from over-exposure and under-exposure situations where important details are lost. Therefore, it is critical to recover the lost details of such images in order to improve visual experience for observers and performance for possible computer vision processing. In this paper, we propose a reformulated Laplacian pyramid and a convolutional neural network (CNN) model to enhance and recover the lost detail of a degraded image. The reformulated Laplacian first decomposes the image into two sub-images that contain global and local image features, respectively. The global features and local features are processed by the proposed CNN model to manipulate the global luminance terrain and enhance local details. The final image is obtained by reconstructing the CNN generated local and global features. Various experiments have been conducted. The results demonstrate that the proposed model outperforms the state-of-the-art methods.
Mengchen Lin, Jie Yang 0033, Orly Yadid-Pecht
AVSS2
2018 A Heterogeneous Parallel Processor for High-Speed Vision Chip
abstract
This paper proposes a heterogeneous parallel processor for high-speed vision chip. It contains four levels of processors with different parallelisms and complexities: processing element (PE) array processor, patch processing unit (PPU) array processor, self-organizing map (SOM) neural network processor, and dual-core microprocessor unit (MPU). The fine-grained PE array processor, middle-grained PPU array processor, and SOM neural network processor carry out image processing in pixel-parallel, patch-parallel, and distributed-parallel fashions, respectively. The MPU controls the overall system and executes some serial algorithms. The processor can improve the total system performance from low-level to high-level image processing significantly. A prototype is implemented with$64 \times 64$PE array,$8 \times 8$PPU array,$16 \times 24$SOM network, and a dual-core MPU. The proposed heterogeneous parallel processor introduces a new degree of parallelism, namely, patch parallel, which is for parallel local-feature extraction and feature detection. It can flexibly perform the state-of-the-art computer vision as well as various image processing algorithms at high speed. Various complicated applications, including feature extraction, face detection, and high-speed tracking, are demonstrated.
Jie Yang 0033, Yongxing Yang, Jian Liu 0021, Nanjian Wu
IEEE Trans. Circuits Syst. Video Technol.1
2017 Multi-Scale histogram tone mapping algorithm enables better object detection in wide dynamic range images
abstract
In this paper, we present a novel tone mapping algorithm based on multi-scale histograms and fusion (MS-Hist), for displaying wide dynamic range (WDR) images and better detection of objects such as human faces. The proposed algorithm tone maps pixels based on multiple scale local histograms, where small scales are used to preserve local contrast and large scales allow to maintain the global brightness consistency. A database of WDR images of humans depicted in high-contrast light conditions was created to validate and compare the performance of various algorithms for face detection in tasks such as biometric based identification. Our experimental results show that the proposed MS-Hist algorithm preserves image detail, brightness and high local contrast, and can benefit tasks such as face detection in WDR images.
Jie Yang 0033, Alain Horé, Ulian Shahnovich, Kenneth Lai, Svetlana N. Yanushkevich, Orly Yadid-Pecht
AVSS1
2017 High-speed visual target tracking with mixed rotation invariant description and skipping searching
Yongxing Yang, Jie Yang 0033, Nanjian Wu
Sci. China Inf. Sci.2
2017 High-Speed Target Tracking System Based on a Hierarchical Parallel Vision Processor and Gray-Level LBP Algorithm
abstract
Visual target tracking has made significant advances in past decades. However, fast and robust vision target tracking systems are still greatly demanded. This paper proposes a novel high-speed target tracking system based on hierarchical parallel vision processor architecture. This system contains three main parts: 1) a CMOS image sensor; 2) a vision processor; and 3) an actuator with two degrees of freedom. The vision processor integrates a pixel-parallel processing element (PE) array, a row-parallel row processor (RP) array, dual-core microprocessor unit (MPU) and motor controller. The PE array and RP array can speed up low-level and middle-level image processing operations by${O(M^{2})}$and${O}$(${M}$), respectively. The MPU is responsible for the high-level image processing and the overall chip management. A novel tracking algorithm based on a gray-level local binary pattern descriptor is proposed. The descriptor describes not only local texture feature but also distribution of luminance. The algorithm increases the robustness of the tracking system under low resolution scenery and complex background. It can be carried out by the vision processor with very high efficiency. Experiment results demonstrate that the system can track a fast moving target under complex conditions and the vision processor can achieve over 2000 frames/s processing speed of the target tracking algorithm with$ {750\times 480}$image resolution.
Yongxing Yang, Jie Yang 0033, Nanjian Wu
IEEE Trans. Syst. Man Cybern. Syst.2
2014 A massively parallel keypoint detection and description (MP-KDD) algorithm for high-speed vision chip
Cong Shi 0003, Jie Yang 0033, Nanjian Wu, Zhihua Wang 0001
Sci. China Inf. Sci.2
2014 A high speed multi-level-parallel array processor for vision chips
Cong Shi 0003, Jie Yang 0033, Nanjian Wu, Zhihua Wang 0001
Sci. China Inf. Sci.2
2004 On the performance of a novel multi-hop packet relaying ad hoc cellular system
abstract
A novel architecture for connectionless packet-based communication in wireless mobile cellular system is studied. A certain number of data relaying stations (DRS), which can relay packets from one cell to another in a manner similar to the ad-hoc multi-hop relaying mechanism, are placed in the cellular system to balance traffic load between neighboring cells. The placement of DRS is carefully studied to provide a full relaying coverage of the cell area with the minimum number of DRS. According to a distributed delay-sensitive relaying control strategy, mobile stations (MS) can locally decide whether to divert their traffic to neighboring cells without any central control. Simulation results demonstrate that the proposed method can significantly improve the system throughput and resource utilization and reduce the average packet delay.
Jie Yang 0033, Zhaoyang Zhang 0001, Zhen-zhou Tang
PIMRC1