VLDB 2026 Research / reviewers in the wild / expert
Chang Gao 0002
dblp:11/8376-2
· DBLP profile ↗
23ranked-venue papers
6as first author
18since 2021 · last 2026
0000-0002-3284-4078ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 5 first-author · 16 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | JaneEye: A 12-nm 2K-FPS 18.9-μJ/Frame Event-based Eye Tracking AcceleratorabstractEye tracking has become a key technology for gaze-based interactions in Extended Reality (XR). However, conventional frame-based eye-tracking systems often fall short of XR’s stringent requirements for high accuracy, low latency, and energy efficiency. Event cameras present a compelling alternative, offering ultra-high temporal resolution and low power consumption. In this paper, we present JaneEye, an energy-efficient event-based eye-tracking hardware accelerator designed specifically for wearable devices, leveraging sparse, high-temporal-resolution event data. We introduce an ultra-lightweight neural network architecture featuring a novel ConvJANET layer, which simplifies the traditional ConvLSTM by retaining only the forget gate, thereby halving computational complexity without sacrificing temporal modeling capability. Our proposed model achieves high accuracy with a pixel error of 2.45 on the 3ET+ dataset, using only 17.6 K parameters, with up to 1250 Hz event frame rate. To further enhance hardware efficiency, we employ custom linear approximations of activation functions (HardSigmoid and Hard-Tanh) and fixed-point quantization. Through software-hardware co-design, our 12-nm ASIC implementation operates at 400 MHz, delivering an end-to-end latency of 0.5 ms (equivalent to 2000 Frames Per Second (FPS)) at an energy efficiency of 18.9 μJ/frame. JaneEye sets a new benchmark in low-power, high-performance eye-tracking solutions suitable for integration into next-generation XR wearables. Qinyu Chen, Chang Gao 0002 |
ASP-DAC | 4 |
| 2026 | SlimTTS: Parameter- and Compute-Efficient Neural Text-to-Speech Synthesis for Real-Time Inference
Arne de Beer, Chang Gao 0002 |
ISCAS | 2 |
| 2025 | SparseDPD: A Sparse Neural Network-Based Digital Predistortion FPGA Accelerator for RF Power Amplifier LinearizationabstractDigital predistortion (DPD) is crucial for linearizing radio frequency (RF) power amplifiers (PAs), improving signal integrity and efficiency in wireless systems. Neural network (NN)-based DPD methods surpass traditional polynomial models but face computational challenges limiting their practical deployment. This paper introduces SparseDPD, an FPGA accelerator employing a spatially sparse phase-normalized time-delay neural network (PNTDNN), optimized through unstructured pruning to reduce computational load without accuracy loss. Implemented on a Xilinx Zynq-7Z010 FPGA, SparseDPD operates at 170 MHz, achieving exceptional linearization performance (ACPR: −59.4 dBc, EVM: −54.0 dBc, NMSE: −48.2 dB) with only 241 mW dynamic power, using 64 parameters with 74% sparsity. This work demonstrates FPGA-based acceleration, making NN-based DPD practical and efficient for real-time wireless communication applications. Code is publicly available at https://github.com/MannoVersluis/SparseDPD. Manno Versluis, Yizhuo Wu, Chang Gao 0002 |
FPL | 3 |
| 2025 | FACET: Fast and Accurate Event-Based Eye Tracking Using Ellipse Modeling for Extended RealityabstractEye tracking is a key technology for gaze-based interactions in Extended Reality (XR), but traditional frame-based systems struggle to meet XR's demands for high accuracy, low latency, and power efficiency. Event cameras offer a promising alternative due to their high temporal resolution and low power consumption. In this paper, we present FACET (Fast and Accurate Event-based Eye Tracking), an end-to-end neural network that directly outputs pupil ellipse parameters from event data, optimized for real-time XR applications. The ellipse output can be directly used in subsequent ellipse-based pupil trackers. We enhance the EV-Eye dataset by expanding annotated data and converting original mask labels to ellipse-based annotations to train the model. Besides, a novel trigonometric loss is adopted to address angle discontinuities and a fast causal event volume event representation method is put forward. On the enhanced EV-Eye test set, FACET achieves an average pupil center error of$\mathbf{0. 2 0}$pixels and an inference time of 0.53 ms, reducing pixel error and inference time by$1.6 \times$and$1.8 \times$compared to the prior art, EV-Eye, with$4.4 \times$and$11.7 \times$less parameters and arithmetic operations. The code is available at https://github.com/DeanJY/FACET. Junyuan Ding, Chang Gao 0002, Qinyu Chen |
ICRA | 3 |
| 2025 | FPGA Hardware Neural Control of CartPole and F1TENTH Race CarabstractLatency and computational cost often limit the use of Nonlinear Model Predictive Control (NMPC) in real-time robotics. To address this limitation, our work investigates FPGA-implemented Neural Controllers (NC) trained through supervised learning, mimicking NMPC. We show that inexpensive embedded FPGA hardware is sufficient to implement these neural controllers for high-frequency control of robotic systems. We demonstrate kilohertz control rates for a cartpole and offload control to the FPGA hardware on the F1TENTH race car. The FPGA NC outperforms NMPC on the cartpole, due to the faster control rate afforded by faster NC inference. The code and hardware implementation for this paper are available at https://github.com/SensorsINI/Neural-Control-Tools. Marcin Paluch, Florian Bolli, Antonio Rios-Navarro, Chang Gao 0002, Tobi Delbruck |
IROS | 5 |
| 2025 | CleanUMamba: A Compact Mamba Network for Speech Denoising using Channel PruningabstractThis paper presents CleanUMamba, a time-domain neural network architecture designed for real-time causal audio denoising directly applied to raw waveforms. CleanUMamba leverages a U-Net encoder-decoder structure, incorporating the Mamba state-space model in the bottleneck layer. By replacing conventional self-attention and LSTM mechanisms with Mamba, our architecture offers superior denoising performance while maintaining a constant memory footprint, enabling streaming operation. To enhance efficiency, we applied structured channel pruning, achieving an 8X reduction in model size without compromising audio quality. Our model demonstrates strong results in the Interspeech 2020 Deep Noise Suppression challenge. Specifically, CleanUMamba achieves a PESQ score of 2.42 and STOI of 95.1% with only 442K parameters and 468M MACs, matching or outperforming larger models in real-time performance. Code will be available at: https://github.com/lab-emi/CleanUMamba Sjoerd Groot, Qinyu Chen, Jan C. van Gemert, Chang Gao 0002 |
ISCAS | 4 |
| 2025 | DPD-NeuralEngine: A 22-nm 6.6-TOPS/W/mm2 Recurrent Neural Network Accelerator for Wideband Power Amplifier Digital Pre-DistortionabstractThe increasing adoption of Deep Neural Network (DNN)-based Digital Pre-distortion (DPD) in modern communication systems necessitates efficient hardware implementations. This paper presents DPD-NeuralEngine, an ultra-fast, tiny-area, and power-efficient DPD accelerator based on a Gated Recurrent Unit (GRU) neural network (NN). Leveraging a co-designed software and hardware approach, our 22 nm CMOS implementation operates at 2 GHz, capable of processing I/Q signals up to 250 MSps. Experimental results demonstrate a throughput of 256.5 GOPS and power efficiency of 1.32 TOPS/W with DPD linearization performance measured in Adjacent Channel Power Ratio (ACPR) of -45.3 dBc and Error Vector Magnitude (EVM) of -39.8 dB. To our knowledge, this work represents the first AI-based DPD application-specific integrated circuit (ASIC) accelerator, achieving a power-area efficiency (PAE) of 6.6 TOPS/W/mm2. Yizhuo Wu, Qinyu Chen, Leo C. N. de Vreede, Chang Gao 0002 |
ISCAS | 6 |
| 2025 | SlimSeiz: Efficient Channel-Adaptive Seizure Prediction Using a Mamba-Enhanced NetworkabstractEpileptic seizures cause abnormal brain activity, and their unpredictability can lead to accidents, underscoring the need for long-term seizure prediction. Although seizures can be predicted by analyzing electroencephalogram (EEG) signals, existing methods often require too many channels or larger models, limiting mobile usability. This paper introduces a SlimSeiz framework that utilizes adaptive channel selection with a lightweight neural network model. SlimSeiz operates in two states: the first stage selects the optimal channel set for seizure prediction using machine learning algorithms, and the second stage employs a lightweight neural network based on convolution and Mamba for prediction. On the Children’s Hospital Boston-MIT (CHB-MIT) EEG dataset, SlimSeiz can reduce channels from 22 to 8 while claiming a satisfactory result of 94.8% accuracy, 95.5% sensitivity, and 94.0% specificity with only 21.2 K model parameters, matching or outperforming larger models’ performance. We also validate SlimSeiz on a new EEG dataset, SRH-LEI, collected from Shanghai Renji Hospital, demonstrating its effectiveness across different patients. The code and SRH-LEI dataset are available at https://github.com/guoruilu/SlimSeiz. Guorui Lu, Bingyuan Huang, Chang Gao 0002, Todor Stefanov, Qinyu Chen |
ISCAS | 4 |
| 2025 | HengNet: An Ultra-lightweight Model with Two-level Reuse Algorithm for Seizure Detection and PredictionabstractTraditional models based on electroencephalographic (EEG) signals for seizure monitoring encounter difficulties in simultaneously optimizing accuracy, response latency, and computational load. These challenges hinder their deployment in edge computing environments, where real-time local inference is critical. To address these issues, we introduce a novel network architecture, designated as HengNet. This architecture integrates a Two-level Reuse Algorithm (TRA), which strategically reutilizes outputs from intermediate layers, considerably reducing the average computational load per inference—vital for scenarios requiring frequent inferences. When tested on the CHB-MIT dataset, this patient-specific model attains classification accuracies of 95.67% and 99.60% for seizure prediction and detection, respectively. Notably, it maintains an average computational load of merely 0.05 million multiply-accumulate operations (MACs) per inference and has a compact model size of 6.87 K parameters. These results represent a significant advancement compared with existing methods. Operating at a rate of 32 inferences per second, the computational load of the model for seizure prediction has been reduced by more than 19.4 times, and for seizure detection, by more than 6.4 times. Heng Zhang 0025, Linxiang Wang, Wenjie Fan 0004, Zhenglin Gu, Youbin Luo, Xingjie Zou, Chang Gao 0002, Qinyu Chen, Li Li 0003 |
ISCAS | 7 |
| 2024 | Exploiting Symmetric Temporally Sparse BPTT for Efficient RNN TrainingabstractRecurrent Neural Networks (RNNs) are useful in temporal sequence tasks. However, training RNNs involves dense matrix multiplications which require hardware that can support a large number of arithmetic operations and memory accesses. Implementing online training of RNNs on the edge calls for optimized algorithms for an efficient deployment on hardware. Inspired by the spiking neuron model, the Delta RNN exploits temporal sparsity during inference by skipping over the update of hidden states from those inactivated neurons whose change of activation across two timesteps is below a defined threshold. This work describes a training algorithm for Delta RNNs that exploits temporal sparsity in the backward propagation phase to reduce computational requirements for training on the edge. Due to the symmetric computation graphs of forward and backward propagation during training, the gradient computation of inactivated neurons can be skipped. Results show a reduction of ∼80% in matrix operations for training a 56k parameter Delta LSTM on the Fluent Speech Commands dataset with negligible accuracy loss. Logic simulations of a hardware accelerator designed for the training algorithm show 2-10X speedup in matrix computations for an activation sparsity range of 50%-90%. Additionally, we show that the proposed Delta RNN training will be useful for online incremental learning on edge devices with limited computing resources. Chang Gao 0002, Zuowen Wang, Longbiao Cheng, Shih-Chii Liu, Tobi Delbruck |
AAAI | 2 |
| 2024 | Epilepsy Seizure Detection and Prediction using an Approximate Spiking Convolutional TransformerabstractEpilepsy is a common disease of the nervous system. Timely prediction of seizures and intervention treatment can significantly reduce the accidental injury of patients and protect the life and health of patients. This paper presents a tiny neuromorphic Spiking Convolutional Transformer, named Spiking Conformer, to detect and predict epileptic seizure segments from scalped long-term electroencephalogram (EEG) recordings. We report evaluation results from the Spiking Conformer model using the Boston Children’s Hospital-MIT (CHB-MIT) EEG dataset. By leveraging spike-based addition operations, the Spiking Conformer significantly reduces the classification computational cost compared to the non-spiking model. Additionally, we introduce an approximate spiking neuron layer to further reduce spike-triggered neuron updates by nearly 38% without sacrificing accuracy. Using raw EEG data as input, the proposed Spiking Conformer achieved an average sensitivity rate of 94.9% and a specificity rate of 99.3% for the seizure detection task, and 96.8%, 89.5% for the seizure prediction task, and needs >10x fewer operations compared to the non-spiking equivalent model. Qinyu Chen, Congyi Sun, Chang Gao 0002, Shih-Chii Liu |
ISCAS | 3 |
| 2024 | OpenDPD: An Open-Source End-to-End Learning & Benchmarking Framework for Wideband Power Amplifier Modeling and Digital Pre-DistortionabstractWith the rise in communication capacity, deep neural networks (DNN) for digital pre-distortion (DPD) to correct non-linearity in wideband power amplifiers (PAs) have become prominent. Yet, there is a void in open-source and measurement-setup-independent platforms for fast DPD exploration and objective DPD model comparison. This paper presents an open-source framework, OpenDPD, crafted in PyTorch, with an associated dataset for PA modeling and DPD learning. We introduce a Dense Gated Recurrent Unit (DGRU)-DPD, trained via a novel end-to-end learning architecture, outperforming previous DPD models on a digital PA (DPA) in the new digital transmitter (DTX) architecture with unconventional transfer characteristics compared to analog PAs. Measurements show our DGRU-DPD achieves an ACPR of -44.69/-44.47dBc and an EVM of -35.22dB for 200MHz OFDM signals. OpenDPD code, datasets and documentation are publicly available at https://github.com/lab-emi/OpenDPD Yizhuo Wu, Gagan Deep Singh, Mohammadreza Beikmirza, Leo C. N. de Vreede, Morteza S. Alavi, Chang Gao 0002 |
ISCAS | 6 |
| 2024 | HAS-RL: A Hierarchical Approximate Scheme Optimized With Reinforcement Learning for NoC-Based NN AcceleratorsabstractNetwork-on-Chip (NoC) is a scalable on-chip communication architecture for the NN accelerator, but with the increase in the number of nodes, the communication delay becomes higher. Applications such as machine learning have a certain resilience to noisy/erroneous transmitted data. Therefore, approximate communication becomes a promising solution to improving performance by reducing traffic loads under the constraint of the acceptable maximum accuracy loss of neural networks. It is a key issue to balance the result quality and the communication delay for approximate NoC systems. The traditional approximate NoC only considers the node-to-node approximation-based dynamic traffic regulation. However, the dynamically changing traffic patterns across different nodes, different times, and different applications lead to a huge search space, which makes it hard to explore an optimal global approximation solution. In this paper, we propose a quality model for different neural networks, which presents the relationship between the quality loss and the data approximate rate. Then, a hierarchical approximate scheme optimized with reinforcement learning (HAS-RL) is proposed and we reduce the complexity of the HAS-RL by reducing the state space and action space, which will reduce the resource overhead as well. After that, we embed a global approximate controller in the NoC system, in which we deploy a policy network trained with the offline reinforcement learning algorithm to adjust the data approximate rates of each node at run time. Compared with the state-of-the-art method, the proposed scheme reduces the average network delay by 13.5% while their accuracies are similar. The proposed HAS-RL only causes an additional area overhead of 1.24% and power consumption of 0.77% compared with the traditional router design. Shize Zhou, Yongqi Xue, Wenjie Fan 0004, Tong Cheng, Jinlun Ji, Chenyang Dai, Wenqing Song, Qinyu Chen, Chang Gao 0002, Li Li 0003 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 10 |
| 2024 | Spartus: A 9.4 TOp/s FPGA-Based LSTM Accelerator Exploiting Spatio-Temporal SparsityabstractLong short-term memory (LSTM) recurrent networks are frequently used for tasks involving time-sequential data, such as speech recognition. Unlike previous LSTM accelerators that either exploit spatial weight sparsity or temporal activation sparsity, this article proposes a new accelerator called "Spartus" that exploits spatio-temporal sparsity to achieve ultralow latency inference. Spatial sparsity is induced using a new column-balanced targeted dropout (CBTD) structured pruning method, producing structured sparse weight matrices for a balanced workload. The pruned networks running on Spartus hardware achieve weight sparsity levels of up to 96% and 94% with negligible accuracy loss on the TIMIT and the Librispeech datasets. To induce temporal sparsity in LSTM, we extend the previous DeltaGRU method to the DeltaLSTM method. Combining spatio-temporal sparsity with CBTD and DeltaLSTM saves on weight memory access and associated arithmetic operations. The Spartus architecture is scalable and supports real-time online speech recognition when implemented on small and large FPGAs. Spartus per-sample latency for a single DeltaLSTM layer of 1024 neurons averages 1 μ s. Exploiting spatio-temporal sparsity on our test LSTM network using the TIMIT dataset leads to 46 × speedup of Spartus over its theoretical hardware performance to achieve 9.4-TOp/s effective batch-1 throughput and 1.1-TOp/s/W power efficiency. Chang Gao 0002, Tobi Delbruck, Shih-Chii Liu |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | An Area-Efficient Ultra-Low-Power Time-Domain Feature Extractor for Edge Keyword SpottingabstractKeyword spotting (KWS) is an important task on edge low-power audio devices. A typical edge KWS system consists of a front-end feature extractor which outputs mel-scale frequency cepstral coefficients (MFCC) features followed by a back-end neural network classifier. KWS edge designs aim for the best power-performance-area metrics. This work proposes an area-efficient ultra-low-power time-domain infinite impulse response (IIR) filter-based feature extractor for a KWS system. It uses a serial architecture, and the architecture is further optimized for a low-cost computing structure and mixed-precision bit selection of the IIR coefficients while maintaining good KWS accuracy. Using a 65 nm process technology and a back-end neural network classifier, this simulated feature extractor has an area of 0.02 mm2and achieves$\mathbf{3.3}\mu \mathbf{W}$@ 1.2 V, and achieves 92.5% accuracy on a 10-keyword, 12-class KWS task using the GSCD dataset. Qinyu Chen, Yaoxing Chang, Kwantae Kim, Chang Gao 0002, Shih-Chii Liu |
ISCAS | 4 |
| 2022 | Skydiver: A Spiking Neural Network Accelerator Exploiting Spatio-Temporal Workload BalanceabstractSpiking neural networks (SNNs) are developed as a promising alternative to artificial neural networks (ANNs) due to their more realistic brain-inspired computing models. SNNs have sparse neuron firing over time, i.e., spatio-temporal sparsity; thus, they are useful to enable energy-efficient hardware inference. However, exploiting spatio-temporal sparsity of SNNs in hardware leads to unpredictable and unbalanced workloads, degrading the energy efficiency. In this work, we propose an FPGA-based convolutional SNN accelerator called Skydiver that exploits spatio-temporal workload balance. We propose the approximate proportional relation construction (APRC) method that can predict the relative workload channel-wisely and a channel-balanced workload schedule (CBWS) method to increase the hardware workload balance ratio to over 90%. Skydiver was implemented on a Xilinx XC7Z045 FPGA and verified on image segmentation and MNIST classification tasks. Results show improved throughput by$1.4\times $and$1.2\times $for the two tasks. Skydiver achieved 22.6KFPS throughput, and$42.4~\mu \text{J}$/image prediction energy on the classification task with 98.5% accuracy. Qinyu Chen, Chang Gao 0002, Xinyuan Fang, Haitao Luan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | Spiking Cochlea With System-Level Local Automatic Gain ControlabstractIncluding local automatic gain control (AGC) circuitry into a silicon cochlea design has been challenging because of transistor mismatch and model complexity. To address this, we present an alternative system-level algorithm that implements channel-specific AGC in a silicon spiking cochlea by measuring the output spike activity of individual channels. The bandpass filter gain of a channel is adapted dynamically to the input amplitude so that the average output spike rate stays within a defined range. Because this AGC mechanism only needs counting and adding operations, it can be implemented at low hardware cost in a future design. We evaluate the impact of the local AGC algorithm on a classification task where the input signal varies over 32dB input range. Two classifier types receiving cochlea spike features were tested on a speech versus noise classification task. The logistic regression classifier achieves an average of 6% improvement and 40.8% relative improvement in accuracy when the AGC is enabled. The deep neural network classifier shows a similar improvement for the AGC case and achieves a higher mean accuracy of 96% compared to the best accuracy of 91% from the logistic regression classifier. Ilya Kiselev, Chang Gao 0002, Shih-Chii Liu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | Cerebron: A Reconfigurable Architecture for Spatiotemporal Sparse Spiking Neural NetworksabstractSpiking neural networks (SNNs) are promising alternatives to artificial neural networks (ANNs) since they are more realistic brain-inspired computing models. SNNs have sparse neuron firing over time, i.e., spatiotemporal sparsity; thus, they are helpful in enabling energy-efficient hardware inference. However, exploiting the spatiotemporal sparsity of SNNs in hardware leads to unpredictable and unbalanced workloads, degrading the energy efficiency. Compared to SNNs with simple fully connected structures, those extensive structures (e.g., standard convolutions, depthwise convolutions, and pointwise convolutions) can deal with more complicated tasks but lead to difficulties in hardware mapping. In this work, we propose a novel reconfigurable architecture, Cerebron, which can fully exploit the spatiotemporal sparsity in SNNs with maximized data reuse and propose optimization techniques to improve the efficiency and flexibility of the hardware. To achieve flexibility, the reconfigurable compute engine is compatible with a variety of spiking layers and supports inter-computing-unit (CU) and intra-CU reconfiguration. The compute engine can exploit data reuse and guarantee parallel data access when processing different convolutions to achieve memory efficiency. A two-step data sparsity exploitation method is introduced to leverage the sparsity of discrete spikes and reduce the computation time. Besides, an online channelwise workload scheduling strategy is designed to reduce the latency further. Cerebron is verified on image segmentation and classification tasks using a variety of state-of-the-art spiking network structures. Experimental results show that Cerebron has achieved at least$17.5\times $prediction energy reduction and$20\times $speedup compared with state-of-the-art field-programmable gate array (FPGA)-based accelerators. Qinyu Chen, Chang Gao 0002 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2020 | Recurrent Neural Network Control of a Hybrid Dynamical Transfemoral Prosthesis with EdgeDRNN AcceleratorabstractLower leg prostheses could improve the life quality of amputees by increasing comfort and reducing energy to locomote, but currently control methods are limited in modulating behaviors based upon the human's experience. This paper describes the first steps toward learning complex controllers for dynamical robotic assistive devices. We provide the first example of behavioral cloning to control a powered transfemoral prostheses using a Gated Recurrent Unit (GRU) based recurrent neural network (RNN) running on a custom hardware accelerator that exploits temporal sparsity. The RNN is trained on data collected from the original prosthesis controller. The RNN inference is realized by a novel EdgeDRNN accelerator in real-time. Experimental results show that the RNN can replace the nominal PD controller to realize end-to-end control of the AMPRO3 prosthetic leg walking on flat ground and unforeseen slopes with comparable tracking accuracy. EdgeDRNN computes the RNN about 240 times faster than real time, opening the possibility of running larger networks for more complex tasks in the future. Implementing an RNN on this real-time dynamical system with impacts sets the ground work to incorporate other learned elements of the human-prosthesis system into prosthesis control. Chang Gao 0002, Rachel Gehlhar, Aaron D. Ames, Shih-Chii Liu, Tobi Delbruck |
ICRA | 1 |
| 2019 | Live Demonstration: Real-Time Spoken Digit Recognition using the DeltaRNN AcceleratorabstractThis demonstration shows a real-time continuous speech recognition hardware system using our previously published DeltaRNN accelerator that enables low latency recurrent neural network (RNN) computation. The network is trained on augmented audio samples from the TIDIGITS dataset to achieve a label error rate (LER) of 2.31%. It is implemented on a Xilinx Zynq-7100 FPGA running at 1 MHz. The incremental RNN power consumption is 30 mW. Visitors interact with the system by speaking digits into a microphone connected to the FPGA system and the classification outputs of the network are continuously displayed on a laptop screen in real time. Chang Gao 0002, Stefan Braun 0005, Ilya Kiselev, Jithendar Anumula, Tobi Delbruck, Shih-Chii Liu |
ISCAS | 1 |
| 2019 | Real-Time Speech Recognition for IoT Purpose using a Delta Recurrent Neural Network AcceleratorabstractThis paper describes a continuous speech recognition hardware system that uses a delta recurrent neural network accelerator (DeltaRNN) implemented on a Xilinx Zynq-7100 FPGA to enable low latency recurrent neural network (RNN) computation. The implemented network consists of a single-layer RNN with 256 gated recurrent unit (GRU) neurons and is driven by input features generated either from the output of a filter bank running on the ARM core of the FPGA in a PmodMic3 microphone setup or from the asynchronous outputs of a spiking silicon cochlea circuit. The microphone setup achieves 7.1 ms minimum latency and 177 frames-per-second (FPS) maximum throughput while the cochlea setup achieves 2.9 ms minimum latency and 345 FPS maximum throughput. The low latency and 70 mW power consumption of the DeltaRNN makes it suitable as an IoT computing platform. Chang Gao 0002, Stefan Braun 0005, Ilya Kiselev, Jithendar Anumula, Tobi Delbruck, Shih-Chii Liu |
ISCAS | 1 |
| 2018 | DeltaRNN: A Power-efficient Recurrent Neural Network AcceleratorabstractRecurrent Neural Networks (RNNs) are widely used in speech recognition and natural language processing applications because of their capability to process temporal sequences. Because RNNs are fully connected, they require a large number of weight memory accesses, leading to high power consumption. Recent theory has shown that an RNN delta network update approach can reduce memory access and computes with negligible accuracy loss. This paper describes the implementation of this theoretical approach in a hardware accelerator called "DeltaRNN" (DRNN). The DRNN updates the output of a neuron only when the neuron»s activation changes by more than a delta threshold. It was implemented on a Xilinx Zynq-7100 FPGA. FPGA measurement results from a single-layer RNN of 256 Gated Recurrent Unit (GRU) neurons show that the DRNN achieves 1.2 TOp/s effective throughput and 164 GOp/s/W power efficiency. The delta update leads to a 5.7x speedup compared to a conventional RNN update because of the sparsity created by the DN algorithm and the zero-skipping ability of DRNN. Chang Gao 0002, Daniel Neil, Enea Ceolini, Shih-Chii Liu, Tobi Delbruck |
FPGA | 1 |
| 2017 | On-chip ID generation for multi-node implantable devices using SA-PUFabstractThis paper presents a 64-bit on-chip identification system featuring low power consumption and randomness compensation for multi-node bio-implantable devices. A sense amplifier based bit-cell is proposed to realize the silicon physical unclonable function, providing a unique value whose probability has a uniform distribution and minimized influence from the temperature and supply variation. The entire system is designed and implemented in a typical 0.35 μm CMOS technology, including an array of 64 bit-cells, readout circuits, and digital controllers for data interfaces. Simulated results show that the proposed bit-cell design achieved a uniformity of 50.24% and a uniqueness of 50.03% for generated IDs. The system achieved an energy consumption of 6.0 pJ per bit with parallel outputs and 17.3 pJ per bit with serial outputs. Chang Gao 0002, Sara S. Ghoreishizadeh, Yan Liu 0016, Timothy G. Constandinou |
ISCAS | 1 |