EDBT 2026 Demo / reviewers in the wild / expert
Shih-Chii Liu
dblp:10/3688
· DBLP profile ↗
117ranked-venue papers
16as first author
24since 2021 · last 2025
0000-0002-7557-045XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 58 · 5 first-author · 10 since 2021Artificial intelligence and machine learning · 48 · 11 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 1 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Modulating State Space Model with SlowFast Framework for Compute-Efficient Ultra Low-Latency Speech EnhancementabstractDeep learning-based speech enhancement (SE) methods often face significant computational challenges when needing to meet low-latency requirements because of the increased number of frames to be processed. This paper introduces the SlowFast framework which aims to reduce computation costs specifically when low-latency enhancement is needed. The framework consists of a slow branch that analyzes the acoustic environment at a low frame rate, and a fast branch that performs SE in the time domain at the needed higher frame rate to match the required latency. Specifically, the fast branch employs a state space model where its state transition process is dynamically modulated by the slow branch. Experiments on a SE task with a 2 ms algorithmic latency requirement using the Voice Bank + Demand dataset show that our approach reduces computation cost by 70% compared to a baseline single-branch network with equivalent parameters, without compromising enhancement performance. Furthermore, by leveraging the SlowFast framework, we implemented a network that achieves an algorithmic latency of just 62.5 μs (one sample point at 16 kHz sample rate) with a computation cost of 100 M MACs/s, while scoring a PESQ-NB of 3.12 and SISNR of 16.62. Longbiao Cheng, Ashutosh Pandey 0004, Buye Xu, Tobi Delbruck, Vamsi K. Ithapu, Shih-Chii Liu |
ICASSP | 6 |
| 2025 | CustomFit: Customizing Task-Optimized and Size-Specific Subnets from a Single Base Network During RuntimeabstractWe address the challenge of efficiently deploying networks across hardware platforms with varying resource constraints and task requirements. While conventional methods train a base network that can be split into subnets of different sizes for deployment, they typically focus on a single task. This paper introduces CustomFit, a framework that trains a base network from which different-sized subnets tailored to multiple tasks can be segmented during inference. CustomFit uses task-specific feature mapping to optimize the base network for each task without altering its weights. Subnet segmentation is guided by the task’s feature mapping process and the desired size. By training a base network with CustomFit on Speech Enhancement (SE) and Spoken Language Understanding (SLU) tasks, we show that 30 subnets of varying sizes can be segmented. Compared to state-of-the-art dynamic networks, CustomFit achieves similar SE scores and improves SLU accuracy by 2.75%, while using 2.9× fewer parameters in the base network. The parameter sizes of task-specific subnets range from 11k to 720k, enabling flexible deployment across a range of edge platforms from microcontroller units to mobile phones. Longbiao Cheng, Shih-Chii Liu |
ISCAS | 3 |
| 2025 | Steerable Zero-Shot Neural Architecture Search for Efficient Edge InferenceabstractRecent advancements in Neural Architecture Search (NAS) have introduced methods capable of identifying optimal neural network architectures in minutes on Graphical Processing Units (GPUs) using zero-shot proxies, but mainly focus on single-objective optimization. NAS is also used to discover efficient architectures for edge devices. However, addressing the diverse hardware and application constraints specific to edge platforms remains a significant challenge. In this paper, we introduce a zero-shot NAS approach designed to generate hardware-aware architectures, combined with a selection technique that allows adaptable model optimization across various deployment scenarios without re-executing the search. We demonstrate the flexibility of our solution by benchmarking the Google Coral Edge Tensor Processing Unit (TPU). Our technique led to the efficient exploration of the architecture space of NAS-Bench-201 (NB201) in under a minute, accelerating the search by 25× compared to previous work while maintaining comparable accuracies of 93.24% on CIFAR-10 and 42.99% on ImageNet16-120. Simon Narduzzi, Rémy Vuagniaux, Kishan Sharma, Shih-Chii Liu, L. Andrea Dunbar |
ISCAS | 4 |
| 2025 | Online Prediction of Core Body Temperature from Sweat Wearable Printed Sensors Using Recurrent Neural NetworkabstractReal-time monitoring of core body temperature (CBT) is important for preventing heat-related physiological problems during work or exercise. This paper shows how sweat biomarkers (sodium and potassium concentrations) as measured by printed sensors on a sweat wearable patch can be used for online continual prediction of CBT. These sensor measurements and additional biomarker data (heart rate and regional sweat rate) were collected from two healthy male athletes during controlled cycling sessions. The biomarker data was used to continuously predict CBT with one of three models: a recurrent neural network (RNN), a multilayer perceptron, and a linear regression model. The results show that with a window size of 30 seconds, sweat sodium and potassium concentrations outperform other biomarker pairs in predicting CBT. Out of all three models, an RNN model with only 70.8 K parameters achieved the lowest prediction error of 0.04 °C using the two sweat biomarkers. These findings support the use of sweat-based non-invasive monitoring systems for reliable online CBT prediction. Silvia Demuru, Céline Lafaye, Brince Paul Kunnel, Cyril Besson, Ilya Kiselev, Vincent Gremeaux, Mathieu Saubade, Danick Briand, Shih-Chii Liu |
ISCAS | 11 |
| 2024 | Exploiting Symmetric Temporally Sparse BPTT for Efficient RNN TrainingabstractRecurrent Neural Networks (RNNs) are useful in temporal sequence tasks. However, training RNNs involves dense matrix multiplications which require hardware that can support a large number of arithmetic operations and memory accesses. Implementing online training of RNNs on the edge calls for optimized algorithms for an efficient deployment on hardware. Inspired by the spiking neuron model, the Delta RNN exploits temporal sparsity during inference by skipping over the update of hidden states from those inactivated neurons whose change of activation across two timesteps is below a defined threshold. This work describes a training algorithm for Delta RNNs that exploits temporal sparsity in the backward propagation phase to reduce computational requirements for training on the edge. Due to the symmetric computation graphs of forward and backward propagation during training, the gradient computation of inactivated neurons can be skipped. Results show a reduction of ∼80% in matrix operations for training a 56k parameter Delta LSTM on the Fluent Speech Commands dataset with negligible accuracy loss. Logic simulations of a hardware accelerator designed for the training algorithm show 2-10X speedup in matrix computations for an activation sparsity range of 50%-90%. Additionally, we show that the proposed Delta RNN training will be useful for online incremental learning on edge devices with limited computing resources. Chang Gao 0002, Zuowen Wang, Longbiao Cheng, Shih-Chii Liu, Tobi Delbruck |
AAAI | 6 |
| 2024 | Regularized Parameter Uncertainty for Improving Generalization in Reinforcement LearningabstractIn order for reinforcement learning (RL) agents to be deployed in real-world environments, they must be able to generalize to unseen environments. However, RL struggles with out-of-distribution generalization, often due to overfitting the particulars of the training environment. Although regularization techniques from supervised learning can be applied to avoid over-fitting, the differences between supervised learning and RL limit their application. To address this, we propose the Signal-to-Noise Ratio regulated Parameter Uncertainty Network (SNR PUN) for RL. We introduce SNR as a new measure of regularizing the parameter uncertainty of a network and provide a formal analysis explaining why SNR regularization works well for RL. We demonstrate the effectiveness of our proposed method to generalize in several simulated environments; and in a physical system showing the possibility of using SNR PUN for applying RL to real-world applications. Pehuen Moure, Longbiao Cheng, Joachim Ott, Zuowen Wang, Shih-Chii Liu |
CVPR | 5 |
| 2024 | Dynamic Gated Recurrent Neural Network for Compute-efficient Speech EnhancementabstractThis paper introduces a new Dynamic Gated Recurrent Neural Network (DG-RNN) for compute-efficient speech enhancement models running on resource-constrained hardware platforms.It leverages the slow evolution characteristic of RNN hidden states over steps, and updates only a selected set of neurons at each step by adding a newly proposed select gate to the RNN model.This select gate allows the computation cost of the conventional RNN to be reduced during network inference.As a realization of the DG-RNN, we further propose the Dynamic Gated Recurrent Unit (D-GRU) which does not require additional parameters.Test results obtained from several state-ofthe-art compute-efficient RNN-based speech enhancement architectures using the DNS challenge dataset, show that the D-GRU based model variants maintain similar speech intelligibility and quality metrics comparable to the baseline GRU based models even with an average 50% reduction in GRU computes. Longbiao Cheng, Ashutosh Pandey 0004, Buye Xu, Tobi Delbruck, Shih-Chii Liu |
INTERSPEECH | 5 |
| 2024 | Epilepsy Seizure Detection and Prediction using an Approximate Spiking Convolutional TransformerabstractEpilepsy is a common disease of the nervous system. Timely prediction of seizures and intervention treatment can significantly reduce the accidental injury of patients and protect the life and health of patients. This paper presents a tiny neuromorphic Spiking Convolutional Transformer, named Spiking Conformer, to detect and predict epileptic seizure segments from scalped long-term electroencephalogram (EEG) recordings. We report evaluation results from the Spiking Conformer model using the Boston Children’s Hospital-MIT (CHB-MIT) EEG dataset. By leveraging spike-based addition operations, the Spiking Conformer significantly reduces the classification computational cost compared to the non-spiking model. Additionally, we introduce an approximate spiking neuron layer to further reduce spike-triggered neuron updates by nearly 38% without sacrificing accuracy. Using raw EEG data as input, the proposed Spiking Conformer achieved an average sensitivity rate of 94.9% and a specificity rate of 99.3% for the seizure detection task, and 96.8%, 89.5% for the seizure prediction task, and needs >10x fewer operations compared to the non-spiking equivalent model. Qinyu Chen, Congyi Sun, Chang Gao 0002, Shih-Chii Liu |
ISCAS | 4 |
| 2024 | DeltaDEQ: Exploiting Heterogeneous Convergence for Accelerating Deep Equilibrium IterationsabstractImplicit neural networks including deep equilibrium models have achieved superior task performance with better parameter efficiency in various applications. However, it is often at the expense of higher computation costs during inference. In this work, we identify a phenomenon named $\textbf{heterogeneous convergence}$ that exists in deep equilibrium models and other iterative methods. We observe much faster convergence of state activations in certain dimensions therefore indicating the dimensionality of the underlying dynamics of the forward pass is much lower than the defined dimension of the states. We thereby propose to exploit heterogeneous convergence by storing past linear operation results (e.g., fully connected and convolutional layers) and only propagating the state activation when its change exceeds a threshold. Thus, for the already converged dimensions, the computations can be skipped. We verified our findings and reached 84\% FLOPs reduction on the implicit neural representation task, 73\% on the Sintel and 76\% on the KITTI datasets for the optical flow estimation task while keeping comparable task accuracy with the models that perform the full update. Zuowen Wang, Longbiao Cheng, Pehuen Moure, Niklas Hahn 0002, Shih-Chii Liu |
NeurIPS | 5 |
| 2024 | Spartus: A 9.4 TOp/s FPGA-Based LSTM Accelerator Exploiting Spatio-Temporal SparsityabstractLong short-term memory (LSTM) recurrent networks are frequently used for tasks involving time-sequential data, such as speech recognition. Unlike previous LSTM accelerators that either exploit spatial weight sparsity or temporal activation sparsity, this article proposes a new accelerator called "Spartus" that exploits spatio-temporal sparsity to achieve ultralow latency inference. Spatial sparsity is induced using a new column-balanced targeted dropout (CBTD) structured pruning method, producing structured sparse weight matrices for a balanced workload. The pruned networks running on Spartus hardware achieve weight sparsity levels of up to 96% and 94% with negligible accuracy loss on the TIMIT and the Librispeech datasets. To induce temporal sparsity in LSTM, we extend the previous DeltaGRU method to the DeltaLSTM method. Combining spatio-temporal sparsity with CBTD and DeltaLSTM saves on weight memory access and associated arithmetic operations. The Spartus architecture is scalable and supports real-time online speech recognition when implemented on small and large FPGAs. Spartus per-sample latency for a single DeltaLSTM layer of 1024 neurons averages 1 μ s. Exploiting spatio-temporal sparsity on our test LSTM network using the TIMIT dataset leads to 46 × speedup of Spartus over its theoretical hardware performance to achieve 9.4-TOp/s effective batch-1 throughput and 1.1-TOp/s/W power efficiency. Chang Gao 0002, Tobi Delbruck, Shih-Chii Liu |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Biologically-Inspired Continual Learning of Human Motion SequencesabstractThis work proposes a model for continual learning on tasks involving temporal sequences, specifically, human motions. It improves on a recently proposed brain-inspired replay model (BI-R) by building a biologically-inspired conditional temporal variational autoencoder (BI-CTVAE), which instantiates a latent mixture-of-Gaussians for class representation. We investigate a novel continual-learning-to-generate (CL2Gen) scenario where the model generates motion sequences of different classes. The generative accuracy of the model is tested over a set of tasks. The final classification accuracy of BI-CTVAE on a human motion dataset after sequentially learning all action classes is 78%, which is 63% higher than using no-replay, and only 5.4% lower than a state-of-the-art offline trained GRU model. Joachim Ott, Shih-Chii Liu |
ICASSP | 2 |
| 2023 | An Area-Efficient Ultra-Low-Power Time-Domain Feature Extractor for Edge Keyword SpottingabstractKeyword spotting (KWS) is an important task on edge low-power audio devices. A typical edge KWS system consists of a front-end feature extractor which outputs mel-scale frequency cepstral coefficients (MFCC) features followed by a back-end neural network classifier. KWS edge designs aim for the best power-performance-area metrics. This work proposes an area-efficient ultra-low-power time-domain infinite impulse response (IIR) filter-based feature extractor for a KWS system. It uses a serial architecture, and the architecture is further optimized for a low-cost computing structure and mixed-precision bit selection of the IIR coefficients while maintaining good KWS accuracy. Using a 65 nm process technology and a back-end neural network classifier, this simulated feature extractor has an area of 0.02 mm2and achieves$\mathbf{3.3}\mu \mathbf{W}$@ 1.2 V, and achieves 92.5% accuracy on a 10-keyword, 12-class KWS task using the GSCD dataset. Qinyu Chen, Yaoxing Chang, Kwantae Kim, Chang Gao 0002, Shih-Chii Liu |
ISCAS | 5 |
| 2023 | A 3.11 μ W 40 n V/ ✓Hz Instrumentation Amplifier for Bio-Impedance Sensors Exploiting Positive-Feedback-Assisted Gain BoostingabstractThis paper presents a low-power and low-noise instrumentation amplifier (IA) for bio-impedance sensing ap-plications. To solve the loop gain reduction problem when the input resistor in a transconductance (TC) stage decreases as low as$< 10\mathrm{k}\Omega$, two positive-feedback-based gain boosting techniques are applied to the flipped voltage follower circuit: (1) a super gain casco de current mirror and (2) a current mirror amplifier with local positive feedback. Simulated in a 65 nm CMOS process, the designed IA ensures a 70 dB DC loop gain within the TC stage while achieving 40$\mathbf{nV}/\sqrt{\mathbf{Hz}}$input referred noise floor and a 369kHz bandwidth at 3.11$\mu \mathrm{W}$power consumption, 1 V supply. The enhanced loop gain leads to only a 0.04 dB (0.46%) IA gain deviation over the PVT variations. Kwantae Kim, Shih-Chii Liu |
ISCAS | 2 |
| 2023 | End-to-End Prediction of Sodium Concentration from Uncalibrated Sodium ISFETsabstractIon-selective field-effect transistors (ISFETs) are widely used for chemical sensing in biomedical and environmental applications. They require calibration before deployment in the field because of individual sensor response variations and temporal drift in their readout. However, calibration can be time-consuming if a large number of ISFETs are to be deployed. This work proposes an end-to-end prediction neural network where individual sensor calibrations are not needed. We train the network to predict the ionic concentration by using a simulated dataset of responses from a physical ISFET model to varying sodium concentrations. The model includes the known non-idealities of real ISFET sensors. Our network also outputs a confidence interval for the prediction which can be useful for determining the quality of the prediction. On a dataset of real sodium ISFET recordings, our end-to-end prediction network gave a decrease of at least 42% in the prediction error of sodium concentration compared to that from ISFETs calibrated using two manual methods. Meritxell Rovira, Yuhuang Hu, Cecilia Jiménez-Jorquera, Shih-Chii Liu |
ISCAS | 5 |
| 2022 | Optimizing The Consumption Of Spiking Neural Networks With Activity RegularizationabstractReducing energy consumption is a critical point for neural network models running on edge devices. In this regard, reducing the number of multiply-accumulate (MAC) operations of Deep Neural Networks (DNNs) running on edge hardware accelerators will reduce the energy consumption during inference. Spiking Neural Networks (SNNs) are an example of bio-inspired techniques that can further save energy by using binary activations, and avoid consuming energy when not spiking. The networks can be configured for equivalent accuracy on a task through DNN-to-SNN conversion frameworks but their conversion is based on rate coding therefore the synaptic operations can be high. In this work, we look into different techniques to enforce sparsity on the neural network activation maps and compare the effect of different training regularizers on the efficiency of the optimized DNNs and SNNs. Simon Narduzzi, Siavash Arjomand Bigdeli, Shih-Chii Liu, L. Andrea Dunbar |
ICASSP | 3 |
| 2022 | T-NGA: Temporal Network Grafting Algorithm for Learning to Process Spiking Audio Sensor EventsabstractSpiking silicon cochlea sensors encode sound as an asynchronous stream of spikes from different frequency channels. The lack of labeled training datasets for spiking cochleas makes it difficult to train deep neural networks on the outputs of these sensors. This work proposes a self-supervised method called Temporal Network Grafting Algorithm (T-NGA), which grafts a recurrent network pretrained on spectrogram features so that the network works with the cochlea event features. T-NGA training requires only temporally aligned audio spectrograms and event features. Our experiments show that the accuracy of the grafted network was similar to the accuracy of a supervised network trained from scratch on a speech recognition task using events from a software spiking cochlea model. Despite the circuit non-idealities of the spiking silicon cochlea, the grafted network accuracy on the silicon cochlea spike recordings was only about 5% lower than the supervised network accuracy using the N-TIDIGITS18 dataset. T-NGA can train networks to process spiking audio sensor events in the absence of large labeled spike datasets. Yuhuang Hu, Shih-Chii Liu |
ICASSP | 3 |
| 2022 | Exploiting Spatial Sparsity for Event Cameras with Visual TransformersabstractEvent cameras report local changes of brightness through an asynchronous stream of output events. Events are spatially sparse at pixel locations with little brightness variation. We propose using a visual transformer (ViT) architecture to leverage its ability to process a variable-length input. The input to the ViT consists of events that are accumulated into time bins and spatially separated into non-overlapping sub-regions called patches. Patches are selected when the number of nonzero pixel locations within a sub-region is above a threshold. We show that by fine-tuning a ViT model on these selected active patches, we can reduce the average number of patches fed into the backbone during the inference by at least 50% with only a minor drop (0.34%) of the classification accuracy on the N-Caltech101 dataset. This reduction translates into a decrease of 51% in Multiply-Accumulate (MAC) operations and an increase of 46% in the inference speed using a server CPU. Zuowen Wang, Yuhuang Hu, Shih-Chii Liu |
ICIP | 3 |
| 2022 | Kernel Modulation: A Parameter-Efficient Method for Training Convolutional Neural NetworksabstractDeep Neural Networks, particularly Convolutional Neural Networks (ConvNets), have achieved incredible success in many vision tasks, but they usually require millions of parameters for good accuracy performance. With increasing applications that use ConvNets, updating hundreds of networks for multiple tasks on an embedded device can be costly in terms of memory, bandwidth, and energy. Approaches to reduce this cost include model compression and parameter-efficient models that adapt a subset of network layers for each new task. This work proposes a novel parameter-efficient kernel modulation (KM) method that adapts all parameters of a base network instead of a subset of layers. KM uses lightweight task-specialized kernel modulators that require only an additional 1.4% of the base network parameters. With multiple tasks, only the task-specialized KM weights are communicated and stored on the end-user device. We applied this method in training ConvNets for Transfer Learning and Meta-Learning scenarios. Our results show that KM delivers up to 9% higher accuracy compared to other parameter-efficient methods on the Transfer Learning benchmark. Yuhuang Hu, Shih-Chii Liu |
ICPR | 2 |
| 2022 | Spiking Cochlea With System-Level Local Automatic Gain ControlabstractIncluding local automatic gain control (AGC) circuitry into a silicon cochlea design has been challenging because of transistor mismatch and model complexity. To address this, we present an alternative system-level algorithm that implements channel-specific AGC in a silicon spiking cochlea by measuring the output spike activity of individual channels. The bandpass filter gain of a channel is adapted dynamically to the input amplitude so that the average output spike rate stays within a defined range. Because this AGC mechanism only needs counting and adding operations, it can be implemented at low hardware cost in a future design. We evaluate the impact of the local AGC algorithm on a classification task where the input signal varies over 32dB input range. Two classifier types receiving cochlea spike features were tested on a speech versus noise classification task. The logistic regression classifier achieves an average of 6% improvement and 40.8% relative improvement in accuracy when the AGC is enabled. The deep neural network classifier shows a similar improvement for the AGC case and achieves a higher mean accuracy of 96% compared to the best accuracy of 91% from the logistic regression classifier. Ilya Kiselev, Chang Gao 0002, Shih-Chii Liu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2022 | Contraction of Dynamically Masked Deep Neural Networks for Efficient Video ProcessingabstractSequential data such as video are characterized by spatio-temporal redundancies. As of yet, few deep learning algorithms exploit them to decrease the often massive cost during inference. This work leverages correlations in video data to reduce the size and run-time cost of deep neural networks. Drawing upon the simplicity of the typically used ReLU activation function, we replace this function by dynamically updating masks. The resulting network is a simple chain of matrix multiplications and bias additions, which can be contracted into a single weight matrix and bias vector. Inference then reduces to an affine transformation of the input sample with these contracted parameters. We show that the method is akin to approximating the neural network with a first-order Taylor expansion around a dynamically updating reference point. For triggering these updates, one static and three data-driven mechanisms are analyzed. We evaluate the proposed algorithm on a range of tasks, including pose estimation on surveillance data, road detection on KITTI driving scenes, object detection on ImageNet videos, as well as denoising MNIST digits, and obtain compression rates up to$3.6\times $. Bodo Rueckauer, Shih-Chii Liu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Predicting Hydration Status Using Machine Learning Models From Physiological and Sweat Biomarkers During Endurance Exercise: A Single Case StudyabstractImproper hydration routines can reduce athletic performance. Recent studies show that data from noninvasive biomarker recordings can help to evaluate the hydration status of subjects during endurance exercise. These studies are usually carried out on multiple subjects. In this work, we present the first study on predicting hydration status using machine learning models from single-subject experiments, which involve 32 exercise sessions of constant moderate intensity performed with and without fluid intake. During exercise, we measured four noninvasive physiological and sweat biomarkers including heart rate, core temperature, sweat sodium concentration, and whole-body sweat rate. Sweat sodium concentration was measured from six body regions using absorbent patches. We used three machine learning models to determine the percentage of body weight loss as an indicator of dehydration with these biomarkers and compared the prediction accuracy. The results on this single subject show that these models gave similar mean absolute errors, while in general the nonlinear models slightly outperformed the linear model in most of the experiments. The prediction accuracy of using the whole-body sweat rate or heart rate was higher than using core temperature or sweat sodium concentration. In addition, the model trained on the sweat sodium concentration collected from the arms gave slightly better accuracy than from the other five body regions. This exploratory work paves the way for the use of these machine learning models to develop personalized health monitoring together with emerging, noninvasive wearable sensor devices. Céline Lafaye, Mathieu Saubade, Cyril Besson, Josep Maria Margarit-Taulé, Vincent Gremeaux, Shih-Chii Liu |
IEEE J. Biomed. Health Informatics | 7 |
| 2021 | Temporal Pattern Coding in Deep Spiking Neural NetworksabstractDeep Artificial Neural Networks (ANNs) employ a simplified analog neuron model that mimics the rate transfer function of integrate-and-fire neurons. In Spiking Neural Networks (SNNs), the predominant information transmission method is based on rate codes. This code is inefficient from a hardware perspective because the number of transmitted spikes is proportional to the encoded analog value. Alternate codes such as temporal codes that are based on single spikes are difficult to scale up for large networks due to their sensitivity to spike timing noise. Here we present a study of an encoding scheme based on temporal spike patterns. This scheme inherits the efficiency of temporal codes but retains the robustness of rate codes. The pattern code is evaluated on MNIST, CIFAR-10, and ImageNet image classification tasks. We compare the network performance of ANNs, rate-coded SNNs, and temporal-coded SNNs, using the classification error and operation count as performance metrics. We also estimate the power consumption of the digital logic needed for the operations associated with each encoding type, and the impact of the bit precision of the weights and activations. On ImageNet, the temporal pattern code achieves up to$35\times$reduction in the estimated power consumption compared to the rate-coded SNN, and$42\times$compared to the ANN. The classification error of the pattern-coded SNN is increased by$< {1\%}$compared to the ANN, and decreased by 2% compared to the rate-coded SNN. Bodo Rueckauer, Shih-Chii Liu |
IJCNN | 2 |
| 2021 | Reducing Latency in a Converted Spiking Video Segmentation NetworkabstractSpiking Neural Networks (SNNs) can be configured to produce almost-equivalent accurate Analog Neural Networks (ANNs) by various ANN-SNN conversion methods. Most of these methods are applied to classification and object detection networks tested on frame-based datasets. In this work, we demonstrate a converted SNN for image segmentation and applied to a natural video dataset. Instead of resetting the network state with each input frame, we capitalize on the temporal redundancy between adjacent frames in a natural scene, and propose an interval reset method where the network state is reset after a fixed number of frames. We studied the trade-off between accuracy and latency with the number of interval reset frames. We also applied layer-specific normalization and early stopping to speed up network convergence and to reduce the latency. Our results show that the SNN achieved a 35.7x increase in convergence speed with only 1.5% accuracy drop using an interval reset of 20 frames. Qinyu Chen, Bodo Rueckauer, Tobi Delbruck, Shih-Chii Liu |
ISCAS | 5 |
| 2021 | Event-Driven Local Gain Control on a Spiking Cochlea SensorabstractIncluding local automatic gain control (AGC) circuitry into a silicon cochlea design can be challenging because of transistor mismatch and model complexity. To address this, we present an alternative system-level algorithm that implements channel-specific AGC by using the output spikes of a spiking silicon cochlea. By measuring the output spike activity of each channel, the bandpass filter gain of a channel is adapted dynamically to the input sound amplitude so that the average output spike rate stays within a defined range. We evaluate the effect of our local AGC algorithm on a classification task where the input signal varies over a large amplitude range. Results on a task to classify speech versus noise show that a classifier trained on spike responses of a cochlea with local AGC maintains an average of 25% higher accuracy over a 32 dB input dynamic range, compared to the case when the AGC is disabled. Ilya Kiselev, Shih-Chii Liu |
ISCAS | 2 |
| 2020 | Learning to Exploit Multiple Vision Modalities by Using Grafted Networks
Yuhuang Hu, Tobi Delbruck, Shih-Chii Liu |
ECCV (16) | 3 |
| 2020 | Recurrent Neural Network Control of a Hybrid Dynamical Transfemoral Prosthesis with EdgeDRNN AcceleratorabstractLower leg prostheses could improve the life quality of amputees by increasing comfort and reducing energy to locomote, but currently control methods are limited in modulating behaviors based upon the human's experience. This paper describes the first steps toward learning complex controllers for dynamical robotic assistive devices. We provide the first example of behavioral cloning to control a powered transfemoral prostheses using a Gated Recurrent Unit (GRU) based recurrent neural network (RNN) running on a custom hardware accelerator that exploits temporal sparsity. The RNN is trained on data collected from the original prosthesis controller. The RNN inference is realized by a novel EdgeDRNN accelerator in real-time. Experimental results show that the RNN can replace the nominal PD controller to realize end-to-end control of the AMPRO3 prosthetic leg walking on flat ground and unforeseen slopes with comparable tracking accuracy. EdgeDRNN computes the RNN about 240 times faster than real time, opening the possibility of running larger networks for more complex tasks in the future. Implementing an RNN on this real-time dynamical system with impacts sets the ground work to incorporate other learned elements of the human-prosthesis system into prosthesis control. Chang Gao 0002, Rachel Gehlhar, Aaron D. Ames, Shih-Chii Liu, Tobi Delbruck |
ICRA | 4 |
| 2020 | Lessons Learned the Hard Wayabstract“Fail often to succeed sooner” is a common mantra that we are told is the secret to success. When reporting research results, however, scholars rarely write about their failed attempts and only focus on the successful ones. Perhaps the source of this disconnect between what we preach and what we do can be found in the underlying assumption that published work is meant to move the field forward and failed attempts supposedly do not. The goal of the confessions presented in this paper is to show that even failed attempts are genuine and valuable contributions to our field provided that we learn from our mistakes and correct them. The 27 confessions span from planning oversights, digital and analog design errors, misunderstanding of devices, overlooked parasitics, LVS errors, and troubles in testing. Tobi Delbruck, Ibrahim M. Elfadel, Shahzad Muzaffar, Germain Haessig, Bo Wang 0012, Amine Bermak, Rui Graca, Luis A. Camuñas-Mesa, Bathiya Senevirathna, Pamela Abshire, Bernabé Linares-Barranco, Saeed Afshar, Shih-Chii Liu, Runchun Wang, Piotr Dudek, Stephen J. Carey, José M. de la Rosa 0001, Marc Dandin, Sheung Lu, Vincent Frick, Teresa Serrano-Gotarredona, Paula López Martinez 0001, Melika Payvand, Advait Madhavan, Eric R. Fossum, Juan Camilo Vasquez Tieck, Yan Liu 0016, Timothy G. Constandinou, Alexander Serb, Ricardo Carmona-Galán, Robert Nawrocki, Walter D. Leon-Salas |
ISCAS | 13 |
| 2020 | Optimal Sampling of Parametric Families: Implications for Machine LearningabstractIt is well known in machine learning that models trained on a training set generated by a probability distribution function perform far worse on test sets generated by a different probability distribution function. In the limit, it is feasible that a continuum of probability distribution functions might have generated the observed test set data; a desirable property of a learned model in that case is its ability to describe most of the probability distribution functions from the continuum equally well. This requirement naturally leads to sampling methods from the continuum of probability distribution functions that lead to the construction of optimal training sets. We study the sequential prediction of Ornstein-Uhlenbeck processes that form a parametric family. We find empirically that a simple deep network trained on optimally constructed training sets using the methods described in this letter can be robust to changes in the test set distribution. Adrian E. G. Huber, Jithendar Anumula, Shih-Chii Liu |
Neural Comput. | 3 |
| 2020 | Evaluating Multi-Channel Multi-Device Speech Separation Algorithms in the Wild: A Hardware-Software SolutionabstractEvaluation methods for multi-channel speech separation algorithms in the real world are becoming increasingly important as the number of applications involving audio assistants and hearing aid devices continues to grow. To make such evaluations easier, this paper presents a multi-microphone hardware platform, WHISPER, built specifically for this purpose and its subsequent use for evaluating speech processing algorithms. The platform can also be constructed as an ad-hoc wireless acoustic sensor network (WASN) with high synchronization precision. Using WHISPER, we describe real-world experiments where an example speech separation algorithm is applied to mixtures of varying number of talkers and signal-to-noise ratios. The results when compared with those from a simulated environment, show the usefulness of WASNs and that simulations tend to underestimate the difficulty of speech separation in real-world scenarios. This work represents an important step towards developing a hardware-software framework for evaluating speech processing algorithms in the wild. Enea Ceolini, Ilya Kiselev, Shih-Chii Liu |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | FaSNet: Low-Latency Adaptive Beamforming for Multi-Microphone Audio ProcessingabstractBeamforming has been extensively investigated for multi-channel audio processing tasks. Recently, learning-based beamforming methods, sometimes called neural beamformers, have achieved significant improvements in both signal quality (e.g. signal-to-noise ratio (SNR)) and speech recognition (e.g. word error rate (WER)). Such systems are generally non-causal and require a large context for robust estimation of inter-channel features, which is impractical in applications requiring low-latency responses. In this paper, we propose filter-and-sum network (FaSNet), a time-domain, filter-based beamforming approach suitable for low-latency scenarios. FaSNet has a two-stage system design that first learns frame-level time-domain adaptive beamforming filters for a selected reference channel, and then calculate the filters for all remaining channels. The filtered outputs at all channels are summed to generate the final output. Experiments show that despite its small model size, FaSNet is able to outperform several traditional oracle beamformers with respect to scale-invariant signal-to-noise ratio (SI-SNR) in reverberant speech enhancement and separation tasks. Moreover, when trained with a frequency-domain objective function on the CHiME-3 dataset, FaSNet achieves 14.3% relative word error rate reduction (RWERR) compared with the baseline model. These results show the efficacy of FaSNet particularly in reverberant and noisy signal conditions. Yi Luo 0004, Cong Han 0001, Nima Mesgarani, Enea Ceolini, Shih-Chii Liu |
ASRU | 5 |
| 2019 | Parameter Uncertainty for End-to-end Speech RecognitionabstractRecent work on neural networks with probabilistic parameters has shown that parameter uncertainty improves network regularization. Parameter-specific signal-to-noise ratio (SNR) levels derived from parameter distributions were further found to have high correlations with task importance. However, most of these studies focus on tasks other than automatic speech recognition (ASR). This work investigates end-to-end models with probabilistic parameters for ASR. We demonstrate that probabilistic networks outperform conventional deterministic networks in pruning and domain adaptation experiments carried out on the Wall Street Journal and CHiME-4 datasets. We use parameter-specific SNR information to select parameters for pruning and to condition the parameter updates during adaptation. Experimental results further show that networks with lower SNR parameters (1) tolerate increased sparsity levels during parameter pruning and (2) reduce catastrophic forgetting during domain adaptation. Stefan Braun 0005, Shih-Chii Liu |
ICASSP | 2 |
| 2019 | Event-driven Pipeline for Low-latency Low-compute Keyword Spotting and Speaker Verification SystemabstractThis work presents an event-driven acoustic sensor processing pipeline to power a low-resource voice-activated smart assistant. The pipeline includes four major steps; namely localization, source separation, keyword spotting (KWS) and speaker verification (SV). The pipeline is driven by a front-end binaural spiking silicon cochlea sensor. The timing information carried by the output spikes of the cochlea provide spatial cues for localization and source separation. Spike features are generated with low latencies from the separated source spikes and are used by both KWS and SV which rely on state-of-the-art deep recurrent neural network architectures with a small memory footprint. Evaluation on a self-recorded event dataset based on TIDIGITS shows accuracies of over 93% and 88% on KWS and SV respectively, with minimum system latency of 5 ms on a limited resource device. Enea Ceolini, Jithendar Anumula, Stefan Braun 0005, Shih-Chii Liu |
ICASSP | 4 |
| 2019 | Attention-driven Multi-sensor SelectionabstractRecent encoder-decoder models for sequence-to-sequence mapping show that integrating both temporal and spatial attention mechanisms into neural networks considerably improve network performance. The use of attention for sensor selection in multi-sensor setups and the benefit of such an attention mechanism is less studied. This work reports on a sensor transformation attention network (STAN) that embeds a sensory attention mechanism to dynamically weigh and combine individual input sensors based on their task-relevant information. We demonstrate the correlation of the attentional signal to changing noise levels of each sensor on the audio-visual GRID dataset and synthetic noise; and on CHiME-4, a multi-microphone real-world noisy dataset. In addition, we demonstrate that the STAN model is able to deal with sensor removal and addition without retraining, and is invariant to channel order. Compared to a two-sensor model that weighs both sensors equally, the equivalent STAN model has a relative parameter increase of only 0.09%, but reduces the relative character error rate (CER) by up to 19.1% on the CHiME-4 dataset. The attentional signal helps to identify a lower SNR sensor with up to 94.2% accuracy. Stefan Braun 0005, Daniel Neil, Jithendar Anumula, Enea Ceolini, Shih-Chii Liu |
IJCNN | 5 |
| 2019 | Live Demonstration: Real-Time Spoken Digit Recognition using the DeltaRNN AcceleratorabstractThis demonstration shows a real-time continuous speech recognition hardware system using our previously published DeltaRNN accelerator that enables low latency recurrent neural network (RNN) computation. The network is trained on augmented audio samples from the TIDIGITS dataset to achieve a label error rate (LER) of 2.31%. It is implemented on a Xilinx Zynq-7100 FPGA running at 1 MHz. The incremental RNN power consumption is 30 mW. Visitors interact with the system by speaking digits into a microphone connected to the FPGA system and the classification outputs of the network are continuously displayed on a laptop screen in real time. Chang Gao 0002, Stefan Braun 0005, Ilya Kiselev, Jithendar Anumula, Tobi Delbruck, Shih-Chii Liu |
ISCAS | 6 |
| 2019 | Real-Time Speech Recognition for IoT Purpose using a Delta Recurrent Neural Network AcceleratorabstractThis paper describes a continuous speech recognition hardware system that uses a delta recurrent neural network accelerator (DeltaRNN) implemented on a Xilinx Zynq-7100 FPGA to enable low latency recurrent neural network (RNN) computation. The implemented network consists of a single-layer RNN with 256 gated recurrent unit (GRU) neurons and is driven by input features generated either from the output of a filter bank running on the ARM core of the FPGA in a PmodMic3 microphone setup or from the asynchronous outputs of a spiking silicon cochlea circuit. The microphone setup achieves 7.1 ms minimum latency and 177 frames-per-second (FPS) maximum throughput while the cochlea setup achieves 2.9 ms minimum latency and 345 FPS maximum throughput. The low latency and 70 mW power consumption of the DeltaRNN makes it suitable as an IoT computing platform. Chang Gao 0002, Stefan Braun 0005, Ilya Kiselev, Jithendar Anumula, Tobi Delbruck, Shih-Chii Liu |
ISCAS | 6 |
| 2019 | Incremental Learning Meets Reduced Precision NetworksabstractHardware accelerators for Deep Neural Networks (DNNs) that use reduced precision parameters are more energy efficient than the equivalent full precision networks. While many studies have focused on reduced precision training methods for supervised networks with the availability of large datasets, less work has been reported on incremental learning algorithms that adapt the network for new classes and the consequence of reduced precision has on these algorithms. This paper presents an empirical study of how reduced precision training methods impact the iCARL incremental learning algorithm. The incremental network accuracies on the CIFAR-100 image dataset show that weights can be quantized to 1 bit (2.39% drop in accuracy) but when activations are quantized to 1 bit, the accuracy drops much more (12.75%). Quantizing gradients from 32 to 8 bits only affects the accuracies of the trained network by less than 1%. These results are encouraging for hardware accelerators that support incremental learning algorithms. Yuhuang Hu, Tobi Delbruck, Shih-Chii Liu |
ISCAS | 3 |
| 2019 | Lip Reading Deep Network Exploiting Multi-Modal Spiking Visual and Auditory SensorsabstractThis work presents a lip reading deep neural network that fuses the asynchronous spiking outputs of two bio-inspired silicon multimodal sensors: the Dynamic Vision Sensor (DVS) and the Dynamic Audio Sensor (DAS). The fusion network is tested on the GRID visual-audio lipreading dataset. Classification is carried out using event-based features generated from the spikes of the DVS and DAS. Networks are trained separately on the two modalities and also jointly trained on both modalities. The jointly trained network when tested on DVS spike frames alone, showed a relative increase in accuracy of around 23% over that of the single DVS modality network. Daniel Neil, Tobi Delbruck, Shih-Chii Liu |
ISCAS | 4 |
| 2019 | Live Demonstration: A Portable Microsensor Fusion System with Real-Time Measurement for On-Site Beverage TastingabstractWe demonstrate a portable multisensor fusion system for the automated analysis of multiple beverages. The system makes use of compact and low-power-consumption electronic equipment to simultaneously read out an array of microsensors formed by six ion-selective field-effect transistors (ISFETs), one conductivity sensor, one redox potential sensor, and two ampero-metric microelectrodes. A custom Python application running on a laptop computer receives real-time multivariate data via USB, and provides chemometric models to classify different varieties and to quantify relevant parameters of mineral water and wine. The software also includes a graphical user interface (GUI) to visualize readouts and analytical estimates. Josep Maria Margarit-Taulé, Pablo Giménez-Gómez, Roger Escudé-Pujol, Manuel Gutiérrez-Capitán, Cecilia Jiménez-Jorquera, Shih-Chii Liu |
ISCAS | 6 |
| 2019 | Filtering of Nonuniformly Sampled Bandlimited FunctionsabstractThe filtering of irregularly sampled bandlimited functions is discussed in this letter. An algorithm is given that enables the filtering of bandlimited functions with bandlimited filters to a desired precision, provided some requirements on the sampling density are met. All operations are carried out on the irregular samples themselves. The resulting algorithm is iterative in nature and converges quickly. It is also computationally tractable as it only requires matrix-vector multiplications. The algorithm is noncausal irrespective of the precise characteristics of the filter used, i.e., a buffer is necessary for implementation. The algorithm is evaluated on synthetic examples for which ground truth functions can be derived analytically. Adrian E. G. Huber, Shih-Chii Liu |
IEEE Signal Process. Lett. | 2 |
| 2019 | NullHop: A Flexible Convolutional Neural Network Accelerator Based on Sparse Representations of Feature MapsabstractConvolutional neural networks (CNNs) have become the dominant neural network architecture for solving many state-of-the-art (SOA) visual processing tasks. Even though graphical processing units are most often used in training and deploying CNNs, their power efficiency is less than 10 GOp/s/W for single-frame runtime inference. We propose a flexible and efficient CNN accelerator architecture called NullHop that implements SOA CNNs useful for low-power and low-latency application scenarios. NullHop exploits the sparsity of neuron activations in CNNs to accelerate the computation and reduce memory requirements. The flexible architecture allows high utilization of available computing resources across kernel sizes ranging from 1×1 to 7×7. NullHop can process up to 128 input and 128 output feature maps per layer in a single pass. We implemented the proposed architecture on a Xilinx Zynq field-programmable gate array (FPGA) platform and presented the results showing how our implementation reduces external memory transfers and compute time in five different CNNs ranging from small ones up to the widely known large VGG16 and VGG19 CNNs. Postsynthesis simulations using Mentor Modelsim in a 28-nm process with a clock frequency of 500 MHz show that the VGG19 network achieves over 450 GOp/s. By exploiting sparsity, NullHop achieves an efficiency of 368%, maintains over 98% utilization of the multiply-accumulate units, and achieves a power efficiency of over 3 TOp/s/W in a core area of 6.3 mm2. As further proof of NullHop's usability, we interfaced its FPGA implementation with a neuromorphic event camera for real-time interactive demonstrations. Alessandro Aimar, Hesham Mostafa, Enrico Calabrese, Antonio Rios-Navarro, Ricardo Tapiador-Morales, Iulia-Alexandra Lungu, Moritz B. Milde, Federico Corradi, Alejandro Linares-Barranco, Shih-Chii Liu, Tobi Delbruck |
IEEE Trans. Neural Networks Learn. Syst. | 10 |
| 2018 | DeltaRNN: A Power-efficient Recurrent Neural Network AcceleratorabstractRecurrent Neural Networks (RNNs) are widely used in speech recognition and natural language processing applications because of their capability to process temporal sequences. Because RNNs are fully connected, they require a large number of weight memory accesses, leading to high power consumption. Recent theory has shown that an RNN delta network update approach can reduce memory access and computes with negligible accuracy loss. This paper describes the implementation of this theoretical approach in a hardware accelerator called "DeltaRNN" (DRNN). The DRNN updates the output of a neuron only when the neuron»s activation changes by more than a delta threshold. It was implemented on a Xilinx Zynq-7100 FPGA. FPGA measurement results from a single-layer RNN of 256 Gated Recurrent Unit (GRU) neurons show that the DRNN achieves 1.2 TOp/s effective throughput and 164 GOp/s/W power efficiency. The delta update leads to a 5.7x speedup compared to a conventional RNN update because of the sparsity created by the DN algorithm and the zero-skipping ability of DRNN. Chang Gao 0002, Daniel Neil, Enea Ceolini, Shih-Chii Liu, Tobi Delbruck |
FPGA | 4 |
| 2018 | On Approximation of Bandlimited Functions with Compressed SensingabstractThe application of Compressed Sensing techniques to bandlimited functions is investigated in this paper. It is shown that under the assumption of sparsity, stable reconstruction of a bandlimited function is possible from finitely many samples, contrary to classical results from signal processing theory. The number of measurements that need to be taken is proportional to the sparsity of the function. In compact intervals, it is shown that the number of pointwise measurements required scales quadratically with the size of the largest expansion coefficient (in a basis in which sparsity is measured) which is sufficient for a faithful function approximation. Adrian E. G. Huber, Shih-Chii Liu |
ICASSP | 2 |
| 2018 | Multi-channel Attention for End-to-End Speech RecognitionabstractRecent end-to-end models for automatic speech recognition use sensory attention to integrate multiple input channels within a single neural network.However, these attention models are sensitive to the ordering of the channels used during training.This work proposes a sensory attention mechanism that is invariant to the channel ordering and only increases the overall parameter count by 0.09%.We demonstrate that even without re-training, our attention-equipped end-to-end model is able to deal with arbitrary numbers of input channels during inference.In comparison to a recent related model with sensory attention, our model when tested on the real noisy recordings from the multichannel CHiME-4 dataset, achieves a relative character error rate (CER) improvement of 40.3% to 42.9%.In a two-channel configuration experiment, the attention signal allows the lower signal-to-noise ratio (SNR) sensor to be identified with 97.7% accuracy. Stefan Braun 0005, Daniel Neil, Jithendar Anumula, Enea Ceolini, Shih-Chii Liu |
INTERSPEECH | 5 |
| 2018 | Speaker Activity Detection and Minimum Variance Beamforming for Source Separation
Enea Ceolini, Jithendar Anumula, Adrian E. G. Huber, Ilya Kiselev, Shih-Chii Liu |
INTERSPEECH | 5 |
| 2018 | An event-driven probabilistic model of sound source localization using cochlea spikesabstractThis work presents a probabilistic model that estimates the location of sound sources using the output spikes of a silicon cochlea such as the Dynamic Audio Sensor. Unlike previous work which estimated the source locations directly from the interaural time differences (ITDs) extracted from the timing of the cochlea spikes, the spikes are used instead to support a distribution model of the ITDs representing possible locations of sound sources. Results on noisy single speaker recordings show average accuracies of approximately 80% on detecting the correct source locations and an estimation lag of <;100ms. Jithendar Anumula, Enea Ceolini, Zhe He 0004, Adrian E. G. Huber, Shih-Chii Liu |
ISCAS | 5 |
| 2018 | Conversion of analog to spiking neural networks using sparse temporal codingabstractThe activations of an analog neural network (ANN) are usually treated as representing an analog firing rate. When mapping the ANN onto an equivalent spiking neural network (SNN), this rate-based conversion can lead to undesired increases in computation cost and memory access, if firing rates are high. This work presents an efficient temporal encoding scheme, where the analog activation of a neuron in the ANN is treated as the instantaneous firing rate given by the time-to-first-spike (TTFS) in the converted SNN. By making use of temporal information carried by a single spike, we show a new spiking network model that uses 7-10× fewer operations than the original rate-based analog model on the MNIST handwritten dataset, with an accuracy loss of <; 1%. Bodo Rueckauer, Shih-Chii Liu |
ISCAS | 2 |
| 2017 | Impact of low-precision deep regression networks on single-channel source separationabstractRecent work on developing training methods for reduced precision Deep Convolutional Networks show that these networks can perform with similar accuracy to full precision networks when tested on a classification task. Reduced precision networks decrease the demand on the memory and computational power capabilities of the computing platform. This paper investigates the impact of reduced precision deep Recurrent Neural Networks (RNNs) when trained on a regression task, in this case, a monaural source separation task. The effect of reduced precision nets is explored for two popular recurrent network architectures: Vanilla RNNs and RNNs using Long-Short Term Memory (LSTM) units. The results show that the performance of the networks as measured by blind source separation metrics and speech intelligibility tests on two datasets, show very little decrease even when the weight precision goes down to 4 bits. Enea Ceolini, Shih-Chii Liu |
ICASSP | 2 |
| 2017 | Delta Networks for Optimized Recurrent Network ComputationabstractMany neural networks exhibit stability in their activation patterns over time in response to inputs from sensors operating under real-world conditions. By capitalizing on this property of natural signals, we propose a Recurrent Neural Network (RNN) architecture called a delta network in which each neuron transmits its value only when the change in its activation exceeds a threshold. The execution of RNNs as delta networks is attractive because their states must be stored and fetched at every timestep, unlike in convolutional neural networks (CNNs). We show that a naive run-time delta network implementation offers modest improvements on the number of memory accesses and computes, but optimized training techniques confer higher accuracy at higher speedup. With these optimizations, we demonstrate a 9X reduction in cost with negligible loss of accuracy for the TIDIGITS audio digit recognition benchmark. Similarly, on the large Wall Street Journal (WSJ) speech recognition benchmark, pretrained networks can also be greatly accelerated as delta networks and trained delta networks show a 5.7x improvement with negligible loss of accuracy. Finally, on an end-to-end CNN-RNN network trained for steering angle prediction in a driving dataset, the RNN cost can be reduced by a substantial 100X. Daniel Neil, Junhaeng Lee, Tobi Delbruck, Shih-Chii Liu |
ICML | 4 |
| 2017 | Live demonstration: Event-driven real-time spoken digit recognition systemabstractSummary form only given. We previously described a deep network system that reached an accuracy of 82% on a digit recognition task using the spike outputs from a Dynamic Audio Sensor (DAS) in response to audio samples from the TIDIGITS database. The audio samples were played directly to the system therefore bypassing the microphones. This work presents an interactive real-time demonstration of this digit recognition system. The system classifies a spoken digit based on the output spikes of the DAS in response to digits spoken into the on-board microphones. Jithendar Anumula, Daniel Neil, Tobi Delbruck, Shih-Chii Liu |
ISCAS | 5 |
| 2017 | Guest Editorial Learning in Neuromorphic Systems and Cyborg IntelligenceabstractNeuromorphic computing has become an important emerging research area in recent years. By emulating computational principles and architecture found in neural systems, neuromorphic computing has led to the development of neuromorphic sensors, processors, and sensory motor systems for robotic agents. It has also led to rapid progress in related areas covering computational theories of sensory coding, synaptic computing, learning, and signal processing algorithms, circuit designs, and implementations. The work in these areas shows neuromorphic approaches with appealing computational advantages over conventional approaches, but at the same time, neuromorphic systems still pose many research challenges. Neuromorphic computing overlaps with another area called cyborg intelligence which is dedicated to integrating artificial intelligence (AI) with biological intelligence closely and deeply by connecting computer systems and biological beings. Cyborg intelligence aims to compensate for the weaknesses of both systems by combining the computational power of machines with the perceptive and cognitive abilities of biological systems. Recently, many of the advances in cyborg intelligence methods, systems, and applications have demonstrated the trend of the rapid integration of cyborg intelligence with neuromorphic computing in both breadth and depth. These areas pose innumerable interesting and significant questions for AI and could fundamentally change the landscape of AI research. Zhaohui Wu 0001, Ryad Benosman, Huajin Tang, Shih-Chii Liu |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2016 | Monaural Source Separation Using a Random Forest ClassifierabstractWe address the problem of separating two audio sources from a single channel mixture recording. A novel method called Multi Layered Random Forest (MLRF) that learns a binary mask for both the sources is presented. Random Forest (RF) classifiers are trained for each frequency band of a source spectrogram. A specialized set of linear transformations are applied to a local time-frequency (T-F) neighborhood of the mixture that captures relevant local statistics. A sampling method is presented that efficiently samples T-F training bins in each frequency band. We draw equal numbers of dominant (more power) training samples from the two sources for RF classifiers that estimate the Ideal Binary Mask (IBM). An estimated IBM in a given layer is used to train a RF classifier in the next higher layer of the MLRF hierarchy. On average, MLRF performs better than deep Recurrent Neural Networks (RNNs) and Non-Negative Sparse Coding (NNSC) in signal-to-noise ratio (SNR) of reconstructed audio, overall T-F bin classification accuracy, as well as PESQ and STOI scores. Additionally, we demonstrate the ability of the MLRF to correctly reconstruct T-F bins of the target even when the latter has lower power in that frequency band. Cosimo Riday, Saurabh Bhargava, Richard H. R. Hahnloser, Shih-Chii Liu |
INTERSPEECH | 4 |
| 2016 | Live demonstration: Event-driven deep neural network hardware system for sensor fusionabstractWe demonstrate an interactive digit recognition system using a spiking Deep Neural Network (DNN) FPGA-based system connected to two event-driven sensors: a Dynamic Vision Sensor (DVS) and a Dynamic Audio Sensor (DAS). Sensor fusion is demonstrated on a digit classification task using a DNN trained on the MNIST dataset supplemented by assignment of a unique pure ton e for each digit. Ilya Kiselev, Daniel Neil, Shih-Chii Liu |
ISCAS | 3 |
| 2016 | Event-driven deep neural network hardware system for sensor fusionabstractThis paper presents a real-time multi-modal spiking Deep Neural Network (DNN) implemented on an FPGA platform. The hardware DNN system, called n-Minitaur, demonstrates a 4-fold improvement in computational speed over the previous DNN FPGA system. The proposed system directly interfaces two different event-based sensors: a Dynamic Vision Sensor (DVS) and a Dynamic Audio Sensor (DAS). The DNN for this bimodal hardware system is trained on the MNIST digit dataset and a set of unique audio tones for each digit. When tested on the spikes produced by each sensor alone, the classification accuracy is around 70% for DVS spikes generated in response to displayed MNIST images, and 60% for DAS spikes generated in response to noisy tones. The accuracy increases to 98% when spikes from both modalities are provided simultaneously. In addition, the system shows a fast latency response of only 5ms. Ilya Kiselev, Daniel Neil, Shih-Chii Liu |
ISCAS | 3 |
| 2016 | Combined frame- and event-based detection and trackingabstractThis paper reports an object tracking algorithm for a moving platform using the dynamic and active-pixel vision sensor (DAVIS). It takes advantage of both the active pixel sensor (APS) frame and dynamic vision sensor (DVS) event outputs from the DAVIS. The tracking is performed in a three step-manner: regions of interest (ROIs) are generated by a cluster-based tracking using the DVS output, likely target locations are detected by using a convolutional neural network (CNN) on the APS output to classify the ROIs as foreground and background, and finally a particle filter infers the target location from the ROIs. Doing convolution only in the ROIs boosts the speed by a factor of 70 compared with full-frame convolutions for the 240×180 frame input from the DAVIS. The tracking accuracy on a predator and prey robot database reaches 90% with a cost of less than 20ms/frame in Matlab on a normal PC without using a GPU. Diederik Paul Moeys, Gautham P. Das, Daniel Neil, Shih-Chii Liu, Tobi Delbruck |
ISCAS | 5 |
| 2016 | Effective sensor fusion with event-based sensors and deep network architecturesabstractThe use of spiking neuromorphic sensors with state-of-art deep networks is currently an active area of research. Still relatively unexplored are the pre-processing steps needed to transform spikes from these sensors and the types of network architectures that can produce high-accuracy performance using these sensors. This paper discusses several methods for preprocessing the spiking data from these sensors for use with various deep network architectures. The outputs of these preprocessing methods are evaluated using different networks including a deep fusion network composed of Convolutional Neural Networks and Recurrent Neural Networks, to jointly solve a recognition task using the MNIST (visual) and TIDIGITS (audio) benchmark datasets. With only 1000 visual input spikes from a spiking hardware retina, the classification accuracy of 64.5% achieved by a particular trained fusion network increases to 98.31% when combined with inputs from a spiking hardware cochlea. Daniel Neil, Shih-Chii Liu |
ISCAS | 2 |
| 2016 | Phased LSTM: Accelerating Recurrent Network Training for Long or Event-based SequencesabstractRecurrent Neural Networks (RNNs) have become the state-of-the-art choice for extracting patterns from temporal sequences. Current RNN models are ill suited to process irregularly sampled data triggered by events generated in continuous time by sensors or other neurons. Such data can occur, for example, when the input comes from novel event-driven artificial sensors which generate sparse, asynchronous streams of events or from multiple conventional sensors with different update intervals. In this work, we introduce the Phased LSTM model, which extends the LSTM unit by adding a new time gate. This gate is controlled by a parametrized oscillation with a frequency range which require updates of the memory cell only during a small percentage of the cycle. Even with the sparse updates imposed by the oscillation, the Phased LSTM network achieves faster convergence than regular LSTMs on tasks which require learning of long sequences. The model naturally integrates inputs from sensors of arbitrary sampling rates, thereby opening new areas of investigation for processing asynchronous sensory events that carry timing information. It also greatly improves the performance of LSTMs in standard RNN applications, and does so with an order-of-magnitude fewer computes. Daniel Neil, Michael Pfeiffer 0001, Shih-Chii Liu |
NIPS | 3 |
| 2015 | Fast-classifying, high-accuracy spiking deep networks through weight and threshold balancingabstractDeep neural networks such as Convolutional Networks (ConvNets) and Deep Belief Networks (DBNs) represent the state-of-the-art for many machine learning and computer vision classification problems. To overcome the large computational cost of deep networks, spiking deep networks have recently been proposed, given the specialized hardware now available for spiking neural networks (SNNs). However, this has come at the cost of performance losses due to the conversion from analog neural networks (ANNs) without a notion of time, to sparsely firing, event-driven SNNs. Here we analyze the effects of converting deep ANNs into SNNs with respect to the choice of parameters for spiking neurons such as firing rates and thresholds. We present a set of optimization techniques to minimize performance loss in the conversion process for ConvNets and fully connected deep networks. These techniques yield networks that outperform all previous SNNs on the MNIST database to date, and many networks here are close to maximum performance after only 20 ms of simulated time. The techniques include using rectified linear units (ReLUs) with zero bias during training, and using a new weight normalization method to help regulate firing rates. Our method for converting an ANN into an SNN enables low-latency classification with high accuracies already after the first output spike, and compared with previous SNN approaches it yields improved performance without increased training time. The presented analysis and optimization techniques boost the value of spiking deep networks as an attractive framework for neuromorphic computing platforms aimed at fast and efficient pattern recognition. Peter U. Diehl, Daniel Neil, Jonathan Binas, Matthew Cook 0001, Shih-Chii Liu, Michael Pfeiffer 0001 |
IJCNN | 5 |
| 2015 | Scalable energy-efficient, low-latency implementations of trained spiking Deep Belief Networks on SpiNNakerabstractDeep neural networks have become the state-of-the-art approach for classification in machine learning, and Deep Belief Networks (DBNs) are one of its most successful representatives. DBNs consist of many neuron-like units, which are connected only to neurons in neighboring layers. Larger DBNs have been shown to perform better, but scaling-up poses problems for conventional CPUs, which calls for efficient implementations on parallel computing architectures, in particular reducing the communication overhead. In this context we introduce a realization of a spike-based variation of previously trained DBNs on the biologically-inspired parallel SpiNNaker platform. The DBN on SpiNNaker runs in real-time and achieves a classification performance of 95% on the MNIST handwritten digit dataset, which is only 0.06% less than that of a pure software implementation. Importantly, using a neurally-inspired architecture yields additional benefits: during network run-time on this task, the platform consumes only 0.3 W with classification latencies in the order of tens of milliseconds, making it suitable for implementing such networks on a mobile platform. The results in this paper also show how the power dissipation of the SpiNNaker platform and the classification latency of a network scales with the number of neurons and layers in the network and the overall spike activity rate. Evangelos Stromatias, Daniel Neil, Francesco Galluppi, Michael Pfeiffer 0001, Shih-Chii Liu, Steve Furber |
IJCNN | 5 |
| 2015 | Scene stitching with event-driven sensors on a robot head platformabstractThis paper describes a robot head platform which holds a pair of event-based Dynamic Vision Sensor (DVS) retinas and microphones connected to an event-based binaural AEREAR2 VLSI cochlea system. The platform has 6 degrees of freedom (DOF): 2 for the neck, and 2 for each of the DVS retinas. Two applications using this platform are described: the first is image stitching of a scene larger than the field of view of the individual retinas as the head pans and tilts and the second is selective image painting of the local visual scene around spatially displaced sound sources. This platform allows for the investigation of event-driven sensory-action models that use the information from multiple event-based sensor modalities in real-time scenarios. Philipp Klein, Jörg Conradt, Shih-Chii Liu |
ISCAS | 3 |
| 2015 | Design of an RGBW color VGA rolling and global shutter dynamic and active-pixel vision sensorabstractThis paper reports the design of a color dynamic and active-pixel vision sensor (C-DAVIS) for robotic vision applications. The C-DAVIS combines monochrome eventgenerating dynamic vision sensor pixels and 5-transistor active pixels sensor (APS) pixels patterned with an RGBW color filter array. The C-DAVIS concurrently outputs rolling or global shutter RGBW coded VGA resolution frames and asynchronous monochrome QVGA resolution temporal contrast events. Hence the C-DAVIS is able to capture spatial details with color and track movements with high temporal resolution while keeping the data output sparse and fast. The C-DAVIS chip is fabricated in TowerJazz 0.18um CMOS image sensor technology. An RGBW 2×2-pixel unit measures 20um × 20um. The chip die measures 8mm × 6.2mm. Cheng-Han Li, Christian Brandli, Raphael Berner, Minhao Yang, Shih-Chii Liu, Tobi Delbruck |
ISCAS | 6 |
| 2015 | Design of a spatiotemporal correlation filter for event-based sensorsabstractThis paper reports the design of a 1mW, 10ns-latency mixed signal system in 0.18μm CMOS which enables filtering out uncorrelated background activity in event-based neuromorphic sensors. Background activity (BA) in the output of dynamic vision sensors is caused by thermal noise and junction leakage current acting on switches connected to floating nodes in the pixels. The reported chip generates a pass flag for spatiotemporally correlated events for post-processing to reduce communication/computation load and improve information rate. A chip with 128×128 array with 20×20μm2cells has been designed. Each filter cell combines programmable spatial subsampling with a temporal window based on current integration. Power-gating is used to minimize the power consumption by only activating the threshold detection and communication circuits in the cell receiving an input event. This correlation filter chip targets embedded neuromorphic visual and auditory systems, where low average power consumption and low latency are critical. Christian Brandli, Cheng-Han Li, Shih-Chii Liu, Tobi Delbruck |
ISCAS | 4 |
| 2015 | Current-mode automated quality control cochlear resonator for bird identity taggingabstractThis paper describes a VLSI automatic quality control pitch detector circuit which can be used for detecting the identity of a unique bird. The detector is based on a previous VLSI model of the local gain control mechanism of the outer hair cells of the biological cochlea. This work presents characterization results from a 20-channel chip fabricated in a 4-metal 2-poly CMOS 0.35 μm technology with estimated dynamic range of 70 dB, power consumption of 825 nW per channel, frequency range covering 0.4-10 kHz and mean Q of 6.31. Results are shown for a pitch detection experiment with a tuned resonator. Diederik Paul Moeys, Tobi Delbruck, Shih-Chii Liu |
ISCAS | 3 |
| 2015 | Live demonstration: Handwritten digit recognition using spiking deep belief networks on SpiNNakerabstractWe demonstrate an interactive handwritten digit recognition system with a spike-based deep belief network running in real-time on SpiNNaker, a biologically inspired many-core architecture. Results show that during the simulation a SpiNNaker chip can deliver spikes in under 1 μs, with a classification latency in the order of tens of milliseconds, while consuming less than 0.3 W. Evangelos Stromatias, Daniel Neil, Francesco Galluppi, Michael Pfeiffer 0001, Shih-Chii Liu, Steve Furber |
ISCAS | 5 |
| 2015 | Linear Methods for Efficient and Fast Separation of Two Sources Recorded with a Single MicrophoneabstractThis letter addresses the problem of separating two speakers from a single microphone recording. Three linear methods are tested for source separation, all of which operate directly on sound spectrograms: (1) eigenmode analysis of covariance difference to identify spectro-temporal features associated with large variance for one source and small variance for the other source; (2) maximum likelihood demixing in which the mixture is modeled as the sum of two gaussian signals and maximum likelihood is used to identify the most likely sources; and (3) suppression-regression, in which autoregressive models are trained to reproduce one source and suppress the other. These linear approaches are tested on the problem of separating a known male from a known female speaker. The performance of these algorithms is assessed in terms of the residual error of estimated source spectrograms, waveform signal-to-noise ratio, and perceptual evaluation of speech quality scores. This work shows that the algorithms compare favorably to nonlinear approaches such as nonnegative sparse coding in terms of simplicity, performance, and suitability for real-time implementations, and they provide benchmark solutions for monaural source separation tasks. Saurabh Bhargava, Florian Blättler, Sepp Kollmorgen, Shih-Chii Liu, Richard H. R. Hahnloser |
Neural Comput. | 4 |
| 2015 | Hardware-Amenable Structural Learning for Spike-Based Pattern Classification Using a Simple Model of Active DendritesabstractThis letter presents a spike-based model that employs neurons with functionally distinct dendritic compartments for classifying high-dimensional binary patterns. The synaptic inputs arriving on each dendritic subunit are nonlinearly processed before being linearly integrated at the soma, giving the neuron the capacity to perform a large number of input-output mappings. The model uses sparse synaptic connectivity, where each synapse takes a binary value. The optimal connection pattern of a neuron is learned by using a simple hardware-friendly, margin-enhancing learning algorithm inspired by the mechanism of structural plasticity in biological neurons. The learning algorithm groups correlated synaptic inputs on the same dendritic branch. Since the learning results in modified connection patterns, it can be incorporated into current event-based neuromorphic systems with little overhead. This work also presents a branch-specific spike-based version of this structural plasticity rule. The proposed model is evaluated on benchmark binary classification problems, and its performance is compared against that achieved using support vector machine and extreme learning machine techniques. Our proposed method attains comparable performance while using 10% to 50% less in computational resource than the other reported techniques. Shaista Hussain, Shih-Chii Liu, Arindam Basu |
Neural Comput. | 2 |
| 2015 | Memory and Information Processing in Neuromorphic SystemsabstractA striking difference between brain-inspired neuromorphic processors and current von Neumann processor architectures is the way in which memory and processing is organized. As information and communication technologies continue to address the need for increased computational power through the increase of cores within a digital processor, neuromorphic engineers and scientists can complement this need by building processor architectures where memory is distributed with the processing. In this paper, we present a survey of brain-inspired processor architectures that support models of cortical networks and deep neural networks. These architectures range from serial clocked implementations of multineuron systems to massively parallel asynchronous ones and from purely digital systems to mixed analog/digital systems which implement more biological-like models of neurons and synapses together with a suite of adaptation and learning mechanisms analogous to the ones found in biological nervous systems. We describe the advantages of the different approaches being pursued and present the challenges that need to be addressed for building artificial neural processing systems that can display the richness of behaviors seen in biological systems. Giacomo Indiveri, Shih-Chii Liu |
Proc. IEEE | 2 |
| 2014 | Live demonstration: The "DAVIS" Dynamic and Active-Pixel Vision SensorabstractThis demonstration will show the features of the Dynamic and Active-Pixel Vision Sensor (DAVIS) reported at the VLSI Symposium and the International Imager Sensor Workshop in 2013. This sensor concurrently outputs conventional CMOS image sensor frames and sparse, low-latency dynamic vision sensor events from the same pixels, sharing the same photodiodes. The setup will allow visitors to explore the advantages of combining of fast and computationally-efficient neuromorphic event-driven vision with the existing body of methods for frame-based computer and machine vision. Christian Brandli, Raphael Berner, Minhao Yang, Shih-Chii Liu, V. Villeneuva, Tobi Delbruck |
ISCAS | 4 |
| 2014 | Improved margin multi-class classification using dendritic neurons with morphological learningabstractWe present an architecture of a spike based multiclass classifier using neurons with non-linear dendrites and sparse synaptic connectivity where each synapse takes a binary value. The learning in this model happens not through weight updates but through structural changes, i.e. a change of connectivity between inputs and dendrites. Hence, it is well suited for implementation in neuromorphic systems using address event representation (AER). We present a new learning rule that allows better generalization of the system to noisy testing data making it feasible to transfer learnt weights in software to a hardware device interfacing with noisy spiking sensors. The new rule improves testing accuracy by 7 - 10% compared to earlier versions. We also present preliminary results for multi-class classification on handwritten digits from the MNIST database and show that our system can attain comparable performance (≈ 3% more error) with other reported spike based classifiers while using at least 50% less synaptic resources. Shaista Hussain, Shih-Chii Liu, Arindam Basu |
ISCAS | 2 |
| 2014 | 1kHz 2D silicon retina motion sensor platformabstractThis paper proposes an optical motion sensor aimed towards small robotic platforms. It incorporates a 20×20 pixel continuous-time CMOS silicon retina vision sensor with pixels that have local gain control and adapt to background lighting and a DSP microcontroller which computes the global optical flow from the sampled sensor output. The system allows the user to validate various motion algorithms suitable for the platform. Measurements are presented that show that the system can compute global 2D translational motion from complex natural scenes using the image interpolation algorithm at a sample rate of 1 kHz and for speeds up to ±1000 pixels/s using <5k instruction cycles per frame. Andreas Steiner 0001, Rico Moeckel, Reto Thurer, Dario Floreano, Tobi Delbruck, Shih-Chii Liu |
ISCAS | 6 |
| 2014 | Comparison of spike encoding schemes in asynchronous vision sensors: Modeling and designabstractTwo in-pixel encoding mechanisms to convert analog input to spike output for vision sensors are modeled and compared with the consideration of feedback delay: one is feedback and reset (FAR), and the other is feedback and subtract (FAS). MATLAB simulations of linear signal reconstruction from spike trains generated by the two encoders show that FAR in general has a lower signal-to-distortion ratio (SDR) compared to FAS due to signal loss during the reset phase and hold period, and the SDR merit of FAS increases as the quantization bit number and input signal frequency increases. A 500 μm2in-pixel circuit implementation of FAS using asynchronous switched capacitors in a UMC 0.18μm 1P6M process is described, and the post-layout simulation results are given to verify the FAS encoding mechanism. Minhao Yang, Shih-Chii Liu, Tobi Delbruck |
ISCAS | 2 |
| 2014 | Minitaur, an Event-Driven FPGA-Based Spiking Network AcceleratorabstractCurrent neural networks are accumulating accolades for their performance on a variety of real-world computational tasks including recognition, classification, regression, and prediction, yet there are few scalable architectures that have emerged to address the challenges posed by their computation. This paper introduces Minitaur, an event-driven neural network accelerator, which is designed for low power and high performance. As an field-programmable gate array-based system, it can be integrated into existing robotics or it can offload computationally expensive neural network tasks from the CPU. The version presented here implements a spiking deep network which achieves 19 million postsynaptic currents per second on 1.5 W of power and supports up to 65 K neurons per board. The system records 92% accuracy on the MNIST handwritten digit classification and 71% accuracy on the 20 newsgroups classification data set. Due to its event-driven nature, it allows for trading off between accuracy and latency. Daniel Neil, Shih-Chii Liu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2013 | Morphological learning: Increased memory capacity of neuromorphic systems with binary synapses exploiting AER based reconfigurationabstractSpiking neurons with lumped nonlinearity representing active dendrites can perform a larger number of input-output mappings than is possible by a neuron with linear synaptic summation of its currents. This is possible due to the additional degree of freedom in such cells-its `morphology' reflected in the number of dendrites and the choice of which inputs form synapses on the same dendrite. We present a hardware friendly algorithm for learning such optimal morphologies utilizing correlations between inputs and dendritic branch activations. We demonstrate the increased memory capacity of neurons with nonlinear dendrites and binary synapses over typically used linearly summing cells with high resolution weights. We have shown that a neuron model with a fixed number of binary weights performs much worse on a pattern classification task when it uses traditional linear dendrites than when it utilizes nonlinear dendrites (19% compared to 9% errors for 1000 patterns). This method allows to trade-off weight resolution, a problem in most current neuromorphic systems, with configurability that is the strength of address event representation (AER) based systems which can store configuration details in an off-chip memory. On a fundamental level, it points to the need of having a higher ratio of nonlinear to linear operations in spiking neural networks than is typically used. Shaista Hussain, Roshan Gopalakrishnan, Arindam Basu, Shih-Chii Liu |
IJCNN | 4 |
| 2012 | A model of attention-driven scene analysisabstractParsing complex acoustic scenes involves an intricate interplay between bottom-up, stimulus-driven salient elements in the scene with top-down, goal-directed, mechanisms that shift our attention to particular parts of the scene. Here, we present a framework for exploring the interaction between these two processes in a simulated cocktail party setting. The model shows improved digit recognition in a multi-talker environment with a goal of tracking the source uttering the highest value. This work highlights the relevance of both data-driven and goal-driven processes in tackling real multi-talker, multi-source sound analysis. Malcolm Slaney, Trevor Agus, Shih-Chii Liu, Emine Merve Kaya, Mounya Elhilali |
ICASSP | 3 |
| 2012 | Real-time speaker identification using the AEREAR2 event-based silicon cochleaabstractThis paper reports a study on methods for real-time speaker identification using the output from an event-based silicon cochlea. These methods are evaluated based on the amount of computation that needs to be performed and the classification performance in a speaker identification task. It uses the binaural AEREAR2 silicon cochlea, with 64 frequency channels and 512 output neurons. Auditory features representing fading histograms of inter-spike intervals and channel activity distributions are extracted from the cochlea spikes. These feature vectors are then classified by a linear Support Vector Machine, which is trained against a subset of 40 speakers (20/20 male/female) from the TIMIT database. Speakers are correctly identified at >90% accuracy during each sentence utterance and with an average latency of 700±200ms from the start of the sentence. Cheng-Han Li, Tobi Delbruck, Shih-Chii Liu |
ISCAS | 3 |
| 2012 | Addressable current reference array with 170dB dynamic rangeabstractConfigurable high-performance bias current reference circuits are useful in complex mixed-signal chips. This paper presents the design of a configurable current reference array with ultra wide dynamic range (DR). A coarse-fine architecture using octal coarse current spacing and 8 bits of fine resolution increases the overall current DR with less area compared with the prior work. Compact current multipliers and dividers also save chip areas. Shifted-source current mirrors and an off-current suppression technique improve the accuracy of generated low currents. A buffer with dual-threshold source followers is used to generate the output biasing voltage with a wide DR input current. Biases are individually addressable and configurable. Measurement results of this design in UMC 0.18μm 1P6M CMOS process suggest that over 170dB DR is achieved at room temperature. Each additional bias occupies an incremental area of 360×22μm2, which is smaller by a factor of 4 compared to the previous design. Minhao Yang, Shih-Chii Liu, Cheng-Han Li, Tobi Delbruck |
ISCAS | 2 |
| 2011 | Estimating the location of a sound source with a spike-timing localization algorithmabstractA binaural spike-based sound localization system suitable for real-time attention systems of robots is presented in this work. This system uses the output spikes from a 64-channel binaural silicon cochlea. The localization algorithm implements a spike-based correlation of these output spikes that measures the interaural time differences (ITDs) between the arrival of a sound to two microphones. The algorithm continuously updates a possible distribution of ITDs in the auditory scene whenever a spike arrives from either ear of the cochlea. Experimental results show that the system can estimate an azimuth angle to below 1 deg resolution. This algorithm is computationally cheaper than conventional generalized cross-correlation methods and is quicker to respond to changes in the scene. Holger Finger, Shih-Chii Liu |
ISCAS | 2 |
| 2011 | Mismatch reduction through dendritic nonlinearities in a 2D silicon dendritic neuron arrayabstractThis paper describes a novel 2D programmable dendritic neuron array consisting of a 3×32 dendritic compartment array and a 1×32 somatic compartment array. Each dendritic compartment contains two types of regenerative nonlinearities: an NMDA nonlinearity and a dendritic spike nonlinearity. The chip supports the programmability of local synaptic weights and the configuration of dendritic morphology for individual neurons through the address-event representation protocol. With a novel local cable circuit between neighboring compartments, different dendritic morphologies can be constructed. From results measured on a chip fabricated in a 4-metal, 2-poly, 0.35μm CMOS technology, we show one instance of how dendritic non- linearities can contribute to neuronal computation: the dendritic spike mechanism dynamically reduces the mismatch-induced coefficient of variation of the somatic response amplitude from approximately 40% to 3.5%. Shih-Chii Liu |
ISCAS | 2 |
| 2010 | Approaches and databases for online calibration of binaural sound localization for robotic headsabstractIn this paper, we evaluate adaptive sound localization algorithms for robotic heads. To this end we built a 3 degree-of-freedom head with two microphones encased in artificial pinnae (outer ears). The geometry of the head and pinnae induce temporal differences in the sound recorded at each microphone. These differences change with the frequency of the sound, location of the sound, and orientation of the robot in a complex manner. To learn the relationship between these auditory differences and the location of a sound source, we applied machine learning methods to a database of different audio source locations and robot head orientations. Our approach achieves a mean error of 2.5 degrees for azimuth and 11 degrees for elevation for estimating the position of an audio source. The impressive results highlight the benefits of a two-stage regression model to make use of the properties of the artificial pinnae for elevation estimation. In this work, the algorithms were trained using ground truth data provided by a motion capture system. We are currently generalizing the approach so that the training signal is provided online based on a real-time face detection and speech detection system. Holger Finger, Shih-Chii Liu, Paul Ruvolo, Javier R. Movellan |
IROS | 2 |
| 2010 | Exploiting spike-based dynamics in a silicon cochlea for speaker identificationabstractLimit-cycle dynamics embedded in neuronal spike-trains can form robust representations for encoding auditory spectral features. In this paper, we present speaker identification experiments based on limit-cycle statistics that were computed using spike-trains obtained from a spike-based silicon cochlea. The features included in this study were: (a) spike-rate; (b) inter-spike-interval distribution; and (c) inter-spike-velocity features, which were then used to design a speaker identification system based on a Gini-support vector machine (SVM) classifier. The results show a strong correlation between the information contained in the spike-rate/interval features and the spike-velocity/acceleration features indicating redundant encoding of auditory features which could be important for achieving noise-robustness in real-world recording conditions. Shantanu Chakrabartty, Shih-Chii Liu |
ISCAS | 2 |
| 2010 | The use of spike-based representations for hardware audition systemsabstractHumans are able to process speech and other sounds effectively in adverse environments, hearing through noise, reverberation, and interference from other speakers. To date, machines have been unable to match human performance. One profound difference between biological and engineering systems comes at the input stage. In machines, an acoustic signal is typically chopped into short equally spaced segments in time. In biological systems, the cochlea outputs asynchronous spikes that respond in real-time to acoustic inputs. In this paper we describe a spiking cochlea implementation and recent experiments in both speaker and speech recognition that use spikes as input. Shih-Chii Liu, Nima Mesgarani, John G. Harris, Hynek Hermansky |
ISCAS | 1 |
| 2010 | Event-based 64-channel binaural silicon cochlea with Q enhancement mechanismsabstractThis paper describes an event-based binaural silicon cochlea aimed at spatial audition and auditory scene analysis. The chip has a matched pair of 64-stage cascaded analog second-order filter banks with 512 pulse-frequency modulated (PFM) address-event representation (AER) outputs. The spectral selectivity is sharpened through 2 different on-chip methods: an on-chip local Q DAC and an on-chip spatial sharpening through nearest neighbour lateral inhibition. The fabricated chip in a 4-metal 2-poly 0.35um CMOS process consumes peak 25mW power for the digital circuits and 33mW for the analog core. Dynamic range to produce PFM output is 36dB (25mVpp to 1500mVpp at microphone preamp output). Event timing jitter is 2us for 250mVpp input. The peak output bandwidth is 10M events per second (eps) but typical speech scenarios show rates of 20keps. Shih-Chii Liu, André van Schaik, Bradley A. Minch, Tobi Delbruck |
ISCAS | 1 |
| 2010 | Motion detection using an aVLSI network of spiking neuronsabstractWe explore the benefits of using a spiking neuronal network to generate direction-selectivity. The experiments are performed using an aVLSI spike-based network configured for two possible cortical architectures. We find that the direction-selective properties are influenced by the various connections in the recurrent network and that the network is sensitive to over almost 3 decades of stimulus speeds. Shih-Chii Liu |
ISCAS | 2 |
| 2010 | Multilayer Processing of Spatiotemporal Spike Patterns in a Neuron with Active DendritesabstractWith the advent of new experimental evidence showing that dendrites play an active role in processing a neuron's inputs, we revisit the question of a suitable abstraction for the computing function of a neuron in processing spatiotemporal input patterns. Although the integrative role of a neuron in relation to the spatial clustering of synaptic inputs can be described by a two-layer neural network, no corresponding abstraction has yet been described for how a neuron processes temporal input patterns on the dendrites. We address this void using a real-time aVLSI (analog very-large-scale-integrated) dendritic compartmental model, which incorporates two widely studied classes of regenerative event mechanisms: one is mediated by voltage-gated ion channels and the other by transmitter-gated NMDA channels. From this model, we find that the response of a dendritic compartment can be described as a nonlinear sigmoidal function of both the degree of input temporal synchrony and the synaptic input spatial clustering. We propose that a neuron with active dendrites can be modeled as a multilayer network that selectively amplifies responses to relevant spatiotemporal input spike patterns. Shih-Chii Liu |
Neural Comput. | 2 |
| 2009 | Implementation of a Time-warping AER MapperabstractIn recent implementations of neuromorphic spike-based sensors, multi-neuron processors, and actuators; the spike traffic between devices is coded in the form of asynchronous spike streams following the address-event-representation protocol. This spike information can be modified during the transmission from one device to another by using a mapper device. In this paper we present a mapper implementation which transforms event addresses and can also delay events in time. We discuss two different architectures for implementing the time delays on an FPGA board (USB-AER), and we present an example of the use of the time delay feature in the mapper in an implementation of a visual elementary motion detection model based on the spike outputs of a temporal contrast retina. Alejandro Linares-Barranco, Francisco Gomez-Rodriguez, Gabriel Jiménez-Moreno, Tobi Delbruck, Raphael Berner, Shih-Chii Liu |
ISCAS | 6 |
| 2009 | Input Evoked Nonlinearities in Silicon Dendritic CircuitsabstractMost VLSI spiking network implementations are constructed using point neurons. However, neurons with extended dendritic structures might offer additional computational advantages. Experimental evidence suggests that dendritic compartments could be considered as independent and parallel computational units. Depending on the synaptic input patterns, the dendritic integration could be either linear or nonlinear. We show the influence of spatio-temporal input patterns on the evoked dendritic integration in an aVLSI neuron chip with programmable dendritic compartments. Shih-Chii Liu |
ISCAS | 2 |
| 2009 | Periodicity Detection and Localization using Spike Timing from the AER EARabstractWe present a system consisting of a spiking cochlea chip and real-time event-based processing software that is able to discriminate between two sets of sounds based on their periodicity content. The periodicity measurements are computed from the spike timing information of asynchronous output spikes from the binaural spiking-cochlea chip. The chip consists of a matched pair of silicon cochlea with an address event interface for the output. Each section of the cochlea is modeled by a second-order low-pass filter followed by a simplified Inner Hair Cell circuit and a Spiking Neuron circuit. We show discrimination results using the periodicity measure for 2 classes of sound and preliminary localization results based on a discriminated sound. Theodore Yu, John G. Harris, Malcolm Slaney, Shih-Chii Liu |
ISCAS | 5 |
| 2009 | Computation with Spikes in a Winner-Take-All NetworkabstractThe winner-take-all (WTA) computation in networks of recurrently connected neurons is an important decision element of many models of cortical processing. However, analytical studies of the WTA performance in recurrent networks have generally addressed rate-based models. Very few have addressed networks of spiking neurons, which are relevant for understanding the biological networks themselves and also for the development of neuromorphic electronic neurons that commmunicate by action potential like address-events. Here, we make steps in that direction by using a simplified Markov model of the spiking network to examine analytically the ability of a spike-based WTA network to discriminate the statistics of inputs ranging from stationary regular to nonstationary Poisson events. Our work extends previous theoretical results showing that a WTA recurrent network receiving regular spike inputs can select the correct winner within one interspike interval. We show first for the case of spike rate inputs that input discrimination and the effects of self-excitation and inhibition on this discrimination are consistent with results obtained from the standard rate-based WTA models. We also extend this discrimination analysis of spiking WTAs to nonstationary inputs with time-varying spike rates resembling statistics of real-world sensory stimuli. We conclude that spiking WTAs are consistent with their continuous counterparts for steady-state inputs, but they also exhibit high discrimination performance with nonstationary inputs. Matthias Oster, Rodney J. Douglas, Shih-Chii Liu |
Neural Comput. | 3 |
| 2009 | CAVIAR: A 45k Neuron, 5M Synapse, 12G Connects/s AER Hardware Sensory-Processing- Learning-Actuating System for High-Speed Visual Object Recognition and TrackingabstractThis paper describes CAVIAR, a massively parallel hardware implementation of a spike-based sensing-processing-learning-actuating system inspired by the physiology of the nervous system. CAVIAR uses the asychronous address-event representation (AER) communication framework and was developed in the context of a European Union funded project. It has four custom mixed-signal AER chips, five custom digital AER interface components, 45k neurons (spiking cells), up to 5M synapses, performs 12G synaptic operations per second, and achieves millisecond object recognition and tracking latencies. Rafael Serrano-Gotarredona, Matthias Oster, Patrick Lichtsteiner, Alejandro Linares-Barranco, Rafael Paz-Vicente, Francisco Gomez-Rodriguez, Luis A. Camuñas-Mesa, Raphael Berner, Manuel Rivas Pérez, Tobi Delbruck, Shih-Chii Liu, Rodney J. Douglas, Philipp Häfliger, Gabriel Jiménez-Moreno, Antonio Abad Civit Balcells, Teresa Serrano-Gotarredona, Antonio J. Acosta 0001, Bernabé Linares-Barranco |
IEEE Trans. Neural Networks | 11 |
| 2008 | Temporally learning floating-gate VLSI synapsesabstractWe present a floating-gate synaptic circuit that updates its weight according to the Spike-Timing-Dependent Plasticity (STDP) rule. The weight (or floating-gate voltage) is updated only if the time difference between the pre- and post-synaptic spikes falls within a learning window. The update is implemented through tunneling and injection mechanisms which can be tuned for very long time constants up to seconds. The novelty of this circuit is that the tunneling and injection mechanisms are turned on only when the correlation of the pre and postsynaptic activity is significant. The additional benefit of this non-volatile technology is that synaptic weights can be stored locally on chip. We present experimental results that show the learning and normalization effects from the fabricated circuits. Shih-Chii Liu, Rico Moeckel |
ISCAS | 1 |
| 2008 | Steering with an aVLSI motion detection chipabstractWe demonstrate the capabilies of a 1D motion detection chip in steering a car in a simulated racing game. The chip implemented in a 1.5 μm CMOS VLSI process, is comprised of 24 motion pixels that extracts optical flow in parallel. The local optical flow values go to a single-layer perceptron whose output is used to steer the car in the racing game. Because of the continuous-time operation of the motion detection chip, the computationally expensive task of generating a control signal for the car based on the visual scene is largely alleviated. Rico Moeckel, Roger Jaeggi, Shih-Chii Liu |
ISCAS | 3 |
| 2007 | AER Auditory Filtering and CPG for Robot ControlabstractAddress-event-representation (AER) is a communication protocol for transferring asynchronous events between VLSI chips, originally developed for bio-inspired processing systems (for example, image processing). The event information in an AER system is transferred using a high-speed digital parallel bus. This paper presents an experiment using AER for sensing, processing and finally actuating a robot. The AER output of a silicon cochlea is processed by an AER filter implemented on a FPGA to produce rhythmic walking in a humanoid robot (Redbot). We have implemented both the AER rhythm detector and the central pattern generator (CPG) on a Spartan II FPGA which is part of a USB-AER platform developed by some of the authors Francisco Gomez-Rodriguez, Alejandro Linares-Barranco, Lourdes Miro-Amarante, Shih-Chii Liu, André van Schaik, Ralph Etienne-Cummings, M. Anthony Lewis |
ISCAS | 4 |
| 2007 | Motion Detection Circuits for a Time-To-Travel AlgorithmabstractWe describe a new motion detection circuit that extracts motion information based on a time-to-travel algorithm. The front-end photoreceptor adapts over 7 decades of background intensity and motion information can be extracted down to a contrast value of 2.5%. Results from the circuits which were fabricated in a 2-metal 2-poly 1.5um CMOS process, show that the motion information can be extracted over 2 decades of speed. Rico Moeckel, Shih-Chii Liu |
ISCAS | 2 |
| 2007 | Quantifying Input and Output Spike Statistics of a Winner-Take-All Network in a Vision SystemabstractEvent-driven spike-based processing systems offer new possibilities for real-time vision. Signals are encoded asynchronously in time thus preserving the time information of the occurrence of an event. The paper examines this form of coding using experimental data from a multi-layered multi-chip system which consists of an artificial retina, a convolution filter bank and a winner-take-all network which detect the position of a moving object. The spike outputs of the convolution stage can be described by an inhomogeneous Poisson distribution of Gaussian profile, although the underlying building blocks are completely deterministic and exhibit only a small amount of variation. The authors discuss a method for measuring the accuracy of the asynchronous spiking representation in both time and value, thereby quantifying the performance of the winner-take-all network in determining the position of a ball rotating in front of the system. Matthias Oster, Rodney J. Douglas, Shih-Chii Liu |
ISCAS | 3 |
| 2007 | A Spike-Based Saccadic Recognition SystemabstractThe paper presents a spike-based saccadic recognition system that uses a temporal-derivative silicon retina on a pan-tilt unit and an aVLSI multi-neuron classifier with a time-to-first-spike output coding. By using the spike information during the last 150 ms of a saccadic movement, we generate a reliable, sparse stimulus representation of image patches. The paper describes a novel classification scheme where the retinal spikes during this time influence the time-to-first spike of classifier neurons which receive the same constant input current. The preferred pattern of the neuron is stored in the synaptic connectivity between the retina and the classifier neuron. The authors demonstrates the robustness and real-time performance of this recognition scheme on a saccadic system which uses analog VLSI components. Matthias Oster, Patrick Lichtsteiner, Tobi Delbruck, Shih-Chii Liu |
ISCAS | 4 |
| 2006 | Spike response properties of an AER EARabstractWe present measured frequency-gain functions and the spike rate outputs of the different sections in a spiking silicon cochlea chip. The chip consists of a matched pair of silicon cochleae with an address event interface for the output. Each section of the cochlea is modelled by a second-order low-pass filter followed by a simplified inner hair cell circuit and a spiking neuron circuit. When the neuron spikes, an address event is generated on the asynchronous data bus. These spike outputs are analogous to the spikes on an auditory nerve, connecting the cochlea with the brain. André van Schaik, Shih-Chii Liu |
ISCAS | 3 |
| 2006 | Feature competition in a spike-based winner-take-all VLSI networkabstractRecurrent networks and hardware analogs that perform a winner-take-all computation have been studied extensively. This computation is rarely demonstrated in a spiking network of neurons receiving input spike trains. In this work, we demonstrate this computation not only within a VLSI network but also across networks of integrate-and-fire neurons in a feature competition task. The chip has four populations of neurons receiving input spike trains that represent the outputs of four feature maps. The connectivity within each population is configured so that all the neurons compete with one another. In addition, a second level of competition, which we call the feature competition, can be introduced between all populations (or feature maps). The two levels of competition are useful in a system that has to select both the locations of relevant features and the best feature map that is coded in an input stimulus. The selection process can be completed as fast as after two input spikes. Shih-Chii Liu, Matthias Oster |
ISCAS | 1 |
| 2006 | Programmable synaptic weights for an aVLSI network of spiking neuronsabstractWe describe a spiking neuronal network which allows local synaptic weights to be assigned to individual synapses. In previous implementations of neuronal networks, the biases that control the parameters of a particular synapse are global to all synapses of the same type regardless of the target neuron. In this new implementation, the parameters for a synapse are set by on-chip digital-analog-converter (DAC) circuits, and the DACs are updated before the selected synapses are activated. Results from the fabricated chip show that the local weights are programmable and the DACs settle in the order of microseconds. These on-chip DACs allow the user to program a selected synaptic weight for connections between neurons and they can also be used for mismatch calibration. Shih-Chii Liu |
ISCAS | 2 |
| 2006 | Attentional Processing on a Spike-Based VLSI Neural NetworkabstractThe neurons of the neocortex communicate by asynchronous events called action potentials (or 'spikes'). However, for simplicity of simulation, most models of processing by cortical neural networks have assumed that the activations of their neurons can be approximated by event rates rather than taking account of individual spikes. The obstacle to exploring the more detailed spike processing of these networks has been reduced considerably in recent years by the development of hybrid analog-digital Very-Large Scale Integrated (hVLSI) neural networks composed of spiking neurons that are able to operate in real-time. In this paper we describe such a hVLSI neural network that performs an interesting task of selective attentional processing that was previously described for a simulated 'pointer-map' rate model by Hahnloser and colleagues. We found that most of the computational features of their rate model can be reproduced in the spiking implementation; but, that spike-based processing requires a modification of the original network architecture in order to memorize a previously attended target. Rodney J. Douglas, Shih-Chii Liu |
NIPS | 3 |
| 2005 | A Hardware/Software Framework for Real-Time Spiking Systems
Matthias Oster, Adrian M. Whatley, Shih-Chii Liu, Rodney J. Douglas |
ICANN (1) | 3 |
| 2005 | Robot Guidance with Neuromorphic Motion SensorsabstractNeuromorphic motion sensors are attractive for use on battery powered robots which require a low payload. Their features include low power consumption, continuous computation, light-weight, and robustness to different light and contrast conditions. Their outputs are not compatible with controllers that require precise measurements from their sensors. We describe a preliminary investigation into neural architectures that can translate information from these type of sensors into an output suitable for controlling the motor outputs of a robot. In this work, we use a neural network to produce an output that is similar to the range measurements of infrared range sensors, and we use this output to guide the behavior of the robot in a collision-avoidance task. Lukas Reichel, David Liechti, Karl Presser, Shih-Chii Liu |
ICRA | 4 |
| 2005 | Spiking Inputs to a Winner-take-all NetworkabstractRecurrent networks that perform a winner-take-all computation have been studied extensively. Although some of these studies include spik- ing networks, they consider only analog input rates. We present results of this winner-take-all computation on a network of integrate-and-fire neurons which receives spike trains as inputs. We show how we can con- figure the connectivity in the network so that the winner is selected after a pre-determined number of input spikes. We discuss spiking inputs with both regular frequencies and Poisson-distributed rates. The robustness of the computation was tested by implementing the winner-take-all network on an analog VLSI array of 64 integrate-and-fire neurons which have an innate variance in their operating parameters. Matthias Oster, Shih-Chii Liu |
NIPS | 2 |
| 2005 | AER Building Blocks for Multi-Layer Multi-Chip Neuromorphic Vision SystemsabstractA 5-layer neuromorphic vision processor whose components communicate spike events asychronously using the address-event- representation (AER) is demonstrated. The system includes a retina chip, two convolution chips, a 2D winner-take-all chip, a delay line chip, a learning classifier chip, and a set of PCBs for computer interfacing and address space remappings. The components use a mixture of analog and digital computation and will learn to classify trajectories of a moving object. A complete experimental setup and measurements results are shown. Rafael Serrano-Gotarredona, Matthias Oster, Patrick Lichtsteiner, Alejandro Linares-Barranco, Rafael Paz-Vicente, Francisco Gomez-Rodriguez, Håvard Kolle Riis, Tobi Delbruck, Shih-Chii Liu, S. Zahnd, Adrian M. Whatley, Rodney J. Douglas, Philipp Häfliger, Gabriel Jiménez-Moreno, Antonio Abad Civit Balcells, Teresa Serrano-Gotarredona, Antonio J. Acosta 0001, Bernabé Linares-Barranco |
NIPS | 9 |
| 2004 | Temporal coding in a silicon network of integrate-and-fire neuronsabstractSpatio-temporal processing of spike trains by neuronal networks depends on a variety of mechanisms distributed across synapses, dendrites, and somata. In natural systems, the spike trains and the processing mechanisms cohere though their common physical instantiation. This coherence is lost when the natural system is encoded for simulation on a general purpose computer. By contrast, analog VLSI circuits are, like neurons, inherently related by their real-time physics, and so, could provide a useful substrate for exploring neuronlike event-based processing. Here, we describe a hybrid analog-digital VLSI chip comprising a set of integrate-and-fire neurons and short-term dynamical synapses that can be configured into simple network architectures with some properties of neocortical neuronal circuits. We show that, despite considerable fabrication variance in the properties of individual neurons, the chip offers a viable substrate for exploring real-time spike-based processing in networks of neurons. Shih-Chii Liu, Rodney J. Douglas |
IEEE Trans. Neural Networks | 1 |
| 2003 | Modeling Short-Term Synaptic Depression in SiliconabstractWe describe a model of short-term synaptic depression that is derived from a circuit implementation. The dynamics of this circuit model is similar to the dynamics of some theoretical models of short-term depression except that the recovery dynamics of the variable describing the depression is nonlinear and it also depends on the presynaptic frequency. The equations describing the steady-state and transient responses of this synaptic model are compared to the experimental results obtained from a fabricated silicon network consisting of leaky integrate-and-fire neurons and different types of short-term dynamic synapses. We also show experimental data demonstrating the possible computational roles of depression. One possible role of a depressing synapse is that the input can quickly bring the neuron up to threshold when the membrane potential is close to the resting potential. Malte Boegershausen, Pascal Suter, Shih-Chii Liu |
Neural Comput. | 3 |
| 2002 | Encoding the Temporal Statistics of Markovian Sequences of Stimuli in Recurrent Neuronal Networks
Alessandro Usseglio Viretta, Stefano Fusi, Shih-Chii Liu |
ICANN | 3 |
| 2002 | Circuit Model of Short-Term Synaptic DynamicsabstractWe describe a model of short-term synaptic depression that is derived from a silicon circuit implementation. The dynamics of this circuit model are similar to the dynamics of some present theoretical models of short- term depression except that the recovery dynamics of the variable de- scribing the depression is nonlinear and it also depends on the presynap- tic frequency. The equations describing the steady-state and transient re- sponses of this synaptic model fit the experimental results obtained from a fabricated silicon network consisting of leaky integrate-and-fire neu- rons and different types of synapses. We also show experimental data demonstrating the possible computational roles of depression. One pos- sible role of a depressing synapse is that the input can quickly bring the neuron up to threshold when the membrane potential is close to the rest- ing potential. Shih-Chii Liu, Malte Boegershausen, Pascal Suter |
NIPS | 1 |
| 2002 | Silicon synaptic adaptation mechanisms for homeostasis and contrast gain controlabstractWe explore homeostasis in a silicon integrate-and-fire neuron. The neuron adapts its firing rate over time periods on the order of seconds or minutes so that it returns to its spontaneous firing rate after a sustained perturbation. Homeostasis is implemented via two schemes. One scheme looks at the presynaptic activity and adapts the synaptic weight depending on the presynaptic spiking rate. The second scheme adapts the synaptic "threshold" depending on the neuron's activity. The threshold is lowered if the neuron's activity decreases over a long time and is increased for prolonged increase in postsynaptic activity. The presynaptic adaptation mechanism models the contrast adaptation responses observed in simple cortical cells. To obtain the long adaptation timescales we require, we used floating-gates. Otherwise, the capacitors we would have to use would be of such a size that we could not integrate them and so we could not incorporate such long-time adaptation mechanisms into a very large-scale integration (VLSI) network of neurons. The circuits for the adaptation mechanisms have been implemented in a 2-/spl mu/m double-poly CMOS process with a bipolar option. The results shown here are measured from a chip fabricated in this process. Shih-Chii Liu, Bradley A. Minch |
IEEE Trans. Neural Networks | 1 |
| 2001 | Orientation-Selective aVLSI Spiking NeuronsabstractWe describe a programmable multi-chip VLSI neuronal system that can be used for exploring spike-based information processing models. The system consists of a silicon retina, a PIC microcontroller, and a transceiver chip whose integrate-and-fire neurons are connected in a soft winner-take-all architecture. The circuit on this multi-neuron chip ap- proximates a cortical microcircuit. The neurons can be configured for different computational properties by the virtual connections of a se- lected set of pixels on the silicon retina. The virtual wiring between the different chips is effected by an event-driven communication pro- tocol that uses asynchronous digital pulses, similar to spikes in a neu- ronal system. We used the multi-chip spike-based system to synthe- size orientation-tuned neurons using both a feedforward model and a feedback model. The performance of our analog hardware spiking model matched the experimental observations and digital simulations of continuous-valued neurons. The multi-chip VLSI system has advantages over computer neuronal models in that it is real-time, and the computa- tional time does not scale with the size of the neuronal network. Shih-Chii Liu, Jörg Kramer, Giacomo Indiveri, Tobi Delbruck, Rodney J. Douglas |
NIPS | 1 |
| 2001 | Orientation-selective aVLSI spiking neurons
Shih-Chii Liu, Jörg Kramer, Giacomo Indiveri, Tobi Delbruck, Thomas Burg, Rodney J. Douglas |
Neural Networks | 1 |
| 2000 | Homeostasis in a Silicon Integrate and Fire NeuronabstractIn this work, we explore homeostasis in a silicon integrate-and-fire neu(cid:173) ron. The neuron adapts its firing rate over long time periods on the order of seconds or minutes so that it returns to its spontaneous firing rate after a lasting perturbation. Homeostasis is implemented via two schemes. One scheme looks at the presynaptic activity and adapts the synaptic weight depending on the presynaptic spiking rate. The second scheme adapts the synaptic "threshold" depending on the neuron's activity. The threshold is lowered if the neuron's activity decreases over a long time and is increased for prolonged increase in postsynaptic activity. Both these mechanisms for adaptation use floating-gate technology. The re(cid:173) sults shown here are measured from a chip fabricated in a 2-J.lm CMOS process. Shih-Chii Liu, Bradley A. Minch |
NIPS | 1 |
| 1999 | A Winner-Take-All Circuit with Controllable Soft Max Property
Shih-Chii Liu |
NIPS | 1 |
| 1997 | Silicon Retina with Adaptive Filtering Properties
Shih-Chii Liu |
NIPS | 1 |
| 1995 | Adaptive Retina with Center-Surround Receptive Field
Shih-Chii Liu, Kwabena Boahen 0001 |
NIPS | 1 |
| 1994 | Continuous-Time Adaptive Delay SystemabstractWe have developed an adaptive delay system that adjusts the delay of a delay element so that it matches the temporal disparity between the onset of two input signals. The delay is controlled either by an external bias voltage, or by an intrinsic signal derived from an adaptive block. The operation of the adaptive delay system is similar to that of a charge-pump phase-lock loop, with an extended lock-in range of more than 5 decades. Standard CMOS transistors are used in their subthreshold region. Experimental results from circuits fabricated in 2 /spl mu/m CMOS technology are in agreement with the analysis.> Shih-Chii Liu, Carver Mead |
ISCAS | 1 |
| 1992 | Object-Based Analog VLSI Vision Circuits
Christof Koch, Binnal Mathur, Shih-Chii Liu, John G. Harris, Massimo Sivilotti |
NIPS | 3 |
| 1992 | Dynamic wires: An alanog VLSI model for object-based processing
Shih-Chii Liu, John G. Harris |
Int. J. Comput. Vis. | 1 |
| 1989 | Generalized smoothing networks in early visionabstractGeneralized smoothing networks have been developed which enforce smoothness constraints for any arbitrary level of derivative of the input data. Furthermore, discontinuities of any order of derivative can be detected by providing for continuous line processes, which selectively inhibit smoothing. Second- and higher-order networks are required for many problems in early vision; first-order networks are often unsatisfactory. Examples in surface interpolation, edge detection, and image segmentation are shown. Solution of these types of problems typically takes a prohibitive amount of time, even on supercomputers. A significant advantage of these proposed networks is that they can be mapped directly to analog VLSI hardware.> Shih-Chii Liu, John G. Harris |
CVPR | 1 |