EDBT 2026 Demo / reviewers in the wild / expert
Tobi Delbruck
dblp:d/TobiDelbruck · also Tobi Delbrueck, Tobi Delbrück, Tobias Delbrück
· DBLP profile ↗
95ranked-venue papers
15as first author
14since 2021 · last 2025
0000-0001-5479-1141ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 63 · 11 first-author · 4 since 2021Artificial intelligence and machine learning · 34 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Modulating State Space Model with SlowFast Framework for Compute-Efficient Ultra Low-Latency Speech EnhancementabstractDeep learning-based speech enhancement (SE) methods often face significant computational challenges when needing to meet low-latency requirements because of the increased number of frames to be processed. This paper introduces the SlowFast framework which aims to reduce computation costs specifically when low-latency enhancement is needed. The framework consists of a slow branch that analyzes the acoustic environment at a low frame rate, and a fast branch that performs SE in the time domain at the needed higher frame rate to match the required latency. Specifically, the fast branch employs a state space model where its state transition process is dynamically modulated by the slow branch. Experiments on a SE task with a 2 ms algorithmic latency requirement using the Voice Bank + Demand dataset show that our approach reduces computation cost by 70% compared to a baseline single-branch network with equivalent parameters, without compromising enhancement performance. Furthermore, by leveraging the SlowFast framework, we implemented a network that achieves an algorithmic latency of just 62.5 μs (one sample point at 16 kHz sample rate) with a computation cost of 100 M MACs/s, while scoring a PESQ-NB of 3.12 and SISNR of 16.62. Longbiao Cheng, Ashutosh Pandey 0004, Buye Xu, Tobi Delbruck, Vamsi K. Ithapu, Shih-Chii Liu |
ICASSP | 4 |
| 2025 | Steering Prediction via a Multi-Sensor System for Autonomous RacingabstractAutonomous racing has rapidly gained research attention. Traditionally, racing cars rely on 2D LiDAR as their primary visual system. In this work, we explore the integration of an event camera with the existing system to provide enhanced temporal information. Our goal is to fuse the 2D LiDAR data with event data in an end-to-end learning framework for steering prediction, which is crucial for autonomous racing. To the best of our knowledge, this is the first study addressing this challenging research topic. We start by creating a multisensor dataset specifically for steering prediction. Using this dataset, we establish a benchmark by evaluating various SOTA fusion methods. Our observations reveal that existing methods often incur substantial computational costs. To address this, we apply low-rank techniques to propose a novel, efficient, and effective fusion design. We introduce a new fusion learning policy to guide the fusion process, enhancing robustness against misalignment. Our fusion architecture provides better steering prediction than LiDAR alone, significantly reducing the RMSE from 7.72 to 1.28. Compared to the second-best fusion method, our work represents only 11% of the learnable parameters while achieving better accuracy. The source code and dataset are publicly available at: https://github.com/ZZY-Zhou/F1Tenth-Steering. Zhuyun Zhou, Zongwei Wu, Florian Bolli, Rémi Boutteau, Fan Yang 0019, Radu Timofte, Dominique Ginhac, Tobi Delbruck |
ICRA | 8 |
| 2025 | FPGA Hardware Neural Control of CartPole and F1TENTH Race CarabstractLatency and computational cost often limit the use of Nonlinear Model Predictive Control (NMPC) in real-time robotics. To address this limitation, our work investigates FPGA-implemented Neural Controllers (NC) trained through supervised learning, mimicking NMPC. We show that inexpensive embedded FPGA hardware is sufficient to implement these neural controllers for high-frequency control of robotic systems. We demonstrate kilohertz control rates for a cartpole and offload control to the FPGA hardware on the F1TENTH race car. The FPGA NC outperforms NMPC on the cartpole, due to the faster control rate afforded by faster NC inference. The code and hardware implementation for this paper are available at https://github.com/SensorsINI/Neural-Control-Tools. Marcin Paluch, Florian Bolli, Antonio Rios-Navarro, Chang Gao 0002, Tobi Delbruck |
IROS | 6 |
| 2024 | Exploiting Symmetric Temporally Sparse BPTT for Efficient RNN TrainingabstractRecurrent Neural Networks (RNNs) are useful in temporal sequence tasks. However, training RNNs involves dense matrix multiplications which require hardware that can support a large number of arithmetic operations and memory accesses. Implementing online training of RNNs on the edge calls for optimized algorithms for an efficient deployment on hardware. Inspired by the spiking neuron model, the Delta RNN exploits temporal sparsity during inference by skipping over the update of hidden states from those inactivated neurons whose change of activation across two timesteps is below a defined threshold. This work describes a training algorithm for Delta RNNs that exploits temporal sparsity in the backward propagation phase to reduce computational requirements for training on the edge. Due to the symmetric computation graphs of forward and backward propagation during training, the gradient computation of inactivated neurons can be skipped. Results show a reduction of ∼80% in matrix operations for training a 56k parameter Delta LSTM on the Fluent Speech Commands dataset with negligible accuracy loss. Logic simulations of a hardware accelerator designed for the training algorithm show 2-10X speedup in matrix computations for an activation sparsity range of 50%-90%. Additionally, we show that the proposed Delta RNN training will be useful for online incremental learning on edge devices with limited computing resources. Chang Gao 0002, Zuowen Wang, Longbiao Cheng, Shih-Chii Liu, Tobi Delbruck |
AAAI | 7 |
| 2024 | Dynamic Gated Recurrent Neural Network for Compute-efficient Speech EnhancementabstractThis paper introduces a new Dynamic Gated Recurrent Neural Network (DG-RNN) for compute-efficient speech enhancement models running on resource-constrained hardware platforms.It leverages the slow evolution characteristic of RNN hidden states over steps, and updates only a selected set of neurons at each step by adding a newly proposed select gate to the RNN model.This select gate allows the computation cost of the conventional RNN to be reduced during network inference.As a realization of the DG-RNN, we further propose the Dynamic Gated Recurrent Unit (D-GRU) which does not require additional parameters.Test results obtained from several state-ofthe-art compute-efficient RNN-based speech enhancement architectures using the DNS challenge dataset, show that the D-GRU based model variants maintain similar speech intelligibility and quality metrics comparable to the baseline GRU based models even with an average 50% reduction in GRU computes. Longbiao Cheng, Ashutosh Pandey 0004, Buye Xu, Tobi Delbruck, Shih-Chii Liu |
INTERSPEECH | 4 |
| 2024 | A Rapid and Robust Tendon-Driven Robotic Hand for Human-Robot Interactions Playing Rock-Paper-ScissorsabstractRapid human-robot interactions require fast hardware platforms with minimal latency and high reliability. In response, we present a cost-effective, electrically actuated, tendon-driven robotic hand. This hand features a unique spool-free actuation mechanism that achieves a limit-to-limit flexion movement in less than 60 ms, matching human speed. To our knowledge, it is among the fastest electric motor tendon-driven robotic hands available today. The high speed of the robotic hand was successfully demonstrated in public by playing Rock Paper Scissors at a science fair. This research work outlines the design methodology and introduces a simulation-optimization framework that allows users to preview the motion of the hand, quantify the actuation performance, and customize the design parameters prior to fabrication. The proposed actuation mechanism, along with the simulation and optimization tools, illustrates design principles and computational methods applicable to other dynamic human-robot applications that require fast reaction times. The Dextra hand design is available at https://sensorsini.github.io/dextra-robot-hand. Stefan Weirich, Robert K. Katzschmann, Tobi Delbruck |
RO-MAN | 4 |
| 2024 | Spartus: A 9.4 TOp/s FPGA-Based LSTM Accelerator Exploiting Spatio-Temporal SparsityabstractLong short-term memory (LSTM) recurrent networks are frequently used for tasks involving time-sequential data, such as speech recognition. Unlike previous LSTM accelerators that either exploit spatial weight sparsity or temporal activation sparsity, this article proposes a new accelerator called "Spartus" that exploits spatio-temporal sparsity to achieve ultralow latency inference. Spatial sparsity is induced using a new column-balanced targeted dropout (CBTD) structured pruning method, producing structured sparse weight matrices for a balanced workload. The pruned networks running on Spartus hardware achieve weight sparsity levels of up to 96% and 94% with negligible accuracy loss on the TIMIT and the Librispeech datasets. To induce temporal sparsity in LSTM, we extend the previous DeltaGRU method to the DeltaLSTM method. Combining spatio-temporal sparsity with CBTD and DeltaLSTM saves on weight memory access and associated arithmetic operations. The Spartus architecture is scalable and supports real-time online speech recognition when implemented on small and large FPGAs. Spartus per-sample latency for a single DeltaLSTM layer of 1024 neurons averages 1 μ s. Exploiting spatio-temporal sparsity on our test LSTM network using the TIMIT dataset leads to 46 × speedup of Spartus over its theoretical hardware performance to achieve 9.4-TOp/s effective batch-1 throughput and 1.1-TOp/s/W power efficiency. Chang Gao 0002, Tobi Delbruck, Shih-Chii Liu |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Deep Polarization Reconstruction with PDAVIS EventsabstractThe polarization event camera PDAVIS is a novel bio-inspired neuromorphic vision sensor that reports both conventional polarization frames and asynchronous, continuously per-pixel polarization brightness changes (polarization events) with fast temporal resolution and large dynamic range. A deep neural network method (Polarization FireNet) was previously developed to reconstruct the polarization angle and degree from polarization events for bridging the gap between the polarization event camera and mainstream computer vision. However, Polarization FireNet applies a network pretrained for normal event-based frame reconstruction independently on each of four channels of polarization events from four linear polarization angles, which ignores the correlations between channels and inevitably introduces content inconsistency between the four reconstructed frames, resulting in unsatisfactory polarization reconstruction performance. In this work, we strive to train an effective, yet efficient, DNN model that directly outputs polarization from the input raw polarization events. To this end, we constructed the first large-scale event-to-polarization dataset, which we subsequently employed to train our events-to-polarization network E2P. E2P extracts rich polarization patterns from input polarization events and enhances features through cross-modality context integration. We demonstrate that E2P outperforms Polarization FireNet by a significant margin with no additional computing cost. Experimental results also show that E2P produces more accurate measurement of polarization than the PDAVIS frames in challenging fast and high dynamic range scenes. Code and data are publicly available at: https://github.com/SensorsINI/e2p. Haiyang Mei, Zuowen Wang, Xin Yang 0011, Xiaopeng Wei, Tobi Delbruck |
CVPR | 5 |
| 2023 | RPGD: A Small-Batch Parallel Gradient Descent Optimizer with Explorative Resampling for Nonlinear Model Predictive ControlabstractNonlinear model predictive control often involves nonconvex optimization for which real-time control systems require fast and numerically stable solutions. This work proposes RPGD, a Resampling Parallel Gradient Descent optimizer designed to exploit small-batch parallelism of modern hardware like neural accelerators or multithreaded microcontrollers. After initialization, it continuously maintains a small population of good control trajectory solution candidates and improves them using gradient information, followed by selection of elite candidates and resampling of the others. In simulation on a cartpole, the OpenAI Gym mountain car, a Dubins car with obstacles, and a high input dimensional 2D arm, it produces similar or lower MPC costs than benchmark cross-entropy and path integral methods. On a physical cartpole, it performs swing-up and cart target following of the pole, using either a differential equation or multilayer perceptron as dynamics model. RPGD drives an F1TENTH simulated race car at near-optimal lap times and a real F1TENTH car in laps around a cluttered room. We study alterations of RPGD's building blocks to justify its composition. RPGD compute time in Python with TensorFlow optimization running on CPU is 2 to 4 times slower than the FORCESPRO commercial embedded solver. Frederik Heetmeyer, Marcin Paluch, Diego Bolliger, Florian Bolli, Ennio Filicicchia, Tobi Delbruck |
ICRA | 7 |
| 2023 | Low Cost and Latency Event Camera Background Activity DenoisingabstractDynamic Vision Sensor (DVS) event camera output includes uninformative background activity (BA) noise events that increase dramatically under dim lighting. Existing denoising algorithms are not effective under these high noise conditions. Furthermore, it is difficult to quantitatively compare algorithm accuracy. This paper proposes a novel framework to better quantify BA denoising algorithms by measuring receiver operating characteristics with known mixtures of signal and noise DVS events. New datasets for stationary and moving camera applications of DVS in surveillance and driving are used to compare 3 new low-cost algorithms: Algorithm 1 checks distance to past events using a tiny fixed size window and removes most of the BA while preserving most of the signal for stationary camera scenarios. Algorithm 2 uses a memory proportional to the number of pixels for improved correlation checking. Compared with existing methods, it removes more noise while preserving more signal. Algorithm 3 uses a lightweight multilayer perceptron classifier driven by local event time surfaces to achieve the best accuracy over all datasets. The code and data are shared with the paper as DND21. Shasha Guo 0001, Tobi Delbruck |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Utility and Feasibility of a Center Surround Event CameraabstractStandard dynamic vision sensor (DVS) event cameras output a stream of spatially-independent log-intensity brightness change events so they cannot suppress spatial redundancy. Nearly all biological retinas use an antagonistic center-surround organization. This paper proposes a practical method of implementing a compact, energy-efficient Center Surround DVS (CSDVS) with a surround smoothing network that uses compact polysilicon resistors for lateral resistance. The paper includes behavioral simulation results for the CSDVS (see sites.google.com/view/csdvs/home). The CSDVS would significantly reduce events caused by low spatial frequencies, but amplify the informative high frequency spatiotemporal events. Tobi Delbruck, Cheng-Han Li, Rui Graca, Brian McReynolds |
ICIP | 1 |
| 2022 | Event-Based Vision: A SurveyabstractEvent cameras are bio-inspired sensors that differ from conventional frame cameras: Instead of capturing images at a fixed rate, they asynchronously measure per-pixel brightness changes, and output a stream of events that encode the time, location and sign of the brightness changes. Event cameras offer attractive properties compared to traditional cameras: high temporal resolution (in the order of μs), very high dynamic range (140 dB versus 60 dB), low power consumption, and high pixel bandwidth (on the order of kHz) resulting in reduced motion blur. Hence, event cameras have a large potential for robotics and computer vision in challenging scenarios for traditional cameras, such as low-latency, high speed, and high dynamic range. However, novel methods are required to process the unconventional output of these sensors in order to unlock their potential. This paper provides a comprehensive overview of the emerging field of event-based vision, with a focus on the applications and the algorithms developed to unlock the outstanding properties of event cameras. We present event cameras from their working principle, the actual sensors that are available and the tasks that they have been used for, from low-level vision (feature detection and tracking, optic flow, etc.) to high-level vision (reconstruction, segmentation, recognition). We also discuss the techniques developed to process events, including learning-based techniques, as well as specialized processors for these novel sensors, such as spiking neural networks. Additionally, we highlight the challenges that remain to be tackled and the opportunities that lie ahead in the search for a more efficient, bio-inspired way for machines to perceive and interact with the world. Guillermo Gallego 0002, Tobi Delbruck, Garrick Orchard, Chiara Bartolozzi, Brian Taba, Andrea Censi, Stefan Leutenegger, Andrew J. Davison, Jörg Conradt, Kostas Daniilidis, Davide Scaramuzza 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | EDFLOW: Event Driven Optical Flow Camera With Keypoint Detection and Adaptive Block MatchingabstractEvent cameras such as the Dynamic Vision Sensor (DVS) are useful because of their low latency, sparse output, and high dynamic range. In this paper, we propose a DVS+FPGA camera platform and use it to demonstrate the hardware implementation of event-based corner keypoint detection and adaptive block-matching optical flow. To adapt sample rate dynamically, events are accumulated in event slices using the area event count slice exposure method. The area event count is feedback controlled by the average optical flow matching distance. Corners are detected by streaks of accumulated events on event slice rings of radius 3 and 4 pixels. Corner detection takes about 6 clock cycles (16 MHz event rate at the 100MHz clock frequency) At the corners, flow vectors are computed in 100 clock cycles (1 MHz event rate). The multiscale block match size is$25\times 25$pixels and the flow vectors span up to 30-pixel match distance. The FPGA processes the sum-of-absolute distance block matching at 123 GOp/s, the equivalent of 1230 Op/clock cycle. EDFLOW is several times more accurate on MVSEC drone and driving optical flow benchmarking sequences than the previous best DVS FPGA optical flow implementation, and achieves similar accuracy to the CNN-based EV-Flownet, although it burns about 100 times less power. The EDFLOW design and benchmarking videos are available athttps://sites.google.com/view/edflow21/home. Min Liu 0031, Tobi Delbruck |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Reducing Latency in a Converted Spiking Video Segmentation NetworkabstractSpiking Neural Networks (SNNs) can be configured to produce almost-equivalent accurate Analog Neural Networks (ANNs) by various ANN-SNN conversion methods. Most of these methods are applied to classification and object detection networks tested on frame-based datasets. In this work, we demonstrate a converted SNN for image segmentation and applied to a natural video dataset. Instead of resetting the network state with each input frame, we capitalize on the temporal redundancy between adjacent frames in a natural scene, and propose an interval reset method where the network state is reset after a fixed number of frames. We studied the trade-off between accuracy and latency with the number of interval reset frames. We also applied layer-specific normalization and early stopping to speed up network convergence and to reduce the latency. Our results show that the SNN achieved a 35.7x increase in convergence speed with only 1.5% accuracy drop using an interval reset of 20 frames. Qinyu Chen, Bodo Rueckauer, Tobi Delbruck, Shih-Chii Liu |
ISCAS | 4 |
| 2020 | Learning to Exploit Multiple Vision Modalities by Using Grafted Networks
Yuhuang Hu, Tobi Delbruck, Shih-Chii Liu |
ECCV (16) | 2 |
| 2020 | Recurrent Neural Network Control of a Hybrid Dynamical Transfemoral Prosthesis with EdgeDRNN AcceleratorabstractLower leg prostheses could improve the life quality of amputees by increasing comfort and reducing energy to locomote, but currently control methods are limited in modulating behaviors based upon the human's experience. This paper describes the first steps toward learning complex controllers for dynamical robotic assistive devices. We provide the first example of behavioral cloning to control a powered transfemoral prostheses using a Gated Recurrent Unit (GRU) based recurrent neural network (RNN) running on a custom hardware accelerator that exploits temporal sparsity. The RNN is trained on data collected from the original prosthesis controller. The RNN inference is realized by a novel EdgeDRNN accelerator in real-time. Experimental results show that the RNN can replace the nominal PD controller to realize end-to-end control of the AMPRO3 prosthetic leg walking on flat ground and unforeseen slopes with comparable tracking accuracy. EdgeDRNN computes the RNN about 240 times faster than real time, opening the possibility of running larger networks for more complex tasks in the future. Implementing an RNN on this real-time dynamical system with impacts sets the ground work to incorporate other learned elements of the human-prosthesis system into prosthesis control. Chang Gao 0002, Rachel Gehlhar, Aaron D. Ames, Shih-Chii Liu, Tobi Delbruck |
ICRA | 5 |
| 2020 | Lessons Learned the Hard Wayabstract“Fail often to succeed sooner” is a common mantra that we are told is the secret to success. When reporting research results, however, scholars rarely write about their failed attempts and only focus on the successful ones. Perhaps the source of this disconnect between what we preach and what we do can be found in the underlying assumption that published work is meant to move the field forward and failed attempts supposedly do not. The goal of the confessions presented in this paper is to show that even failed attempts are genuine and valuable contributions to our field provided that we learn from our mistakes and correct them. The 27 confessions span from planning oversights, digital and analog design errors, misunderstanding of devices, overlooked parasitics, LVS errors, and troubles in testing. Tobi Delbruck, Ibrahim M. Elfadel, Shahzad Muzaffar, Germain Haessig, Bo Wang 0012, Amine Bermak, Rui Graca, Luis A. Camuñas-Mesa, Bathiya Senevirathna, Pamela Abshire, Bernabé Linares-Barranco, Saeed Afshar, Shih-Chii Liu, Runchun Wang, Piotr Dudek, Stephen J. Carey, José M. de la Rosa 0001, Marc Dandin, Sheung Lu, Vincent Frick, Teresa Serrano-Gotarredona, Paula López Martinez 0001, Melika Payvand, Advait Madhavan, Eric R. Fossum, Juan Camilo Vasquez Tieck, Yan Liu 0016, Timothy G. Constandinou, Alexander Serb, Ricardo Carmona-Galán, Robert Nawrocki, Walter D. Leon-Salas |
ISCAS | 1 |
| 2020 | Live Demonstration: CNN Edge Computing for Mobile Robot NavigationabstractThe brain cortex processes visual information to classify it following a scheme that has been mimicked by Convolutional Neural Networks (CNN). Specialised hardware accelerators are currently used as CPU co-processors for mobile applications. These accelerators are getting closer to the sensors for an edge computation of its output towards a faster and lower power consumption improvements. In this demonstration we use a dynamic vision sensor (inspired in the retina neural cells) as a visual source of the NullHop CNN accelerator deployed on a MPSoC FPGA and placed into a mobile robot for edge-computing the visual information and classify it to properly command a Summit-XL mobile robot for a target destiny. The reduced latency of the used CNN accelerator allows to process several histograms before taking a movement decision. A distance sensor mounted on the robot ensures that the direction change is done at the right distance for a proper path following. Enrique Piñero-Fuentes, Antonio Rios-Navarro, Ricardo Tapiador-Morales, Tobi Delbruck, Alejandro Linares-Barranco |
ISCAS | 4 |
| 2020 | Self Calibration of Wide Dynamic Range Bias Current GeneratorsabstractMany neuromorphic chips now include on-chip, digitally programmable bias generator circuits. So far, precision of these generated biases has been designed by transistor sizing and circuit design to ensure tolerable statistical variance due to threshold mismatch. This paper reports the use of an integrated measurement circuit based on spiking neuron and a scheme for calibrating each chip set of biases against the smallest of all the biases from that chip. That way, the averaging across individual biases improves overall matching both within a chip and across chips. This paper includes measurements of generated biases, the method for remapping bias values towards more uniform values, and measurements across 5 sample chips. With the method presented in this paper, 1s mismatch of subthreshold currents is decreased by at least a factor of 3. The firmware implementation completes calibration in about a minute and uses about 1kB of flash storage of calibration data. Zhenming Yu, Tobi Delbruck |
ISCAS | 2 |
| 2019 | EV-IMO: Motion Segmentation Dataset and Learning Pipeline for Event CamerasabstractWe present the first event-based learning approach for motion segmentation in indoor scenes and the first event-based dataset - EV-IMO- which includes accurate pixel-wise motion masks, egomotion and ground truth depth. Our approach is based on an efficient implementation of the SfM learning pipeline using a low parameter neural network architecture on event data. In addition to camera egomotion and a dense depth map, the network estimates independently moving object segmentation at the pixel-level and computes per-object 3D translational velocities of moving objects. We also train a shallow network with just 40k parameters, which is able to compute depth and egomotion. Our EV-IMO dataset features 32 minutes of indoor recording with up to 3 fast moving objects in the camera field of view. The objects and the camera are tracked using a VICON®motion capture system. By 3D scanning the room and the objects, ground truth of the depth map and pixel-wise object masks are obtained. We then train and evaluate our learning pipeline on EV-IMO and demonstrate that it is well suited for scene constrained robotics applications. SUPPLEMENTARY MATERIAL The supplementary video, code, trained models, appendix and a dataset will be made available at http://prg.cs.umd.edu/EV-IMO.html. Anton Mitrokhin, Chengxi Ye, Cornelia Fermüller, Yiannis Aloimonos, Tobi Delbruck |
IROS | 5 |
| 2019 | Live Demonstration: Real-Time Spoken Digit Recognition using the DeltaRNN AcceleratorabstractThis demonstration shows a real-time continuous speech recognition hardware system using our previously published DeltaRNN accelerator that enables low latency recurrent neural network (RNN) computation. The network is trained on augmented audio samples from the TIDIGITS dataset to achieve a label error rate (LER) of 2.31%. It is implemented on a Xilinx Zynq-7100 FPGA running at 1 MHz. The incremental RNN power consumption is 30 mW. Visitors interact with the system by speaking digits into a microphone connected to the FPGA system and the classification outputs of the network are continuously displayed on a laptop screen in real time. Chang Gao 0002, Stefan Braun 0005, Ilya Kiselev, Jithendar Anumula, Tobi Delbruck, Shih-Chii Liu |
ISCAS | 5 |
| 2019 | Real-Time Speech Recognition for IoT Purpose using a Delta Recurrent Neural Network AcceleratorabstractThis paper describes a continuous speech recognition hardware system that uses a delta recurrent neural network accelerator (DeltaRNN) implemented on a Xilinx Zynq-7100 FPGA to enable low latency recurrent neural network (RNN) computation. The implemented network consists of a single-layer RNN with 256 gated recurrent unit (GRU) neurons and is driven by input features generated either from the output of a filter bank running on the ARM core of the FPGA in a PmodMic3 microphone setup or from the asynchronous outputs of a spiking silicon cochlea circuit. The microphone setup achieves 7.1 ms minimum latency and 177 frames-per-second (FPS) maximum throughput while the cochlea setup achieves 2.9 ms minimum latency and 345 FPS maximum throughput. The low latency and 70 mW power consumption of the DeltaRNN makes it suitable as an IoT computing platform. Chang Gao 0002, Stefan Braun 0005, Ilya Kiselev, Jithendar Anumula, Tobi Delbruck, Shih-Chii Liu |
ISCAS | 5 |
| 2019 | Incremental Learning Meets Reduced Precision NetworksabstractHardware accelerators for Deep Neural Networks (DNNs) that use reduced precision parameters are more energy efficient than the equivalent full precision networks. While many studies have focused on reduced precision training methods for supervised networks with the availability of large datasets, less work has been reported on incremental learning algorithms that adapt the network for new classes and the consequence of reduced precision has on these algorithms. This paper presents an empirical study of how reduced precision training methods impact the iCARL incremental learning algorithm. The incremental network accuracies on the CIFAR-100 image dataset show that weights can be quantized to 1 bit (2.39% drop in accuracy) but when activations are quantized to 1 bit, the accuracy drops much more (12.75%). Quantizing gradients from 32 to 8 bits only affects the accuracies of the trained network by less than 1%. These results are encouraging for hardware accelerators that support incremental learning algorithms. Yuhuang Hu, Tobi Delbruck, Shih-Chii Liu |
ISCAS | 2 |
| 2019 | Lip Reading Deep Network Exploiting Multi-Modal Spiking Visual and Auditory SensorsabstractThis work presents a lip reading deep neural network that fuses the asynchronous spiking outputs of two bio-inspired silicon multimodal sensors: the Dynamic Vision Sensor (DVS) and the Dynamic Audio Sensor (DAS). The fusion network is tested on the GRID visual-audio lipreading dataset. Classification is carried out using event-based features generated from the spikes of the DVS and DAS. Networks are trained separately on the two modalities and also jointly trained on both modalities. The jointly trained network when tested on DVS spike frames alone, showed a relative increase in accuracy of around 23% over that of the single DVS modality network. Daniel Neil, Tobi Delbruck, Shih-Chii Liu |
ISCAS | 3 |
| 2019 | CNN-based Object Detection on Low Precision Hardware: Racing Car Case StudyabstractIncreasing interest in deep learning and convolutional neural networks resulted in the last years in multiple techniques aiming to improve their accuracy, training speed, and inference speed. At the same time, their computational cost triggered the design of several dedicated hardware architectures, aiming to handle the elevated number of operations neural networks require with minimal power budget, often exploiting reduced precision arithmetic. In this case study, we analyzed how several techniques can be merged together in the design of a track detector for a self-driving racing car, illustrating a step-by-step procedure required to adapt several theoretical works to a real-world scenario. Compared with the best previous detector, the new Proteins cone detector is optimized for low-precision deep learning accelerators. It runs 50% faster on GPU than the previous detector and at a simulated 272.5 FPS on a 1 W ASIC or at 16.4 FPS on a 12 W FPGA and achieves a detection score 16% higher than the previous implementation. Nicolò De Rita, Alessandro Aimar, Tobi Delbruck |
IV | 3 |
| 2019 | NullHop: A Flexible Convolutional Neural Network Accelerator Based on Sparse Representations of Feature MapsabstractConvolutional neural networks (CNNs) have become the dominant neural network architecture for solving many state-of-the-art (SOA) visual processing tasks. Even though graphical processing units are most often used in training and deploying CNNs, their power efficiency is less than 10 GOp/s/W for single-frame runtime inference. We propose a flexible and efficient CNN accelerator architecture called NullHop that implements SOA CNNs useful for low-power and low-latency application scenarios. NullHop exploits the sparsity of neuron activations in CNNs to accelerate the computation and reduce memory requirements. The flexible architecture allows high utilization of available computing resources across kernel sizes ranging from 1×1 to 7×7. NullHop can process up to 128 input and 128 output feature maps per layer in a single pass. We implemented the proposed architecture on a Xilinx Zynq field-programmable gate array (FPGA) platform and presented the results showing how our implementation reduces external memory transfers and compute time in five different CNNs ranging from small ones up to the widely known large VGG16 and VGG19 CNNs. Postsynthesis simulations using Mentor Modelsim in a 28-nm process with a clock frequency of 500 MHz show that the VGG19 network achieves over 450 GOp/s. By exploiting sparsity, NullHop achieves an efficiency of 368%, maintains over 98% utilization of the multiply-accumulate units, and achieves a power efficiency of over 3 TOp/s/W in a core area of 6.3 mm2. As further proof of NullHop's usability, we interfaced its FPGA implementation with a neuromorphic event camera for real-time interactive demonstrations. Alessandro Aimar, Hesham Mostafa, Enrico Calabrese, Antonio Rios-Navarro, Ricardo Tapiador-Morales, Iulia-Alexandra Lungu, Moritz B. Milde, Federico Corradi, Alejandro Linares-Barranco, Shih-Chii Liu, Tobi Delbruck |
IEEE Trans. Neural Networks Learn. Syst. | 11 |
| 2018 | Adaptive Time-Slice Block-Matching Optical Flow Algorithm for Dynamic Vision Sensors
Min Liu 0031, Tobi Delbruck |
BMVC | 2 |
| 2018 | DeltaRNN: A Power-efficient Recurrent Neural Network AcceleratorabstractRecurrent Neural Networks (RNNs) are widely used in speech recognition and natural language processing applications because of their capability to process temporal sequences. Because RNNs are fully connected, they require a large number of weight memory accesses, leading to high power consumption. Recent theory has shown that an RNN delta network update approach can reduce memory access and computes with negligible accuracy loss. This paper describes the implementation of this theoretical approach in a hardware accelerator called "DeltaRNN" (DRNN). The DRNN updates the output of a neuron only when the neuron»s activation changes by more than a delta threshold. It was implemented on a Xilinx Zynq-7100 FPGA. FPGA measurement results from a single-layer RNN of 256 Gated Recurrent Unit (GRU) neurons show that the DRNN achieves 1.2 TOp/s effective throughput and 164 GOp/s/W power efficiency. The delta update leads to a 5.7x speedup compared to a conventional RNN update because of the sparsity created by the DN algorithm and the zero-skipping ability of DRNN. Chang Gao 0002, Daniel Neil, Enea Ceolini, Shih-Chii Liu, Tobi Delbruck |
FPGA | 5 |
| 2018 | Live Demonstration: Front and Back Illuminated Dynamic and Active Pixel Vision Sensors ComparisonabstractThe demonstration shows the differences between two novel Dynamic and Active Pixel Vision Sensors (DAVIS). While both sensors are based on the same circuits and have the same resolution (346×260), they differ in their manufacturing. The first sensor is a DAVIS with standard Front Side Illuminated (FSI) technology and the second sensor is the first Back Side Illuminated (BSI) DAVIS sensor. Gemma Taverni, Diederik Paul Moeys, Cheng-Han Li, Tobi Delbruck, Celso Cavaco, Vasyl Motsnyi, David San Segundo Bello |
ISCAS | 4 |
| 2018 | Event-Based, 6-DOF Camera Tracking from Photometric Depth MapsabstractEvent cameras are bio-inspired vision sensors that output pixel-level brightness changes instead of standard intensity frames. These cameras do not suffer from motion blur and have a very high dynamic range, which enables them to provide reliable visual information during high-speed motions or in scenes characterized by high dynamic range. These features, along with a very low power consumption, make event cameras an ideal complement to standard cameras for VR/AR and video game applications. With these applications in mind, this paper tackles the problem of accurate, low-latency tracking of an event camera from an existing photometric depth map (i.e., intensity plus depth information) built via classic dense reconstruction pipelines. Our approach tracks the 6-DOF pose of the event camera upon the arrival of each event, thus virtually eliminating latency. We successfully evaluate the method in both indoor and outdoor scenes and show that-because of the technological advantages of the event camera-our pipeline works in scenes characterized by high-speed motion, which are still inaccessible to standard cameras. Guillermo Gallego 0002, Jon E. A. Lund, Elias Mueggler, Henri Rebecq, Tobi Delbruck, Davide Scaramuzza 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2017 | A Low Power, Fully Event-Based Gesture Recognition SystemabstractWe present the first gesture recognition system implemented end-to-end on event-based hardware, using a TrueNorth neurosynaptic processor to recognize hand gestures in real-time at low power from events streamed live by a Dynamic Vision Sensor (DVS). The biologically inspired DVS transmits data only when a pixel detects a change, unlike traditional frame-based cameras which sample every pixel at a fixed frame rate. This sparse, asynchronous data representation lets event-based cameras operate at much lower power than frame-based cameras. However, much of the energy efficiency is lost if, as in previous work, the event stream is interpreted by conventional synchronous processors. Here, for the first time, we process a live DVS event stream using TrueNorth, a natively event-based processor with 1 million spiking neurons. Configured here as a convolutional neural network (CNN), the TrueNorth chip identifies the onset of a gesture with a latency of 105 ms while consuming less than 200 mW. The CNN achieves 96.5% out-of-sample accuracy on a newly collected DVS dataset (DvsGesture) comprising 11 hand gesture categories from 29 subjects under 3 illumination conditions. Arnon Amir, Brian Taba, David J. Berg, Timothy Melano, Jeffrey L. McKinstry, Carmelo di Nolfo, Tapan K. Nayak, Alexander Andreopoulos, Guillaume Garreau, Marcela Mendoza, Jeffrey A. Kusnitz, Michael DeBole, Steven K. Esser, Tobi Delbruck, Myron Flickner, Dharmendra S. Modha |
CVPR | 14 |
| 2017 | Neuromorphic Approach Sensitivity Cell Modeling and FPGA Implementation
Antonio Rios-Navarro, Diederik Paul Moeys, Tobi Delbruck, Alejandro Linares-Barranco |
ICANN (1) | 4 |
| 2017 | Delta Networks for Optimized Recurrent Network ComputationabstractMany neural networks exhibit stability in their activation patterns over time in response to inputs from sensors operating under real-world conditions. By capitalizing on this property of natural signals, we propose a Recurrent Neural Network (RNN) architecture called a delta network in which each neuron transmits its value only when the change in its activation exceeds a threshold. The execution of RNNs as delta networks is attractive because their states must be stored and fetched at every timestep, unlike in convolutional neural networks (CNNs). We show that a naive run-time delta network implementation offers modest improvements on the number of memory accesses and computes, but optimized training techniques confer higher accuracy at higher speedup. With these optimizations, we demonstrate a 9X reduction in cost with negligible loss of accuracy for the TIDIGITS audio digit recognition benchmark. Similarly, on the large Wall Street Journal (WSJ) speech recognition benchmark, pretrained networks can also be greatly accelerated as delta networks and trained delta networks show a 5.7x improvement with negligible loss of accuracy. Finally, on an end-to-end CNN-RNN network trained for steering angle prediction in a driving dataset, the RNN cost can be reduced by a substantial 100X. Daniel Neil, Junhaeng Lee, Tobi Delbruck, Shih-Chii Liu |
ICML | 3 |
| 2017 | Live demonstration: Event-driven real-time spoken digit recognition systemabstractSummary form only given. We previously described a deep network system that reached an accuracy of 82% on a digit recognition task using the spike outputs from a Dynamic Audio Sensor (DAS) in response to audio samples from the TIDIGITS database. The audio samples were played directly to the system therefore bypassing the microphones. This work presents an interactive real-time demonstration of this digit recognition system. The system classifies a spoken digit based on the output spikes of the DAS in response to digits spoken into the on-board microphones. Jithendar Anumula, Daniel Neil, Tobi Delbruck, Shih-Chii Liu |
ISCAS | 4 |
| 2017 | Block-matching optical flow for dynamic vision sensors: Algorithm and FPGA implementationabstractRapid and low power computation of optical flow (OF) is potentially useful in robotics. The dynamic vision sensor (DVS) event camera produces quick and sparse output, and has high dynamic range, but conventional OF algorithms are frame-based and cannot be directly used with event-based cameras. Previous DVS OF methods do not work well with dense textured input and are designed for implementation in logic circuits. This paper proposes a new block-matching based DVS OF algorithm which is inspired by motion estimation methods used for MPEG video compression. The algorithm was implemented both in software and on FPGA. For each event, it computes the motion direction as one of 9 directions. The speed of the motion is set by the sample interval. Results show that the Average Angular Error can be improved by 30% compared with previous methods. The OF can be calculated on FPGA with 50 MHz clock in 0.2 us per event (11 clock cycles), 20 times faster than a Java software implementation running on a desktop PC. Sample data is shown that the method works on scenes dominated by edges, sparse features, and dense texture. Min Liu 0031, Tobi Delbruck |
ISCAS | 2 |
| 2017 | Live demonstration: Convolutional neural network driven by dynamic vision sensor playing RoShamBoabstractThis demonstration presents a convolutional neural network (CNN) playing “RoShamBo” (“rock-paper-scissors”) against human opponents in real time. The network is driven by dynamic and active-pixel vision sensor (DAVIS) events, acquired by accumulating events into fixed event-number frames. Iulia-Alexandra Lungu, Federico Corradi, Tobi Delbruck |
ISCAS | 3 |
| 2017 | Color temporal contrast sensitivity in dynamic vision sensorsabstractThis paper introduces the first simulations and measurements of event data obtained from the first Dynamic and Active Vision Sensors (DAVIS) with RGBW color filters. The absolute quantum efficiency spectral responses of the RGBW photodiodes were measured, the behavior of the color-sensitive DVS pixels were simulated and measured, and reconstruction through color events interpolation was developed. Diederik Paul Moeys, Cheng-Han Li, Julien N. P. Martel, Simeon A. Bamford, Luca Longinotti, Vasyl Motsnyi, David San Segundo Bello, Tobi Delbruck |
ISCAS | 8 |
| 2016 | Combined frame- and event-based detection and trackingabstractThis paper reports an object tracking algorithm for a moving platform using the dynamic and active-pixel vision sensor (DAVIS). It takes advantage of both the active pixel sensor (APS) frame and dynamic vision sensor (DVS) event outputs from the DAVIS. The tracking is performed in a three step-manner: regions of interest (ROIs) are generated by a cluster-based tracking using the DVS output, likely target locations are detected by using a convolutional neural network (CNN) on the APS output to classify the ROIs as foreground and background, and finally a particle filter infers the target location from the ROIs. Doing convolution only in the ROIs boosts the speed by a factor of 70 compared with full-frame convolutions for the 240×180 frame input from the DAVIS. The tracking accuracy on a predator and prey robot database reaches 90% with a cost of less than 20ms/frame in Matlab on a normal PC without using a GPU. Diederik Paul Moeys, Gautham P. Das, Daniel Neil, Shih-Chii Liu, Tobi Delbruck |
ISCAS | 6 |
| 2016 | Retinal ganglion cell software and FPGA model implementation for object detection and trackingabstractThis paper describes the software and FPGA implementation of a Retinal Ganglion Cell model which detects moving objects. It is shown how this processing, in conjunction with a Dynamic Vision Sensor as its input, can be used to extrapolate information about object position. Software-wise, a system based on an array of these of RGCs has been developed in order to obtain up to two trackers. These can track objects in a scene, from a still observer, and get inhibited when saccadic camera motion happens. The entire processing takes on average 1000 ns/event. A simplified version of this mechanism, with a mean latency of 330 ns/event, at 50 MHz, has also been implemented in a Spartan6 FPGA. Diederik Paul Moeys, Tobi Delbruck, Antonio Rios-Navarro, Alejandro Linares-Barranco |
ISCAS | 2 |
| 2016 | Live demonstration: Retinal ganglion cell software and FPGA implementation for object detection and trackingabstractThis demonstration shows how object detection and tracking are possible thanks to a new implementation which takes inspiration from the visual processing of a particular type of ganglion cell in the retina. Diederik Paul Moeys, Tobi Delbruck, Antonio Rios-Navarro, Alejandro Linares-Barranco |
ISCAS | 2 |
| 2015 | Human vs. computer slot car racing using an event and frame-based DAVIS vision sensorabstractThis paper describes an open-source implementation of an event-based dynamic and active pixel vision sensor (DAVIS) for racing human vs. computer on a slot car track. The DAVIS is mounted in "eye-of-god" view. The DAVIS image frames are only used for setup and are subsequently turned off because they are not needed. The dynamic vision sensor (DVS) events are then used to track both the human and computer controlled cars. The precise control of throttle and braking afforded by the low latency of the sensor output enables consistent outperformance of human drivers at a laptop CPU load of <;3% and update rate of 666Hz. The sparse output of the DVS event stream results in a data rate that is about 1000 times smaller than from a frame-based camera with the same resolution and update rate. The scaled average lap speed of the 1/64 scale cars is about 450km/h which is twice as fast as the fastest Formula 1 lap speed. A feedbackcontroller mode allows competitive racing by slowing the computer controlled car when it is ahead of the human. In tests of human vs. computer racing the computer still won more than 80% of the races. Tobi Delbruck, Michael Pfeiffer 0001, R. Juston, Garrick Orchard, Elias Mueggler, Alejandro Linares-Barranco, M. W. Tilden |
ISCAS | 1 |
| 2015 | Design of an RGBW color VGA rolling and global shutter dynamic and active-pixel vision sensorabstractThis paper reports the design of a color dynamic and active-pixel vision sensor (C-DAVIS) for robotic vision applications. The C-DAVIS combines monochrome eventgenerating dynamic vision sensor pixels and 5-transistor active pixels sensor (APS) pixels patterned with an RGBW color filter array. The C-DAVIS concurrently outputs rolling or global shutter RGBW coded VGA resolution frames and asynchronous monochrome QVGA resolution temporal contrast events. Hence the C-DAVIS is able to capture spatial details with color and track movements with high temporal resolution while keeping the data output sparse and fast. The C-DAVIS chip is fabricated in TowerJazz 0.18um CMOS image sensor technology. An RGBW 2×2-pixel unit measures 20um × 20um. The chip die measures 8mm × 6.2mm. Cheng-Han Li, Christian Brandli, Raphael Berner, Minhao Yang, Shih-Chii Liu, Tobi Delbruck |
ISCAS | 7 |
| 2015 | A USB3.0 FPGA event-based filtering and tracking framework for dynamic vision sensorsabstractDynamic vision sensors (DVS) are frame-free sensors with an asynchronous variable-rate output that is ideal for hard real-time dynamic vision applications under power and latency constraints. Post-processing of the digital sensor output can reduce sensor noise, extract low level features, and track objects using simple algorithms that have previously been implemented in software. In this paper we present an FPGA-based framework for event-based processing that allows uncorrelated-event noise removal and real-time tracking of multiple objects, with dynamic capabilities to adapt itself to fast or slow and large or small objects. This framework uses a new hardware platform based on a Lattice FPGA which filters the sensor output and which then transmits the results through a super-speed Cypress FX3 USB microcontroller interface to a host computer. The packets of events and timestamps are transmitted to the host computer at rates of 10 Mega events per second. Experimental results are presented that demonstrate a low latency of 10us for tracking and computing the center of mass of a detected object. Alejandro Linares-Barranco, Francisco Gomez-Rodriguez, Vicente Villanueva, Luca Longinotti, Tobi Delbruck |
ISCAS | 5 |
| 2015 | Design of a spatiotemporal correlation filter for event-based sensorsabstractThis paper reports the design of a 1mW, 10ns-latency mixed signal system in 0.18μm CMOS which enables filtering out uncorrelated background activity in event-based neuromorphic sensors. Background activity (BA) in the output of dynamic vision sensors is caused by thermal noise and junction leakage current acting on switches connected to floating nodes in the pixels. The reported chip generates a pass flag for spatiotemporally correlated events for post-processing to reduce communication/computation load and improve information rate. A chip with 128×128 array with 20×20μm2cells has been designed. Each filter cell combines programmable spatial subsampling with a temporal window based on current integration. Power-gating is used to minimize the power consumption by only activating the threshold detection and communication circuits in the cell receiving an input event. This correlation filter chip targets embedded neuromorphic visual and auditory systems, where low average power consumption and low latency are critical. Christian Brandli, Cheng-Han Li, Shih-Chii Liu, Tobi Delbruck |
ISCAS | 5 |
| 2015 | Current-mode automated quality control cochlear resonator for bird identity taggingabstractThis paper describes a VLSI automatic quality control pitch detector circuit which can be used for detecting the identity of a unique bird. The detector is based on a previous VLSI model of the local gain control mechanism of the outer hair cells of the biological cochlea. This work presents characterization results from a 20-channel chip fabricated in a 4-metal 2-poly CMOS 0.35 μm technology with estimated dynamic range of 70 dB, power consumption of 825 nW per channel, frequency range covering 0.4-10 kHz and mean Q of 6.31. Results are shown for a pitch detection experiment with a tuned resonator. Diederik Paul Moeys, Tobi Delbruck, Shih-Chii Liu |
ISCAS | 2 |
| 2014 | Live demonstration: The "DAVIS" Dynamic and Active-Pixel Vision SensorabstractThis demonstration will show the features of the Dynamic and Active-Pixel Vision Sensor (DAVIS) reported at the VLSI Symposium and the International Imager Sensor Workshop in 2013. This sensor concurrently outputs conventional CMOS image sensor frames and sparse, low-latency dynamic vision sensor events from the same pixels, sharing the same photodiodes. The setup will allow visitors to explore the advantages of combining of fast and computationally-efficient neuromorphic event-driven vision with the existing body of methods for frame-based computer and machine vision. Christian Brandli, Raphael Berner, Minhao Yang, Shih-Chii Liu, V. Villeneuva, Tobi Delbruck |
ISCAS | 6 |
| 2014 | Real-time, high-speed video decompression using a frame- and event-based DAVIS sensorabstractDynamic and active pixel vision sensors (DAVISs) are a new type of sensor that combine a frame-based intensity readout with an event-based temporal contrast readout. This paper demonstrates that these sensors inherently perform high-speed, video compression in each pixel by describing the first decompression algorithm for this data. The algorithm performs an online optimization of the event decoding in real time. Example scenes were recorded by the 240×180 pixel sensor at sub-Hz frame rates and successfully decompressed yielding an equivalent frame rate of 2kHz. A quantitative analysis of the compression quality resulted in an average pixel error of 0.5DN intensity resolution for non-saturating stimuli. The system exhibits an adaptive compression ratio which depends on the activity in a scene; for stationary scenes it can go up to 1862. The low data rate and power consumption of the proposed video compression system make it suitable for distributed sensor networks. Christian Brandli, Lorenz Müller, Tobi Delbruck |
ISCAS | 3 |
| 2014 | Integration of dynamic vision sensor with inertial measurement unit for electronically stabilized event-based visionabstractNeuromorphic spike event-based dynamic vision sensors (DVS) offer the possibility of fast, computationally efficient visual processing for navigation in mobile robotics. To extract motion parallax cues relating to 3D scene structure, the uninformative camera rotation must be removed from the visual input to allow the un-blurred features and informative relative optical flow to be analyzed. Here we describe the integration of an inertial measurement unit (IMU) with a 240×180 pixel DVS. The algorithm for electronic stabilization of the visual input against camera rotation is described. Examples are presented showing the stabilization performance of the system. Tobi Delbruck, Vicente Villanueva, Luca Longinotti |
ISCAS | 1 |
| 2014 | 1kHz 2D silicon retina motion sensor platformabstractThis paper proposes an optical motion sensor aimed towards small robotic platforms. It incorporates a 20×20 pixel continuous-time CMOS silicon retina vision sensor with pixels that have local gain control and adapt to background lighting and a DSP microcontroller which computes the global optical flow from the sampled sensor output. The system allows the user to validate various motion algorithms suitable for the platform. Measurements are presented that show that the system can compute global 2D translational motion from complex natural scenes using the image interpolation algorithm at a sample rate of 1 kHz and for speeds up to ±1000 pixels/s using <5k instruction cycles per frame. Andreas Steiner 0001, Rico Moeckel, Reto Thurer, Dario Floreano, Tobi Delbruck, Shih-Chii Liu |
ISCAS | 5 |
| 2014 | Comparison of spike encoding schemes in asynchronous vision sensors: Modeling and designabstractTwo in-pixel encoding mechanisms to convert analog input to spike output for vision sensors are modeled and compared with the consideration of feedback delay: one is feedback and reset (FAR), and the other is feedback and subtract (FAS). MATLAB simulations of linear signal reconstruction from spike trains generated by the two encoders show that FAR in general has a lower signal-to-distortion ratio (SDR) compared to FAS due to signal loss during the reset phase and hold period, and the SDR merit of FAS increases as the quantization bit number and input signal frequency increases. A 500 μm2in-pixel circuit implementation of FAS using asynchronous switched capacitors in a UMC 0.18μm 1P6M process is described, and the post-layout simulation results are given to verify the FAS encoding mechanism. Minhao Yang, Shih-Chii Liu, Tobi Delbruck |
ISCAS | 3 |
| 2014 | Retinomorphic Event-Based Vision Sensors: Bioinspired Cameras With Spiking OutputabstractState-of-the-art image sensors suffer from significant limitations imposed by their very principle of operation. These sensors acquire the visual information as a series of “snapshot” images, recorded at discrete points in time. Visual information gets time quantized at a predetermined frame rate which has no relation to the dynamics present in the scene. Furthermore, each recorded frame conveys the information from all pixels, regardless of whether this information, or a part of it, has changed since the last frame had been acquired. This acquisition method limits the temporal resolution, potentially missing important information, and leads to redundancy in the recorded image data, unnecessarily inflating data rate and volume. Biology is leading the way to a more efficient style of image acquisition. Biological vision systems are driven by events happening within the scene in view, and not, like image sensors, by artificially created timing and control signals. Translating the frameless paradigm of biological vision to artificial imaging systems implies that control over the acquisition of visual information is no longer being imposed externally to an array of pixels but the decision making is transferred to the single pixel that handles its own information individually. In this paper, recent developments in bioinspired, neuromorphic optical sensing and artificial vision are presented and discussed. It is suggested that bioinspired vision systems have the potential to outperform conventional, frame-based vision systems in many application fields and to establish new benchmarks in terms of redundancy suppression and data compression, dynamic range, temporal resolution, and power efficiency. Demanding vision tasks such as real-time 3-D mapping, complex multiobject tracking, or fast visual feedback loops for sensory-motor action, tasks that often pose severe, sometimes insurmountable, challenges to conventional artificial vision systems, are in reach using bioinspired vision sensing and processing techniques. Christoph Posch, Teresa Serrano-Gotarredona, Bernabé Linares-Barranco, Tobi Delbruck |
Proc. IEEE | 4 |
| 2014 | Real-Time Gesture Interface Based on Event-Driven Processing From Stereo Silicon RetinasabstractWe propose a real-time hand gesture interface based on combining a stereo pair of biologically inspired event-based dynamic vision sensor (DVS) silicon retinas with neuromorphic event-driven postprocessing. Compared with conventional vision or 3-D sensors, the use of DVSs, which output asynchronous and sparse events in response to motion, eliminates the need to extract movements from sequences of video frames, and allows significantly faster and more energy-efficient processing. In addition, the rate of input events depends on the observed movements, and thus provides an additional cue for solving the gesture spotting problem, i.e., finding the onsets and offsets of gestures. We propose a postprocessing framework based on spiking neural networks that can process the events received from the DVSs in real time, and provides an architecture for future implementation in neuromorphic hardware devices. The motion trajectories of moving hands are detected by spatiotemporally correlating the stereoscopically verged asynchronous events from the DVSs by using leaky integrate-and-fire (LIF) neurons. Adaptive thresholds of the LIF neurons achieve the segmentation of trajectories, which are then translated into discrete and finite feature vectors. The feature vectors are classified with hidden Markov models, using a separate Gaussian mixture model for spotting irrelevant transition gestures. The disparity information from stereovision is used to adapt LIF neuron parameters to achieve recognition invariant of the distance of the user to the sensor, and also helps to filter out movements in the background of the user. Exploiting the high dynamic range of DVSs, furthermore, allows gesture recognition over a 60-dB range of scene illuminance. The system achieves recognition rates well over 90% under a variety of variable conditions with static and dynamic backgrounds with naïve users. Junhaeng Lee, Tobi Delbruck, Michael Pfeiffer 0001, Paul K. J. Park, Chang-Woo Shin, Hyunsurk Ryu, Byung-Chang Kang |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2013 | Low-latency localization by active LED markers tracking using a dynamic vision sensorabstractAt the current state of the art, the agility of an autonomous flying robot is limited by its sensing pipeline, because the relatively high latency and low sampling frequency limit the aggressiveness of the control strategies that can be implemented. To obtain more agile robots, we need faster sensing pipelines. A Dynamic Vision Sensor (DVS) is a very different sensor than a normal CMOS camera: rather than providing discrete frames like a CMOS camera, the sensor output is a sequence of asynchronous timestamped events each describing a change in the perceived brightness at a single pixel. The latency of such sensors can be measured in the microseconds, thus offering the theoretical possibility of creating a sensing pipeline whose latency is negligible compared to the dynamics of the platform. However, to use these sensors we must rethink the way we interpret visual data. This paper presents a method for low-latency pose tracking using a DVS and Active Led Markers (ALMs), which are LEDs blinking at high frequency (>1 KHz). The sensor's time resolution allows distinguishing different frequencies, thus avoiding the need for data association. This approach is compared to traditional pose tracking based on a CMOS camera. The DVS performance is not affected by fast motion, unlike the CMOS camera, which suffers from motion blur. Andrea Censi, Jonas Strubel, Christian Brandli, Tobi Delbruck, Davide Scaramuzza 0001 |
IROS | 4 |
| 2012 | Touchless hand gesture UI with instantaneous responsesabstractIn this paper we present a simple technique for real-time touchless hand gesture user interface (UI) for mobile devices based on a biologically inspired vision sensor, the dynamic vision sensor (DVS). The DVS can detect a moving object in a fast and cost effective way by outputting events asynchronously on edges of the object. The output events are spatiotemporally correlated by using novel event-driven processing algorithms based on leaky integrate-and-fire neurons to track a finger tip or to infer directions of hand swipe motions. The experimental results show that the proposed technique can achieve graphic UI capable finger tip tracking with milliseconds intervals and accurate hand swipe motion detection with negligible latency. Junhaeng Lee, Paul K. J. Park, Chang-Woo Shin, Hyunsurk Ryu, Byung-Chang Kang, Tobi Delbruck |
ICIP | 6 |
| 2012 | Live demonstration: Behavioural emulation of event-based vision sensorsabstractThis demonstration shows how an inexpensive high frame-rate USB camera is used to emulate existing and proposed activity-driven event-based vision sensors. A PS3-Eye camera which runs at a maximum of 125 frames/second with colour QVGA (320×240) resolution is used to emulate several event-based vision sensors, including a Dynamic Vision Sensor (DVS), a colour-change sensitive DVS (cDVS), and a hybrid vision sensor with DVS+cDVS pixels. The emulator is integrated into the jAER software project for event-based real-time vision and is used to study use cases for future vision sensor designs. Matthew L. Katz, Konstantin Nikolic, Tobi Delbruck |
ISCAS | 3 |
| 2012 | Live demonstration: Gesture-based remote control using stereo pair of dynamic vision sensorsabstractThis demonstration shows a natural gesture interface for console entertainment devices using as input a stereo pair of dynamic vision sensors. The event-based processing of the sparse sensor output allows fluid interaction at a laptop processor load of less than 3%. Junhaeng Lee, Tobi Delbruck, Paul K. J. Park, Michael Pfeiffer 0001, Chang-Woo Shin, Hyunsurk Ryu, Byung-Chang Kang |
ISCAS | 2 |
| 2012 | Real-time speaker identification using the AEREAR2 event-based silicon cochleaabstractThis paper reports a study on methods for real-time speaker identification using the output from an event-based silicon cochlea. These methods are evaluated based on the amount of computation that needs to be performed and the classification performance in a speaker identification task. It uses the binaural AEREAR2 silicon cochlea, with 64 frequency channels and 512 output neurons. Auditory features representing fading histograms of inter-spike intervals and channel activity distributions are extracted from the cochlea spikes. These feature vectors are then classified by a linear Support Vector Machine, which is trained against a subset of 40 speakers (20/20 male/female) from the TIMIT database. Speakers are correctly identified at >90% accuracy during each sentence utterance and with an average latency of 700±200ms from the start of the sentence. Cheng-Han Li, Tobi Delbruck, Shih-Chii Liu |
ISCAS | 2 |
| 2012 | Addressable current reference array with 170dB dynamic rangeabstractConfigurable high-performance bias current reference circuits are useful in complex mixed-signal chips. This paper presents the design of a configurable current reference array with ultra wide dynamic range (DR). A coarse-fine architecture using octal coarse current spacing and 8 bits of fine resolution increases the overall current DR with less area compared with the prior work. Compact current multipliers and dividers also save chip areas. Shifted-source current mirrors and an off-current suppression technique improve the accuracy of generated low currents. A buffer with dual-threshold source followers is used to generate the output biasing voltage with a wide DR input current. Biases are individually addressable and configurable. Measurement results of this design in UMC 0.18μm 1P6M CMOS process suggest that over 170dB DR is achieved at room temperature. Each additional bias occupies an incremental area of 360×22μm2, which is smaller by a factor of 4 compared to the previous design. Minhao Yang, Shih-Chii Liu, Cheng-Han Li, Tobi Delbruck |
ISCAS | 4 |
| 2012 | Asynchronous Event-Based Binocular Stereo MatchingabstractWe present a novel event-based stereo matching algorithm that exploits the asynchronous visual events from a pair of silicon retinas. Unlike conventional frame-based cameras, recent artificial retinas transmit their outputs as a continuous stream of asynchronous temporal events, in a manner similar to the output cells of the biological retina. Our algorithm uses the timing information carried by this representation in addressing the stereo-matching problem on moving objects. Using the high temporal resolution of the acquired data stream for the dynamic vision sensor, we show that matching on the timing of the visual events provides a new solution to the real-time computation of 3-D objects when combined with geometric constraints using the distance to the epipolar lines. The proposed algorithm is able to filter out incorrect matches and to accurately reconstruct the depth of moving objects despite the low spatial resolution of the sensor. This brief sets up the principles for further event-based vision processing and demonstrates the importance of dynamic information and spike timing in processing asynchronous streams of visual events. Paul Rogister, Ryad Benosman, Sio-Hoi Ieng, Patrick Lichtsteiner, Tobi Delbruck |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2011 | Confession session: Learning from others mistakesabstractPeople rarely put in their papers the things that didn't work, the mistakes they made, and how they found out what went wrong. Such confessions can help others learn how to avoid similar mistakes. Twenty-six confessions were collected to form the bulk of this paper. Themes that arise are errors that result from not understanding the limitations of simulation tools in modeling physical reality, chip verification errors that result from lack of clear communication between designers, and projects that are considered in their own isolated environment of technical challenges rather than the broader context of their environment or application. Pamela Abshire, Amine Bermak, Raphael Berner, Gert Cauwenberghs, Shoushun Chen, Jennifer Blain Christen, Timothy G. Constandinou, Eugenio Culurciello, Marc Dandin, Timir Datta, Tobi Delbruck, Piotr Dudek, Amir Eftekhar, Ralph Etienne-Cummings, Giacomo Indiveri, Matthew K. Law, Bernabé Linares-Barranco, Jonathan Tapson, Wei Tang 0002, Yiming Zhai |
ISCAS | 11 |
| 2010 | Event-based color change pixel in standard CMOSabstractThis paper describes a novel dichromatic spiking pixel circuit that reacts to color change but not to intensity change. It is built in standard CMOS using a buried double junction to sense wavelength information. The pixel can detect light wavelength changes of about 14nm, while not responding to intensity steps of at least a factor of three. The pixel is suitable for integration into an array and can easily be combined with a temporal log intensity contrast change pixel. Raphael Berner, Tobi Delbruck |
ISCAS | 2 |
| 2010 | Temporal contrast AER pixel with 0.3%-contrast event thresholdabstractBioluminescence analysis methods require measurement of small temporal variations of scene brightness over a 2d spatial array. Conventional solutions use low-noise frame-based image sensors and high resolution ADCs to achieve the required sensitivity to small fluctuations of less than 1%. Here we report a pixel design for address event representation (AER) temporal contrast detection that is optimized for detecting small relative changes of intensity. Detected changes exceeding a threshold asynchronously generate events. The pixel uses two successive gain stages to memorize and amplify changes of intensity, allowing detection of changes as small as 0.3% of intensity; a factor of about 50 better than prior capabilities for event-based temporal contrast sensors. Tobi Delbruck, Raphael Berner |
ISCAS | 1 |
| 2010 | 32-bit Configurable bias current generator with sub-off-current capabilityabstractA fully configurable bias current reference is described. The output of the current reference is a gate voltage which produces a desired current. For each daisy-chained bias, 32 bits of configuration are divided into 22 bits of bias current, 6 bits of active-mirror buffer current, and 4 bits of other configuration. Configuration of each bias allows specifying the type of transistor (nfet or pfet), whether the bias is enabled or weakly pulled to the rail, whether the bias is for a cascode, and whether the bias transistor uses a shifted source (SS) voltage for sub-off-current biasing. In addition, the current reference integrates a pair of voltage regulators that generate stable voltage sources near the rails, suitable for the SS current references. Measurements from fabricated current references built in 180 nm CMOS show that the reference achieves at least 110 dB (22-bit) dynamic range and reaches 160dB when power-rail gate biasing is included. Generated bias currents reach at least 30x smaller current than the transistor off-current. Each current reference occupies an area of 620 × 50 um2. The design kit schematics and layout are open-sourced. Tobi Delbruck, Raphael Berner, Patrick Lichtsteiner, Carlos Dualibe |
ISCAS | 1 |
| 2010 | Fully integrated 500uW speech detection wake-up circuitabstractSpeech analysis requires substantial computation. It is desirable to run this analysis only when needed and at other times to go to a low power state. Here we propose a self-biased low power speech detection wake up circuit which interfaces directly to standard electret microphones. The speech detector includes a microphone preamplifier, a power extraction squaring circuit, a bandpass filter passing power of the modulation spectrum in the speech band from 2-12 Hz, a half rectifier which extracts this phoneme band power, and a PFM silicon neuron which emits spikes indicating phoneme-rate modulation of the audio spectrum. The output of the speech detector circuit is an asynchronous stream of digital spikes at a rate of 1Hz to 20Hz whose temporal structure indicates the presence of speech. A subsequent conventional processor will go to sleep between spikes and only wake up for full power speech analysis when the temporal structure indicates speech. The circuit is built in 1.6um 2P-2M CMOS and consumes 500uW with a 3V supply when attached to a standard electret microphone. Tobi Delbruck, Raphael Berner, Hynek Hermansky |
ISCAS | 1 |
| 2010 | Activity-driven, event-based vision sensorsabstractThe four chips presented in the special session on "Activity-driven, event-based vision sensors" quickly output compressed digital data in the form of events. These sensors reduce redundancy and latency and increase dynamic range compared with conventional imagers. The digital sensor output is easily interfaced to conventional digital post processing, where it reduces the latency and cost of post processing compared to imagers. The asynchronous data could spawn a new area of DSP that breaks from conventional Nyquist rate signal processing. This paper reviews the rationale and history of this event-based approach, introduces sensor functionalities, and gives an overview of the papers in this session. The paper concludes with a brief discussion on open questions. Tobi Delbruck, Bernabé Linares-Barranco, Eugenio Culurciello, Christoph Posch |
ISCAS | 1 |
| 2010 | Event-based 64-channel binaural silicon cochlea with Q enhancement mechanismsabstractThis paper describes an event-based binaural silicon cochlea aimed at spatial audition and auditory scene analysis. The chip has a matched pair of 64-stage cascaded analog second-order filter banks with 512 pulse-frequency modulated (PFM) address-event representation (AER) outputs. The spectral selectivity is sharpened through 2 different on-chip methods: an on-chip local Q DAC and an on-chip spatial sharpening through nearest neighbour lateral inhibition. The fabricated chip in a 4-metal 2-poly 0.35um CMOS process consumes peak 25mW power for the digital circuits and 33mW for the analog core. Dynamic range to produce PFM output is 36dB (25mVpp to 1500mVpp at microphone preamp output). Event timing jitter is 2us for 250mVpp input. The peak output bandwidth is 10M events per second (eps) but typical speech scenarios show rates of 20keps. Shih-Chii Liu, André van Schaik, Bradley A. Minch, Tobi Delbruck |
ISCAS | 4 |
| 2009 | A Pencil Balancing Robot using a Pair of AER Dynamic Vision SensorsabstractBalancing a normal pencil on its tip requires rapid feedback control with latencies on the order of milliseconds. This demonstration shows how a pair of spike-based silicon retina dynamic vision sensors (DVS) is used to provide fast visual feedback for controlling an actuated table to balance an ordinary pencil. Two DVSs view the pencil from right angles. Movements of the pencil cause spike address-events (AEs) to be emitted from the DVSs. These AEs are transmitted to a PC over USB interfaces and are processed procedurally in real time. The PC updates its estimate of the pencil's location and angle in 3d space upon each incoming AE, applying a novel tracking method based on spike-driven fitting to a model of the vertical shape of the pencil. A PD-controller adjusts X-Y-position and velocity of the table to maintain the pencil balanced upright. The controller also minimizes the deviation of the pencil's base from the center of the table. The actuated table is built using ordinary high-speed hobby servos which have been modified to obtain feedback from linear position encoders via a microcontroller. Our system can balance any small, thin object such as a pencil, pen, chop-stick, or rod for many minutes. Balancing is only possible when incoming AEs are processed as they arrive from the sensors, typically at intervals below millisecond ranges. Controlling at normal image sensor sample rates (e.g. 60 Hz) results in too long latencies for a stable control loop. Jörg Conradt, Matthew Cook 0001, Raphael Berner, Patrick Lichtsteiner, Rodney J. Douglas, Tobi Delbruck |
ISCAS | 6 |
| 2009 | Live Demonstration: A Pencil Balancing Robot using a Pair of AER Dynamic Vision SensorsabstractBalancing a normal pencil on its tip requires rapid feedback control with latencies on the order of milliseconds. This demonstration shows how a pair of spike-based silicon retina dynamic vision sensors (DVS) is used to provide fast visual feedback for controlling an actuated table to balance an ordinary pencil. Two DVSs view the pencil from right angles. Movements of the pencil cause spike address-events (AEs) to be emitted from the DVSs. These AEs are transmitted to a PC over USB interfaces and are processed procedurally in real time. The PC updates its estimate of the pencil's location and angle in 3d space upon each incoming AE, applying a novel tracking method based on spike-driven fitting to a model of the vertical shape of the pencil. A PD-controller adjusts X-Y-position and velocity of the table to maintain the pencil balanced upright. The controller also minimizes the deviation of the pencil's base from the center of the table. The actuated table is built using ordinary high-speed hobby servos which have been modified to obtain feedback from linear position encoders via a microcontroller. Our system can balance any small, thin object such as a pencil, pen, chop-stick, or rod for many minutes. Balancing is only possible when incoming AEs are processed as they arrive from the sensors, typically at intervals below millisecond ranges. Controlling at normal image sensor sample rates (e.g. 60 Hz) results in too long latencies for a stable control loop. Jörg Conradt, Matthew Cook 0001, Raphael Berner, Patrick Lichtsteiner, Rodney J. Douglas, Tobi Delbruck |
ISCAS | 6 |
| 2009 | Implementation of a Time-warping AER MapperabstractIn recent implementations of neuromorphic spike-based sensors, multi-neuron processors, and actuators; the spike traffic between devices is coded in the form of asynchronous spike streams following the address-event-representation protocol. This spike information can be modified during the transmission from one device to another by using a mapper device. In this paper we present a mapper implementation which transforms event addresses and can also delay events in time. We discuss two different architectures for implementing the time delays on an FPGA board (USB-AER), and we present an example of the use of the time delay feature in the mapper in an implementation of a visual elementary motion detection model based on the spike outputs of a temporal contrast retina. Alejandro Linares-Barranco, Francisco Gomez-Rodriguez, Gabriel Jiménez-Moreno, Tobi Delbruck, Raphael Berner, Shih-Chii Liu |
ISCAS | 4 |
| 2009 | Computing Spike-based Convolutions on GPUsabstractIn spiking neural networks, asynchronous spike events are processed in parallel by neurons. Emulations of such networks are traditionally computed by CPUs or realized using dedicated neuromorphic hardware. In many neuromorphic systems, the Address-Event-Representation (AER) is used for spike communication. In this paper we present the acceleration of AER based spike processing using a Graphics Processing Unit (GPU). In our experiment we interface a 128×128 pixel AER vision sensor to a spiking neural network implemented on a GPU for real-time convolution-based nonlinear feature extraction with convolution kernel sizes ranging from 48×48 to 112×112 pixels. We show parallelism-performance trade-offs on GPUs for single spike per thread, multiple spikes per thread, and multiple objects parallelism techniques. Our implementation can achieve a kernel speedup of up to 35× on a single NVIDIA GTX280 board when compared to a CPU-only implementation. Jayram Moorkanikara Nageswaran, Nikil Dutt, Tobi Delbruck |
ISCAS | 4 |
| 2009 | Live Demonstration: Computing Spike-based Convolutions on GPUsabstractThis demonstration shows the first implementation of a real-time spike-based convolution processing system which combines a spike based dynamic vision sensor (DVS) with parallel graphics processor unit (GPU) computation. Moving objects with different features (shape and size) are presented to the system. In the first demo, the system responses in real time to recognize and keep track of one user specified object and ignore the others. In the second one, the system concurrently extracts several features, and labels the outputs with different colors. Users will enjoy the real-time response and learn about using spike-based sensors combined with conventional procedural processing. Jayram Moorkanikara Nageswaran, Nikil Dutt, Tobi Delbruck |
ISCAS | 4 |
| 2009 | Getting to Know Your Neighbors: Unsupervised Learning of Topography from Real-World, Event-Based InputabstractBiological neural systems must grow their own connections and maintain topological relations between elements that are related to the sensory input surface. Artificial systems have traditionally prewired such maps, but the sensor arrangement is not always known and can be expensive to specify before run time. Here we present a method for learning and updating topographic maps in systems comprising modular, event-based elements. Using an unsupervised neural spike-timing-based learning rule combined with Hebbian learning, our algorithm uses the spatiotemporal coherence of the external world to train its network. It improves on existing algorithms by not assuming a known topography of the target map and includes a novel method for automatically detecting edge elements. We show how, for stimuli that are small relative to the sensor resolution, the temporal learning window parameters can be determined without using any user-specified constants. For stimuli that are larger relative to the sensor resolution, we provide a parameter extraction method that generally outperforms the small-stimulus method but requires one user-specified constant. The algorithm was tested on real data from a 64 x 64-pixel section of an event-based temporal contrast silicon retina and a 360-tile tactile luminous floor. It learned 95.8% of the correct neighborhood relations for the silicon retina within about 400 seconds of real-world input from a driving scene and 98.1% correct for the sensory floor after about 160 minutes of human pedestrian traffic. Residual errors occurred in regions receiving little or ambiguous input, and the learned topological representations were able to update automatically in response to simulated damage. Our algorithm has applications in the design of modular autonomous systems in which the interfaces between components are learned during operation rather than at design time. Martin Boerlin, Tobi Delbruck, Kynan Eng |
Neural Comput. | 2 |
| 2009 | CAVIAR: A 45k Neuron, 5M Synapse, 12G Connects/s AER Hardware Sensory-Processing- Learning-Actuating System for High-Speed Visual Object Recognition and TrackingabstractThis paper describes CAVIAR, a massively parallel hardware implementation of a spike-based sensing-processing-learning-actuating system inspired by the physiology of the nervous system. CAVIAR uses the asychronous address-event representation (AER) communication framework and was developed in the context of a European Union funded project. It has four custom mixed-signal AER chips, five custom digital AER interface components, 45k neurons (spiking cells), up to 5M synapses, performs 12G synaptic operations per second, and achieves millisecond object recognition and tracking latencies. Rafael Serrano-Gotarredona, Matthias Oster, Patrick Lichtsteiner, Alejandro Linares-Barranco, Rafael Paz-Vicente, Francisco Gomez-Rodriguez, Luis A. Camuñas-Mesa, Raphael Berner, Manuel Rivas Pérez, Tobi Delbruck, Shih-Chii Liu, Rodney J. Douglas, Philipp Häfliger, Gabriel Jiménez-Moreno, Antonio Abad Civit Balcells, Teresa Serrano-Gotarredona, Antonio J. Acosta 0001, Bernabé Linares-Barranco |
IEEE Trans. Neural Networks | 10 |
| 2008 | Self-timed vertacolor dichromatic vision sensor for low power pattern detectionabstractThis paper proposes a simple focal plane pattern detector architecture using a novel pixel sensor based on the dichromatic vertacolor structure. Additionally, the sensor transfers dichromatic intensity values using a self-timed time-to- first-spike scheme, which provides high dynamic range imaging. The intensity information is transmitted using the address event representation protocol. The spectral information is sampled automatically at each intensity reading in a ratioed way that maintains high dynamic range. A test chip consisting of 20 pixels has been fabricated in 1.5 um 2P 2M CMOS and characterized. The combined pattern detector/ imager core consumes 45 uA at 5 V supply voltage. Raphael Berner, Patrick Lichtsteiner, Tobi Delbruck |
ISCAS | 3 |
| 2008 | Fall detection using an address-event temporal contrast vision sensorabstractIn this paper we describe an address-event vision system designed to detect accidental falls in elderly home care applications. The system raises an alarm when a fall hazard is detected. We use an asynchronous temporal contrast vision sensor which features sub-millisecond temporal resolution. A lightweight algorithm computes an instantaneous motion vector and reports fall events. We are able to distinguish fall events from normal human behavior, such as walking, crouching down, and sitting down. Our system is robust to the monitored person’s spatial position in a room and presence of pets. Zhengming Fu, Eugenio Culurciello, Patrick Lichtsteiner, Tobi Delbruck |
ISCAS | 4 |
| 2007 | A 5 Meps $100 USB2.0 Address-Event Monitor-Sequencer InterfaceabstractThis paper describes a high-speed USB2.0 address-event representation (AER) interface that allows simultaneous monitoring and sequencing of precisely timed AER data. This low-cost (<$100), two chip, bus powered interface can achieve sustained AER event rates of 5 megaevents per second (Meps). Several boards can be electrically synchronized, allowing simultaneous synchronized capture from multiple devices. It has three parallel AER ports, one for sequencing, one for monitoring and one for passing through the monitored events. This paper also describes the host software infrastructure that makes the board usable for a heterogeneous mixture of AER devices and that allows recording and playback of recorded data. Raphael Berner, Tobi Delbruck, Antonio Abad Civit Balcells, Alejandro Linares-Barranco |
ISCAS | 2 |
| 2007 | Fast sensory motor control based on event-based hybrid neuromorphic-procedural systemabstractFast sensory-motor processing is challenging when using traditional frame-based cameras and computers. Here the authors show how a hybrid neuromorphic-procedural system consisting of an address-event silicon retina, a computer, and a servo motor can be used to implement a fast sensory-motor reactive controller to track and block balls shot at a goal. The system consists of a 128times128 retina that asynchronously reports scene reflectance changes, a laptop PC, and a servo motor controller. Components are interconnected by USB. The retina looks down onto the field in front of the goal. Moving objects are tracked by an event-driven cluster tracker algorithm that detects the ball as the nearest object that is approaching the goal. The ball's position and velocity are used to control the servo motor. Running under Windows XP, the reaction latency is 2.8plusmn0.5 ms at a CPU load of1 million events per second (Meps), although fast balls only create ~30 keps. This system demonstrates the advantages of hybrid event-based sensory motor processing Tobi Delbruck, Patrick Lichtsteiner |
ISCAS | 1 |
| 2007 | Dichromatic spectral measurement circuit in vanilla CMOSabstractThe circuit described in this paper uses a "verta-color" stacked two-diode structure to measure relative long and short wavelength spectral content. The p-type source-drain to nwell forms the top diode and the nwell-psubstrate diode forms the bottom diode. The circuit output is a digital PWM signal whose frequency encodes absolute intensity and whose duty cycle encodes the relative photodiode current. This signal is formed by a self-timed circuit that alternately discharges the top and bottom photodiodes. This circuit was fabricated in a standard 3M 2P 0.5μm CMOS process. Monochromatic stimulation shows that the duty cycle varies between 50% and 7% as the photon wavelength is varied between 400 nm to 750 nm. The output frequency is 150 Hz at incident irradiance of 1.7 W/m2. Chip-to-chip variation of PWM duty cycle and frequency is about 1% measured over 5 chips. Power consumption is 20μW. A modified version of this circuit could form the basis for simple color vision sensors built in widely-available vanilla CMOS. Daniel Bernhard Fasnacht, Tobi Delbruck |
ISCAS | 2 |
| 2007 | Using FPGA for visuo-motor control with a silicon retina and a humanoid robotabstractThe address-event representation (AER) is a neuromorphic communication protocol for transferring asynchronous events between VLSI chips. The event information is transferred using a high speed digital parallel bus. This paper present an experiment based on AER for visual sensing, processing and finally actuating a robot. The AER output of a silicon retina is processed by an AER filter implemented into a FPGA to produce a mimicking behaviour in a humanoid robot (The RoboSapiens V2). We have implemented the visual filter into the Spartan II FPGA of the USB-AER platform and the central pattern generator (CPG) into the Spartan 3 FPGA of the AER-Robot platform, both developed by authors. Alejandro Linares-Barranco, Francisco Gomez-Rodriguez, Angel Jiménez-Fernandez, Tobi Delbruck, P. Lichtensteiner |
ISCAS | 4 |
| 2007 | A Spike-Based Saccadic Recognition SystemabstractThe paper presents a spike-based saccadic recognition system that uses a temporal-derivative silicon retina on a pan-tilt unit and an aVLSI multi-neuron classifier with a time-to-first-spike output coding. By using the spike information during the last 150 ms of a saccadic movement, we generate a reliable, sparse stimulus representation of image patches. The paper describes a novel classification scheme where the retinal spikes during this time influence the time-to-first spike of classifier neurons which receive the same constant input current. The preferred pattern of the neuron is stored in the synaptic connectivity between the retina and the classifier neuron. The authors demonstrates the robustness and real-time performance of this recognition scheme on a saccadic system which uses analog VLSI components. Matthias Oster, Patrick Lichtsteiner, Tobi Delbruck, Shih-Chii Liu |
ISCAS | 3 |
| 2006 | Modeling orientation selectivity using a neuromorphic multi-chip systemabstractThe growing interest in pulse-mode processing by neural networks is encouraging the development of hardware implementations of massively parallel, distributed networks of integrate-and-fire (I&F) neurons. We have developed a reconfigurable multi-chip neuronal system for modeling feature selectivity and applied it to oriented visual stimuli. Our system comprises a temporally differentiating imager and a VLSI competitive network of neurons which use an asynchronous address event representation (AER) for communication. Here we describe the overall system, and present experimental data demonstrating the effect of recurrent connectivity on the pulse-based orientation selectivity Elisabetta Chicca, Patrick Lichtsteiner, Tobi Delbruck, Giacomo Indiveri, Rodney J. Douglas |
ISCAS | 3 |
| 2006 | Fully programmable bias current generator with 24 bit resolution per biasabstractThis paper describes an on-chip programmable bias current generator, intended for mixed signal chips requiring a wide ranging set of currents. The individual generators share a master current reference. A serial digital interface to the chip controls the biases by bits loaded into a 24-bit shift register. These bits control the steering of current from a current splitter. The summed current splitter output is actively mirrored to a broadcasted bias voltage. Measurements from an implementation in 0.35/spl mu/ 4M-2P CMOS show a total range of bias current of over 6 decades (>120 dB) ranging from a few times the off-current up to the master reference current. For currents larger than the minimum, the generator has resolution spanning nearly its full 24 bit range (144dB), e.g. for a master current of 10 /spl mu/A, any bias current can be varied by as little as 0.5 pA with the caveat that the code is not guaranteed monotonic. Each bias occupies an area of 0.026 mm/sup 2/, which is about 65% of the bonding pad that it replaces. Measured variation in generated currents is <10% in strong inversion and about 20-30% in weak inversion. Tobi Delbruck, Patrick Lichtsteiner |
ISCAS | 1 |
| 2006 | A 100dB dynamic range high-speed dual-line optical transient sensor with asynchronous readoutabstractWe present a 100dB dynamic range 2times64 pixel dual-line optical sensor with asynchronous event-based readout. Each individual pixel of the sensor operates autonomously and responds with <100mus latency to relative intensity changes. It operates largely independent of overall scene illumination, directly encodes object reflectance, and greatly reduces redundancy while preserving precise timing information. The line sensor was fabricated using a 0.35mum standard CMOS technology. Results of the measurement performance of the sensor are presented. The intended application area is precision timing measurement under variable lighting conditions Patrick Lichtsteiner, Tobi Delbruck, Christoph Posch |
ISCAS | 2 |
| 2005 | AER Building Blocks for Multi-Layer Multi-Chip Neuromorphic Vision SystemsabstractA 5-layer neuromorphic vision processor whose components communicate spike events asychronously using the address-event- representation (AER) is demonstrated. The system includes a retina chip, two convolution chips, a 2D winner-take-all chip, a delay line chip, a learning classifier chip, and a set of PCBs for computer interfacing and address space remappings. The components use a mixture of analog and digital computation and will learn to classify trajectories of a moving object. A complete experimental setup and measurements results are shown. Rafael Serrano-Gotarredona, Matthias Oster, Patrick Lichtsteiner, Alejandro Linares-Barranco, Rafael Paz-Vicente, Francisco Gomez-Rodriguez, Håvard Kolle Riis, Tobi Delbruck, Shih-Chii Liu, S. Zahnd, Adrian M. Whatley, Rodney J. Douglas, Philipp Häfliger, Gabriel Jiménez-Moreno, Antonio Abad Civit Balcells, Teresa Serrano-Gotarredona, Antonio J. Acosta 0001, Bernabé Linares-Barranco |
NIPS | 8 |
| 2003 | Ada -intelligent space: an artificial creature for the swiss Expo.02abstractAda is an entertainment exhibit that is able to interact with many people simultaneously, using a language of light and sound. "She " received 553,700 visitors over 5 months during the Swiss Expo.02 in 2002. In this paper we present the broad motivations, design and technologies behind Ada, and a first overview of the outcomes of the exhibit. Kynan Eng, Andreas Bäbler, Ulysses Bernardet, Mark Blanchard, Márcio O. Costa, Tobi Delbruck, Rodney J. Douglas, Klaus Hepp, David Klein 0002, Jônatas Manzolli 0001, Matti Mintz, Fabian Roth, Ueli Rutishauser, Klaus Wassermann, Adrian M. Whatley, Aaron Wittmann, Reto Wyss, Paul F. M. J. Verschure |
ICRA | 6 |
| 2003 | Ada: a Playful Interactive Space
Tobi Delbruck, Kynan Eng, Andreas Bäbler, Ulysses Bernardet, Mark Blanchard, Adam Briska, Márcio O. Costa, Rodney J. Douglas, Klaus Hepp, David Klein 0002, Jônatas Manzolli 0001, Matti Mintz, Fabian Roth, Ueli Rutishauser, Klaus Wassermann, Aaron Wittmann, Adrian M. Whatley, Reto Wyss, Paul F. M. J. Verschure |
INTERACT | 1 |
| 2002 | Ada: constructing a synthetic organismabstractDespite immense progress in neuroscience, we remain restricted in our ability to construct autonomous behaving robots that match the competence of even simple animals. The barriers to the realisation of this goal include: the lack of knowledge of system integration issues, engineering limitations and organisational constraints common to many research laboratories. In this paper we describe our approach to addressing these issues by constructing an artificial organism within the framework of the Ada project - a large-scale public exhibit for the Swiss Expo.02 national exhibition. Kynan Eng, Andreas Bäbler, Ulysses Bernardet, Mark Blanchard, Adam Briska, Jörg Conradt, Márcio O. Costa, Tobi Delbruck, Rodney J. Douglas, Klaus Hepp, David Klein 0002, Jônatas Manzolli 0001, Matti Mintz, Thomas Netter, Fabian Roth, Ueli Rutishauser, Klaus Wassermann, Adrian M. Whatley, Aaron Wittmann, Reto Wyss, Paul F. M. J. Verschure |
IROS | 8 |
| 2001 | Orientation-Selective aVLSI Spiking NeuronsabstractWe describe a programmable multi-chip VLSI neuronal system that can be used for exploring spike-based information processing models. The system consists of a silicon retina, a PIC microcontroller, and a transceiver chip whose integrate-and-fire neurons are connected in a soft winner-take-all architecture. The circuit on this multi-neuron chip ap- proximates a cortical microcircuit. The neurons can be configured for different computational properties by the virtual connections of a se- lected set of pixels on the silicon retina. The virtual wiring between the different chips is effected by an event-driven communication pro- tocol that uses asynchronous digital pulses, similar to spikes in a neu- ronal system. We used the multi-chip spike-based system to synthe- size orientation-tuned neurons using both a feedforward model and a feedback model. The performance of our analog hardware spiking model matched the experimental observations and digital simulations of continuous-valued neurons. The multi-chip VLSI system has advantages over computer neuronal models in that it is real-time, and the computa- tional time does not scale with the size of the neuronal network. Shih-Chii Liu, Jörg Kramer, Giacomo Indiveri, Tobi Delbruck, Rodney J. Douglas |
NIPS | 4 |
| 2001 | Orientation-selective aVLSI spiking neurons
Shih-Chii Liu, Jörg Kramer, Giacomo Indiveri, Tobi Delbruck, Thomas Burg, Rodney J. Douglas |
Neural Networks | 4 |
| 2000 | Silicon retina for autofocusabstractThis paper describes a silicon retina that measures image sharpness. The idea is to use this sensor in the image plane of an autofocusing camera system. The chip has 25/spl times/26 pixels, each (60 /spl mu/m)/sup 2/, is fabricated on a (2.2 mm)/sup 2/ 1.2 /spl mu/m CMOS process, and consumes 100 /spl mu/A with 5 V supply. Tobi Delbruck |
ISCAS | 1 |
| 1994 | Adaptive Photoreceptor with Wide Dynamic RangeabstractWe describe a photoreceptor circuit that can be used in massively parallel analog VLSI silicon chips, in conjunction with other local circuits, to perform initial analog visual information processing. The receptor provides a continuous-time output that has low gain for static signals (including circuit mismatches), and high gain for transient signals that are centered around the adaptation point. The response is logarithmic, which makes the response to a fixed image contrast invariant to absolute light intensity. The 5-transistor receptor can be fabricated in an area of about 70 /spl mu/m by 70 /spl mu/m in a 2-/spl mu/m single-poly CMOS technology. It has a dynamic range of 1-2 decades at a single adaptation level, and a total dynamic range of more than 6 decades. Several technical improvements in the circuit yield an additional 1-2 decades dynamic range over previous designs without sacrificing signal quality. The lower limit of the dynamic range, defined arbitrarily as the illuminance at which the bandwidth of the receptor is 60 Hz, is at approximately 1 lux, which is the border between rod and cone vision and also the limit of current consumer video cameras. The receptor uses an adaptive element that is resistant to excess minority carrier diffusion. The continuous and logarithmic transduction process makes the bandwidth scale with intensity. As a result, the total AC RMS receptor noise is constant, independent of intensity. The spectral density of the noise is within a factor of two of pure photon shot noise and varies inversely with intensity. The connection between shot and thermal noise in a system governed by Boltzmann statistics is beautifully illustrated.> Tobi Delbruck, Carver Mead |
ISCAS | 1 |
| 1993 | Silicon retina with correlation-based, velocity-tuned pixelsabstractA functional two-dimensional silicon retina that computes a complete set of local direction-selective outputs is reported. The chip motion computation uses unidirectional delay lines as tuned filters for moving edges. Photoreceptors detect local changes in image intensity, and the outputs from these photoreceptors are coupled into the delay line, where they propagate with a particular speed in one direction. If the velocity of the moving edges matches that of the delay line, then the signal on the delay line is reinforced. The output of each pixel is the power in the delay line signal, computed within each pixel. This power computation provides the essential nonlinearity for velocity selectivity. The delay line architecture differs from the usual pairwise correlation models in that motion information is aggregated over an extended spatiotemporal range. As a result, the detectors are sensitive to motion over a wide range of spatial frequencies. The design of functional one- and two-dimensional silicon retinas with direction-selective, velocity-tuned pixels is described. It is shown that pixels with three hexagonal directions of motion selectivity are approximately (225 mum)(2) in area in a 2-mum CMOS technology and consume less than 5 muW of power. Tobi Delbruck |
IEEE Trans. Neural Networks | 1 |
| 1991 | Direction Selective Silicon Retina that Uses Null Inhibition
Ronald G. Benson, Tobi Delbruck |
NIPS | 2 |
| 1988 | An Electronic Photoreceptor Sensitive to Small Changes in Intensity
Tobi Delbruck, Carver Mead |
NIPS | 1 |
| 1988 | An analog VLSI implementation of the marr-poggio stereo correspondence algorithm
Misha Mahowald, Tobi Delbruck |
Neural Networks | 2 |