Emre Neftci

dblp:62/5283 · also Emre O. Neftci, Emre Ozgur Neftci · DBLP profile ↗
← Back
38ranked-venue papers
7as first author
17since 2021 · last 2026
0000-0002-0332-3273ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 3 first-author · 11 since 2021Systems, architecture and hardware · 14 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Hybrid Guided Variational Autoencoder for Visual Place Recognition
Zihan You, Emre Neftci, Thorben Schoepe
ICPR (15)3
2026 Zero-shot temporal resolution domain adaptation for spiking neural networks
abstract
Spiking Neural Networks (SNNs) are biologically-inspired deep neural networks that efficiently extract temporal information while offering promising gains in terms of energy efficiency and latency when deployed on neuromorphic devices. SNN parameters are sensitive to temporal resolution, leading to significant performance drops when the temporal resolution of target data during deployment is not the same as that of the source data used for training, especially when fine-tuning with the target data is not possible during deployment. To address this challenge, we propose three novel domain adaptation methods for adapting neuron parameters to account for the change in time resolution without re-training on target time resolution. The proposed methods are based on a mapping between neuron dynamics in SNNs and State Space Models (SSMs) and are applicable to general neuron models. We evaluate the proposed methods under spatio-temporal data tasks, namely the audio keyword spotting datasets SHD and MSWC, and the neuromorphic image NMINST dataset. Our methods provide an alternative to-and in most cases significantly outperform-the existing reference method that consists of scaling only the time constant. Notably, when the temporal resolution of the target data is double that of the source data, applying one of our proposed methods instead of the benchmark achieves classification accuracy of 89.5 % instead of 53.0 % on SHD, 93.6 % instead of 38.8 % on MSWC and 98.5 % instead of 97.2 % on NMNIST. Moreover, our results show that high accuracy on high temporal resolution data can be obtained by time-efficient training on lower temporal resolution data.
Sanja Karilanova, Maxime Fabre, Emre Neftci, Ayça Özçelikkale
Neural Networks3
2025 Efficient Event-Based Object Detection: A Hybrid Neural Network with Spatial and Temporal Attention
abstract
Event cameras offer high temporal resolution and dynamic range with minimal motion blur, making them promising for robust object detection. While Spiking Neural Networks (SNNs) on neuromorphic hardware are often considered for energy efficient and low latency event-based data processing, they often fall short of Artificial Neural Networks (ANNs) in accuracy and flexibility. Here, we introduce Attention-based Hybrid SNN-ANN backbones for event-based object detection to leverage the strengths of both SNN and ANN architectures. A novel Attention-based SNN-ANN bridge module captures sparse spatial and temporal relations from the SNN layer and converts them into dense feature maps for the ANN part of the backbone. Additionally, we present a variant that integrates DWConvLSTMs to the ANN blocks to capture slower dynamics. This multi-timescale network combines fast SNN processing for short timesteps with long-term dense RNN processing, effectively capturing both fast and slow dynamics. Experimental results demonstrate that our proposed method surpasses SNN-based approaches by significant margins, with results comparable to existing ANN and RNN-based methods. Unlike ANN-only networks, the hybrid setup allows us to implement the SNN blocks on digital neuromorphic hardware to investigate the feasibility of our approach. Extensive ablation studies and implementation on neuromorphic hardware confirm the effectiveness of our proposed modules and architectural choices. Our hybrid SNN-ANN architectures pave the way for ANN-like performance at a drastically reduced parameter, latency, and power budget.
Soikat Hasan Ahmed, Jan Finkbeiner, Emre Neftci
CVPR3
2025 Gain Cell-based Analog Content Addressable Memory for Dynamic Associative Tasks in AI
abstract
analog Content Addressable Memories (aCAMs) have proven useful for associative Compute-in-Memory (CIM) applications like Decision Trees, Finite State Machines, and Hyper-dimensional Computing. While non-volatile implementations using FeFETs and ReRAM devices offer speed, power, and area advantages, they suffer from slow write speeds and limited write cycles, making them less suitable for computations involving fully dynamic data patterns. To address these limitations, in this work, we propose a capacitor gain cell-based aCAM designed for dynamic processing, where frequent memory updates are required. Our system compares analog input voltages to boundaries stored in capacitors, enabling efficient dynamic tasks. We demonstrate the application of aCAM within transformer attention mechanisms by replacing the softmax-scaled dot-product similarity with aCAM similarity, achieving competitive results. Circuit simulations on a TSMC 28 nm node show promising performance in terms of energy efficiency, precision, and latency, making it well-suited for fast, dynamic AI applications.
Paul-Philipp Manea, Nathan Leroux, Emre Neftci, John Paul Strachan
ISCAS3
2025 Contrastive Consolidation of Top-Down Modulations Achieves Sparsely Supervised Continual Learning
abstract
Biological brains learn continually from a stream of unlabeled data, while integrating specialized information from sparsely labeled examples without compromising their ability to generalize. Meanwhile, machine learning methods are susceptible to catastrophic forgetting in this natural learning setting, as supervised specialist fine-tuning degrades performance on the original task. We introduce task-modulated contrastive learning (TMCL), which takes inspiration from the biophysical machinery in the neocortex, using predictive coding principles to integrate top-down information continually and without supervision. We follow the idea that these principles build a view-invariant representation space, and that this can be implemented using a contrastive loss. Then, whenever labeled samples of a new class occur, new affine modulations are learned that improve separation of the new class from all others, without affecting feedforward weights. By co-opting the view-invariance learning mechanism, we then train feedforward weights to match the unmodulated representation of a data sample to its modulated counterparts. This introduces modulation invariance into the representation space, and, by also using past modulations, stabilizes it. Our experiments show improvements in both class-incremental and transfer learning over state-of-the-art unsupervised approaches, as well as over comparable supervised approaches, using as few as 1% of available labels. Taken together, our work suggests that top-down modulations play a crucial role in balancing stability and plasticity.
Viet Anh Khoa Tran, Emre Neftci, Willem A. M. Wybo
NeurIPS2
2024 Understanding and Improving Optimization in Predictive Coding Networks
abstract
Backpropagation (BP), the standard learning algorithm for artificial neural networks, is often considered biologically implausible. In contrast, the standard learning algorithm for predictive coding (PC) models in neuroscience, known as the inference learning algorithm (IL), is a promising, bio-plausible alternative. However, several challenges and questions hinder IL's application to real-world problems. For example, IL is computationally demanding, and without memory-intensive optimizers like Adam, IL may converge to poor local minima. Moreover, although IL can reduce loss more quickly than BP, the reasons for these speedups or their robustness remains unclear. In this paper, we tackle these challenges by 1) altering the standard implementation of PC circuits to substantially reduce computation, 2) developing a novel optimizer that improves the convergence of IL without increasing memory usage, and 3) establishing theoretical results that help elucidate the conditions under which IL is sensitive to second and higher-order information.
Nicholas Alonso, Jeffrey L. Krichmar, Emre Neftci
AAAI3
2024 Harnessing Manycore Processors with Distributed Memory for Accelerated Training of Sparse and Recurrent Models
abstract
Current AI training infrastructure is dominated by single instruction multiple data (SIMD) and systolic array architectures, such as Graphics Processing Units (GPUs) and Tensor Processing Units (TPUs), that excel at accelerating parallel workloads and dense vector matrix multiplications. Potentially more efficient neural network models utilizing sparsity and recurrence cannot leverage the full power of SIMD processor and are thus at a severe disadvantage compared to today's prominent parallel architectures like Transformers and CNNs, thereby hindering the path towards more sustainable AI. To overcome this limitation, we explore sparse and recurrent model training on a massively parallel multiple instruction multiple data (MIMD) architecture with distributed local memory. We implement a training routine based on backpropagation though time (BPTT) for the brain-inspired class of Spiking Neural Networks (SNNs) that feature binary sparse activations. We observe a massive advantage in using sparse activation tensors with a MIMD processor, the Intelligence Processing Unit (IPU) compared to GPUs. On training workloads, our results demonstrate 5-10x throughput gains compared to A100 GPUs and up to 38x gains for higher levels of activation sparsity, without a significant slowdown in training convergence or reduction in final model performance. Furthermore, our results show highly promising trends for both single and multi IPU configurations as we scale up to larger model sizes. Our work paves the way towards more efficient, non-standard models via AI training hardware beyond GPUs, and competitive large scale SNN models.
Jan Finkbeiner, Thomas Gmeinder, Mark Pupilli, Alexander Titterton, Emre Neftci
AAAI5
2024 Optimizing Automatic Differentiation with Deep Reinforcement Learning
abstract
Computing Jacobians with automatic differentiation is ubiquitous in many scientific domains such as machine learning, computational fluid dynamics, robotics and finance. Even small savings in the number of computations or memory usage in Jacobian computations can already incur massive savings in energy consumption and runtime. While there exist many methods that allow for such savings, they generally trade computational efficiency for approximations of the exact Jacobian. In this paper, we present a novel method to optimize the number of necessary multiplications for Jacobian computation by leveraging deep reinforcement learning (RL) and a concept called cross-country elimination while still computing the exact Jacobian. Cross-country elimination is a framework for automatic differentiation that phrases Jacobian accumulation as ordered elimination of all vertices on the computational graph where every elimination incurs a certain computational cost. Finding the optimal elimination order that minimizes the number of necessary multiplications can be seen as a single player game which in our case is played by an RL agent. We demonstrate that this method achieves up to 33% improvements over state-of-the-art methods on several relevant tasks taken from relevant domains. Furthermore, we show that these theoretical gains translate into actual runtime improvements by providing a cross-country elimination interpreter in JAX that can execute the obtained elimination orders.
Jamie Lohoff, Emre Neftci
NeurIPS2
2023 Achieving efficient interpretability of reinforcement learning via policy distillation and selective input gradient regularization
Jinwei Xing, Takashi Nagata, Xinyun Zou, Emre Neftci, Jeffrey L. Krichmar
Neural Networks4
2023 Training Spiking Neural Networks Using Lessons From Deep Learning
abstract
The brain is the perfect place to look for inspiration to develop more efficient neural networks. The inner workings of our synapses and neurons provide a glimpse at what the future of deep learning might look like. This article serves as a tutorial and perspective showing how to apply the lessons learned from several decades of research in deep learning, gradient descent, backpropagation, and neuroscience to biologically plausible spiking neural networks (SNNs). We also explore the delicate interplay between encoding data as spikes and the learning process; the challenges and solutions of applying gradient-based learning to SNNs; the subtle link between temporal backpropagation and spike timing-dependent plasticity; and how deep learning might move toward biologically plausible online learning. Some ideas are well accepted and commonly used among the neuromorphic engineering community, while others are presented or justified for the first time here. A series of companion interactive tutorials complementary to this article using our Python package,snnTorch, are also made available: https://snntorch.readthedocs.io/en/latest/tutorials/index.html.
Jason Kamran Eshraghian, Max Ward 0001, Emre Neftci, Xinxin Wang 0002, Gregor Lenz, Girish Dwivedi, Mohammed Bennamoun, Doo Seok Jeong, Wei Lu 0003
Proc. IEEE3
2023 HyperSpikeASIC: Accelerating Event-Based Workloads With HyperDimensional Computing and Spiking Neural Networks
abstract
Today’s machine learning (ML) systems, running workloads, such as deep neural networks, which require billions of parameters and many hours to train a model, consume a significant amount of energy. Due to the complexity of computation and topology, even the quantized models are hard to deploy on edge devices under energy constraints. To combat this, researchers have been focusing on new emerging neuromorphic computing models. Two of those models are hyperdimensional computing (HDC) and spiking neural networks (SNNs), both with their own benefits. HDC has various desirable properties that other ML algorithms lack, such as robustness to noise, simple operations, and high parallelism. SNNs are able to process event-based signal data in an efficient manner. This work develops$\mathsf {HyperSpike}$, which utilizes a single, randomly initialized, and untrained SNN layer as a feature extractor connected to a trained HDC classifier. HDC is used to enable more efficient classification as well as provide robustness to errors. We experimentally show that$\mathsf {HyperSpike}$is on average$31.5\times $more robust to errors than traditional SNNs. On Intel’s Loihi (Davies et al., 2018),$\mathsf {HyperSpike}$is$10\times $faster and$2.6\times $more energy efficient over traditional SNN networks. We further develop$\mathsf {HyperSpikeASIC}$, a customized accelerator for$\mathsf {HyperSpike}$. By decoupling the neuron and synapses,$\mathsf {HyperSpikeASIC}$skips the inactive neurons and limits the neuron state updating to once per time step at most.$\mathsf {HyperSpikeASIC}$is$601\times $faster and$3467\times $more energy efficient than$\mathsf {HyperSpike}$running on Intel’s Loihi for SNN acceleration, and$12.2\times $faster and$211\times $more energy efficient than the state-of-the-art SNN ASIC implementation (Wang et al., 2022).
Justin Morris, Kenneth Michael Stewart, Hin Wai Lui, Behnam Khaleghi, Anthony Thomas, Thiago Goncalves-Marback, Baris Aksanli, Emre Neftci, Tajana Rosing
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.9
2022 HyperSpike: HyperDimensional Computing for More Efficient and Robust Spiking Neural Networks
abstract
Today's Machine Learning(ML) systems, especially those running in server farms running workloads such as Deep Neural Networks, which require billions of parameters and many hours to train a model, consume a significant amount of energy. To combat this, researchers have been focusing on new emerging neuromorphic computing models. Two of those models are Hyperdimensional Computing (HDC) and Spiking Neural Networks (SNNs), both with their own benefits. HDC has various desirable properties that other Machine Learning (ML) algorithms lack such as: robustness to noise in the system, simple operations, and high parallelism. SNNs are able to process event based signal data in an efficient manner. In this paper, we create HyperSpike, which utilizes a single, randomly initialized and untrained SNN layer as feature extractor connected to a trained HDC classifier. HDC is used to enable more efficient classification as well as provide robustness to errors. We experimentally show that HyperSpike is on average 31.5× more robust to errors than traditional SNNs. We also implement HyperSpike in hardware, and show that it is 10x faster and 2.6× more energy efficient over traditional SNN networks run on Intel's Loihi [1].
Justin Morris, Hin Wai Lui, Kenneth Michael Stewart, Behnam Khaleghi, Anthony Thomas, Thiago Goncalves-Marback, Baris Aksanli, Emre Neftci, Tajana Rosing
DATE8
2022 Uncertainty Aware Model Integration on Reinforcement Learning
abstract
Model-based reinforcement learning is an effective approach to reducing sample complexity by adding more data from the model. Dyna is a well-known architecture that contains model-based reinforcement learning and integrates learning from interactions with an environment and a model of the environment. Although the model can greatly help to speed up the agent's learning, acquiring an accurate model is a hard problem in spite of the recent great success of function approximation using neural networks. A wrong model causes degradation of the agent's performance and raises another question: to which extent should an agent rely on the model to update its policy? In this paper, we propose to use the confidence of the model simulations to the integrated learning process so that the agent avoids updating its policy based on uncertain simulations by the model. To obtain confidence, we apply the Monte Carlo dropout technique to the state transition model. We show that this approach contributes to improving early-stage training, thus helping speed up the agent to reach reasonable performance. We conduct experiments on simulated robotic locomotion tasks to demonstrate the effectiveness of our approach.
Takashi Nagata, Jinwei Xing, Tsutomu Kumazawa, Emre Neftci
IJCNN4
2022 Skipper: Enabling efficient SNN training through activation-checkpointing and time-skipping
abstract
Spiking neural networks (SNNs) are a highly efficient signal processing mechanism in biological systems that have inspired a plethora of research efforts aimed at translating their energy efficiency to computational platforms. Efficient training approaches are critical for the successful deployment of SNNs. Compared to mainstream deep neural networks (ANNs), training SNNs is far more challenging due to complex neural dynamics that evolve with time and their discrete, binary computing paradigm. Back-propagation-through-time (BPTT) with surrogate gradients has recently emerged as an effective technique to train deep SNNs directly. SNN-BPTT, however, has a major drawback in that it has a high memory requirement that increases with the number of timesteps. SNNs generally result from the discretization of Ordinary Differential Equations, due to which the sequence length must be typically longer than RNNs, compounding the time dependence problem. It, therefore, becomes hard to train deep SNNs on a single or multi-GPU setup with sufficiently large batch sizes or timesteps, and extended periods of training are required to achieve reasonable network performance. In this work, we reduce the memory requirements of BPTT in SNNs to enable the training of deeper SNNs with more timesteps (T). For this, we leverage the notion of activation re-computation in the context of SNN training that enables the GPU memory to scale sub-linearly with increasing time-steps. We observe that naively deploying the re-computation based approach leads to a considerable computational overhead. To solve this, we propose a time-skipped BPTT approximation technique, called Skipper, for SNNs, that not only alleviates this computation overhead, but also lowers memory consumption further with little to no loss of accuracy. We show the efficacy of our proposed technique by comparing it against a popular method for memory footprint reduction during training. Our evaluations on 5 state-of-the-art networks and 4 datasets show that for a constant batch size and time-steps, skipper reduces memory usage by 3.3× to 8.4× (6.7× on average) over baseline SNN-BPTT. It also achieves a speedup of 29% to 70% over the checkpointed approach and of 4% to 40% over the baseline approach. For a constant memory budget, skipper can scale to an order of magnitude higher timesteps compared to baseline SNN-BPTT.
Sonali Singh, Anup Sarma, Sen Lu, Abhronil Sengupta, Mahmut T. Kandemir, Emre Neftci, Narayanan Vijaykrishnan, Chita R. Das
MICRO6
2022 A Theoretical Framework for Inference Learning
abstract
Backpropagation (BP) is the most successful and widely used algorithm in deep learning. However, the computations required by BP are challenging to reconcile with known neurobiology. This difficulty has stimulated interest in more biologically plausible alternatives to BP. One such algorithm is the inference learning algorithm (IL). IL trains predictive coding models of neural circuits and has achieved equal performance to BP on supervised and auto-associative tasks. In contrast to BP, however, the mathematical foundations of IL are not well-understood. Here, we develop a novel theoretical framework for IL. Our main result is that IL closely approximates an optimization method known as implicit stochastic gradient descent (implicit SGD), which is distinct from the explicit SGD implemented by BP. Our results further show how the standard implementation of IL can be altered to better approximate implicit SGD. Our novel implementation considerably improves the stability of IL across learning rates, which is consistent with our theory, as a key property of implicit SGD is its stability. We provide extensive simulation results that further support our theoretical interpretations and find IL achieves quicker convergence when trained with mini-batch size one while performing competitively with BP for larger mini-batches when combined with Adam.
Nick Alonso, Beren Millidge, Jeffrey L. Krichmar, Emre Neftci
NeurIPS4
2021 Domain Adaptation In Reinforcement Learning Via Latent Unified State Representation
abstract
Despite the recent success of deep reinforcement learning (RL), domain adaptation remains an open problem. Although the generalization ability of RL agents is critical for the real-world applicability of Deep RL, zero-shot policy transfer is still a challenging problem since even minor visual changes could make the trained agent completely fail in the new task. To address this issue, we propose a two-stage RL agent that first learns a latent unified state representation (LUSR) which is consistent across multiple domains in the first stage, and then do RL training in one source domain based on LUSR in the second stage. The cross-domain consistency of LUSR allows the policy acquired from the source domain to generalize to other target domains without extra training. We first demonstrate our approach in variants of CarRacing games with customized manipulations, and then verify it in CARLA, an autonomous driving simulator with more complex and realistic visual observations. Our results show that this approach can achieve state-of-the-art domain adaptation performance in related RL tasks and outperforms prior approaches based on latent-representation based RL and image-to-image translation.
Jinwei Xing, Takashi Nagata, Kexin Chen 0002, Xinyun Zou, Emre Neftci, Jeffrey L. Krichmar
AAAI5
2021 Brain-Inspired Learning on Neuromorphic Substrates
abstract
Neuromorphic hardware strives to emulate brain-like neural networks and thus holds the promise for scalable, low-power information processing on temporal data streams. Yet, to solve real-world problems, these networks need to be trained. However, training on neuromorphic substrates creates significant challenges due to the offline character and the required nonlocal computations of gradient-based learning algorithms. This article provides a mathematical framework for the design of practical online learning algorithms for neuromorphic substrates. Specifically, we show a direct connection between real-time recurrent learning (RTRL), an online algorithm for computing gradients in conventional recurrent neural networks (RNNs), and biologically plausible learning rules for training spiking neural networks (SNNs). Furthermore, we motivate a sparse approximation based on block-diagonal Jacobians, which reduces the algorithm's computational complexity, diminishes the nonlocal information requirements, and empirically leads to good learning performance, thereby improving its applicability to neuromorphic substrates. In summary, our framework bridges the gap between synaptic plasticity and gradient-based approaches from deep learning and lays the foundations for powerful information processing on future neuromorphic hardware systems.
Friedemann Zenke, Emre Neftci
Proc. IEEE2
2020 Building a Better Lie Detector with BERT: The Difference Between Truth and Lies
abstract
Detecting lies or deceptive statements in text is a valuable skill. This is partly because the patterns that underlie deceptive text are not known. The aim of this work is to identify patterns that characterize deceptive text. A key step in this approach is to train a classifier based on the BERT (Bidirectional Encoder Representations from Transformers) network. BERT beats the state of the art in deception classification accuracy on the Ott Deceptive Opinion Spam corpus. The results of our ablation study indicate that certain components of the input, such as some parts of speech, are more informative to the classifier than others. Further part-of-speech analysis in "swing" sentences that are considered important to BERT's classification indicates that deceptive text is more formulaic and less varied than truthful text. We expanded our classifier into a new Generative Adversarial Network based on BERT to create exemplars of deceptive and truthful text that further showed the differences between truth and deception, reinforcing the underlying similarity of deceptive text in terms of part-of-speech makeup.
Dan Barsever, Sameer Singh 0001, Emre Neftci
IJCNN3
2020 Memory Organization for Energy-Efficient Learning and Inference in Digital Neuromorphic Accelerators
abstract
The energy efficiency of neuromorphic hardware is greatly affected by the energy of storing, accessing, and updating synaptic parameters. Various methods of memory organisation targeting energy-efficient digital accelerators have been investigated in the past, however, they do not completely encapsulate the energy costs at a system level. To address this shortcoming and to account for various overheads, we synthesize the controller and memory for different encoding schemes and extract the energy costs from these synthesized blocks. Additionally, we introduce functional encoding for structured connectivity such as the connectivity in convolutional layers. Functional encoding offers a 58% reduction in the energy to implement a backward pass and weight update in such layers compared to existing index-based solutions. We show that for a 2 layer spiking neural network trained to retain a spatio-temporal pattern, bitmap (PB-BMP) based organization can encode the sparser networks more efficiently. This form of encoding delivers a 1.37× improvement in energy efficiency coming at the cost of a 4% degradation in network retention accuracy as measured by the van Rossum distance.
Clemens JS Schaefer, Patrick Faley, Emre Neftci, Siddharth Joshi 0001
ISCAS3
2020 Terrain Classification with a Reservoir-Based Network of Spiking Neurons
abstract
Terrain classification is important for outdoor path planning, mapping, and navigation. We developed a reservoir-based spiking neural network (r-SNN) to classify three terrain types (i.e. grass, dirt, and road) in a botanical garden. It included a recurrent layer and a supervised layer. The input spike trains to the recurrent layer were generated from linear accelerometer and gyroscope sensor signals as well as camera frames from an Android smartphone that controlled a ground robot. Compared to a Support Vector Machine (SVM) model and a 3-layer (3L) logistic regression model, our r-SNN method generated better prediction accuracy without reliance on a time window of data. Using both images and sensors as input, the test accuracy of the r-SNN was over 95%, which was significantly better than the SVM and the 3L logistic regression. Because the r-SNN is compatible with neuromorphic hardware, our proposed method could be part of a biologically-inspired power-efficient autonomous robot navigation system.
Xinyun Zou, Tiffany Hwu, Jeffrey L. Krichmar, Emre Neftci
ISCAS4
2019 Inherent Weight Normalization in Stochastic Neural Networks
abstract
Multiplicative stochasticity such as Dropout improves the robustness and gener- alizability deep neural networks. Here, we further demonstrate that always-on multiplicative stochasticity combined with simple threshold neurons provide a suf- ficient substrate for deep learning machines. We call such models Neural Sampling Machines (NSM). We find that the probability of activation of the NSM exhibits a self-normalizing property that mirrors Weight Normalization, a previously studied mechanism that fulfills many of the features of Batch Normalization in an online fashion. The normalization of activities during training speeds up convergence by preventing internal covariate shift caused by changes in the distribution of inputs. The always-on stochasticity of the NSM confers the following advantages: the network is identical in the inference and learning phases, making the NSM a suitable substrate for continual learning, it can exploit stochasticity inherent to a physical substrate such as analog non-volatile memories for in memory computing, and it is suitable for Monte Carlo sampling, while requiring almost exclusively addition and comparison operations. We demonstrate NSMs on standard classification benchmarks (MNIST and CIFAR) and event-based classification benchmarks (N-MNIST and DVS Gestures). Our results show that NSMs perform comparably or better than conventional artificial neural networks with the same architecture.
Georgios Detorakis, Abhishek Khanna, Matthew Jerry, Suman Datta, Emre Neftci
NeurIPS6
2019 Contrastive Hebbian learning with random feedback weights
Georgios Detorakis, Travis Bartley, Emre Neftci
Neural Networks3
2018 A Recurrent Neural Network Based Model of Predictive Smooth Pursuit Eye Movement in Primates
abstract
A predictive mechanism in the brain enables primates to visually track a target with almost zero lag smooth pursuit eye movements, overcoming the delays in processing retinal inputs. Interestingly, it also allows pursuit of occluded targets with nonlinear motion patterns. We propose a recurrent neural network (RNN) model that rapidly learns the target velocity sequence and generates eye velocity signals to eliminate the initial lag between target and eye velocities, and to track occluded targets with nonlinear velocity. Moreover, the model is able to adapt to unpredictable perturbation and phase shift of target velocity and qualitatively reproduce the initial pursuit acceleration in experimentally observed timescales. We propose that the frontal eye field (FEF) region of the primate brain is homologous to the proposed RNN based on its persistent predictive activities during pursuit and location on the pursuit pathway.
Hirak J. Kashyap, Georgios Detorakis, Nikil Dutt, Jeffrey L. Krichmar, Emre Neftci
IJCNN5
2017 Event-driven random backpropagation: Enabling neuromorphic deep learning machines
abstract
An ongoing challenge in neuromorphic computing is to devise general and computationally efficient models of inference and learning which are compatible with the spatial and temporal constraints of the brain. The gradient descent back-propagation rule is a powerful algorithm that is ubiquitous in deep learning, but it relies on the immediate availability of network-wide information stored with high-precision memory. However, recent work shows that exact backpropagated weights are not essential for learning deep representations. Here, we demonstrate an event-driven random backpropagation (eRBP) rule that uses an error-modulated synaptic plasticity rule for learning deep representations in neuromorphic computing hardware. The rule is very suitable for implementation in neuromorphic hardware using a two-compartment leaky integrate & fire neuron and a membrane-voltage modulated, spike-driven plasticity rule. Our results show that using eRBP, deep representations are rapidly learned without using backpropagated gradients, achieving nearly identical classification accuracies compared to artificial neural network simulations on GPUs, while being robust to neural and synaptic state quantizations during learning.
Emre Neftci, Charles Augustine, Somnath Paul, Georgios Detorakis
ISCAS1
2016 Stochastic neuromorphic learning machines for weakly labeled data
abstract
At learning tasks where humans typically outperform computers, neuromorphic learning machines can have potential advantages in learning in terms of power and complexity compared to mainstream technologies. Here, we present Synaptic Sampling Machines (S2M), a class of stochastic neural networks that use stochasticity at the connections (synapses) to implement energy efficient semi- and unsupervised learning for weakly or unlabeled data. Stochastic synapses play the dual role of a regularizer during learning and a mechanism for implementing stochasticity in neural networks. We show a S2M network architecture that is well suited for a dedicated digital implementation, that is potentially hundredfold more energy efficient compared to equivalent algorithms operating on GPUs.
Emre Neftci
ICCD1
2016 TrueHappiness: Neuromorphic emotion recognition on TrueNorth
abstract
We present an approach to constructing a neuromorphic device that responds to language input by producing neuron spikes in proportion to the strength of the appropriate positive or negative emotional response. Specifically, we perform a fine-grained sentiment analysis task with implementations on two different systems: one using conventional spiking neural network (SNN) simulators and the other one using IBM's Neurosynaptic System TrueNorth. Input words are projected into a high-dimensional semantic space and processed through a fully-connected neural network (FCNN) containing rectified linear units (ReLU) trained via backpropagation. After training, this FCNN is converted to a SNN by substituting the ReLUs with integrate-and-fire neurons. We show that there is practically no performance loss due to conversion to a spiking network on a sentiment analysis test set, i.e. correlations with human annotations differ by less than 0.02 between the original DNN and its spiking equivalent. Additionally, we show that the SNN generated with this technique can be mapped to existing neuromorphic hardware - in our case, the TrueNorth chip. Mapping to the chip involves 4-bit synaptic weight discretization and adjustment of the neuron thresholds. The resulting end-to-end system can take a user input, i.e. a word in a vocabulary of over 300,000 words, and estimate its sentiment on TrueNorth with a power consumption of approximately 50 μW.
Peter U. Diehl, Bruno U. Pedroni, Andrew S. Cassidy, Paul Merolla, Emre Neftci, Guido Zarrella
IJCNN5
2016 Stochastic synaptic plasticity with memristor crossbar arrays
abstract
Memristive devices have been shown to exhibit slow and stochastic resistive switching behavior under low-voltage, low-current operating conditions. Here we explore such mechanisms to emulate stochastic plasticity in memristor crossbar synapse arrays. Interfaced with integrate-and-fire spiking neurons, the memristive synapse arrays are capable of implementing stochastic forms of spike-timing dependent plasticity which parallel mean-rate models of stochastic learning with binary synapses. We present theory and experiments with spike-based stochastic learning in memristor crossbar arrays, including simplified modeling as well as detailed physical simulation of memristor stochastic resistive switching characteristics due to voltage and current induced filament formation and collapse.
Rawan Naous, Maruan Al-Shedivat, Emre Neftci, Gert Cauwenberghs, Khaled N. Salama
ISCAS3
2016 Synaptic sampling in hardware spiking neural networks
abstract
Using a neural sampling approach, networks of stochastic spiking neurons, interconnected with plastic synapses, have been used to construct computational machines such as Restricted Boltzmann Machines (RBMs). Previous work towards building such networks achieved lower performances than traditional RBMs. More recently, Synaptic Sampling Machines (SSMs) were shown to outperform equivalent RBMs. In Synaptic Sampling Machines (SSMs), the stochasticity for the sampling is generated at the synapse. Stochastic synapses play the dual role of a regularizer during learning and an efficient mechanism for implementing stochasticity in neural networks over a wide dynamic range. In this paper we show that SSMs with stochastic synapses implemented in FPGA-based spiking neural networks can obtain a high accuracy in classifying MNIST handwritten digit database. We compare classification accuracy for different bit precision for stochastic and non-stochastic synapses and further argue that stochastic synapses have the same effect as synapses with higher bit precision but require significantly lower computational resources.
Sadique Sheik, Somnath Paul, Charles Augustine, Chinnikrishna Kothapalli, Muhammad M. Khellah, Gert Cauwenberghs, Emre Neftci
ISCAS7
2015 Learning of Chunking Sequences in Cognition and Behavior
abstract
We often learn and recall long sequences in smaller segments, such as a phone number 858 534 22 30 memorized as four segments. Behavioral experiments suggest that humans and some animals employ this strategy of breaking down cognitive or behavioral sequences into chunks in a wide variety of tasks, but the dynamical principles of how this is achieved remains unknown. Here, we study the temporal dynamics of chunking for learning cognitive sequences in a chunking representation using a dynamical model of competing modes arranged to evoke hierarchical Winnerless Competition (WLC) dynamics. Sequential memory is represented as trajectories along a chain of metastable fixed points at each level of the hierarchy, and bistable Hebbian dynamics enables the learning of such trajectories in an unsupervised fashion. Using computer simulations, we demonstrate the learning of a chunking representation of sequences and their robust recall. During learning, the dynamics associates a set of modes to each information-carrying item in the sequence and encodes their relative order. During recall, hierarchical WLC guarantees the robustness of the sequence order when the sequence is not too long. The resulting patterns of activities share several features observed in behavioral experiments, such as the pauses between boundaries of chunks, their size and their duration. Failures in learning chunking sequences provide new insights into the dynamical causes of neurological disorders such as Parkinson's disease and Schizophrenia.
Jordi Fonollosa, Emre Neftci, Mikhail I. Rabinovich
PLoS Comput. Biol.2
2013 Neuromorphic adaptations of restricted Boltzmann machines and deep belief networks
abstract
Restricted Boltzmann Machines (RBMs) and Deep Belief Networks (DBNs) have been demonstrated to perform efficiently on a variety of applications, such as dimensionality reduction and classification. Implementation of RBMs on neuromorphic platforms, which emulate large-scale networks of spiking neurons, has significant advantages from concurrency and low-power perspectives. This work outlines a neuromorphic adaptation of the RBM, which uses a recently proposed neural sampling algorithm (Buesing et al. 2011), and examines its algorithmic efficiency. Results show the feasibility of such alterations, which will serve as a guide for future implementation of such algorithms in neuromorphic very large scale integration (VLSI) platforms.
Bruno U. Pedroni, Srinjoy Das, Emre Neftci, Kenneth Kreutz-Delgado, Gert Cauwenberghs
IJCNN3
2012 Function approximation with uncertainty propagation in a VLSI spiking neural network
abstract
The brain combines and integrates multiple cues to take coherent, context-dependent action using distributed, event-based computational primitives. Computational models that use these principles in software simulations of recurrently coupled spiking neural networks have been demonstrated in the past, but their implementation in hybrid analog/digital Very Large Scale Integration (VLSI) spiking neural networks remains challenging. Here, we demonstrate a distributed spiking neural network architecture comprising multiple neuromorphic VLSI chips able to reproduce these types of cue combination and integration operations. This is achieved by encoding cues as population activities of input nodes in a network of recurrently coupled VLSI Integrate-and-Fire (I&F) neurons. The value of the cue is place-encoded, while its uncertainty is represented by the width of the population activity profile. Relationships among different cues are specified through bidirectional connectivity matrices, shared between the individual input node populations and an intermediate node population. The resulting network dynamics bidirectionally relate not only the values of three variables according to a specified relation, but also their uncertainties. When cues on two populations are specified, the standard deviation of the activity in the unspecified population varies approximately linearly with the widths of the two input cues, and has less than 6% error in position compared to the value specified by the inputs. The results suggest a mechanism for recurrently relating cues such that missing information can both be recovered and assigned a level of certainty.
Dane S. Corneil, Daniel Sonnleithner, Emre Neftci, Elisabetta Chicca, Matthew Cook 0001, Giacomo Indiveri, Rodney J. Douglas
IJCNN3
2012 Real-time inference in a VLSI spiking neural network
abstract
The ongoing motor output of the brain depends on its remarkable ability to rapidly transform and fuse a variety of sensory streams in real-time. The brain processes these data using networks of neurons that communicate by asynchronous spikes, a technology that is dramatically different from conventional electronic systems. We report here a step towards constructing electronic systems with analogous performance to the brain. Our VLSI spiking neural network combines in real-time three distinct sources of input data; each is place-encoded on an individual neuronal population that expresses soft Winner-Take-All dynamics. These arrays are combined according to a user-specified function that is embedded in the reciprocal connections between the soft Winner-Take-All populations and an intermediate shared population. The overall network is able to perform function approximation (missing data can be inferred from the available streams) and cue integration (when all input streams are present they enhance one another synergistically). The network performs these tasks with about 80% and 90% reliability, respectively. Our results suggest that with further technical improvement, it may be possible to implement more complex probabilistic models such as Bayesian networks in neuromorphic electronic systems.
Dane S. Corneil, Daniel Sonnleithner, Emre Neftci, Elisabetta Chicca, Matthew Cook 0001, Giacomo Indiveri, Rodney J. Douglas
ISCAS3
2012 Dynamic State and Parameter Estimation Applied to Neuromorphic Systems
abstract
Neuroscientists often propose detailed computational models to probe the properties of the neural systems they study. With the advent of neuromorphic engineering, there is an increasing number of hardware electronic analogs of biological neural systems being proposed as well. However, for both biological and hardware systems, it is often difficult to estimate the parameters of the model so that they are meaningful to the experimental system under study, especially when these models involve a large number of states and parameters that cannot be simultaneously measured. We have developed a procedure to solve this problem in the context of interacting neural populations using a recently developed dynamic state and parameter estimation (DSPE) technique. This technique uses synchronization as a tool for dynamically coupling experimentally measured data to its corresponding model to determine its parameters and internal state variables. Typically experimental data are obtained from the biological neural system and the model is simulated in software; here we show that this technique is also efficient in validating proposed network models for neuromorphic spike-based very large-scale integration (VLSI) chips and that it is able to systematically extract network parameters such as synaptic weights, time constants, and other variables that are not accessible by direct observation. Our results suggest that this method can become a very useful tool for model-based identification and configuration of neuromorphic multichip VLSI systems.
Emre Neftci, Bryan A. Toth, Giacomo Indiveri, Henry D. I. Abarbanel
Neural Comput.1
2011 Systematic configuration and automatic tuning of neuromorphic systems
abstract
In the past recent years several research groups have proposed neuromorphic Very Large Scale Integration (VLSI) devices that implement event-based sensors or biophysically realistic networks of spiking neurons. It has been argued that these devices can be used to build event-based systems, for solving real-world applications in real-time, with efficiencies and robustness that cannot be achieved with conventional computing technologies. In order to implement complex event-based neuromorphic systems it is necessary to interface the neuromorphic VLSI sensors and devices among each other, to robotic platforms, and to workstations (e.g. for data-logging and analysis). This apparently simple goal requires painstaking work that spans multiple levels of complexity and disciplines: from the custom layout of microelectronic circuits and asynchronous printed circuit boards, to the development of object oriented classes and methods in software; from electrical engineering and physics for analog/digital circuit design to neuroscience and computer science for neural computation and spike-based learning methods. Within this context, we present a framework we developed to simplify the configuration of multi-chip neuromorphic VLSI systems, and automate the mapping of neural network model parameters to neuromorphic circuit bias values.
Sadique Sheik, Fabio Stefanini, Emre Neftci, Elisabetta Chicca, Giacomo Indiveri
ISCAS3
2011 A Systematic Method for Configuring VLSI Networks of Spiking Neurons
abstract
An increasing number of research groups are developing custom hybrid analog/digital very large scale integration (VLSI) chips and systems that implement hundreds to thousands of spiking neurons with biophysically realistic dynamics, with the intention of emulating brainlike real-world behavior in hardware and robotic systems rather than simply simulating their performance on general-purpose digital computers. Although the electronic engineering aspects of these emulation systems is proceeding well, progress toward the actual emulation of brainlike tasks is restricted by the lack of suitable high-level configuration methods of the kind that have already been developed over many decades for simulations on general-purpose computers. The key difficulty is that the dynamics of the CMOS electronic analogs are determined by transistor biases that do not map simply to the parameter types and values used in typical abstract mathematical models of neurons and their networks. Here we provide a general method for resolving this difficulty. We describe a parameter mapping technique that permits an automatic configuration of VLSI neural networks so that their electronic emulation conforms to a higher-level neuronal simulation. We show that the neurons configured by our method exhibit spike timing statistics and temporal dynamics that are the same as those observed in the software simulated neurons and, in particular, that the key parameters of recurrent VLSI neural networks (e.g., implementing soft winner-take-all) can be precisely tuned. The proposed method permits a seamless integration between software simulations with hardware emulations and intertranslatability between the parameters of abstract neuronal models and their emulation counterparts. Most important, our method offers a route toward a high-level task configuration language for neuromorphic VLSI systems.
Emre Neftci, Elisabetta Chicca, Giacomo Indiveri, Rodney J. Douglas
Neural Comput.1
2010 Live demonstration: State-dependent sensory processing in networks of VLSI spiking neurons
abstract
This demonstration will show a distributed VLSI neuromorphic system implementing the soft Winner-Take-All (WTA) operation using spiking neurons. It also shows how recurrently connected instances of them can have persistent activity states, which can used for state-dependent computation. The live demonstration of this network will show that the position of a localized stimulus can be tracked and remembered along a trajectory initially encoded in the system. The visitors will experience the real-time, fast state-dependent processing of the sensory input occurring in the network.
Emre Neftci, Elisabetta Chicca, Matthew Cook 0001, Giacomo Indiveri, Rodney J. Douglas
ISCAS1
2010 State-dependent sensory processing in networks of VLSI spiking neurons
abstract
An increasing number of research groups develop dedicated hybrid analog/digital very large scale integration (VLSI) devices implementing hundreds of spiking neurons with bio-physically realistic dynamics. However, despite the significant progress in their design, there is still little insight in translating circuitry of neural assemblies into desired (non-trivial) function. In this work, we propose to use neural circuits implementing the soft Winner-Take-All (WTA) function. By showing that recurrently connected instances of them can have persistent activity states, which can be used as a form of working memory, we argue that such circuits can perform state-dependent computation. We demonstrate such a network in a distributed neuromorphic system consisting of two multi-neuron chips implementing soft WTA, stimulated by an event-based vision sensor. The resulting network is able to track and remember the position of a localized stimulus along a trajectory previously encoded in the system.
Emre Neftci, Elisabetta Chicca, Matthew Cook 0001, Giacomo Indiveri, Rodney J. Douglas
ISCAS1
2007 Contraction Properties of VLSI Cooperative Competitive Neural Networks of Spiking Neurons
abstract
A non–linear dynamic system is called contracting if initial conditions are for- gotten exponentially fast, so that all trajectories converge to a single trajectory. We use contraction theory to derive an upper bound for the strength of recurrent connections that guarantees contraction for complex neural networks. Specifi- cally, we apply this theory to a special class of recurrent networks, often called Cooperative Competitive Networks (CCNs), which are an abstract representation of the cooperative-competitive connectivity observed in cortex. This specific type of network is believed to play a major role in shaping cortical responses and se- lecting the relevant signal among distractors and noise. In this paper, we analyze contraction of combined CCNs of linear threshold units and verify the results of our analysis in a hybrid analog/digital VLSI CCN comprising spiking neurons and dynamic synapses.
Emre Neftci, Elisabetta Chicca, Giacomo Indiveri, Jean-Jacques E. Slotine, Rodney J. Douglas
NIPS1