Abhronil Sengupta

dblp:12/10604 · DBLP profile ↗
← Back
31ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0002-5545-4494ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 20 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 10 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Software engineering, systems software and programming languages · 3Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 P-Spikessm: Harnessing Probabilistic Spiking State Space Models for Long-Range Dependency Tasks
abstract
Spiking neural networks (SNNs) are posited as a computationally efficient and biologically plausible alternative to conventional neural architectures, with their core computational framework primarily using the leaky integrate-and-fire (LIF) neuron model. However, the limited hidden state representation of LIF neurons, characterized by a scalar membrane potential, and sequential spike generation process, poses challenges for effectively developing scalable spiking models to address long-range dependencies in sequence learning tasks. In this study, we develop a scalable probabilistic spiking learning framework for long-range dependency tasks leveraging the fundamentals of state space models. Unlike LIF neurons that rely on the deterministic Heaviside function for a sequential process of spike generation, we introduce a SpikeSampler layer that samples spikes stochastically based on an SSM-based neuronal model while allowing parallel computations. To address non-differentiability of the spiking operation and enable effective training, we also propose a surrogate function tailored for the stochastic nature of the SpikeSampler layer. To enhance inter-neuron communication, we introduce the SpikeMixer block, which integrates spikes from neuron populations in each layer. This is followed by a ClampFuse layer, incorporating a residual connection to capture complex dependencies, enabling scalability of the model. Our models attain state-of-the-art performance among SNN models across diverse long-range dependency tasks, encompassing the Long Range Arena benchmark, permuted sequential MNIST, and the Speech Command dataset and demonstrate sparse spiking pattern highlighting its computational efficiency.
Malyaban Bal, Abhronil Sengupta
ICLR2
2024 SpikingBERT: Distilling BERT to Train Spiking Language Models Using Implicit Differentiation
abstract
Large language Models (LLMs), though growing exceedingly powerful, comprises of orders of magnitude less neurons and synapses than the human brain. However, it requires significantly more power/energy to operate. In this work, we propose a novel bio-inspired spiking language model (LM) which aims to reduce the computational cost of conventional LMs by drawing motivation from the synaptic information flow in the brain. In this paper, we demonstrate a framework that leverages the average spiking rate of neurons at equilibrium to train a neuromorphic spiking LM using implicit differentiation technique, thereby overcoming the non-differentiability problem of spiking neural network (SNN) based algorithms without using any type of surrogate gradient. The steady-state convergence of the spiking neurons also allows us to design a spiking attention mechanism, which is critical in developing a scalable spiking LM. Moreover, the convergence of average spiking rate of neurons at equilibrium is utilized to develop a novel ANN-SNN knowledge distillation based technique wherein we use a pre-trained BERT model as “teacher” to train our “student” spiking architecture. While the primary architecture proposed in this paper is motivated by BERT, the technique can be potentially extended to different kinds of LLMs. Our work is the first one to demonstrate the performance of an operational spiking LM architecture on multiple different tasks in the GLUE benchmark. Our implementation source code is available at https://github.com/NeuroCompLab-psu/SpikingBERT.
Malyaban Bal, Abhronil Sengupta
AAAI2
2024 Equilibrium-Based Learning Dynamics in Spiking Architectures
abstract
This paper delves into methodologies that treat spiking architectures as continuously evolving dynamical systems, revealing intriguing parallels with the learning dynamics in the brain. The methods discussed in this paper addresses multiple challenges of training spiking architectures and highlights the necessity for bio-plausible local learning and increasing model scalability in spiking architectures. We begin by exploring an energy-based learning mechanism, namely Equilibrium Propagation (EP), which emphasizes the attainment of stable states by converging to energy minimas at each training phase, thus allowing for formulation of spatially and temporally local state and weight update rules. Subsequently, we delve into the synergy achieved by integrating the underlying energy-based convergent RNN architecture with a different energy-based model, namely modern Hopfield networks, thereby amplifying the capabilities of the resultant model. We further explore an efficient learning framework rooted in the convergence of the average spiking rates of neurons, which can be leveraged to advance the creation of highly scalable spiking architectures. The methodologies discussed allows spiking architectures to transition beyond simple vision-related tasks and develop solutions for complex sequence learning problems. Moreover, both the frameworks can be used to develop spiking architectures which can be deployed in neuromorphic hardware to realize their energy/power efficiency.
Malyaban Bal, Abhronil Sengupta
ISCAS2
2023 Astromorphic Self-Repair of Neuromorphic Hardware Systems
abstract
While neuromorphic computing architectures based on Spiking Neural Networks (SNNs) are increasingly gaining interest as a pathway toward bio-plausible machine learning, attention is still focused on computational units like the neuron and synapse. Shifting from this neuro-synaptic perspective, this paper attempts to explore the self-repair role of glial cells, in particular, astrocytes. The work investigates stronger correlations with astrocyte computational neuroscience models to develop macro-models with a higher degree of bio-fidelity that accurately captures the dynamic behavior of the self-repair process. Hardware-software co-design analysis reveals that bio-morphic astrocytic regulation has the potential to self-repair hardware realistic faults in neuromorphic hardware systems with significantly better accuracy and repair convergence for unsupervised learning tasks on the MNIST and F-MNIST datasets. Our implementation source code and trained models are available at https://github.com/NeuroCompLab-psu/Astromorphic_Self_Repair.
Zhuangyu Han, A. N. M. Nafiul Islam, Abhronil Sengupta
AAAI3
2023 Sequence Learning Using Equilibrium Propagation
abstract
Equilibrium Propagation (EP) is a powerful and more bio-plausible alternative to conventional learning frameworks such as backpropagation. The effectiveness of EP stems from the fact that it relies only on local computations and requires solely one kind of computational unit during both of its training phases, thereby enabling greater applicability in domains such as bio-inspired neuromorphic computing. The dynamics of the model in EP is governed by an energy function and the internal states of the model consequently converge to a steady state following the state transition rules defined by the same. However, by definition, EP requires the input to the model (a convergent RNN) to be static in both the phases of training. Thus it is not possible to design a model for sequence classification using EP with an LSTM or GRU like architecture. In this paper, we leverage recent developments in modern hopfield networks to further understand energy based models and develop solutions for complex sequence classification tasks using EP while satisfying its convergence criteria and maintaining its theoretical similarities with recurrent backpropagation. We explore the possibility of integrating modern hopfield networks as an attention mechanism with convergent RNN models used in EP, thereby extending its applicability for the first time on two different sequence classification tasks in natural language processing viz. sentiment analysis (IMDB dataset) and natural language inference (SNLI dataset). Our implementation source code is available at https://github.com/NeuroCompLab-psu/EqProp-SeqLearning.
Malyaban Bal, Abhronil Sengupta
IJCAI2
2023 Leveraging Probabilistic Switching in Superparamagnets for Temporal Information Encoding in Neuromorphic Systems
abstract
Brain-inspired computing—leveraging neuroscientific principles underpinning the unparalleled efficiency of the brain in solving cognitive tasks—is emerging to be a promising pathway to solve several algorithmic and computational challenges faced by deep learning today. Nonetheless, current research in neuromorphic computing is driven by our well-developed notions of running deep learning algorithms on computing platforms that perform deterministic operations. In this article, we argue that taking a different route of performing temporal information encoding in probabilistic neuromorphic systems may help solve some of the current challenges in the field. The article considers superparamagnetic tunnel junctions as a potential pathway to enable a new generation of brain-inspired computing that combines the facets and associated advantages of two complementary insights from computational neuroscience: 1) how information is encoded and 2) how computing occurs in the brain. The hardware-algorithm co-design analysis demonstrates 97.41% accuracy of a state-compressed 3-layer spintronics-enabled stochastic spiking network on the MNIST dataset with high spiking sparsity due to temporal information encoding.
Kezhou Yang, Dhuruva Priyan G. M, Abhronil Sengupta
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2022 Skipper: Enabling efficient SNN training through activation-checkpointing and time-skipping
abstract
Spiking neural networks (SNNs) are a highly efficient signal processing mechanism in biological systems that have inspired a plethora of research efforts aimed at translating their energy efficiency to computational platforms. Efficient training approaches are critical for the successful deployment of SNNs. Compared to mainstream deep neural networks (ANNs), training SNNs is far more challenging due to complex neural dynamics that evolve with time and their discrete, binary computing paradigm. Back-propagation-through-time (BPTT) with surrogate gradients has recently emerged as an effective technique to train deep SNNs directly. SNN-BPTT, however, has a major drawback in that it has a high memory requirement that increases with the number of timesteps. SNNs generally result from the discretization of Ordinary Differential Equations, due to which the sequence length must be typically longer than RNNs, compounding the time dependence problem. It, therefore, becomes hard to train deep SNNs on a single or multi-GPU setup with sufficiently large batch sizes or timesteps, and extended periods of training are required to achieve reasonable network performance. In this work, we reduce the memory requirements of BPTT in SNNs to enable the training of deeper SNNs with more timesteps (T). For this, we leverage the notion of activation re-computation in the context of SNN training that enables the GPU memory to scale sub-linearly with increasing time-steps. We observe that naively deploying the re-computation based approach leads to a considerable computational overhead. To solve this, we propose a time-skipped BPTT approximation technique, called Skipper, for SNNs, that not only alleviates this computation overhead, but also lowers memory consumption further with little to no loss of accuracy. We show the efficacy of our proposed technique by comparing it against a popular method for memory footprint reduction during training. Our evaluations on 5 state-of-the-art networks and 4 datasets show that for a constant batch size and time-steps, skipper reduces memory usage by 3.3× to 8.4× (6.7× on average) over baseline SNN-BPTT. It also achieves a speedup of 29% to 70% over the checkpointed approach and of 4% to 40% over the baseline approach. For a constant memory budget, skipper can scale to an order of magnitude higher timesteps compared to baseline SNN-BPTT.
Sonali Singh, Anup Sarma, Sen Lu, Abhronil Sengupta, Mahmut T. Kandemir, Emre Neftci, Narayanan Vijaykrishnan, Chita R. Das
MICRO4
2021 Gesture-SNN: Co-optimizing accuracy, latency and energy of SNNs for neuromorphic vision sensors
abstract
As originally published figures in the document were missing. A corrected replacement file was provided by the authors.Spiking neural networks (SNNs) are recently gaining popularity due to their low-power, spatio-temporal computing paradigm as opposed to more conventional deep learning approaches that mainly focus on spatial characteristics of data. When paired with biologically-inspired asynchronous event sensors, they can create energy-efficient near-sensor systems that are ideal for mobile, resource-constrained and embedded-computing scenarios. Training deep SNNs, however, is challenging due to their discrete nature. The most successful method so far involves training deep artificial neural networks (ANNs) using Gradient-Descent and then converting them to SNNs. The ANN-to-SNN conversion technique has mostly been evaluated on standard static image datasets using rate-based encoding of spikes. In this work, we find that a direct application of the ANN-to-SNN conversion technique to process event data via SNNs leads to arbitrary accuracy losses. Through insights gained from theoretical analyses as well as empirical observations, we propose three novel techniques to restore the conversion accuracy on event data and show proof-of-concept results, comparable to the state-of-the-art, on the IBM DVS Gesture dataset. Further exploration of the SNN design space reveals additional insights to fine-tune the accuracy-latency-peak power trade-off. Finally, we evaluate our proposed schemes on an existing neuromorphic accelerator and show that our best-performing model is $\sim 38$% more accurate with $\sim 35$% lower energy and $\sim 55$% lower EDP compared to its traditional SNN counterpart.
Sonali Singh, Anup Sarma, Sen Lu, Abhronil Sengupta, Narayanan Vijaykrishnan, Chita R. Das
ISLPED4
2021 RxNN: A Framework for Evaluating Deep Neural Networks on Resistive Crossbars
abstract
Resistive crossbars designed with nonvolatile memory devices have emerged as promising building blocks for deep neural network (DNN) hardware, due to their ability to compactly and efficiently realize vector-matrix multiplication (VMM), the dominant computational kernel in DNNs. However, a key challenge with resistive crossbars is that they suffer from a range of device and circuit level nonidealities, such as driver resistance, sensing resistance, sneak paths, interconnect parasitics, nonlinearities in the peripheral circuits, stochastic write operations, and process variations. These nonidealities can lead to errors in VMMs, eventually degrading the DNN's accuracy. It is therefore critical to study the impact of crossbar nonidealities on the accuracy of large-scale DNNs (with millions of neurons and billions of synaptic connections). However, this is challenging because the existing device and circuit models are too slow to use in application-level evaluations. We present RxNN, a fast and accurate simulation framework to evaluate large-scale DNNs on resistive crossbar systems. RxNN splits and maps the computations involved in each DNN layer into crossbar operations, and evaluates them using a fast crossbar model (FCM) that accurately captures the errors arising due to crossbar nonidealities while being four-to-five orders of magnitude faster than circuit simulation. FCM models a crossbar-based VMM operation using three stages-nonlinear models for the input and output peripheral circuits (digital-to-analog and analog-to-digital converters), and an equivalent nonideal conductance matrix for the core crossbar array. We implement RxNN by extending the Caffe machine learning framework and use it to evaluate a suite of six large-scale DNNs developed for the ImageNet Challenge (ILSVRC). Our experiments reveal that resistive crossbar nonidealities can lead to significant accuracy degradations (9.6%-32%) for these large-scale DNNs. To the best of our knowledge, this article is the first quantitative evaluation of the accuracy of large-scale DNNs on resistive crossbar-based hardware. We also demonstrate that RxNN enables fast model-in-the-loop retraining of DNNs to partially mitigate the accuracy degradation.
Shubham Jain 0004, Abhronil Sengupta, Kaushik Roy 0001, Anand Raghunathan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2020 Training Deep Spiking Neural Networks for Energy-Efficient Neuromorphic Computing
abstract
Spiking Neural Networks (SNNs), widely known as the third generation of neural networks, encode input information temporally using sparse spiking events, which can be harnessed to achieve higher computational efficiency for cognitive tasks. However, considering the rapid strides in accuracy enabled by state-of-the-art Analog Neural Networks (ANNs), SNN training algorithms are much less mature, leading to accuracy gap between SNNs and ANNs. In this paper, we propose different SNN training methodologies, varying in degrees of biofidelity, and evaluate their efficacy on complex image recognition datasets. First, we present biologically plausible Spike Timing Dependent Plasticity (STDP) based deterministic and stochastic algorithms for unsupervised representation learning in SNNs. Our analysis on the CIFAR-10 dataset indicates that STDP-based learning rules enable the convolutional layers to self-learn low-level input features using fewer training examples. However, STDP-based learning is limited in applicability to shallow SNNs (≤4 layers) while yielding considerably lower than state-of-the-art accuracy. In order to scale the SNNs deeper and improve the accuracy further, we propose conversion methodology to map off-the-shelf trained ANN to SNN for energy-efficient inference. We demonstrate 69.96% accuracy for VGG16-SNN on ImageNet. However, ANN-to-SNN conversion leads to high inference latency for achieving the best accuracy. In order to minimize the inference latency, we propose spike-based error backpropagation algorithm using differentiable approximation for the spiking neuron. Our preliminary experiments on CIFAR-10 show that spike-based error backpropagation effectively captures temporal statistics to reduce the inference latency by up to 8× compared to converted SNNs while yielding comparable accuracy
Gopalakrishnan Srinivasan, Chankyu Lee, Abhronil Sengupta, Priyadarshini Panda, Syed Shakib Sarwar, Kaushik Roy 0001
ICASSP3
2020 NEBULA: A Neuromorphic Spin-Based Ultra-Low Power Architecture for SNNs and ANNs
abstract
Brain-inspired cognitive computing has so far followed two major approaches - one uses multi-layered artificial neural networks (ANNs) to perform pattern-recognition-related tasks, whereas the other uses spiking neural networks (SNNs) to emulate biological neurons in an attempt to be as efficient and fault-tolerant as the brain. While there has been considerable progress in the former area due to a combination of effective training algorithms and acceleration platforms, the latter is still in its infancy due to the lack of both. SNNs have a distinct advantage over their ANN counterparts in that they are capable of operating in an event-driven manner, thus consuming very low power. Several recent efforts have proposed various SNN hardware design alternatives, however, these designs still incur considerable energy overheads.In this context, this paper proposes a comprehensive design spanning across the device, circuit, architecture and algorithm levels to build an ultra low-power architecture for SNN and ANN inference. For this, we use spintronics-based magnetic tunnel junction (MTJ) devices that have been shown to function as both neuro-synaptic crossbars as well as thresholding neurons and can operate at ultra low voltage and current levels. Using this MTJ-based neuron model and synaptic connections, we design a low power chip that has the flexibility to be deployed for inference of SNNs, ANNs as well as a combination of SNN-ANN hybrid networks - a distinct advantage compared to prior works. We demonstrate the competitive performance and energy efficiency of the SNNs as well as hybrid models on a suite of workloads. Our evaluations show that the proposed design, NEBULA, is up to 7.9× more energy efficient than a state-of-the-art design, ISAAC, in the ANN mode. In the SNN mode, our design is about 45× more energy-efficient than a contemporary SNN architecture, INXS. Power comparison between NEBULA ANN and SNN modes indicates that the latter is at least 6.25× more power-efficient for the observed benchmarks.
Sonali Singh, Anup Sarma, Nicholas Jao, Ashutosh Pattnaik, Sen Lu, Kezhou Yang, Abhronil Sengupta, Narayanan Vijaykrishnan, Chita R. Das
ISCA7
2020 TraNNsformer: Clustered Pruning on Crossbar-Based Architectures for Energy-Efficient Neural Networks
abstract
Implementation of neuromorphic systems using memristive crossbar array (MCA) has emerged as a promising solution to enable low-power acceleration of neural networks. However, the recent trend to design deep neural networks (DNNs) for achieving human-like cognitive abilities poses significant challenges toward the scalable design of neuromorphic systems (due to the increase in computation/storage demands). Network pruning is a powerful technique to remove redundant connections for designing optimally connected (maximally sparse) DNNs. However, such pruning techniques induce irregular connections that are incoherent to the crossbar structure. Eventually, they produce DNNs with highly inefficient hardware realizations (in terms of area and energy). In this article, we propose TraNNsformer-an integrated training framework that transforms DNNs to enable their efficient realization on MCA-based systems. TraNNsformer first prunes the connectivity matrix while forming clusters with the remaining connections. Subsequently, it retrains the network to fine-tune the connections and reinforce the clusters. This is done iteratively to transform the original connectivity into an optimally pruned and maximally clustered mapping. We evaluated the proposed framework by transforming networks of different complexity based on multilayer perceptron (MLP) and convolutional neural network (CNN) topologies on a wide range of datasets (MNIST, SVHN, CIFAR10, and ImageNet) and executing them on MCA-based systems to analyze the area and energy benefits. Without accuracy loss, TraNNsformer reduces the area (energy) consumption by 28%-55% (49%-67%)of MLP networks and by 28%-48% (3%-39%) of CNN networks with respect to the original network implementations. Compared to network pruning, TraNNsformer achieves 28%-49% (15%-29%) area (energy) savings for MLP networks and 20%-44% (1%-11%) area (energy) saving for CNN networks. Furthermore, TraNNsformer is a technology-aware framework that allows mapping a given DNN to any MCA size permissible by the memristive technology for reliable operations.
Aayush Ankit, Timur Ibrayev, Abhronil Sengupta, Kaushik Roy 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2020 Revisiting Stochastic Computing in the Era of Nanoscale Nonvolatile Technologies
abstract
In this era of nanoscale technologies, the inherent characteristics of some nonvolatile devices, such as resistive random access memory (ReRAM), phase-change material (PCM), and spintronics, can emulate stochastic functionalities. Traditionally, these devices have been engineered to suppress the stochastic switching behavior as it poses reliability concerns for memory storage and logic applications. However, leveraging stochasticity in such devices led to a renewed interest in hardware-software codesign of stochastic algorithms since the CMOS-based implementations of stochastic algorithms involve cumbersome circuitry to generate “stochastic bits.” In this article, we consider two classes of problems: deep neural networks (DNNs) and combinatorial optimization. The rapidly growing demands of artificial intelligence (AI) have sparked an interest in energy-efficient implementations of large DNNs, with binary representations of synaptic weights and neuronal activities. Stochasticity plays an important role in leveraging the benefits of these binary representations, leading to model compression and optimization during training. In combinatorial optimization, such as graph coloring or traveling salesman problems, stochastic algorithms, such as the Ising computing model, have been shown to be effective. These problems require exhaustive computational procedures, and the Ising model uses a natural annealing agent to achieve near-optimal solutions in a reasonable timescale, without getting stuck in “local minima.” In this article, we present a broad review of stochastic computing utilizing the stochastic switching characteristics of devices based on nanoscale nonvolatile technologies. We show how to codesign of the devices and algorithms that can enable optimal solutions for both combinatorial problems and binary neural networks for local learning and inference. Directly mapping the nonvolatile device characteristics to the stochastic algorithms without the need for storing the bits in a separate memory leads to efficient use of hardware.
Amogh Agrawal, Indranil Chakraborty, Deboleena Roy, Utkarsh Saxena, Saima Sharmin, Minsuk Koo, Yong Shim, Gopalakrishnan Srinivasan, Chamika M. Liyanagedera, Abhronil Sengupta, Kaushik Roy 0001
IEEE Trans. Very Large Scale Integr. Syst.10
2019 Programmable Non-Volatile Memory Design Featuring Reconfigurable In-Memory Operations
abstract
With data volume growing exponentially in today's era, modern computing systems are increasingly bottlenecked and consistently burdened by the costs of data movement. Driven by the development of emerging non-volatile memory (NVM) technologies and by the increasing demand for high throughput in big data applications, considerable research effort has gone into embedding computing in memory and exploiting parallelism in data-intensive workloads to address the “memory wall” bottleneck. In this work, we propose a non-volatile memory design which leverages run-time reconfigurability of peripheral circuits to perform various in-memory computations like that of a field-programmable gate array (FPGA). Our architecture allows this intelligent storage system to operate as both a main memory and an accelerator for memory-intensive applications such as matrix multiplication, database query and artificial neural networks.
Nicholas Jao, Akshay Krishna Ramanathan, Abhronil Sengupta, Jack Sampson, Narayanan Vijaykrishnan
ISCAS3
2017 RESPARC: A Reconfigurable and Energy-Efficient Architecture with Memristive Crossbars for Deep Spiking Neural Networks
abstract
Neuromorphic computing using post-CMOS technologies is gaining immense popularity due to its promising abilities to address the memory and power bottlenecks in von-Neumann computing systems. In this paper, we propose RESPARC - a reconfigurable and energy efficient architecture built-on Memristive Crossbar Arrays (MCA) for deep Spiking Neural Networks (SNNs). Prior works were primarily focused on device and circuit implementations of SNNs on crossbars. RESPARC advances this by proposing a complete system for SNN acceleration and its subsequent analysis. RESPARC utilizes the energy-efficiency of MCAs for inner-product computation and realizes a hierarchical reconfigurable design to incorporate the data-flow patterns in an SNN in a scalable fashion. We evaluate the proposed architecture on different SNNs ranging in complexity from 2k-230k neurons and 1.2M-5.5M synapses. Simulation results on these networks show that compared to the baseline digital CMOS architecture, RESPARC achieves 500x (15x) efficiency in energy benefits at 300x (60x) higher throughput for multi-layer perceptrons (deep convolutional networks). Furthermore, RESPARC is a technology-aware architecture that maps a given SNN topology to the most optimized MCA size for the given crossbar technology.
Aayush Ankit, Abhronil Sengupta, Priyadarshini Panda, Kaushik Roy 0001
DAC2
2017 Magnetic tunnel junction enabled all-spin stochastic spiking neural network
abstract
Biologically-inspired spiking neural networks (SNNs) have attracted significant research interest due to their inherent computational efficiency in performing classification and recognition tasks. The conventional CMOS-based implementations of large-scale SNNs are power intensive. This is a consequence of the fundamental mismatch between the technology used to realize the neurons and synapses, and the neuroscience mechanisms governing their operation, leading to area-expensive circuit designs. In this work, we present a three-terminal spintronic device, namely, the magnetic tunnel junction (MTJ)-heavy metal (HM) heterostructure that is inherently capable of emulating the neuronal and synaptic dynamics. We exploit the stochastic switching behavior of the MTJ in the presence of thermal noise to mimic the probabilistic spiking of cortical neurons, and the conditional change in the state of a binary synapse based on the pre- and post-synaptic spiking activity required for plasticity. We demonstrate the efficacy of a crossbar organization of our MTJ-HM based stochastic SNN in digit recognition using a comprehensive device-circuit-system simulation framework. The energy efficiency of the proposed system stems from the ultra-low switching energy of the MTJ-HM device, and the in-memory computation rendered possible by the localized arrangement of the computational units (neurons) and non-volatile synaptic memory in such crossbar architectures.
Gopalakrishnan Srinivasan, Abhronil Sengupta, Kaushik Roy 0001
DATE2
2017 TraNNsformer: Neural network transformation for memristive crossbar based neuromorphic system design
abstract
Implementation of Neuromorphic Systems using post Complementary Metal-Oxide-Semiconductor (CMOS) technology based Memristive Crossbar Array (MCA) has emerged as a promising solution to enable low-power acceleration of neural networks. However, the recent trend to design Deep Neural Networks (DNNs) for achieving human-like cognitive abilities poses significant challenges towards the scalable design of neuromorphic systems (due to the increase in computation/storage demands). Network pruning [7] is a powerful technique to remove redundant connections for designing optimally connected (maximally sparse) DNNs. However, such pruning techniques induce irregular connections that are incoherent to the crossbar structure. Eventually they produce DNNs with highly inefficient hardware realizations (in terms of area and energy). In this work, we propose TraNNsformer - an integrated training framework that transforms DNNs to enable their efficient realization on MCA-based systems. TraNNsformer first prunes the connectivity matrix while forming clusters with the remaining connections. Subsequently, it retrains the network to fine tune the connections and reinforce the clusters. This is done iteratively to transform the original connectivity into an optimally pruned and maximally clustered mapping. We evaluated the proposed framework by transforming different Multi-Layer Perceptron (MLP) based Spiking Neural Networks (SNNs) on a wide range of datasets (MNIST, SVHN and CIFAR10) and executing them on MCA-based systems to analyze the area and energy benefits. Without accuracy loss, TraNNsformer reduces the area (energy) consumption by 28%-55% (49%-67%) with respect to the original network. Compared to network pruning, TraNNsformer achieves 28%-49% (15%-29%) area (energy) savings. Furthermore, TraNNsformer is a technology-aware framework that allows mapping a given DNN to any MCA size permissible by the memristive technology for reliable operations.
Aayush Ankit, Abhronil Sengupta, Kaushik Roy 0001
ICCAD2
2017 Performance analysis and benchmarking of all-spin spiking neural networks (Special session paper)
abstract
Spiking Neural Network based brain-inspired computing paradigms are becoming increasingly popular tools for various cognitive tasks. The sparse event-driven processing capability enabled by such networks can be potentially appealing for implementation of low-power neural computing platforms. However, the parallel and memory-intensive computations involved in such algorithms is in complete contrast to the sequential fetch, decode, execute cycles of conventional von-Neumann processors. Recent proposals have investigated the design of spintronic “in-memory” crossbar based computing architectures driving “spin neurons” that can potentially alleviate the memory-access bottleneck of CMOS based systems and simultaneously offer the prospect of low-power inner product computations. In this article, we perform a rigorous system-level simulation study of such All-Spin Spiking Neural Networks on a benchmark suite of 6 recognition problems ranging in network complexity from 10k-7.4M synapses and 195-9.2k neurons. System level simulations indicate that the proposed spintronic architecture can potentially achieve ~1292× energy efficiency and ~ 235× speedup on average over the benchmark suite in comparison to an optimized CMOS implementation at 45nm technology node.
Abhronil Sengupta, Aayush Ankit, Kaushik Roy 0001
IJCNN1
2017 Energy-Efficient and Improved Image Recognition with Conditional Deep Learning
abstract
Deep-learning neural networks have proven to be very successful for a wide range of recognition tasks across modern computing platforms. However, the computational requirements associated with such deep nets can be quite high, and hence their energy-efficient implementation is of great interest. Although, traditionally, the entire network is utilized for the recognition of all inputs, we observe that the classification difficulty varies widely across inputs in real-world datasets; only a small fraction of inputs requires the full computational effort of a network, while a large majority can be classified correctly with very low effort. In this article, we propose Conditional Deep Learning (CDL), where the convolutional layer features are used to identify the variability in the difficulty of input instances and conditionally activate the deeper layers of the network. We achieve this by cascading a linear network of output neurons for each convolutional layer and monitoring the output of the linear network to decide whether classification can be terminated at the current stage or not. The proposed methodology thus enables the network to dynamically adjust the computational effort depending on the difficulty of the input data while maintaining competitive classification accuracy. The overall energy benefits for MNIST/CIFAR10/Tiny ImageNet datasets with state-of-the-art deep-learning architectures are 1.84 × /2.83 × /4.02 × , respectively. We further employ the conditional approach to train deep-learning networks from scratch with integrated supervision from the additional output neurons appended at the intermediate convolutional layers. Our proposed integrated CDL training leads to an improvement in the gradient convergence behavior giving substantial error rate reduction on MNIST/CIFAR-10, resulting in improved classification over state-of-the-art baseline networks.
Priyadarshini Panda, Abhronil Sengupta, Kaushik Roy 0001
ACM J. Emerg. Technol. Comput. Syst.2
2017 Energy-Efficient Object Detection Using Semantic Decomposition
abstract
In this brief, we present a new approach to optimize energy efficiency of object detection tasks using semantic decomposition to build a hierarchical classification framework. We observe that certain semantic information like color/texture is common across various images in real-world data sets for object detection applications. We exploit these common semantic features to distinguish the objects of interest from the remaining inputs (nonobjects of interest) in a data set at a lower computational effort. We propose a 2-stage hierarchical classification framework, with increasing levels of complexity, wherein the first stage is trained to recognize the broad representative semantic features relevant to the object of interest. The first stage rejects the input instances that do not have the representative features and passes only the relevant instance to the second stage. Our methodology thus allows us to reject certain information at lower complexity and utilize the full computational effort of a network only on a smaller fraction of inputs resulting in energy-efficient detection.
Priyadarshini Panda, Swagath Venkataramani, Abhronil Sengupta, Anand Raghunathan, Kaushik Roy 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2016 Prospects of efficient neural computing with arrays of magneto-metallic neurons and synapses
abstract
Non-von Neumann computing models, like Artificial and Spiking Neural Networks, inspired from the functionalities of the human brain, would require devices that can offer a direct mapping to the underlying neuroscience mechanisms for energy-efficient and compact hardware implementation. To that effect, spin-transfer torque phenomena in devices based on lateral spin valves, domain wall motion in magnets and magnetic tunnel junctions can potentially pave the way for spintronic neural computing systems, where spintronic neurons interfaced with spintronic synapses, can directly mimic biological neural and synaptic functionalities. We explore various device structures suitable for such non-Boolean functionalities and demonstrate the potential benefits of such neural computing based on arrays of magneto-metallic neurons and synapses.
Abhronil Sengupta, Karthik Yogendra, Deliang Fan, Kaushik Roy 0001
ASP-DAC1
2016 Invited - Cross-layer approximations for neuromorphic computing: from devices to circuits and systems
abstract
Neuromorphic algorithms are being increasingly deployed across the entire computing spectrum from data centers to mobile and wearable devices to solve problems involving recognition, analytics, search and inference. For example, large-scale artificial neural networks (popularly called deep learning) now represent the state-of-the art in a wide and ever-increasing range of video/image/audio/text recognition problems. However, the growth in data sets and network complexities have led to deep learning becoming one of the most challenging workloads across the computing spectrum. We posit that approximate computing can play a key role in the quest for energy-efficient neuromorphic systems. We show how the principles of approximate computing can be applied to the design of neuromorphic systems at various layers of the computing stack. At the algorithm level, we present techniques to significantly scale down the computational requirements of a neural network with minimal impact on its accuracy. At the circuit level, we show how approximate logic and memory can be used to implement neurons and synapses in an energy-efficient manner, while still meeting accuracy requirements. A fundamental limitation to the efficiency of neuromorphic computing in traditional implementations (software and custom hardware alike) is the mismatch between neuromorphic algorithms and the underlying computing models such as von Neumann architecture and Boolean logic. To overcome this limitation, we describe how emerging spintronic devices can offer highly efficient, approximate realization of the building blocks of neuromorphic computing systems.
Priyadarshini Panda, Abhronil Sengupta, Syed Shakib Sarwar, Gopalakrishnan Srinivasan, Swagath Venkataramani, Anand Raghunathan, Kaushik Roy 0001
DAC2
2016 Low-power approximate convolution computing unit with domain-wall motion based "spin-memristor" for image processing applications
abstract
Convolution serves as the basic computational primitive for various associative computing tasks ranging from edge detection to image matching. CMOS implementation of such computations entails significant bottlenecks in area and energy consumption due to the large number of multiplication and addition operations involved. In this paper, we propose an ultra-low power and compact hybrid spintronic-CMOS design for the convolution computing unit. Low-voltage operation of domain-wall motion based magneto-metallic "Spin-Memristor"s interfaced with CMOS circuits is able to perform the convolution operation with reasonable accuracy. Simulation results of Gabor filtering for edge detection reveal ~ 2.5× lower energy consumption compared to a baseline 45nm-CMOS implementation.
Yong Shim, Abhronil Sengupta, Kaushik Roy 0001
DAC2
2016 Conditional Deep Learning for energy-efficient and enhanced pattern recognition
Priyadarshini Panda, Abhronil Sengupta, Kaushik Roy 0001
DATE2
2016 On the energy benefits of spiking deep neural networks: A case study
abstract
Deep learning neural networks have achieved success in a large number of visual processing tasks and are currently utilized for many real-world applications like image search and speech recognition among others. However, in spite of achieving high accuracy in such classification problems, they involve significant computational resources. Over the past few years, artificial neural network models have evolved into the biologically realistic and event-driven spiking neural networks. Recent research efforts have been directed at developing mechanisms to convert traditional deep artificial nets to spiking nets where the neurons communicate by means of spikes. However, there have been limited studies providing insights on the specific power, area and energy benefits offered by deep spiking neural nets in comparison to their non-spiking counterparts. In this paper, we perform a case study for a hardware implementation of a spiking/non-spiking deep net on the MNIST dataset and clearly outline the design prospects involved in implementing neural computing platforms in the spiking mode of operation.
Bing Han 0006, Abhronil Sengupta, Kaushik Roy 0001
IJCNN2
2016 Spintronic devices for ultra-low power neuromorphic computation (Special session paper)
abstract
Emerging spin-transfer torque mechanisms in devices like vertical spin valves, lateral spin valves, domain wall motion based devices, spin-torque oscillators and spin-orbit torque based devices have opened up new possibilities of mimicking various neural and synaptic functionalities by the underlying device physics. In this paper, we review various spintronic device structures that can provide a compact and area-efficient implementation of artificial neurons and synapses. Neuromorphic architectures based on such spintronic devices can potentially provide ~ 10-100× lower energy consumption in comparison to a baseline CMOS implementation.
Abhronil Sengupta, Karthik Yogendra, Kaushik Roy 0001
ISCAS1
2016 Spin-Transfer Torque Devices for Logic and Memory: Prospects and Perspectives
abstract
As CMOS technology begins to face significant scaling challenges, considerable research efforts are being directed to investigate alternative device technologies that can serve as a replacement for CMOS. Spintronic devices, which utilize the spin of electrons as the state variable for computation, have recently emerged as one of the leading candidates for post-CMOS technology. Recent experiments have shown that a nano-magnet can be switched by a spin-polarized current and this has led to a number of novel device proposals over the past few years. In this paper, we provide a review of different mechanisms that manipulate the state of a nano-magnet using current-induced spin-transfer torque and demonstrate how such mechanisms have been engineered to develop device structures for energy-efficient on-chip memory and logic.
Xuanyao Fong, Yusung Kim 0002, Karthik Yogendra, Deliang Fan, Abhronil Sengupta, Anand Raghunathan, Kaushik Roy 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2016 Hierarchical Temporal Memory Based on Spin-Neurons and Resistive Memory for Energy-Efficient Brain-Inspired Computing
abstract
Hierarchical temporal memory (HTM) tries to mimic the computing in cerebral neocortex. It identifies spatial and temporal patterns in the input for making inferences. This may require a large number of computationally expensive tasks, such as dot product evaluations. Nanodevices that can provide direct mapping for such primitives are of great interest. In this paper, we propose that the computing blocks for HTM can be mapped using low-voltage, magnetometallic spin-neurons combined with an emerging resistive crossbar network, which involves a comprehensive design at algorithm, architecture, circuit, and device levels. Simulation results show the possibility of more than 200× lower energy as compared with a 45-nm CMOS ASIC design.
Deliang Fan, Mrigank Sharad, Abhronil Sengupta, Kaushik Roy 0001
IEEE Trans. Neural Networks Learn. Syst.3
2015 Spin-Transfer Torque Magnetic neuron for low power neuromorphic computing
abstract
Neuromorphic computing attempts to emulate the remarkable efficiency of the human brain in vision, perception and cognition related tasks. Nanoscale devices that offer a direct mapping to the underlying neural computations have emerged as a promising candidate for such neuromorphic architectures. In this paper, a Magnetic Tunneling Junction (MTJ) has been proposed to perform the thresholding operation of a biological neuron. A crossbar array consisting of programmable resistive synapses generates an excitatory / inhibitory charge current input to the neuron. The magnetization of the free layer of the MTJ is manipulated by Spin-Transfer Torque generated by the net synaptic current. Algorithm, device and circuit co-simulation framework suggest the possibility of ∼ 1.63 – 1.79x power savings in comparison to a 45nm digital CMOS implementation.
Abhronil Sengupta, Kaushik Roy 0001
IJCNN1
2012 An Adaptive Memetic Algorithm using a synergy of Differential Evolution and Learning Automata
abstract
In recent years there has been a growing trend in the application of Memetic Algorithms for solving numerical optimization problems. They are population based search heuristics that integrate the benefits of natural and cultural evolution. In this paper, we propose an Adaptive Memetic Algorithm, named LA-DE which employs a competitive variant of Differential Evolution for global search and Learning Automata as the local search technique. During evolution Stochastic Automata Learning helps to balance the exploration and exploitation capabilities of DE resulting in local refinement. The proposed algorithm has been evaluated on a test-suite of 25 benchmark functions provided by CEC 2005 special session on real parameter optimization. Experimental results indicate that LA-DE outperforms several existing DE variants in terms of solution quality.
Abhronil Sengupta, Tathagata Chakraborti, Amit Konar, Atulya K. Nagar
IEEE Congress on Evolutionary Computation1
2012 A heuristic approach to 3D face modelling for efficient face recognition
abstract
This article provides a swarm intelligence approach to 3D face recognition. A parametric evolutionary face model is proposed and the optimal parameters are determined by minimizing an error function. The extracted parameters are employed in the recognition phase for classification. Experimental validations have been performed on the neutral face scans of the CASIA Face Database, a challenging database for face recognition purposes and the results demonstrate the efficacy of the approach.
Tathagata Chakraborti, Abhronil Sengupta, Amit Konar, Ramadoss Janarthanan
HIS2